Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

Bridge Defense

Join to apply for the Lead Site Reliability Engineer role at Bridge Defense About Bridge Defense Bridge Defense is redefining how modern defense technology is delivered. Based in Washington, D.C., we are built for the dynamic mission environment facing the Department of Defense, the Intelligence Community, and federal law enforcement agencies. We provide full‑stream national security solutions that combine secure infrastructure, cleared talent, and mission‑ready software to meet evolving defense challenges. Our services include secure software development in classified environments and the design and implementation of advanced IT and cybersecurity capabilities ranging from secure cloud architectures and enterprise infrastructure to data center operations, scientific analysis, and cutting‑edge cyber defense. We are led by technologists and veterans with firsthand mission experience, which enables us to understand both the operational realities and the innovation needed to succeed. Our approach is agile and outcome‑based, delivering results in weeks rather than months whenever possible. At Bridge Defense we value people, integrity, and excellence. We foster an environment where innovation thrives in support of traditional mission requirements. Our team members receive competitive compensation, robust benefits, professional development and certification opportunities, and clear paths for growth while working on the nation’s most critical projects. Core Values Innovation & Responsiveness: We push beyond legacy models with efficient, tech‑led solutions built to scale and evolve. Trusted Performance: Security, compliance, and deep experience in delivering to demanding environments guides all we do. Mission Focused Expertise: From veteran leadership to cleared engineers, our people understand both the technology and the mission. About The Role As the Lead Site Reliability Engineer for our ComputeBridge Engagement, you’ll be responsible for the reliability, scalability, and performance of one of the largest hardware and AI infrastructure efforts in the U.S. defense sector. You will lead the deployment, management, and automation of a high‑performance computing mesh across multiple secure environments, ensuring operational excellence and mission continuity for a 9‑figure government program. This is a hands‑on engineering leadership role that bridges physical infrastructure and modern DevOps automation, ideal for someone who thrives at the intersection of hardware systems, distributed computing, and AI/ML workflows. What You’ll Do Lead infrastructure design, deployment, and operations for ComputeBridge hardware clusters across secure and distributed environments Install and configure physical systems, including high‑density GPU servers, networking gear, and storage arrays Build and deploy secure Linux images and containerized workloads using OpenShift and other orchestration platforms Develop and manage automation pipelines for provisioning, configuration management, and monitoring using modern DevOps toolchains (Ansible, Terraform, etc.) Operate and maintain distributed networking meshes across multiple classified and unclassified domains Implement and manage out‑of‑band management tools (IMPI, iDRAC, BMC, etc.) for remote troubleshooting and control Integrate and optimize NVIDIA GPU infrastructure for AI/ML training and inference workloads Collaborate with mission engineers, software teams, and government operators to ensure system readiness and performance Provide on‑site technical leadership for deployments, troubleshooting, and continuous improvement Mentor junior engineers and establish operational best practices across the ComputeBridge program as the contract grows What You’ll Bring 3+ years of experience in site reliability, systems engineering, or hardware operations roles Deep expertise with physical infrastructure: server racking, cabling, diagnostics, and troubleshooting Strong experience with Linux systems administration, imaging, and automated deployment Hands‑on experience managing large‑scale clusters or distributed systems in OpenShift or Kubernetes environments Familiarity with DevOps automation (Ansible, Terraform, CI/CD pipelines) Experience configuring and managing networking and mesh architectures Direct experience with NVIDIA GPUs, CUDA, and related AI/ML frameworks Proficiency with out‑of‑band management and IMPI/iDRAC tooling Certifications: Linux+ and Security+ (required or in‑progress) Excellent communication, documentation, and problem‑solving skills Clearance: Active TS/SCI required or ability to obtain Bonus Points For Experience operating in secure DoD or intelligence environments Familiarity with Palantir platforms or other government data systems Prior experience supporting AI/ML infrastructure in production or tactical settings Experience with performance tuning and monitoring of HPC or GPU‑accelerated clusters General Factors Depending on project requirements, may be required to work within a compressed schedule; overtime should be expected when schedules demand it. Willing to travel, if needed. No Relocation. Why Bridge Defense Shape how advanced computing supports national security missions at scale Lead engineering for a major government program with direct mission impact Competitive compensation, benefits, and growth opportunities in a mission‑driven environment Bridge Defense is committed to building a collaborative and mission‑focused team. Bridge Defense reserves the right to modify job duties or requirements at any time. Employment with Bridge Defense is at‑will. Candidates must be eligible to work in the United States and complete any required background checks or security clearance processes as a condition of employment. Seniority level: Mid‑Senior level Employment type: Full‑time Job function: Engineering and Information Technology Industries: Defense and Space Manufacturing Referrals increase your chances of interviewing at Bridge Defense by 2x #J-18808-Ljbffr

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Washington DC vacancy
  •  ...Communication : Excellent communicator. Expected to actively lead and triage proactively identified issues/incidents where...  ...a new job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago Seattle... 
    Suggested
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    2 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA....  ...Remote unavailable. Modality: On‑Site only. Must live within...  ...this role, you will take the lead on designing, deploying, and...  ...scalability, performance, and reliability across environments. What You... 
    Suggested
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    2 days ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology...  ...address performance and availability issues proactively. Lead the strategic effort to eliminate toil, identifying and... 
    Suggested
    Remote work

    Noctua Technology

    Washington DC
    4 days ago
  • $131k - $227.13k

     ...Description: The 1LMX MES COE is seeking an engineer who will own infrastructure‑as‑code, cloud platform, and reliability for the Apriso environment on AWS. This role blends full‑stack development, DevOps, and Site Reliability Engineering (SRE) practices to deliver... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Work at office
    Remote work
    Relocation
    Flexible hours
    Shift work
    3 days per week

    Lockheed Martin Corporation

    Bethesda, MD
    3 days ago
  • $121.4k - $218.6k

     ...and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages...  ...Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company... 
    Suggested
    Work experience placement
    Work at office

    Akamai

    Washington DC
    1 day ago
  • $153k - $185k

     ...Senior Site Reliability Engineer El Segundo, California, United States About Varda Low Earth orbit is open for business. Varda is accelerating the development of commercial space infrastructure, from in-orbit pharmaceutical processing to reliable and economical... 
    Permanent employment
    Full time
    Immediate start
    Relocation package
    Flexible hours
    Weekend work

    Varda Space Industries

    Washington DC
    4 days ago
  • $112k - $179k

     ...The Role Peraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in...  ...extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise... 
    Contract work
    Worldwide
    Shift work

    Peraton

    Washington DC
    3 days ago
  • $135k - $150k

    Senior Site Reliability Engineer Job number: 884 This is a remote position. Ad Hoc is a technology company that empowers organizations...  ...across metrics, logging, tracing, and alerting Leading incident response and on-call practices, including escalation... 
    Remote work
    Flexible hours

    Ad Hoc LLC

    Silver Spring, MD
    1 day ago
  •  ...including, with hands-on Development and Systems engineering background ~3-5 years of experience in a Site Reliability Engineering role ~ Experience with Enterprise...  ...Cloud journey and as a member of our team help lead Software automation and reliability for our platform... 
    Temporary work
    Immediate start

    Samprasoft

    Washington DC
    3 days ago
  • $166k - $220k

     ...Senior Site Reliability Engineer Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities...  ...through analysis, design and code. They are comfortable leading large, focused projects. They lead in the development of... 
    Full time
    Work experience placement
    Immediate start

    anduril

    Washington DC
    1 day ago
  • $165k - $230k

     ...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts... 
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    3 days ago
  • $95k - $171k

     ...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for:...  ...Akamai powers and protects life online. Leading companies worldwide choose Akamai to... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Washington DC
    1 day ago
  • $81.1k - $187k

     ...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role...  ...future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Washington DC
    2 days ago
  • $106.3k - $221.1k

     ...more. Join us to drive positive, lasting change that moves missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key... 
    Live in
    Work at office
    Local area

    Accenture

    Arlington, VA
    3 days ago
  • $51.9 per hour

     ...This job is responsible for the reliability, availability, and...  ...efficiency. This role blends software engineering, clinical engineering, and...  ...cross-functionally with AHN site leaders and teams to navigate...  ...drills and exercises, as needed. Leads or participates in post-... 
    For contractors
    Local area

    Highmark Health

    Washington DC
    1 day ago
  • $84.9k - $209.5k

     ...unencumbered and will need your contribution to make it a special engineering center with the focus on excellence. Health Data...  ...a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives... 
    Temporary work
    Immediate start
    Flexible hours

    Oracle

    Washington DC
    3 days ago
  • $131k - $164k

     ...Staff Site Reliability Engineer New York, New York, United States Position Overview We are seeking a highly skilled Staff Site Reliability...  ...they need to drive greater impact and accountability – to lead with purpose. Our employees are passionate, smart, and... 
    Work at office
    Local area
    Flexible hours

    Diligent

    Washington DC
    4 days ago
  • $220k - $250k

     ...Staff Site Reliability Engineer Yugabyte is the company behind YugabyteDB, the AI-ready, multi-modal, distributed PostgreSQL database for cloud...  ..., our hard-working team of experts and our industry-leading technology are uniquely positioned to meet the demands of modern... 
    H1b
    Local area
    Worldwide
    Visa sponsorship

    YugaByte

    Washington DC
    4 days ago
  •  ...Associate Software Engineer The purpose of the Associate Software Engineer position is to provide technical assistance directly to...  ...finthrive. Award-winning Culture of Customer-centricity and Reliability At FinThrive we're proud of our agile and committed... 
    Contract work
    Work experience placement
    Casual work
    Local area

    FinThrive

    Washington DC
    3 days ago
  • $126k - $248k

     ..., you will partner with SRE leaders and engineers to scale the platform that underpins all...  ...program execution, strengthen production reliability practices, and coordinate cross-...  ...Strengthen Production Reliability – Lead change management and launch readiness programs... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Washington DC
    2 days ago
  • $103.5k - $150k

     ...Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the management of experiences, insights, and...  ...Bring your whole self. The Role and Team The Site Reliability Engineering organization at Medallia brings together the infrastructure... 
    Temporary work
    Work experience placement
    Local area
    3 days per week

    Medallia

    McLean, VA
    1 day ago
  • $80k - $133k

     ...degree, Four (4) years additional experience will be needed. * Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud infrastructure and enterprise systems. * One(1)+ years of experience... 
    Permanent employment
    Contract work
    Temporary work
    Flexible hours

    Guidehouse Careers

    McLean, VA
    7 hours ago
  • $17 - $27.75 per hour

     ...deliver an exceptional customer experience * Serves as a Brand Ambassador embodying of Coach values and increasing brand awareness * Leads implementation of Company initiatives and support full operation of the business * Maintain a growth mindset for business and... 
    Minimum wage
    Shift work

    Tapestry

    Oxon Hill, MD
    1 day ago
  •  ...Title: Sr. IT Application Solutions Architect /SRE Engineer Important Note : We have shifted to adopting SAFe and 1. Encourage...  ...Overview We are seeking a highly skilled DevSecOps Engineer to lead the integration of security into our cloud-native development and... 
    For contractors
    Remote work
    Shift work

    Lumen Solutions Group, Inc.

    Washington DC
    2 days ago
  •  ...Site Reliability Engineer II Join the leader in providing smarter solutions for a safer world. The property technology space is growing rapidly, and Kastle Systems is leading the way. Kastle Systems is the leader in managed security, with a track record of introducing... 
    Remote work

    Kastle Systems

    Falls Church, VA
    5 hours ago
  •  ...Site Reliability Engineer Mc Lean, VA Long Term Client's Enterprise Data Machine Learning (EDML) employs innovative minds like yourself to design and develop software-systems that can meet the demand of our ever-growing customer base. Like a... 
    Immediate start

    Maintec Technologies

    McLean, VA
    4 hours ago
  •  ...SRE Engineer Location: Washington, DC (Onsite) Duration: 08-17-2026 - 07-30-2027 Key...  ...using ITIL frameworks and ServiceNow; lead troubleshooting efforts, conduct deep root...  ...knowledge base articles. Reliability Engineering: Champion SRE metrics including... 

    Georgia IT Inc

    Washington DC
    2 days ago
  • $75 - $85 per hour

     ...premier client in the Washington, D.C. area to find a talented Site Reliability Engineer (SRE) to champion system availability, performance, and...  ...production on-call responder using ITIL frameworks and ServiceNow; lead troubleshooting efforts, conduct deep root-cause analysis (... 
    Hourly pay
    Contract work
    Temporary work
    Work experience placement

    Randstad

    Washington DC
    3 days ago
  • $147k - $202k

     ...(TechOps) team, we live this mission by building the most reliable and performant systems on the planet. We empower organizations...  ...The Role We are looking for an experienced Senior Site Reliability Engineer (SRE) who thrives on the challenge of managing large-scale... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    a month ago
  •  ...Description Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the...  ...Management: Act as a primary responder in on-call rotations, leading the technical resolution of production outages.... 
    Local area

    Tiger Analytics Inc.

    Washington DC
    21 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!