Lead Site Reliability Engineer
Bridge Defense
Lead Site Reliability Engineer
Bridge Defense is redefining how modern defense technology is delivered. Based in Washington, D.C., we are built for the dynamic mission environment facing the Department of Defense, the Intelligence Community, and federal law enforcement agencies. We provide full-spectrum national security solutions that combine secure infrastructure, cleared talent, and mission-ready software to meet evolving defense challenges. Our services include secure software development in classified environments and the design and implementation of advanced IT and cybersecurity capabilities ranging from secure cloud architectures and enterprise infrastructure to data center operations, scientific analysis, and cutting-edge cyber defense.
We are led by technologists and veterans with firsthand mission experience, which enables us to understand both the operational realities and the innovation needed to succeed. Our approach is agile and outcome-based, delivering results in weeks rather than months whenever possible.
At Bridge Defense we value people, integrity, and excellence. We foster an environment where innovation thrives in support of traditional mission requirements. Our team members receive competitive compensation, robust benefits, professional development and certification opportunities, and clear paths for growth while working on the nation's most critical projects.
Core Values
- Innovation & Responsiveness: We push beyond legacy models with efficient, tech-led solutions built to scale and evolve.
- Trusted Performance: Security, compliance, and deep experience in delivering to demanding environments guides all we do.
- Mission Focused Expertise: From veteran leadership to cleared engineers, our people understand both the technology and the mission.
About the Role
As the Lead Site Reliability Engineer for our ComputeBridge Engagement, you'll be responsible for the reliability, scalability, and performance of one of the largest hardware and AI infrastructure efforts in the U.S. defense sector. You will lead the deployment, management, and automation of a high-performance computing mesh across multiple secure environments, ensuring operational excellence and mission continuity for a 9-figure government program.
This is a hands-on engineering leadership role that bridges physical infrastructure and modern DevOps automation, ideal for someone who thrives at the intersection of hardware systems, distributed computing, and AI/ML workflows.
What You'll Do
- Lead infrastructure design, deployment, and operations for ComputeBridge hardware clusters across secure and distributed environments
- Install and configure physical systems, including high-density GPU servers, networking gear, and storage arrays
- Build and deploy secure Linux images and containerized workloads using OpenShift and other orchestration platforms
- Develop and manage automation pipelines for provisioning, configuration management, and monitoring using modern DevOps toolchains (Ansible, Terraform, etc.)
- Operate and maintain distributed networking meshes across multiple classified and unclassified domains
- Implement and manage out-of-band management tools (IMPI, iDRAC, BMC, etc.) for remote troubleshooting and control
- Integrate and optimize NVIDIA GPU infrastructure for AI/ML training and inference workloads
- Collaborate with mission engineers, software teams, and government operators to ensure system readiness and performance
- Provide on-site technical leadership for deployments, troubleshooting, and continuous improvement
- Mentor junior engineers and establish operational best practices across the ComputeBridge program as the contract grows
What You'll Bring
- 3+ years of experience in site reliability, systems engineering, or hardware operations roles
- Deep expertise with physical infrastructure: server racking, cabling, diagnostics, and troubleshooting
- Strong experience with Linux systems administration, imaging, and automated deployment
- Hands-on experience managing large-scale clusters or distributed systems in OpenShift or Kubernetes environments
- Familiarity with DevOps automation (Ansible, Terraform, CI/CD pipelines)
- Experience configuring and managing networking and mesh architectures
- Direct experience with NVIDIA GPUs, CUDA, and related AI/ML frameworks
- Proficiency with out-of-band management and IMPI/iDRAC tooling
- Certifications: Linux+ and Security+ (required or in-progress)
- Excellent communication, documentation, and problem-solving skills
- Clearance: Active TS/SCI required or ability to obtain
Bonus Points For
- Experience operating in secure DoD or intelligence environments
- Familiarity with Palantir platforms or other government data systems
- Prior experience supporting AI/ML infrastructure in production or tactical settings
- Experience with performance tuning and monitoring of HPC or GPU-accelerated clusters
General Factors
- Depending on project requirements, may be required to work within a compressed schedule; overtime should be expected when schedules demand it.
- Willing to travel, if needed.
- No Relocation.
Why Bridge Defense
- Shape how advanced computing supports national security missions at scale
- Lead engineering for a major government program with direct mission impact
- Competitive compensation, benefits, and growth opportunities in a mission-driven environment
Bridge Defense is committed to building a collaborative and mission-focused team. Bridge Defense reserves the right to modify job duties or requirements at any time. Employment with Bridge Defense is at-will. Candidates must be eligible to work in the United States and complete any required background checks or security clearance processes as a condition of employment.
- ...hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical... ...technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident response in a constrained environment...SuggestedFull timeWork at officeRemote workFlexible hours
- ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III - DevOps Engineer at JPMorgan Chase within the Commercial and Investment Bank, you will solve complex and broad business...Suggested
$166k - $220k
...Site Reliability Engineer (SRE) Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities... ...through analysis, design and code. They are comfortable leading large, focused projects. They lead in the development of...SuggestedFull timeWork experience placementImmediate start$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role... ...future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives...SuggestedTemporary workImmediate startFlexible hoursShift work$104.9k - $174.7k
...scale, 24x7, distributed and fault-tolerant systems within agreed reliability objectives, whilst enabling the fast flow of feature and... ...strong automation skills. About team; This diverse team of Engineers in assisting multiple product teams as we continue to innovate...SuggestedLocal areaImmediate startWorldwide$121.4k - $218.6k
...for ensuring best-in-class uptime and reliability of our AI hardware infrastructure... ...when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing... ...building technical runbooks, leading complex incident response bridges, and...Work experience placementWork at office- ...including, with hands-on Development and Systems engineering background ~3-5 years of experience in a Site Reliability Engineering role ~ Experience with Enterprise... ...Cloud journey and as a member of our team help lead Software automation and reliability for our platform...Temporary workImmediate start
$106.3k - $221.1k
...more. Join us to drive positive, lasting change that moves missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key...Live inWork at officeLocal area$125k - $135k
...Site Reliability Engineer Job number: 880 This is a remote position. Ad Hoc is a technology company that empowers organizations to deliver scalable, impactful digital services. Using modern, agile methods, our team creates products that meet people's needs...Remote workFlexible hours$165k - $230k
...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts...Permanent employmentTemporary workImmediate startWeekend work$95k - $171k
...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for:... ...Akamai powers and protects life online. Leading companies worldwide choose Akamai to...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$112k - $179k
...The Role Peraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in... ...extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise...Contract workWorldwideShift work$153k - $185k
...Senior Site Reliability Engineer El Segundo, California, United States About Varda Low Earth orbit is open for business. Varda is accelerating the development of commercial space infrastructure, from in-orbit pharmaceutical processing to reliable and economical...Permanent employmentFull timeImmediate startRelocation packageFlexible hoursWeekend work$131k - $227.13k
...Description: The 1LMX MES COE is seeking an engineer who will own infrastructure‑as‑code, cloud platform, and reliability for the Apriso environment on AWS. This role blends full‑stack development, DevOps, and Site Reliability Engineering (SRE) practices to deliver...Full timeTemporary workWork experience placementWork at officeRemote workRelocationFlexible hoursShift work3 days per week- ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments, predicting, and resolving issues, and supporting the production environment...Work experience placement
- ...Site Reliability Engineer (SRE) Randstad is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our client in the Washington D.C. area, focusing on optimizing the availability, performance, and scalability of critical production services. The ideal...
$114.6k - $190.2k
...with MANTECH! ***This is for a future opportunity*** MANTECHseeks motivated, career, and customer-oriented Site Reliability Engineer (SRE) for a new initiative. This effort supports the rapid design, deployment, operation, and sustainment of enterprise-...Hourly payContract workTemporary workWork experience placementWork at officeLocal areaRemote work$131k - $164k
...Staff Site Reliability Engineer New York, New York, United States Position Overview We are seeking a highly skilled Staff Site Reliability... ...they need to drive greater impact and accountability – to lead with purpose. Our employees are passionate, smart, and...Work at officeLocal areaFlexible hours$51.9 per hour
...This job is responsible for the reliability, availability, and... ...efficiency. This role blends software engineering, clinical engineering, and... ...cross-functionally with AHN site leaders and teams to navigate... ...drills and exercises, as needed. Leads or participates in post-...For contractorsLocal area$178k - $213k
...Partners 2022 Cybersecurity Excellence Award for MDR Manager, Site Reliability Engineering Reports to: VP, Product Engineering Location: While... ...to remote candidates who can support the Eastern Time Zone. Lead the architecture, automation, and reliability of secure, scalable...Permanent employmentWork experience placementWork at officeRemote workWork from homeHome officeFlexible hours$121.5k - $306.4k
...and provides input on best practices for reliability and functionality. Establishes direction... ..., executing improvements, building site reliability knowledge, and providing clear... ...Collaboration & Partnership: -Role models leading cross-functional collaborative efforts...Temporary workFlexible hours- ...Associate Software Engineer The purpose of the Associate Software Engineer position is to provide technical assistance directly to... ...finthrive. Award-winning Culture of Customer-centricity and Reliability At FinThrive we're proud of our agile and committed...Contract workWork experience placementCasual workLocal area
$84.9k - $209.5k
...unencumbered and will need your contribution to make it a special engineering center with the focus on excellence. Health Data... ...a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives...Temporary workImmediate startFlexible hours- ...Principal Site Reliability Engineer The Principal Site Reliability Engineer will be a critical technical leader responsible for driving the... ...SRE), including defining SLOs, managing error budgets, and leading incident response. You will mentor cross-functional teams,...
$220k - $250k
...Staff Site Reliability Engineer Yugabyte is the company behind YugabyteDB, the AI-ready, multi-modal, distributed PostgreSQL database for cloud... ..., our hard-working team of experts and our industry-leading technology are uniquely positioned to meet the demands of modern...H1bLocal areaWorldwideVisa sponsorship- Overview MANTECH seeks motivated, career, and customer-oriented Site Reliability Engineer (SRE) for a new initiative. This effort supports the rapid design, deployment, operation, and sustainment of enterprise-scale AI, data, and mission platform capabilities across cloud...Work at office
$100.2k - $203.4k
As a Site Reliability Engineer, you will play a pivotal role in advancing operational AI adoption within a cutting‑edge Hub‑and‑Spoke architecture... ...AI systems within a modern Hub‑and‑Spoke architecture. Lead incident response efforts to minimize downtime and maintain...Live inLocal area$149.4k - $202k
Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC... ...address performance and availability issues proactively. Lead the strategic effort to eliminate toil, identifying and championing...Remote work$116.9k - $234.1k
Site Reliability Engineer The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key Performance Indicators and Service Level Objectives, identify and resolve performance bottlenecks...Local area- ...Communication : Excellent communicator. Expected to actively lead and triage proactively identified issues/incidents where... ...a new job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago Seattle...Contract workRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead maintenance engineer Washington DC
- lead operating engineer Washington DC
- lead industrial engineer Washington DC
- lead engineer Washington DC
- lead infrastructure engineer Washington DC
- lead network engineer Washington DC
- lead system engineer Washington DC
- lead web developer Washington DC
- site reliability engineer Washington DC
- site reliability engineer remote Washington DC


