Senior Site Reliability Engineer
ClearanceJobs
On behalf of our Federal Contracting client, ClearanceJobs Talent Solutions Team is seeking a DevSecOps & Site Reliability Engineer to keep our client’s production deployed applications reliable, observable, and continuously improving in the field. You will own the full lifecycle of platform operations: building resilient deployment pipelines, instrumenting systems with deep telemetry, responding to incidents, and shepherding new releases from developer commits through regression testing and into production. You will sit alongside GEOINT analysts, and engineers, and your work will directly determine whether warfighters and intelligence professionals get the answers they need, when they need them. An active Top Secret security clearance with SCI eligibility is required (active TS/SCI with CI Poly is preferred). Work is performed 100% on-site in Springfield, VA. There is a 24x7 on-call rotation associated with this position and potential for CONUS/OCONUS travel for deployments, exercises, and customer engagements. Responsibilities Build and operate resilient platforms. Design, deploy, and maintain containerized services on OpenShift and Kubernetes across AWS and on-premise edge hardware, across multiple classification environments, including air-gapped environments. Run distributed deployments. Execute and continuously improve the distributed deployment, configuration, and lifecycle management of applications across enterprise data centers and forward edge nodes. Observe and troubleshoot . Build, tune, and operate the Prometheus and Grafana observability stack - metrics, dashboards, alerts, and SLOs, and apply AIOps techniques to detect anomalies, correlate signals, and shorten time-to-resolution. Deliver releases under pressure. Rapidly install new software releases into development and test environments, execute formal regression test protocols, and coordinate scheduled deployments to production with zero or minimal user impact. Test, deploy, and scale services. Containerize, deploy, and horizontally scale services as part of an enterprise- or edge-ready architecture using Docker and Kubernetes across cloud and on-premise edge deployable hardware. Integrate identity and access . Implement and maintain integrations with enterprise identity providers including GEOAxIS, Microsoft Entra ID, and other PKI/SAML/OIDC infrastructure. Run the help desk loop. Triage, document, and resolve user-reported incidents using ticketing systems; maintain runbooks, postmortems, and knowledge-based articles that raise the floor for the whole support team. Collaborate across the mission. Work shoulder-to-shoulder with GEOINT analysts, engineers and researchers to translate operational needs into a deployable capability, and to translate field issues back into engineering fixes. Harden and secure. Apply DevSecOps practices end-to-end: vulnerability scanning, hardened base images, secrets management, STIG compliance, and continuous accreditation maintenance for ATO-sustained environments. Support 24/7 worldwide users. Participate in an on-call rotation supporting users across multiple time zones and combatant commands; respond decisively to outages, degradations, and high-priority operational events. Required Education / Experience 5+ years of professional experience in DevOps, SRE, DevSecOps, or production O&M roles supporting distributed software systems. Active TS/SCI clearance Container platforms: Strong hands-on experience with OpenShift and Kubernetes, workloads, operators, ingress, networking, storage, and upgrades. Cloud: Practical AWS experience (EC2, EKS, S3, IAM, VPC, CloudWatch); experience working in GovCloudstyle restricted environments. Distributed deployment & O&M: Demonstrated ownership of production deployments, configuration management, patching, and lifecycle operations for distributed systems. Observability: Production-grade Prometheus and Grafana experience, exporters, recording rules, alerting, dashboards, and SLO-driven operations. AIOps: Familiarity with anomaly detection, log/metric correlation, and automated remediation patterns to reduce MTTR. Containers & CI/CD: Fluency with Docker, image hardening, and CI/CD pipelines (GitLab CI, Jenkins, GitHub Actions, or equivalent). Identity management: Working familiarity with GEOAxIS, Microsoft Entra ID, PKI/CAC, SAML, and OIDC integration patterns. Ticketing & incident management: Experience operating a help desk / incident workflow. Linux & scripting: Strong Linux administration skills and proficiency in Bash plus at least one of Python, Go, or equivalent. Preferred Qualifications Direct experience supporting GEOINT, IMINT, or all-source intelligence production environments. Experience with edge or tactical deployments operating under DIL/DDIL constraints. Knowledge of ICD 503, RMF, and ATO sustainment processes. Hands-on with Kafka, OpenSearch, Elasticsearch, or similar distributed data platforms. Experience deploying or operating ML/AI inference workloads, including GPU-accelerated services. Infrastructure-as-Code: Terraform, Ansible, or Helm at production scale. On-site work at a customer facility is required. Participation in a 24/7 on-call rotation. Occasional travel to additional CONUS and OCONUS sites for deployments, exercises, and customer engagements. Work is performed on classified networks; standard security and handling protocols apply. #J-18808-Ljbffr ClearanceJobs
$135k - $150k
Senior Site Reliability Engineer Job number: 884 This is a remote position. Ad Hoc is a technology company that empowers organizations to deliver scalable, impactful digital services. Using modern, agile methods, our team creates products that meet people's...SeniorRemote workFlexible hours- ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments, predicting, and resolving issues, and supporting the production environment...SeniorWork experience placement
$106.3k - $221.1k
...more. Join us to drive positive, lasting change that moves missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key...SeniorLive inWork at officeLocal area$104.9k - $174.7k
...immediately hire a highly skilled and proactive Senior SRE to join our dynamic team. You will... ...fault-tolerant systems within agreed reliability objectives, whilst enabling the fast... ...skills. About team; This diverse team of Engineers in assisting multiple product teams as...SeniorLocal areaImmediate startWorldwide$121.4k - $218.6k
...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner... ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling...SeniorWork experience placementWork at office$166k - $220k
...Site Reliability Engineer (SRE) Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative...SeniorFull timeWork experience placementImmediate start$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection...SeniorTemporary workImmediate startFlexible hoursShift work$153k - $185k
...Senior Site Reliability Engineer El Segundo, California, United States About Varda Low Earth orbit is open for business. Varda is accelerating the development of commercial space infrastructure, from in-orbit pharmaceutical processing to reliable and economical...SeniorPermanent employmentFull timeImmediate startRelocation packageFlexible hoursWeekend work$116.9k - $234.1k
Site Reliability Engineer The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key Performance Indicators and Service Level Objectives, identify and resolve performance bottlenecks...SeniorLocal area$175k - $250k
Senior Cloud Infrastructure Engineer Location: San Francisco, CA (On‑site only) — must live within commuting distance or be willing to relocate. Compensation: $175,0... ...while ensuring scalability, performance, and reliability across environments. What You’ll Do Design,...SeniorFull timeRelocation$175k - $250k
Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or be willing to... ...ensuring scalability, performance, and reliability across environments. What You’ll Do...SeniorFull timeRemote workRelocationRelocation package- ...Azure, Oracle, Cassandra, SQL Server, My SQL and Mongo DB Seniority level Seniority level Mid-Senior level Employment type Employment... ...new job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago...SeniorContract workRemote work
$149.4k - $202k
Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on...SeniorRemote work- ...impact. If you’re excited to be part of a team that’s leading the way in AI-powered collaboration, we’d love to meet you. Senior Site Reliability Engineer As a Senior Site Reliability Engineer, you will use your advanced development and operations knowledge to run...SeniorWork experience placementWork at officeWorldwide
$165k - $230k
...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts...SeniorPermanent employmentTemporary workImmediate startWeekend work- ...Senior Geospatial Systems Software Engineer Lead the future of geospatial innovation. GDIT is your place to lead, innovate, and drive strategic change. We are seeking a Senior Geospatial Systems Software Engineer with proven experience delivering mission-critical integration...SeniorWork from home
- ...ears, and hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical platform running... ...of deep technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident...Full timeWork at officeRemote workFlexible hours
- Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient, scalable, and highly performant...SeniorLocal area
- ...Overview Abile Group has an exciting and challenging opportunity for a Senior Software Engineer on a 10 year contract providing User Facing and Data Center Services supporting an Intelligence Community customer. All the personnel on the team will work together to support...SeniorContract workWork at officeWorldwide
- ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III - DevOps Engineer at JPMorgan Chase within the Commercial and Investment Bank, you will solve complex and broad business...
$147k - $202k
...TechOps) team, we live this mission by building the most reliable and performant systems on the planet. We empower... ...need. The Role We are looking for an experienced Senior Site Reliability Engineer (SRE) who thrives on the challenge of managing large-scale...SeniorPermanent employmentLocal areaWorldwideFlexible hours$86.8k - $198k
...production-ready.This role is more than just coding. As a Sof tware Engineer at Booz Allen, you’ll use your passion to master new tools and... ...our total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.Salary at Booz Allen is...SeniorFull timeContract workPart timeWork at officeLocal areaRemote work- Modern Technology Solutions, Inc. (MTSI) is hiring a Senior Software Engineer in Springfield, VA. Qualified candidates must be U.S. citizens... ...evaluation consistency, operational scalability, and workflow reliability. Support automation, orchestration, and optimization of...Senior
$175k - $185k
...Impact. Deliver simple solutions to complex problems as a Senior DevOps Engineer at D2. In this role, you’ll design and support modern DevOps... ...intelligence systems. You’ll focus on automation, reliability, and scalability—while we focus on investing in your growth...Senior$180k - $202.4k
Job Title: Senior Software Engineer Location: Springfield, VA Security Clearance: Active DoD Top Secret clearance with SCI eligibility Omni Federal, founded in 2017 and headquartered in Washington, DC, is a specialized software solutions provider with key locations...Senior$140k - $160k
Zachary Piper Solutions is seeking a talented Senior Software Engineer to join our team in Springfield, VA. As a Senior Software Engineer, you... ...using server instances hosted on AWS EC2, ensuring the reliability and stability of the applications. Utilize AWS S3 to securely...SeniorWork experience placement$148k - $171k
...Overview Senior DevOps Engineer Springfield, VA or St. Louis, MO Active TS (SCI eligibility) clearance and eligibility to obtain... ...documentation Enforce best practices for security and reliability, and drive security initiatives, like access control and vulnerability...Senior$207k - $284.9k
...This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Senior Manager, Site Reliability Engineering Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI...SeniorPermanent employmentLocal areaWorldwideFlexible hoursDay shift- ...Principal Site Reliability Engineer The Principal Site Reliability Engineer will be a critical technical leader responsible for driving the... ...for a key Randstad client in the Washington D.C. area. This senior role merges deep expertise in infrastructure automation (IaC...
- Abile Group has an exciting and challenging opportunity for a Senior Software Engineer on a 10 year contract providing User Facing and Data Center Services supporting an Intelligence Community customer. All the personnel on the team will work together to support innovative...SeniorContract workWork at officeWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- senior implementation project manager Springfield, VA
- senior commercial counsel Springfield, VA
- senior director diversity & inclusion Springfield, VA
- senior software engineer remote Springfield, VA
- senior living Springfield, VA
- senior resident engineer Springfield, VA
- senior infrastructure engineer Springfield, VA
- senior analyst Springfield, VA
- senior consulting engineer Springfield, VA
- remote senior salesforce administrator Springfield, VA


