Cloud Site Reliability Engineer (SRE)
$130k - $180kECS Federal
Job DescriptionEverforth ECS is seeking a Cloud Site Reliability Engineer (SRE)to work in our Arlington, VA office/remotely. Our Philosophy We believe the job of an SRE is to engineer the cloud to run itself. That means writing software and automation that lets systems detect and recover from failure on their own, rather than relying on someone to notice an alert and manually fix it. When something breaks, self-healing comes first, deep root-cause debugging happens after service is restored, not instead of it. We’re looking for someone who automates the operational task by default, not documents the runbook for doing it by hand. About the Role This role owns reliability and operational readiness for production systems across our federal cloud platform (AWS GovCloud, IL5 zero-trust). You’ll define what “reliable enough” looks like for our services, build the automation that gets us there, and do it all on an infrastructure-as-code (IaC) foundation. Responsibilities Self-Healing Operations Design and build automated remediation so systems detect, respond to, and recover from failure without manual intervention Shift the team’s posture from “is it running, how do we fix it” to “how do we make it fix itself” Automate service restoration first; investigate root cause after Uptime Goals & Reliability Define reasonable, data-driven SLOs and error budgets for critical services alongside the teams that own them Use live metrics to decide what’s “reliable enough” and where to invest next Infrastructure Enforce infrastructure-as-code and configuration-as-code, no manual tech change Own Terraform standards and reusable modules adopted across programs Drive a containerization-first approach with production-scale Kubernetes (multi-tenancy, security policies, advanced scheduling) Set CI/CD and pipeline-as-code standards, including progressive delivery Observability & Incidents Build monitoring, logging, alerting, and tracing (Datadog, Splunk) that gives automation the signal it needs to self-correct Own the incident framework: escalation, restoration, root cause analysis, and post-incident review that closes the loop with more automation Collaboration & Leadership Partner with development and contractor teams leads to embed reliability and automation across the software Mentor engineers toward this same automation-first philosophy Support ATO/RMF and FedRAMP High compliance as it relates to infrastructure and automation Salary Range: $130,000 - $180,000 General Description of BenefitJob RequirementsBachelor’s degree in Computer Science, Information Technology, or related field (or equivalent practical experience) 5+ years of SRE experience (or equivalent), with demonstrated technical leadership 10 years of general work experienceTrack record building self-healing/auto-remediating systems, not just dashboards Jenkins experience Expert AWS knowledge, GovCloud experience strongly preferred Deep Kubernetes and Terraform expertise at production scale Strong software engineering background (Python and/or Go) Experience operating observability platforms (Grafana, Splunk, Prometheus, Loki, etc.) Proven incident command and postmortem experience Strong communication skills across technical and federal leadership audiences Ability to obtain/maintain required government clearance or suitability (CAC/PIV as applicable) US Citizenship Job DetailsJob Type: Full-timeCategory: Software Development, Engineering & ApplicationsSalaried: Salaried
- ...Site Reliability Engineer (SRE) Reston, VA Site Reliability Engineer (SRE) Position: Site Reliability Engineer (SRE) Work Authorization... ...Hands-On experience creating, configuring and maintaining cloud-based applications and infrastructure for the rapid development...CloudContract work
- ...join our talented Team. Job Title: SRE / DevOps Engineer Job Location: Mclean, VA... ...Job Description: We are seeking a Site Reliability Engineer (SRE) with strong expertise... ...client ecosystem and deep knowledge of cloud-native infrastructure. The ideal candidate...Cloud
- ...Title: Sr. IT Application Solutions Architect /SRE Engineer Important Note : We have shifted to adopting SAFe and 1. Encourage... ...DevSecOps Engineer to lead the integration of security into our cloud-native development and operations workflows. This role requires...CloudFor contractorsShift work
- ...SRE Engineer Location: Washington, DC (Onsite) Duration: 08-17-2026 - 07-30-2027 Key... ...Pipeline, or Jenkins; provision scalable cloud infrastructure using Terraform,... ...comprehensive knowledge base articles. Reliability Engineering: Champion SRE metrics including...Cloud
- ...DescriptionThe AI Inference Engineer plays a critical role in the... ...capabilities, ensuring enterprise-grade reliability by leveraging hardware... ...Docker, Kubernetes, and cloud platforms such as AWS, GCP, and... ...kernels). Background in MLOps or SRE roles focused on high-...CloudFull timeLocal areaImmediate start
$95k - $171k
...Do you want to build your SRE career on one of the most exciting platforms in cloud computing? Join the Akamai... ...infrastructure, Kubernetes, and ensuring reliability for AI workloads within... ...inference platform. As an Site Reliability Engineer II, you will be responsible...CloudPermanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...Site Reliability Engineer Location- Wilmington De, Washington DC, Dallas, TX (Onsite Position) Full... ...Minimum of 5 years of experience in an SRE role, with specific experience in... ...Kubernetes. Strong understanding of cloud platforms (AWS preferred, Azure, GCP)....CloudFull time
- ...Description Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the... ...Pulumi to provision and manage consistent, version-controlled cloud environments. CI/CD & GitOps: Design and optimize...CloudLocal area
$119.8k - $234.7k
...yearEmployment type: Full-TimeWork site: 3 days / week in-... ...: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft... ...the limit thinking in a cloud-enabled world.Microsoft’s Azure Data engineering team is leading the... ...is what sets the SRE’s in Cosmos DB apart. If...Cloud3 days per week$210k - $230k
...GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient... ..., reliability, and performance optimization across multi-cloud environments. This position is located in Arlington, VA and...CloudCurrently hiringRemote work$112k - $179k
...seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers... ...applications and infrastructure. The SRE will drive automation initiatives,... ...operations, and security teams.Support cloud and hybrid-cloud environments.QualificationsRequired...CloudContract workWorldwideShift work- ...Site Reliability Engineer (SRE) Randstad is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our client in the Washington... ...as-Code (IaC), incident response, and deep experience with cloud platforms, preferably AWS. The SRE will collaborate across...Cloud
$166k - $258k
...for this position.Nordstrom is looking for a Senior Engineer 2 to join our Site Reliability Engineering (SRE) team — and we think that person could be you.You'... ...languages such as Java, Go, PythonExperience with cloud providers such as AWS, GCP, and AzureStrong networking...CloudFull timeWork at office$107k - $220k
...The Site Reliability Engineer (SRE) will ensure the reliability, performance, and scalability of the WDP System. This person will define and track... ...Requirements SECRET Security Clearance Required Cloud+, GICSP, GSEC, Security+, or SSCP cerification (for Junior...CloudFull timeContract workTemporary workWork at officeVisa sponsorshipWork visa- ...Site Reliability Engineer Qualifications: ~10+ years of overall experience in IT including, with hands-on Development... ...role ~ Experience with Enterprise Cloud transformation efforts ~ Experience with SRE principles and transformation ~3+ years of experience...CloudTemporary workImmediate start
$160k - $210k
...growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure... ...environment as we continue expanding our hybrid cloud footprint. Past that, we are a rapidly... ...closely with our datacenter-focused SRE, so while deep, hands-on AWS expertise...CloudWork at officeImmediate startRemote workWork from home- ...Site Reliability Engineer Description: We are looking for a dynamic Site Reliability Engineer (SRE) with a Top Secret clearance to join our team! The Site Reliability Engineer... ...have a deep understanding of Kubernetes, Cloud Infrastructure, and Infrastructure as Code...Cloud
$75.7k - $136.3k
...building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services. Our SRE teams solve reliability, security, and usability at...CloudWork experience placementWork at office$128.5k - $190k
...winning SaaS platform, Medallia Experience Cloud, leads the market in the management of... ...self. The Role and Team The Site Reliability Engineering organization at Medallia brings... ...engineering practices. Drive adoption of SRE principles, reliability standards, and...CloudTemporary workWork experience placementLocal area$117.2k - $176.7k
...clearance required for this role.Overview of the Role:Join our Site Reliability Engineering (SRE) team, where you'll work alongside Infrastructure and Research & Development (R&D) partners to keep Salesforce cloud services available for customers around the clock. In this...CloudFull timeWork experience placement$121.4k - $218.6k
...building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services. Our SRE teams solve reliability, security, and usability at...CloudWork experience placementWork at office- ...Principal Site Reliability Engineer The Principal Site Reliability Engineer will be a critical technical... ...(IaC), CI/CD architecture, and cloud security with the foundational principles of Site Reliability Engineering (SRE), including defining SLOs, managing error...Cloud
$160k - $180k
...Site Reliability Engineer Location: Hybrid – Washington DC/Virginia/Maryland metro with the ability... ...code deployments in all environments (cloud, on-premises) Implement tools to monitor... ...~4-5 years hands-on experience in a SRE/DevOps role supporting production systems...CloudFull timeTemporary workLocal areaRemote workFlexible hours- ...Quadrant for Supplier Risk Management. Site Reliability Engineer Location: U.S. (Hybrid) This... ...role. You will help stand up the SRE function at Exiger: setting the standards... ...operating systems. ~ Familiarity with cloud platforms (AWS) and secure system integration...CloudWork at officeWork from homeFlexible hours
$90k - $130k
...innovation meets mission. Our AI, cloud, cyber, and modernization... ...an immediate opening for a Site Reliability SME who has hands-on experience... ...as a Cloud Operations Engineer with experience in IT operations... ...Preferred ~5+years of AWS SRE experience ~ Strong Windows...CloudTemporary workWork experience placementImmediate startWorldwide$61k - $101k
...or certification in software engineering concepts, along with 5+ years... ...vector search. We need solid cloud infrastructure experience,... ...or high-scale systems where reliability is critical. Preferred experience... ...Preferred knowledge includes SRE concepts such as SLOs/SLIs,...CloudFull time$105.79k - $141.05k
...high‑performance connectivity across cloud, edge, and AI workloads for... ...demand networking at scale. As Lead SRE, you'll own the reliability of that platform — partnering with operations... ...'ll coordinate across architecture, engineering, and systems development...CloudTemporary workRemote work$207k - $284.9k
...are too, let's talk. Senior Manager, Site Reliability Engineering Secure Every Identity, from AI to Human... ...leader who understands both the SRE discipline and the unique demands of federal... ...~ Experience operating large-scale cloud infrastructure in FedRAMP-authorized environments...CloudPermanent employmentLocal areaWorldwideFlexible hoursDay shift$194k - $267k
...highly technical Staff Observability Site Reliability Engineer with a specialty in Splunk to own and... ...Observability Platform that enables our SRE teams and business partners. You will... ...Experience scaling and managing Splunk Cloud at scale (1000+ SVCs), including Workload...CloudPermanent employmentWork at officeLocal areaWorldwideFlexible hours$174k - $238k
...mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the... ...to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace...CloudLocal areaWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Cloud Site Reliability Engineer (SRE). Be the first to apply!
- senior cloud data engineer Arlington, VA
- cloud engineer Arlington, VA
- cloud engineer remote Arlington, VA
- principal cloud computing engineer Arlington, VA
- big data cloud engineer Arlington, VA
- aws cloud architect Arlington, VA
- aws cloud security engineer Arlington, VA
- cloud developer Arlington, VA
- informatica cloud developer Arlington, VA
- senior cloud infrastructure engineer Arlington, VA


