Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Cloud Site Reliability Engineer (SRE)

$130k - $180k

ECS Federal

Job DescriptionEverforth ECS is seeking a Cloud Site Reliability Engineer (SRE)to work in our Arlington, VA office/remotely. Our Philosophy We believe the job of an SRE is to engineer the cloud to run itself. That means writing software and automation that lets systems detect and recover from failure on their own, rather than relying on someone to notice an alert and manually fix it. When something breaks, self-healing comes first, deep root-cause debugging happens after service is restored, not instead of it. We’re looking for someone who automates the operational task by default, not documents the runbook for doing it by hand. About the Role This role owns reliability and operational readiness for production systems across our federal cloud platform (AWS GovCloud, IL5 zero-trust). You’ll define what “reliable enough” looks like for our services, build the automation that gets us there, and do it all on an infrastructure-as-code (IaC) foundation. Responsibilities Self-Healing Operations Design and build automated remediation so systems detect, respond to, and recover from failure without manual intervention Shift the team’s posture from “is it running, how do we fix it” to “how do we make it fix itself” Automate service restoration first; investigate root cause after Uptime Goals & Reliability Define reasonable, data-driven SLOs and error budgets for critical services alongside the teams that own them Use live metrics to decide what’s “reliable enough” and where to invest next Infrastructure Enforce infrastructure-as-code and configuration-as-code, no manual tech change Own Terraform standards and reusable modules adopted across programs Drive a containerization-first approach with production-scale Kubernetes (multi-tenancy, security policies, advanced scheduling) Set CI/CD and pipeline-as-code standards, including progressive delivery Observability & Incidents Build monitoring, logging, alerting, and tracing (Datadog, Splunk) that gives automation the signal it needs to self-correct Own the incident framework: escalation, restoration, root cause analysis, and post-incident review that closes the loop with more automation Collaboration & Leadership Partner with development and contractor teams leads to embed reliability and automation across the software Mentor engineers toward this same automation-first philosophy Support ATO/RMF and FedRAMP High compliance as it relates to infrastructure and automation Salary Range: $130,000 - $180,000 General Description of BenefitJob RequirementsBachelor’s degree in Computer Science, Information Technology, or related field (or equivalent practical experience) 5+ years of SRE experience (or equivalent), with demonstrated technical leadership 10 years of general work experienceTrack record building self-healing/auto-remediating systems, not just dashboards Jenkins experience Expert AWS knowledge, GovCloud experience strongly preferred Deep Kubernetes and Terraform expertise at production scale Strong software engineering background (Python and/or Go) Experience operating observability platforms (Grafana, Splunk, Prometheus, Loki, etc.) Proven incident command and postmortem experience Strong communication skills across technical and federal leadership audiences Ability to obtain/maintain required government clearance or suitability (CAC/PIV as applicable) US Citizenship Job DetailsJob Type: Full-timeCategory: Software Development, Engineering & ApplicationsSalaried: Salaried

Vacancy posted 16 hours ago
Similar jobs that could be interesting for youBased on the Cloud Site Reliability Engineer (SRE) in Arlington, VA vacancy
  • $114.6k - $252.1k

    Job Title: SRE Platform EngineerJob Category: Information TechnologyTime Type...  ...:CACI is seeking a seasoned Site Reliability (SRE) Platform Engineer to support the Department of Homeland...  ...practices Proven expertise with Azure cloud services including compute, storage,... 
    Cloud
    Contract work
    Work experience placement
    Work at office
    Flexible hours

    CACI International

    Washington DC
    3 days ago
  •  ...MUST HAVES: Minimum of 8 years of experience as a Site Reliability Engineer with a strong understanding of SRE principles for highly scalable and reliable...  ...containers Previous experience with commercial cloud (e.g. AWS, Azure) Can establish and maintain a high... 
    Cloud
    Local area
    Relocation package
    3 days per week

    Beyond SOF

    Vienna, VA
    3 days ago
  •  ...join our talented Team. Job Title: SRE / DevOps Engineer Job Location: Mclean, VA...  ...Job Description: We are seeking a Site Reliability Engineer (SRE) with strong expertise...  ...client ecosystem and deep knowledge of cloud-native infrastructure. The ideal candidate... 
    Cloud

    Ampcus

    McLean, VA
    3 days ago
  •  ...SRE Engineer Location: Washington, DC (Onsite) Duration: 08-17-2026 - 07-30-2027 Key...  ...Pipeline, or Jenkins; provision scalable cloud infrastructure using Terraform,...  ...comprehensive knowledge base articles. Reliability Engineering: Champion SRE metrics including... 
    Cloud

    Georgia IT Inc

    Washington DC
    4 days ago
  • $166k - $220k

     ...expectations. Our systems integration engineers internalize the nuances of each...  ...ABOUT THE JOB We are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team...  ...determine the technical direction of cloud deployments and deliver with speed through... 
    Cloud
    Full time
    Work experience placement

    Mosaic

    Washington DC
    3 days ago
  •  ...interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and...  ...availability, and performance across regions and clouds.Build, automate, and maintain Kubernetes-based infrastructure... 
    Cloud

    Kong

    Washington DC
    1 day ago
  • $115.5k - $164.8k

     ...issues with our ecosystem of devices and cloud software. Like our products, we work...  ...where you matter.Your ImpactAs an engineer on the APX SRE CloudOps team, you will spend a significant...  ...required human intervention with reliable, tested automation. You will also participate... 
    Cloud
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    2 days ago
  • $210k - $230k

     ...GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient...  ..., reliability, and performance optimization across multi-cloud environments. This position is located in Arlington, VA and... 
    Cloud
    Currently hiring
    Remote work

    GovCIO

    Arlington, VA
    2 days ago
  •  ...Job Description Job Description Site Reliability Engineer II Metro DC · Hybrid · 24/7 FedRAMP Operations...  ...and act on it, on a FedRAMP-authorised cloud platform that genuinely can't afford...  .... If you're a couple of years into SRE or cloud support and want your next role... 
    Cloud
    Hourly pay
    Contract work
    For contractors
    Shift work
    Night shift
    Weekend work

    C-Serv

    Arlington, VA
    a month ago
  •  ...Description Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the...  ...Pulumi to provision and manage consistent, version-controlled cloud environments. CI/CD & GitOps: Design and optimize... 
    Cloud
    Local area

    Tiger Analytics Inc.

    Washington DC
    a month ago
  • $103.5k - $150k

     ...winning SaaS platform, Medallia Experience Cloud, leads the market in the management of...  ...self. The Role and Team The Site Reliability Engineering organization at Medallia brings...  ...reliable global SaaS platform. As an SRE II, you will help operate and improve... 
    Cloud
    Temporary work
    Work experience placement
    Local area
    3 days per week

    Medallia

    McLean, VA
    3 days ago
  • $160k - $180k

     ...Site Reliability Engineer Location: Hybrid – Washington DC/Virginia/Maryland metro with the ability...  ...code deployments in all environments (cloud, on-premises) Implement tools to monitor...  ...~4-5 years hands-on experience in a SRE/DevOps role supporting production systems... 
    Cloud
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    RiseMe

    Washington DC
    3 days ago
  •  ...Site Reliability Engineer Qualifications: ~10+ years of overall experience in IT including, with hands-on Development...  ...role ~ Experience with Enterprise Cloud transformation efforts ~ Experience with SRE principles and transformation ~3+ years of experience... 
    Cloud
    Temporary work
    Immediate start

    Samprasoft

    Washington DC
    22 hours ago
  • $169.3k - $304.7k

     ...our Network Infrastructure SRE team! Our team designs, develops...  ..., efficient, scalable, and reliable routing software and...  ...delivery to Akamai Connected Cloud customers. Partner with the...  ...platform. As a Principal Site Reliability Engineer - Network, you will be responsible... 
    Cloud
    Work experience placement
    Work at office

    Akamai

    Washington DC
    4 days ago
  • $107k - $220k

     ...The Site Reliability Engineer (SRE) will ensure the reliability, performance, and scalability of the WDP System. This person will define and track...  ...Requirements SECRET Security Clearance Required Cloud+, GICSP, GSEC, Security+, or SSCP cerification (for Junior... 
    Cloud
    Full time
    Contract work
    Temporary work
    Work at office
    Visa sponsorship
    Work visa

    Avalore, LLC

    Arlington, VA
    2 days ago
  • $100k - $160k

     ...innovative and trusted results. We are looking for a dynamic Site Reliability Engineer (SRE) with a Top Secret clearance to join our team! The Site...  ...candidate will have a deep understanding of Kubernetes, Cloud Infrastructure, and Infrastructure as Code (IaC) practices... 
    Cloud
    Temporary work

    Cathexis

    McLean, VA
    3 days ago
  • $61k - $101k

     ...or certification in software engineering concepts, along with 5+ years...  ...vector search. We need solid cloud infrastructure experience,...  ...or high-scale systems where reliability is critical. Preferred experience...  ...Preferred knowledge includes SRE concepts such as SLOs/SLIs,... 
    Cloud
    Full time

    J.P. Morgan

    Washington DC
    10 days ago
  • $207k - $284.9k

     ...are too, let's talk. Senior Manager, Site Reliability Engineering Secure Every Identity, from AI to Human...  ...leader who understands both the SRE discipline and the unique demands of federal...  ...~ Experience operating large-scale cloud infrastructure in FedRAMP-authorized environments... 
    Cloud
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    3 days ago
  • $174k - $238k

     ...mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the...  ...to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace... 
    Cloud
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    3 days ago
  • $174k - $239k

     ...across functions to drive scale, reliability, and innovation through technology. The Staff Site Reliability Engineer Opportunity Okta Federal,...  ..., improve, and maintain our cloud platform services that help...  ...service and advocate for SRE and DevOps practices across... 
    Cloud
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    26 days ago
  • $182k - $250.8k

     ...too, let's talk. The SRE Leadership Team The...  ...backbone of our platform's reliability and operational...  ...forward-thinking group of engineers and leaders who believe...  ...worldwide. As a Manager, Site Reliability Engineer, you...  ...roles within cloud-native environments, combined... 
    Cloud
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    Washington DC
    3 days ago
  • $160k - $180k

     ...innovation meets mission. Our AI, cloud, cyber, and modernization...  ...an immediate opening for a Site Reliability SME who has hands-on experience...  ...as a Cloud Operations Engineer with experience in IT operations...  ...Preferred ~5+years of AWS SRE experience ~ Strong Windows... 
    Cloud
    Temporary work
    Work experience placement
    Immediate start
    Worldwide

    Credence

    McLean, VA
    2 days ago
  • $204k - $306k

     ...If you are too, let's talk. Manager, Site Reliability Engineering San Francisco, California Secure Every...  ...you’ll be doing  Managing a team of SRE’s supporting various workloads and...  ...constraints. Improve SDLC processes for Cloud infrastructure as a code, including... 
    Cloud
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    Washington DC
    3 days ago
  • $174k - $238k

     ...mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the...  ...to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace... 
    Cloud
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    3 days ago
  •  ...seeking a highly motivated and intellectually curious Senior Site Reliability Engineer to join our team working with a Federal client. The...  ...Preferred Qualifications Experience operating AWS or other cloud platforms Familiarity with Spark, JupyterHub, and Hue in... 
    Cloud
    Remote work

    VALID8 Financial

    Washington DC
    3 days ago
  • $119.8k - $234.7k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...for a Senior Site Reliability Engineer (SRE) to join the Azure Silver and...  ...spanning data transmission across clouds. Our team works across all facets... 
    Cloud
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Reston, VA
    1 day ago
  •  ...SRE/DevOps Engineer Location: McLean, VA (5 Days mandatory) - Only locals/nearby F2F interview mandatory Developing appropriate DevOps channels throughout the organization. Evaluating, implementing and streamlining DevOps practices. Establishing a continuous... 
    Local area

    E-Solutions

    McLean, VA
    22 hours ago
  • $230k - $250k

     ...GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability...  ...Qualifications Bachelor's with 12+ years of infrastructure/cloud engineering experience (or commensurate experience)... 
    Cloud
    Remote work

    GovCIO

    Arlington, VA
    2 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or be willing...  ...ensuring scalability, performance, and reliability across environments. What You’ll Do Design... 
    Cloud
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    4 days ago
  •  ...Site Reliability Engineer III (AI Platform) Location: Mount Laurel, NJ (Onsite) Duration: Contract...  ...seeking a Site Reliability Engineer (SRE) III to support a cutting-edge AI Platform...  ...systems, Kubernetes environments, and cloud-native platforms that power next-... 
    Cloud
    Contract work

    GCS Recruitment

    Laurel, MD
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Cloud Site Reliability Engineer (SRE). Be the first to apply!