Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Cloud Site Reliability Engineer (SRE)

$130k - $180k

ECS Federal

Job DescriptionEverforth ECS is seeking a Cloud Site Reliability Engineer (SRE)to work in our Arlington, VA office/remotely. Our Philosophy We believe the job of an SRE is to engineer the cloud to run itself. That means writing software and automation that lets systems detect and recover from failure on their own, rather than relying on someone to notice an alert and manually fix it. When something breaks, self-healing comes first, deep root-cause debugging happens after service is restored, not instead of it. We’re looking for someone who automates the operational task by default, not documents the runbook for doing it by hand. About the Role This role owns reliability and operational readiness for production systems across our federal cloud platform (AWS GovCloud, IL5 zero-trust). You’ll define what “reliable enough” looks like for our services, build the automation that gets us there, and do it all on an infrastructure-as-code (IaC) foundation. Responsibilities Self-Healing Operations Design and build automated remediation so systems detect, respond to, and recover from failure without manual intervention Shift the team’s posture from “is it running, how do we fix it” to “how do we make it fix itself” Automate service restoration first; investigate root cause after Uptime Goals & Reliability Define reasonable, data-driven SLOs and error budgets for critical services alongside the teams that own them Use live metrics to decide what’s “reliable enough” and where to invest next Infrastructure Enforce infrastructure-as-code and configuration-as-code, no manual tech change Own Terraform standards and reusable modules adopted across programs Drive a containerization-first approach with production-scale Kubernetes (multi-tenancy, security policies, advanced scheduling) Set CI/CD and pipeline-as-code standards, including progressive delivery Observability & Incidents Build monitoring, logging, alerting, and tracing (Datadog, Splunk) that gives automation the signal it needs to self-correct Own the incident framework: escalation, restoration, root cause analysis, and post-incident review that closes the loop with more automation Collaboration & Leadership Partner with development and contractor teams leads to embed reliability and automation across the software Mentor engineers toward this same automation-first philosophy Support ATO/RMF and FedRAMP High compliance as it relates to infrastructure and automation Salary Range: $130,000 - $180,000 General Description of BenefitJob RequirementsBachelor’s degree in Computer Science, Information Technology, or related field (or equivalent practical experience) 5+ years of SRE experience (or equivalent), with demonstrated technical leadership 10 years of general work experienceTrack record building self-healing/auto-remediating systems, not just dashboards Jenkins experience Expert AWS knowledge, GovCloud experience strongly preferred Deep Kubernetes and Terraform expertise at production scale Strong software engineering background (Python and/or Go) Experience operating observability platforms (Grafana, Splunk, Prometheus, Loki, etc.) Proven incident command and postmortem experience Strong communication skills across technical and federal leadership audiences Ability to obtain/maintain required government clearance or suitability (CAC/PIV as applicable) US Citizenship Job DetailsJob Type: Full-timeCategory: Software Development, Engineering & ApplicationsSalaried: Salaried

Vacancy posted 9 days ago
Similar jobs that could be interesting for youBased on the Cloud Site Reliability Engineer (SRE) in Arlington, VA vacancy
  •  ...Site Reliability Engineer (SRE) Reston, VA Site Reliability Engineer (SRE) Position: Site Reliability Engineer (SRE) Work Authorization...  ...Hands-On experience creating, configuring and maintaining cloud-based applications and infrastructure for the rapid development... 
    Cloud
    Contract work

    Knack Solutions

    Reston, VA
    3 days ago
  •  ...join our talented Team. Job Title: SRE / DevOps Engineer Job Location: Mclean, VA...  ...Job Description: We are seeking a Site Reliability Engineer (SRE) with strong expertise...  ...client ecosystem and deep knowledge of cloud-native infrastructure. The ideal candidate... 
    Cloud

    Ampcus

    McLean, VA
    4 days ago
  •  ...Title: Sr. IT Application Solutions Architect /SRE Engineer Important Note : We have shifted to adopting SAFe and 1. Encourage...  ...DevSecOps Engineer to lead the integration of security into our cloud-native development and operations workflows. This role requires... 
    Cloud
    For contractors
    Shift work

    Lumen Solutions Group, Inc.

    Washington DC
    5 days ago
  •  ...SRE Engineer Location: Washington, DC (Onsite) Duration: 08-17-2026 - 07-30-2027 Key...  ...Pipeline, or Jenkins; provision scalable cloud infrastructure using Terraform,...  ...comprehensive knowledge base articles. Reliability Engineering: Champion SRE metrics including... 
    Cloud

    Staffing the Universe

    Washington DC
    4 days ago
  •  ...DescriptionThe AI Inference Engineer plays a critical role in the...  ...capabilities, ensuring enterprise-grade reliability by leveraging hardware...  ...Docker, Kubernetes, and cloud platforms such as AWS, GCP, and...  ...kernels). Background in MLOps or SRE roles focused on high-... 
    Cloud
    Full time
    Local area
    Immediate start

    F5 Networks

    Washington DC
    15 days ago
  • $95k - $171k

     ...Do you want to build your SRE career on one of the most exciting platforms in cloud computing? Join the Akamai...  ...infrastructure, Kubernetes, and ensuring reliability for AI workloads within...  ...inference platform. As an Site Reliability Engineer II, you will be responsible... 
    Cloud
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Washington DC
    4 days ago
  •  ...Site Reliability Engineer Location- Wilmington De, Washington DC, Dallas, TX (Onsite Position) Full...  ...Minimum of 5 years of experience in an SRE role, with specific experience in...  ...Kubernetes. Strong understanding of cloud platforms (AWS preferred, Azure, GCP).... 
    Cloud
    Full time

    Yochana

    Washington DC
    2 days ago
  •  ...Description Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the...  ...Pulumi to provision and manage consistent, version-controlled cloud environments. CI/CD & GitOps: Design and optimize... 
    Cloud
    Local area

    Tiger Analytics Inc.

    Washington DC
    a month ago
  • $119.8k - $234.7k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-...  ...: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft...  ...the limit thinking in a cloud-enabled world.Microsoft’s Azure Data engineering team is leading the...  ...is what sets the SRE’s in Cosmos DB apart. If... 
    Cloud
    3 days per week

    Microsoft

    Washington DC
    3 days ago
  • $210k - $230k

     ...GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient...  ..., reliability, and performance optimization across multi-cloud environments. This position is located in Arlington, VA and... 
    Cloud
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    a month ago
  • $112k - $179k

     ...seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers...  ...applications and infrastructure. The SRE will drive automation initiatives,...  ...operations, and security teams.Support cloud and hybrid-cloud environments.QualificationsRequired... 
    Cloud
    Contract work
    Worldwide
    Shift work

    Peraton Corporation

    Washington DC
    a month ago
  •  ...Site Reliability Engineer (SRE) Randstad is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our client in the Washington...  ...as-Code (IaC), incident response, and deep experience with cloud platforms, preferably AWS. The SRE will collaborate across... 
    Cloud

    Software Technology Inc

    Washington DC
    5 days ago
  • $166k - $258k

     ...for this position.Nordstrom is looking for a Senior Engineer 2 to join our Site Reliability Engineering (SRE) team — and we think that person could be you.You'...  ...languages such as Java, Go, PythonExperience with cloud providers such as AWS, GCP, and AzureStrong networking... 
    Cloud
    Full time
    Work at office

    Nordstrom

    Washington DC
    3 days ago
  • $107k - $220k

     ...The Site Reliability Engineer (SRE) will ensure the reliability, performance, and scalability of the WDP System. This person will define and track...  ...Requirements SECRET Security Clearance Required Cloud+, GICSP, GSEC, Security+, or SSCP cerification (for Junior... 
    Cloud
    Full time
    Contract work
    Temporary work
    Work at office
    Visa sponsorship
    Work visa

    Avalore, LLC

    Arlington, VA
    2 days ago
  •  ...Site Reliability Engineer Qualifications: ~10+ years of overall experience in IT including, with hands-on Development...  ...role ~ Experience with Enterprise Cloud transformation efforts ~ Experience with SRE principles and transformation ~3+ years of experience... 
    Cloud
    Temporary work
    Immediate start

    Samprasoft

    Washington DC
    1 day ago
  • $160k - $210k

     ...growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure...  ...environment as we continue expanding our hybrid cloud footprint. Past that, we are a rapidly...  ...closely with our datacenter-focused SRE, so while deep, hands-on AWS expertise... 
    Cloud
    Work at office
    Immediate start
    Remote work
    Work from home

    Cognitiv

    Washington DC
    26 days ago
  •  ...Site Reliability Engineer Description: We are looking for a dynamic Site Reliability Engineer (SRE) with a Top Secret clearance to join our team! The Site Reliability Engineer...  ...have a deep understanding of Kubernetes, Cloud Infrastructure, and Infrastructure as Code... 
    Cloud

    STEM Solutions

    McLean, VA
    1 day ago
  • $75.7k - $136.3k

     ...building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services. Our SRE teams solve reliability, security, and usability at... 
    Cloud
    Work experience placement
    Work at office

    Akamai

    Washington DC
    4 days ago
  • $128.5k - $190k

     ...winning SaaS platform, Medallia Experience Cloud, leads the market in the management of...  ...self. The Role and Team The Site Reliability Engineering organization at Medallia brings...  ...engineering practices. Drive adoption of SRE principles, reliability standards, and... 
    Cloud
    Temporary work
    Work experience placement
    Local area

    Medallia

    McLean, VA
    2 days ago
  • $117.2k - $176.7k

     ...clearance required for this role.Overview of the Role:Join our Site Reliability Engineering (SRE) team, where you'll work alongside Infrastructure and Research & Development (R&D) partners to keep Salesforce cloud services available for customers around the clock. In this... 
    Cloud
    Full time
    Work experience placement

    Salesforce

    Herndon, VA
    a month ago
  • $121.4k - $218.6k

     ...building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services. Our SRE teams solve reliability, security, and usability at... 
    Cloud
    Work experience placement
    Work at office

    Akamai

    Washington DC
    4 days ago
  •  ...Principal Site Reliability Engineer The Principal Site Reliability Engineer will be a critical technical...  ...(IaC), CI/CD architecture, and cloud security with the foundational principles of Site Reliability Engineering (SRE), including defining SLOs, managing error... 
    Cloud

    Software Technology Inc

    Washington DC
    5 days ago
  • $160k - $180k

     ...Site Reliability Engineer Location: Hybrid – Washington DC/Virginia/Maryland metro with the ability...  ...code deployments in all environments (cloud, on-premises) Implement tools to monitor...  ...~4-5 years hands-on experience in a SRE/DevOps role supporting production systems... 
    Cloud
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    Fortress Information Security

    Washington DC
    4 days ago
  •  ...Quadrant for Supplier Risk Management. Site Reliability Engineer Location: U.S. (Hybrid) This...  ...role. You will help stand up the SRE function at Exiger: setting the standards...  ...operating systems. ~ Familiarity with cloud platforms (AWS) and secure system integration... 
    Cloud
    Work at office
    Work from home
    Flexible hours

    Exiger

    McLean, VA
    1 day ago
  • $90k - $130k

     ...innovation meets mission. Our AI, cloud, cyber, and modernization...  ...an immediate opening for a Site Reliability SME who has hands-on experience...  ...as a Cloud Operations Engineer with experience in IT operations...  ...Preferred ~5+years of AWS SRE experience ~ Strong Windows... 
    Cloud
    Temporary work
    Work experience placement
    Immediate start
    Worldwide

    Credence

    McLean, VA
    13 days ago
  • $61k - $101k

     ...or certification in software engineering concepts, along with 5+ years...  ...vector search. We need solid cloud infrastructure experience,...  ...or high-scale systems where reliability is critical. Preferred experience...  ...Preferred knowledge includes SRE concepts such as SLOs/SLIs,... 
    Cloud
    Full time

    J.P. Morgan

    Washington DC
    6 days ago
  • $105.79k - $141.05k

     ...high‑performance connectivity across cloud, edge, and AI workloads for...  ...demand networking at scale. As Lead SRE, you'll own the reliability of that platform — partnering with operations...  ...'ll coordinate across architecture, engineering, and systems development... 
    Cloud
    Temporary work
    Remote work

    Lumen Inc

    Washington DC
    5 days ago
  • $207k - $284.9k

     ...are too, let's talk. Senior Manager, Site Reliability Engineering Secure Every Identity, from AI to Human...  ...leader who understands both the SRE discipline and the unique demands of federal...  ...~ Experience operating large-scale cloud infrastructure in FedRAMP-authorized environments... 
    Cloud
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    21 days ago
  • $194k - $267k

     ...highly technical Staff Observability Site Reliability Engineer with a specialty in Splunk to own and...  ...Observability Platform that enables our SRE teams and business partners. You will...  ...Experience scaling and managing Splunk Cloud at scale (1000+ SVCs), including Workload... 
    Cloud
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    21 days ago
  • $174k - $238k

     ...mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the...  ...to build highly reliable, scalable, and secure cloud services that our customers can trust. We embrace... 
    Cloud
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    21 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Cloud Site Reliability Engineer (SRE). Be the first to apply!