Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Full-time

TalentDome Staffing

Senior Site Reliability Engineer (SRE)

Location: Seattle, hybrid - 2 times a week in the office

Job Type: Full-time, direct hire

Industry: High-Growth Technology / SaaS

About the Role

We are seeking a highly skilled Senior Site Reliability Engineer to drive the reliability, scalability, and performance of our client's production systems. The ideal candidate combines deep software engineering ability with system-level expertise, applying core SRE principles to reduce toil, minimize downtime, and build self-healing infrastructure across complex, high-scale environments.

Key Responsibilities

  • Reliability Engineering: Define and drive adoption of SLIs, SLOs, and error budgets across services, using them to guide engineering priorities and release decisions.
  • Automation & Toil Reduction: Build tools and automation (Python, Go, Bash) to eliminate manual operational work and enable self-service capabilities for engineering teams.
  • Infrastructure as Code (IaC): Design and maintain scalable, resilient infrastructure on AWS using Terraform, CloudFormation, or Pulumi, ensuring consistency and repeatability.
  • Observability: Architect monitoring, logging, tracing, and alerting systems (Prometheus, Grafana, Datadog, Splunk, OpenTelemetry) that give clear, actionable signals into system health.
  • Incident Management: Act as an incident commander during major outages, lead blameless postmortems, and drive systemic fixes to prevent recurrence.
  • Capacity Planning & Performance: Forecast growth, execute load/chaos testing, and tune systems proactively to stay ahead of scaling bottlenecks.
  • CI/CD & Deployment Safety: Partner with engineering teams to build safe, progressive delivery pipelines (canary, blue/green, feature flags) using tools like ArgoCD, Jenkins, or GitLab CI.
  • Security & Compliance: Embed security best practices into infrastructure and deployment pipelines, including access control, network segmentation, and vulnerability management.
  • On-Call Leadership: Participate in and help evolve on-call rotations, escalation policies, and runbooks to reduce alert fatigue and improve response times.
  • Mentorship & Culture: Champion SRE best practices across the organization, mentor engineers on reliability thinking, and influence upstream architecture decisions.

Qualifications & Experience

  • 6+ years of hands-on experience in Site Reliability Engineering, DevOps, or Backend/Systems Engineering with a track record of owning production reliability at scale.
  • Strong Software Engineering Background: Proficiency in Python, Go , or similar languages—focused on building maintainable services and tooling, not just basic scripting.
  • AWS Expertise: Deep technical knowledge of AWS core services (EC2, Lambda, RDS, VPC, IAM, S3, Terraform/CloudFormation) alongside cost and performance optimization.
  • Linux Systems Mastery: Demonstrated proficiency in performance tuning, kernel/network troubleshooting, and security hardening.
  • Container Orchestration: Hands-on experience operating Kubernetes/EKS in production environments at scale.
  • SRE Frameworks: Proven experience defining and operationalizing SLIs, SLOs, and error budgets.
  • Distributed Systems: Strong understanding of consistency, fault tolerance, failover strategies, and graceful degradation.
  • Observability & Incident Command: Experience building full-stack observability pipelines and leading major incident response efforts.

Nice to Have

  • Experience with multi-cloud environments (AWS, GCP, Azure).
  • Chaos engineering experience (Gremlin, Chaos Mesh, or custom fault-injection tooling).
  • Active AWS Certifications (DevOps Engineer Professional, Solutions Architect).
  • Familiarity with compliance frameworks (SOC2, ISO 27001, HIPAA).
  • Experience with service mesh technologies (Istio, Linkerd).
  • Background in internal platform/developer experience (DevEx) teams.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Washington DC vacancy
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Costa... 
    Suggested
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    1 day ago
  • $112k - $179k

     ...delivery of system, network, software, and security solutions.About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems... 
    Suggested
    Contract work
    Worldwide
    Shift work

    Peraton Corporation

    Washington DC
    4 days ago
  • $135k - $154k

     ...where you matter.Your ImpactAs a contributor in the APX platform engineering organization on the CloudNet team, you are passionate about...  .... You are also obsessed about achieving the high quality and reliability our customers demand. You will work closely with sovereign... 
    Suggested
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    1 day ago
  • $230k - $250k

     ...GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is... 
    Suggested
    Remote work

    Govcio

    Arlington, VA
    4 days ago
  • $165k - $270k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts.... 
    Suggested
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    1 day ago
  • $210k - $230k

     ...GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    1 day ago
  •  ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS... 

    Kong

    Washington DC
    10 hours ago
  •  ...customers depend on every day. We're hiring a senior, hands-on engineer to own the reliability, availability, security, and performance of that platform...  ...Who you are: ~5+ years of hands-on Cloud Operations and Site Reliability Engineering, operating production-scale SaaS (... 
    Full time

    MangoApps

    Washington DC
    3 days ago
  •  ...Lead Site Reliability Engineer Defense Tech / National Security US Defense Tech Startup The Company Early-stage defense technology leader building modern software for air-gapped, high-side, and accredited environments. Backed by a $99M defense contract to... 
    Full time
    Contract work

    Attis

    Washington DC
    3 days ago
  •  ...Washington D.C., District of Columbia, United States About the job Sr. Site Reliability Engineer Our Client is currently hiring a full-time Sr. Site Reliability Engineer (SRE), who will play a vital role in continuously driving improvements in observability, performance... 
    Full time
    Currently hiring
    3 days per week

    CruitZi

    Washington DC
    4 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or...  ...while ensuring scalability, performance, and reliability across environments. What You’ll Do Design, build... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    3 days ago
  • $141.8k - $195k

     ...best work, grow fast, and bring their full selves to the herd. Why You'll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl... 
    Temporary work
    Remote work

    Cribl

    Washington DC
    5 days ago
  •  ...plan Paid maternity leave 401(k) Get notified when a new job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago Seattle, WA $115,000.00-$175,000.00 5 months ago Senior ServiceNow... 
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    3 days ago
  •  ...Site Reliability Engineer III (AI Platform) Location: Mount Laurel, NJ (Onsite) Duration: Contract Experience: 4+ years About the Role We are seeking a Site Reliability Engineer (SRE) III to support a cutting-edge AI Platform Engineering team responsible... 
    Contract work

    GCS Recruitment

    Laurel, MD
    2 days ago
  •  ...ears, and hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical platform running...  ...of deep technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Twenty Inc.

    Arlington, VA
    1 day ago
  •  ...technology solutions using a tailored Agile methodology. We are seeking a highly motivated and intellectually curious Senior Site Reliability Engineer to join our team working with a Federal client. The position will be a remote role open to US citizens residing in the... 
    Remote work

    VALID8 Financial

    Washington DC
    2 days ago
  • $130k - $160k

     ...customers. Our customers and partners trust us to deliver reliable, first-to-market solutions and safeguard the data we receive...  ...pioneers of market-changing solutions. We are seeking a Site Reliability Engineer to design, build, and maintain highly available systems and... 
    Local area

    LE038 Second Sight Solutions, LLC

    Washington DC
    2 days ago
  •  ...Join the Site Reliability Engineering (SRE) team to support high-engagement multimodal applications. You will be responsible for managing infrastructure, observability solutions, and platform automation while ensuring the reliability and security of hosted applications... 

    NextGen | GTA: A Kelly Telecom Company

    Laurel, MD
    3 days ago
  •  ...Washington, District of Columbia, United States Contractor | On-site Job Description We are seeking an experienced Site Reliability Engineer (SRE) to help build and maintain highly reliable, scalable, and secure technology platforms. The SRE will combine software... 
    For contractors

    Mybridge

    Washington DC
    3 days ago
  • $125k - $185k

    Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir builds the world’s leading software for data-driven decisions...  ...locate missing children, and more.The RoleWe’re looking for Site Reliability Engineers who can help us build, operate, and maintain high-... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    1 day ago
  • $125k - $185k

     ...lifesaving drugs, forecast supply chain disruptions, locate missing children, and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for our production... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    4 days ago
  • $182k - $250.8k

     ...Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great...  ...for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    Washington DC
    10 hours ago
  • $174k - $239k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for an experienced Staff TDI Site Reliability... 
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    4 days ago
  • $174k - $238k

     ...work. We're all in on this mission. If you are too, let's talk.The Federal SRE TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable,... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    2 days ago
  • $207k - $284.9k

     ...on this mission. If you are too, let's talk.Senior Manager, Site Reliability EngineeringSecure Every Identity, from AI to HumanIdentity is...  ...mission. If you are too, let's talk.The Federal Operations Engineering GroupOkta's Federal Operations team supports government customers... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    1 day ago
  • $106.5k - $177.5k

     ...Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the... 
    Remote work

    Noctua Technology

    Washington DC
    3 days ago
  •  ...Manager, Site Reliability Engineering Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work... 

    Okta, Inc.

    Washington DC
    2 days ago
  • $140k - $210k

     ...Our Mission As the world’s number 1 job site*, our mission is to help people get jobs. We strive to cultivate an...  ...Comscore, Total Visits, March 2026) Day to Day As an Engineering Manager in Site Reliability Engineering at Indeed, you will manage and grow a team that... 
    Work experience placement
    Local area

    Indeed

    Washington DC
    1 day ago
  • Role Summary The Senior Site Reliability Engineer (SRE) is a hands-on role responsible for the availability, performance, and end-to-end observability of QSR digital platforms across Mobile (iOS/Android), Web, and POS systems. This role is part of the Observability... 
    Flexible hours

    Donato Technologies, Inc

    Washington DC
    2 days ago
  •  ...Job Description Job Description Description: Onsite in Washington, DC   our client seeks a Sr. Site Reliability Engineer III to design, automate, and operate mission-critical systems for federal environments. The role focuses on Kubernetes or VMWare platforms,... 
    Hourly pay
    Permanent employment
    Full time
    Local area
    Immediate start

    Eliassen Group

    Washington DC
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!