Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

TalentDome Staffing

Senior Site Reliability Engineer (SRE)

Location: Seattle, hybrid - 2 times a week in the office

Job Type: Full-time, direct hire

Industry: High-Growth Technology / SaaS

About the Role

We are seeking a highly skilled Senior Site Reliability Engineer to drive the reliability, scalability, and performance of our client's production systems. The ideal candidate combines deep software engineering ability with system-level expertise, applying core SRE principles to reduce toil, minimize downtime, and build self-healing infrastructure across complex, high-scale environments.

Key Responsibilities

  • Reliability Engineering: Define and drive adoption of SLIs, SLOs, and error budgets across services, using them to guide engineering priorities and release decisions.
  • Automation & Toil Reduction: Build tools and automation (Python, Go, Bash) to eliminate manual operational work and enable self-service capabilities for engineering teams.
  • Infrastructure as Code (IaC): Design and maintain scalable, resilient infrastructure on AWS using Terraform, CloudFormation, or Pulumi, ensuring consistency and repeatability.
  • Observability: Architect monitoring, logging, tracing, and alerting systems (Prometheus, Grafana, Datadog, Splunk, OpenTelemetry) that give clear, actionable signals into system health.
  • Incident Management: Act as an incident commander during major outages, lead blameless postmortems, and drive systemic fixes to prevent recurrence.
  • Capacity Planning & Performance: Forecast growth, execute load/chaos testing, and tune systems proactively to stay ahead of scaling bottlenecks.
  • CI/CD & Deployment Safety: Partner with engineering teams to build safe, progressive delivery pipelines (canary, blue/green, feature flags) using tools like ArgoCD, Jenkins, or GitLab CI.
  • Security & Compliance: Embed security best practices into infrastructure and deployment pipelines, including access control, network segmentation, and vulnerability management.
  • On-Call Leadership: Participate in and help evolve on-call rotations, escalation policies, and runbooks to reduce alert fatigue and improve response times.
  • Mentorship & Culture: Champion SRE best practices across the organization, mentor engineers on reliability thinking, and influence upstream architecture decisions.

Qualifications & Experience

  • 6+ years of hands-on experience in Site Reliability Engineering, DevOps, or Backend/Systems Engineering with a track record of owning production reliability at scale.
  • Strong Software Engineering Background: Proficiency in Python, Go , or similar languages—focused on building maintainable services and tooling, not just basic scripting.
  • AWS Expertise: Deep technical knowledge of AWS core services (EC2, Lambda, RDS, VPC, IAM, S3, Terraform/CloudFormation) alongside cost and performance optimization.
  • Linux Systems Mastery: Demonstrated proficiency in performance tuning, kernel/network troubleshooting, and security hardening.
  • Container Orchestration: Hands-on experience operating Kubernetes/EKS in production environments at scale.
  • SRE Frameworks: Proven experience defining and operationalizing SLIs, SLOs, and error budgets.
  • Distributed Systems: Strong understanding of consistency, fault tolerance, failover strategies, and graceful degradation.
  • Observability & Incident Command: Experience building full-stack observability pipelines and leading major incident response efforts.

Nice to Have

  • Experience with multi-cloud environments (AWS, GCP, Azure).
  • Chaos engineering experience (Gremlin, Chaos Mesh, or custom fault-injection tooling).
  • Active AWS Certifications (DevOps Engineer Professional, Solutions Architect).
  • Familiarity with compliance frameworks (SOC2, ISO 27001, HIPAA).
  • Experience with service mesh technologies (Istio, Linkerd).
  • Background in internal platform/developer experience (DevEx) teams.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Seattle, WA vacancy
  • $143k - $194k

     ...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental...  ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    15 hours ago
  • Company DescriptionComtech LLC is a woman-owned small business focused on delivering end-to-end solutions and products. Since 1998, we have successfully serviced enterprises across the public and private sectors, and the Department of Defense. Our services span all aspects...
    Suggested

    Comtech

    Seattle, WA
    2 days ago
  • $134.25k - $214.8k

     ...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance... 
    Suggested
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Seattle, WA
    15 hours ago
  • $95k - $134k

     ...disrupting the industry it helped build. Job Application Deadline: 10/31/2026 The Opportunity DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work hybrid in Seattle, WA or Portland, OR Candidate profile DAT... 
    Suggested
    Temporary work
    For contractors
    Work experience placement
    Work at office
    Local area
    Immediate start
    Flexible hours

    DAT Freight & Analytics

    Seattle, WA
    15 hours ago
  •  ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers...  ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them... 
    Suggested
    Worldwide
    Home office
    Flexible hours

    Superhuman

    Seattle, WA
    15 hours ago
  •  ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems. You'll work on observability, incident...  ...management Nice to have Experience with chaos engineering Knowledge of distributed systems Background in high-scale... 
    Remote work
    Flexible hours

    ACI Infotech

    Seattle, WA
    3 days ago
  •  ...This is an engineering-first Senior SRE role. We’re looking for senior engineers who have: Built and shipped significant backend...  ...services end-to-end in production (design → launch → on-call → reliability improvements) Led incident response and driven durable... 

    Practice by Numbers

    Bellevue, WA
    15 hours ago
  •  ...to our people, and the incredible connections we get to make in every community we are in. About this team Site Reliability Engineering We are looking for a motivated engineer to join the Foundations team which is responsibility for observability and... 

    Kaav Inc.

    Seattle, WA
    3 days ago
  •  ...A Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of an organization's software systems and cloud infrastructure. The role combines software engineering with IT operations to automate processes, monitor system health... 

    Mybridge

    Seattle, WA
    15 hours ago
  • $160k - $250k

     ...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content...  ...machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering... 

    Hive

    Seattle, WA
    4 days ago
  • Job Title Technical/Functional Skills: Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities: Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper...

    The Dignify Solutions, LLC

    Bellevue, WA
    3 days ago
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    2 days ago
  •  ...where you can push the limits of what's possible.As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,...  .... These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare... 

    JP Morgan Chase

    Seattle, WA
    3 days ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    1 day ago
  • $204k - $306k

     ...all in on this mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure Every Identity, from...  ...in our San Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and provisions millions... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    Bellevue, WA
    15 hours ago
  • $55k - $151.47k

     ...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our... 
    Full time
    H1b

    PwC

    Seattle, WA
    7 hours ago
  • $194k - $267k

     ...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    15 hours ago
  •  ...you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,...  .... These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare... 

    Fairygodboss

    Seattle, WA
    4 days ago
  • $94k - $142.3k

     ...on Salesforce, customer and partner enablement, applications engineering, infrastructure, collaboration, enterprise operations,...  ...Salesforce products delivered globally, at scale, sustainably.As a Site Reliability Operations Engineer you'll be part of our internal DET Site... 
    Full time
    Shift work

    Salesforce

    Seattle, WA
    1 day ago
  •  ...you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,...  .... These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare... 

    Next Frontier Capital

    Seattle, WA
    1 day ago
  •  ...Lead Software Engineer We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology, Infrastructure Platforms team, you are an... 

    Hackajob

    Seattle, WA
    2 days ago
  • $232k - $319k

     ...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and...  ...enabled with self-serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    3 days ago
  •  ...Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core...  ...not about building features, but about engineering the resilience and performance of the...  ...to maintain system stability.As a Site Reliability Engineer, you will be on the... 

    TikTok

    Seattle, WA
    2 days ago
  •  ...where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap...  ...in advancing F5’s AI capabilities, ensuring enterprise-grade reliability by leveraging hardware acceleration, designing scalable... 
    Full time
    Local area
    Immediate start

    F5 Networks

    Seattle, WA
    1 day ago
  • SRE / DevOps EngineerSeattle based client. Seattle-WA (3 days onsite).U.S. Citizens and those authorized to work in the U.S. are encouraged to apply. We are unable to sponsor currently.Must Have SkillsF5 load balancerChef and Terraform - must have most critical skillVMware...

    Georgia IT Inc

    Seattle, WA
    1 day ago
  • $165k - $206k

     ...coordinate support and resolve platform issues across CPU/radio SoCs, MCU/PIC, NPU/GPU, and peripheral devices.Support hardware engineering teams with deep technical debugging and contribute to OS/platform modernization efforts.What You’ll NeedBasic Qualifications:Bachelor... 
    Full time
    Temporary work
    Work at office
    Immediate start
    Visa sponsorship
    Work visa

    Sonos

    Seattle, WA
    3 days ago
  • $160k - $200k

     ...change and achieving remarkable growth in a rapidly evolving industry. Now, we're growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve service management across Cognitiv. Our immediate challenge is to scale... 
    Work at office
    Immediate start
    Remote work
    Work from home

    Cognitiv

    Bellevue, WA
    1 day ago
  • $143.7k - $194.4k

    As a Software Development Engineer II in the SPCL (Selling Partner Close Looping) team, you will design and build innovative solutions...  ...Debug and resolve production issues while maintaining service reliability* Monitor system performance and optimize for scale and... 
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  • Role: Sr. Software Engineer, Embedded Systems Control Type: Direct Hire (Full-time) Location: Seattle, WA (Onsite)About: As a Sr. Software...  ...might include going to the farm - to ensure our customers have reliable and safe products.What you’ll do:Partner with Engineering teams... 
    Full time
    For contractors

    Spectraforce Technologies

    Seattle, WA
    4 days ago
  • $145k - $200k

     ...children, and more.The RoleWe are seeking an experienced Software Engineer to join a newly-formed team focused on developing and deploying...  ...capabilities, performance, and reliabilityEngineer scalable, reliable, and fail-safe systems capable of functioning in high-stakes... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Seattle, WA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!