Site Reliability Engineer
TalentDome Staffing
Senior Site Reliability Engineer (SRE)
Location: Seattle, hybrid - 2 times a week in the office
Job Type: Full-time, direct hire
Industry: High-Growth Technology / SaaS
About the Role
We are seeking a highly skilled Senior Site Reliability Engineer to drive the reliability, scalability, and performance of our client's production systems. The ideal candidate combines deep software engineering ability with system-level expertise, applying core SRE principles to reduce toil, minimize downtime, and build self-healing infrastructure across complex, high-scale environments.
Key Responsibilities
- Reliability Engineering: Define and drive adoption of SLIs, SLOs, and error budgets across services, using them to guide engineering priorities and release decisions.
- Automation & Toil Reduction: Build tools and automation (Python, Go, Bash) to eliminate manual operational work and enable self-service capabilities for engineering teams.
- Infrastructure as Code (IaC): Design and maintain scalable, resilient infrastructure on AWS using Terraform, CloudFormation, or Pulumi, ensuring consistency and repeatability.
- Observability: Architect monitoring, logging, tracing, and alerting systems (Prometheus, Grafana, Datadog, Splunk, OpenTelemetry) that give clear, actionable signals into system health.
- Incident Management: Act as an incident commander during major outages, lead blameless postmortems, and drive systemic fixes to prevent recurrence.
- Capacity Planning & Performance: Forecast growth, execute load/chaos testing, and tune systems proactively to stay ahead of scaling bottlenecks.
- CI/CD & Deployment Safety: Partner with engineering teams to build safe, progressive delivery pipelines (canary, blue/green, feature flags) using tools like ArgoCD, Jenkins, or GitLab CI.
- Security & Compliance: Embed security best practices into infrastructure and deployment pipelines, including access control, network segmentation, and vulnerability management.
- On-Call Leadership: Participate in and help evolve on-call rotations, escalation policies, and runbooks to reduce alert fatigue and improve response times.
- Mentorship & Culture: Champion SRE best practices across the organization, mentor engineers on reliability thinking, and influence upstream architecture decisions.
Qualifications & Experience
- 6+ years of hands-on experience in Site Reliability Engineering, DevOps, or Backend/Systems Engineering with a track record of owning production reliability at scale.
- Strong Software Engineering Background: Proficiency in Python, Go , or similar languages—focused on building maintainable services and tooling, not just basic scripting.
- AWS Expertise: Deep technical knowledge of AWS core services (EC2, Lambda, RDS, VPC, IAM, S3, Terraform/CloudFormation) alongside cost and performance optimization.
- Linux Systems Mastery: Demonstrated proficiency in performance tuning, kernel/network troubleshooting, and security hardening.
- Container Orchestration: Hands-on experience operating Kubernetes/EKS in production environments at scale.
- SRE Frameworks: Proven experience defining and operationalizing SLIs, SLOs, and error budgets.
- Distributed Systems: Strong understanding of consistency, fault tolerance, failover strategies, and graceful degradation.
- Observability & Incident Command: Experience building full-stack observability pipelines and leading major incident response efforts.
Nice to Have
- Experience with multi-cloud environments (AWS, GCP, Azure).
- Chaos engineering experience (Gremlin, Chaos Mesh, or custom fault-injection tooling).
- Active AWS Certifications (DevOps Engineer Professional, Solutions Architect).
- Familiarity with compliance frameworks (SOC2, ISO 27001, HIPAA).
- Experience with service mesh technologies (Istio, Linkerd).
- Background in internal platform/developer experience (DevEx) teams.
$143k - $194k
...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental... ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and...SuggestedFull timeTemporary workWork experience placementImmediate start- Company DescriptionComtech LLC is a woman-owned small business focused on delivering end-to-end solutions and products. Since 1998, we have successfully serviced enterprises across the public and private sectors, and the Department of Defense. Our services span all aspects...Suggested
$134.25k - $214.8k
...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance...SuggestedWork experience placementWork at officeRemote workFlexible hours$95k - $134k
...disrupting the industry it helped build. Job Application Deadline: 10/31/2026 The Opportunity DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work hybrid in Seattle, WA or Portland, OR Candidate profile DAT...SuggestedTemporary workFor contractorsWork experience placementWork at officeLocal areaImmediate startFlexible hours- ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers... ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them...SuggestedWorldwideHome officeFlexible hours
- ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems. You'll work on observability, incident... ...management Nice to have Experience with chaos engineering Knowledge of distributed systems Background in high-scale...Remote workFlexible hours
- ...This is an engineering-first Senior SRE role. We’re looking for senior engineers who have: Built and shipped significant backend... ...services end-to-end in production (design → launch → on-call → reliability improvements) Led incident response and driven durable...
- ...to our people, and the incredible connections we get to make in every community we are in. About this team Site Reliability Engineering We are looking for a motivated engineer to join the Foundations team which is responsibility for observability and...
- ...A Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of an organization's software systems and cloud infrastructure. The role combines software engineering with IT operations to automate processes, monitor system health...
$160k - $250k
...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content... ...machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering...- Job Title Technical/Functional Skills: Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities: Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper...
$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As...Work at officeLocal areaRemote workWorldwideFlexible hours- ...where you can push the limits of what's possible.As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... .... These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare...
$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$204k - $306k
...all in on this mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure Every Identity, from... ...in our San Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and provisions millions...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week$55k - $151.47k
...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our...Full timeH1b$194k - $267k
...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... .... These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare...
$94k - $142.3k
...on Salesforce, customer and partner enablement, applications engineering, infrastructure, collaboration, enterprise operations,... ...Salesforce products delivered globally, at scale, sustainably.As a Site Reliability Operations Engineer you'll be part of our internal DET Site...Full timeShift work- ...you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... .... These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare...
- ...Lead Software Engineer We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology, Infrastructure Platforms team, you are an...
$232k - $319k
...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and... ...enabled with self-serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and...Permanent employmentLocal areaWorldwideFlexible hours- ...Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core... ...not about building features, but about engineering the resilience and performance of the... ...to maintain system stability.As a Site Reliability Engineer, you will be on the...
- ...where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap... ...in advancing F5’s AI capabilities, ensuring enterprise-grade reliability by leveraging hardware acceleration, designing scalable...Full timeLocal areaImmediate start
- SRE / DevOps EngineerSeattle based client. Seattle-WA (3 days onsite).U.S. Citizens and those authorized to work in the U.S. are encouraged to apply. We are unable to sponsor currently.Must Have SkillsF5 load balancerChef and Terraform - must have most critical skillVMware...
$165k - $206k
...coordinate support and resolve platform issues across CPU/radio SoCs, MCU/PIC, NPU/GPU, and peripheral devices.Support hardware engineering teams with deep technical debugging and contribute to OS/platform modernization efforts.What You’ll NeedBasic Qualifications:Bachelor...Full timeTemporary workWork at officeImmediate startVisa sponsorshipWork visa$160k - $200k
...change and achieving remarkable growth in a rapidly evolving industry. Now, we're growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve service management across Cognitiv. Our immediate challenge is to scale...Work at officeImmediate startRemote workWork from home$143.7k - $194.4k
As a Software Development Engineer II in the SPCL (Selling Partner Close Looping) team, you will design and build innovative solutions... ...Debug and resolve production issues while maintaining service reliability* Monitor system performance and optimize for scale and...InternshipFlexible hours- Role: Sr. Software Engineer, Embedded Systems Control Type: Direct Hire (Full-time) Location: Seattle, WA (Onsite)About: As a Sr. Software... ...might include going to the farm - to ensure our customers have reliable and safe products.What you’ll do:Partner with Engineering teams...Full timeFor contractors
$145k - $200k
...children, and more.The RoleWe are seeking an experienced Software Engineer to join a newly-formed team focused on developing and deploying... ...capabilities, performance, and reliabilityEngineer scalable, reliable, and fail-safe systems capable of functioning in high-stakes...Full timeWork experience placementWork at officeRemote workWork from homeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!


