Site Reliability Engineer
TalentDome Staffing
Senior Site Reliability Engineer (SRE)
Location: Seattle, hybrid - 2 times a week in the office
Job Type: Full-time, direct hire
Industry: High-Growth Technology / SaaS
About the Role
We are seeking a highly skilled Senior Site Reliability Engineer to drive the reliability, scalability, and performance of our client's production systems. The ideal candidate combines deep software engineering ability with system-level expertise, applying core SRE principles to reduce toil, minimize downtime, and build self-healing infrastructure across complex, high-scale environments.
Key Responsibilities
- Reliability Engineering: Define and drive adoption of SLIs, SLOs, and error budgets across services, using them to guide engineering priorities and release decisions.
- Automation & Toil Reduction: Build tools and automation (Python, Go, Bash) to eliminate manual operational work and enable self-service capabilities for engineering teams.
- Infrastructure as Code (IaC): Design and maintain scalable, resilient infrastructure on AWS using Terraform, CloudFormation, or Pulumi, ensuring consistency and repeatability.
- Observability: Architect monitoring, logging, tracing, and alerting systems (Prometheus, Grafana, Datadog, Splunk, OpenTelemetry) that give clear, actionable signals into system health.
- Incident Management: Act as an incident commander during major outages, lead blameless postmortems, and drive systemic fixes to prevent recurrence.
- Capacity Planning & Performance: Forecast growth, execute load/chaos testing, and tune systems proactively to stay ahead of scaling bottlenecks.
- CI/CD & Deployment Safety: Partner with engineering teams to build safe, progressive delivery pipelines (canary, blue/green, feature flags) using tools like ArgoCD, Jenkins, or GitLab CI.
- Security & Compliance: Embed security best practices into infrastructure and deployment pipelines, including access control, network segmentation, and vulnerability management.
- On-Call Leadership: Participate in and help evolve on-call rotations, escalation policies, and runbooks to reduce alert fatigue and improve response times.
- Mentorship & Culture: Champion SRE best practices across the organization, mentor engineers on reliability thinking, and influence upstream architecture decisions.
Qualifications & Experience
- 6+ years of hands-on experience in Site Reliability Engineering, DevOps, or Backend/Systems Engineering with a track record of owning production reliability at scale.
- Strong Software Engineering Background: Proficiency in Python, Go , or similar languages—focused on building maintainable services and tooling, not just basic scripting.
- AWS Expertise: Deep technical knowledge of AWS core services (EC2, Lambda, RDS, VPC, IAM, S3, Terraform/CloudFormation) alongside cost and performance optimization.
- Linux Systems Mastery: Demonstrated proficiency in performance tuning, kernel/network troubleshooting, and security hardening.
- Container Orchestration: Hands-on experience operating Kubernetes/EKS in production environments at scale.
- SRE Frameworks: Proven experience defining and operationalizing SLIs, SLOs, and error budgets.
- Distributed Systems: Strong understanding of consistency, fault tolerance, failover strategies, and graceful degradation.
- Observability & Incident Command: Experience building full-stack observability pipelines and leading major incident response efforts.
Nice to Have
- Experience with multi-cloud environments (AWS, GCP, Azure).
- Chaos engineering experience (Gremlin, Chaos Mesh, or custom fault-injection tooling).
- Active AWS Certifications (DevOps Engineer Professional, Solutions Architect).
- Familiarity with compliance frameworks (SOC2, ISO 27001, HIPAA).
- Experience with service mesh technologies (Istio, Linkerd).
- Background in internal platform/developer experience (DevEx) teams.
- ...enterprise solutions within the Information Security team. The successful candidate will apply an engineering-based approach to solving complex security and reliability challenges, leveraging machine data analytics and log analysis across the enterprise. Designing systems...SuggestedFull timeInternshipSummer internshipWork at officeLocal areaRemote workFlexible hours
$134.25k - $214.8k
...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed... ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,...SuggestedWork experience placementWork at officeRemote work$134.25k - $214.8k
...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance...SuggestedWork experience placementWork at officeRemote workFlexible hours- Company DescriptionComtech LLC is a woman-owned small business focused on delivering end-to-end solutions and products. Since 1998, we have successfully serviced enterprises across the public and private sectors, and the Department of Defense. Our services span all aspects...Suggested
$143k - $194k
...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental... ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and...SuggestedFull timeTemporary workWork experience placementImmediate start- ...customers depend on every day. We're hiring a senior, hands-on engineer to own the reliability, availability, security, and performance of that platform... ...Who you are: ~5+ years of hands-on Cloud Operations and Site Reliability Engineering, operating production-scale SaaS (...Full time
- ...We're seeking an SRE to ensure the reliability and performance of our clients' critical systems. You'll work on observability, incident... ...management Nice to have Experience with chaos engineering Knowledge of distributed systems Background in high-scale...Remote workFlexible hours
- ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers... ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them...WorldwideHome officeFlexible hours
- ...certification), ISO 27001:2005 Information Security Management System (ISMS), and CMMI-DEV Level 3. Job Description Sr. Site Reliability Engineer Location – Seattle, WA Duration – 12 months Interview – in-person if local or Phone + Skype Minimum...Local areaWorldwide
$95k - $134k
...helped build. For more information, visit Job Application Deadline: 10/31/2026 The Opportunity DAT is looking for a Site Reliability Engineer to join our SRE platform team. This position will work hybrid in Seattle, WA or Portland, OR Candidate profile DAT...Temporary workFor contractorsWork experience placementWork at officeLocal areaImmediate startFlexible hours- ...A Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of an organization's software systems and cloud infrastructure. The role combines software engineering with IT operations to automate processes, monitor system health...
- ...This is an engineering-first Senior SRE role. We’re looking for senior engineers who have: Built and shipped significant backend... ...services end-to-end in production (design → launch → on-call → reliability improvements) Led incident response and driven durable...
$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As...Work at officeLocal areaRemote workWorldwideFlexible hours$151.2k - $204.6k
Would you like to be an engineer who builds the systems that power advertising at scale,... ...advertising queries every day, where latency, reliability, and quality translate directly into... ...Development Engineer, operating as a Site Reliability Engineer, to raise the reliability...Flexible hours$194k - $267k
...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...where you can push the limits of what's possible.As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... .... These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare...
$204k - $306k
...all in on this mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure Every Identity, from... ...in our San Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and provisions millions...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week$55k - $151.47k
...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our...Full timeH1b$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...future together. We are responsible for the reliability of all the company's major data warehouse products, services, and query engines. We serve business needs across domains... ...practices, and emerging technologies related to site reliability and infrastructure engineering....
$94k - $142.3k
...on Salesforce, customer and partner enablement, applications engineering, infrastructure, collaboration, enterprise operations,... ...Salesforce products delivered globally, at scale, sustainably.As a Site Reliability Operations Engineer you'll be part of our internal DET Site...Full timeShift work- ...you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... .... These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare...
$145k - $160k
...We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep...Temporary workRemote workFlexible hours- Overview Site Reliability Engineer, Compute - USDS TikTok is the leading destination for short-form mobile video. U.S. Data Security (USDS) is a subsidiary of TikTok in the U.S. This security-first division was created to bring heightened focus and governance to data protection...Work experience placement
- ...you can push the limits of what's possible. As a Lead Software Engineer at JPMorganChase within the Enterprise Technology,... .... These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare...
- The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that power... ...Message Queue. Our work is not about building features, but about engineering the resilience and performance of the underlying platform...
$232k - $319k
...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and... ...enabled with self-serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and...Permanent employmentLocal areaWorldwideFlexible hours- ...Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core... ...not about building features, but about engineering the resilience and performance of the... ...to maintain system stability.As a Site Reliability Engineer, you will be on the...
$120k - $170k
Sr. Manager/Manager Site Reliability Engineering Join to apply for the Sr. Manager/Manager Site Reliability Engineering role at Aritzia Sr. Manager/Manager Site Reliability Engineering 1 day ago Be among the first 25 applicants Join to apply for the Sr. Manager/Manager...Full timeWork at officeRemote workFlexible hours$160k - $180k
...Socure is seeking a Site Reliability Engineer to own AWS infrastructure and improve Kubernetes platforms. Ideal candidates will have deep AWS expertise, strong Kubernetes fundamentals, and the ability to write production-quality code in Go or Python. The position emphasizes...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!



