Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Site Reliability Engineer [Remote]

$177k - $240k
Full-time

jobgether

United States
  • Remote job

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer based in Austria.

This is a high-impact reliability leadership role within a fully remote engineering organization operating globally.
You will be the first dedicated SRE, helping establish reliability practices across multiple engineering teams and critical production systems.
The role combines hands-on engineering with organization-wide influence, covering observability, incident response, operational readiness, and resilience.
You will work closely with engineering leadership, infrastructure specialists, architects, and product teams to make reliability measurable and actionable.
A major focus will be embedding SRE principles into engineering culture rather than simply owning individual services.
You will also help shape how AI is used for incident investigation, operational tooling, observability, and safe system operations.
The position offers substantial autonomy to define standards, coach engineers, and build practices that scale with the organization.

Accountabilities

  • Define and implement SLIs and SLOs for critical production request paths, ensuring reliability objectives are visible, measurable, reviewed, and connected to engineering decisions.

  • Introduce and champion error budgets as a practical framework for balancing reliability investments with product and feature delivery.

  • Establish and maintain the reliability metrics used by engineering leadership to evaluate progress and identify areas requiring investment.

  • Strengthen the complete incident management lifecycle, including detection, response, communication, escalation, postmortems, and follow-up actions.

  • Improve alert quality, anomaly detection, escalation processes, and shared operational tooling in collaboration with infrastructure teams.

  • Lead reliability assessments for high-risk changes and new services, covering production readiness, capacity, failure modes, rollback strategies, and operational risks.

  • Introduce deliberate failure testing, game days, and chaos exercises to identify weaknesses and validate safe operational limits before incidents occur.

  • Work directly with engineering teams on complex reliability challenges through focused engagements, leaving behind stronger practices and clear ownership.

  • Coach Staff and Lead engineers to become reliability advocates within their respective teams and help establish distributed SRE ownership.

  • Develop lightweight, repeatable operational standards covering production readiness, on-call practices, runbooks, change safety, and service operability.

  • Partner with architects and technical leads to ensure reliability and failure tolerance are incorporated into system design rather than addressed after deployment.

  • Remain hands-on during production incidents and investigations, building tooling, dashboards, automation, and reference implementations where appropriate.

  • Promote effective use of AI for incident investigation, telemetry analysis, postmortem development, runbook creation, observability, and reliability tooling.

  • Help structure operational data, alerts, dashboards, and runbooks so that both engineers and AI agents can safely interpret and act on production signals.

  • Contribute production fixes and improvements directly through code and infrastructure changes rather than limiting the role to recommendations and reviews.

Requirements

  • 10+ years of engineering experience, including at least 3 years in SRE, production engineering, or a reliability-focused Staff Engineer role operating across multiple teams.

  • Demonstrated experience owning reliability at a platform or organizational level rather than only for an individual service.

  • Deep practical experience designing and implementing SLIs, SLOs, and error budgets, including successfully driving adoption across product and engineering teams.

  • Strong incident leadership experience, including managing high-severity, customer-facing incidents and leading effective postmortems that result in measurable improvements.

  • Advanced understanding of distributed-system failure modes, including database and cache saturation, cascading failures, retry storms, capacity constraints, graceful degradation, and load shedding.

  • Strong hands-on experience with Kubernetes, AWS, and modern observability platforms such as Datadog or comparable technologies.

  • Ability to read and write production code in Go, TypeScript, or a similar language, as well as work with infrastructure as code.

  • Demonstrated ability to influence teams without direct authority and successfully change engineering practices across an organization.

  • Strong coaching and mentoring skills, with evidence of developing engineers into effective reliability owners.

  • Exceptional written and verbal communication skills, with the ability to clearly communicate incidents, risks, technical trade-offs, and reliability priorities to both engineers and executives.

  • Strong preference for asynchronous, documented decision-making and clear technical communication.

  • Practical experience using AI tools for incident investigation, telemetry analysis, runbook and postmortem development, and engineering tooling.

  • Understanding of how operational data, alerts, dashboards, and runbooks should be structured to support safe AI-assisted diagnosis and operations.

  • Pragmatic approach to reliability, with the ability to balance operational risk, engineering investment, delivery speed, and business priorities.

  • Experience in fraud detection, identity, payments, or other real-time and adversarial environments is an asset.

  • Experience with multi-region architectures, cell-based architectures, or failure-isolation strategies is a plus.

  • Experience operating Elasticsearch, Redis, DynamoDB, or Kafka at scale and understanding their failure modes is beneficial.

  • Familiarity with FinOps and cloud infrastructure cost-versus-reliability trade-offs is an advantage.

  • Must be authorized to work from the hiring location; visa sponsorship is not provided.

Benefits

  • Fully remote working environment.

  • Opportunity to become the first dedicated Site Reliability Engineer and establish organization-wide reliability practices.

  • High level of autonomy and direct influence over engineering standards, operational practices, and platform reliability.

  • Opportunity to work across multiple engineering teams and critical production systems.

  • Close collaboration with engineering leadership, architects, infrastructure teams, and technical leads.

  • Opportunity to shape AI-assisted reliability practices and the future of production operations.

  • Strong focus on professional growth, technical leadership, coaching, and knowledge sharing.

  • Inclusive, globally distributed engineering environment that values diverse perspectives and backgrounds.

  • For US-based employees, the stated cash compensation range is $177,000–$240,000 USD , with actual offers varying according to factors such as experience, skills, education, certifications, and market conditions. Compensation may differ for other hiring locations.

  • Remote work eligibility is subject to applicable regulatory and security requirements in the candidate's location.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer [Remote] in United States vacancy
  •  ...The Team Platform Engineering is the department within SRE that is responsible for a range...  ...role in developing and maintaining the reliable and globally connected multi-cloud network...  ...Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong... 
    Suggested
    Full time
    Work at office
    Remote work
    Worldwide

    Mongodb

    United States
    7 days ago
  •  ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Overview-The ProCOM team is looking for a Site Reliability Engineering (SRE) who can help us solve problems, build our... 
    Suggested
    Full time
    Part time
    Immediate start
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    4 days ago
  • $96k - $163k

     ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer, Performance Engineering Senior Site Reliability Engineer, Performance Engineering Payment Optimization unifies... 
    Suggested
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    4 days ago
  • $96k - $163k

     ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Overview The BizOps team is looking for a Senior Site Reliability Engineer who can help us solve problems and... 
    Suggested
    Full time
    Part time
    Worldwide
    Flexible hours
    Shift work

    Mastercard

    O Fallon, MO
    4 days ago
  • $76k - $127k

     ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer II Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that... 
    Suggested
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    4 days ago
  • $96k - $163k

     ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that... 
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    4 days ago
  • $25 - $30 per hour

     ...and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineering Intern, Summer 2027 – St. Louis, MO, US Why join Mastercard's internship program? Mastercard's internship program... 
    Full time
    Part time
    Summer work
    Internship
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    4 days ago
  •  ...Mass General Brigham in Massachusetts is seeking an Epic Data Courier Administrator, Systems Engineer to manage data migrations and change control across Epic environments. The role sits at the center of deployment operations and collaborates with Epic analysts and application... 

    Mass General Brigham

    Somerville, MA
    23 hours ago
  • $184k - $264.5k

     ...diverse and supportive environment, where We are seeking a Site Reliability Operations Technical Lead to serve as the senior technical...  ...platform teams, and provides technical leadership to site support engineers. Be responsible for the hardest issues across Active... 
    Permanent employment
    Work at office
    Local area
    Relocation

    NVIDIA Gruppe

    Durham, NC
    1 day ago
  •  ...Lambda Inc. in San Francisco is seeking a Storage Engineer to own the reliability, performance, and capacity of our production storage fleet across multiple data centers, using a software-defined data plane. You will build monitoring, dashboards, and alerting for storage... 

    Lambda

    San Francisco, CA
    23 hours ago
  • $145k - $175k

     ...straightforward communication and clinical domain expertise, Commence cuts straight to better care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and operational health of our mission-critical healthcare data... 
    Full time
    Remote work

    GrabJobs

    Tempe, AZ
    23 hours ago
  •  ...The Site Reliability Engineer (SRE) / Subject Matter Expert (SME) – Computer Systems Engineer/Architect will provide senior-level reach-back expertise to support the reliability, scalability, performance, and operational resilience of the GEOMAP platform in secure cloud... 
    Full time
    Contract work
    For contractors
    For subcontractor
    Remote work

    Diné Development

    United States
    1 day ago
  • $145k - $193k

     ...entertainment, we want to talk to you. About the Role & Team The SRE team at PENN Entertainment is looking for a Senior Site Reliability Engineer to help build and operate the infrastructure behind a large-scale sports betting and media platform. You'll own critical... 
    Remote work

    Penn Interactive

    United States
    2 days ago
  •  ...organization across multiple locations in the US, South America, and India. Location: Remote (US-Based Candidates Only) Site Reliability Engineer II (SRE) Position Overview We are seeking a Site Reliability Engineer II (SRE) to join our growing Site... 
    Remote work
    Flexible hours
    Shift work
    Weekday work

    NationsBenefits, LLC

    United States
    1 day ago
  •  ...About The Role: We're looking for a Senior Site Reliability Engineer to help us mature and scale the infrastructure behind our multi-cloud SaaS platform. Most of our footprint runs on Microsoft Azure, built from the ground up around cloud architecture principles:... 
    Remote work
    Flexible hours

    Dental Intelligence

    United States
    23 hours ago
  •  ...SitusAMC in Annapolis, MD, is seeking a seasoned Cloud Reliability Engineer to lead AWS-based deployments and SaaS reliability initiatives. You will optimize CI/CD pipelines, implement IaC, and drive observability across complex microservices. Join a collaborative... 
    Remote work

    SitusAMC

    United States
    23 hours ago
  • $113.3k - $205.52k

     ...important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development velocity with the reliability our customers depend on. You'll partner with engineering... 
    Work at office
    Remote work
    Worldwide
    Flexible hours

    GrabJobs

    Orlando, FL
    1 day ago
  • $194k - $237k

    ## Principal Site Reliability EngineerApplylocations: Scottsdaletime type: Full timeposted on: Posted 5 Days Agojob requisition id: REQ2026...  ...sponsorship.**Overall Purpose**The Principal Site Reliability Engineer partners with development teams by designing availability and... 
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    Scottsdale, AZ
    23 hours ago
  •  ...troubleshooting, providing rubric-based written feedback. This role requires hands-on Kubernetes expertise in EKS/GKE/AKS or self-managed clusters, with strong scripting in Go, Python, or TypeScript, and ability to document findings clearly for engineering #J-18808-Ljbffr

    Obsidian

    San Francisco, CA
    2 days ago
  • $350k

     ...with leading AI companies and infrastructure providers to build reliable, high-performance platforms supporting next-generation AI workloads. This opportunity is for a Staff Site Reliability Engineer to lead the reliability of large-scale GPU infrastructure, covering... 

    Hamilton Barnes Associates Limited

    San Francisco, CA
    10 hours ago
  • $120k - $175k

     ...level of sports fandom. Ready to reimagine the DFS industry together? We are seeking a highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate about delivering cutting-edge solutions and pushing the boundaries of what's possible.... 
    Full time
    Remote work
    Work visa
    Flexible hours

    GrabJobs

    Greensboro, NC
    4 days ago
  • $200k - $240k

     ...systems across all product teams. You will collaborate closely with engineering leadership, product managers, and cross-functional teams to...  ...and Helm ~ Understand the importance of performant and reliable systems ~ Education - Ideally looking for a B.A. / B.S. degree... 
    Work at office
    Immediate start
    3 days per week

    Altruist

    San Francisco, CA
    4 days ago
  •  ...Site Reliability Engineer Company: Quzara Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: Azure, Terraform, Bicep, Ansible, Azure Monitor, Azure Automation, Azure Policy, Azure Site Recovery, TLS/SSL Requirements: 4+ years in SRE... 
    Full time
    Remote work

    Quzara LLC

    United States
    1 day ago
  • $160k - $180k

     ...big impact. See Arkestro in action at arkestro.com. About the Role Arkestro is hiring for a Senior SRE Engineer to manage our performance and reliability for our software platform and infrastructure. The right candidate will own and develop our infrastructural... 
    Local area
    Remote work

    Arkestro

    United States
    1 day ago
  •  ...JPMorganChase in Seattle seeks a Lead Software Engineer to join the Enterprise Technology, Infrastructure Platforms team. You will act as a core technical contributor, delivering trusted, scalable technology across multiple domains, while guiding AI-assisted engineering... 

    Fairygodboss

    Seattle, WA
    3 days ago
  • $165k - $280k

     ...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARLINK) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s... 
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    InvestedintheMission

    Palo Alto, CA
    23 hours ago
  •  ...Nscale is seeking a Senior Site Reliability Engineer to own the reliability bar for AI infrastructure operations. You will tackle the hardest reliability challenges, mentor peers, and shape SRE practices across the platform. You will work closely with teams running... 

    Nscale

    Seattle, WA
    1 day ago
  •  ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service, from inception and design, through to deployment, operation and refinement. Ensure reliable, fault-tolerant, efficiently scalable and cost-effective data... 

    Tik Tok

    Mountain View, CA
    4 days ago
  •  ...Site Reliability Engineer Company: Crunchafi Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Azure, AKS, Azure Kubernetes Service, Terraform, Bicep, ARM templates, GitHub Actions, Azure DevOps, Kubernetes, Docker, App Insights... 
    Full time
    Remote work

    Crunchafi

    United States
    1 day ago
  •  ...customers globally. As part of the ongoing investment in the reliability and modernization of these systems, client is making a multi-...  ...delivery of change without impact to reliability. Integration Engineering, ensuring that client is consuming the right infrastructure... 
    Work experience placement
    3 days per week

    Syntricate Technologies

    Irving, TX
    10 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer [Remote]. Be the first to apply!