Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior/ Staff Site Reliability Engineer (SRE)

Robert Half

DescriptionWe are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco, California. This role will focus on improving service health, refining observability, and partnering with engineering teams to build systems that perform consistently under real-world demand. The ideal candidate brings deep production experience, a strong automation mindset, and a practical approach to incident response and continuous improvement.Responsibilities:• Establish measurable reliability standards for critical services by creating and maintaining service indicators, objectives, and error budget practices.• Take ownership of production stability by monitoring uptime, latency, and availability, and driving improvements that reduce operational risk.• Lead live incident response efforts, coordinate troubleshooting during outages, and ensure issues are resolved efficiently and thoroughly.• Run blameless post-incident reviews, document findings clearly, and track corrective actions through completion.• Design and enhance observability across logs, metrics, and distributed tracing using tools such as Datadog, CloudWatch, Grafana, OpenTelemetry, and Sentry.• Improve alert quality and dashboard design so engineering teams can quickly identify meaningful system issues without unnecessary noise.• Evaluate system behavior under load, uncover performance constraints, and recommend changes that improve scalability and resource efficiency.• Build automation and internal tooling that streamline operational work, strengthen deployment safety, and support incident management, debugging, and capacity planning.• Contribute to infrastructure and delivery workflows across AWS, Terraform, Ansible, Linux, and GitHub Actions with a focus on dependable releases and resilient systems.• Partner with security and compliance stakeholders to support operational standards, audit readiness, and the integration of monitoring into broader engineering practicesRequirements• 7+ years of experience in Site Reliability Engineering, infrastructure engineering, or a closely related production environment role.• Strong background operating, supporting, and troubleshooting distributed systems at scale.• Hands-on experience with observability platforms such as Datadog, Grafana, OpenTelemetry, CloudWatch, or similar tools.• Proven involvement in on-call operations, incident management, and reliability-focused problem resolution.• Practical experience defining and using SLIs, SLOs, and error budgets to guide service reliability decisions.• Familiarity with AWS environments, including serverless and container-based architectures.• Experience working with relational databases such as Postgres and performance analysis in production systems.• Ability to write automation scripts or lightweight tooling in languages such as Python or Bash, with strong judgment around failure modes and system designJob typePerm

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior/ Staff Site Reliability Engineer (SRE) in San Francisco, CA vacancy
  • $167.7k - $245.2k

     ...requiring approximately 2 days per week on-site at Cisco offices in either San...  ...behave as intended, improving reliability and reducing risks. This unified approach...  ...observability and control.As a Senior Site Reliability Engineer (SRE), you will build, operate, and continuously... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    San Francisco, CA
    4 days ago
  •  ...We are seeking a Sr. Site Reliability Engineer to join our team and run critical infrastructure for our blockchain and web applications. You’ll learn...  ...to streamline development processes. DevOps Engineer/SRE Transitioning to Blockchain An experienced DevOps Engineer... 
    Senior
    Remote work

    Blockchain Works

    San Francisco, CA
    1 day ago
  •  ...better than we found it. The Apple Service Engineering (ASE) team builds and provides systems...  ...The ASE Compute team is looking for a senior SRE software engineer to own the technical...  ...infrastructure, strengthen the reliability of our Kubernetes services, and engage... 
    Senior

    Socket.dev

    San Francisco, CA
    2 days ago
  • $300k

     ...experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and...  .... Skills / Must Have: ~7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting... 
    Senior
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services...  ...define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role includes... 
    Senior

    Socket.dev

    San Francisco, CA
    2 days ago
  • $186.9k - $267.7k

     ...requiring approximately 2 days per week on-site at Cisco offices in either San...  ...behave as intended, improving reliability and reducing risks. This unified approach...  ...enhanced observability and control.As a Staff Site Reliability Engineer (SRE), you will provide technical... 
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    San Francisco, CA
    12 hours ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the... 
    Senior

    Alembic

    San Francisco, CA
    12 hours ago
  • $152.5k - $205k

     ...everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and...  ...public-cloud environments. This role is for an experienced SRE or infrastructure engineer who enjoys solving hard distributed... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    4 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support...  ...alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    2 days ago
  • $117k - $209.33k

     ...73Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure...  ...services for Autodesk GovCloud products.As part of a new SRE team supporting Autodesk GovCloud, you will have a unique... 
    Senior
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    1 day ago
  •  ...what’s next.About the teamThe Engineering team at Airwallex is a...  ...together to build scalable, reliable, and secure products that empower...  ...to grow without borders.Our SRE team is breaking new engineering...  ...services.What you’ll doAs a Senior Site Reliability Engineer, you’ll... 
    Senior
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    12 hours ago
  • $165k - $225.6k

     ...we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability...  ...engineering teams to champion DevOps and SRE best practices, deliver excellent internal... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago
  • $148.5k - $223.9k

     ...are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with...  ...cloud and our customers protected. The ExperienceAs an SRE, you will be a technical leader of the team driving... 
    Senior
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    4 days ago
  • $165k - $241.4k

     ...portfolios Your ImpactThe FedRAMP SRE team is focused on our...  ...effective.We’re looking for talented engineers with a software or operations...  ...teams to ensure the reliability, performance and security of...  ...Please see the Cisco careers site to discover more benefits and... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    4 days ago
  • $220k - $235k

     ...We are seeking a strategic, high-output Staff/Senior Staff SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role...  ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud... 
    Senior
    Full time
    Contract work
    Work at office

    Ironclad

    San Francisco, CA
    4 days ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates... 
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    12 hours ago
  •  ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer...  ...compliance, and uptime requirements. ~ Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident... 
    Senior

    Kody

    San Francisco, CA
    a month ago
  • $165k - $241.4k

     ...within Cisco’s Networking, Security, Collaboration, and Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale,... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    1 day ago
  •  ...Invisible Technologies is looking for a Principal Software Engineer (SRE/DevOps) to work remotely. The ideal candidate will possess dual expertise in application engineering and infrastructure, contributing to a variety of technical initiatives. This role includes overseeing... 
    Remote work

    Invisible Technologies

    San Francisco, CA
    1 day ago
  • $180.5k - $236.91k

    Hi, we're Oscar. We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering team.Oscar...  ....You will report into a Staff/Senior Staff Engineer.Work Location...  ...technical domains such as DevOps, site reliability, and cloud best practicesLead the... 
    Senior
    Full time
    Work at office
    Remote work

    Oscar Health Insurance

    San Francisco, CA
    4 days ago
  • $15k

     ...beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to...  ...work to provide the best experience to our technical staff. You will leverage IaC, Automation, and SRE principles... 
    Senior
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    3 days ago
  •  ...A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves... 
    Senior

    Breakout Tools

    San Francisco, CA
    5 days ago
  •  ...getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'... 
    Senior

    Alembic

    San Francisco, CA
    2 days ago
  •  ...Partner with software developers, platform engineers, and IT staff to improve system design, operability,...  ...requirements, service quality, reliability, security, and compliance needs. Drive...  ...Skills Required: 8+ years of experience in Site Reliability Engineering, DevOps,... 
    Senior
    Work at office
    Remote work

    GrabJobs

    San Francisco, CA
    2 days ago
  • $215k - $275k

     ...by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing... 
    Senior
    Work at office

    Anyscale

    San Francisco, CA
    2 days ago
  • $175k - $250k

     .../yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting...  ...ensuring scalability, performance, and reliability across environments. What You’ll... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    1 day ago
  • $250k

     ...across Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves... 
    Senior
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...is currently Tuesday. Lambda's Storage Engineering team is the backbone behind our world-...  ...workloads in the industry, which means reliability and performance aren't just goals—they're...  ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc... 
    Senior
    Full time
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    1 day ago
  • $210.8k - $272.8k

    About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user... 
    Senior
    Local area

    Thumbtack

    San Francisco, CA
    12 hours ago
  •  ...builds the platforms and tooling that help engineering teams develop, deploy, and operate...  ...default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll...  ...professional experience in backend systems, SRE, or platform engineering roles.Proven... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior/ Staff Site Reliability Engineer (SRE). Be the first to apply!