Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Happy Robot

About HappyRobot

HappyRobot is the infrastructure for enterprises to build and orchestrate AI workforces. Our AI workers don't just communicate - they make decisions, take action, and run operations autonomously across voice, email, and enterprise systems. Born in Y Combinator (S23) and backed by a16z and Base10 with over $60M raised, we power critical operations for global enterprises worldwide.

Our platform is battle-tested in the most demanding environments - where AI has real consequences. We started in logistics, built our own voice stack, models, and orchestration layer from the ground up, and are now bringing that infrastructure to every enterprise that runs the real economy. Learn more about our vision in our manifesto.

About the Role

We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You'll own the stability, observability, and debugging workflows that keep our systems running smoothly. You'll be the go-to person for untangling complex failures in real time, designing tools that turn chaos into clarity, and helping us shift from reactive to proactive operations.

This is a high-impact, high-trust role where you'll shape how reliability is done - reducing incident load, building internal tooling, and directly improving developer focus and system uptime. If you love getting to the root of hard problems and making systems (and teams) stronger, this is your moment.

Must-Have
  • 3+ years of hands-on experience debugging production systems (logs, traces, incidents, etc.)
  • Strong problem-solving skills and ability to dive into unfamiliar backend codebases
  • Strong Go and Kubernetes experience.
  • Familiarity with observability and monitoring tools (e.g., Grafana, Prometheus, Sentry)
  • Clear, calm communication under pressure - especially during live incidents
Nice-to-Have
  • Experience working with distributed systems or services at scale
  • Built or maintained internal tooling for on-call teams or reliability workflows
  • Familiarity with deployment pipelines, CI/CD, or infra-as-code
  • Experience improving system observability (e.g., custom metrics, traces, log pipelines)
Why join us?
  • Opportunity to work at a high-growth AI startup , backed by top investors.
  • Fast Growth - Backed by a16z and YC , on track for double-digit ARR .
  • Top-Tier Compensation - Competitive salary + equity in a high-growth startup.
  • Ownership & Autonomy - Take full ownership of projects and ship fast.
  • Work With the Best - Join a world-class team of engineers and builders.

Our Operating Principles

Extreme Ownership

We take full responsibility for our work, outcomes, and team success. No excuses, no blame-shifting - if something needs fixing, we own it and make it better. This means stepping up, even when it's not "your job." If a ball is dropped, we pick it up. If a customer is unhappy, we fix it. If a process is broken, we redesign it. We don't wait for someone else to solve it - we lead with accountability and expect the same from those around us.

Craftsmanship

Putting care and intention into every task, striving for excellence, and taking deep ownership of the quality and outcome of your work. Craftsmanship means never settling for "just fine." We sweat the details because details compound. Whether it's a product feature, an internal doc, or a sales call - we treat it as a reflection of our standards. We aim to deliver jaw-dropping customer experiences by being curious, meticulous, and proud of what we build - even when nobody's watching.

We are "majos"
Be friendly & have fun with your coworkers. Always be genuine & honest, but kind. "Majo" is our way of saying: be a good human. Be approachable, helpful, and warm. We're building something ambitious, and it's easier (and more fun) when we enjoy the ride together. We give feedback with kindness, challenge each other with respect, and celebrate wins together without ego.

Urgency with Focus
Create the highest impact in the shortest amount of time. Move fast, but in the right direction. We operate with speed because time is our most limited resource. But speed without focus is chaos. We prioritize ruthlessly, act decisively, and stay aligned. We aim for high leverage: the biggest results from the simplest, smartest actions. We're running a high-speed marathon - not a sprint with no strategy.

Talent Density and Meritocracy
Hire only people who can raise the average; 'exceptional performance is the passing grade.' Ability trumps seniority. We believe the best teams are built on talent density - every hire should raise the bar. We reward contribution, not titles or tenure. We give ownership to those who earn it, and we all hold each other to a high standard. A-players want to work with other A-players - that's how we win.

First-Principles Thinking
Strip a problem to physics-level facts, ignore industry dogma, rebuild the solution from scratch. We don't copy-paste solutions. We go back to basics, ask why things are the way they are, and rebuild from the ground up if needed. This mindset pushes us to innovate, challenge stale assumptions, and move faster than incumbents. It's how we build what others think is impossible.

The personal data provided in your application and during the selection process will be processed by Happyrobot, Inc., acting as Data Controller.

By sending us your CV, you consent to the processing of your personal data for the purpose of evaluating and selecting you as a candidate for the position. Your personal data will be treated confidentially and will only be used for the recruitment process of the selected job offer.

In relation to the period of conservation of your personal data, these will be eliminated after three months of inactivity in compliance with the GDPR and legislation on the protection of personal data.

If you wish to exercise your rights of access, rectification, deletion, portability or opposition in relation to your personal data, you can do so through View email address on click.appcast.io subject to the GDPR.

For more information, visit

By submitting your request, you confirm that you have read and understood this clause and that you agree to the processing of your personal data as described.
Vacancy posted 19 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in San Francisco, CA vacancy
  • $350k

     ...and novel use-cases. We’re hiring to grow the platform alongside the Tinker community. About the Role We're looking for a Site Reliability Engineer to drive the reliability of Tinker end-to-end. You'll work alongside the engineers building the platform and research... 
    Suggested
    Full time
    Visa sponsorship
    Work visa
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    2 days ago
  • $139.76k - $287.75k

     ...their business.We are seeking a Senior Site ReliabilityEngineer to help operate, scale...  ...will be instrumental in advancing the reliability, scalability, automation, observability,...  ...The ideal candidate is a highly hands-on engineer with strong production experience and a... 
    Suggested
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    2 days ago
  • $148.5k - $223.9k

     ...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations,... 
    Suggested
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    3 days ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Suggested

    Alembic

    San Francisco, CA
    4 days ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,... 
    Suggested
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $113.4k - $162k

     ...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at... 
    Temporary work

    TextNow

    San Francisco, CA
    2 days ago
  • $152.5k - $205k

     ...work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical... 
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  • $117k - $209.33k

    Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    15 hours ago
  • $165k - $241.4k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    3 days ago
  •  ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    4 days ago
  •  ...let’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of...  ..., working together to build scalable, reliable, and secure products that empower businesses...  ...services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work closely... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    4 days ago
  •  ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely...  ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    3 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    1 day ago
  • $140k - $205k

    Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The Senior Technology Site Reliability Engineer (“SRE”) is responsible for ensuring the reliability... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    Weekend work

    Cooley

    San Francisco, CA
    4 days ago
  •  ...A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves... 

    Breakout Tools

    San Francisco, CA
    4 days ago
  • $120k - $168.49k

     ...Site Reliability Engineer, Cloud Infrastructure About Quizlet At Quizlet, our mission is to help every learner achieve their outcomes in the most effective and delightful way. Our $1B+ learning platform serves tens of millions of students every month, including two-thirds... 
    Internship
    Work at office
    3 days per week

    Quizlet

    San Francisco, CA
    4 days ago
  •  ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering...  ...You will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role includes... 

    Socket

    San Francisco, CA
    19 hours ago
  • $155k - $222.6k

     ...global cloud platform. As a team of six engineers distributed across the US, Canada, and the...  ...with a strong focus on automation, reliability, and operational excellence. We are one...  ...Qualifications ~2+ years of experience in Site Reliability Engineering, DevOps,... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Cisco

    Daly City, CA
    19 hours ago
  •  ...Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was first-to...  ...expect us to hit our SLAs. What? We're looking for a Senior Site level Reliability Engineer as part of Infrastructure team to: Own... 
    Contract work
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    IVO Inc

    San Francisco, CA
    2 days ago
  •  ...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built... 
    Relocation package

    1872 Consulting

    San Francisco, CA
    3 days ago
  • $163.71k - $306k

     ...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system...  ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially... 

    Retool

    San Francisco, CA
    4 days ago
  • $86k - $105k

     ...generation of application infrastructure and to be responsible for reliability, automation and scalability using and the latest best...  ...certifications. Minimum of 2 years prior DevOps, software engineering or related experience. Must be able to work different schedules... 
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    San Francisco, CA
    18 hours ago
  • $230k - $310k

     ...millions of daily users while enabling our engineering teams to ship fast. You'll own the...  ...building automation and tooling that improves reliability and partnering with engineering to...  ...What You'll Bring ~5+ years in site reliability engineering, DevOps, or systems... 
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    2 days ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,... 
    Work at office
    Local area
    1 day per week

    Mithril

    San Francisco, CA
    3 days ago
  •  ...human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to...  ...The Role We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible... 
    Remote work
    1 day per week

    Runloop AI, Inc

    San Francisco, CA
    3 days ago
  •  ...Arena Intelligence Engineer Arena Intelligence is looking for an engineer to build the core infrastructure that sits beneath our online...  ...foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at... 
    Permanent employment
    Shift work

    Arena AI

    San Francisco, CA
    2 days ago
  •  ...globe. Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and scalability of our systems. You... 
    Temporary work
    Worldwide

    Air Apps

    San Francisco, CA
    19 hours ago
  •  ...Engineering Hiring Sprint We're growing our engineering team and are accelerating hiring through a focused Engineering Hiring Sprint...  ...: Platform Engineers Database Engineers Site Reliability Engineers Extensibility API Engineers AI Agents Engineers... 
    Work at office
    Local area
    Flexible hours

    Airbyte

    San Francisco, CA
    4 days ago
  • A leading technology firm is looking for a Manager to expand their Cloud Site Reliability team. The ideal candidate will have extensive Linux administration experience, a passion for automation, and be comfortable in a remote, diverse workplace. This position emphasizes... 
    Remote work

    mbhsobana

    San Francisco, CA
    4 days ago
  • $200k - $300k

     ...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture Round — Healthcare, AI Office Type: Onsite Salary: $200K–$300K Company Description We're representing a dynamic company... 
    Work at office

    Recruiting from Scratch

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!