Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE)

$170k - $250k

Recruiting from Scratch

Site Reliability Engineer (SRE)

Location: San Francisco, CA / Palo Alto, CA
Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised)
Office Type: Onsite (4 Days Per Week)
Salary: $170,000-$250,000 + Competitive Equity
Company Description

We're representing a rapidly growing AI infrastructure company building a next-generation GPU cloud platform for enterprises, startups, and AI researchers. Their platform provides flexible access to GPU compute through intelligent reservation, marketplace, and consumption models that help customers optimize performance, availability, and cost.

Backed by Sequoia Capital and Lightspeed with more than $80 million in funding, the company has achieved 6x revenue growth over the past year. As demand for AI infrastructure accelerates, they're investing heavily in reliability engineering to build the automation, observability, and platform infrastructure that powers their multi-cloud GPU marketplace at scale.
What You Will Do
  • Design, build, and own the observability platform supporting a large-scale, multi-cloud GPU infrastructure.
  • Develop monitoring, distributed tracing, dashboards, and alerting systems using modern observability tooling.
  • Define and implement SLIs, SLOs, and operational metrics across customer-facing APIs and internal platform services.
  • Build automation that eliminates repetitive operational work and improves platform reliability.
  • Develop production tooling in Python or Go for infrastructure management, health checks, reconciliation, and capacity optimization.
  • Design and maintain Infrastructure-as-Code using Terraform, Pulumi, and Kubernetes.
  • Improve platform resiliency through incident response, root cause analysis, and long-term reliability improvements.
  • Partner closely with Platform, Product, and Engineering teams to ensure new services are designed for operational excellence.
  • Help establish infrastructure engineering standards, reliability practices, and operational processes as the company scales.
  • Participate in production on-call rotations while continuously reducing operational burden through automation.
Ideal Background
  • 3-10 years of experience in Site Reliability Engineering, Production Engineering, Infrastructure Engineering, or Platform Engineering.
  • Strong experience building production automation and operational tooling rather than solely responding to incidents.
  • Proven experience designing and operating large-scale Kubernetes environments.
  • Strong cloud infrastructure experience across AWS, GCP, Azure, or multi-cloud environments.
  • Experience designing distributed systems with a strong understanding of networking fundamentals.
  • Proficiency with Python and/or Go for building production-grade infrastructure tooling.
  • Experience implementing observability platforms using Prometheus, Grafana, OpenTelemetry, or similar technologies.
  • Strong understanding of Linux systems, containers, Docker, and production operations.
  • Excellent communication skills with the ability to collaborate across engineering teams.
Preferred
  • Experience supporting AI infrastructure, GPU clusters, machine learning platforms, or accelerated compute environments.
  • Familiarity with Terraform, Pulumi, Infrastructure-as-Code, and cloud automation.
  • Experience designing reliability standards, operational playbooks, and incident management processes.
  • Background at high-growth startups or major cloud infrastructure organizations.
  • Strong understanding of distributed systems, capacity planning, and performance optimization.
  • Experience building greenfield infrastructure rather than maintaining legacy systems.
  • Passion for automation, reducing operational toil, and continuously improving developer experience.
  • Ability to thrive in fast-paced startup environments with significant ownership and autonomy.
Compensation and Benefits
  • Base salary: $170,000-$250,000.
  • Competitive equity package.
  • Visa transfer sponsorship available.
  • Four-day onsite schedule across San Francisco and Palo Alto offices (all engineers collaborate in Palo Alto on Mondays).
  • Opportunity to help define the reliability and operational foundation of one of the fastest-growing AI infrastructure platforms.
  • Significant ownership over observability, automation, and production infrastructure.
  • Work alongside experienced engineers solving large-scale distributed systems and cloud infrastructure challenges.
  • Join a high-growth, venture-backed company building the infrastructure powering the next generation of AI applications.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) in San Francisco, CA vacancy
  •  ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge... 
    Suggested
    Work at office
    Weekend work

    Fluix AI

    San Francisco, CA
    2 days ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,... 
    Suggested
    Work at office
    Local area
    1 day per week

    Mithril

    San Francisco, CA
    2 days ago
  •  ...startups across the US. We’re building a pool of world-class Site Reliability Engineers for current roles and for upcoming opportunities. You will...  ...into one of our partner startups or added to our vetted SRE network for future projects. This role is ideal for engineers... 
    Suggested
    Local area

    Breakout Tools

    San Francisco, CA
    4 days ago
  • $163.71k - $306k

     ...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system...  ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially... 
    Suggested

    Retool

    San Francisco, CA
    3 days ago
  • DescriptionWe are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco, California. This role will focus on improving service health, refining observability, and... 
    Suggested

    Robert Half

    San Francisco, CA
    4 days ago
  •  ...better than we found it. The Apple Service Engineering (ASE) team builds and provides systems...  ...Compute team is looking for a senior SRE software engineer to own the technical direction...  ...infrastructure, strengthen the reliability of our Kubernetes services, and engage with... 

    Socket.dev

    San Francisco, CA
    5 days ago
  • $300k

     ...experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and...  .... Skills / Must Have: ~7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting... 
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services...  ...define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role includes... 

    Socket.dev

    San Francisco, CA
    5 days ago
  • $165k - $225.6k

     ...we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability...  ...engineering teams to champion DevOps and SRE best practices, deliver excellent internal... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $117k - $209.33k

     ...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,...  ...cloud services for Autodesk GovCloud products.As part of a new SRE team supporting Autodesk GovCloud, you will have a unique... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    4 days ago
  • $113.4k - $162k

     ...conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd,...  ...practices. Contribute to the design and implementation of new SRE best practices.You'll be a great fit if you have:Experienced... 
    Temporary work

    TextNow

    San Francisco, CA
    1 day ago
  •  ...what’s next.About the teamThe Engineering team at Airwallex is a diverse...  ...working together to build scalable, reliable, and secure products that...  ...to grow without borders.Our SRE team is breaking new engineering...  ....What you’ll doAs a Senior Site Reliability Engineer, you’ll work... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    3 days ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the... 

    Alembic

    San Francisco, CA
    3 days ago
  • $148.5k - $223.9k

     ...future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts...  ...and our customers protected. The ExperienceAs an SRE, you will be a technical leader of the team driving... 
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    2 days ago
  • $152.5k - $205k

     ...is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate...  ...public-cloud environments. This role is for an experienced SRE or infrastructure engineer who enjoys solving hard distributed... 
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  •  ...OpportunityTo achieve our ambitious goals, we’re looking for an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth.... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    3 days ago
  • $167.7k - $245.2k

     ...portfolios Your ImpactThe FedRAMP SRE team is focused on our...  ...effective.We’re looking for talented engineers with a software or operations...  ...teams to ensure the reliability, performance and security of...  ...Please see the Cisco careers site to discover more benefits and... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    2 days ago
  • $150k

     ...Site Reliability Engineer San Francisco, CA About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security... 

    VantageScore®

    San Francisco, CA
    3 days ago
  •  ...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built... 
    Relocation package

    1872 Consulting

    San Francisco, CA
    2 days ago
  •  ...healthcare, we'd love to meet you. Apply now to join our growing team. About the Role Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating... 
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    2 days ago
  •  ...Engineering Hiring Sprint We're growing our engineering team and are accelerating hiring...  ...Engineers Database Engineers Site Reliability Engineers Extensibility API Engineers...  ...infrastructure, platform engineering, SRE, or DevOps. ~ Hands-on ownership of Kubernetes... 
    Work at office
    Local area
    Flexible hours

    Airbyte

    San Francisco, CA
    3 days ago
  •  ...would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet,...  ...Role We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability, observability... 
    Remote work
    1 day per week

    Runloop AI, Inc

    San Francisco, CA
    2 days ago
  • $155k - $222.6k

     ....S. soil. Meet the Team The SRE Fleet team is responsible for...  ...cloud platform. As a team of six engineers distributed across the US,...  ...strong focus on automation, reliability, and operational excellence....  ...~2+ years of experience in Site Reliability Engineering, DevOps... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Cisco

    San Francisco, CA
    3 days ago
  •  ...builds the platforms and tooling that help engineering teams develop, deploy, and operate...  ...default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll...  ...experience in backend systems, SRE, or platform engineering roles.Proven track... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    2 days ago
  •  ...SRE Location: San Francisco, CA (5 Days In-Office) You are the infrastructure...  ...treatment. What We Look for in a Great Engineer You have the intensity and...  ...feature release while maintaining the highest reliability. DevX Support: Support Developer Experience... 
    Work at office

    Latent

    San Francisco, CA
    4 days ago
  •  ...customers live their lives. A bank for all of us.About the roleVaro’s SRE team is well established, designing, building, and running large...  ...build and maintain a rapid-feedback platform that enables our engineers to accomplish their own goals instead of creating friction.... 
    For contractors

    Varo Money

    San Francisco, CA
    5 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support...  ...alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    5 days ago
  • $220k - $235k

     ...a strategic, high-output Staff/Senior Staff SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role, you...  ...technical leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud... 
    Full time
    Contract work
    Work at office

    Ironclad

    San Francisco, CA
    2 days ago
  • $160k - $200k

    Job Purpose:BTIG seeks a DevOps/Site Reliability Engineer to join our technology team. This role is central to improving developer velocity by handling...  ...& Qualifications:• 3-5+ years of experience in a DevOps, SRE, platform engineering, or production support role• Strong... 
    Full time

    BTIG

    San Francisco, CA
    1 day ago
  • $204k - $306k

     ...mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco,...  ...Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and...  ...What you’ll be doing Managing a team of SRE’s supporting various workloads and teams... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!