Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE)

$170k - $230k

Mithril

Site Reliability Engineer (SRE)

Palo Alto / San Francisco Bay Area

About Mithril

Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community, including LG AI Research, Saronic, and the Broad Institute (among many others). Founded by a former Google DeepMind research scientist and Stanford CS PhD, Mithril has raised $80M across seed and Series A funding led by Sequoia Capital & Lightspeed Venture Partners. Platform revenue has grown >6x over the past year, and we were recently recognized by Fast Company as the 8th Most Innovative Company in Artificial Intelligence for 2026.

Our engineering team is lean and high-impact. This role is a core infrastructure hire that will shape how Mithril scales its platform across a heterogeneous, multi-cloud environment.

About the Opportunity

You will be a core contributor to the stability and performance of Mithril's global GPU orchestration platform. This is not a 'keep the lights on' role — you will build the automation, observability, and tooling that allows Mithril to coordinate advanced compute across multiple cloud providers at scale, ensuring customers have fast, reliable access to the infrastructure they need.

You will work directly with our founding team on high-impact infrastructure decisions — from SLO design to capacity orchestration — with clear visibility into the technical and commercial dynamics of the business.

What Makes This Role Different

Most SRE roles at this stage are primarily reactive — on-call, incident management, keeping services stable. At Mithril, you'll also be building infrastructure that directly shapes how our marketplace operates: how supply is sourced, allocated, and monitored across providers. The systems you build will have a direct line to customer experience and company revenue.

You'll have real ownership, a short feedback loop to leadership, and the opportunity to define how infrastructure engineering operates as the company scales.

Core Responsibilities

Platform reliability and infrastructure automation are the primary focus of this role (~70–75% of time).

Reliability & SLOs

  • Implement and own SLIs and SLOs across Mithril's API layer and internal orchestration services to ensure customer commitments are met.
  • Partner with Product and Platform teams to ensure new features are designed for operability, reliability, and performance from the start.
  • Participate in an on-call rotation; drive root cause analysis (RCA) for production incidents and implement durable fixes to prevent recurrence.

Observability & Monitoring

  • Build and maintain dashboards, alerts, and distributed tracing within Mithril's monitoring stack to provide high-granularity visibility into our multi-cloud infrastructure.
  • Own the tooling that surfaces supply-side signal — GPU availability, provider health, reservation fill rates — to both engineering and operations teams.

Infrastructure as Code & Automation

  • Develop and maintain Terraform/Pulumi modules and Kubernetes configurations to manage Mithril's growing multi-cloud provider footprint.
  • Write clean, maintainable Python (or Go) to automate repetitive operational tasks — from provider API reconciliation to automated health checks and capacity rebalancing.

Capacity Support

  • Assist in managing GPU capacity across providers, ensuring the marketplace can dynamically respond to supply fluctuations and customer demand signals.
  • Surface capacity constraints and failure modes early; contribute to the tooling that enables Mithril to make fast, data-driven supply decisions.
Requirements

The profile we're hiring for combines strong systems instincts with the ability to build and own tooling end-to-end. Candidates should be able to point to infrastructure or automation they've built that is still running in production.

  • 3+ years of experience in SRE, Production Engineering, or Infrastructure roles at a high-growth technology company.
  • Hands-on Kubernetes experience: comfortable managing clusters, deployments, and troubleshooting production incidents in a multi-tenant environment.
  • Cloud proficiency in at least one major provider (AWS, GCP, or Azure), including practical understanding of cloud networking fundamentals (VPC, DNS, load balancing, security groups).
  • Coding ability: proficiency in Python or equivalent (Go, Rust, etc.) — you build tools and services, not just scripts. Willing to pick up new languages as needed.
  • Linux fundamentals: strong command of Linux systems, TCP/IP networking, and security best practices.
  • Disciplined troubleshooter: calm under pressure during production incidents, with a rigorous approach to RCA and long-term remediation.
  • Clear communicator: able to document processes and explain technical trade-offs to engineering and non-engineering teammates.
Nice to Have
  • Experience with GPU/TPU-accelerated workloads or AI/ML infrastructure.
  • Exposure to multi-cloud deployments or niche/specialized cloud providers (e.g., CoreWeave, Lambda Labs, Nebius).
  • Familiarity with distributed systems concepts — service discovery, circuit breakers, consensus protocols.
  • Production experience with Prometheus, Grafana, or OpenTelemetry.
  • Prior experience in a high-growth startup environment where infrastructure scope expands faster than headcount.
Benefits
  • Health, dental, and vision coverage for you and your dependents
  • 401k Plan with 4% company match
  • 21 days of PTO & 14 company holidays; including 2 floating holidays
Salary Range Information

In consideration of market analysis and various pertinent factors, the remuneration bracket for this role is set between $170,000 and $230,000. Nevertheless, adjustments beyond this range could be warranted for candidates whose qualifications substantially deviate from those delineated in the job description.

In-Office Requirement

At Mithril, we take our work extremely seriously, though not always ourselves. We recognize that we are striving to achieve something substantial—an all-too-rare and elusive counterfactual contribution. Our work is not easy, so we seek out any lever that can accelerate our progress and increase the likelihood of realizing our full ambitions. Working collaboratively in person is one such lever.

Our headquarters is in Palo Alto (very near the Caltrain station) and we recently opened a second office in the Jackson Square area of San Francisco. We expect team members to primarily work from their local office (Palo Alto or SF), with everyone gathering at HQ a minimum one day a week while our team remains small and cross-collaboration is critical.

This approach is built on trust. We take our mission seriously and are committed to fostering an environment where you can make impactful decisions and drive success. We also understand that life can present challenges, and if extenuating circumstances arise, we're here to support you.

Ultimately, we believe this guidance helps us be as effective as possible while maintaining the spirit of teamwork and flexibility.

Equal Opportunity Employer

Mithril maintains a strict commitment to Equal Opportunity employment practices. All applicants are evaluated without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

We emphasize that candidates need not fulfill every expectation listed to be eligible for this position. Our objective is to cultivate a diverse team encompassing a spectrum of backgrounds, experiences, and skill sets.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) in Palo Alto, CA vacancy
  • $175k - $229k

     ...DevOps Engineer Instrumental builds the manufacturing acceleration platform behind the...  ...Requirements: ~5 or more years of DevOps or SRE experience deploying and operating...  ...KPIs to ensure ongoing performance, reliability and efficiency. ~ Network/application... 
    Suggested

    Instrumental Inc

    Palo Alto, CA
    2 days ago
  • $170k - $250k

     ...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000–$250,000 + Competitive Equity We're representing a rapidly... 
    Suggested
    Work at office
    Visa sponsorship
    Flexible hours

    Recruiting from Scratch

    Palo Alto, CA
    2 days ago
  •  .... Overview We are seeking a highly motivated Systems Reliability Engineer (SRE) to lead the design and implementation of operational excellence...  ...company supporting sensitive and cleared workforces. The Site Reliability Engineer (SRE) - SecOps will embrace our... 
    Suggested
    For contractors
    Work at office
    Flexible hours

    Arkenstone

    Menlo Park, CA
    6 days ago
  • $100k - $200k

    OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Suggested
    Full time

    OPPO

    Palo Alto, CA
    3 days ago
  •  ...Title: Site Reliability Engineer (SRE) Location: Location: Sunnyvale, CA (3x/ week onsite) Contract Responsibilities: Engage with our product teams to understand requirements, design and implement resilient and scalable infrastructure... 
    Suggested
    Contract work

    AceStack LLC

    Sunnyvale, CA
    4 days ago
  • $101k - $161k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...  ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine strong software... 

    Arista Networks

    Santa Clara, CA
    20 hours ago
  •  ...Position: Site Reliability Engineering (SRE) Location: Santa Clara, CA (Onsite) Duration: W2 / C2C Contract Experience: 10+ Years Job Description: • WS application and CI/CD pipelines, Microsoft Server admin and workload support (Data Center and AWS)... 
    Contract work
    Immediate start

    Syntricate Technologies

    Santa Clara, CA
    1 day ago
  • $61k - $101k

     ...,000 per year Requirements: We require formal training or certification in site reliability engineering, along with 5+ years of hands-on experience. We need advanced knowledge of SRE culture and principles, with proven ability to apply them in an application or... 
    Full time

    J.P. Morgan

    Palo Alto, CA
    6 days ago
  •  ...development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes...  ...DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    Remote job
    For contractors

    YO AI Labs

    Palo Alto, CA
    17 days ago
  • $165k - $190k

     ...term growth and IPO readiness.About the DevOps / SRE TeamThe DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and...  ...complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate with... 
    Work from home

    Obsidian Security

    Palo Alto, CA
    4 days ago
  • $70 - $100 per hour

     ...Job Title: Cloud SRE Engineer - Mandarin Bilingual Position Type: Contract (12 months) Location: Palo Alto, CA Salary Rate: $7...  ...team is looking for a skilled Cloud SRE Engineer to own the reliability, stability, and continuous improvement of core cloud services... 
    Hourly pay
    Contract work
    Temporary work
    Work experience placement

    IntelliPro Group Inc.

    Palo Alto, CA
    more than 2 months ago
  • $207k - $300k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.Own end-to-end availability...  ...or Engineering.Experience with Large Language Model.Site Reliability Engineering (SRE) combines software and systems engineering to build and... 

    Google

    Sunnyvale, CA
    2 days ago
  • $100k - $200k

    A leading technology firm in Palo Alto is seeking a skilled Site Reliability Engineer (SRE) to ensure the stability and performance of application systems. Responsibilities include managing cloud platforms, implementing CI/CD pipelines, and maintaining Linux systems. The... 

    OPPO

    Palo Alto, CA
    3 days ago
  • $140k - $165k

     ...Site Reliability Engineer Instrumental builds the manufacturing acceleration platform behind the world's most complex electronics. We capture...  ...high-growth B2B SaaS environment. Experience implementing SRE practices such as SLIs, SLOs, and error budgets. Experience... 

    Instrumental Inc

    Palo Alto, CA
    1 day ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally...  ...yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer...  ...open-source projects, particularly in SRE, observability, or AI/ML domains, and certifications... 

    JP Morgan Chase

    Palo Alto, CA
    2 days ago
  • $222k - $300.5k

     ...OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational...  .... The Fintech Platform Systems Engineering team builds and operates the AWS-based...  ...engineering, product, security, and other SRE/infrastructure leaders across Intuit to... 
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    20 hours ago
  • $262k - $364k

     ...within the AViD ecosystem have reliability and uptime appropriate to...  ...and performance.Build creative engineering solutions to operations and infrastructure...  ...and contribute to the cross-SRE AI Ops program, driving the...  ...in a strategic way.Site Reliability Engineering (SRE)... 

    Google

    Mountain View, CA
    1 day ago
  •  ...SRE Engineer St Louis, MO (Onsite from day 1) Client Required Skills: • Bachelor's Degree in Computer Science, Computer Systems, Information Technology or related. Equivalent experience is acceptable. • Experience with web applications and distributed systems... 

    Omega Solutions

    Santa Clara, CA
    1 day ago
  • $200k - $260k

     ...every company. About the Role: Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical...  ...and eliminating work through automation. On the SRE team, you'll have the opportunity to manage the complex challenges... 
    Work at office
    Home office
    Flexible hours

    Glean.info

    Mountain View, CA
    3 days ago
  • $207k - $300k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.Own end-to-end availability...  ...or Engineering.1 year of people management experience. Site Reliability Engineering (SRE) combines software and systems engineering to build and... 

    Google

    Mountain View, CA
    1 day ago
  • $252k - $308k

     ...Staff Site Reliability Engineer Mountain View, US About EarnIn As one of the first pioneers of earned wage access, our passion at EarnIn...  ...heroics, tribal knowledge, manual investigation, or isolated SRE expertise. We must embed reliability practices that scale across... 
    Full time
    Work at office
    2 days per week

    Earnin

    Mountain View, CA
    5 days ago
  •  ...Enterprise Technologies Inc. is a recognized provider of professional IT Consulting services in the US. We are actively seeking SRE Devops Engineer Fulltime Role for one of our direct client. Role: SRE Devops Engineer Location :- Santa Clara,CA (Remote... 
    Full time
    Local area
    Remote work

    Rootshell Enterprise Technologies

    Santa Clara, CA
    1 day ago
  •  ...everyone. Role Summary We are seeking an experienced Site Reliability Engineer to help design, build, and operate the infrastructure that...  ...of relevant experience in Platform Engineering, DevOps, or SRE roles. ~ Proficiency with Infrastructure as code, preferably... 
    Full time
    Contract work

    Rivian and Volkswagen Group Technologies

    Palo Alto, CA
    2 days ago
  • A leading tech company is seeking a Site Reliability Engineer for their U.S. Data Security division. You will improve service lifecycle and ensure data services are reliable and scalable. The ideal candidate has a Bachelor's degree in Computer Science and experience with... 

    TikTok

    Mountain View, CA
    5 days ago
  •  ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2 Positions) We are...  ...platforms. This role is focused on infrastructure, reliability, and automation , with Java exposure as a supporting skill.... 

    Eitacies Inc

    Santa Clara, CA
    15 days ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining,...  ...background in infrastructure automation, system reliability, and a SRE mindset of continuous improvement.Key Responsibilities:... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    4 days ago
  •  ...simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure....  ...Who You AreExperienced Architect: 5+ years of experience in SRE, DevOps, or Systems Engineering, with a proven track record... 
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    2 days ago
  • $128k - $216k

     ...millions of times a day - quickly, reliably, and securely. Any time you...  ...at Fiserv.Job TitleSr. Site Reliability EngineerAbout CloverClover...  ...Senior Site Reliability Engineer do at Fiserv?As a Senior Site...  ...Design Reviews, mentor teams on SRE principles, and bridge the gap... 
    Full time
    Worldwide

    Fiserv

    Sunnyvale, CA
    2 days ago
  • $148k - $235.75k

     ...on the world.Join our team of innovative engineers who are building an AI Data Center AIOps...  ...that turns raw, high-volume telemetry into reliable, job-centric insights and automation for...  ...operating production distributed systems as SRE/DevOps/Platform Ops.Proven ownership of... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $165k - $280k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most... 
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!