Site Reliability Engineer (SRE)
$170k - $230kMithril
Site Reliability Engineer (SRE)
Palo Alto / San Francisco Bay Area
About Mithril
Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community, including LG AI Research, Saronic, and the Broad Institute (among many others). Founded by a former Google DeepMind research scientist and Stanford CS PhD, Mithril has raised $80M across seed and Series A funding led by Sequoia Capital & Lightspeed Venture Partners. Platform revenue has grown >6x over the past year, and we were recently recognized by Fast Company as the 8th Most Innovative Company in Artificial Intelligence for 2026.
Our engineering team is lean and high-impact. This role is a core infrastructure hire that will shape how Mithril scales its platform across a heterogeneous, multi-cloud environment.
About the Opportunity
You will be a core contributor to the stability and performance of Mithril's global GPU orchestration platform. This is not a 'keep the lights on' role — you will build the automation, observability, and tooling that allows Mithril to coordinate advanced compute across multiple cloud providers at scale, ensuring customers have fast, reliable access to the infrastructure they need.
You will work directly with our founding team on high-impact infrastructure decisions — from SLO design to capacity orchestration — with clear visibility into the technical and commercial dynamics of the business.
What Makes This Role Different
Most SRE roles at this stage are primarily reactive — on-call, incident management, keeping services stable. At Mithril, you'll also be building infrastructure that directly shapes how our marketplace operates: how supply is sourced, allocated, and monitored across providers. The systems you build will have a direct line to customer experience and company revenue.
You'll have real ownership, a short feedback loop to leadership, and the opportunity to define how infrastructure engineering operates as the company scales.
Core Responsibilities
Platform reliability and infrastructure automation are the primary focus of this role (~70–75% of time).
Reliability & SLOs
- Implement and own SLIs and SLOs across Mithril's API layer and internal orchestration services to ensure customer commitments are met.
- Partner with Product and Platform teams to ensure new features are designed for operability, reliability, and performance from the start.
- Participate in an on-call rotation; drive root cause analysis (RCA) for production incidents and implement durable fixes to prevent recurrence.
Observability & Monitoring
- Build and maintain dashboards, alerts, and distributed tracing within Mithril's monitoring stack to provide high-granularity visibility into our multi-cloud infrastructure.
- Own the tooling that surfaces supply-side signal — GPU availability, provider health, reservation fill rates — to both engineering and operations teams.
Infrastructure as Code & Automation
- Develop and maintain Terraform/Pulumi modules and Kubernetes configurations to manage Mithril's growing multi-cloud provider footprint.
- Write clean, maintainable Python (or Go) to automate repetitive operational tasks — from provider API reconciliation to automated health checks and capacity rebalancing.
Capacity Support
- Assist in managing GPU capacity across providers, ensuring the marketplace can dynamically respond to supply fluctuations and customer demand signals.
- Surface capacity constraints and failure modes early; contribute to the tooling that enables Mithril to make fast, data-driven supply decisions.
Requirements
The profile we're hiring for combines strong systems instincts with the ability to build and own tooling end-to-end. Candidates should be able to point to infrastructure or automation they've built that is still running in production.
- 3+ years of experience in SRE, Production Engineering, or Infrastructure roles at a high-growth technology company.
- Hands-on Kubernetes experience: comfortable managing clusters, deployments, and troubleshooting production incidents in a multi-tenant environment.
- Cloud proficiency in at least one major provider (AWS, GCP, or Azure), including practical understanding of cloud networking fundamentals (VPC, DNS, load balancing, security groups).
- Coding ability: proficiency in Python or equivalent (Go, Rust, etc.) — you build tools and services, not just scripts. Willing to pick up new languages as needed.
- Linux fundamentals: strong command of Linux systems, TCP/IP networking, and security best practices.
- Disciplined troubleshooter: calm under pressure during production incidents, with a rigorous approach to RCA and long-term remediation.
- Clear communicator: able to document processes and explain technical trade-offs to engineering and non-engineering teammates.
Nice to Have
- Experience with GPU/TPU-accelerated workloads or AI/ML infrastructure.
- Exposure to multi-cloud deployments or niche/specialized cloud providers (e.g., CoreWeave, Lambda Labs, Nebius).
- Familiarity with distributed systems concepts — service discovery, circuit breakers, consensus protocols.
- Production experience with Prometheus, Grafana, or OpenTelemetry.
- Prior experience in a high-growth startup environment where infrastructure scope expands faster than headcount.
Benefits
- Health, dental, and vision coverage for you and your dependents
- 401k Plan with 4% company match
- 21 days of PTO & 14 company holidays; including 2 floating holidays
Salary Range Information
In consideration of market analysis and various pertinent factors, the remuneration bracket for this role is set between $170,000 and $230,000. Nevertheless, adjustments beyond this range could be warranted for candidates whose qualifications substantially deviate from those delineated in the job description.
In-Office Requirement
At Mithril, we take our work extremely seriously, though not always ourselves. We recognize that we are striving to achieve something substantial—an all-too-rare and elusive counterfactual contribution. Our work is not easy, so we seek out any lever that can accelerate our progress and increase the likelihood of realizing our full ambitions. Working collaboratively in person is one such lever.
Our headquarters is in Palo Alto (very near the Caltrain station) and we recently opened a second office in the Jackson Square area of San Francisco. We expect team members to primarily work from their local office (Palo Alto or SF), with everyone gathering at HQ a minimum one day a week while our team remains small and cross-collaboration is critical.
This approach is built on trust. We take our mission seriously and are committed to fostering an environment where you can make impactful decisions and drive success. We also understand that life can present challenges, and if extenuating circumstances arise, we're here to support you.
Ultimately, we believe this guidance helps us be as effective as possible while maintaining the spirit of teamwork and flexibility.
Equal Opportunity Employer
Mithril maintains a strict commitment to Equal Opportunity employment practices. All applicants are evaluated without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
We emphasize that candidates need not fulfill every expectation listed to be eligible for this position. Our objective is to cultivate a diverse team encompassing a spectrum of backgrounds, experiences, and skill sets.
$175k - $229k
...DevOps Engineer Instrumental builds the manufacturing acceleration platform behind the... ...Requirements: ~5 or more years of DevOps or SRE experience deploying and operating... ...KPIs to ensure ongoing performance, reliability and efficiency. ~ Network/application...Suggested$170k - $250k
...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000–$250,000 + Competitive Equity We're representing a rapidly...SuggestedWork at officeVisa sponsorshipFlexible hours- .... Overview We are seeking a highly motivated Systems Reliability Engineer (SRE) to lead the design and implementation of operational excellence... ...company supporting sensitive and cleared workforces. The Site Reliability Engineer (SRE) - SecOps will embrace our...SuggestedFor contractorsWork at officeFlexible hours
$100k - $200k
OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about...SuggestedFull time- ...Title: Site Reliability Engineer (SRE) Location: Location: Sunnyvale, CA (3x/ week onsite) Contract Responsibilities: Engage with our product teams to understand requirements, design and implement resilient and scalable infrastructure...SuggestedContract work
$101k - $161k
...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,... ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s... ...-as-a-Service (CVaaS) global SRE team. SREs at Arista combine strong software...- ...Position: Site Reliability Engineering (SRE) Location: Santa Clara, CA (Onsite) Duration: W2 / C2C Contract Experience: 10+ Years Job Description: • WS application and CI/CD pipelines, Microsoft Server admin and workload support (Data Center and AWS)...Contract workImmediate start
$61k - $101k
...,000 per year Requirements: We require formal training or certification in site reliability engineering, along with 5+ years of hands-on experience. We need advanced knowledge of SRE culture and principles, with proven ability to apply them in an application or...Full time- ...development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... ...workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes... ...DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform...Remote jobFor contractors
$165k - $190k
...term growth and IPO readiness.About the DevOps / SRE TeamThe DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and... ...complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate with...Work from home$70 - $100 per hour
...Job Title: Cloud SRE Engineer - Mandarin Bilingual Position Type: Contract (12 months) Location: Palo Alto, CA Salary Rate: $7... ...team is looking for a skilled Cloud SRE Engineer to own the reliability, stability, and continuous improvement of core cloud services...Hourly payContract workTemporary workWork experience placement$207k - $300k
Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.Own end-to-end availability... ...or Engineering.Experience with Large Language Model.Site Reliability Engineering (SRE) combines software and systems engineering to build and...$100k - $200k
A leading technology firm in Palo Alto is seeking a skilled Site Reliability Engineer (SRE) to ensure the stability and performance of application systems. Responsibilities include managing cloud platforms, implementing CI/CD pipelines, and maintaining Linux systems. The...$140k - $165k
...Site Reliability Engineer Instrumental builds the manufacturing acceleration platform behind the world's most complex electronics. We capture... ...high-growth B2B SaaS environment. Experience implementing SRE practices such as SLIs, SLOs, and error budgets. Experience...- Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally... ...yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer... ...open-source projects, particularly in SRE, observability, or AI/ML domains, and certifications...
$222k - $300.5k
...OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational... .... The Fintech Platform Systems Engineering team builds and operates the AWS-based... ...engineering, product, security, and other SRE/infrastructure leaders across Intuit to...WorldwideShift work$262k - $364k
...within the AViD ecosystem have reliability and uptime appropriate to... ...and performance.Build creative engineering solutions to operations and infrastructure... ...and contribute to the cross-SRE AI Ops program, driving the... ...in a strategic way.Site Reliability Engineering (SRE)...- ...SRE Engineer St Louis, MO (Onsite from day 1) Client Required Skills: • Bachelor's Degree in Computer Science, Computer Systems, Information Technology or related. Equivalent experience is acceptable. • Experience with web applications and distributed systems...
$200k - $260k
...every company. About the Role: Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical... ...and eliminating work through automation. On the SRE team, you'll have the opportunity to manage the complex challenges...Work at officeHome officeFlexible hours$207k - $300k
Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.Own end-to-end availability... ...or Engineering.1 year of people management experience. Site Reliability Engineering (SRE) combines software and systems engineering to build and...$252k - $308k
...Staff Site Reliability Engineer Mountain View, US About EarnIn As one of the first pioneers of earned wage access, our passion at EarnIn... ...heroics, tribal knowledge, manual investigation, or isolated SRE expertise. We must embed reliability practices that scale across...Full timeWork at office2 days per week- ...Enterprise Technologies Inc. is a recognized provider of professional IT Consulting services in the US. We are actively seeking SRE Devops Engineer Fulltime Role for one of our direct client. Role: SRE Devops Engineer Location :- Santa Clara,CA (Remote...Full timeLocal areaRemote work
- ...everyone. Role Summary We are seeking an experienced Site Reliability Engineer to help design, build, and operate the infrastructure that... ...of relevant experience in Platform Engineering, DevOps, or SRE roles. ~ Proficiency with Infrastructure as code, preferably...Full timeContract work
- A leading tech company is seeking a Site Reliability Engineer for their U.S. Data Security division. You will improve service lifecycle and ensure data services are reliable and scalable. The ideal candidate has a Bachelor's degree in Computer Science and experience with...
- ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2 Positions) We are... ...platforms. This role is focused on infrastructure, reliability, and automation , with Java exposure as a supporting skill....
$170k - $200k
We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining,... ...background in infrastructure automation, system reliability, and a SRE mindset of continuous improvement.Key Responsibilities:...Full timeWorldwide- ...simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure.... ...Who You AreExperienced Architect: 5+ years of experience in SRE, DevOps, or Systems Engineering, with a proven track record...Full timeWork at office2 days per week
$128k - $216k
...millions of times a day - quickly, reliably, and securely. Any time you... ...at Fiserv.Job TitleSr. Site Reliability EngineerAbout CloverClover... ...Senior Site Reliability Engineer do at Fiserv?As a Senior Site... ...Design Reviews, mentor teams on SRE principles, and bridge the gap...Full timeWorldwide$148k - $235.75k
...on the world.Join our team of innovative engineers who are building an AI Data Center AIOps... ...that turns raw, high-volume telemetry into reliable, job-centric insights and automation for... ...operating production distributed systems as SRE/DevOps/Platform Ops.Proven ownership of...Full time$165k - $280k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most...Permanent employmentTemporary workWorldwideWeekend work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!
- site reliability engineer Palo Alto, CA
- site reliability engineer sre Palo Alto, CA
- junior website developer Palo Alto, CA
- construction site safety Palo Alto, CA
- site services specialist Palo Alto, CA
- website content developer Palo Alto, CA
- on-site clinical research associate (traveling/remote) Palo Alto, CA
- historic site Palo Alto, CA
- official site Palo Alto, CA
- site leader Palo Alto, CA


