Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$150k - $200k

Runpod Inc

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.

We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.

Learn more in our CEO's funding announcement:

The Reliability team owns the availability, performance, and operational excellence of Runpod's global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions.

This team is responsible for:
  • Defining and enforcing reliability standards across engineering
  • Designing incident response processes and improving recovery times
  • Building observability systems and reliability tooling
  • Driving SLO adoption and production readiness reviews
  • Reducing operational toil through automation
The Reliability team works cross-functionally with Infrastructure, Product Engineering, and Support to ensure our systems remain stable and performant as we scale rapidly. We value proactive problem solving, automation-first thinking, and strong ownership of production systems.

As a Site Reliability Engineer on the Reliability team, you will focus on ensuring the stability and resilience of Runpod's distributed platform. You will partner with engineering teams to improve system design, strengthen observability, and prevent incidents before they happen.

This role blends software engineering with production operations. You'll work on reliability frameworks, SLO design, automation, and production hardening, reducing errors and improving performance across different services and infrastructure.

This is a high-impact role central to maintaining trust with developers running critical AI workloads on Runpod.

Your Impact
  • Increase platform uptime and reduce incident frequency and duration
  • Establish and operationalize SLIs/SLOs across services
  • Improve MTTR through better tooling, automation, and runbooks
  • Strengthen production readiness standards
  • Drive long-term systemic reliability improvements
You will influence how reliability is defined and measured across Runpod and help build the operational backbone of the company.

Responsibilities:

Reliability Engineering
  • Define and implement SLIs/SLOs for critical services
  • Lead incident response and coordinate cross-team mitigation efforts
  • Conduct blameless postmortems and ensure corrective actions are completed
  • Perform production readiness reviews for new services and features
  • Identify systemic risks and drive preventative improvements
Observability & Monitoring
  • Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.)
  • Improve signal-to-noise ratio in alerts and reduce alert fatigue
  • Build internal tooling for reliability tracking and reporting
  • Improve visibility into GPU performance and distributed systems health
Automation & Toil Reduction
  • Automate recurring operational workflows
  • Build tools and scripts (Python, Go, Bash) to eliminate manual processes
  • Improve deployment safety through automation and guardrails
  • Strengthen CI/CD reliability and release processes
Cross-Functional Reliability Advocacy
  • Partner with engineering teams to improve system resilience
  • Provide guidance on fault tolerance, scalability, and failure handling
  • Contribute to architectural discussions with a reliability-first mindset
Requirements:
  • 5+ years of experience in SRE, Reliability Engineering, or Production Engineering
  • Strong Linux systems and Networking expertise
  • Experience managing containerized production systems
  • Strong understanding of distributed systems and failure modes
  • Experience defining and managing SLIs/SLOs
  • Proven incident response and postmortem leadership experience
  • Strong scripting or programming skills
  • Experience with monitoring and alerting systems
  • Excellent written communication skills
  • Successful completion of a background check
Preferred:
  • Experience with GPU infrastructure or AI/ML platforms
  • Experience improving reliability in high-growth or large scale environments
  • Familiarity with GPU observability tooling
  • Experience with Infrastructure as Code
  • Experience working in startup environments
  • Experience building internal reliability platforms or frameworks
What You'll Receive:
  • The competitive base pay for this position ranges from $150,000- $200,000 usd. This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate's experience, qualifications, and location
  • Meaningful equity in a fast-growing company- everyone on the team receives stock options - your impact drives our growth, and you share in the upside.
  • Generous medical, dental & vision plans
  • Flexible PTO- take the time you need to recharge
  • Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication
  • Join a passionate team on the cutting edge of AI infrastructure - where culture, learning, and ownership are at the heart of how we scale.

Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Irvine, CA vacancy
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Costa... 
    Suggested
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    3 days ago
  • $166k - $220k

     ...globally. We work with mission partners and operators to deploy reliable and robust capabilities on operationally-relevant fielding...  ...scalable deployment solution must be reached. As a Senior Software Engineer, you will re-imagine the infrastructure pipeline required to convert... 
    Suggested
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    1 day ago
  • $143k - $191k

     ...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental...  ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    1 day ago
  • $98.58k - $138.02k

     ...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company...  ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,... 
    Suggested
    Full time
    Work at office

    Restaurant 365

    Irvine, CA
    21 hours ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank, Global Payments team, you will solve complex and broad business... 
    Suggested

    JP Morgan Chase

    Irvine, CA
    3 days ago
  • $166k - $220k

     ...technology to the military in months, not years. ABOUT THE TEAM We are seeking a highly skilled and mission-driven Site Reliability Engineer (SRE) to join our Mission Autonomy team. In this critical role, you will be responsible for ensuring the reliability,... 
    Full time
    Work experience placement
    Immediate start
    Remote work

    Anduril Industries

    Costa Mesa, CA
    4 days ago
  •  ...The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As... 
    Work at office
    Local area
    Remote work
    Worldwide

    GrabJobs

    Irvine, CA
    21 hours ago
  •  ...SchoolsFirst FCU seeks an experienced Splunk Administrator/Site Reliability Engineer to deploy, optimize, and monitor our IT enterprise platforms. You will onboard data, craft SPL searches, and build dashboards powering IT applications, security operations, and observability... 

    SchoolsFirst Federal Credit Union

    Tustin, CA
    1 day ago
  • $90k - $105k

    Site Reliability Engineer At ICEYE, we design, build and operate the largest fleet of Synthetic Aperture Radar (SAR) satellites in the world. Using advanced technology, our constellation collects topographical data about any location on Earth, day or night, through any... 
    Full time
    Work at office
    Local area
    Flexible hours
    Night shift

    ICEYE US

    Irvine, CA
    2 days ago
  • $166k - $250k

     ...to serve a wide variety of defense, IC and commercial customers in US and international markets. ABOUT THE JOB As a Senior Site Reliability Engineer on the Undersea Dominance team, you will build and operate the infrastructure that keeps our operational and production... 
    Full time
    Work experience placement

    Anduril

    Costa Mesa, CA
    15 hours ago
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...end solutions we ship. ABOUT THE JOB We are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    2 days ago
  •  ...our industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital-...  ...for people around the world. We’re looking for a Senior Site Reliability Engineer Engineer to take ownership of building and evolving... 
    Full time
    Remote work
    Work from home

    GrabJobs

    Irvine, CA
    2 days ago
  •  ...Site Reliability Engineer (SRE) Develop and provide operational support for full-stack software applications. Collaborate with development operations staff to create, monitor, and troubleshoot the system infrastructure. Increase system resilience and serve larger customer... 

    InterSources

    Irvine, CA
    21 hours ago
  • $100k - $140k

     ...and software services. Our team of passionate engineers are constantly innovating, engineering...  ...user experience with simpler, smarter, and more reliable connectivity. We're looking for a passionate and experienced  Site Reliability Engineer  to join our team and play... 
    Local area
    Worldwide
    Weekend work

    TP-Link Systems Inc.

    Irvine, CA
    2 days ago
  •  ...Job Description One of Insight Global's customers is looking to onboard a Sr. Site Reliability Engineer with strong expertise in modern DevOps practices, cloud infrastructure, observability, and platform security. This role partners directly with product teams to support... 
    Contract work

    Insight Global

    Irvine, CA
    4 days ago
  • $191k - $253k

     ...TEAM:CorpTech Platform is the internal engineering force multiplier behind Anduril’s corporate...  ...THE JOB:This Staff SRE role sets the reliability architecture for the systems that run Anduril...  ...:10+ years of experience in site reliability engineering, production engineering... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    2 days ago
  • $160k - $240k

     ...passionate about building unified IT solutions that simplify the way IT organizations work. We are currently looking for a Senior Site Reliability Engineer to join our SRE team in the Platform Engineering organization and help us scale our products to millions of end-users. We... 
    Permanent employment
    Full time
    Remote work
    Work from home
    Relocation
    Flexible hours

    GrabJobs

    Irvine, CA
    4 days ago
  • $102.8k - $190.2k

    Team Name:Battle.net & Online ProductsJob Title:Senior Site Reliability Engineer, Data & AnalyticsRequisition ID:R027436Job Description:This Senior Site Reliability Engineer role is on our Data & Analytics team, partnering with data, analytics, ML, and platform engineering... 
    Full time
    Temporary work
    Part time
    Local area
    Remote work
    Relocation package

    Blizzard Entertainment

    Irvine, CA
    2 days ago
  • $145.7k - $218.5k

     ...synonymous with entertainment excellence and creativity.Service Reliability EngineerDo you want to use transformative technologies to...  ...scalability and efficiency? Do you want a career that combines your engineering skills and your passion for video gaming? Are you fascinated... 
    Work experience placement
    Shift work

    Sony Interactive Entertainment America

    Aliso Viejo, CA
    2 days ago
  • $210k - $220k

     ...secure and private by design, it’s popular with security, IT, engineering, finance, and other security-focused teams. At Tines, we're...  ...we’re looking for others to join us on our journey. Senior Site Reliability Engineer - Government Cloud You'll join the team responsible... 
    Work at office
    Remote work

    GrabJobs

    Santa Ana, CA
    21 hours ago
  • $114k - $148k

     ...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based... 
    Full time
    Temporary work
    Work experience placement
    Remote work

    GrabJobs

    Santa Ana, CA
    3 days ago
  • $185k - $227k

     ...professionals. If the opportunity to build your career is compelling, read on for more details. ROLE AND RESPONSIBILITIES: A Senior Site Reliability Engineer (SRE) is expected to own the operational stability and performance ofJuul’s hybrid cloud infrastructure (Nutanix, AWS/GCP)... 
    Remote work

    GrabJobs

    Santa Ana, CA
    1 day ago
  •  ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.   As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank, Global Payments team, you will solve complex and broad... 

    JPMorgan Chase & Co.

    Irvine, CA
    10 days ago
  • Manager - Production Operations & Site Reliability EngineeringAt Alcon, we are driven by the meaningful work we do to help people see brilliantly...  ...for diverse, talented people to join Alcon. As a Principal Engineer you will provide technical leadership for the reliability,... 
    Full time
    Temporary work

    Alcon Pharma

    Lake Forest, CA
    1 day ago
  • $166k - $220k

     ...autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.ABOUT THE TEAMThe Reliability Engineering team partners across Anduril's engineering, manufacturing, and operations organizations to ensure our autonomous systems... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    2 days ago
  • $166k - $220k

     ..., AI, computer vision, sensor fusion, and networking technology to the military in months, not years.WHY WE'RE HERE: The systems engineering team is looking for experienced systems engineers to support a Group 5 aircraft program at Anduril. The ideal candidate will have... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    21 hours ago
  • $102k - $177.1k

     ...Locations:Irvine, California, United States of AmericaJob Description:RELIABILITY MANAGER:The Reliability Manager leads initiatives to ensure...  ...failures, drive corrective actions, and lead a team of engineers to build a "zero defect" culture. The role focuses on improving... 
    Full time
    Immediate start

    Johnson & Johnson

    Irvine, CA
    3 days ago
  • $191k - $253k

     ...priorities, we want you to join Anduril’s Maritime Division and help us build the future of defense capability.About the JobSr. Software Engineers independently drive the delivery of a variety of software integrated in to our products. This includes autonomy, simulation, data... 
    Full time
    Work experience placement
    Immediate start
    Remote work
    Flexible hours

    Anduril Industries

    Costa Mesa, CA
    2 days ago
  • $166k - $220k

     ...in this contested warfighting domain.ABOUT THE JOBAs a Software Engineer, Systems Test for our Space team, you will be instrumental in...  ...threads. We work with mission partners and operators to deploy reliable and robust capabilities on operationally relevant fielding timelines... 
    Full time
    Work experience placement
    Immediate start
    Remote work

    Anduril Industries

    Costa Mesa, CA
    4 days ago
  • $166k - $220k

     ...ABOUT THE JOB:We are seeking a highly skilled and driven Software Engineer to join a fast-paced team dedicated to the continued...  ...-quality code that enhances the performance, scalability, and reliability of the platformEngage in code reviews, architectural discussions... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!