Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$150k - $200k

Runpod Inc

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.

We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.

Learn more in our CEO's funding announcement:

The Reliability team owns the availability, performance, and operational excellence of Runpod's global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions.

This team is responsible for:
  • Defining and enforcing reliability standards across engineering
  • Designing incident response processes and improving recovery times
  • Building observability systems and reliability tooling
  • Driving SLO adoption and production readiness reviews
  • Reducing operational toil through automation
The Reliability team works cross-functionally with Infrastructure, Product Engineering, and Support to ensure our systems remain stable and performant as we scale rapidly. We value proactive problem solving, automation-first thinking, and strong ownership of production systems.

As a Site Reliability Engineer on the Reliability team, you will focus on ensuring the stability and resilience of Runpod's distributed platform. You will partner with engineering teams to improve system design, strengthen observability, and prevent incidents before they happen.

This role blends software engineering with production operations. You'll work on reliability frameworks, SLO design, automation, and production hardening, reducing errors and improving performance across different services and infrastructure.

This is a high-impact role central to maintaining trust with developers running critical AI workloads on Runpod.

Your Impact
  • Increase platform uptime and reduce incident frequency and duration
  • Establish and operationalize SLIs/SLOs across services
  • Improve MTTR through better tooling, automation, and runbooks
  • Strengthen production readiness standards
  • Drive long-term systemic reliability improvements
You will influence how reliability is defined and measured across Runpod and help build the operational backbone of the company.

Responsibilities:

Reliability Engineering
  • Define and implement SLIs/SLOs for critical services
  • Lead incident response and coordinate cross-team mitigation efforts
  • Conduct blameless postmortems and ensure corrective actions are completed
  • Perform production readiness reviews for new services and features
  • Identify systemic risks and drive preventative improvements
Observability & Monitoring
  • Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.)
  • Improve signal-to-noise ratio in alerts and reduce alert fatigue
  • Build internal tooling for reliability tracking and reporting
  • Improve visibility into GPU performance and distributed systems health
Automation & Toil Reduction
  • Automate recurring operational workflows
  • Build tools and scripts (Python, Go, Bash) to eliminate manual processes
  • Improve deployment safety through automation and guardrails
  • Strengthen CI/CD reliability and release processes
Cross-Functional Reliability Advocacy
  • Partner with engineering teams to improve system resilience
  • Provide guidance on fault tolerance, scalability, and failure handling
  • Contribute to architectural discussions with a reliability-first mindset
Requirements:
  • 5+ years of experience in SRE, Reliability Engineering, or Production Engineering
  • Strong Linux systems and Networking expertise
  • Experience managing containerized production systems
  • Strong understanding of distributed systems and failure modes
  • Experience defining and managing SLIs/SLOs
  • Proven incident response and postmortem leadership experience
  • Strong scripting or programming skills
  • Experience with monitoring and alerting systems
  • Excellent written communication skills
  • Successful completion of a background check
Preferred:
  • Experience with GPU infrastructure or AI/ML platforms
  • Experience improving reliability in high-growth or large scale environments
  • Familiarity with GPU observability tooling
  • Experience with Infrastructure as Code
  • Experience working in startup environments
  • Experience building internal reliability platforms or frameworks
What You'll Receive:
  • The competitive base pay for this position ranges from $150,000- $200,000 usd. This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate's experience, qualifications, and location
  • Meaningful equity in a fast-growing company- everyone on the team receives stock options - your impact drives our growth, and you share in the upside.
  • Generous medical, dental & vision plans
  • Flexible PTO- take the time you need to recharge
  • Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication
  • Join a passionate team on the cutting edge of AI infrastructure - where culture, learning, and ownership are at the heart of how we scale.

Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Santa Ana, CA vacancy
  • $100k - $110k

     ...for new hire onboarding and occasional in-person team meetings and company events. We are seeking an operational-focused Site Reliability Engineer (SRE) to maximize the availability, performance, and resilience of our production healthcare systems. In this role, you will... 
    Suggested
    Permanent employment
    Remote work
    Flexible hours

    GrabJobs

    Santa Ana, CA
    4 days ago
  •  ...SchoolsFirst FCU seeks an experienced Splunk Administrator/Site Reliability Engineer to deploy, optimize, and monitor our IT enterprise platforms. You will onboard data, craft SPL searches, and build dashboards powering IT applications, security operations, and observability... 
    Suggested

    SchoolsFirst Federal Credit Union

    Tustin, CA
    3 days ago
  • $113.3k - $205.52k

     ...important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development velocity with the reliability our customers depend on. You'll partner with engineering... 
    Suggested
    Work at office
    Remote work
    Worldwide
    Flexible hours

    GrabJobs

    Santa Ana, CA
    3 days ago
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine... 
    Suggested
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    2 days ago
  • $166k - $250k

     ...to serve a wide variety of defense, IC and commercial customers in US and international markets.ABOUT THE JOBAs a Senior Site Reliability Engineer on the Undersea Dominance team, you will build and operate the infrastructure that keeps our operational and production systems... 
    Suggested
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    1 day ago
  • $166k - $220k

     ...globally. We work with mission partners and operators to deploy reliable and robust capabilities on operationally-relevant fielding...  ...scalable deployment solution must be reached. As a Senior Software Engineer, you will re-imagine the infrastructure pipeline required to convert... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    3 days ago
  • $143k - $191k

     ...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental...  ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and... 
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    3 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank, Global Payments team, you will solve complex and broad business... 

    JP Morgan Chase

    Irvine, CA
    16 hours ago
  • $98.58k - $138.02k

     ...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company...  ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,... 
    Full time
    Work at office

    Restaurant 365

    Irvine, CA
    2 days ago
  • $100k - $140k

     ...and software services. Our team of passionate engineers are constantly innovating, engineering...  ...user experience with simpler, smarter, and more reliable connectivity. We're looking for a passionate and experienced  Site Reliability Engineer  to join our team and play... 
    Local area
    Worldwide
    Weekend work

    TP-Link Systems Inc.

    Irvine, CA
    4 days ago
  • $131.5k - $175.5k

     ...functionality of Tenable cloud products and ensuring they’re reliable and highly available in cloud environments Responsible for responding...  ...with peers on complex projects Collaboration with cloud engineers in understanding new cloud technologies, assessing impact to security... 
    Work experience placement
    H1b
    Local area
    Remote work
    Flexible hours

    GrabJobs

    Anaheim, CA
    2 days ago
  • $114k - $148k

     ...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based... 
    Full time
    Temporary work
    Work experience placement
    Remote work

    GrabJobs

    Irvine, CA
    3 days ago
  •  ...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and...  ...vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to deploy at massive scale. To deliver on that... 

    GrabJobs

    Anaheim, CA
    3 days ago
  •  ...Site Reliability Engineer (SRE) Develop and provide operational support for full-stack software applications. Collaborate with development operations staff to create, monitor, and troubleshoot the system infrastructure. Increase system resilience and serve larger customer... 

    InterSources

    Irvine, CA
    2 days ago
  •  ...our industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital-...  ...for people around the world. We’re looking for a Senior Site Reliability Engineer Engineer to take ownership of building and evolving... 
    Full time
    Remote work
    Work from home

    GrabJobs

    Irvine, CA
    1 day ago
  •  ...Job Description One of Insight Global's customers is looking to onboard a Sr. Site Reliability Engineer with strong expertise in modern DevOps practices, cloud infrastructure, observability, and platform security. This role partners directly with product teams to support... 
    Contract work

    Insight Global

    Irvine, CA
    1 day ago
  • $165k - $190k

     ...DevOps / SRE Team The DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-...  ...security platform Address complex challenges around scalability, reliability, observability, and cost efficiency Collaborate with... 
    Work from home
    Flexible hours

    Obsidian Security

    Newport Beach, CA
    a month ago
  • $166k - $220k

     ...Senior Site Reliability Engineer Costa Mesa, California, United States Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business... 
    Full time
    Work experience placement
    Immediate start
    Remote work

    anduril

    Costa Mesa, CA
    2 days ago
  • $191k - $253k

     ...TEAM:CorpTech Platform is the internal engineering force multiplier behind Anduril’s corporate...  ...THE JOB:This Staff SRE role sets the reliability architecture for the systems that run Anduril...  ...:10+ years of experience in site reliability engineering, production engineering... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    4 days ago
  •  ...scale, we invite you to bring your talents to Zscaler to help shape the future of cybersecurity. Role We are looking for a Site Reliability Engineer-SkillBridge Intern (San JosA Ca or Bellevue WA) to join our Zero Trust Exchange team. This is a remote role based in San... 
    Internship
    Work at office
    Local area
    Remote work
    Worldwide

    GrabJobs

    Anaheim, CA
    1 day ago
  •  ...Senior Site Reliability Engineer (Enterprise Platform) Location: Remote - US - Open to Europe if happy to overlap with EST Compensation: Competitive We are a high-growth software company supporting the development of a premier open-source, EVM-compatible public ledger... 
    Contract work
    Currently hiring
    Remote work

    GrabJobs

    Anaheim, CA
    5 days ago
  • $130k - $180k

     ...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and...  ...an in-house AI R&D team. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to... 
    Temporary work
    Work at office
    Immediate start
    Remote work
    Flexible hours

    GrabJobs

    Anaheim, CA
    4 days ago
  • $185k - $227k

     ...professionals. If the opportunity to build your career is compelling, read on for more details. ROLE AND RESPONSIBILITIES: A Senior Site Reliability Engineer (SRE) is expected to own the operational stability and performance ofJuul’s hybrid cloud infrastructure (Nutanix, AWS/GCP)... 
    Remote work

    GrabJobs

    Irvine, CA
    5 days ago
  • $102.8k - $190.2k

    Team Name:Battle.net & Online ProductsJob Title:Senior Site Reliability Engineer, Data & AnalyticsRequisition ID:R027436Job Description:This Senior Site Reliability Engineer role is on our Data & Analytics team, partnering with data, analytics, ML, and platform engineering... 
    Full time
    Temporary work
    Part time
    Local area
    Remote work
    Relocation package

    Blizzard Entertainment

    Irvine, CA
    4 days ago
  • Manager - Production Operations & Site Reliability EngineeringAt Alcon, we are driven by the meaningful work we do to help people see brilliantly...  ...for diverse, talented people to join Alcon. As a Principal Engineer you will provide technical leadership for the reliability,... 
    Full time
    Temporary work

    Alcon Pharma

    Lake Forest, CA
    3 days ago
  • $87 per hour

    Join a long-standing IT services provider with decades of experience. The company has built a reputation for delivering award-winning digital transformation projects, securing top-tier partnerships with Microsoft, Cisco, and other leading technology vendors. The organisation...
    Hourly pay
    Temporary work
    Fountain Valley, CA
    more than 2 months ago
  • $83.94k - $120.03k

     ...10062 – Software Engineer II Location – Fountain Valley (5-days Onsite) PURPOSE The Software Engineer II is responsible for analyzing, designing, developing, and implementing software projects and tasks. Directly interfacing with user clients about their systems... 
    Full time

    Hyundai Autoever America

    Fountain Valley, CA
    16 hours ago
  • $172k - $220k

     ...Location : Santa Ana, California — on-site Employment type : Full-time Salary...  ...company. We are expanding our engineering team with two appointments covering our...  ...integrity, tenant security and overall platform reliability underpinning created.ai. The role is... 
    Full time
    Work at office
    Local area

    eJam

    Santa Ana, CA
    15 days ago
  •  ...Masters, and Product Owners. They seek to remove barriers and empower teams to effectively self-organize. Skills of a Release Train Engineer: Leadership Skills: A strong leader with the capacity to motivate and influence teams is required for an effective RTE. They... 
    Immediate start

    TechDigital Group

    Fountain Valley, CA
    16 hours ago
  • $129.3k - $172.3k

     ...visionary technology leaders at First American—where history meets innovation, and tradition fuels transformation. As a Senior Software Engineer - Salesforce, you'll play a key role in designing and delivering the enterprise Salesforce capabilities that enable customer... 
    Full time
    Local area
    Remote work

    First American Financial Corporation

    Santa Ana, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!