Site Reliability Engineer
$150k - $200kRunpod Inc
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.
We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale. Learn more in our CEO's funding announcement: The Reliability team owns the availability, performance, and operational excellence of Runpod's global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions. This team is responsible for:- Defining and enforcing reliability standards across engineering
- Designing incident response processes and improving recovery times
- Building observability systems and reliability tooling
- Driving SLO adoption and production readiness reviews
- Reducing operational toil through automation
- Increase platform uptime and reduce incident frequency and duration
- Establish and operationalize SLIs/SLOs across services
- Improve MTTR through better tooling, automation, and runbooks
- Strengthen production readiness standards
- Drive long-term systemic reliability improvements
- Define and implement SLIs/SLOs for critical services
- Lead incident response and coordinate cross-team mitigation efforts
- Conduct blameless postmortems and ensure corrective actions are completed
- Perform production readiness reviews for new services and features
- Identify systemic risks and drive preventative improvements
- Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.)
- Improve signal-to-noise ratio in alerts and reduce alert fatigue
- Build internal tooling for reliability tracking and reporting
- Improve visibility into GPU performance and distributed systems health
- Automate recurring operational workflows
- Build tools and scripts (Python, Go, Bash) to eliminate manual processes
- Improve deployment safety through automation and guardrails
- Strengthen CI/CD reliability and release processes
- Partner with engineering teams to improve system resilience
- Provide guidance on fault tolerance, scalability, and failure handling
- Contribute to architectural discussions with a reliability-first mindset
- 5+ years of experience in SRE, Reliability Engineering, or Production Engineering
- Strong Linux systems and Networking expertise
- Experience managing containerized production systems
- Strong understanding of distributed systems and failure modes
- Experience defining and managing SLIs/SLOs
- Proven incident response and postmortem leadership experience
- Strong scripting or programming skills
- Experience with monitoring and alerting systems
- Excellent written communication skills
- Successful completion of a background check
- Experience with GPU infrastructure or AI/ML platforms
- Experience improving reliability in high-growth or large scale environments
- Familiarity with GPU observability tooling
- Experience with Infrastructure as Code
- Experience working in startup environments
- Experience building internal reliability platforms or frameworks
- The competitive base pay for this position ranges from $150,000- $200,000 usd. This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate's experience, qualifications, and location
- Meaningful equity in a fast-growing company- everyone on the team receives stock options - your impact drives our growth, and you share in the upside.
- Generous medical, dental & vision plans
- Flexible PTO- take the time you need to recharge
- Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication
- Join a passionate team on the cutting edge of AI infrastructure - where culture, learning, and ownership are at the heart of how we scale.
$166k - $220k
...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the... ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Costa...SuggestedFull timeWork experience placementImmediate start$166k - $220k
...globally. We work with mission partners and operators to deploy reliable and robust capabilities on operationally-relevant fielding... ...scalable deployment solution must be reached. As a Senior Software Engineer, you will re-imagine the infrastructure pipeline required to convert...SuggestedFull timeWork experience placementImmediate start$143k - $191k
...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental... ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and...SuggestedFull timeTemporary workWork experience placementImmediate start$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...SuggestedFull timeWork at office- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank, Global Payments team, you will solve complex and broad business...Suggested
$166k - $220k
...technology to the military in months, not years. ABOUT THE TEAM We are seeking a highly skilled and mission-driven Site Reliability Engineer (SRE) to join our Mission Autonomy team. In this critical role, you will be responsible for ensuring the reliability,...Full timeWork experience placementImmediate startRemote work- ...The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As...Work at officeLocal areaRemote workWorldwide
- ...SchoolsFirst FCU seeks an experienced Splunk Administrator/Site Reliability Engineer to deploy, optimize, and monitor our IT enterprise platforms. You will onboard data, craft SPL searches, and build dashboards powering IT applications, security operations, and observability...
$90k - $105k
Site Reliability Engineer At ICEYE, we design, build and operate the largest fleet of Synthetic Aperture Radar (SAR) satellites in the world. Using advanced technology, our constellation collects topographical data about any location on Earth, day or night, through any...Full timeWork at officeLocal areaFlexible hoursNight shift$166k - $250k
...to serve a wide variety of defense, IC and commercial customers in US and international markets. ABOUT THE JOB As a Senior Site Reliability Engineer on the Undersea Dominance team, you will build and operate the infrastructure that keeps our operational and production...Full timeWork experience placement$166k - $220k
...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the... ...end solutions we ship. ABOUT THE JOB We are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in...Full timeWork experience placementImmediate start- ...our industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital-... ...for people around the world. We’re looking for a Senior Site Reliability Engineer Engineer to take ownership of building and evolving...Full timeRemote workWork from home
- ...Site Reliability Engineer (SRE) Develop and provide operational support for full-stack software applications. Collaborate with development operations staff to create, monitor, and troubleshoot the system infrastructure. Increase system resilience and serve larger customer...
$100k - $140k
...and software services. Our team of passionate engineers are constantly innovating, engineering... ...user experience with simpler, smarter, and more reliable connectivity. We're looking for a passionate and experienced Site Reliability Engineer to join our team and play...Local areaWorldwideWeekend work- ...Job Description One of Insight Global's customers is looking to onboard a Sr. Site Reliability Engineer with strong expertise in modern DevOps practices, cloud infrastructure, observability, and platform security. This role partners directly with product teams to support...Contract work
$191k - $253k
...TEAM:CorpTech Platform is the internal engineering force multiplier behind Anduril’s corporate... ...THE JOB:This Staff SRE role sets the reliability architecture for the systems that run Anduril... ...:10+ years of experience in site reliability engineering, production engineering...Full timeWork experience placementImmediate start$160k - $240k
...passionate about building unified IT solutions that simplify the way IT organizations work. We are currently looking for a Senior Site Reliability Engineer to join our SRE team in the Platform Engineering organization and help us scale our products to millions of end-users. We...Permanent employmentFull timeRemote workWork from homeRelocationFlexible hours$102.8k - $190.2k
Team Name:Battle.net & Online ProductsJob Title:Senior Site Reliability Engineer, Data & AnalyticsRequisition ID:R027436Job Description:This Senior Site Reliability Engineer role is on our Data & Analytics team, partnering with data, analytics, ML, and platform engineering...Full timeTemporary workPart timeLocal areaRemote workRelocation package$145.7k - $218.5k
...synonymous with entertainment excellence and creativity.Service Reliability EngineerDo you want to use transformative technologies to... ...scalability and efficiency? Do you want a career that combines your engineering skills and your passion for video gaming? Are you fascinated...Work experience placementShift work$210k - $220k
...secure and private by design, it’s popular with security, IT, engineering, finance, and other security-focused teams. At Tines, we're... ...we’re looking for others to join us on our journey. Senior Site Reliability Engineer - Government Cloud You'll join the team responsible...Work at officeRemote work$114k - $148k
...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based...Full timeTemporary workWork experience placementRemote work$185k - $227k
...professionals. If the opportunity to build your career is compelling, read on for more details. ROLE AND RESPONSIBILITIES: A Senior Site Reliability Engineer (SRE) is expected to own the operational stability and performance ofJuul’s hybrid cloud infrastructure (Nutanix, AWS/GCP)...Remote work- ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank, Global Payments team, you will solve complex and broad...
- Manager - Production Operations & Site Reliability EngineeringAt Alcon, we are driven by the meaningful work we do to help people see brilliantly... ...for diverse, talented people to join Alcon. As a Principal Engineer you will provide technical leadership for the reliability,...Full timeTemporary work
$166k - $220k
...autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.ABOUT THE TEAMThe Reliability Engineering team partners across Anduril's engineering, manufacturing, and operations organizations to ensure our autonomous systems...Full timeWork experience placementImmediate start$166k - $220k
..., AI, computer vision, sensor fusion, and networking technology to the military in months, not years.WHY WE'RE HERE: The systems engineering team is looking for experienced systems engineers to support a Group 5 aircraft program at Anduril. The ideal candidate will have...Full timeWork experience placementImmediate start$102k - $177.1k
...Locations:Irvine, California, United States of AmericaJob Description:RELIABILITY MANAGER:The Reliability Manager leads initiatives to ensure... ...failures, drive corrective actions, and lead a team of engineers to build a "zero defect" culture. The role focuses on improving...Full timeImmediate start$191k - $253k
...priorities, we want you to join Anduril’s Maritime Division and help us build the future of defense capability.About the JobSr. Software Engineers independently drive the delivery of a variety of software integrated in to our products. This includes autonomy, simulation, data...Full timeWork experience placementImmediate startRemote workFlexible hours$166k - $220k
...in this contested warfighting domain.ABOUT THE JOBAs a Software Engineer, Systems Test for our Space team, you will be instrumental in... ...threads. We work with mission partners and operators to deploy reliable and robust capabilities on operationally relevant fielding timelines...Full timeWork experience placementImmediate startRemote work$166k - $220k
...ABOUT THE JOB:We are seeking a highly skilled and driven Software Engineer to join a fast-paced team dedicated to the continued... ...-quality code that enhances the performance, scalability, and reliability of the platformEngage in code reviews, architectural discussions...Full timeWork experience placementImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre Irvine, CA
- site reliability engineer Irvine, CA
- on-site clinical research associate (traveling/remote) Irvine, CA
- website coordinator Irvine, CA
- junior website developer Irvine, CA
- site leader Irvine, CA
- historic site Irvine, CA
- website content developer Irvine, CA
- construction site safety Irvine, CA
- official site Irvine, CA


