Site Reliability Engineer
$150k - $200kRunpod Inc
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.
We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale. Learn more in our CEO's funding announcement: The Reliability team owns the availability, performance, and operational excellence of Runpod's global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions. This team is responsible for:- Defining and enforcing reliability standards across engineering
- Designing incident response processes and improving recovery times
- Building observability systems and reliability tooling
- Driving SLO adoption and production readiness reviews
- Reducing operational toil through automation
- Increase platform uptime and reduce incident frequency and duration
- Establish and operationalize SLIs/SLOs across services
- Improve MTTR through better tooling, automation, and runbooks
- Strengthen production readiness standards
- Drive long-term systemic reliability improvements
- Define and implement SLIs/SLOs for critical services
- Lead incident response and coordinate cross-team mitigation efforts
- Conduct blameless postmortems and ensure corrective actions are completed
- Perform production readiness reviews for new services and features
- Identify systemic risks and drive preventative improvements
- Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.)
- Improve signal-to-noise ratio in alerts and reduce alert fatigue
- Build internal tooling for reliability tracking and reporting
- Improve visibility into GPU performance and distributed systems health
- Automate recurring operational workflows
- Build tools and scripts (Python, Go, Bash) to eliminate manual processes
- Improve deployment safety through automation and guardrails
- Strengthen CI/CD reliability and release processes
- Partner with engineering teams to improve system resilience
- Provide guidance on fault tolerance, scalability, and failure handling
- Contribute to architectural discussions with a reliability-first mindset
- 5+ years of experience in SRE, Reliability Engineering, or Production Engineering
- Strong Linux systems and Networking expertise
- Experience managing containerized production systems
- Strong understanding of distributed systems and failure modes
- Experience defining and managing SLIs/SLOs
- Proven incident response and postmortem leadership experience
- Strong scripting or programming skills
- Experience with monitoring and alerting systems
- Excellent written communication skills
- Successful completion of a background check
- Experience with GPU infrastructure or AI/ML platforms
- Experience improving reliability in high-growth or large scale environments
- Familiarity with GPU observability tooling
- Experience with Infrastructure as Code
- Experience working in startup environments
- Experience building internal reliability platforms or frameworks
- The competitive base pay for this position ranges from $150,000- $200,000 usd. This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate's experience, qualifications, and location
- Meaningful equity in a fast-growing company- everyone on the team receives stock options - your impact drives our growth, and you share in the upside.
- Generous medical, dental & vision plans
- Flexible PTO- take the time you need to recharge
- Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication
- Join a passionate team on the cutting edge of AI infrastructure - where culture, learning, and ownership are at the heart of how we scale.
$100k - $110k
...for new hire onboarding and occasional in-person team meetings and company events. We are seeking an operational-focused Site Reliability Engineer (SRE) to maximize the availability, performance, and resilience of our production healthcare systems. In this role, you will...SuggestedPermanent employmentRemote workFlexible hours- ...SchoolsFirst FCU seeks an experienced Splunk Administrator/Site Reliability Engineer to deploy, optimize, and monitor our IT enterprise platforms. You will onboard data, craft SPL searches, and build dashboards powering IT applications, security operations, and observability...Suggested
$113.3k - $205.52k
...important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development velocity with the reliability our customers depend on. You'll partner with engineering...SuggestedWork at officeRemote workWorldwideFlexible hours$166k - $220k
...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the... ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine...SuggestedFull timeWork experience placementImmediate start$166k - $250k
...to serve a wide variety of defense, IC and commercial customers in US and international markets.ABOUT THE JOBAs a Senior Site Reliability Engineer on the Undersea Dominance team, you will build and operate the infrastructure that keeps our operational and production systems...SuggestedFull timeWork experience placementImmediate start$166k - $220k
...globally. We work with mission partners and operators to deploy reliable and robust capabilities on operationally-relevant fielding... ...scalable deployment solution must be reached. As a Senior Software Engineer, you will re-imagine the infrastructure pipeline required to convert...Full timeWork experience placementImmediate start$143k - $191k
...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental... ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and...Full timeTemporary workWork experience placementImmediate start- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank, Global Payments team, you will solve complex and broad business...
$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...Full timeWork at office$100k - $140k
...and software services. Our team of passionate engineers are constantly innovating, engineering... ...user experience with simpler, smarter, and more reliable connectivity. We're looking for a passionate and experienced Site Reliability Engineer to join our team and play...Local areaWorldwideWeekend work$131.5k - $175.5k
...functionality of Tenable cloud products and ensuring they’re reliable and highly available in cloud environments Responsible for responding... ...with peers on complex projects Collaboration with cloud engineers in understanding new cloud technologies, assessing impact to security...Work experience placementH1bLocal areaRemote workFlexible hours$114k - $148k
...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based...Full timeTemporary workWork experience placementRemote work- ...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and... ...vision, audio, and emerging multimodal architectures — fast, reliable, and effortless to deploy at massive scale. To deliver on that...
- ...Site Reliability Engineer (SRE) Develop and provide operational support for full-stack software applications. Collaborate with development operations staff to create, monitor, and troubleshoot the system infrastructure. Increase system resilience and serve larger customer...
- ...our industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital-... ...for people around the world. We’re looking for a Senior Site Reliability Engineer Engineer to take ownership of building and evolving...Full timeRemote workWork from home
- ...Job Description One of Insight Global's customers is looking to onboard a Sr. Site Reliability Engineer with strong expertise in modern DevOps practices, cloud infrastructure, observability, and platform security. This role partners directly with product teams to support...Contract work
$165k - $190k
...DevOps / SRE Team The DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-... ...security platform Address complex challenges around scalability, reliability, observability, and cost efficiency Collaborate with...Work from homeFlexible hours$166k - $220k
...Senior Site Reliability Engineer Costa Mesa, California, United States Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business...Full timeWork experience placementImmediate startRemote work$191k - $253k
...TEAM:CorpTech Platform is the internal engineering force multiplier behind Anduril’s corporate... ...THE JOB:This Staff SRE role sets the reliability architecture for the systems that run Anduril... ...:10+ years of experience in site reliability engineering, production engineering...Full timeWork experience placementImmediate start- ...scale, we invite you to bring your talents to Zscaler to help shape the future of cybersecurity. Role We are looking for a Site Reliability Engineer-SkillBridge Intern (San JosA Ca or Bellevue WA) to join our Zero Trust Exchange team. This is a remote role based in San...InternshipWork at officeLocal areaRemote workWorldwide
- ...Senior Site Reliability Engineer (Enterprise Platform) Location: Remote - US - Open to Europe if happy to overlap with EST Compensation: Competitive We are a high-growth software company supporting the development of a premier open-source, EVM-compatible public ledger...Contract workCurrently hiringRemote work
$130k - $180k
...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and... ...an in-house AI R&D team. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to...Temporary workWork at officeImmediate startRemote workFlexible hours$185k - $227k
...professionals. If the opportunity to build your career is compelling, read on for more details. ROLE AND RESPONSIBILITIES: A Senior Site Reliability Engineer (SRE) is expected to own the operational stability and performance ofJuul’s hybrid cloud infrastructure (Nutanix, AWS/GCP)...Remote work$102.8k - $190.2k
Team Name:Battle.net & Online ProductsJob Title:Senior Site Reliability Engineer, Data & AnalyticsRequisition ID:R027436Job Description:This Senior Site Reliability Engineer role is on our Data & Analytics team, partnering with data, analytics, ML, and platform engineering...Full timeTemporary workPart timeLocal areaRemote workRelocation package- Manager - Production Operations & Site Reliability EngineeringAt Alcon, we are driven by the meaningful work we do to help people see brilliantly... ...for diverse, talented people to join Alcon. As a Principal Engineer you will provide technical leadership for the reliability,...Full timeTemporary work
$87 per hour
Join a long-standing IT services provider with decades of experience. The company has built a reputation for delivering award-winning digital transformation projects, securing top-tier partnerships with Microsoft, Cisco, and other leading technology vendors. The organisation...Hourly payTemporary work$83.94k - $120.03k
...10062 – Software Engineer II Location – Fountain Valley (5-days Onsite) PURPOSE The Software Engineer II is responsible for analyzing, designing, developing, and implementing software projects and tasks. Directly interfacing with user clients about their systems...Full time$172k - $220k
...Location : Santa Ana, California — on-site Employment type : Full-time Salary... ...company. We are expanding our engineering team with two appointments covering our... ...integrity, tenant security and overall platform reliability underpinning created.ai. The role is...Full timeWork at officeLocal area- ...Masters, and Product Owners. They seek to remove barriers and empower teams to effectively self-organize. Skills of a Release Train Engineer: Leadership Skills: A strong leader with the capacity to motivate and influence teams is required for an effective RTE. They...Immediate start
$129.3k - $172.3k
...visionary technology leaders at First American—where history meets innovation, and tradition fuels transformation. As a Senior Software Engineer - Salesforce, you'll play a key role in designing and delivering the enterprise Salesforce capabilities that enable customer...Full timeLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- on-site clinical research associate (traveling/remote) Santa Ana, CA
- junior website developer Santa Ana, CA
- site leader Santa Ana, CA
- construction site safety Santa Ana, CA
- official site Santa Ana, CA
- site services specialist Santa Ana, CA
- site safety Santa Ana, CA
- IT site lead Santa Ana, CA
- site reliability engineer remote
- site reliability engineer sre



