Remote Site Reliability Engineer
$150k - $200kGrabJobs
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on. We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale. Learn more in our CEO's funding announcement: . The Reliability team owns the availability, performance, and operational excellence of Runpod’s global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions. This team is responsible for: Defining and enforcing reliability standards across engineering Designing incident response processes and improving recovery times Building observability systems and reliability tooling Driving SLO adoption and production readiness reviews Reducing operational toil through automation The Reliability team works cross-functionally with Infrastructure, Product Engineering, and Support to ensure our systems remain stable and performant as we scale rapidly. We value proactive problem solving, automation-first thinking, and strong ownership of production systems. As a Site Reliability Engineer on the Reliability team, you will focus on ensuring the stability and resilience of Runpod’s distributed platform. You will partner with engineering teams to improve system design, strengthen observability, and prevent incidents before they happen. This role blends software engineering with production operations. You’ll work on reliability frameworks, SLO design, automation, and production hardening, reducing errors and improving performance across different services and infrastructure. This is a high-impact role central to maintaining trust with developers running critical AI workloads on Runpod. Your Impact Increase platform uptime and reduce incident frequency and duration Establish and operationalize SLIs/SLOs across services Improve MTTR through better tooling, automation, and runbooks Strengthen production readiness standards Drive long-term systemic reliability improvements You will influence how reliability is defined and measured across Runpod and help build the operational backbone of the company. Responsibilities: Reliability Engineering Define and implement SLIs/SLOs for critical services Lead incident response and coordinate cross-team mitigation efforts Conduct blameless postmortems and ensure corrective actions are completed Perform production readiness reviews for new services and features Identify systemic risks and drive preventative improvements Observability & Monitoring Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.) Improve signal-to-noise ratio in alerts and reduce alert fatigue Build internal tooling for reliability tracking and reporting Improve visibility into GPU performance and distributed systems health Automation & Toil Reduction Automate recurring operational workflows Build tools and scripts (Python, Go, Bash) to eliminate manual processes Improve deployment safety through automation and guardrails Strengthen CI/CD reliability and release processes Cross-Functional Reliability Advocacy Partner with engineering teams to improve system resilience Provide guidance on fault tolerance, scalability, and failure handling Contribute to architectural discussions with a reliability-first mindset Requirements: 5+ years of experience in SRE, Reliability Engineering, or Production Engineering Strong Linux systems and Networking expertise Experience managing containerized production systems Strong understanding of distributed systems and failure modes Experience defining and managing SLIs/SLOs Proven incident response and postmortem leadership experience Strong scripting or programming skills Experience with monitoring and alerting systems Excellent written communication skills Successful completion of a background check Preferred: Experience with GPU infrastructure or AI/ML platforms Experience improving reliability in high-growth or large scale environments Familiarity with GPU observability tooling Experience with Infrastructure as Code Experience working in startup environments Experience building internal reliability platforms or frameworks What You’ll Receive: The competitive base pay for this position ranges from $150,000- $200,000 usd. This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside. Generous medical, dental & vision plans Flexible PTO- take the time you need to recharge Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale. Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.
- We're seeking a skilled and proactive Site Reliability Engineer to join our team, ensuring the stability, security, and efficiency of our technological... ...-edge AI solutions to the government. This is a fully remote position for candidates in the continental U.S., with work...Remote work
$72.8k - $130k
...technology systems in accordance with modern design standardsThe Site Reliability Engineer will architect, develop, and maintain Optum Serve's cloud... ...cloud infrastructure.You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough...Remote workMinimum wageFull timeWork experience placementWork at officeLocal area- ...and Seattle). Squad as a whole is responsible for the core platform services, split into 2 teams: 1 is Platform Engineering, and the other is Site Reliability Engineering. This is for the Site Reliability Engineering team. • Need in-depth knowledge of Linux, Windows...Remote workLocal area
$150k - $200k
...depend on. We're a small, remote-first team. We take ownership... ...funding announcement: The Reliability team owns the availability,... ...reliability standards across engineering Designing incident response... ...production systems. As a Site Reliability Engineer on the Reliability...Remote workVisa sponsorshipWork visaFlexible hours- ...such as public cloud, data science, AI, engineering innovation, and IoT. Our customers... ...profitable, and growing. We are hiring a Site Reliability Engineer Our goal is to perfect... ...metrics and code. Location: Globally remote role The role We deploy and run OpenStack...Remote workFull timeWork at officeLocal areaWork from homeWorldwide
- ...such as public cloud, data science, AI, engineering innovation, and IoT. Our customers... ...profitable, and growing. We are hiring a Site Reliability / Gitops Engineer to our Information... ...Location : This role is available remotely in any timezone. As a Site...Remote workFull timeWork at officeWork from homeFlexible hours
$35 - $44 per hour
DescriptionKforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a Site Reliability Engineer (SRE) to support a large-scale system modernization and legacy platform retirement initiative. This role will focus on maintaining and...Remote work$125k - $185k
Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir builds the... ...children, and more.The RoleWe’re looking for Site Reliability Engineers who can help us build,... ...there are a few roles that allow for “Remote” work on an exceptional basis. If you are...Remote workFull timeWork experience placementWork at officeWork from homeRelocation package- ...education. Client is currently seeking a talented Software Engineer who is able to work into the Site Reliability Engineer role. This candidate is expected to work... ...in production and pre-production environments. Remote to start due to covid, but will go back on site Duration...Remote work
$230k - $250k
GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of... .... This role is based in Arlington, VA, as a hybrid/remote position.ResponsibilitiesResponsibilities:Design and maintain...Remote work$125k - $185k
...children, and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-... ...Based on business need, there are a few roles that allow for “Remote” work on an exceptional basis. If you are applying for one...Remote workFull timeWork experience placementWork at officeWork from homeRelocation package- ...GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8 to 15 yearsSkillsKubernetes... ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate...Remote work
$210k - $230k
GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient... ...This position is located in Arlington, VA and is a hybrid remote/onsite position.ResponsibilitiesKey Responsibilities:...Remote workCurrently hiring$78k - $124.75k
...bonus + benefitsJob Function: Engineering & ArchitectureSchedule: Full... ...American ExpressDescriptionSite Reliability Engineer I enhances system... ...platformsKnowledge of cloud‑based Site Reliability Engineering (SRE)... ..., firewalls, and secure remote accessLicenses & CertificationsCertification...Remote work- ...role (three days in the office/two days remote).Job Summary:With a "document first" approach... ...the availability, scalability, and reliability of systems and applications.What will be... ...Terraform or CloudFormation.Mentor junior engineers and provide technical guidance.Stay up-...Remote workWork at office
- ...cloud-native platforms to advanced release engineering practices, our teams are redefining how... ...: Hybrid - 2 days onsite, 3 days remote per weekSponsorship Notice: At this time... ...that accelerate development and improve reliability. Your work will directly influence how...Remote workH1bWork at officeVisa sponsorshipFlexible hours2 days per week
$190.8k - $267.1k
...information, visit .This role is remote friendly. Reddit has a... ...Reddit grow its business. The reliability of our Ads systems directly impacts... ...partners closely with Ads Engineering teams to improve reliability,... ....We're looking for a Staff Site Reliability Engineer who will...Remote workFor contractorsWork experience placementFlexible hours$150k - $180k
...operates through three business units: Remote Sensing (the data), Space Systems (the... ...are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale... ...organization.This position is based on-site in either our Arlington, VA office, Reston...Remote workPermanent employmentFull timeWork at officeLocal areaWorldwide- Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems... ...Three or more years of experience in Site Reliability Engineering, platform engineering... ...that integrate cloud platforms with remote sites, field equipment, industrial...Remote work
$91.7k - $163.7k
...Join us to start Caring. Connecting. Growing together. The Site Reliability Engineer will architect, develop, and maintain Optum Serve's cloud... ...cloud infrastructure. You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough...Remote workMinimum wageFull timeWork experience placementWork at officeLocal area$90k - $180k
...people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale,... ...performance, and operational excellence of Merlin.net — a remote monitoring platform designed to help doctors, cardiologists...Remote work$80k - $133k
...Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud... ...obligated to pay a placement fee.SummaryLocation: US - TX, San Antonio; US - VA, McLean; US - Remote (Any location)Type: Full timeRemote workPermanent employmentFull timeContract workFlexible hours$86.8k - $198k
Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether... ...expected to have their cameras on during meetings.Remote: If this position is listed as remote, there may still...Remote workFull timeContract workPart timeWork at officeLocal area$15k
...office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to... ...law.Compensation Range: $205K - $235KLocationBerkeley, CA; Remote, United StatesEmployment TypeFull timeLocation TypeRemoteDepartmentSoftwareCompensationBase...Remote workWork at officeLocal area$117k - $209.33k
...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,... ...(not on this external site).SummaryLocation: Idaho, USA - Remote; AMER - United States - Texas - PlanoType: Full timeRemote workFull timeFor contractors$146k - $194k
...Anduril as a lead provider of specialized engineering and products for Intelligence Community... ...requirements.ABOUT THE JOBAs a Site Reliability Engineer, your primary mission is to ensure... ...supporting edge. Patch/update management and remote management tooling.DevOps improvements:...Remote workFull timeWork experience placementImmediate start$87.12k - $151.25k
...be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Digital Site Reliability Sr Engineer - Remote to join our team in Memphis, Tennessee (US-TN), United States (US).Digital Site Reliability Senior EngineerWe are seeking...Remote workTemporary workWork at officeFlexible hours$91.7k - $163.7k
...technology systems in accordance with modern design standards.As a Site Reliability Engineer (SRE), you will play a key role in ensuring the... ...our customers (OSIT).You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough...Remote workMinimum wageFull timeWork experience placementLocal area- ...cloud-native platforms to advanced release engineering practices, our teams are redefining how... ...: Hybrid - 2 days onsite, 3 days remote per weekSponsorship Notice: At this time... ...office#LI-KC1#GMFjobsAbout The Role: The Site Reliability Engineer under the general direction...Remote workWork experience placementH1bWork at officeVisa sponsorshipFlexible hoursShift work2 days per week
$102.1k - $202.2k
...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole... ...EngineeringDiscipline: Site Reliability EngineeringCompany:... ...the cloud. Our work enables remote computing experiences that are... ...team, you will collaborate with engineers across disciplines to deliver...Remote workOngoing contractWork experience placementLocal area3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Remote Site Reliability Engineer. Be the first to apply!
- site reliability engineer Sunnyvale, CA
- site reliability engineer sre Sunnyvale, CA
- senior network engineer remote Sunnyvale, CA
- remote legal intern Sunnyvale, CA
- remote customer service chat Sunnyvale, CA
- junior designer remote Sunnyvale, CA
- remote software sales Sunnyvale, CA
- full time remote Sunnyvale, CA
- remote startup Sunnyvale, CA
- remote technician Sunnyvale, CA


