Remote Site Reliability Engineer
$150k - $200kGrabJobs
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on. We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale. Learn more in our CEO's funding announcement: . The Reliability team owns the availability, performance, and operational excellence of Runpod’s global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions. This team is responsible for: Defining and enforcing reliability standards across engineering Designing incident response processes and improving recovery times Building observability systems and reliability tooling Driving SLO adoption and production readiness reviews Reducing operational toil through automation The Reliability team works cross-functionally with Infrastructure, Product Engineering, and Support to ensure our systems remain stable and performant as we scale rapidly. We value proactive problem solving, automation-first thinking, and strong ownership of production systems. As a Site Reliability Engineer on the Reliability team, you will focus on ensuring the stability and resilience of Runpod’s distributed platform. You will partner with engineering teams to improve system design, strengthen observability, and prevent incidents before they happen. This role blends software engineering with production operations. You’ll work on reliability frameworks, SLO design, automation, and production hardening, reducing errors and improving performance across different services and infrastructure. This is a high-impact role central to maintaining trust with developers running critical AI workloads on Runpod. Your Impact Increase platform uptime and reduce incident frequency and duration Establish and operationalize SLIs/SLOs across services Improve MTTR through better tooling, automation, and runbooks Strengthen production readiness standards Drive long-term systemic reliability improvements You will influence how reliability is defined and measured across Runpod and help build the operational backbone of the company. Responsibilities: Reliability Engineering Define and implement SLIs/SLOs for critical services Lead incident response and coordinate cross-team mitigation efforts Conduct blameless postmortems and ensure corrective actions are completed Perform production readiness reviews for new services and features Identify systemic risks and drive preventative improvements Observability & Monitoring Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.) Improve signal-to-noise ratio in alerts and reduce alert fatigue Build internal tooling for reliability tracking and reporting Improve visibility into GPU performance and distributed systems health Automation & Toil Reduction Automate recurring operational workflows Build tools and scripts (Python, Go, Bash) to eliminate manual processes Improve deployment safety through automation and guardrails Strengthen CI/CD reliability and release processes Cross-Functional Reliability Advocacy Partner with engineering teams to improve system resilience Provide guidance on fault tolerance, scalability, and failure handling Contribute to architectural discussions with a reliability-first mindset Requirements: 5+ years of experience in SRE, Reliability Engineering, or Production Engineering Strong Linux systems and Networking expertise Experience managing containerized production systems Strong understanding of distributed systems and failure modes Experience defining and managing SLIs/SLOs Proven incident response and postmortem leadership experience Strong scripting or programming skills Experience with monitoring and alerting systems Excellent written communication skills Successful completion of a background check Preferred: Experience with GPU infrastructure or AI/ML platforms Experience improving reliability in high-growth or large scale environments Familiarity with GPU observability tooling Experience with Infrastructure as Code Experience working in startup environments Experience building internal reliability platforms or frameworks What You’ll Receive: The competitive base pay for this position ranges from $150,000- $200,000 usd. This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside. Generous medical, dental & vision plans Flexible PTO- take the time you need to recharge Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale. Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.
- We're seeking a skilled and proactive Site Reliability Engineer to join our team, ensuring the stability, security, and efficiency of our technological... ...-edge AI solutions to the government. This is a fully remote position for candidates in the continental U.S., with work...Remote work
$150k - $200k
...will depend on. We're a small, remote-first team. We take ownership... ...funding announcement: . The Reliability team owns the availability,... ...reliability standards across engineering Designing incident response processes... ...of production systems. As a Site Reliability Engineer on the...Remote workVisa sponsorshipWork visaFlexible hours$130k - $180k
...experienced and innovative leaders and engineers in the field. Where we work Headquartered... ...team. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure... ...systems. Working conditions: Primarily remote Occasional travel to data centers required...Remote workTemporary workWork at officeImmediate startFlexible hours$114k - $148k
...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based...Remote workFull timeTemporary workWork experience placement$104.9k - $174.7k
...link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability... ...you may work a hybrid schedule. If not, this role is fully remote. We do not restrict applicants based on job site or posting...Remote workFull timeWork at officeLocal areaWork from home$230k - $250k
GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of... .... This role is based in Arlington, VA, as a hybrid/remote position.ResponsibilitiesResponsibilities:Design and maintain...Remote work$125k - $185k
...children, and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-... ...Based on business need, there are a few roles that allow for “Remote” work on an exceptional basis. If you are applying for one...Remote workFull timeWork experience placementWork at officeWork from homeRelocation package$125k - $185k
Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir builds the... ...children, and more.The RoleWe’re looking for Site Reliability Engineers who can help us build,... ...there are a few roles that allow for “Remote” work on an exceptional basis. If you are...Remote workFull timeWork experience placementWork at officeWork from homeRelocation package$65 - $75 per hour
DescriptionKforce has a client seeking a remote Senior Site Reliability Engineer to be a l be a leading member of the team working with a diverse range of technologies. You will enjoy working in a friendly environment and benefit from our investment in staff. The role also...Remote work- ...GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8 to 15 yearsSkillsKubernetes... ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate...Remote work
- ...role (three days in the office/two days remote).Job Summary:With a "document first" approach... ...the availability, scalability, and reliability of systems and applications.What will be... ...Terraform or CloudFormation.Mentor junior engineers and provide technical guidance.Stay up-...Remote workWork at office
$96.8k - $145.2k
...thinking organization, apply now.We are currently seeking a Site Reliability Engineer (Onsite Hybrid) to join our team in Plano, Texas (US-TX),... ...tailored to each client’s needs. While many positions offer remote or hybrid work options, these arrangements are subject to change...Remote workTemporary workWork at officeFlexible hours$67.2k - $100.8k
...Locations and Workstyle:Blue Bell, PA: Primarily remote; candidates should be within commuting... ...partners (IT, Security, DevOps, Engineering) to improve operational health and apply SRE best practicesSupport the reliability, availability, scalability, and performance...Remote workTemporary workH1bWork at office$146k - $194k
...Anduril as a lead provider of specialized engineering and products for Intelligence Community... ...requirements.ABOUT THE JOBAs a Site Reliability Engineer, your primary mission is to ensure... ...supporting edge. Patch/update management and remote management tooling.DevOps improvements:...Remote workFull timeWork experience placementImmediate start$80k - $133k
...Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud... ...obligated to pay a placement fee.SummaryLocation: US - TX, San Antonio; US - VA, McLean; US - Remote (Any location)Type: Full timeRemote workPermanent employmentFull timeContract workFlexible hours- ...thinking organization, apply now.We are currently seeking a Site Reliability Engineer to join our team in Westlake, Texas (US-TX), United States... ...compensation for specific roles. The starting pay range for this remote role is 85,000-140,000 This range reflects the minimum and...Remote workFull timeTemporary workWork at officeFlexible hours
- ...cloud-native platforms to advanced release engineering practices, our teams are redefining how... ...: Hybrid - 2 days onsite, 3 days remote per weekSponsorship Notice: At this time... ...that accelerate development and improve reliability. Your work will directly influence how...Remote workH1bWork at officeVisa sponsorshipFlexible hours2 days per week
$62k - $141k
Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether... ...expected to have their cameras on during meetings.Remote: If this position is listed as remote, there may still...Remote workFull timeContract workPart timeWork at officeLocal area$90k - $180k
...people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale,... ...performance, and operational excellence of Merlin.net — a remote monitoring platform designed to help doctors, cardiologists...Remote work$15k
...office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to... ...law.Compensation Range: $205K - $235KLocationBerkeley, CA; Remote, United StatesEmployment TypeFull timeLocation TypeRemoteDepartmentSoftwareCompensationBase...Remote workWork at officeLocal area$138.1k - $198.2k
...of applications are received.This is a remote role based out of the US.The successful... ...technology that simply works. The SRE Engineering Enablement Team supports our CI Platforms... ...engineers at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter...Remote workPermanent employmentFull timeTemporary workWork experience placementLocal areaFlexible hours$117k - $209.33k
...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,... ...(not on this external site).SummaryLocation: Idaho, USA - Remote; AMER - United States - Texas - PlanoType: Full timeRemote workFull timeFor contractors$87.12k - $151.25k
...be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Digital Site Reliability Sr Engineer - Remote to join our team in Memphis, Tennessee (US-TN), United States (US).Digital Site Reliability Senior EngineerWe are seeking...Remote workTemporary workWork at officeFlexible hours$104.43k - $156.65k
..., Comcast prefers to have employees on-site collaborating unless the team has been... ...than 100 miles from the office for the remote option.)Job SummaryCOMCAST Technology Solutions... ..., Paramount+, and many others.Our Site Reliability Engineering (SRE) team is at the heart of our...Remote workPermanent employmentFull timeWork at officeWorldwideFlexible hours$140k - $150k
WORK OPTION: Remote_________________The NBA is hiring a Senior Site Reliability Engineer (SRE) - Messaging & Collaboration to ensure the availability, performance, and reliability of enterprise messaging and collaboration platforms, including Microsoft Exchange Online (...Remote workFull timeTemporary workLocal areaWeekend work$150k - $180k
...operates through three business units: Remote Sensing (the data), Space Systems (the... ...are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale... ...organization.This position is based on-site in either our Arlington, VA office, Reston...Remote workPermanent employmentFull timeWork at officeLocal areaWorldwide$110k - $120k
...expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading... ...protected by applicable discrimination laws.SummaryLocation: Remote - New York, US; Remote - Kansas, US; Remote - Pennsylvania...Remote workOngoing contractFull timeCasual workFlexible hours- Site Reliability Engineers are responsible for ensuring the availability, reliability, scalability, and performance of the firm’s most critical customer... ....This is an on-site position located in Springfield, MO. Remote work is not an option for this position.Primary...Remote workLocal areaFlexible hoursShift work
- ...cloud-native platforms to advanced release engineering practices, our teams are redefining how... ...: Hybrid - 2 days onsite, 3 days remote per weekSponsorship Notice: At this time... ...office#LI-KC1#GMFjobsAbout The Role: The Site Reliability Engineer under the general direction...Remote workWork experience placementH1bWork at officeVisa sponsorshipFlexible hoursShift work2 days per week
$102.1k - $202.2k
...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole... ...EngineeringDiscipline: Site Reliability EngineeringCompany:... ...the cloud. Our work enables remote computing experiences that are... ...team, you will collaborate with engineers across disciplines to deliver...Remote workOngoing contractWork experience placementLocal area3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Remote Site Reliability Engineer. Be the first to apply!
- operations coordinator remote Nashville, TN
- remote coordinator Nashville, TN
- remote broker Nashville, TN
- remote design intern Nashville, TN
- remote accounts receivable Nashville, TN
- remote b2b sales Nashville, TN
- immediate hire remote Nashville, TN
- remote senior business analyst Nashville, TN
- remote insurance Nashville, TN
- remote customer service chat Nashville, TN

