Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Remote Site Reliability Engineer

$150k - $200k

GrabJobs

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on. We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale. Learn more in our CEO's funding announcement: . The Reliability team owns the availability, performance, and operational excellence of Runpod’s global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions. This team is responsible for: Defining and enforcing reliability standards across engineering Designing incident response processes and improving recovery times Building observability systems and reliability tooling Driving SLO adoption and production readiness reviews Reducing operational toil through automation The Reliability team works cross-functionally with Infrastructure, Product Engineering, and Support to ensure our systems remain stable and performant as we scale rapidly. We value proactive problem solving, automation-first thinking, and strong ownership of production systems. As a Site Reliability Engineer on the Reliability team, you will focus on ensuring the stability and resilience of Runpod’s distributed platform. You will partner with engineering teams to improve system design, strengthen observability, and prevent incidents before they happen. This role blends software engineering with production operations. You’ll work on reliability frameworks, SLO design, automation, and production hardening, reducing errors and improving performance across different services and infrastructure. This is a high-impact role central to maintaining trust with developers running critical AI workloads on Runpod. Your Impact Increase platform uptime and reduce incident frequency and duration Establish and operationalize SLIs/SLOs across services Improve MTTR through better tooling, automation, and runbooks Strengthen production readiness standards Drive long-term systemic reliability improvements You will influence how reliability is defined and measured across Runpod and help build the operational backbone of the company. Responsibilities: Reliability Engineering Define and implement SLIs/SLOs for critical services Lead incident response and coordinate cross-team mitigation efforts Conduct blameless postmortems and ensure corrective actions are completed Perform production readiness reviews for new services and features Identify systemic risks and drive preventative improvements Observability & Monitoring Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.) Improve signal-to-noise ratio in alerts and reduce alert fatigue Build internal tooling for reliability tracking and reporting Improve visibility into GPU performance and distributed systems health Automation & Toil Reduction Automate recurring operational workflows Build tools and scripts (Python, Go, Bash) to eliminate manual processes Improve deployment safety through automation and guardrails Strengthen CI/CD reliability and release processes Cross-Functional Reliability Advocacy Partner with engineering teams to improve system resilience Provide guidance on fault tolerance, scalability, and failure handling Contribute to architectural discussions with a reliability-first mindset Requirements: 5+ years of experience in SRE, Reliability Engineering, or Production Engineering Strong Linux systems and Networking expertise Experience managing containerized production systems Strong understanding of distributed systems and failure modes Experience defining and managing SLIs/SLOs Proven incident response and postmortem leadership experience Strong scripting or programming skills Experience with monitoring and alerting systems Excellent written communication skills Successful completion of a background check Preferred: Experience with GPU infrastructure or AI/ML platforms Experience improving reliability in high-growth or large scale environments Familiarity with GPU observability tooling Experience with Infrastructure as Code Experience working in startup environments Experience building internal reliability platforms or frameworks What You’ll Receive: The competitive base pay for this position ranges from $150,000- $200,000 usd. This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside. Generous medical, dental & vision plans Flexible PTO- take the time you need to recharge Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale. Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Remote Site Reliability Engineer in Sunnyvale, CA vacancy
  • We're seeking a skilled and proactive Site Reliability Engineer to join our team, ensuring the stability, security, and efficiency of our technological...  ...-edge AI solutions to the government. This is a fully remote position for candidates in the continental U.S., with work... 
    Remote work

    Knexus

    Vienna, VA
    4 days ago
  • $72.8k - $130k

     ...technology systems in accordance with modern design standardsThe Site Reliability Engineer will architect, develop, and maintain Optum Serve's cloud...  ...cloud infrastructure.You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough... 
    Remote work
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area

    UnitedHealth Group

    Eden Prairie, MN
    14 hours ago
  •  ...and Seattle). Squad as a whole is responsible for the core platform services, split into 2 teams: 1 is Platform Engineering, and the other is Site Reliability Engineering. This is for the Site Reliability Engineering team. • Need in-depth knowledge of Linux, Windows... 
    Remote work
    Local area

    My3Tech Inc

    United States
    5 days ago
  • $150k - $200k

     ...depend on. We're a small, remote-first team. We take ownership...  ...funding announcement: The Reliability team owns the availability,...  ...reliability standards across engineering Designing incident response...  ...production systems. As a Site Reliability Engineer on the Reliability... 
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    GrabJobs

    Plano, TX
    1 day ago
  •  ...such as public cloud, data science, AI, engineering innovation, and IoT. Our customers...  ...profitable, and growing. We are hiring a Site Reliability Engineer Our goal is to perfect...  ...metrics and code. Location: Globally remote role The role We deploy and run OpenStack... 
    Remote work
    Full time
    Work at office
    Local area
    Work from home
    Worldwide

    Canonical

    Remote
    16 days ago
  •  ...such as public cloud, data science, AI, engineering innovation, and IoT. Our customers...  ...profitable, and growing. We are hiring a  Site Reliability / Gitops Engineer to our Information...  ...Location : This role is available remotely in any timezone. As a Site... 
    Remote work
    Full time
    Work at office
    Work from home
    Flexible hours

    Canonical

    Remote
    16 days ago
  • $35 - $44 per hour

    DescriptionKforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a Site Reliability Engineer (SRE) to support a large-scale system modernization and legacy platform retirement initiative. This role will focus on maintaining and... 
    Remote work

    KForce

    Atlanta, GA
    3 days ago
  • $125k - $185k

    Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir builds the...  ...children, and more.The RoleWe’re looking for Site Reliability Engineers who can help us build,...  ...there are a few roles that allow for “Remote” work on an exceptional basis. If you are... 
    Remote work
    Full time
    Work experience placement
    Work at office
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    1 day ago
  •  ...education. Client is currently seeking a talented Software Engineer who is able to work into the Site Reliability Engineer role. This candidate is expected to work...  ...in production and pre-production environments. Remote to start due to covid, but will go back on site Duration... 
    Remote work

    Intelliswift

    Durham, NC
    4 days ago
  • $230k - $250k

    GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of...  .... This role is based in Arlington, VA, as a hybrid/remote position.ResponsibilitiesResponsibilities:Design and maintain... 
    Remote work

    Govcio

    Arlington, VA
    4 days ago
  • $125k - $185k

     ...children, and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-...  ...Based on business need, there are a few roles that allow for “Remote” work on an exceptional basis. If you are applying for one... 
    Remote work
    Full time
    Work experience placement
    Work at office
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    4 days ago
  •  ...GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8 to 15 yearsSkillsKubernetes...  ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate... 
    Remote work

    SRI Tech

    Plano, TX
    4 days ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient...  ...This position is located in Arlington, VA and is a hybrid remote/onsite position.ResponsibilitiesKey Responsibilities:... 
    Remote work
    Currently hiring

    Govcio

    Arlington, VA
    1 day ago
  • $78k - $124.75k

     ...bonus + benefitsJob Function: Engineering & ArchitectureSchedule: Full...  ...American ExpressDescriptionSite Reliability Engineer I enhances system...  ...platformsKnowledge of cloud‑based Site Reliability Engineering (SRE)...  ..., firewalls, and secure remote accessLicenses & CertificationsCertification... 
    Remote work

    American Express

    Sunrise, FL
    14 hours ago
  •  ...role (three days in the office/two days remote).Job Summary:With a "document first" approach...  ...the availability, scalability, and reliability of systems and applications.What will be...  ...Terraform or CloudFormation.Mentor junior engineers and provide technical guidance.Stay up-... 
    Remote work
    Work at office

    Interactive Brokers

    Greenwich, CT
    3 days ago
  •  ...cloud-native platforms to advanced release engineering practices, our teams are redefining how...  ...: Hybrid - 2 days onsite, 3 days remote per weekSponsorship Notice: At this time...  ...that accelerate development and improve reliability. Your work will directly influence how... 
    Remote work
    H1b
    Work at office
    Visa sponsorship
    Flexible hours
    2 days per week

    GM Financial

    Arlington, TX
    4 days ago
  • $190.8k - $267.1k

     ...information, visit .This role is remote friendly. Reddit has a...  ...Reddit grow its business. The reliability of our Ads systems directly impacts...  ...partners closely with Ads Engineering teams to improve reliability,...  ....We're looking for a Staff Site Reliability Engineer who will... 
    Remote work
    For contractors
    Work experience placement
    Flexible hours

    Reddit

    San Francisco, CA
    3 days ago
  • $150k - $180k

     ...operates through three business units: Remote Sensing (the data), Space Systems (the...  ...are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale...  ...organization.This position is based on-site in either our Arlington, VA office, Reston... 
    Remote work
    Permanent employment
    Full time
    Work at office
    Local area
    Worldwide

    Umbra

    Arlington, VA
    4 days ago
  • Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems...  ...Three or more years of experience in Site Reliability Engineering, platform engineering...  ...that integrate cloud platforms with remote sites, field equipment, industrial... 
    Remote work

    Patterson-UTI

    Houston, TX
    3 days ago
  • $91.7k - $163.7k

     ...Join us to start Caring. Connecting. Growing together. The Site Reliability Engineer will architect, develop, and maintain Optum Serve's cloud...  ...cloud infrastructure. You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough... 
    Remote work
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area

    UnitedHealth Group

    Eden Prairie, MN
    14 hours ago
  • $90k - $180k

     ...people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale,...  ...performance, and operational excellence of Merlin.net — a remote monitoring platform designed to help doctors, cardiologists... 
    Remote work

    Abbott

    Sunnyvale, CA
    1 day ago
  • $80k - $133k

     ...Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud...  ...obligated to pay a placement fee.SummaryLocation: US - TX, San Antonio; US - VA, McLean; US - Remote (Any location)Type: Full time
    Remote work
    Permanent employment
    Full time
    Contract work
    Flexible hours

    Guidehouse

    San Antonio, TX
    1 day ago
  • $86.8k - $198k

    Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether...  ...expected to have their cameras on during meetings.Remote: If this position is listed as remote, there may still... 
    Remote work
    Full time
    Contract work
    Part time
    Work at office
    Local area

    Booz Allen Hamilton

    McLean, VA
    4 days ago
  • $15k

     ...office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to...  ...law.Compensation Range: $205K - $235KLocationBerkeley, CA; Remote, United StatesEmployment TypeFull timeLocation TypeRemoteDepartmentSoftwareCompensationBase... 
    Remote work
    Work at office
    Local area

    The Voleon Group

    Berkeley, CA
    2 days ago
  • $117k - $209.33k

     ...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,...  ...(not on this external site).SummaryLocation: Idaho, USA - Remote; AMER - United States - Texas - PlanoType: Full time
    Remote work
    Full time
    For contractors

    Autodesk

    Plano, TX
    14 hours ago
  • $146k - $194k

     ...Anduril as a lead provider of specialized engineering and products for Intelligence Community...  ...requirements.ABOUT THE JOBAs a Site Reliability Engineer, your primary mission is to ensure...  ...supporting edge. Patch/update management and remote management tooling.DevOps improvements:... 
    Remote work
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Reston, VA
    2 days ago
  • $87.12k - $151.25k

     ...be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Digital Site Reliability Sr Engineer - Remote to join our team in Memphis, Tennessee (US-TN), United States (US).Digital Site Reliability Senior EngineerWe are seeking... 
    Remote work
    Temporary work
    Work at office
    Flexible hours

    NTT DATA

    Memphis, TN
    1 day ago
  • $91.7k - $163.7k

     ...technology systems in accordance with modern design standards.As a Site Reliability Engineer (SRE), you will play a key role in ensuring the...  ...our customers (OSIT).You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough... 
    Remote work
    Minimum wage
    Full time
    Work experience placement
    Local area

    UnitedHealth Group

    Eden Prairie, MN
    14 hours ago
  •  ...cloud-native platforms to advanced release engineering practices, our teams are redefining how...  ...: Hybrid - 2 days onsite, 3 days remote per weekSponsorship Notice: At this time...  ...office#LI-KC1#GMFjobsAbout The Role: The Site Reliability Engineer under the general direction... 
    Remote work
    Work experience placement
    H1b
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work
    2 days per week

    GM Financial

    Arlington, TX
    3 days ago
  • $102.1k - $202.2k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...the cloud. Our work enables remote computing experiences that are...  ...team, you will collaborate with engineers across disciplines to deliver... 
    Remote work
    Ongoing contract
    Work experience placement
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    14 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Remote Site Reliability Engineer. Be the first to apply!