Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Remote Site Reliability Engineer

$150k - $200k

GrabJobs

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on. We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale. Learn more in our CEO's funding announcement: . The Reliability team owns the availability, performance, and operational excellence of Runpod’s global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions. This team is responsible for: Defining and enforcing reliability standards across engineering Designing incident response processes and improving recovery times Building observability systems and reliability tooling Driving SLO adoption and production readiness reviews Reducing operational toil through automation The Reliability team works cross-functionally with Infrastructure, Product Engineering, and Support to ensure our systems remain stable and performant as we scale rapidly. We value proactive problem solving, automation-first thinking, and strong ownership of production systems. As a Site Reliability Engineer on the Reliability team, you will focus on ensuring the stability and resilience of Runpod’s distributed platform. You will partner with engineering teams to improve system design, strengthen observability, and prevent incidents before they happen. This role blends software engineering with production operations. You’ll work on reliability frameworks, SLO design, automation, and production hardening, reducing errors and improving performance across different services and infrastructure. This is a high-impact role central to maintaining trust with developers running critical AI workloads on Runpod. Your Impact Increase platform uptime and reduce incident frequency and duration Establish and operationalize SLIs/SLOs across services Improve MTTR through better tooling, automation, and runbooks Strengthen production readiness standards Drive long-term systemic reliability improvements You will influence how reliability is defined and measured across Runpod and help build the operational backbone of the company. Responsibilities: Reliability Engineering Define and implement SLIs/SLOs for critical services Lead incident response and coordinate cross-team mitigation efforts Conduct blameless postmortems and ensure corrective actions are completed Perform production readiness reviews for new services and features Identify systemic risks and drive preventative improvements Observability & Monitoring Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.) Improve signal-to-noise ratio in alerts and reduce alert fatigue Build internal tooling for reliability tracking and reporting Improve visibility into GPU performance and distributed systems health Automation & Toil Reduction Automate recurring operational workflows Build tools and scripts (Python, Go, Bash) to eliminate manual processes Improve deployment safety through automation and guardrails Strengthen CI/CD reliability and release processes Cross-Functional Reliability Advocacy Partner with engineering teams to improve system resilience Provide guidance on fault tolerance, scalability, and failure handling Contribute to architectural discussions with a reliability-first mindset Requirements: 5+ years of experience in SRE, Reliability Engineering, or Production Engineering Strong Linux systems and Networking expertise Experience managing containerized production systems Strong understanding of distributed systems and failure modes Experience defining and managing SLIs/SLOs Proven incident response and postmortem leadership experience Strong scripting or programming skills Experience with monitoring and alerting systems Excellent written communication skills Successful completion of a background check Preferred: Experience with GPU infrastructure or AI/ML platforms Experience improving reliability in high-growth or large scale environments Familiarity with GPU observability tooling Experience with Infrastructure as Code Experience working in startup environments Experience building internal reliability platforms or frameworks What You’ll Receive: The competitive base pay for this position ranges from $150,000- $200,000 usd. This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside. Generous medical, dental & vision plans Flexible PTO- take the time you need to recharge Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale. Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Remote Site Reliability Engineer in Nashville, TN vacancy
  • We're seeking a skilled and proactive Site Reliability Engineer to join our team, ensuring the stability, security, and efficiency of our technological...  ...-edge AI solutions to the government. This is a fully remote position for candidates in the continental U.S., with work... 
    Remote work

    Knexus

    Vienna, VA
    3 days ago
  • $150k - $200k

     ...will depend on. We're a small, remote-first team. We take ownership...  ...funding announcement: . The Reliability team owns the availability,...  ...reliability standards across engineering Designing incident response processes...  ...of production systems. As a Site Reliability Engineer on the... 
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    GrabJobs

    Henderson, NV
    2 days ago
  • $130k - $180k

     ...experienced and innovative leaders and engineers in the field. Where we work Headquartered...  ...team. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure...  ...systems. Working conditions: Primarily remote Occasional travel to data centers required... 
    Remote work
    Temporary work
    Work at office
    Immediate start
    Flexible hours

    GrabJobs

    Anchorage, AK
    16 hours ago
  • $114k - $148k

     ...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based... 
    Remote work
    Full time
    Temporary work
    Work experience placement

    GrabJobs

    Lubbock, TX
    16 hours ago
  • $104.9k - $174.7k

     ...link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability...  ...you may work a hybrid schedule. If not, this role is fully remote. We do not restrict applicants based on job site or posting... 
    Remote work
    Full time
    Work at office
    Local area
    Work from home

    RELX Group

    Allen, TX
    2 days ago
  • $230k - $250k

    GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of...  .... This role is based in Arlington, VA, as a hybrid/remote position.ResponsibilitiesResponsibilities:Design and maintain... 
    Remote work

    Govcio

    Arlington, VA
    3 days ago
  • $125k - $185k

     ...children, and more.The RoleWe’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-...  ...Based on business need, there are a few roles that allow for “Remote” work on an exceptional basis. If you are applying for one... 
    Remote work
    Full time
    Work experience placement
    Work at office
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    3 days ago
  • $125k - $185k

    Washington, D.C.Engineering /Full-time /HybridA World-Changing CompanyPalantir builds the...  ...children, and more.The RoleWe’re looking for Site Reliability Engineers who can help us build,...  ...there are a few roles that allow for “Remote” work on an exceptional basis. If you are... 
    Remote work
    Full time
    Work experience placement
    Work at office
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    16 hours ago
  • $65 - $75 per hour

    DescriptionKforce has a client seeking a remote Senior Site Reliability Engineer to be a l be a leading member of the team working with a diverse range of technologies. You will enjoy working in a friendly environment and benefit from our investment in staff. The role also... 
    Remote work

    KForce

    Boca Raton, FL
    2 days ago
  •  ...GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8 to 15 yearsSkillsKubernetes...  ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate... 
    Remote work

    SRI Tech

    Plano, TX
    3 days ago
  •  ...role (three days in the office/two days remote).Job Summary:With a "document first" approach...  ...the availability, scalability, and reliability of systems and applications.What will be...  ...Terraform or CloudFormation.Mentor junior engineers and provide technical guidance.Stay up-... 
    Remote work
    Work at office

    Interactive Brokers

    Greenwich, CT
    2 days ago
  • $96.8k - $145.2k

     ...thinking organization, apply now.We are currently seeking a Site Reliability Engineer (Onsite Hybrid) to join our team in Plano, Texas (US-TX),...  ...tailored to each client’s needs. While many positions offer remote or hybrid work options, these arrangements are subject to change... 
    Remote work
    Temporary work
    Work at office
    Flexible hours

    NTT DATA

    Plano, TX
    16 hours ago
  • $67.2k - $100.8k

     ...Locations and Workstyle:Blue Bell, PA: Primarily remote; candidates should be within commuting...  ...partners (IT, Security, DevOps, Engineering) to improve operational health and apply SRE best practicesSupport the reliability, availability, scalability, and performance... 
    Remote work
    Temporary work
    H1b
    Work at office

    ADT Worldwide

    Blue Bell, PA
    3 days ago
  • $146k - $194k

     ...Anduril as a lead provider of specialized engineering and products for Intelligence Community...  ...requirements.ABOUT THE JOBAs a Site Reliability Engineer, your primary mission is to ensure...  ...supporting edge. Patch/update management and remote management tooling.DevOps improvements:... 
    Remote work
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Reston, VA
    1 day ago
  • $80k - $133k

     ...Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud...  ...obligated to pay a placement fee.SummaryLocation: US - TX, San Antonio; US - VA, McLean; US - Remote (Any location)Type: Full time
    Remote work
    Permanent employment
    Full time
    Contract work
    Flexible hours

    Guidehouse

    San Antonio, TX
    16 hours ago
  •  ...thinking organization, apply now.We are currently seeking a Site Reliability Engineer to join our team in Westlake, Texas (US-TX), United States...  ...compensation for specific roles. The starting pay range for this remote role is 85,000-140,000 This range reflects the minimum and... 
    Remote work
    Full time
    Temporary work
    Work at office
    Flexible hours

    NTT DATA

    Texas
    2 days ago
  •  ...cloud-native platforms to advanced release engineering practices, our teams are redefining how...  ...: Hybrid - 2 days onsite, 3 days remote per weekSponsorship Notice: At this time...  ...that accelerate development and improve reliability. Your work will directly influence how... 
    Remote work
    H1b
    Work at office
    Visa sponsorship
    Flexible hours
    2 days per week

    GM Financial

    Arlington, TX
    3 days ago
  • $62k - $141k

    Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether...  ...expected to have their cameras on during meetings.Remote: If this position is listed as remote, there may still... 
    Remote work
    Full time
    Contract work
    Part time
    Work at office
    Local area

    Booz Allen Hamilton

    Aurora, CO
    2 days ago
  • $90k - $180k

     ...people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale,...  ...performance, and operational excellence of Merlin.net — a remote monitoring platform designed to help doctors, cardiologists... 
    Remote work

    Abbott

    Sunnyvale, CA
    11 hours ago
  • $15k

     ...office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to...  ...law.Compensation Range: $205K - $235KLocationBerkeley, CA; Remote, United StatesEmployment TypeFull timeLocation TypeRemoteDepartmentSoftwareCompensationBase... 
    Remote work
    Work at office
    Local area

    The Voleon Group

    Berkeley, CA
    1 day ago
  • $138.1k - $198.2k

     ...of applications are received.This is a remote role based out of the US.The successful...  ...technology that simply works.  The SRE Engineering Enablement Team supports our CI Platforms...  ...engineers at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter... 
    Remote work
    Permanent employment
    Full time
    Temporary work
    Work experience placement
    Local area
    Flexible hours

    CISCO Systems

    New York, NY
    2 days ago
  • $117k - $209.33k

     ...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,...  ...(not on this external site).SummaryLocation: Idaho, USA - Remote; AMER - United States - Texas - PlanoType: Full time
    Remote work
    Full time
    For contractors

    Autodesk

    Plano, TX
    4 days ago
  • $87.12k - $151.25k

     ...be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Digital Site Reliability Sr Engineer - Remote to join our team in Memphis, Tennessee (US-TN), United States (US).Digital Site Reliability Senior EngineerWe are seeking... 
    Remote work
    Temporary work
    Work at office
    Flexible hours

    NTT DATA

    Memphis, TN
    16 hours ago
  • $104.43k - $156.65k

     ..., Comcast prefers to have employees on-site collaborating unless the team has been...  ...than 100 miles from the office for the remote option.)Job SummaryCOMCAST Technology Solutions...  ..., Paramount+, and many others.Our Site Reliability Engineering (SRE) team is at the heart of our... 
    Remote work
    Permanent employment
    Full time
    Work at office
    Worldwide
    Flexible hours

    Comcast

    Centennial, CO
    16 hours ago
  • $140k - $150k

    WORK OPTION: Remote_________________The NBA is hiring a Senior Site Reliability Engineer (SRE) - Messaging & Collaboration to ensure the availability, performance, and reliability of enterprise messaging and collaboration platforms, including Microsoft Exchange Online (... 
    Remote work
    Full time
    Temporary work
    Local area
    Weekend work

    National Basketball Association

    Secaucus, NJ
    4 days ago
  • $150k - $180k

     ...operates through three business units: Remote Sensing (the data), Space Systems (the...  ...are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale...  ...organization.This position is based on-site in either our Arlington, VA office, Reston... 
    Remote work
    Permanent employment
    Full time
    Work at office
    Local area
    Worldwide

    Umbra

    Arlington, VA
    2 days ago
  • $110k - $120k

     ...expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading...  ...protected by applicable discrimination laws.SummaryLocation: Remote - New York, US; Remote - Kansas, US; Remote - Pennsylvania... 
    Remote work
    Ongoing contract
    Full time
    Casual work
    Flexible hours

    SS&C Technologies

    Pennsylvania
    11 hours ago
  • Site Reliability Engineers are responsible for ensuring the availability, reliability, scalability, and performance of the firm’s most critical customer...  ....This is an on-site position located in Springfield, MO. Remote work is not an option for this position.Primary... 
    Remote work
    Local area
    Flexible hours
    Shift work

    O'Reilly Auto Parts

    Springfield, MO
    4 days ago
  •  ...cloud-native platforms to advanced release engineering practices, our teams are redefining how...  ...: Hybrid - 2 days onsite, 3 days remote per weekSponsorship Notice: At this time...  ...office#LI-KC1#GMFjobsAbout The Role: The Site Reliability Engineer under the general direction... 
    Remote work
    Work experience placement
    H1b
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work
    2 days per week

    GM Financial

    Arlington, TX
    2 days ago
  • $102.1k - $202.2k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...the cloud. Our work enables remote computing experiences that are...  ...team, you will collaborate with engineers across disciplines to deliver... 
    Remote work
    Ongoing contract
    Work experience placement
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Remote Site Reliability Engineer. Be the first to apply!