Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$150k - $200k

Runpod

Site Reliability Engineer

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.

The Reliability team owns the availability, performance, and operational excellence of Runpod's global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions.

This team is responsible for:

  • Defining and enforcing reliability standards across engineering
  • Designing incident response processes and improving recovery times
  • Building observability systems and reliability tooling
  • Driving SLO adoption and production readiness reviews
  • Reducing operational toil through automation

The Reliability team works cross-functionally with Infrastructure, Product Engineering, and Support to ensure our systems remain stable and performant as we scale rapidly. We value proactive problem solving, automation-first thinking, and strong ownership of production systems.

As a Site Reliability Engineer on the Reliability team, you will focus on ensuring the stability and resilience of Runpod's distributed platform. You will partner with engineering teams to improve system design, strengthen observability, and prevent incidents before they happen.

This role blends software engineering with production operations. You'll work on reliability frameworks, SLO design, automation, and production hardening, reducing errors and improving performance across different services and infrastructure.

This is a high-impact role central to maintaining trust with developers running critical AI workloads on Runpod.

Your impact will include:

  • Increase platform uptime and reduce incident frequency and duration
  • Establish and operationalize SLIs/SLOs across services
  • Improve MTTR through better tooling, automation, and runbooks
  • Strengthen production readiness standards
  • Drive long-term systemic reliability improvements

You will influence how reliability is defined and measured across Runpod and help build the operational backbone of the company.

Responsibilities include:

  • Reliability Engineering
    • Define and implement SLIs/SLOs for critical services
    • Lead incident response and coordinate cross-team mitigation efforts
    • Conduct blameless postmortems and ensure corrective actions are completed
    • Perform production readiness reviews for new services and features
    • Identify systemic risks and drive preventative improvements
  • Observability & Monitoring
    • Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.)
    • Improve signal-to-noise ratio in alerts and reduce alert fatigue
    • Build internal tooling for reliability tracking and reporting
    • Improve visibility into GPU performance and distributed systems health
  • Automation & Toil Reduction
    • Automate recurring operational workflows
    • Build tools and scripts (Python, Go, Bash) to eliminate manual processes
    • Improve deployment safety through automation and guardrails
    • Strengthen CI/CD reliability and release processes
  • Cross-Functional Reliability Advocacy
    • Partner with engineering teams to improve system resilience
    • Provide guidance on fault tolerance, scalability, and failure handling
    • Contribute to architectural discussions with a reliability-first mindset

Requirements include:

  • 5+ years of experience in SRE, Reliability Engineering, or Production Engineering
  • Strong Linux systems and Networking expertise
  • Experience managing containerized production systems
  • Strong understanding of distributed systems and failure modes
  • Experience defining and managing SLIs/SLOs
  • Proven incident response and postmortem leadership experience
  • Strong scripting or programming skills
  • Experience with monitoring and alerting systems
  • Excellent written communication skills
  • Successful completion of a background check

Preferred:

  • Experience with GPU infrastructure or AI/ML platforms
  • Experience improving reliability in high-growth or large scale environments
  • Familiarity with GPU observability tooling
  • Experience with Infrastructure as Code
  • Experience working in startup environments
  • Experience building internal reliability platforms or frameworks

What you'll receive:

  • The competitive base pay for this position ranges from $150,000- $200,000 usd. This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate's experience, qualifications, and location
  • Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside.
  • Generous medical, dental & vision plans
  • Flexible PTO- take the time you need to recharge
  • Most roles are remote work first with an inclusive, collaborative teams utilizing Slack as the main form of internal communication
  • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.

Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in United States vacancy
  • $125k - $145k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER, GNCSpaceX’s mission is to make humanity multiplanetary by developing fully and rapidly reusable launch systems capable of... 
    Suggested
    Permanent employment
    Temporary work
    Flexible hours
    Weekend work

    SpaceX

    Hawthorne, CA
    5 days ago
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Costa... 
    Suggested
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    5 days ago
  • Company DescriptionComtech LLC is a woman-owned small business focused on delivering end-to-end solutions and products. Since 1998, we have successfully serviced enterprises across the public and private sectors, and the Department of Defense. Our services span all aspects...
    Suggested

    Comtech

    Seattle, WA
    5 days ago
  • $81.1k - $187k

     ...architect infrastructure and service to ensure reliability and functionality. Forecasts demands and...  ...impact and develops knowledge of site reliability trends.Only Oracle brings together...  ...guidance and mentorship to junior engineers. Communicate status, risks, blockers, and... 
    Suggested
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    5 days ago
  • $138.4k - $173k

     ...infrastructure as well as help improve the reliability, quality of services and overall...  ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the reliability...  ...about our locations by visiting our site.Compensation & BenefitsThe base salary that... 
    Suggested
    Full time
    Flexible hours

    AppFolio

    Dallas, TX
    3 days ago
  • $143k - $191k

     ...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental...  ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and... 
    Full time
    Temporary work
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    3 days ago
  •  ...thousands of companies. Join us as we help people all over the world thrive at work.Location: Salt Lake City, UTAs a Senior Site Reliability Engineer, you will help define the future of reliability for our world-class employee recognition platform. You'll leverage... 
    Full time
    Shift work

    O.C. Tanner

    Salt Lake City, UT
    3 days ago
  • $166k - $220k

     ...globally. We work with mission partners and operators to deploy reliable and robust capabilities on operationally-relevant fielding...  ...scalable deployment solution must be reached. As a Senior Software Engineer, you will re-imagine the infrastructure pipeline required to convert... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    3 days ago
  • $35 - $45 per hour

    DescriptionKforce has a client that is seeking a remote Site Reliability Engineer to join their team.Summary:The team consists of systems that can track lead management, job management and sales management. It is built on Salesforce but underpinned by a lot of Java/API'... 
    Remote work

    KForce

    Atlanta, GA
    3 days ago
  • $170k - $220k

    Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating... 

    Supio

    Seattle, WA
    3 days ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure... 
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    LexisNexis Risk Solutions Group

    Alpharetta, GA
    2 days ago
  • $86.6k - $144.4k

     ...platforms and using automation to solve complex security and reliability challenges?Do you enjoy shaping the future of security...  ...You can learn more about LexisNexis Risk at our TeamOur Site Reliability Engineering (SRE) team plays a critical role in ensuring the... 
    Full time
    Local area

    LexisNexis Risk Solutions Group

    Alpharetta, GA
    1 day ago
  • $146.6k - $263.6k

     ...Job Title: Senior Site Reliability Engineer Work Location: 145 Broadway, Cambridge, MA 02142 Job Description: Akamai Technologies, Inc. is hiring for the following role in Cambridge, MA (multiple openings): Senior Site Reliability Engineer Perform site reliability... 
    Full time
    Work experience placement
    Work at office
    Remote work

    Akamai

    Remote
    6 days ago
  • $96k - $163k

     ...realize their greatest potential. Title and Summary Senior Site Reliability EngineerOverview What we create today will define tomorrow....  ...Ops) team is seeking a Business Operations Site Reliability Engineer (SRE). The role of Business Operations Organization is to be... 
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    2 days ago
  • $166k - $220k

     ...through to the field. When something breaks in a deployed environment, we fix it. About the Role We're looking for a Site Reliability Engineer to join the Imaging team. This is not a product development role, and it isn't a traditional cloud-SRE role either. You... 
    Full time
    Work experience placement
    Immediate start
    Remote work
    Weekend work
    Day shift

    Anduril Industries

    Waltham, MA
    6 days ago
  • $160k - $208k

     ...through early diagnosis and longitudinal care management of chronic conditions. We are looking for an experienced Site Reliability and Infrastructure Engineer to join our engineering team. You will support Counterpart Health’s existing technology infrastructure by... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Clover Health

    Remote
    8 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO) AI/ML & Data Platforms team,, you will solve complex... 
    Work at office

    JP Morgan Chase

    Jersey City, NJ
    5 days ago
  •  ...to physicians, providing critical information about the right treatments for the right patients, at the right time.The Site Reliability Engineering team works with all departments and business units to provide dependable cloud infrastructure solutions, along with support... 
    Full time

    Tempus

    Chicago, IL
    5 days ago
  • $80k - $133k

     ...degree, Four (4) years additional experience will be needed.Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud infrastructure and enterprise systems.One(1)+ years of experience deploying and... 
    Permanent employment
    Full time
    Contract work
    Remote work
    Flexible hours

    Guidehouse

    San Antonio, TX
    5 days ago
  • $165k - $280k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most... 
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    4 days ago
  • Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate...  ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization,... 
    Work at office
    Local area

    Realtor.com

    Austin, TX
    4 days ago
  • Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service...  ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud... 
    Remote work

    Patterson-UTI

    Houston, TX
    1 day ago
  • $139k - $257.55k

    The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses... 
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    4 days ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $158.5k - $172k

     ...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and...  .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    Chicago, IL
    3 days ago
  • $125k - $150k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (RAPTOR)SpaceX is looking for a Site Reliability Engineer with a strong drive to solve challenging problems in the Raptor... 
    Permanent employment
    Temporary work

    SpaceX

    Hawthorne, CA
    3 days ago
  • $230k - $250k

    GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is... 
    Remote work

    Govcio

    Arlington, VA
    3 days ago
  • $128.6k - $184.9k

     ...global cloud platform. As a team of six engineers distributed across the US, Canada, and the...  ...with a strong focus on automation, reliability, and operational excellence. We are one...  ...Qualifications7+ years of experience in Site Reliability Engineering, DevOps, Infrastructure... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    CISCO Systems

    Richardson, TX
    2 days ago
  • $130k - $200k

    IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion... 
    Full time
    Work at office
    Immediate start

    IXL Learning

    San Mateo, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!