Site Reliability Engineer
$150k - $200kRunpod
Site Reliability Engineer
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.
The Reliability team owns the availability, performance, and operational excellence of Runpod's global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions.
This team is responsible for:
- Defining and enforcing reliability standards across engineering
- Designing incident response processes and improving recovery times
- Building observability systems and reliability tooling
- Driving SLO adoption and production readiness reviews
- Reducing operational toil through automation
The Reliability team works cross-functionally with Infrastructure, Product Engineering, and Support to ensure our systems remain stable and performant as we scale rapidly. We value proactive problem solving, automation-first thinking, and strong ownership of production systems.
As a Site Reliability Engineer on the Reliability team, you will focus on ensuring the stability and resilience of Runpod's distributed platform. You will partner with engineering teams to improve system design, strengthen observability, and prevent incidents before they happen.
This role blends software engineering with production operations. You'll work on reliability frameworks, SLO design, automation, and production hardening, reducing errors and improving performance across different services and infrastructure.
This is a high-impact role central to maintaining trust with developers running critical AI workloads on Runpod.
Your impact will include:
- Increase platform uptime and reduce incident frequency and duration
- Establish and operationalize SLIs/SLOs across services
- Improve MTTR through better tooling, automation, and runbooks
- Strengthen production readiness standards
- Drive long-term systemic reliability improvements
You will influence how reliability is defined and measured across Runpod and help build the operational backbone of the company.
Responsibilities include:
- Reliability Engineering
- Define and implement SLIs/SLOs for critical services
- Lead incident response and coordinate cross-team mitigation efforts
- Conduct blameless postmortems and ensure corrective actions are completed
- Perform production readiness reviews for new services and features
- Identify systemic risks and drive preventative improvements
- Observability & Monitoring
- Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.)
- Improve signal-to-noise ratio in alerts and reduce alert fatigue
- Build internal tooling for reliability tracking and reporting
- Improve visibility into GPU performance and distributed systems health
- Automation & Toil Reduction
- Automate recurring operational workflows
- Build tools and scripts (Python, Go, Bash) to eliminate manual processes
- Improve deployment safety through automation and guardrails
- Strengthen CI/CD reliability and release processes
- Cross-Functional Reliability Advocacy
- Partner with engineering teams to improve system resilience
- Provide guidance on fault tolerance, scalability, and failure handling
- Contribute to architectural discussions with a reliability-first mindset
Requirements include:
- 5+ years of experience in SRE, Reliability Engineering, or Production Engineering
- Strong Linux systems and Networking expertise
- Experience managing containerized production systems
- Strong understanding of distributed systems and failure modes
- Experience defining and managing SLIs/SLOs
- Proven incident response and postmortem leadership experience
- Strong scripting or programming skills
- Experience with monitoring and alerting systems
- Excellent written communication skills
- Successful completion of a background check
Preferred:
- Experience with GPU infrastructure or AI/ML platforms
- Experience improving reliability in high-growth or large scale environments
- Familiarity with GPU observability tooling
- Experience with Infrastructure as Code
- Experience working in startup environments
- Experience building internal reliability platforms or frameworks
What you'll receive:
- The competitive base pay for this position ranges from $150,000- $200,000 usd. This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate's experience, qualifications, and location
- Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside.
- Generous medical, dental & vision plans
- Flexible PTO- take the time you need to recharge
- Most roles are remote work first with an inclusive, collaborative teams utilizing Slack as the main form of internal communication
- Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.
Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.
$125k - $145k
...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER, GNCSpaceX’s mission is to make humanity multiplanetary by developing fully and rapidly reusable launch systems capable of...SuggestedPermanent employmentTemporary workFlexible hoursWeekend work$166k - $220k
...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the... ...-to-end solutions we ship.ABOUT THE JOBWe are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Costa...SuggestedFull timeWork experience placementImmediate start- Company DescriptionComtech LLC is a woman-owned small business focused on delivering end-to-end solutions and products. Since 1998, we have successfully serviced enterprises across the public and private sectors, and the Department of Defense. Our services span all aspects...Suggested
$81.1k - $187k
...architect infrastructure and service to ensure reliability and functionality. Forecasts demands and... ...impact and develops knowledge of site reliability trends.Only Oracle brings together... ...guidance and mentorship to junior engineers. Communicate status, risks, blockers, and...SuggestedTemporary workFlexible hours$138.4k - $173k
...infrastructure as well as help improve the reliability, quality of services and overall... ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the reliability... ...about our locations by visiting our site.Compensation & BenefitsThe base salary that...SuggestedFull timeFlexible hours$143k - $191k
...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental... ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and...Full timeTemporary workWork experience placementImmediate start- ...thousands of companies. Join us as we help people all over the world thrive at work.Location: Salt Lake City, UTAs a Senior Site Reliability Engineer, you will help define the future of reliability for our world-class employee recognition platform. You'll leverage...Full timeShift work
$166k - $220k
...globally. We work with mission partners and operators to deploy reliable and robust capabilities on operationally-relevant fielding... ...scalable deployment solution must be reached. As a Senior Software Engineer, you will re-imagine the infrastructure pipeline required to convert...Full timeWork experience placementImmediate start$35 - $45 per hour
DescriptionKforce has a client that is seeking a remote Site Reliability Engineer to join their team.Summary:The team consists of systems that can track lead management, job management and sales management. It is built on Salesforce but underpinned by a lot of Java/API'...Remote work$170k - $220k
Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating...$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure...Flexible hours$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...Full timeWork at officeLocal areaRemote workWork from home$86.6k - $144.4k
...platforms and using automation to solve complex security and reliability challenges?Do you enjoy shaping the future of security... ...You can learn more about LexisNexis Risk at our TeamOur Site Reliability Engineering (SRE) team plays a critical role in ensuring the...Full timeLocal area$146.6k - $263.6k
...Job Title: Senior Site Reliability Engineer Work Location: 145 Broadway, Cambridge, MA 02142 Job Description: Akamai Technologies, Inc. is hiring for the following role in Cambridge, MA (multiple openings): Senior Site Reliability Engineer Perform site reliability...Full timeWork experience placementWork at officeRemote work$96k - $163k
...realize their greatest potential. Title and Summary Senior Site Reliability EngineerOverview What we create today will define tomorrow.... ...Ops) team is seeking a Business Operations Site Reliability Engineer (SRE). The role of Business Operations Organization is to be...Full timePart timeWorldwideFlexible hours$166k - $220k
...through to the field. When something breaks in a deployed environment, we fix it. About the Role We're looking for a Site Reliability Engineer to join the Imaging team. This is not a product development role, and it isn't a traditional cloud-SRE role either. You...Full timeWork experience placementImmediate startRemote workWeekend workDay shift$160k - $208k
...through early diagnosis and longitudinal care management of chronic conditions. We are looking for an experienced Site Reliability and Infrastructure Engineer to join our engineering team. You will support Counterpart Health’s existing technology infrastructure by...Full timeWork experience placementWork at officeRemote workFlexible hours- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO) AI/ML & Data Platforms team,, you will solve complex...Work at office
- ...to physicians, providing critical information about the right treatments for the right patients, at the right time.The Site Reliability Engineering team works with all departments and business units to provide dependable cloud infrastructure solutions, along with support...Full time
$80k - $133k
...degree, Four (4) years additional experience will be needed.Minimum Four (4) years of experience in IT administration, software engineering, or platform engineering, with a focus on AWS cloud infrastructure and enterprise systems.One(1)+ years of experience deploying and...Permanent employmentFull timeContract workRemote workFlexible hours$165k - $280k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most...Permanent employmentTemporary workWorldwideWeekend work- Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate... ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization,...Work at officeLocal area
- Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service... ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud...Remote work
$139k - $257.55k
The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses...Full timeTemporary workLocal areaRemote workWorldwide$165k - $225.6k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,...Permanent employmentLocal areaWorldwideFlexible hours$158.5k - $172k
...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and... .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology...Full timeTemporary workWork at officeFlexible hours3 days per week$125k - $150k
...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (RAPTOR)SpaceX is looking for a Site Reliability Engineer with a strong drive to solve challenging problems in the Raptor...Permanent employmentTemporary work$230k - $250k
GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is...Remote work$128.6k - $184.9k
...global cloud platform. As a team of six engineers distributed across the US, Canada, and the... ...with a strong focus on automation, reliability, and operational excellence. We are one... ...Qualifications7+ years of experience in Site Reliability Engineering, DevOps, Infrastructure...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours$130k - $200k
IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion...Full timeWork at officeImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer remote United States
- lead site reliability engineer United States
- site reliability engineer United States
- site reliability engineer sre United States
- site reliability engineering manager United States
- junior website developer United States
- website content developer United States
- on site coordinator United States
- after school site coordinator United States
- website coordinator United States
