Site Reliability Engineer
$150k - $200kGrabJobs
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on. We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale. Learn more in our CEO's funding announcement: . The Reliability team owns the availability, performance, and operational excellence of Runpod’s global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions. This team is responsible for: Defining and enforcing reliability standards across engineering Designing incident response processes and improving recovery times Building observability systems and reliability tooling Driving SLO adoption and production readiness reviews Reducing operational toil through automation The Reliability team works cross-functionally with Infrastructure, Product Engineering, and Support to ensure our systems remain stable and performant as we scale rapidly. We value proactive problem solving, automation-first thinking, and strong ownership of production systems. As a Site Reliability Engineer on the Reliability team, you will focus on ensuring the stability and resilience of Runpod’s distributed platform. You will partner with engineering teams to improve system design, strengthen observability, and prevent incidents before they happen. This role blends software engineering with production operations. You’ll work on reliability frameworks, SLO design, automation, and production hardening, reducing errors and improving performance across different services and infrastructure. This is a high-impact role central to maintaining trust with developers running critical AI workloads on Runpod. Your Impact Increase platform uptime and reduce incident frequency and duration Establish and operationalize SLIs/SLOs across services Improve MTTR through better tooling, automation, and runbooks Strengthen production readiness standards Drive long-term systemic reliability improvements You will influence how reliability is defined and measured across Runpod and help build the operational backbone of the company. Responsibilities: Reliability Engineering Define and implement SLIs/SLOs for critical services Lead incident response and coordinate cross-team mitigation efforts Conduct blameless postmortems and ensure corrective actions are completed Perform production readiness reviews for new services and features Identify systemic risks and drive preventative improvements Observability & Monitoring Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.) Improve signal-to-noise ratio in alerts and reduce alert fatigue Build internal tooling for reliability tracking and reporting Improve visibility into GPU performance and distributed systems health Automation & Toil Reduction Automate recurring operational workflows Build tools and scripts (Python, Go, Bash) to eliminate manual processes Improve deployment safety through automation and guardrails Strengthen CI/CD reliability and release processes Cross-Functional Reliability Advocacy Partner with engineering teams to improve system resilience Provide guidance on fault tolerance, scalability, and failure handling Contribute to architectural discussions with a reliability-first mindset Requirements: 5+ years of experience in SRE, Reliability Engineering, or Production Engineering Strong Linux systems and Networking expertise Experience managing containerized production systems Strong understanding of distributed systems and failure modes Experience defining and managing SLIs/SLOs Proven incident response and postmortem leadership experience Strong scripting or programming skills Experience with monitoring and alerting systems Excellent written communication skills Successful completion of a background check Preferred: Experience with GPU infrastructure or AI/ML platforms Experience improving reliability in high-growth or large scale environments Familiarity with GPU observability tooling Experience with Infrastructure as Code Experience working in startup environments Experience building internal reliability platforms or frameworks What You’ll Receive: The competitive base pay for this position ranges from $150,000- $200,000 usd. This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside. Generous medical, dental & vision plans Flexible PTO- take the time you need to recharge Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale. Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.
$81.1k - $187k
...architect infrastructure and service to ensure reliability and functionality. Forecasts demands and... ...impact and develops knowledge of site reliability trends.Only Oracle brings together... ...guidance and mentorship to junior engineers. Communicate status, risks, blockers, and...SuggestedTemporary workFlexible hours$81.1k - $187k
...architect infrastructure and service to ensure reliability and functionality. Forecasts demands and... ...impact and develops knowledge of site reliability trends.Only Oracle brings together... ...Science, Information Technology, Engineering, or a related field, or equivalent practical...SuggestedTemporary workFlexible hours- ...scale, we invite you to bring your talents to Zscaler to help shape the future of cybersecurity. Role We are looking for a Site Reliability Engineer-SkillBridge Intern (Virginia) to join our Zero Trust Exchange team. This is an onsite role based in Crystal City, Virginia...SuggestedInternshipWork at officeLocal areaRemote workNight shift
$84.9k - $209.5k
This role combines strategic architecture with practical systems engineering, deployment, automation, patching, troubleshooting, incident response, and compliance support. The Principal Site Reliability Engineer will work across Windows, Linux, Oracle Cloud Infrastructure...SuggestedTemporary workFlexible hours$96.3k - $264.1k
...infrastructure and service, ensuring alignment with reliability and functionality standards. Takes full... ...tools and provides expertise in site reliability trends.Only Oracle brings... ...LeadershipDefine and drive the site reliability engineering strategy for large-scale, distributed,...SuggestedTemporary workFlexible hours- ...meaningful products that make a real impact on children's education and literacy. About the Role We're looking for a Senior Site Reliability Engineer to drive the stability, observability, and reliability of Epic's platform as we grow. You are an experienced engineer who...Remote work
- ...scale, we invite you to bring your talents to Zscaler to help shape the future of cybersecurity. Role We are looking for a Site Reliability Engineer-SkillBridge Intern (San JosA Ca or Bellevue WA) to join our Zero Trust Exchange team. This is a remote role based in San...InternshipWork at officeLocal areaRemote workWorldwide
$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity...Work at officeLocal areaVisa sponsorshipFlexible hours3 days per week$160.8k - $214.1k
...observability needs of modern infrastructure. The Customer Reliability Engineering team is the deep technical escalation tier for Cisco Hypershield... .../fix and reliability cases escalated by Cisco TAC, applying Site Reliability Engineering practices across the full stack: the...Full timeTemporary workLocal areaRemote workFlexible hours- Nashville, TennesseeOperations - Engineering /Full-Time /On-siteHeadquartered in Nashville,... ...global health, come grow with August!The Reliability Engineer is the system owner for equipment... ...with GDP, data-integrity, and site procedures. Support development and revision...Permanent employmentFull timeContract workFor contractorsWork at officeLocal area
$122k - $127k
...seasonal items at everyday low prices in convenient neighborhood locations. Learn more about Dollar General at .A Senior Software Engineer, working independently or with limited supervision, translates high-level business requirements into technical designs, proposes design...Work experience placementSeasonal work$180k - $210k
...Job Description Job Description Staff Site Reliability Engineer, Platforms Location: Nashville, TN (Hybrid, 3 days in office) Compensation: $180,000 - $210,000 base salary Eligibility: U.S. residents only About Us At 360 Privacy, we build proactive...Work at officeLocal area$152k - $157k
...How would you like to Serve? Join the Dollar General Journey and see how your career can thrive. Company Overview A Lead Software Engineer (LSE) is recognized as an expert in their strategic functional area and applies advanced technical expertise and a structured approach...Work experience placementSeasonal work$126.2k - $264.1k
>>This position will be full-time on-site at Oracle's offices located in Nashville, TN.<< Relocation assistance may be available... ...with Oracle’s relocation policies.As Senior Manager - Reliability Engineering, you will lead the teams, methods, and programs responsible...Full timeTemporary workRelocationRelocation packageFlexible hours$94.9k - $135.6k
...development, testing, operations, and platform teams to deliver value safely and efficiently. Cardinal Health is seeking a Release Engineer to lead iteration and release management activities supporting mission critical warehouse transformation initiatives on Program...Temporary workLocal areaImmediate startFlexible hours$125k - $191.7k
...Job Description Hybrid: This role is categorized as hybrid/Remote Role: As a Senior Software Systems Engineer on the Software Validation team within the AV organization, you will play a critical role in leading the strategy and execution of validation efforts...Local areaRemote workWork from homeFlexible hours- ...Technology team is seeking a Nashville, TN based AI Solutions Engineer to build and deliver AI‑powered tools that integrate directly with... ...prompt strategies, orchestration logic, and guardrails for reliable AI behavior.Develop agent‑style workflows that combine reasoning...Full time
$96k - $152k
...AI, Microsoft Power Platform, Databricks, and modern software engineering to deliver intelligent applications, automation, and data-driven... ...Data Operations and business teams to ensure solutions are reliable, secure, and aligned with organizational goals.Collaborate with...Full timeWork at officeImmediate startWork visaRelocation package3 days per week- DescriptionWe are seeking an innovative RCM AI Solutions Engineer to lead the development and implementation of AI-powered solutions that transform revenue cycle operations. This individual will identify automation opportunities, design intelligent workflows, and deploy...
$114.6k - $234.6k
...creates durable fixes and preventive controls. Designs performance, reliability, and fault-tolerance improvements for drivers, services, and... ...development lifecycle; provides guidance and coaching to engineers to drive improvements. Utilizes advanced knowledge to develop...Temporary workFlexible hoursShift work$92.5k - $209.5k
...production support for platform services.· Improve workflow reliability, deployment safety, scalability, observability, and developer... ...Break down complex platform problems into clear, maintainable engineering solutions.· Contribute to high standards for code quality, system...Temporary workFlexible hours$92.5k - $209.5k
...autonomy and support to do your best work. It is a dynamic and flexible workplace where you’ll belong and be encouraged.As Software Engineer, you will work with a team of software engineers responsible for the software design, development, and operations for our new and...Temporary workFlexible hours$92.5k - $209.5k
...improvements, ensure automation, testing, and debugging of systems to ensure service/product availability, health, support, and reliability.Core ResponsibilitiesPlanning & Execution:Independently manages work, monitoring timelines and deliverables to ensure projects or...Temporary workImmediate startFlexible hoursShift work- ...L3Harris Technologies in Nashville is seeking a Lead Cybersecurity Engineer to own cybersecurity implementation and validation across the program and product lines. You will translate contractual, mission, and regulatory requirements into secure architectures and verifiable...
$208.2k - $264.4k
...unrelenting focus on our customers' success, we are Cisco's growth engine and shape the company’s future. Our values of Customer-Driven... ...coverage, and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...Full timeTemporary workLocal areaRelocationFlexible hours$122.5k - $423.78k
...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelDirectorJob Description & SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and techniques to design and develop robust data solutions for clients. They play a...Full timeTemporary workH1bRemote work$165k - $216.56k
...the future of how work gets done.We are looking for a Solution Engineer who is accustomed to solving customer’s most complex problems and... ...United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake....Remote work$100k - $150k
Software EngineerNashville, TNA company operating in a Microsoft technology environment is seeking a Software Engineer to design, develop, and support enterprise web and mobile applications. You will contribute to modernization initiatives, including migrating legacy applications...Full time- ...strategic success. Job Summary:BLR is seeking a remote Senior Software Engineer (PHP) to help maintain, improve and extend our established... ...with product and engineering teammates to keep the platform reliable and continuously improving. This is a hands-on role for an engineer...Immediate startRemote work
$89.2k - $209.5k
...simplify and accelerate cloud adoption by delivering scalable, reliable, and secure migration solutions that reduce complexity and... ...workloads efficiently and confidently. You will collaborate with engineers across OCI to solve complex technical challenges while helping...Temporary workWork experience placementRelocation packageFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- on-site clinical research associate (traveling/remote) Nashville, TN
- website coordinator Nashville, TN
- junior website developer Nashville, TN
- site leader Nashville, TN
- historic site Nashville, TN
- website content developer Nashville, TN
- construction site safety Nashville, TN
- official site Nashville, TN
- site services specialist Nashville, TN
- on site coordinator Nashville, TN

