Staff Site Reliability Engineer - AI Platform Runtime
$168k - $270.25kNvidia
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices. This is a highly specialized discipline which demands knowledge across different systems, networking, coding, database, capacity management, continuous delivery and deployment, open source cloud enabling technologies like Kubernetes and Public Cloud. SRE at NVIDIA ensures that our internal and external facing services run maximum reliability and uptime as promised to the users and at the same time enabling developers to make changes to the existing system through careful preparation and planning while keeping an eye on capacity, latency and performance. SRE is also a mindset and a set of engineering approaches to running better production systems and optimizations. Much of our software development focuses on building components to eliminate manual work through automation, performance tuning and growing efficiency of production systems.As SREs are responsible for the big picture of how our systems relate to each other, we use a breadth of tools and approaches to tackle a broad spectrum of problems. Practices such as limiting time spent on reactive operational work, blameless postmortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and interesting dynamic day-to-day work. SRE's culture of diversity, intellectual curiosity, problem solving and openness is important to our success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to build an environment that provides the support and mentorship needed to learn and grow.What you’ll be doing:Lead the technical strategy and roadmap for large-scale, cross-functional SRE initiatives that improve reliability, scalability, and developer productivity across enterprise systems.Design, and build resilient distributed systems that power NVIDIA’s next-generation AI-driven enterprise products and services.Architect and develop AI Agents, AI Skills to accelerate platform operationsDrive automation and observability improvements, using metrics and analytics to enhance performance, reliability, and efficiency.Collaborate across Cloud, Platform, Security, and AI/ML teams to implement modern SRE components that ensure high availability and secure operations.Analyze and troubleshoot complex systems, championing best practices in system design, incident management, and postmortem analysis.Mentor and influence engineers across teams, fostering technical excellence and a culture of reliability engineering.What we need to see:10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics), or equivalent experienceStrong proficiency in programming languages such as Python, Typescript, JavaScript, or Go, with a focus on automation and infrastructure-as-code.Experience with infrastructure-as-code such as AWS CDK, AWS CloudFormation, Terraform or CrossPlaneSolid understanding of OpenTelemetry or other Observability implementation at scale.Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS, Azure, or GCP).Outstanding problem-solving, communication, and teamwork skills, with the ability to influence across technical and interpersonal boundaries.Ways to stand out from the crowd:Passion for and experience with Public Cloud or large-scale automation systems.Demonstrated ability to drive technical strategy and deliver measurable reliability outcomes in complex environments.A strong sense of ownership, curiosity, and innovation, you thrive in ambiguity and turn challenges into opportunities.NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables outstanding creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence.#LI-HybridYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 12, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
- ...the world's largest AI chip, 56 times... ...by the Wafer-Scale Engine (WSE). This team will... ...-class, ultra-reliable inference infrastructure... ...labs.As a Staff SRE, you will lead... ...and mentor them as platform engineers.You will... ...including model serving runtimes, GPU or wafer-...SuggestedShift work
$209.7k - $266.8k
...on in the future. Very few people in AI can say this. Every role here,... ...join a motivated and talented team of engineers to deliver a reliable, stable and flexible software stack... ...success of Wayve’s mission. The Runtime Platform team equips all Wayve teams with the...SuggestedFull timeFlexible hours$165.2k - $223.6k
...Software Development Engineer for the Neuron Runtime Team, you will be... ...learning applications and AI accelerators. You... ...across hardware platforms such as Trainium and... ...ensuring scalability, reliability, and usability. You... ...employees, supervisors, and staff; adhere to standards...SuggestedInternshipLocal areaWork from homeFlexible hours$150k - $180k
Santa Clara, CAUS Research and Development - Runtime Platform /Full-time /HybridPlusAI is a Physical AI company pioneering AI-based virtual driver software for factory-built autonomous trucks. Headquartered in Silicon Valley with operations in the United States and Europe...SuggestedFull time$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building... ...system performance, build scalable platforms, and continuously strengthen the... ...technical direction of NVIDIA’s AI Platform Runtime and lead reliability engineering initiatives...SuggestedFull time$120k - $200k
...Research and Development - Runtime Platform /Full-time /HybridYou will develop... ...that facilitates reliable, low-latency execution of on... ...driving — you will also equip engineers with the tools needed to analyze... ...use artificial intelligence (AI) tools to support parts of...Full time$130k - $220k
Santa Clara, CASoftware Engineering - Motion Planning /... ...HybridPlusAI is a Physical AI company pioneering AI-... ...the production runtime pipeline for ML-based... ...-constrained embedded platforms.Develop high-performance... ...while maintaining system reliability and performance.Drive...Full time$143.4k - $165.6k
...and existing systems, including patterns for reliability and scaling. We need hands-on programming experience... ...: We develop and maintain high-performance runtime libraries and drivers for machine learning applications and AI accelerators. We lead the design,...Full timeInternshipFlexible hours$267k - $356k
...Cloud, is a leader in AI cloud infrastructure serving... ....Lambda's Storage Engineering team is the backbone... ...spectrum of Lambda's data platform services—from low-... ...industry, which means reliability and performance aren't... ...across new and existing sites using tools such as...Work experience placementWork at officeLocal areaWork from homeFlexible hours$152k - $241.5k
...generation of our global services platform. At NVIDIA, you’ll keep... ...You’ll harness the power of AI to deliver groundbreaking solutions... ...lifecycle management, fleet reliability/auto-healing, E2E... ...Perl, or Ruby.Mentored other engineers and influenced technical direction...Full time$168k - $270.25k
...artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing,... ..., 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering...Full time- ...Cloud, is a leader in AI cloud infrastructure serving... ...is currently Tuesday.Engineering at Lambda is... ...tenant cloud networking platform and SDN infrastructureOperate... ...to improve service reliability and deployment... ...years of experience in Site Reliability Engineering...Work at officeLocal areaWork from homeFlexible hours
- ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving... ...day is currently Tuesday.Engineering at Lambda is responsible for... ...automate the validation of platform quality.Design, build, and... ...services, workloads, and platform reliability.You6+ years of experience in...Work at officeLocal areaWork from homeFlexible hours
$230k - $250k
...foundation for autonomous networking, giving engineers and AI agents the ability to know the impact... ..., building a groundbreaking platform that transforms how teams run and secure... ...been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a...Night shift$148k - $235.75k
...tapping into the unlimited potential of AI to define the next era of computing.... ...the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for...Full time- ...responsibilities of a Technical Support Engineer within a SaaS (Software as a... ...environment with a growing focus on Site Reliability Engineering (SRE).The ideal candidate... ...to run, support, and scale an AI Security Public SaaS platform, operating AI inference workloads at...Full timeLocal area
$145k - $175k
...care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability... ...our mission-critical healthcare data platform. You will bridge the gap between... ...workflows, streaming systems). Exposure to AI/ML infrastructure and the reliability...Full timeRemote work$207.4k - $259.2k
...an end-to-end advanced air mobility platform that delivers air taxis, unmanned aircraft... ...physical artificial intelligence (“AI”) solutions, and other technologies... ...highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In...Permanent employmentLocal areaWorldwideVisa sponsorship$122.5k - $175k
...efficient, resilient, and secure. As an AI-forward enterprise, we are constantly... ...our cloud-native Zero Trust Exchange platform. This innovation protects our... ...cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role...Full timeWork at officeLocal area3 days per week$132.6k - $214.5k
...and Inclusion. We weave AI into the fabric of... ...most advanced SecOps platform, consisting of XDR, XSIAM... ...collaborate closely with our engineering teams to develop... .... As a Senior Staff SRE with the Cortex Observability... ...and ensure the reliability and availability of our...Full timeWork at officeVisa sponsorshipWork visa$192.4k - $275.8k
...demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service... ...automation and frameworks that make the whole platform more resilient. If you are the kind of... ...and protect organizations in the AI era - and beyond. We’ve been...Full timeTemporary workLocal areaFlexible hours- ...Powered by the Illumio AI Security Graph, our breach containment platform identifies and contains... ...running. Location: 5 on-site days a week in... ...Team's Vision: Our Engineering team is shaping the future... ...experienced Senior Site Reliability Engineer (SRE) with a strong...Work experience placementImmediate start
$272k - $431.25k
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning... ..., and Android. It supports hardware platforms including NVIDIA GPUs and Tegra... ...optimize the speed and cost efficiency of AI development and testing systems....Full timeWork experience placementWorldwide$272k - $431.25k
...Principal System Software Engineer to drive next-... ...innovations in automotive platform software, system architecture... ...architecture, kernel, AI, middleware, and... ...drivers, middleware, runtime frameworks, and platform... ...improve performance, reliability, determinism, and...Full time- ...Integrity, and Inclusion. We weave AI into the fabric of everything we do... ...Lead, mentor, and develop a team of Site Reliability/Production Engineers, providing technical direction, coaching... ...Strong technical knowledge of cloud platforms, preferably Google Cloud Platform (...Full timeWork at officeVisa sponsorshipWork visa
$119k - $170k
...secure. The Zscaler Zero Trust Exchange️ platform protects thousands of customers from... ...the future of work is Human + AI and are building an AI-native enterprise... ...at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team...Full timeWork at officeLocal areaRemote workShift work3 days per week$207k - $300k
...team of Software/Systems Engineers on projects for users... ...and responsibly applying AI tooling and workflows... ...establishing sustainable multi-site on-call rotations... ...expertise in Site Reliability Engineering practices,... ...next-generation of Google platforms, we make Google's product...$160k - $225k
...the world's best data and AI infrastructure platform so our customers can use deep... ...' Mosaic AI mission. AI Runtime (AIR) is our managed platform... ...As a Senior Software Engineer for AI Runtime, you will play... ...large-scale training fast, reliable, and effortless. You will drive...Full timeLocal areaWorldwide$179.2k - $268.8k
...Latitude AI (lat.ai) is building the future of... ...developed hands-free ADAS platform will debut on the all-... ..., systems and safety engineering – all dedicated to... ...to come and join the Runtime Services team at Latitude... ...develop, and test the reliable and high-performance...Permanent employmentFull timeWork at officeImmediate startVisa sponsorship$224k - $356.5k
...into the unlimited potential of AI to define the next era of... ...best experience. We own the platform — performance, CI/CD pipelines... ...in Computer Science, Computer Engineering, Electrical Engineering, or equivalent... ...builds, layer optimization, runtime configuration, NVIDIA...Full timeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer - AI Platform Runtime. Be the first to apply!
- senior staff engineer Santa Clara, CA
- senior staff systems engineer Santa Clara, CA
- engineering aide Santa Clara, CA
- software engineer staff Santa Clara, CA
- assistant engineer Santa Clara, CA
- technology administrator Santa Clara, CA
- staff engineer Santa Clara, CA
- site reliability engineer sre Santa Clara, CA
- site reliability engineer Santa Clara, CA
- platform engineer Santa Clara, CA


