Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Site Reliability Engineer - AI Platform Runtime

$168k - $270.25k

Nvidia

Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices. This is a highly specialized discipline which demands knowledge across different systems, networking, coding, database, capacity management, continuous delivery and deployment, open source cloud enabling technologies like Kubernetes and Public Cloud. SRE at NVIDIA ensures that our internal and external facing services run maximum reliability and uptime as promised to the users and at the same time enabling developers to make changes to the existing system through careful preparation and planning while keeping an eye on capacity, latency and performance. SRE is also a mindset and a set of engineering approaches to running better production systems and optimizations. Much of our software development focuses on building components to eliminate manual work through automation, performance tuning and growing efficiency of production systems.As SREs are responsible for the big picture of how our systems relate to each other, we use a breadth of tools and approaches to tackle a broad spectrum of problems. Practices such as limiting time spent on reactive operational work, blameless postmortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and interesting dynamic day-to-day work. SRE's culture of diversity, intellectual curiosity, problem solving and openness is important to our success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to build an environment that provides the support and mentorship needed to learn and grow.What you’ll be doing:Lead the technical strategy and roadmap for large-scale, cross-functional SRE initiatives that improve reliability, scalability, and developer productivity across enterprise systems.Design, and build resilient distributed systems that power NVIDIA’s next-generation AI-driven enterprise products and services.Architect and develop AI Agents, AI Skills to accelerate platform operationsDrive automation and observability improvements, using metrics and analytics to enhance performance, reliability, and efficiency.Collaborate across Cloud, Platform, Security, and AI/ML teams to implement modern SRE components that ensure high availability and secure operations.Analyze and troubleshoot complex systems, championing best practices in system design, incident management, and postmortem analysis.Mentor and influence engineers across teams, fostering technical excellence and a culture of reliability engineering.What we need to see:10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics), or equivalent experienceStrong proficiency in programming languages such as Python, Typescript, JavaScript, or Go, with a focus on automation and infrastructure-as-code.Experience with infrastructure-as-code such as AWS CDK, AWS CloudFormation, Terraform or CrossPlaneSolid understanding of OpenTelemetry or other Observability implementation at scale.Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS, Azure, or GCP).Outstanding problem-solving, communication, and teamwork skills, with the ability to influence across technical and interpersonal boundaries.Ways to stand out from the crowd:Passion for and experience with Public Cloud or large-scale automation systems.Demonstrated ability to drive technical strategy and deliver measurable reliability outcomes in complex environments.A strong sense of ownership, curiosity, and innovation, you thrive in ambiguity and turn challenges into opportunities.NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables outstanding creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence.#LI-HybridYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 12, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 3 hours ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer - AI Platform Runtime in Santa Clara, CA vacancy
  •  ...the world's largest AI chip, 56 times...  ...by the Wafer-Scale Engine (WSE). This team will...  ...-class, ultra-reliable inference infrastructure...  ...labs.As a Staff SRE, you will lead...  ...and mentor them as platform engineers.You will...  ...including model serving runtimes, GPU or wafer-... 
    Suggested
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $209.7k - $266.8k

     ...on in the future. Very few people in AI can say this. Every role here,...  ...join a motivated and talented team of engineers to deliver a reliable, stable and flexible software stack...  ...success of Wayve’s mission. The Runtime Platform team equips all Wayve teams with the... 
    Suggested
    Full time
    Flexible hours

    Wayve

    Sunnyvale, CA
    1 day ago
  • $165.2k - $223.6k

     ...Software Development Engineer for the Neuron Runtime Team, you will be...  ...learning applications and AI accelerators. You...  ...across hardware platforms such as Trainium and...  ...ensuring scalability, reliability, and usability. You...  ...employees, supervisors, and staff; adhere to standards... 
    Suggested
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $150k - $180k

    Santa Clara, CAUS Research and Development - Runtime Platform /Full-time /HybridPlusAI is a Physical AI company pioneering AI-based virtual driver software for factory-built autonomous trucks. Headquartered in Silicon Valley with operations in the United States and Europe... 
    Suggested
    Full time

    Plus.ai

    Santa Clara, CA
    1 day ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building...  ...system performance, build scalable platforms, and continuously strengthen the...  ...technical direction of NVIDIA’s AI Platform Runtime and lead reliability engineering initiatives... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 hours ago
  • $120k - $200k

     ...Research and Development - Runtime Platform /Full-time /HybridYou will develop...  ...that facilitates reliable, low-latency execution of on...  ...driving — you will also equip engineers with the tools needed to analyze...  ...use artificial intelligence (AI) tools to support parts of... 
    Full time

    Plus.ai

    Santa Clara, CA
    1 day ago
  • $130k - $220k

    Santa Clara, CASoftware Engineering - Motion Planning /...  ...HybridPlusAI is a Physical AI company pioneering AI-...  ...the production runtime pipeline for ML-based...  ...-constrained embedded platforms.Develop high-performance...  ...while maintaining system reliability and performance.Drive... 
    Full time

    Plus.ai

    Santa Clara, CA
    2 days ago
  • $143.4k - $165.6k

     ...and existing systems, including patterns for reliability and scaling. We need hands-on programming experience...  ...: We develop and maintain high-performance runtime libraries and drivers for machine learning applications and AI accelerators. We lead the design,... 
    Full time
    Internship
    Flexible hours

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    4 days ago
  • $267k - $356k

     ...Cloud, is a leader in AI cloud infrastructure serving...  ....Lambda's Storage Engineering team is the backbone...  ...spectrum of Lambda's data platform services—from low-...  ...industry, which means reliability and performance aren't...  ...across new and existing sites using tools such as... 
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $152k - $241.5k

     ...generation of our global services platform. At NVIDIA, you’ll keep...  ...You’ll harness the power of AI to deliver groundbreaking solutions...  ...lifecycle management, fleet reliability/auto-healing, E2E...  ...Perl, or Ruby.Mentored other engineers and influenced technical direction... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $168k - $270.25k

     ...artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing,...  ..., 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Cloud, is a leader in AI cloud infrastructure serving...  ...is currently Tuesday.Engineering at Lambda is...  ...tenant cloud networking platform and SDN infrastructureOperate...  ...to improve service reliability and deployment...  ...years of experience in Site Reliability Engineering... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving...  ...day is currently Tuesday.Engineering at Lambda is responsible for...  ...automate the validation of platform quality.Design, build, and...  ...services, workloads, and platform reliability.You6+ years of experience in... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    14 hours ago
  • $230k - $250k

     ...foundation for autonomous networking, giving engineers and AI agents the ability to know the impact...  ..., building a groundbreaking platform that transforms how teams run and secure...  ...been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a... 
    Night shift

    Forward Networks

    Santa Clara, CA
    1 day ago
  • $148k - $235.75k

     ...tapping into the unlimited potential of AI to define the next era of computing....  ...the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...responsibilities of a Technical Support Engineer within a SaaS (Software as a...  ...environment with a growing focus on Site Reliability Engineering (SRE).The ideal candidate...  ...to run, support, and scale an AI Security Public SaaS platform, operating AI inference workloads at... 
    Full time
    Local area

    F5 Networks

    San Jose, CA
    2 days ago
  • $145k - $175k

     ...care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability...  ...our mission-critical healthcare data platform. You will bridge the gap between...  ...workflows, streaming systems). Exposure to AI/ML infrastructure and the reliability... 
    Full time
    Remote work

    GrabJobs

    Santa Clara, CA
    3 days ago
  • $207.4k - $259.2k

     ...an end-to-end advanced air mobility platform that delivers air taxis, unmanned aircraft...  ...physical artificial intelligence (“AI”) solutions, and other technologies...  ...highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In... 
    Permanent employment
    Local area
    Worldwide
    Visa sponsorship

    Archer Aviation

    San Jose, CA
    4 days ago
  • $122.5k - $175k

     ...efficient, resilient, and secure. As an AI-forward enterprise, we are constantly...  ...our cloud-native Zero Trust Exchange platform. This innovation protects our...  ...cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    2 days ago
  • $132.6k - $214.5k

     ...and Inclusion. We weave AI into the fabric of...  ...most advanced SecOps platform, consisting of XDR, XSIAM...  ...collaborate closely with our engineering teams to develop...  .... As a Senior Staff SRE with the Cortex Observability...  ...and ensure the reliability and availability of our... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    4 days ago
  • $192.4k - $275.8k

     ...demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service...  ...automation and frameworks that make the whole platform more resilient. If you are the kind of...  ...and protect organizations in the AI era - and beyond. We’ve been... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    2 days ago
  •  ...Powered by the Illumio AI Security Graph, our breach containment platform identifies and contains...  ...running. Location: 5 on-site days a week in...  ...Team's Vision: Our Engineering team is shaping the future...  ...experienced Senior Site Reliability Engineer (SRE) with a strong... 
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    2 days ago
  • $272k - $431.25k

    NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning...  ..., and Android. It supports hardware platforms including NVIDIA GPUs and Tegra...  ...optimize the speed and cost efficiency of AI development and testing systems.... 
    Full time
    Work experience placement
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...Principal System Software Engineer to drive next-...  ...innovations in automotive platform software, system architecture...  ...architecture, kernel, AI, middleware, and...  ...drivers, middleware, runtime frameworks, and platform...  ...improve performance, reliability, determinism, and... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Integrity, and Inclusion. We weave AI into the fabric of everything we do...  ...Lead, mentor, and develop a team of Site Reliability/Production Engineers, providing technical direction, coaching...  ...Strong technical knowledge of cloud platforms, preferably Google Cloud Platform (... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    1 day ago
  • $119k - $170k

     ...secure. The Zscaler Zero Trust Exchange️ platform protects thousands of customers from...  ...the future of work is Human + AI and are building an AI-native enterprise...  ...at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team... 
    Full time
    Work at office
    Local area
    Remote work
    Shift work
    3 days per week

    Zscaler

    San Jose, CA
    3 hours ago
  • $207k - $300k

     ...team of Software/Systems Engineers on projects for users...  ...and responsibly applying AI tooling and workflows...  ...establishing sustainable multi-site on-call rotations...  ...expertise in Site Reliability Engineering practices,...  ...next-generation of Google platforms, we make Google's product... 

    Google

    San Jose, CA
    2 days ago
  • $160k - $225k

     ...the world's best data and AI infrastructure platform so our customers can use deep...  ...' Mosaic AI mission. AI Runtime (AIR) is our managed platform...  ...As a Senior Software Engineer for AI Runtime, you will play...  ...large-scale training fast, reliable, and effortless. You will drive... 
    Full time
    Local area
    Worldwide

    Databricks

    Mountain View, CA
    3 days ago
  • $179.2k - $268.8k

     ...Latitude AI (lat.ai) is building the future of...  ...developed hands-free ADAS platform will debut on the all-...  ..., systems and safety engineering – all dedicated to...  ...to come and join the Runtime Services team at Latitude...  ...develop, and test the reliable and high-performance... 
    Permanent employment
    Full time
    Work at office
    Immediate start
    Visa sponsorship

    Latitude AI

    Palo Alto, CA
    1 day ago
  • $224k - $356.5k

     ...into the unlimited potential of AI to define the next era of...  ...best experience. We own the platform — performance, CI/CD pipelines...  ...in Computer Science, Computer Engineering, Electrical Engineering, or equivalent...  ...builds, layer optimization, runtime configuration, NVIDIA... 
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer - AI Platform Runtime. Be the first to apply!