Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer, DGX Cloud

$168k - $270.25k

NVIDIA

NVIDIA DGX Cloud is developing and managing large-scale GPU infrastructure for AI research and production workloads. We are looking for Senior Reliability Engineers to help build the automation, tooling, and operational systems that make GPU clusters reliable, scalable, and safe to run. This role is part of a production engineering team passionate about Kubernetes-based infrastructure, GPU cluster operations, reliability, automation, GitOps, and Day 2 operability across DGX Cloud environments.What you’ll be doing:Build and operate automation for large-scale Kubernetes clusters across NVIDIA Cloud Partners (NCP) and on-prem environments.Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations.Improve Day 0 / Day 1 / Day 2 workflows for cluster bringup, handoff, and production operations.Define SLOs/SLIs, monitor error allowances, and streamline reportingReduce manual production touches through APIs, GitOps, automation, and agent-assisted workflows.Participate in on-call, incident response, debugging, and durable follow-up work.Partner with platform, storage, networking, security, and workload teams to make infrastructure production-ready.What we need to see:8+ years of experience building or operating production infrastructure.Strong programming skills in Python, Go, or similar.Expert-level knowledge with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation.Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management.Ability to troubleshoot distributed systems in production.Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.Clear communication and ability to work across teams.BS/MS in Computer Science or equivalent experience.Ways to stand out from the crowd:Experience with GPU infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, or fleet automation.Experience with SLOs, on-call, incident response, observability, and reliability practices.Experience operating and resolving problems in production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis.NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hard-working people on the planet working for us. If you're creative, hard-working and self-motivated, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 22, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Remote; US, CA, Remote; US, WY, RemoteType: Full time

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer, DGX Cloud in Santa Clara, CA vacancy
  • $168k - $270.25k

    NVIDIA is driving AI and high-performance computing forward. DGX Cloud aims to deliver a fully managed AI platform on major...  ...infrastructure. Work with NVIDIA's DGX Cloud team as a Senior Site Reliability Engineer to maintain high-performance DGX Cloud clusters for AI researchers... 
    Senior
    Full time
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  • $200k - $322k

    The DGX Cloud organization bridges customer success and cloud infrastructure engineering, partnering directly with NVIDIA's internal research and product teams to accelerate AI workload development. As a Customer Success Engineer, you'll embed deeply with internal customers... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    22 days ago
  • $140k - $224.25k

    NVIDIA DGX Cloud provides the infrastructure and software platform that enables enterprises...  ...management is critical to delivering reliable customer experiences while maximizing...  ...infrastructure.We are looking for a Senior Software Engineer to design and build the systems that... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    7 days ago
  • $152k - $241.5k

    As part of DGX Cloud, the Attestation & Trust Services team builds...  ...traffic scale. We want hands-on engineers with strong systems...  ...production support.Improve reliability and security in the components...  ...targeted experiments, working with senior engineers on problems that... 
    Senior
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...lasting impact on the world. The DGX Cloud organization at NVIDIA brings...  ...a group of forward‑thinking engineers tackling some of the globe’s...  ...lives. We’re searching for a Senior Systems Software Engineer...  ...and related projects to enable reliable operation at hyperscale cluster... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    a month ago
  • $136k - $224.25k

     ...NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA, from Graphics Drivers to Autonomous Vehicles and Artificial Intelligence... 
    Senior
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    a month ago
  • $184k - $287.5k

    NVIDIA is seeking a Senior Software Engineer to help us develop distributed storage services for AI/ML...  ...NVIDIA team to design and build a reliable, scalable, and efficient storage-as-a-...  ...teams, and external customers to deliver Cloud services.Automating distributed storage... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    7 days ago
  • $168k - $264.5k

     ...NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure. The goal is to craft a reliable, scalable and efficient network to support NVIDIA software development workflows and tools, including CI/CD pipelines, compute resource management... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $184k - $287.5k

     ...We are looking for a Senior Software Engineer to become part of our storage management plane team. The management plane is a web-based application crafted to provide our storage customers the capabilities to handle and supervise our distributed storage infrastructure.... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $224k - $356.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means advancing the performance, efficiency...  ...software stacks. We are seeking a Senior Performance Engineer to characterize workloads, establish...  ...raise the performance and reliability of AI workloads. Join our technically... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation... 
    Senior

    Oracle

    Santa Clara, CA
    3 days ago
  •  ...management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence...  ...the AIOps substrate The remediation-actuator and workflow engine land here — you make the control plane safe for automated action... 
    Senior
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    2 days ago
  • $145k - $175k

     ...Commence cuts straight to better care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability,...  ...unique to model serving. Familiarity with additional cloud platforms (Azure, Google Cloud). Contributions to open... 
    Senior
    Full time
    Remote work

    GrabJobs

    Santa Clara, CA
    7 hours ago
  • $187.04k - $359.72k

     ...systems by pushing for changes that improve reliability and velocity. Qualifications Minimum...  ...degree in Computer Science, Electrical Engineering, Computer Engineering or related areas....  ...Product Ops, Corporate Functions and more. On-site presence across teams allows the company... 
    Senior
    Temporary work
    Local area
    Overseas
    Shift work

    Tik Tok

    San Jose, CA
    3 days ago
  • $184k - $287.5k

    Joining NVIDIA's DGX Cloud Lepton Team means contributing to the leading...  ...AI infrastructure software engineer to join our team. You'll be...  ...Agentic AI in production.As a senior DGX Cloud AI Infrastructure...  ...Define meaningful and actionable reliability metrics to track and improve... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $176k - $276k

    Production engineering is a field that involves crafting, building, and maintaining large-scale...  ...and deployment, along with open-source cloud-enabling technologies such as Kubernetes...  ...ensuring storage architectures are reliable, scalable, and efficient. They optimize... 
    Senior
    Full time
    Flexible hours

    Nvidia

    Santa Clara, CA
    a month ago
  • $272k - $431.25k

     ...Join our dynamic team at NVIDIA as a Principal Software Engineer in Networking within the DGX Cloud division. Embrace this outstanding opportunity to...  ...attention to detail.Performing code reviews to ensure reliable and correct feature implementation.Collaborating with... 
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    11 days ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud...  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...networking teams to improve service reliability and deployment...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    a month ago
  • $168k - $258.75k

    NVIDIA's DGX Cloud (DGXC) powers AI for strategic research and product...  .... The company seeks a Senior Technical Program Manager (TPM...  ...focus is on enabling scalable, reliable, and supportable software for...  ...responsible for managing high-impact engineering programs within a dynamic,... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    7 days ago
  •  ...world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is designed for a... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    a month ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens...  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building...  ...Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    22 days ago
  • $168k - $270.25k

     ...of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing...  ...(HPC) storage solutions while harnessing the power of cloud computing. You will be responsible for crafting and deploying... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    25 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $160k - $240k

     ...times a day - quickly, reliably, and securely. Any time...  ...Fiserv.Job TitleSenior Site Reliability EngineerWhat...  ...Site Reliability Engineer do at Fiserv?You will join...  ...improvement across our cloud-native environments.What...  ...or DevOps at a mid-to-senior level.Strong shell scripting... 
    Senior
    Full time

    Fiserv

    Sunnyvale, CA
    18 days ago
  • $267k - $356k

    Lambda, The Superintelligence Cloud, is a leader in AI cloud...  ...currently Tuesday.Lambda's Storage Engineering team is the backbone behind...  ...in the industry, which means reliability and performance aren't just goals...  ...across new and existing sites using tools such as Ansible,... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    25 days ago
  • $184k - $287.5k

    At NVIDIA, the DGX Cloud division merges fresh hardware and software...  .... Our team of skilled engineers is committed to addressing major...  ...world!We are looking for a Senior Systems Software Engineer with...  ...needed to maintain cluster reliability at frontier AI scale. In this... 
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...NVIDIA DGX Cloud is scaling GPU infrastructure across internal, partner...  ...for Principal Software Engineers to help shape the technical direction...  ...operations, automation, and reliability across large-scale GPU clusters.This role is for senior technical leaders who can define... 
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $262k - $364k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible...  ...or Engineering, or a related field.Site Reliability Engineering (SRE) combines software and...  ...Software Engineer chose to join SRE.As the Senior Engineering Manager for Collaboration... 
    Senior

    Google

    Sunnyvale, CA
    14 days ago
  • $184k - $287.5k

     ...make a lasting impact on the world.We are looking for an experienced Senior Software Engineer for the Embedded Platform team. This is an outstanding opportunity to accelerate the pace of Jetson, IGX and DGX Spark Product Software system development within NVIDIA. Using... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    6 days ago
  • $184k - $287.5k

    We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution...  ...and production-ready.Operational Reliability: Implement atomic update semantics...  ...(NVIDIA Cloud Functions) or DGX Cloud.Deep familiarity with Isaac... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    14 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer, DGX Cloud. Be the first to apply!