Senior Systems Software Engineer, Kubernetes Node Lifecycle - DGX Cloud (Santa Clara)
$184k - $287.5kNvidia
At NVIDIA, the DGX Cloud division merges fresh hardware and software innovations to offer leading accelerated computing solutions for the most challenging AI workloads worldwide. Our team of skilled engineers is committed to addressing major global issues, consistently advancing technology, and making a difference in millions of lives around the world!We are looking for a Senior Systems Software Engineer with strong experience in Kubernetes node engineering, OS image packaging, and cloud infrastructure. The ideal candidate will possess deep hyperscaler-level knowledge across the entire node lifecycle. This covers CAPI providers, bring-your-own-node onboarding, OS image build pipelines, packaging, and nodepool management. They must have the technical depth needed to maintain cluster reliability at frontier AI scale. In this vital role, you will manage the node layer within NVIDIA Kubernetes Engine (NKE). Your work will ensure it scales to fulfill DGX Cloud's two main goals: supporting internal researchers and enabling NCPs. Are you prepared to innovate?What you'll be doing:Direct the building and refinement of CAPI providers for NVIDIA Kubernetes Engine, maintaining steady, consistent, and scalable node provisioning across DGX Cloud and NCP environments.Develop and maintain bring-your-own-node workflows that allow customers to integrate different NVIDIA hardware into NKE clusters while ensuring high operational consistency.Coordinate OS image generation, packaging, deployment, and update processes for NKE nodes. Ensure images are fine-tuned for NVIDIA GPU workloads and satisfy enterprise- and cloud-grade security and compliance criteria.Develop and sustain node image hardening pipelines, incorporating CIS benchmarks, automated CVE remediation, and promotion gates connected to security posture.Develop and maintain automated test suites for node images. These tests verify accuracy across Kubernetes versions and NVIDIA hardware configurations. This process occurs prior to production deployment and facilitates continuous validation through modern CI/CD pipelines.Handle nodepool lifecycle at scale, including provisioning, upgrades, drain and cordon workflows, and seamless node replacement across very large clusters with diverse NVIDIA hardware.Examine, resolve, and determine underlying causes of node-layer faults in production NKE clusters, such as those involving image configuration, driver packaging, kubelet operation, and hardware activation, and review and optimize the node layer in real-world high-scale scenarios.Partner with upstream communities including Cluster API, Kubernetes, and CNCF projects to establish node provisioning and lifecycle standards in accordance with NKE requirements. Communicate your progress and findings at internal and external gatherings such as KubeCon and GTC.What we need to see:8 years of experience with a background in systems software, cloud infrastructure, or Kubernetes node engineering.Bachelor’s or Master’s degree in Engineering (Electrical, Computer Engineering, Computer Science) or equivalent experience.Deep expertise in Cluster API (CAPI), including provider development and full machine lifecycle from provisioning to deletion.Extensive experience with OS image build pipelines, node image packaging, and delivery systems for Kubernetes nodes (for example image-builder, containerd, cloud-init, packer).Practical experience with bring-your-own-node models and integrating diverse hardware into live Kubernetes environments, including large-scale nodepool lifecycle management and upgrades.Strong understanding of kubelet configuration, node bootstrap, and the Kubernetes node registration lifecycle.Experience with node image security, including vulnerability scanning, patch automation, and compliance gating as part of image build pipelines.Proficiency in Golang and/or Python, and hands-on experience with at least one major public cloud provider (GCP, AWS, Azure, OCI or equivalent).Ways to stand out from the crowd:Direct experience building or maintaining node image pipelines for a hyperscaler Kubernetes distribution (GKE, EKS, AKS, OKE, or equivalent).Experience with supply chain security and hardening for node images, including image signing, provenance attestation, SBOM generation, CIS benchmark consistency, and automated CVE remediation.Experience with automated node provisioning and optimal sizing at scale (for example Karpenter, GKE NAP or similar) and how these interact with GPU workload scheduling.Strong operational experience working with immutable OS image distributions (such as Flatcar, Bottlerocket, Azure Linux) and debugging node-layer failures in large Kubernetes clusters.Proven background of upstream contributions to Cluster API, Kubernetes or related CNCF projects, combined with excellent communication and interpersonal abilities.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 14, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, WA, SeattleType: Full time
$184k - $287.5k
...the world. The DGX Cloud organization... ...hardware and software innovation to... ...forward‑thinking engineers tackling some... ...searching for a Senior Systems Software... ...distributed systems, Kubernetes, containers,... ...Network Operator, node-feature-... ...SummaryLocation: US, CA, Santa Clara; US, WA,...CloudSeniorFull timePart timeRemote work$184k - $287.5k
...The DGX Cloud organization at NVIDIA... ...hardware and software innovation to... ...of innovative engineers dedicated to solving... ...outstanding Senior Systems Software Engineer... ...such as Kubernetes and containers... ...to ultra-large node and object countsDemonstrated... ...: US, CA, Santa Clara; US, WA,...CloudSeniorFull timePart timeWorldwide$184k - $287.5k
...are looking for a Senior Software Engineer to join our DGX Cloud team and build the foundational systems that drive NVIDIA’... ...streamline infrastructure lifecycle processes.... ...orchestration tools like Kubernetes and observability... ...: US, CA, Santa Clara; US, RemoteType: Full...CloudSeniorFull timePart time$184k - $287.5k
...NVIDIA DGX Cloud is building and operating... ...are looking for Senior Software Engineers to help build the... ...and operational systems that make GPU... ...team focused on Kubernetes-based infrastructure... ..., and cluster lifecycle operations.Improve... ...: US, CA, Santa Clara; US, RemoteType:...CloudSeniorFull timePart time$184k - $287.5k
...Senior Systems Software Engineer, Observability and Telemetry Platform at... ...deployment and open source cloud enabling technologies like Kubernetes and OpenStack. Senior... ...improve the whole lifecycle of services—from... ...SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full...CloudSeniorFull timePart time$224k - $356.5k
...how the world uses AI, cloud, and accelerated... ...large-scale distributed systems. We partner closely with... ...native platforms such as Kubernetes and containers, CI/CD,... ...TLS/mTLS, certificate lifecycle, signing, secrets... ...SummaryLocation: US, CA, Santa Clara; US, TX, Remote; US, WA...CloudSeniorFull timePart timeRemote work$208k - $327.75k
...world-class Senior Product... ...NVIDIA DGX is the undisputed... ...a 1,000-node private... ...public cloud? The... ...own the software-defined blueprint... ...On-Prem Lifecycle: Define... ...to Kubernetes: Lead the... ...integration of DGX systems into the... ...multiple engineering fields.... ...US, CA, Santa ClaraType...CloudSeniorFull timePart timeNight shift$176k - $276k
...Cloud Foundations Reliability (CFR... ...and operate the Kubernetes-based platform... ...and lifecycle of this platform... ...enablement. We build software and automation... ...a hands-on senior engineer to own the lifecycle... ...GNI network systems. You will also... ...: US, CA, Santa Clara; US, IL, Remote...CloudSeniorFull timePart timeRemote workWeekend work$184k - $287.5k
...NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements... ...team, focusing on system software for... ...responsibilities to enable cloud service providers... ...virtualization, Kubernetes, and cloud-native architectures... ...: US, CA, Santa Clara; US, TX, Austin; US,...CloudSeniorFull timePart time$224k - $356.5k
...DGX Station (Galaxy) is NVIDIA’s workstation-class AI computer—built on GB300 Blackwell... ...We are looking for a deeply technical systems software engineer who will own AI stack readiness on DGX... ...protected by law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time...SeniorFull timePart timeLocal area$224k - $356.5k
...redefining AI, HPC, and cloud computing. To... ..., our diagnostic systems need to evolve across... ...leader to engineer and propel innovation... ...involve hardware and software tools to develop... ...across the server lifecycle.Drive hardware validation... ...: US, CA, Santa Clara; US, OR, HillsboroType...CloudSeniorFull timePart time$168k - $264.5k
...play. NVIDIA is seeking a Senior Network Deployment Engineer to help build and scale our... ...inventory and ticketing systems.Support circuit testing/troubleshooting... ...plusExperience with multi-cloud networking (AWS VPC, Azure... ....SummaryLocation: US, CA, Santa Clara; US, TX, Remote; US, IL,...CloudSeniorFull timeContract workPart timeWork experience placementRemote workShift work$272k - $431.25k
...for a Principal Software Engineer to join our DGX Cloud team and build... ...the foundational systems that drive... ...automate fleet lifecycle operations at a... ...and encouraging senior engineers, elevating... ...(e.g., Kubernetes, Slurm) and innovative... ...: US, CA, Santa Clara; US, WA, SeattleType...CloudFull timePart time$184k - $287.5k
...passionate about building world-class reliability systems? Join NVIDIA as a Senior Software Engineer - Resilience Engineering, DGX Cloud, and be a pivotal part of a team that... ...protected by law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time...CloudSeniorFull timePart time$240k - $379.5k
...in generative models, autonomous systems, and large-scale research. The DGX Cloud organization builds and operates... ...workstreams.Partner with engineering, product, operations, security,... ...by law.SummaryLocation: US, CA, Santa Clara; US, WA, SeattleType: Full time...CloudSeniorFull timePart time$200k - $322k
...As a Senior Technical Program Manager... ...about Cloud Security, you... ...will drive the DGX Cloud infrastructure... ...roadmaps and the software development lifecycle. It aligns... ...Compliance, SRE, and Engineering to continually... ..., including system diagrams,... ...SummaryLocation: US, CA, Santa ClaraType:...CloudSeniorFull timePart time$272k - $431.25k
...researchers and engineers to reliably... ...own the full software development and... ...service delivery lifecycle — from roadmap... ...to senior leadership, providing... ...distributed storage systems, object... ...or large-scale cloud data services;... ...SummaryLocation: US, CA, Santa Clara; US, CA,...CloudSeniorFull timePart time$184k - $287.5k
...We are looking for a Senior Software Engineer to become part of our storage management plane team.... ...You Will Be Doing:Maintain and develop Kubernetes operators and our Container Storage... ...characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time...CloudSeniorFull timePart time$168k - $258.75k
...our NVIDIA DGX Cloud team. This... ...internal teams (Software Engineering, Production... ...lifecycle.Fleet Bring... ...maintenance, node security, and... ...and align senior multi-functional... ..., managed-Kubernetes platform layers... ...systems.Significant... ...SummaryLocation: US, CA, Santa Clara; US, WA,...CloudSeniorFull timePart timeRemote work$184k - $287.5k
...Joining NVIDIA's DGX Cloud AI Efficiency Team means... ...an AI infrastructure software engineer to join our team. You... ...implementing software and systems engineering practices... ...of AI systems.As a senior DGX Cloud AI... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US,...CloudSeniorFull timePart timeRemote work$184k - $287.5k
...world.We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU... ...stacks (CUDA).Experience with modern cloud and container-based enterprise computing... ...by law.SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US, TX, Remote; US...CloudSeniorFull timePart timeRemote work$184k - $287.5k
...best work. The video software team is seeking... ...passionate about system software development... ...ultra-low latency cloud gaming, video... ...through the whole lifecycle from requirements... ...Bachelors in Electrical Engineering or Computer... ...SummaryLocation: US, CA, Santa ClaraType: Full...CloudSeniorFull timePart timeWork experience placement$184k - $287.5k
...GeForce NOW is NVIDIA's cloud gaming platform, delivering ultra-low-latency,... ...connected mobility.We're looking for a Senior Systems Software Engineer with strong C++ experience and... ...protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, AustinType: Full time...CloudSeniorFull timePart time$184k - $287.5k
...world.We are looking for a senior systems software engineer to advance the Continuous... ...understanding and experience with Kubernetes administration and... ...security requirements in cloud and on-prem environments.Experience... ...law.SummaryLocation: US, CA, Santa ClaraType: Full time...CloudSeniorFull timePart time$272k - $431.25k
...NVIDIA DGX Cloud is scaling GPU infrastructure... ...Principal Software Engineers to help shape... ...engineering, Kubernetes-based... ...role is for senior technical leaders... ...critical systems, and turn ambiguous... ...for cluster lifecycle, validation,... ...: US, CA, Santa Clara; US, RemoteType...CloudFull timePart time$224k - $356.5k
...GeForce NOW is Nvidia’s Cloud Gaming service,... ...Nvidia proprietary software, GeForce NOW transforms... ...are looking for a Senior System Software Engineer for Cloud who sees the... ...best practices in Kubernetes, observability, and... ...SummaryLocation: US, CA, Santa ClaraType: Full time...CloudSeniorFull timePart timeLocal area$184k - $287.5k
...is seeking a Senior System Architect:... ...heterogeneous nodes, a single... ...We need an engineer to develop and... ...instability, and software defects.What... ...Slurm or Kubernetes clusters.... ...for HPC or cloud-scale environments... ...manage job lifecycle and signal... ...SummaryLocation: US, CA, Santa Clara; US, MA,...CloudSeniorFull timePart time$168k - $258.75k
...the enterprise software stack for next... ..., Kubernetes-native services... ...motivated team behind DGX systems and DGX SuperPOD... ...complete solution lifecycle—from solution... ...management or engineering experience on high-tech, cloud, AI/ML,... ...SummaryLocation: US, CA, Santa ClaraType:...CloudSeniorFull timePart time$184k - $287.5k
...seeking an NCX Senior Engineer to join our... ...intricate distributed systems, and ensure... ...NCP and Neo Cloud platforms,... ...workloads across DGX Cloud, NCP data... ...using Kubernetes, containers, and... ...and enterprise software connectivity.Build... ...SummaryLocation: US, CA, Santa Clara; US, Remote;...CloudSeniorFull timePart timeRemote work$184k - $287.5k
...designs. From single node HGX/DGX systems all the way up... ...enterprise and cloud provider... ...NVIDIA AI and HPC software stack. We are searching... ...motivated engineer to lead performance... ...tools (Docker, Kubernetes, SLURM). Understanding... ...: US, CA, Santa Clara; US, RemoteType:...CloudSeniorFull timePart time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Systems Software Engineer, Kubernetes Node Lifecycle - DGX Cloud (Santa Clara). Be the first to apply!
- IT system engineer Santa Clara, CA
- systems software developer Santa Clara, CA
- system programmer Santa Clara, CA
- react node js developer Santa Clara, CA
- react node js developer (remote) Santa Clara, CA
- senior lighting artist Santa Clara, CA
- senior hvac project manager Santa Clara, CA
- senior technical product manager Santa Clara, CA
- senior medical science liaison Santa Clara, CA
- senior accountant remote Santa Clara, CA

