Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Systems Software Engineer, Kubernetes Node Lifecycle - DGX Cloud

$184k - $287.5k

Nvidia

At NVIDIA, the DGX Cloud division merges fresh hardware and software innovations to offer leading accelerated computing solutions for the most challenging AI workloads worldwide. Our team of skilled engineers is committed to addressing major global issues, consistently advancing technology, and making a difference in millions of lives around the world!We are looking for a Senior Systems Software Engineer with strong experience in Kubernetes node engineering, OS image packaging, and cloud infrastructure. The ideal candidate will possess deep hyperscaler-level knowledge across the entire node lifecycle. This covers CAPI providers, bring-your-own-node onboarding, OS image build pipelines, packaging, and nodepool management. They must have the technical depth needed to maintain cluster reliability at frontier AI scale. In this vital role, you will manage the node layer within NVIDIA Kubernetes Engine (NKE). Your work will ensure it scales to fulfill DGX Cloud's two main goals: supporting internal researchers and enabling NCPs. Are you prepared to innovate?What you'll be doing:Direct the building and refinement of CAPI providers for NVIDIA Kubernetes Engine, maintaining steady, consistent, and scalable node provisioning across DGX Cloud and NCP environments.Develop and maintain bring-your-own-node workflows that allow customers to integrate different NVIDIA hardware into NKE clusters while ensuring high operational consistency.Coordinate OS image generation, packaging, deployment, and update processes for NKE nodes. Ensure images are fine-tuned for NVIDIA GPU workloads and satisfy enterprise- and cloud-grade security and compliance criteria.Develop and sustain node image hardening pipelines, incorporating CIS benchmarks, automated CVE remediation, and promotion gates connected to security posture.Develop and maintain automated test suites for node images. These tests verify accuracy across Kubernetes versions and NVIDIA hardware configurations. This process occurs prior to production deployment and facilitates continuous validation through modern CI/CD pipelines.Handle nodepool lifecycle at scale, including provisioning, upgrades, drain and cordon workflows, and seamless node replacement across very large clusters with diverse NVIDIA hardware.Examine, resolve, and determine underlying causes of node-layer faults in production NKE clusters, such as those involving image configuration, driver packaging, kubelet operation, and hardware activation, and review and optimize the node layer in real-world high-scale scenarios.Partner with upstream communities including Cluster API, Kubernetes, and CNCF projects to establish node provisioning and lifecycle standards in accordance with NKE requirements. Communicate your progress and findings at internal and external gatherings such as KubeCon and GTC.What we need to see:8 years of experience with a background in systems software, cloud infrastructure, or Kubernetes node engineering.Bachelor’s or Master’s degree in Engineering (Electrical, Computer Engineering, Computer Science) or equivalent experience.Deep expertise in Cluster API (CAPI), including provider development and full machine lifecycle from provisioning to deletion.Extensive experience with OS image build pipelines, node image packaging, and delivery systems for Kubernetes nodes (for example image-builder, containerd, cloud-init, packer).Practical experience with bring-your-own-node models and integrating diverse hardware into live Kubernetes environments, including large-scale nodepool lifecycle management and upgrades.Strong understanding of kubelet configuration, node bootstrap, and the Kubernetes node registration lifecycle.Experience with node image security, including vulnerability scanning, patch automation, and compliance gating as part of image build pipelines.Proficiency in Golang and/or Python, and hands-on experience with at least one major public cloud provider (GCP, AWS, Azure, OCI or equivalent).Ways to stand out from the crowd:Direct experience building or maintaining node image pipelines for a hyperscaler Kubernetes distribution (GKE, EKS, AKS, OKE, or equivalent).Experience with supply chain security and hardening for node images, including image signing, provenance attestation, SBOM generation, CIS benchmark consistency, and automated CVE remediation.Experience with automated node provisioning and optimal sizing at scale (for example Karpenter, GKE NAP or similar) and how these interact with GPU workload scheduling.Strong operational experience working with immutable OS image distributions (such as Flatcar, Bottlerocket, Azure Linux) and debugging node-layer failures in large Kubernetes clusters.Proven background of upstream contributions to Cluster API, Kubernetes or related CNCF projects, combined with excellent communication and interpersonal abilities.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 14, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, WA, SeattleType: Full time

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Systems Software Engineer, Kubernetes Node Lifecycle - DGX Cloud in Santa Clara, CA vacancy
  • $184k - $287.5k

     ...on the world. The DGX Cloud organization at...  ...edge hardware and software innovation to deliver...  ...forward‑thinking engineers tackling some of...  ...searching for a Senior Systems Software Engineer...  ...systems, Kubernetes, containers, and...  ...Network Operator, node-feature-discovery... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $200k - $322k

    NVIDIA’s DGX Cloud team helps some of the most advanced AI builders...  ...customers across the full lifecycle of their DGX Cloud usage,...  ...friction.gainsightWork across Engineering, Product, Operations, and...  ...working with distributed systems, Kubernetes, schedulers, or large-scale... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...Senior Systems Software Engineer This role has been designed as "Onsite" with an expectation that you...  ...Packard Enterprise is the global edge-to-cloud company advancing the way people...  ...Cisco is a strong plus Knowledge of Kubernetes and associated technologies... 
    Cloud
    Senior
    Work experience placement
    Work at office
    2 days per week

    Hewlett Packard Enterprise

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

    We are looking for a Senior Software Engineer to join our DGX Cloud team and build the foundational systems that drive NVIDIA’s high-performance...  ...streamline infrastructure lifecycle processes.Collaborate...  ...container orchestration tools like Kubernetes and observability systems... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...Vehicles Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's autonomous...  ...support the robust deployment and lifecycle management of next-generation...  ...pipelines (e.g., Jenkins, Docker, Kubernetes) in cloud-native or hybrid environments for... 
    Cloud
    Senior
    Full time

    2100 NVIDIA USA

    Santa Clara, CA
    5 days ago
  • $184k

     ...impact on the world. We are looking for a dedicated engineer for the Senior Systems Software Engineer role, focusing on GPU Performance at Scale....  ...computing software stacks (CUDA). * Experience with modern cloud and container-based enterprise computing architectures... 
    Cloud
    Senior
    Full time

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    NVIDIA DGX Cloud is building and operating large-scale...  ...We are looking for Senior Software Engineers to help build the...  ...tooling, and operational systems that make GPU...  ...engineering team focused on Kubernetes-based infrastructure...  ...repair, and cluster lifecycle operations.Improve... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $200k - $322k

    NVIDIA's DGX Cloud (DGXC) powers AI for...  ...seeks a Senior Technical Program...  ...generation AI software platforms. In...  ...infrastructure, and system integration....  ...high-impact engineering programs within...  ...software development lifecycle.What We Need...  ...systems, Kubernetes-based environments... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $208k - $327.75k

     ...world-class Senior Product...  ...the NVIDIA DGX is the undisputed...  ...a 1,000-node private cluster...  ...the public cloud? The...  ...role, own the software-defined blueprint...  ...On-Prem Lifecycle: Define the...  ...to Kubernetes: Lead the integration...  ...of DGX systems into the...  ...of multiple engineering fields. As... 
    Cloud
    Senior
    Full time
    Night shift

    Nvidia

    Santa Clara, CA
    4 hours ago
  • $229.9k - $262.4k

     ...Senior Lead Software Engineer, Distributed Systems (Golang + Python on Kubernetes) Do you love building and pioneering in the technology...  ...managers, and deliver robust cloud-based solutions that drive powerful...  ...: Python, Golang, or Node.js ~4+ years of experience... 
    Cloud
    Senior
    Full time
    Part time
    Internship
    Local area

    Capital One

    San Jose, CA
    5 days ago
  • $224k - $356.5k

     ...transforming how the world uses AI, cloud, and accelerated computing,...  ...and large-scale distributed systems. We partner closely with...  ...cloud-native platforms such as Kubernetes and containers, CI/CD,...  ...including PKI, TLS/mTLS, certificate lifecycle, signing, secrets management,... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...motivated Performance engineer to influence the...  ..., PCIe) within a node and with high-...  ...understanding of computer system architecture, HW-...  ...(aka systems software fundamentals)Implement...  ...with containers, cloud provisioning and scheduling tools (Kubernetes, SLURM, Ansible,... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...best work. The video software team is seeking someone...  ...and passionate about system software development....  ...include ultra-low latency cloud gaming, video...  ...features through the whole lifecycle from requirements and...  ...Bachelors in Electrical Engineering or Computer Science (or... 
    Cloud
    Senior
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    4 hours ago
  • $200k - $322k

    NVIDIA is seeking a Senior Technical Program...  ...programs for DGX Cloud. DGX Cloud powers...  ...security, compliance, engineering execution, and...  ..., platform, and software teams.Establish program...  ..., distributed systems, or...  ...product or software lifecycle processes, including... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

    GeForce NOW is NVIDIA's cloud gaming platform, delivering ultra-low-latency, high...  ...mobility.We're looking for a Senior Systems Software Engineer with strong C++ experience and familiarity...  ...architecture, peer connection lifecycle, ICE handshake, and real-time protocols... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $272k - $431.25k

     ...looking for a Principal Software Engineer to join our DGX Cloud team and build the foundational systems that drive NVIDIA’s...  ...that automate fleet lifecycle operations at a...  ...mentoring, and encouraging senior engineers, elevating...  ...schedulers (e.g., Kubernetes, Slurm) and... 
    Cloud
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...impact on the world.We are looking for a senior systems software engineer to advance the Continuous...  ...Strong understanding and experience with Kubernetes administration and deployment.Familiarity...  ...and security requirements in cloud and on-prem environments.Experience... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $272k - $431.25k

     ...Principal Systems Software Engineer at NVIDIA is an engineering discipline to design, build...  ...and deployment and open source cloud enabling technologies like Kubernetes and OpenStack. SRE at NVIDIA...  ...Engage in and improve the whole lifecycle of services—from inception and... 
    Cloud

    NVIDIA Gruppe

    Santa Clara, CA
    1 day ago
  • $195.2k - $391.2k

     ...OpportunityWe are looking for a Senior Engineering Manager to lead the...  ...of a next-generation Kubernetes platform powering...  ..., including cluster lifecycle, fleet management,...  ..., and storage systems.About the TeamWe are...  ...-on familiarity with cloud platforms, infrastructure... 
    Cloud
    Senior
    Work at office
    Local area
    Remote work
    Relocation package
    3 days per week

    Nutanix

    San Jose, CA
    4 hours ago
  • $160k - $180k

     ...What You’ll Do As a Senior AI Systems Engineer, you will architect,...  ...the AI development lifecycle. Compute & Inference...  ...utilization, managing multi-cloud compute scheduling...  ...AI researchers and Software Engineers to...  ...grade orchestration (Kubernetes), paired with cloud‑... 
    Cloud
    Senior
    Local area

    Archer56

    San Jose, CA
    2 days ago
  • $152k - $241.5k

     ...is hiring experienced software engineers to help scale up its AI...  ...operator development, node health monitoring and working...  ...You will be part of an DGX Cloud team responsible for production systems that enable large...  ...cluster management systems (Kubernetes, Slurm, Base Command... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA is seeking a Senior Software Engineer to join our CSP Engagements team, focusing on system software for Datacenter products such...  ...facing responsibilities to enable cloud service providers with next-...  ...with virtualization, Kubernetes, and cloud-native architectures... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

    GeForce NOW is Nvidia’s Cloud Gaming service, streaming games...  ...and Nvidia proprietary software, GeForce NOW transforms the...  ...see We are looking for a Senior System Software Engineer for Cloud who sees the big...  ...drive best practices in Kubernetes, observability, and infrastructure... 
    Cloud
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    4 hours ago
  • $176k - $276k

    Cloud Foundations Reliability (CFR) is part...  ...and operate the Kubernetes-based platform...  ...architecture and lifecycle of this platform,...  ...enablement. We build software and automation to...  ...for a hands-on senior engineer to own the...  ...supporting GNI network systems. You will also provide... 
    Cloud
    Senior
    Full time
    Remote work
    Weekend work

    Nvidia

    Santa Clara, CA
    4 hours ago
  • $153k - $204k

     ...CoreWeave is The Essential Cloud for AI™. Built for pioneers...  ...at What You'll Do As a Senior Engineer in Compute Services, you will...  ...automated tooling to provision Kubernetes control planes on bare-metal...  ...operators • Perform day 2 lifecycle tasks and maintenance on... 
    Cloud
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    14 days ago
  • $184k - $287.5k

    We are looking for a senior systems software engineer to improve the operation and user experience of distributed system infrastructure using AI....  ...Background with container technologies such as Docker and Kubernetes.Experience with DevOps tools such as Ansible, Terraform,... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...is hiring experienced software engineers to help scale up its AI...  ...operator development, node health monitoring and working...  ...You will be part of an DGX Cloud team responsible for production systems that enable large...  ...cluster management systems (Kubernetes, Slurm, Base Command... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 hours ago
  • $176k - $276k

    Production engineering is a field that involves...  ...-scale production systems with high efficiency...  ...areas, including software and systems engineering...  ...with open-source cloud-enabling technologies such as Kubernetes, containers, and...  ....Improve the lifecycle of storage services... 
    Cloud
    Senior
    Full time
    Flexible hours

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

    DGX Station (Galaxy) is NVIDIA’s workstation-class AI computer—built on GB300 Blackwell GPUs with NVLink interconnect, delivering...  ...of this platform.We are looking for a deeply technical systems software engineer who will own AI stack readiness on DGX Station. You will profile... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    4 hours ago
  • $151.8k - $265.35k

     ...CA, where your engineering skills will...  ...platforms, agentic systems, or developer-...  ...such as Node.js, React, or...  ...Experience with Kubernetes and modern deployment...  ...of major cloud platforms such...  ...understanding of software development...  ...software development lifecycle. Technical... 
    Cloud
    Senior
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe

    San Jose, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Systems Software Engineer, Kubernetes Node Lifecycle - DGX Cloud. Be the first to apply!