Senior Platform Engineer, Network Infrastructure - DGX Cloud
$176k - $276kNVIDIA
Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, and operate NVIDIA’s global network across data centers, colocation facilities, and cloud environments. The team owns the architecture and lifecycle of this platform, including cluster provisioning and upgrades, GitOps delivery, observability, capacity, and service enablement. We build software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments.We are looking for a hands-on senior engineer to own the lifecycle and automation of the Kubernetes platform supporting GNI network systems. You will also provide production support for network services running on the platform, partnering with their engineering owners when issues or changes cross the platform boundary. You will take complex problems from design through production and remain accountable for the outcome. You will bring deep Kubernetes expertise and help establish consistent engineering practices across the US and Bangalore teams. This is a senior individual contributor role with end-to-end ownership and production responsibility.What You’ll Be Doing:Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments.Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps.Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features.Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. Drive issues from initial signal through verified resolution.Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery.Participate in CFR’s production on-call rotation, including scheduled after-hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion.What We Need to See:Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems.Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.Proficiency in at least one general-purpose programming language, such as Go or Python.Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.Experience deploying and supporting network automation or telemetry services on Kubernetes.Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion.Ways to Stand Out From the Crowd:Strong knowledge of IP routing, data center fabrics, and cloud networking is a great plus.Experience designing and operating large, multi-region Kubernetes fleets, including fleet-wide upgrades and recovery.Hands-on experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle, machine remediation, and upgrades.Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns.Experience designing or operating network automation and telemetry services on Kubernetes at global scale.Contributions to Cluster API, Metal3, or other open-source Kubernetes infrastructure projects.NVIDIA’s deep learning platforms have made major impact to various fields is broadly used across leading academic institutions, start-ups, and industry, including the world’s largest Internet companies. We need passionate, hard-working and creative people to help us take on more of these unique opportunities in deep learning cloud solutions. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD - 276,000 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 3, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, IL, Remote; US, WA, Remote; US, Remote; US, CA, Remote; US, MA, RemoteType: Full time
$208k - $327.75k
...a world-class Senior Product Manager... ...the NVIDIA DGX is the undisputed... ...as the public cloud? The mission... ...provisioning and network fabric... ...self-healing infrastructure. Thoughtfully... ...experience.The "Platform-First" Approach... ...intersection of multiple engineering fields. As you...CloudSeniorNetworkFull timeNight shift- Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of... ...day is currently Tuesday.Engineering at Lambda is responsible for... ...standards for resource management, networking, and RBAC across the platform.Lead incident response, root...CloudSeniorNetworkWork at officeLocal areaWork from homeFlexible hours
$174k - $253k
...the impact on hardware, network, or service operations... .... Google's software engineers develop the next-generation... ...solutions.The AI and Infrastructure team is redefining... ...include Googlers, Google Cloud customers, and... ...providing the essential platforms that enable developers...CloudSeniorNetworkWorldwide$174k - $253k
...the impact on hardware, network, or service operations... ...large-scale infrastructure, distributed systems or... ...technologies. Google's software engineers develop the next-... ...include Googlers, Googler Cloud customers, and... ...services, and offers platforms that developers use to...CloudSeniorNetworkWorldwide$174k - $253k
...code developed by other engineers and provide feedback... ...the impact on hardware, network, or service operations... ...solutions.The AI and Infrastructure team is redefining... ...include Googlers, Google Cloud customers, and billions... ...the essential platforms that enable developers...CloudSeniorNetworkWorldwide$178k - $321k
...capability: a multi-agent platform (Hive Mind), agentic... ...harness: a resilient cloud platform, the agentic... ...governed data and AI infrastructure everything else... ...This is a two-person engineering team: you deploy, debug... ...code, CI/CD, identity, networking, observability, and...CloudSeniorNetwork$184k - $287.5k
...workloads. We are looking for a Senior Software Engineer to lead the bring-up,... ...across NVIDIA GPU platforms at the largest scales we run... ...large-scale AI clusters, infrastructure, and end-to-end workloads,... ...performance across compute, memory, networking, and communication layers...CloudSeniorNetworkFull timeRemote work$152k - $241.5k
The Autonomous Vehicles Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's... ...-in-the-loop (HIL) infrastructure to support the robust... ...Docker, Kubernetes) in cloud-native or hybrid environments... ...in automotive networking protocols (Ethernet, CAN...CloudSeniorNetworkFull time$159k - $231k
...system-level power delivery network solutions across silicon,... ...Bachelor’s degree in Electrical Engineering, Computer Engineering,... ...Integrity Engineer within Platforms Infrastructure Engineering, you will play... ...customers include Googlers, Google Cloud customers, and billions of...CloudSeniorNetworkWorldwide$174k - $253k
...the impact on hardware, network, or service operations... ...large-scale infrastructure, distributed systems or... ...Experience developing Cloud or SaaS products.Expertise... ...projects.Google's software engineers develop the next-... ...We focus on building a platform that allows enterprises...CloudSeniorNetwork$280k - $380k
...#1 TV streaming platform in the U.S., Canada... ...-active, multi-cloud platform on AWS and... ...and automation, engineering systems that perform... ...turning complex infrastructure into reliable,... ...Reliability Engineering) Senior Software Engineer... ...understanding of networking, security, and...CloudSeniorNetworkWork at officeLocal areaRemote workMonday to ThursdayFlexible hours$150k - $218k
...test fixture for better engineering efficiency and data... ...powerful computing infrastructures with custom-built machines... ...Googlers, Google Cloud customers, and billions... ...the essential platforms that enable developers... ...Cloud, Google Global Networking, Data Center operations...CloudSeniorNetworkWorldwide$159k - $231k
...initiatives and roadmap, engineering execution with... ...to join our Physical infrastructure Robotics team. This team... ...include Googlers, Google Cloud customers, and... ...providing the essential platforms that enable developers... ...Cloud, Google Global Networking, Data Center...CloudSeniorNetworkContract workWorldwide$236k - $330k
...initiatives and roadmap, engineering execution with... ...liquid-cooled AI hardware infrastructure.Oversee the end-to-... ...bipedal/wheeled humanoid platforms for flexible, human-... ...Googlers, Google Cloud customers, and billions... ...Cloud, Google Global Networking, Data Center...CloudSeniorNetworkContract workRemote workWorldwideFlexible hours$184k - $287.5k
We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline... ...addressing function-to-function networking, gRPC bottlenecks, and in-cluster... ...NVCF (NVIDIA Cloud Functions) or DGX Cloud.Deep familiarity with Isaac...CloudSeniorNetworkFull timeLocal area$200k - $322k
...NVIDIA’s DGX Cloud team helps some of the most advanced... ..., and long-term platform success. The work is... ...translating complex infrastructure into practical outcomes... ...recommendations across compute, networking, storage, and cloud... ...gainsightWork across Engineering, Product, Operations,...CloudSeniorNetworkFull time$184k - $287.5k
NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure for AI research and production workloads. We are looking for Senior Software Engineers to help build the automation, tooling... ...up work.Partner with platform, storage, networking, security, and...CloudSeniorNetworkFull time$224k - $356.5k
Joining NVIDIA's DGX Cloud AI Efficiency Team means... ...AI researchers and platform teams understand... ...behavior across GPUs, networking, storage, and... .... We are seeking a Senior Performance Engineer to characterize workloads... ...diverse team of infrastructure experts to unlock more...CloudSeniorNetworkFull timeRemote work$184k - $287.5k
...on the world. The DGX Cloud organization at... ...forward‑thinking engineers tackling some of... ...searching for a Senior Systems Software... ...and major cloud platforms. You’ll own hard... ...help shape how AI infrastructure runs in production... ...as GPU Operator, Network Operator, node-feature...CloudSeniorNetworkFull timeRemote work$184k - $287.5k
Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative... ...software engineer to join our team. You... ...availability of AI systems.As a senior DGX Cloud AI Infrastructure... ...NVIDIA's AI platforms.Define meaningful and...CloudSeniorFull timeRemote work$272k - $431.25k
...NVIDIA DGX Cloud is the AI supercomputing-as-a-service substrate... .... As a Security Data Engineer within our Infrastructure Security Engineering organization... ...autonomous action our platform takes stands on the data foundation... ..., data classification, network isolation, and verifiable...CloudNetworkFull timeRemote work$184k - $287.5k
...technological advancement. We’re hiring a Senior Backend/Platform Engineer to build and maintain the core infrastructure behind NVIDIA Brev. You’ll develop reliable cloud services, control planes, and... ...Linux, Kubernetes, containers, networking, and public cloud...CloudSeniorNetworkFull timeRemote work$255k - $340k
..., The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands... ...Tuesday.Hardware Engineering at Lambda is responsible... ...concept for state-of-the-art platforms, new product introduction... ...purpose compute, storage, and network hardware into Lambda’s...CloudSeniorNetworkWork at officeLocal areaWork from homeFlexible hours$209k - $235k
...Role We're looking for a Senior Software Engineer to join our ML Infrastructure team and support the... ...powers Gridmatic. Our platform challenges are shaped by... ...infrastructure on a public cloud platform (we run on GCP... ...and comfort with cloud networking, IAM, and secrets...CloudSeniorNetworkFull timeHome officeFlexible hours- ...Synopsys is the leader in engineering solutions from silicon to systems... ...Are You are a strong platform engineer with a passion for... ...that improve how complex infrastructure is observed, understood,... ...compute infrastructure, storage, networking, cloud services, and business-...CloudSeniorNetwork
$100k
...software models, compilers, platforms, networking, and semiconductors.... ...contributors of all seniorities.Tenstorrent’s AI Software Infrastructure team builds the... ...backend or infrastructure engineer with experience building... ...differs from cloud-native environments at...CloudNetworkPermanent employment$224k - $356.5k
...transforming how the world uses AI, cloud, and accelerated computing,... ...show customers their NVIDIA platforms are healthy, resilient, and... ...across data center, AI, networking, and partner environments.... ...Kubernetes and containers, CI/CD, infrastructure as code, observability tools...CloudSeniorNetworkFull timeRemote work$168k - $264.5k
NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure. The goal is to craft a reliable, scalable and efficient network to support NVIDIA software development workflows and tools, including CI/CD pipelines, compute resource management...CloudSeniorNetworkFull time$168k - $264.5k
NVIDIA is seeking a Sr Network Security Engineer to implement and maintain robust security across on-premise and cloud environments - enabling business verticals that span Graphics... ...and management of critical security infrastructure, including next-generation firewalls,...CloudSeniorNetworkFull time$184k - $287.5k
...looking for an experienced infrastructure Solutions Architect. Do you... ...field? We are looking for a networking savvy Solutions Architect to... ...you’ll be doing:Working with Cloud Providers and Hyperscalers to... ...recommendations to business and engineering teams on product strategy....CloudSeniorNetworkFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Platform Engineer, Network Infrastructure - DGX Cloud. Be the first to apply!
- platform engineer Santa Clara, CA
- platform developer Santa Clara, CA
- network engineer level Santa Clara, CA
- network infrastructure engineer Santa Clara, CA
- network engineer - transport Santa Clara, CA
- data center network engineer Santa Clara, CA
- network implementation engineer Santa Clara, CA
- principal network engineer Santa Clara, CA
- ip network engineer Santa Clara, CA
- network applications engineer Santa Clara, CA

