Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Platform Engineer, Network Infrastructure - DGX Cloud

$176k - $276k

NVIDIA

Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, and operate NVIDIA’s global network across data centers, colocation facilities, and cloud environments. The team owns the architecture and lifecycle of this platform, including cluster provisioning and upgrades, GitOps delivery, observability, capacity, and service enablement. We build software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments.We are looking for a hands-on senior engineer to own the lifecycle and automation of the Kubernetes platform supporting GNI network systems. You will also provide production support for network services running on the platform, partnering with their engineering owners when issues or changes cross the platform boundary. You will take complex problems from design through production and remain accountable for the outcome. You will bring deep Kubernetes expertise and help establish consistent engineering practices across the US and Bangalore teams. This is a senior individual contributor role with end-to-end ownership and production responsibility.What You’ll Be Doing:Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments.Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps.Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features.Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. Drive issues from initial signal through verified resolution.Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery.Participate in CFR’s production on-call rotation, including scheduled after-hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion.What We Need to See:Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems.Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.Proficiency in at least one general-purpose programming language, such as Go or Python.Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.Experience deploying and supporting network automation or telemetry services on Kubernetes.Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion.Ways to Stand Out From the Crowd:Strong knowledge of IP routing, data center fabrics, and cloud networking is a great plus.Experience designing and operating large, multi-region Kubernetes fleets, including fleet-wide upgrades and recovery.Hands-on experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle, machine remediation, and upgrades.Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns.Experience designing or operating network automation and telemetry services on Kubernetes at global scale.Contributions to Cluster API, Metal3, or other open-source Kubernetes infrastructure projects.​NVIDIA’s deep learning platforms have made major impact to various fields is broadly used across leading academic institutions, start-ups, and industry, including the world’s largest Internet companies. We need passionate, hard-working and creative people to help us take on more of these unique opportunities in deep learning cloud solutions. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD - 276,000 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 3, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, IL, Remote; US, WA, Remote; US, Remote; US, CA, Remote; US, MA, RemoteType: Full time

Vacancy posted 8 hours ago
Similar jobs that could be interesting for youBased on the Senior Platform Engineer, Network Infrastructure - DGX Cloud in Santa Clara, CA vacancy
  • $208k - $327.75k

     ...a world-class Senior Product Manager...  ...the NVIDIA DGX is the undisputed...  ...as the public cloud? The mission...  ...provisioning and network fabric...  ...self-healing infrastructure. Thoughtfully...  ...experience.The "Platform-First" Approach...  ...intersection of multiple engineering fields. As you... 
    Cloud
    Senior
    Network
    Full time
    Night shift

    Nvidia

    Santa Clara, CA
    8 hours ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of...  ...day is currently Tuesday.Engineering at Lambda is responsible for...  ...standards for resource management, networking, and RBAC across the platform.Lead incident response, root... 
    Cloud
    Senior
    Network
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    8 hours ago
  • $174k - $253k

     ...the impact on hardware, network, or service operations...  .... Google's software engineers develop the next-generation...  ...solutions.The AI and Infrastructure team is redefining...  ...include Googlers, Google Cloud customers, and...  ...providing the essential platforms that enable developers... 
    Cloud
    Senior
    Network
    Worldwide

    Google

    Sunnyvale, CA
    8 hours ago
  • $174k - $253k

     ...the impact on hardware, network, or service operations...  ...large-scale infrastructure, distributed systems or...  ...technologies. Google's software engineers develop the next-...  ...include Googlers, Googler Cloud customers, and...  ...services, and offers platforms that developers use to... 
    Cloud
    Senior
    Network
    Worldwide

    Google

    Sunnyvale, CA
    8 hours ago
  • $174k - $253k

     ...code developed by other engineers and provide feedback...  ...the impact on hardware, network, or service operations...  ...solutions.The AI and Infrastructure team is redefining...  ...include Googlers, Google Cloud customers, and billions...  ...the essential platforms that enable developers... 
    Cloud
    Senior
    Network
    Worldwide

    Google

    Sunnyvale, CA
    8 hours ago
  • $178k - $321k

     ...capability: a multi-agent platform (Hive Mind), agentic...  ...harness: a resilient cloud platform, the agentic...  ...governed data and AI infrastructure everything else...  ...This is a two-person engineering team: you deploy, debug...  ...code, CI/CD, identity, networking, observability, and... 
    Cloud
    Senior
    Network

    OKX

    San Jose, CA
    3 days ago
  • $184k - $287.5k

     ...workloads. We are looking for a Senior Software Engineer to lead the bring-up,...  ...across NVIDIA GPU platforms at the largest scales we run...  ...large-scale AI clusters, infrastructure, and end-to-end workloads,...  ...performance across compute, memory, networking, and communication layers... 
    Cloud
    Senior
    Network
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

    The Autonomous Vehicles Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's...  ...-in-the-loop (HIL) infrastructure to support the robust...  ...Docker, Kubernetes) in cloud-native or hybrid environments...  ...in automotive networking protocols (Ethernet, CAN... 
    Cloud
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $159k - $231k

     ...system-level power delivery network solutions across silicon,...  ...Bachelor’s degree in Electrical Engineering, Computer Engineering,...  ...Integrity Engineer within Platforms Infrastructure Engineering, you will play...  ...customers include Googlers, Google Cloud customers, and billions of... 
    Cloud
    Senior
    Network
    Worldwide

    Google

    Sunnyvale, CA
    8 hours ago
  • $174k - $253k

     ...the impact on hardware, network, or service operations...  ...large-scale infrastructure, distributed systems or...  ...Experience developing Cloud or SaaS products.Expertise...  ...projects.Google's software engineers develop the next-...  ...We focus on building a platform that allows enterprises... 
    Cloud
    Senior
    Network

    Google

    Sunnyvale, CA
    2 days ago
  • $280k - $380k

     ...#1 TV streaming platform in the U.S., Canada...  ...-active, multi-cloud platform on AWS and...  ...and automation, engineering systems that perform...  ...turning complex infrastructure into reliable,...  ...Reliability Engineering) Senior Software Engineer...  ...understanding of networking, security, and... 
    Cloud
    Senior
    Network
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  • $150k - $218k

     ...test fixture for better engineering efficiency and data...  ...powerful computing infrastructures with custom-built machines...  ...Googlers, Google Cloud customers, and billions...  ...the essential platforms that enable developers...  ...Cloud, Google Global Networking, Data Center operations... 
    Cloud
    Senior
    Network
    Worldwide

    Google

    Sunnyvale, CA
    2 days ago
  • $159k - $231k

     ...initiatives and roadmap, engineering execution with...  ...to join our Physical infrastructure Robotics team. This team...  ...include Googlers, Google Cloud customers, and...  ...providing the essential platforms that enable developers...  ...Cloud, Google Global Networking, Data Center... 
    Cloud
    Senior
    Network
    Contract work
    Worldwide

    Google

    Sunnyvale, CA
    4 days ago
  • $236k - $330k

     ...initiatives and roadmap, engineering execution with...  ...liquid-cooled AI hardware infrastructure.Oversee the end-to-...  ...bipedal/wheeled humanoid platforms for flexible, human-...  ...Googlers, Google Cloud customers, and billions...  ...Cloud, Google Global Networking, Data Center... 
    Cloud
    Senior
    Network
    Contract work
    Remote work
    Worldwide
    Flexible hours

    Google

    Sunnyvale, CA
    2 days ago
  • $184k - $287.5k

    We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline...  ...addressing function-to-function networking, gRPC bottlenecks, and in-cluster...  ...NVCF (NVIDIA Cloud Functions) or DGX Cloud.Deep familiarity with Isaac... 
    Cloud
    Senior
    Network
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    4 days ago
  • $200k - $322k

     ...NVIDIA’s DGX Cloud team helps some of the most advanced...  ..., and long-term platform success. The work is...  ...translating complex infrastructure into practical outcomes...  ...recommendations across compute, networking, storage, and cloud...  ...gainsightWork across Engineering, Product, Operations,... 
    Cloud
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure for AI research and production workloads. We are looking for Senior Software Engineers to help build the automation, tooling...  ...up work.Partner with platform, storage, networking, security, and... 
    Cloud
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means...  ...AI researchers and platform teams understand...  ...behavior across GPUs, networking, storage, and...  .... We are seeking a Senior Performance Engineer to characterize workloads...  ...diverse team of infrastructure experts to unlock more... 
    Cloud
    Senior
    Network
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...on the world. The DGX Cloud organization at...  ...forward‑thinking engineers tackling some of...  ...searching for a Senior Systems Software...  ...and major cloud platforms. You’ll own hard...  ...help shape how AI infrastructure runs in production...  ...as GPU Operator, Network Operator, node-feature... 
    Cloud
    Senior
    Network
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative...  ...software engineer to join our team. You...  ...availability of AI systems.As a senior DGX Cloud AI Infrastructure...  ...NVIDIA's AI platforms.Define meaningful and... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...NVIDIA DGX Cloud is the AI supercomputing-as-a-service substrate...  .... As a Security Data Engineer within our Infrastructure Security Engineering organization...  ...autonomous action our platform takes stands on the data foundation...  ..., data classification, network isolation, and verifiable... 
    Cloud
    Network
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...technological advancement. We’re hiring a Senior Backend/Platform Engineer to build and maintain the core infrastructure behind NVIDIA Brev. You’ll develop reliable cloud services, control planes, and...  ...Linux, Kubernetes, containers, networking, and public cloud... 
    Cloud
    Senior
    Network
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $255k - $340k

     ..., The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands...  ...Tuesday.Hardware Engineering at Lambda is responsible...  ...concept for state-of-the-art platforms, new product introduction...  ...purpose compute, storage, and network hardware into Lambda’s... 
    Cloud
    Senior
    Network
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    8 hours ago
  • $209k - $235k

     ...Role We're looking for a Senior Software Engineer to join our ML Infrastructure team and support the...  ...powers Gridmatic. Our platform challenges are shaped by...  ...infrastructure on a public cloud platform (we run on GCP...  ...and comfort with cloud networking, IAM, and secrets... 
    Cloud
    Senior
    Network
    Full time
    Home office
    Flexible hours

    Gridmatic Inc

    Cupertino, CA
    2 days ago
  •  ...Synopsys is the leader in engineering solutions from silicon to systems...  ...Are You are a strong platform engineer with a passion for...  ...that improve how complex infrastructure is observed, understood,...  ...compute infrastructure, storage, networking, cloud services, and business-... 
    Cloud
    Senior
    Network

    Synopsys

    Sunnyvale, CA
    3 days ago
  • $100k

     ...software models, compilers, platforms, networking, and semiconductors....  ...contributors of all seniorities.Tenstorrent’s AI Software Infrastructure team builds the...  ...backend or infrastructure engineer with experience building...  ...differs from cloud-native environments at... 
    Cloud
    Network
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    8 hours ago
  • $224k - $356.5k

     ...transforming how the world uses AI, cloud, and accelerated computing,...  ...show customers their NVIDIA platforms are healthy, resilient, and...  ...across data center, AI, networking, and partner environments....  ...Kubernetes and containers, CI/CD, infrastructure as code, observability tools... 
    Cloud
    Senior
    Network
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $168k - $264.5k

    NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure. The goal is to craft a reliable, scalable and efficient network to support NVIDIA software development workflows and tools, including CI/CD pipelines, compute resource management... 
    Cloud
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    8 hours ago
  • $168k - $264.5k

    NVIDIA is seeking a Sr Network Security Engineer to implement and maintain robust security across on-premise and cloud environments - enabling business verticals that span Graphics...  ...and management of critical security infrastructure, including next-generation firewalls,... 
    Cloud
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    8 hours ago
  • $184k - $287.5k

     ...looking for an experienced infrastructure Solutions Architect. Do you...  ...field? We are looking for a networking savvy Solutions Architect to...  ...you’ll be doing:Working with Cloud Providers and Hyperscalers to...  ...recommendations to business and engineering teams on product strategy.... 
    Cloud
    Senior
    Network
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Platform Engineer, Network Infrastructure - DGX Cloud. Be the first to apply!