Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Platform Engineer, Network Infrastructure - DGX Cloud

$176k - $276k

NVIDIA

Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, and operate NVIDIA’s global network across data centers, colocation facilities, and cloud environments. The team owns the architecture and lifecycle of this platform, including cluster provisioning and upgrades, GitOps delivery, observability, capacity, and service enablement. We build software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments. We are looking for a hands‑on senior engineer to own the lifecycle and automation of the Kubernetes platform supporting GNI network systems. You will also provide production support for network services running on the platform, partnering with their engineering owners when issues or changes cross the platform boundary. You will take complex problems from design through production and remain accountable for the outcome. You will bring deep Kubernetes expertise and help establish consistent engineering practices across the US and Bangalore teams. This is a senior individual contributor role with end‑to‑end ownership and production responsibility. What You’ll Be Doing Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments. Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery. Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi‑cluster delivery through GitOps. Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features. Diagnose complex Kubernetes platform and hosted‑service failures involving control‑plane health, cluster networking, storage, scheduling, workload placement, and multi‑cluster dependencies. Drive issues from initial signal through verified resolution. Define production‑readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery. Participate in CFR’s production on‑call rotation, including scheduled after‑hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion. What We Need To See Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience. 8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems. Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery. Proficiency in at least one general‑purpose programming language, such as Go or Python. Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery. Experience deploying and supporting network automation or telemetry services on Kubernetes. Experience with production on‑call, incident response, root‑cause analysis, and driving corrective actions to completion. Ways To Stand Out From The Crowd Strong knowledge of IP routing, data center fabrics, and cloud networking is a great plus. Experience designing and operating large, multi‑region Kubernetes fleets, including fleet‑wide upgrades and recovery. Hands‑on experience with Cluster API (CAPI) and Metal3 for bare‑metal provisioning, cluster lifecycle, machine remediation, and upgrades. Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns. Experience designing or operating network automation and telemetry services on Kubernetes at global scale. Contributions to Cluster API, Metal3, or other open‑source Kubernetes infrastructure projects. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD – 276,000 USD for Level 4, and 208,000 USD – 333,500 USD for Level 5. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until July 20, 2026. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. #J-18808-Ljbffr NVIDIA

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Platform Engineer, Network Infrastructure - DGX Cloud in Santa Clara, CA vacancy
  • $152k - $241.5k

     ...outstanding, passionate, and dedicated Senior AI Infrastructure Engineer to join our DGX Cloud group. This engineering role...  ...across different systems, networking, coding, databases, capacity management...  ...AI training and inferencing platform built on top of cloud infrastructure... 
    Cloud
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $208k - $327.75k

     ...a world-class Senior Product Manager...  ...the NVIDIA DGX is the undisputed...  ...as the public cloud? The mission...  ...provisioning and network fabric...  ...self‑healing infrastructure. Thoughtfully...  ...experience. The "Platform‑First" Approach...  ...intersection of multiple engineering fields. As you... 
    Cloud
    Senior
    Network
    Night shift

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $166k - $244k

    Senior Software Engineer, Infrastructure, Platforms Infrastructure Engineering Mid Experience driving progress, solving...  ..., distributed systems or networks, or experience with compute technologies...  ...include Googlers, Googler Cloud customers, and billions of Google... 
    Cloud
    Senior
    Network
    Full time
    Worldwide

    Google Inc.

    Sunnyvale, CA
    5 days ago
  • $203.3k - $305.6k

     ...Summary The Apple Services Engineering team is one of the...  .... The Apple Cloud Services Infrastructure team is seeking a senior engineering program manager...  ...development of our AI platform software stack. You’ll...  ...building compute, storage, networking, and platform services... 
    Cloud
    Senior
    Network
    Relocation package
    Flexible hours

    Apple

    Cupertino, CA
    4 days ago
  • NVIDIA is hiring engineers to scale up its AI Infrastructure. We expect you to have a strong programming...  ..., operations, and networking, familiarity with...  ...systems on a comprehensive platform that automates GPU asset...  ...lifecycle management across cloud providers. Implement monitoring... 
    Cloud
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...workloads. We are looking for a Senior Software Engineer to lead the bring-up,...  ...across NVIDIA GPU platforms at the largest scales we run...  ...large‑scale AI clusters, infrastructure, and end‑to‑end workloads,...  ...performance across compute, memory, networking, and communication layers... 
    Cloud
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    1 day ago
  • Netris is a leading Network Automation,...  ...Multi-Tenancy (NAAM) platform powering next-generation AI infrastructure operators worldwide. Neoclouds, GPU cloud providers, Telcos,...  ...highly technical engineering team solving complex...  ...for an exceptional Senior DevOps/Integration... 
    Cloud
    Senior
    Network
    Worldwide

    Netris

    Santa Clara, CA
    1 day ago
  • NVIDIA is seeking a Senior AI Infrastructure Engineer to design, build, and maintain large-scale production systems for the DGX Cloud group in California. The role covers software and systems...  ...engineering practices across Linux, networks, storage, and cloud-native tooling... 
    Cloud
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $174k - $253k

     ...experience developing large-scale infrastructure, distributed systems or networks, or experience with compute...  ...About the job Google's software engineers develop the next-generation technologies...  ...the next generation of Google platforms, we make Google's product portfolio... 
    Cloud
    Senior
    Network

    Google

    Sunnyvale, CA
    1 day ago
  • $174k - $253k

     ...Job Google's software engineers develop the next-...  ...scale system design, networking and data storage, security...  ...team develops Google Cloud’s managed service for...  ...focus on building a platform that allows...  ...software and Google’s core infrastructure. Our mission is to solve... 
    Cloud
    Senior
    Network

    Google

    Sunnyvale, CA
    1 day ago
  • $159k - $231k

    Senior Robotics Automation Engineer, Platforms Infrastructure This is a specialized role which requires physical interaction...  ...customers include Googlers, Google Cloud customers, and billions of...  ...for Google Cloud, Google Global Networking, Data Center operations, systems... 
    Cloud
    Senior
    Network
    Contract work
    Worldwide

    Google

    Sunnyvale, CA
    1 day ago
  • $236k - $330k

     ...development and processing of engineering hardware must be...  .../wheeled humanoid platforms for flexible, human-...  ...latency tele-operation infrastructure, haptic feedback...  ...include Googlers, Google Cloud customers, and...  ...Cloud, Google Global Networking, Data Center operations... 
    Cloud
    Senior
    Network
    Contract work
    Remote work
    Worldwide
    Flexible hours

    Google

    Sunnyvale, CA
    1 day ago
  • $165k - $242k

     ...CoreWeave is The Essential Cloud for AI™. Built for...  ...CoreWeave delivers a platform of technology, tools,...  ...combines superior infrastructure performance with deep...  ...platform that gives network engineers, fleet engineers, and...  ...hundreds of sites. As a senior backend engineer on... 
    Cloud
    Senior
    Network
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    more than 2 months ago
  • $147k - $237.5k

    Palo Alto Networks in Santa Clara, CA is seeking a skilled developer...  ...and Python to enhance cloud applications and deliver security...  ...collaborating with senior engineers throughout the software development...  ...understanding of cloud infrastructure. The position offers a... 
    Cloud
    Senior
    Network

    Palo Alto Networks

    Santa Clara, CA
    4 days ago
  •  ...located in Sunnyvale, California, is seeking a Principal Engineer to lead the reliability and scalability of its cloud infrastructure. The role involves ensuring operational excellence across compute, storage, and networking systems, while mentoring engineers and driving... 
    Cloud
    Senior
    Network

    Crusoe Energy Systems

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

    Joining NVIDIA’s DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative...  ...software engineer to join our team. You...  ...of AI systems. As a senior DGX Cloud AI Infrastructure...  ...underpinning NVIDIA’s AI platforms. Define meaningful and... 
    Cloud
    Senior

    NVIDIA

    Santa Clara, CA
    3 days ago
  • Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distributed...  ...MGX modular architectures, and DGX Grace Hopper platforms—combined with cloud‑native software. These...  ...distributed AI Radio Access Network (AI‑RAN) deployments, where AI... 
    Cloud
    Senior
    Network

    Intelliswift - An LTTS Company

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

     ...on the world. The DGX Cloud organization at...  ...forward‑thinking engineers tackling some of...  ...searching for a Senior Systems Software...  ...and major cloud platforms. You’ll own hard...  ...help shape how AI infrastructure runs in production...  ...as GPU Operator, Network Operator, node-feature... 
    Cloud
    Senior
    Network
    Remote work

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $208k - $327.75k

     ...components for groundbreaking AI technologies while collaborating with diverse teams and customers. Your expertise in cloud infrastructure and networking will greatly influence our product strategies, as you engage with partners and internal teams to highlight NVIDIA’s... 
    Cloud
    Senior
    Network

    PMs for Hire

    Santa Clara, CA
    1 day ago
  • You Are You are a strong platform engineer with a passion for building platforms and services that improve how complex infrastructure is observed, understood, and operated. You bring...  ..., compute infrastructure, storage, networking, cloud services, and business-critical... 
    Cloud
    Senior
    Network

    Synopsys

    Sunnyvale, CA
    1 day ago
  •  ...iPronics, we are redefining intra-datacenter networking for the AI era with our Optical Network Engine — enabling full optical bandwidth switching for next-generation AI and cloud infrastructure. We’re looking for a Senior AI Infrastructure & Networking Solutions Architect... 
    Cloud
    Senior
    Network

    iPRONICS

    Santa Clara, CA
    1 day ago
  • $176k - $276k

     ...invites applications for a Senior DevOps Platform Engineer skilled in Platform and...  ...maintaining foundational infrastructure and CI/CD systems that run...  ...systems administration, networking, and distributed systems....  ...in BareMetal and hybrid cloud (AWS, GCP, Azure) environment... 
    Cloud
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    6 days ago
  •  ...technical, creative, and Senior AI Platform Engineer to build, support, and maintain...  ...to collaborate with Cloud and AI/ML teams in a multifaceted...  ...Define and lead AI‑native infrastructure roadmaps and cross‑...  ...practices, such as identity/auth, network segmentation, supply chain... 
    Cloud
    Senior
    Network

    NVIDIA AI

    Santa Clara, CA
    1 day ago
  • $176k - $333.5k

    We are seeking a Senior Infrastructure System Software Engineer with profound expertise in High-Performance Computing...  ...and scaling sophisticated AI and cloud environments using workload...  ...system software challenges in compute, networking, and storage to enhance the overall... 
    Cloud
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    1 day ago
  • The Role We are looking for a senior backend engineer to help design, own, and build...  ...plane services behind the Netris platform. Our platform automates and manages the networking infrastructure powering modern AI factories, GPU clusters, cloud providers, and large-scale... 
    Cloud
    Senior
    Network

    Netris

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...skilled and experienced Senior DevOps Engineer to join our dynamic...  ...in managing CI/CD infrastructure. You will work on...  ...across on prem and cloud environments. Use infrastructure...  ...spanning power, networking, OS, containers, CI...  ...or similar cloud platforms for CI or compute... 
    Cloud
    Senior
    Network
    Remote work

    Nvidia Corporation

    Santa Clara, CA
    4 days ago
  • The Field Engineering organization at CoreWeave is dedicated to ensuring...  ...This team supports the infrastructure that powers the AI...  ...maintain the integrity of our cloud platform Field Engineering aligns closely...  ...up (IT service, break‑fix, network, and firmware), and... 
    Cloud
    Senior
    Network
    Contract work
    Work at office

    CoreWeave

    Sunnyvale, CA
    1 day ago
  •  ...on, action‑oriented Senior Solutions Architect...  ...execution across engineering, product, and sales...  ...debugging datacenter platforms. You will serve as...  ...performance, distributed AI infrastructure on‑prem or in the cloud built with the...  ...on GPU and networking products, directly... 
    Cloud
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    5 days ago
  •  ...GPU-based hyperscale cloud inference services....  ...with Hardware Engineering, Inference Engineering...  ...Engineering AI Cloud Infrastructure & Operations Network & Storage...  ...operational risks to senior leadership Required...  ...Hardware-centric platforms Proven ability to... 
    Cloud
    Senior
    Network

    Cerebras

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

     ...accelerated computing platforms for AI and HPC. Because...  ..., researchers, and engineers can push the boundaries...  ...Partner with NVIDIA Cloud Partners in GPU cluster design and networking and convey architecture...  ..., AI clusters, or HPC infrastructure. Ability to translate... 
    Cloud
    Senior
    Network

    NVIDIA

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Platform Engineer, Network Infrastructure - DGX Cloud. Be the first to apply!