Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Software Engineer, DGX Cloud AI Infrastructure

$184k - $287.5k

NVIDIA

NVIDIA is at the forefront of the generative AI revolution, building the software and systems that power the world’s most advanced large language model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at the largest scales we run.In this role you will set technical direction across communication libraries, model frameworks, and inference/training stacks to ensure state-of-the-art LLM workloads run efficiently and reliably at scale. You will lead deep performance and reliability investigations on multi-GPU and multi-node deployments, define how we benchmark and qualify new platforms, and build the resilience and failure-attribution capabilities that keep large clusters productive. This is a hands-on senior individual-contributor role for an engineer who operates at the intersection of deep learning systems, GPU performance, distributed computing, and large-scale operations — and who raises the bar for the engineers around them.What you’ll be doing:Lead bring-up, validation, and debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the standard for how the team operates.Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks.Profile and optimize end-to-end workload performance across compute, memory, networking, and communication layers using tools such as Nsight Systems, NCCL tests, and custom microbenchmarks.Analyze scaling efficiency for distributed LLM workloads using data, tensor, pipeline, and expert parallelism across modern GPU clusters, and translate findings into concrete tuning guidance.Own root-cause analysis of complex failures — hangs, performance regressions, topology sensitivity in large distributed environments.Define and build the resilience and failure-attribution stack: detecting, triaging, and attributing node, fabric, and workload failures across the cluster at scale.Build repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms.Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams.Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization.Mentor engineers, drive technical standards, and act as a force multiplier across the broader performance and infrastructure organization.What we need to see:Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience).8+ years of experience developing software infrastructure for large-scale AI or HPC systems, including a track record of technical leadership.Expertise debugging and triaging AI applications across the full stack — from the application layer down to the hardware.Deep hands-on experience with NCCL, CUDA-aware distributed execution, and debugging multi-GPU and multi-node workloads at scale.Proven track record of architecting, debugging, and scaling large-scale distributed systems.Expert-level Python and C/C++ programming skills.Experience operating workloads in scheduled, containerized cluster environments.Excellent analytical, debugging, and communication skills, with the ability to influence across teams.Ways to stand out from the crowd:Demonstrated experience debugging and optimizing AI workloads at large scale.Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric).Strong knowledge of GPU cluster fabrics and topology, including NVLink, NVSwitch, PCIe, RoCE, and InfiniBand.Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms.Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative, autonomous, and love a challenge, we want to hear from you.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 8, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR, Remote; US, WA, Remote; US, WA, RedmondType: Full time

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Senior Software Engineer, DGX Cloud AI Infrastructure in Santa Clara, CA vacancy
  • $184k - $287.5k

    Joining NVIDIA's DGX Cloud Lepton Team means contributing...  ...powers innovative AI research and...  ...developing scalable AI infrastructure services globally. We...  ...an AI infrastructure software engineer to join our team. You...  ...AI in production.As a senior DGX Cloud AI Infrastructure... 
    Senior
    Software
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $140k - $224.25k

    NVIDIA DGX Cloud provides the infrastructure and software platform that enables enterprises to build, train, and deploy AI at scale. As demand for accelerated computing grows, effective...  ....We are looking for a Senior Software Engineer to design and build the systems... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    6 days ago
  • $224k - $356.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means advancing the performance...  ..., networking, storage, and software stacks. We are seeking a Senior Performance Engineer to characterize workloads,...  ...our technically diverse team of infrastructure experts to unlock more efficient... 
    Senior
    Software
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    a month ago
  • $152k - $241.5k

    As part of DGX Cloud, the Attestation & Trust...  ...We want hands-on engineers with strong systems...  ...verification flows, software development kits...  ...experiments, working with senior engineers on...  ...software, SDKs, infrastructure platforms, or customer...  .... NVIDIA uses AI tools in its recruiting... 
    Senior
    Software
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $168k - $264.5k

     ...NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure. The goal is to craft a reliable, scalable...  ...network to support NVIDIA software development workflows and tools,...  ...an existing vacancy. NVIDIA uses AI tools in its recruiting processes... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $168k - $270.25k

    NVIDIA is driving AI and high-performance computing forward. DGX Cloud aims to deliver a fully...  ...-performance NVIDIA infrastructure. Work with NVIDIA's DGX Cloud team as a Senior Site Reliability Engineer to maintain high-...  ...consulting, developing software tools, platforms and... 
    Senior
    Software
    Full time
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  • $136k - $224.25k

     ...NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA, from Graphics...  ...vacancy. NVIDIA uses AI tools in its recruiting... 
    Senior
    Software
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    a month ago
  • $183k - $240k

     ...About The Role   Antora Energy is seeking an engineer to own the real-time cloud infrastructure connecting our software systems to physical hardware. Every five minutes...  ...in mind from the beginning. Advance an AI-First Software Development Lifecycle Our engineers... 
    Senior
    Software
    Remote work
    Flexible hours

    Antora Energy

    San Jose, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure for AI research and production workloads. We are looking for Senior Software Engineers to help build the automation, tooling, and operational systems that make GPU clusters reliable, scalable, and... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $143.5k - $212.85k

     ...spanning all phases of the Software Development Lifecycle...  ..., guiding junior engineers, operating with little...  ...key contributor to our AI-powered tooling initiative...  ...velocity of Venmo's Cloud Infrastructure and DevOps engineering...  ...tests. As a Senior Engineer on the AI Tooling... 
    Senior
    Software
    Full time
    Work at office
    Local area
    Immediate start
    Flexible hours

    Paypal

    San Jose, CA
    11 hours ago
  • $208k - $327.75k

     ...seeking a world-class Senior Product Manager...  ...of Enterprise AI. While the NVIDIA DGX is the...  ...invisible as the public cloud? The mission is...  ...this role, own the software-defined...  ...to self-healing infrastructure. Thoughtfully define...  ...intersection of multiple engineering fields. As you... 
    Senior
    Software
    Full time
    Night shift

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $155k - $230k

     ...data spreads across various clouds and devices, traditional...  ..., and confidential AI solutions.   As data...  ...We are looking for a Senior/Staff Infrastructure & Platform Engineer to help architect, build,...  ..., networking, CI/CD, and software development. You will help... 
    Senior
    Software
    Temporary work
    H1b
    Worldwide

    Fortanix

    Santa Clara, CA
    16 days ago
  • $159k - $230k

     ...functionally with Hardware, Software, Mechanical, Thermal,...  ...degree in Electrical Engineering, Computer Engineering,...  ...and integration.As a Senior Hardware Engineer in...  ..., you will work on ML/AI hardware systems...  ...center.Our Platforms Infrastructure Engineering team designs... 
    Senior
    Software
    Worldwide

    Google

    Sunnyvale, CA
    2 days ago
  • $143k - $197.44k

     ...ranging from Level 2 to Level 4, addressing mobility, logistics, and urban services. WeRide.ai is looking for a Software Engineer to build powerful and efficient world-leading cloud infra platforms for autonomous driving, inlcuding PaaS platforms, such as Big Data, AI,... 
    Senior
    Software

    WeRide.ai

    San Jose, CA
    23 days ago
  • $224k - $356.5k

     ...tapping into the unlimited potential of AI to define the next era of...  ...Seeking seasoned leader to manage a software engineering team for NVIDIA DGX Cloud Kubernetes (DGXC-K8s) organization...  ...networking subsystems ~ Understanding of infrastructure, networking, storage, and DevOps... 
    Software
    Temporary work

    Jobleads-US

    Santa Clara, CA
    1 day ago
  •  ...Trener is building the software foundation for a new generation...  .... We combine advanced AI, intuitive programming,...  ...build software that connects cloud infrastructure, developer platforms, and...  ...AWS footprint. We need a senior infrastructure engineer to help establish a... 
    Senior
    Software

    Trener Robotics

    San Jose, CA
    5 days ago
  • $184k - $287.5k

    We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline...  ...supports structural validation, AI-driven runtime behavioral testing,...  ...NVCF (NVIDIA Cloud Functions) or DGX Cloud.Deep familiarity with Isaac... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    13 days ago
  • $174k - $252k

     ...years of experience with software development in one or...  ...large-scale infrastructure, distributed systems or...  ...technologies.Google's software engineers develop the next-...  ...technology forward.As Senior Software Engineer, you...  ...the Google Distributed Cloud Hosted (GDCH) Node Platform... 
    Senior
    Software

    Google

    Sunnyvale, CA
    6 days ago
  • $200k - $322k

    NVIDIA is seeking a Senior Technical Program Manager...  ...Services programs for DGX Cloud. DGX Cloud powers large-scale AI infrastructure across NVIDIA, cloud service...  ...security, compliance, engineering execution, and partner...  ..., platform, and software teams.Establish program... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $184k - $287.5k

     ...We are looking for a Senior Software Engineer to become part of our storage management plane team....  ...and supervise our distributed storage infrastructure. Our team is continually dedicated to...  ...recently, GPU deep learning ignited modern AI — the next era of computing — with... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $174k - $253k

     ...maintaining, or launching software products, and 1 year of experience...  .... Google's software engineers develop the next-generation...  ...enhance software solutions.The AI and Infrastructure team is redefining what’s...  ...include Googlers, Google Cloud customers, and billions of... 
    Senior
    Software
    Worldwide

    Google

    Sunnyvale, CA
    a month ago
  • $184k - $287.5k

    NVIDIA is looking for a Senior Software Engineer in Object Storage to design, implement...  ...that is critical to NVIDIA AI/ML research teams creating...  ...levelsAutomating storage infrastructure end-to-end including...  ...with building and delivering cloud services, with specific focus... 
    Senior
    Software
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $262k - $364k

     ...coach a distributed team of engineers.Facilitate alignment and clarity...  ..., and enhance large scale software solutions.Minimum...  ...and developing large-scale infrastructure, distributed systems or networks...  ...software solutions. Google Cloud accelerates every organization... 
    Senior
    Software

    Google

    Sunnyvale, CA
    3 days ago
  • $178k - $321k

     ...an early but working AI-native capability: a multi...  ...harness: a resilient cloud platform, the agentic...  ...governed data and AI infrastructure everything else depends...  ...This is a two-person engineering team: you deploy, debug...  ...workflows. Working software wins arguments; migrate... 
    Senior
    Software

    OKX

    San Jose, CA
    a month ago
  • $159k - $230k

     ...issues.Support design engineers with debug, component...  ...test equipment. As a Senior Hardware Test...  ...sustaining efforts.The AI and Infrastructure team is redefining what...  ...include Googlers, Google Cloud customers, and...  ...build the future. From software to hardware our teams... 
    Senior
    Software
    Worldwide

    Google

    Sunnyvale, CA
    22 days ago
  • $184k - $287.5k

     ...unlimited potential of AI to define the next...  ...on the world. The DGX Cloud organization at...  ...edge hardware and software innovation to...  ...forward‑thinking engineers tackling some of the...  ...re searching for a Senior Systems Software Engineer...  ...help shape how AI infrastructure runs in production... 
    Senior
    Software
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    a month ago
  • $176k - $276k

    Production engineering is a field that involves crafting...  ...various areas, including software and systems...  ...along with open-source cloud-enabling technologies...  ...data access for HPC and AI/ML workloads.Storage Production...  ...production storage infrastructure by supervising availability... 
    Senior
    Software
    Full time
    Flexible hours

    Nvidia

    Santa Clara, CA
    2 days ago
  • $168k - $270.25k

     ...generation architecture and software for managing storage...  ...to support our engineers, architects with NVIDIA...  ...VLSI, Corporate, and AI/ML teams. You will build...  ...of NVIDIA engineering infrastructure services and various business...  ...)Experience with cloud infrastructure - AWS,... 
    Senior
    Software
    Full time

    Nvidia

    Santa Clara, CA
    6 days ago
  • $147k - $237.5k

     ...Python, etc) Solid understanding of infrastructure and cloud environments (experience with k8s infra...  ...willingness and ability to leverage AI tools and agents to improve productivity...  ...role involves working across the full software development lifecycle, including... 
    Senior
    Software

    Jobleads-US

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

    GeForce NOW is Nvidia’s Cloud Gaming service,...  ...Nvidia proprietary software, GeForce NOW transforms...  ...are looking for a Senior System Software Engineer for Cloud who sees the...  ...observability, and infrastructure automation.What we need...  ...using the latest AI tools like Codex and... 
    Senior
    Software
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Software Engineer, DGX Cloud AI Infrastructure. Be the first to apply!