Senior Software Engineer, DGX Cloud AI Infrastructure
$184k - $287.5kNVIDIA
NVIDIA is at the forefront of the generative AI revolution, building the software and systems that power the world’s most advanced large language model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at the largest scales we run.In this role you will set technical direction across communication libraries, model frameworks, and inference/training stacks to ensure state-of-the-art LLM workloads run efficiently and reliably at scale. You will lead deep performance and reliability investigations on multi-GPU and multi-node deployments, define how we benchmark and qualify new platforms, and build the resilience and failure-attribution capabilities that keep large clusters productive. This is a hands-on senior individual-contributor role for an engineer who operates at the intersection of deep learning systems, GPU performance, distributed computing, and large-scale operations — and who raises the bar for the engineers around them.What you’ll be doing:Lead bring-up, validation, and debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the standard for how the team operates.Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks.Profile and optimize end-to-end workload performance across compute, memory, networking, and communication layers using tools such as Nsight Systems, NCCL tests, and custom microbenchmarks.Analyze scaling efficiency for distributed LLM workloads using data, tensor, pipeline, and expert parallelism across modern GPU clusters, and translate findings into concrete tuning guidance.Own root-cause analysis of complex failures — hangs, performance regressions, topology sensitivity in large distributed environments.Define and build the resilience and failure-attribution stack: detecting, triaging, and attributing node, fabric, and workload failures across the cluster at scale.Build repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms.Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams.Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization.Mentor engineers, drive technical standards, and act as a force multiplier across the broader performance and infrastructure organization.What we need to see:Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience).8+ years of experience developing software infrastructure for large-scale AI or HPC systems, including a track record of technical leadership.Expertise debugging and triaging AI applications across the full stack — from the application layer down to the hardware.Deep hands-on experience with NCCL, CUDA-aware distributed execution, and debugging multi-GPU and multi-node workloads at scale.Proven track record of architecting, debugging, and scaling large-scale distributed systems.Expert-level Python and C/C++ programming skills.Experience operating workloads in scheduled, containerized cluster environments.Excellent analytical, debugging, and communication skills, with the ability to influence across teams.Ways to stand out from the crowd:Demonstrated experience debugging and optimizing AI workloads at large scale.Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric).Strong knowledge of GPU cluster fabrics and topology, including NVLink, NVSwitch, PCIe, RoCE, and InfiniBand.Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms.Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative, autonomous, and love a challenge, we want to hear from you.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until October 3, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR, Remote; US, WA, Remote; US, WA, RedmondType: Full time
$184k - $287.5k
Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This... ...seeking an AI infrastructure software engineer to join our team. You'll be... ...of AI systems.As a senior DGX Cloud AI Infrastructure software...SeniorSoftwareFull timeRemote work$184k - $287.5k
NVIDIA DGX Cloud delivers AI services and endpoints for research and production workloads. We are looking for a Senior Production Engineer to build software and automation that make those services reliable... ...; and the GPU/CPU compute infrastructure where inference and...SeniorSoftwareFull time$200k - $322k
NVIDIA DGX Cloud provides the infrastructure and software platform that enables enterprises to build, train, and deploy AI at scale. As demand for accelerated computing grows, effective... ....We are looking for a Senior Software Engineer to design and build the systems...SeniorSoftwareFull time$224k - $356.5k
Joining NVIDIA's DGX Cloud AI Efficiency Team means advancing the performance... ..., networking, storage, and software stacks. We are seeking a Senior Performance Engineer to characterize workloads,... ...our technically diverse team of infrastructure experts to unlock more efficient...SeniorSoftwareFull timeRemote work$152k - $241.5k
NVIDIA is hiring experienced software engineers with kubernetes experience to help scale up its AI Infrastructure. We expect you to have significant software engineering experience... ...you will be doing:You will be part of an DGX Cloud team responsible for production systems...SeniorSoftwareFull time$152k - $241.5k
NVIDIA DGX Cloud builds and operates large-scale GPU infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering experience who have worked hands-on with bare-metal NVIDIA systems. This team builds the software and operational...SeniorSoftwarePermanent employmentFull time$168k - $264.5k
NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure. The goal is to craft a reliable, scalable... ...network to support NVIDIA software development workflows and tools,... ...an existing vacancy. NVIDIA uses AI tools in its recruiting processes...SeniorSoftwareFull time$224k - $356.5k
NVIDIA is transforming how the world uses AI, cloud, and accelerated computing, and trust is at the center of that mission. Our Attestation... ...-native platforms such as Kubernetes and containers, CI/CD, infrastructure as code, observability tools, and at least one major cloud...SeniorSoftwareFull timeRemote work$136k - $224.25k
NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA, from Graphics... ...vacancy. NVIDIA uses AI tools in its recruiting...SeniorSoftwareFull timeRemote workShift work$184k - $287.5k
NVIDIA is looking for a Senior Software Engineer in Object Storage to design, implement... ...that is critical to NVIDIA AI/ML research teams creating... ...levelsAutomating storage infrastructure end-to-end including... ...with building and delivering cloud services, with specific focus...SeniorSoftwareFull timeRemote work$174k - $252k
...years of experience with software development in C++, C,... ...large-scale infrastructure, distributed systems or... ...technologies. Google's software engineers develop the next-... ...software solutions.The AI and Infrastructure team... ...Googlers, Google Cloud customers, and billions...SeniorSoftwareWorldwide$174k - $252k
...years of experience with software development in C++, C,... ...large-scale infrastructure, distributed systems or... ...technologies. Google's software engineers develop the next-... ...experience possible.The AI and Infrastructure... ...include Googlers, Google Cloud customers, and billions...SeniorSoftwareWorldwide$224k - $356.5k
...securely and powers the largest AI workloads worldwide across cloud, neocloud, and on-prem setups. With... ...looking for a hands-on Storage Software Engineer to join the storage team as an individual... ...-performance file systems on GPU infrastructure, and help operators and internal...SeniorSoftwareFull timeWorldwide$176k - $276k
Production engineering is a field that involves crafting... ...various areas, including software and systems... ...along with open-source cloud-enabling technologies... ...data access for HPC and AI/ML workloads.Storage Production... ...production storage infrastructure by supervising availability...SeniorSoftwareFull timeFlexible hours$183k - $240k
...About The Role Antora Energy is seeking an engineer to own the real-time cloud infrastructure connecting our software systems to physical hardware. Every five minutes... ...in mind from the beginning. Advance an AI-First Software Development Lifecycle Our engineers...SeniorSoftwareRemote workFlexible hours$152k - $241.5k
...tapping into the unlimited potential of AI to define the next era of computing. An... ...impact on the world. NVIDIA seeks a driven Software Engineer to advance our Kubernetes and AI... ...Kubernetes.Join the core group working on Cloud Native technologies, improving NVIDIA accelerators...SeniorSoftwareFull timeWork experience placement$272k - $431.25k
...gaming, and advanced computing, as a Principal Software Engineer for DGX Cloud. At NVIDIA, our innovation legacy drives... .... Are you passionate about Kubernetes and AI and want to help build the best platform for ML/AI infrastructure? Do you thrive when your work directly...SoftwareFull timeWork experience placement$184k - $287.5k
NVIDIA is seeking a Senior Software Engineer to help us develop distributed storage services for AI/ML. In this role you will work closely with the broader NVIDIA team to design... ...teams, and external customers to deliver Cloud services.Automating distributed storage service...SeniorSoftwareFull time$174k - $252k
...5 years of experience with software development in C++, C, or Python... ...developing large-scale infrastructure, distributed systems or networking... .... Google's software engineers develop the next-generation... ...enhance software solutions.Google Cloud accelerates every...SeniorSoftware$184k - $287.5k
We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline... ...supports structural validation, AI-driven runtime behavioral testing,... ...NVCF (NVIDIA Cloud Functions) or DGX Cloud.Deep familiarity with Isaac...SeniorFull timeLocal area$174k - $253k
...maintaining, or launching software products, and 1 year of experience... .... Google's software engineers develop the next-generation... ...enhance software solutions.The AI and Infrastructure team is redefining what’s... ...include Googlers, Google Cloud customers, and billions of...SeniorSoftwareWorldwide$184k - $287.5k
We are looking for a Senior Software Engineer to become part of our storage management plane team.... ...and supervise our distributed storage infrastructure. Our team is continually dedicated to... ...recently, GPU deep learning ignited modern AI — the next era of computing — with...SeniorSoftwareFull time$272k - $431.25k
...into the unlimited potential of AI to define the next era of computing... ...team at NVIDIA as a Principal Software Engineer in Networking within the DGX Cloud division. Embrace this outstanding... ...advance software-defined networking infrastructure by using EVPN and BGP...SoftwareFull timeWork experience placement$184k - $287.5k
At NVIDIA, the DGX Cloud division merges fresh hardware and software innovations to offer... ...most challenging AI workloads worldwide... ...Our team of skilled engineers is committed to... ...are looking for a Senior Systems Software Engineer... ..., and cloud infrastructure. The ideal candidate...SeniorSoftwareFull timeWorldwide$262k - $364k
...coach a distributed engineering team, fostering... ...and deploy scalable software and advanced... ...RDMA, storage, and AI/ML, staying ahead... ...Machine Learning Infrastructure.Google's software... ...Google and Google Cloud .TheHPN team is at... ...Storage (GDS).As a Senior Staff Software Engineer...SeniorSoftwareRemote workWorldwide$168k - $270.25k
As a Senior Software Engineer on NVIDIA’s Global Network Visibility (GNV) team within NVIDIA's Global Network Infrastructure (GNI) organization, you will lead the platform that turns network topology... ...incident triage, and bring AI, storage, backbone, and edge infrastructure...SeniorSoftwareFull time$178k - $321k
...an early but working AI-native capability: a multi... ...harness: a resilient cloud platform, the agentic... ...governed data and AI infrastructure everything else depends... ...This is a two-person engineering team: you deploy, debug... ...workflows. Working software wins arguments; migrate...SeniorSoftware$174k - $252k
...Review code developed by other engineers and provide feedback to... ...maintaining, or launching software products, and 1 year of experience... ...software solutions.The AI and Infrastructure team is redefining what’s... ...customers include Googlers, Google Cloud customers, and billions of...SeniorSoftwareWorldwide$272k - $431.25k
NVIDIA DGX Cloud is scaling GPU infrastructure across internal, partner, and cloud environments... ...looking for Principal Software Engineers to help shape the... ...clusters.This role is for senior technical leaders who can... ...Experience with GPU clusters, AI/ML infrastructure, Kubernetes...SoftwareFull time$143.5k - $212.85k
...spanning all phases of the Software Development Lifecycle... ..., guiding junior engineers, operating with little... ...key contributor to our AI-powered tooling initiative... ...velocity of Venmo's Cloud Infrastructure and DevOps engineering... ...synthetic tests. As a Senior Engineer on the AI...SeniorSoftwareFull timeWork at officeLocal areaImmediate startFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software Engineer, DGX Cloud AI Infrastructure. Be the first to apply!
- software engineer full time Santa Clara, CA
- graduate software developer no experience Santa Clara, CA
- software engineer healthcare Santa Clara, CA
- network software engineer Santa Clara, CA
- software engineer internship Santa Clara, CA
- senior software engineer Santa Clara, CA
- software system engineer Santa Clara, CA
- software developer Santa Clara, CA
- ngo software engineer Santa Clara, CA
- startup software engineer Santa Clara, CA


