Senior Software Engineer, DGX Cloud AI Infrastructure
$184k - $287.5kNVIDIA
NVIDIA is at the forefront of the generative AI revolution, building the software and systems that power the world’s most advanced large language model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at the largest scales we run.In this role you will set technical direction across communication libraries, model frameworks, and inference/training stacks to ensure state-of-the-art LLM workloads run efficiently and reliably at scale. You will lead deep performance and reliability investigations on multi-GPU and multi-node deployments, define how we benchmark and qualify new platforms, and build the resilience and failure-attribution capabilities that keep large clusters productive. This is a hands-on senior individual-contributor role for an engineer who operates at the intersection of deep learning systems, GPU performance, distributed computing, and large-scale operations — and who raises the bar for the engineers around them.What you’ll be doing:Lead bring-up, validation, and debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the standard for how the team operates.Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks.Profile and optimize end-to-end workload performance across compute, memory, networking, and communication layers using tools such as Nsight Systems, NCCL tests, and custom microbenchmarks.Analyze scaling efficiency for distributed LLM workloads using data, tensor, pipeline, and expert parallelism across modern GPU clusters, and translate findings into concrete tuning guidance.Own root-cause analysis of complex failures — hangs, performance regressions, topology sensitivity in large distributed environments.Define and build the resilience and failure-attribution stack: detecting, triaging, and attributing node, fabric, and workload failures across the cluster at scale.Build repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms.Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams.Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization.Mentor engineers, drive technical standards, and act as a force multiplier across the broader performance and infrastructure organization.What we need to see:Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience).8+ years of experience developing software infrastructure for large-scale AI or HPC systems, including a track record of technical leadership.Expertise debugging and triaging AI applications across the full stack — from the application layer down to the hardware.Deep hands-on experience with NCCL, CUDA-aware distributed execution, and debugging multi-GPU and multi-node workloads at scale.Proven track record of architecting, debugging, and scaling large-scale distributed systems.Expert-level Python and C/C++ programming skills.Experience operating workloads in scheduled, containerized cluster environments.Excellent analytical, debugging, and communication skills, with the ability to influence across teams.Ways to stand out from the crowd:Demonstrated experience debugging and optimizing AI workloads at large scale.Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric).Strong knowledge of GPU cluster fabrics and topology, including NVLink, NVSwitch, PCIe, RoCE, and InfiniBand.Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms.Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative, autonomous, and love a challenge, we want to hear from you.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 8, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR, Remote; US, WA, Remote; US, WA, RedmondType: Full time
$184k - $287.5k
Joining NVIDIA's DGX Cloud Lepton Team means contributing... ...powers innovative AI research and... ...developing scalable AI infrastructure services globally. We... ...an AI infrastructure software engineer to join our team. You... ...AI in production.As a senior DGX Cloud AI Infrastructure...SeniorSoftwareFull timeRemote work$176k - $276k
Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate... .... We build software and automation to standardize... ...for a hands-on senior engineer to own the lifecycle... .... NVIDIA uses AI tools in its recruiting...SeniorSoftwareFull timeRemote workWeekend work$184k - $287.5k
Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This... ...seeking an AI infrastructure software engineer to join our team. You'll be... ...of AI systems.As a senior DGX Cloud AI Infrastructure software...SeniorSoftwareFull timeRemote work$184k - $287.5k
...into the unlimited potential of AI to define the next era of... ...reliability systems? Join NVIDIA as a Senior Software Engineer - Resilience Engineering, DGX Cloud, and be a pivotal part of a... ...operating GPU, HPC, or AI training infrastructure with outstanding failure modes....SeniorSoftwareFull time$184k - $287.5k
...about Kubernetes and AI and want to help build... ...best platform for ML/AI infrastructure? Do you thrive when... ...team within NVIDIA's DGX Cloud organization - a collaborative... ...of cloud platform engineers, architects, and SREs... ...decommissioning. We build software in the open, with a...SeniorSoftwareFull timeWork experience placement$224k - $356.5k
Joining NVIDIA's DGX Cloud AI Efficiency Team means advancing the performance... ..., networking, storage, and software stacks. We are seeking a Senior Performance Engineer to characterize workloads,... ...our technically diverse team of infrastructure experts to unlock more efficient...SeniorSoftwareFull timeRemote work$184k - $287.5k
NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure for AI research and production workloads. We are looking for Senior Software Engineers to help build the automation, tooling, and operational systems that make GPU clusters reliable, scalable, and...SeniorSoftwareFull time$208k - $327.75k
...seeking a world-class Senior Product Manager... ...of Enterprise AI. While the NVIDIA DGX is the... ...invisible as the public cloud? The mission is... ...this role, own the software-defined... ...to self-healing infrastructure. Thoughtfully define... ...intersection of multiple engineering fields. As you...SeniorSoftwareFull timeNight shift$168k - $264.5k
NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure. The goal is to craft a reliable, scalable... ...network to support NVIDIA software development workflows and tools,... ...an existing vacancy. NVIDIA uses AI tools in its recruiting processes...SeniorSoftwareFull time$224k - $356.5k
NVIDIA is transforming how the world uses AI, cloud, and accelerated computing, and trust is at the center of that mission. Our Attestation... ...-native platforms such as Kubernetes and containers, CI/CD, infrastructure as code, observability tools, and at least one major cloud...SeniorSoftwareFull timeRemote work$136k - $224.25k
NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA, from Graphics... ...vacancy. NVIDIA uses AI tools in its recruiting...SeniorSoftwareFull timeRemote workShift work$163k - $237k
...degree in Electrical Engineering, Computer Engineering,... ...to shape the future of AI/ML hardware acceleration... ...scenarios.The AI and Infrastructure team is redefining... ...include Googlers, Google Cloud customers, and billions... ...the future. From software to hardware our teams...SeniorSoftwareWorldwide$174k - $252k
...years of experience with software development in C++, C,... ...large-scale infrastructure, distributed systems or... ...technologies. Google's software engineers develop the next-... ...software solutions.The AI and Infrastructure team... ...Googlers, Google Cloud customers, and billions...SeniorSoftwareWorldwide$174k - $252k
....5 years of experience with software development in one or more programming... ...and developing large-scale infrastructure or distributed systems.1... ..., Python).Google's software engineers develop the next-generation... ...forward.The Google Cloud AI Research team addresses AI challenges...SeniorSoftware$168k - $270.25k
...into the unlimited potential of AI to define the next era of... ...provisioning, workload management, and infrastructure monitoring. It provides all... ...! Sr Site Reliability Engineer in this role will... ...reliability engineering and/or software development roles.Fluency in...SeniorSoftwareFull timeWorldwide- Design, build, and maintain AI-powered software infrastructure APIs and services to support millions of requests across thousands of servers. Collaborate... ...a bachelor's degree and over 8 years of software engineering experience, including 2 years specifically with...SeniorSoftware
$176k - $276k
Production engineering is a field that involves crafting... ...various areas, including software and systems... ...along with open-source cloud-enabling technologies... ...data access for HPC and AI/ML workloads.Storage Production... ...production storage infrastructure by supervising availability...SeniorSoftwareFull timeFlexible hours- ...Inclusion. We weave AI into the fabric of everything... ...Networks, Secure Cloud and AI infrastructure is the foundation of... ...‑class Principal Engineer (Sr Manager‑equivalent... ...elevate our standards for software quality, and unlock... ...platforms, mentoring senior engineers and...SeniorSoftwareFull timeWork at office3 days per week
$272k - $431.25k
...looking for a Principal Software Engineer to join our DGX Cloud team and build the foundational... ...’s high-performance GPU infrastructure. You will play a... ...that fuels the future of AI and cloud computing.What... ...mentoring, and encouraging senior engineers, elevating the...SoftwareFull time$200k - $322k
...TPM) to join our NVIDIA DGX Cloud team. This is a... ...extensive experience in cloud infrastructure bring-up and relationship... ...with companies and engineering teams internally to help build AI capacity and infrastructure... ..., Infrastructure, Software teams and their leadership...SeniorSoftwareFull time$152k - $241.5k
...Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's... ...hardware-in-the-loop (HIL) infrastructure to support the robust deployment... ....Build and evolve agentic AI tools and frameworks to... ..., Docker, Kubernetes) in cloud-native or hybrid environments...SeniorSoftwareFull time$124k - $195.5k
As a member of our NICo Software Development Team, you will be responsible... ...lifecycle, including using AI to automate the process of... ...BS or MS degree in Electrical Engineering, Computer Engineering or Computer... ...for advancing the state of Cloud Operations and Open-Source Cloud...SoftwareFull time$200k - $322k
...NVIDIA’s DGX Cloud team helps some of the most advanced AI builders in the world move from idea to production faster... ...who enjoy translating complex infrastructure into practical outcomes.Customer... ...reduce friction.gainsightWork across Engineering, Product, Operations, and...SeniorFull time$200k - $322k
...manager for NVIDIA's DGX Cloud. We want... ...providers and NVIDIA engineering teams, building outstanding cloud infrastructure and user experiences... ...of large programs, software engineering projects... ...alignment across senior and executive leaders... .... NVIDIA uses AI tools in its recruiting...SeniorSoftwareFull time$174k - $252k
...connectivity solution for hybrid/multi cloud, involving data plane and... ....5 years of experience with software development in C++, C, or... ...developing large-scale infrastructure, distributed systems or networking... ..., etc.).Google's software engineers develop the next-generation...SeniorSoftware$236k - $329k
...Robotics initiatives and roadmap, engineering execution with business... ...specialized liquid-cooled AI hardware infrastructure.Oversee the end-to-end... ...collaborate closely with hardware, software, and system engineers to... ...include Googlers, Google Cloud customers, and billions of...SeniorSoftwareContract workRemote workWorldwideFlexible hours$159k - $231k
...degree in Electrical Engineering, Computer Engineering,... ...Engineer within Platforms Infrastructure Engineering, you will... ...our next generation AI hardware for AL/ML, server... ...Googlers, Google Cloud customers, and billions... ...build the future. From software to hardware our teams...SeniorSoftwareWorldwide$150k - $218k
...system, firmware, and software tools.Improve the testing... ...fixture for better engineering efficiency and data accuracy... ...powerful computing infrastructures with custom-built... ...communication skills.The AI and Infrastructure... ...include Googlers, Google Cloud customers, and billions...SeniorSoftwareWorldwide$200k - $322k
NVIDIA is seeking a Senior Technical Program Manager... ...Services programs for DGX Cloud. DGX Cloud powers large-scale AI infrastructure across NVIDIA, cloud service... ...security, compliance, engineering execution, and partner... ..., platform, and software teams.Establish program...SeniorSoftwareFull time$174k - $252k
...5 years of experience with software development in C++, C, or Python... ...developing large-scale infrastructure, distributed systems or networks... .... Google's software engineers develop the next-generation... ...experience possible.Google Cloud accelerates every organization...SeniorSoftware
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software Engineer, DGX Cloud AI Infrastructure. Be the first to apply!
- cybersecurity software engineer Santa Clara, CA
- graduate software engineer Santa Clara, CA
- software developer fintech Santa Clara, CA
- new graduate software engineer Santa Clara, CA
- senior robotics software engineer Santa Clara, CA
- software engineer visa sponsorship Santa Clara, CA
- software qa engineer Santa Clara, CA
- network software engineer Santa Clara, CA
- software engineer remote Santa Clara, CA
- part time software developer remote Santa Clara, CA
