Senior Software Engineer, DGX Cloud AI Infrastructure
$184k - $356.5kNVIDIA
NVIDIA is at the forefront of the generative AI revolution, building the software and systems that power the world’s most advanced large language model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at the largest scales we run.
In this role you will set technical direction across communication libraries, model frameworks, and inference/training stacks to ensure state-of-the-art LLM workloads run efficiently and reliably at scale. You will lead deep performance and reliability investigations on multi-GPU and multi-node deployments, define how we benchmark and qualify new platforms, and build the resilience and failure-attribution capabilities that keep large clusters productive. This is a hands-on senior individual-contributor role for an engineer who operates at the intersection of deep learning systems, GPU performance, distributed computing, and large-scale operations — and who raises the bar for the engineers around them.
What you’ll be doing:
Lead bring-up, validation, and debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the standard for how the team operates.
Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks.
Profile and optimize end-to-end workload performance across compute, memory, networking, and communication layers using tools such as Nsight Systems, NCCL tests, and custom microbenchmarks.
Analyze scaling efficiency for distributed LLM workloads using data, tensor, pipeline, and expert parallelism across modern GPU clusters, and translate findings into concrete tuning guidance.
Own root-cause analysis of complex failures — hangs, performance regressions, topology sensitivity in large distributed environments.
Define and build the resilience and failure-attribution stack: detecting, triaging, and attributing node, fabric, and workload failures across the cluster at scale.
Build repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms.
Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams.
Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization.
Mentor engineers, drive technical standards, and act as a force multiplier across the broader performance and infrastructure organization.
What we need to see:
Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience).
8+ years of experience developing software infrastructure for large-scale AI or HPC systems, including a track record of technical leadership.
Expertise debugging and triaging AI applications across the full stack — from the application layer down to the hardware.
Deep hands-on experience with NCCL, CUDA-aware distributed execution, and debugging multi-GPU and multi-node workloads at scale.
Proven track record of architecting, debugging, and scaling large-scale distributed systems.
Expert-level Python and C/C++ programming skills.
Experience operating workloads in scheduled, containerized cluster environments.
Excellent analytical, debugging, and communication skills, with the ability to influence across teams.
Ways to stand out from the crowd:
Demonstrated experience debugging and optimizing AI workloads at large scale.
Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric).
Strong knowledge of GPU cluster fabrics and topology, including NVLink, NVSwitch, PCIe, RoCE, and InfiniBand.
Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms.
Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.
NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative, autonomous, and love a challenge, we want to hear from you.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until October 3, 2026.This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.$184k - $356.5k
...Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research.... ...seeking an AI infrastructure software engineer to join our team. You'll be instrumental... ...of AI systems. As a senior DGX Cloud AI Infrastructure...SeniorSoftwareFull time$116k - $224.25k
...at the forefront of the generative AI revolution, building the software and systems that power the world’s... .... We are looking for a Software Engineer focused on bring-up, triage, benchmarking... ...debug large-scale AI clusters, infrastructure, and end-to-end workloads....SoftwareFull time$176k - $220k
...for a powerful suite of software applications, helping... ...and expanding our infrastructure stack and adopting SRE... ...Have built and operated cloud infrastructure (AWS) based... ..., forward-deployed engineers, and customer success... ...Spearheaded the adoption of AI-driven development...SeniorSoftwareFor contractorsWork experience placementWork at officeLocal areaRemote workFlexible hours$152k - $205k
...Our partner is looking for a Senior Infrastructure Engineer based in United States.... ...improvement. The position combines cloud architecture, distributed... ..., and deploy production software and developer-facing tools... ...Jobgether works: We use an AI-powered matching process...SeniorSoftwareRemote workWork from homeFree visa$240k - $280k
...network for trusted trade. Our AI-powered product network... ...at Altana The Cloud Engineering team is hiring a Senior Engineering Manager to lead... ...Azure, along with the IaC (Infrastructure-as-Code), Kubernetes, and... ...maintenance ~ Strong software engineering fundamentals....SeniorSoftwareFull timeTemporary workWork experience placementRemote workFlexible hours$160k - $190k
...Our partner is looking for a Senior Software Engineer - Core Platform (Ruby/Rails... ..., frontend, databases, infrastructure, and external integrations,... ...through responsible use of AI-assisted development tools.... ...~ Experience with AWS cloud environments is preferred....SeniorSoftwareFull timeWork at officeRemote workHome office$160k - $190k
...Reports to: Engineering Manager Location: Remote US... ...for an impact-oriented Senior Software Engineer with... ..., backend, database, infrastructure) to deliver functional... ...proactive approach to AI-assisted development,... ...Experience with AWS Cloud Environments Experience...SeniorSoftwareFull timeRemote workWorldwideHome office- ...We are looking for a Senior Data Platform Engineer to lead high-... ...robust data pipelines, software codebases, and automated... ..., multi-team data infrastructure software projects from... ...data systems across cloud providers. Strong... ..., auditable AI and advanced analytics...SeniorSoftwareRemote work
$253.9k - $298.7k
...more about working at Coinbase. As a Senior Staff Software Engineer on the Data Platform team within... ...technical strategy for Coinbase's data infrastructure, spanning ingestion, transformation,... ...systems, data engineering, and AI-readiness, reporting to the Senior Director...SeniorSoftwareLocal area$135k - $180k
...Our partner is looking for a Senior Full Stack Engineer, Client Platform based in... ...least 5 years of professional software engineering experience.... ...processing, asynchronous event infrastructure, open-source contributions,... ...works: We use an AI-powered matching process to...SeniorSoftwareRemote work$152k - $205k
...partner is looking for a Senior Data Warehouse Engineer based in United... ...opportunity to shape the data infrastructure supporting a high-... ...work with modern cloud technologies... ...reviews, and other software engineering practices... ...works: We use an AI-powered matching process...SeniorSoftwareRemote workWork from homeFree visa$186.07k - $218.9k
...working at Coinbase. As a Senior Software Engineer on the Data Platform team within... ...data services spanning cloud data warehouses, data lakes... ...running on our platform infrastructure. Architect end-to-end data... ...monitoring. Utilizes generative AI responsibly, maintaining...SeniorSoftwareLocal area$130k - $165k
...expert consulting services software development, digital transformation... ...CTG is seeking a Senior Data Engineer to design, build, and maintain... ...databases . Build AI-enabled data solutions... ...vector indexes . Develop cloud infrastructure using CloudFormation (...SeniorSoftwareFull timeTemporary workRemote work$186.07k - $218.9k
...in-person working sessions called “surges.”learn more about working at Coinbase. As a Senior Software Engineer on the Platform Core Automation team, you'll build the AI infrastructure and agentic systems that automate customer support and compliance operations at...SeniorSoftwareLocal area$235k - $285k
...Our partner is looking for a Senior Staff Software Engineer, Data based in United... ...semantic layers, analytics, and AI/ML enablement. You will also help build AI-ready infrastructure supporting LLMs, agents,... ...data and analytics systems in cloud environments. ~ Deep...SeniorSoftwareLocal areaRemote workHome officeFlexible hours$107k - $148k
...Program Manager – Data Center Infrastructure to lead the development and... ...work with cross-functional engineering and delivery teams to see... ...standard business productivity software (e.g., Google Sheets, Slides... ...center infrastructure for AI, cloud, and hybrid cloud and advances...SoftwareContract workWork at officeLocal areaRemote workWorldwideShift work$142.4k - $213.6k
...lower range. The Senior Manager, Engineering & Enterprise... ...direction of applications, infrastructure, data platforms, and... ...on-premises and cloud environments. The role... ...to meet emerging AI and machine learning... ...knowledge of modern software development practices...SeniorSoftwareRemote workFlexible hours- ...Description Job Description Senior AI Engineer Austin, TX At IPT... ...FMs, wiring deterministic software systems to non-... ...certifications (e.g., Google Cloud Professional Machine Learning... ...Engineer Associate) Cloud & Infrastructure: Demonstrated experience consuming...SeniorSoftware
- ...Senior or Staff Software Engineer, Infrastructure AcuityMD is the AI platform for MedTech trusted by over 500 MedTech companies – including 16 of the top 20. Commercial... .... We deliver direct customer value through our cloud capabilities and security. We impact internal...SeniorSoftwareFull time
$130k - $170k
...partner is looking for a Senior Full Stack DevOps Engineer based in the United States... ...BSS/OSS, monitoring, and infrastructure platforms. Working closely... ...provide technical guidance on software development, DevOps,... ...Jobgether works: We use an AI-powered matching process...SeniorSoftwareRemote work$141.8k - $195k
...SAAS data observability software. Join the company that's building the telemetry infrastructure for the AI era. At Cribl, we partner... ...Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where... ...in our production cloud services and drive teams...SeniorSoftwareTemporary workRemote work- ...on premises, in the cloud, or through a hybrid... ...business value with AI. What You'll Do... ...strong relationships with engineering, product, and business... ...product leadership and senior engineering audiences... ...years in enterprise software, cloud infrastructure, or data analytics....SeniorSoftwarePermanent employmentFlexible hours
$145k - $193k
...Our partner is looking for a Senior Software Engineer, Promotions based in United... ...event-handling infrastructure such as Apache Kafka, RabbitMQ... ...RabbitMQ, AWS SQS/SNS, or Google Cloud Pub/Sub. Experience building... ...works: We use an AI-powered matching process...SeniorSoftwareRemote work$170k - $235k
...internal, external, cloud, and hybrid cloud... ..., startup engineers, and formerly frustrated... ...ecosystem. This is a senior individual... ...pieces of platform infrastructure while helping set... ...to build their own AI-powered workflows... ...years of professional software engineering...SeniorSoftwareFull timeWork at officeRemote workFlexible hours$148.7k - $199.4k
...global organization of engineers, product developers, designers... ...and most complex cloud and technology spend... ...ecosystem of enterprise software vendors. As a Senior Software Engineer, you... ...build and operate the data infrastructure, automation, and AI-driven tooling that...SeniorSoftwareFull timeWork experience placement$180k - $195k
...partner is looking for a Senior Site Reliability Engineer, NetBox Delivery... ...widely used network infrastructure platform reliable,... ...everything between core software releases and healthy... ...work hands-on with cloud infrastructure,... ...operating systems within AI-augmented...SeniorSoftwareTemporary workRemote work$160k - $200k
...steps. Our partner is looking for a Senior Software Engineer based in United States. This role... ...a strong focus on backend systems, cloud infrastructure, scalability, and reliability. You will... ...and workflows that may incorporate AI, experimentation, model evaluation,...SeniorSoftwareRemote work$110k - $140k
...Oracle Cloud Infrastructure Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established...SoftwareFull timeH1bLocal areaImmediate startRemote workVisa sponsorship$136.9k - $181.9k
...is looking for a Senior Manager, Global Digital... ...Systems Engineering based in United States... ...role spans hybrid cloud, SaaS, high-performance... ..., and emerging AI and generative AI... ...IT, cybersecurity, infrastructure, identity, cloud,... ...Experience applying software lifecycle...SeniorSoftwareRemote workFlexible hours$90k - $141k
...B2B SAAS data observability software.Join the company that’s building the telemetry infrastructure for the AI era. At Cribl, we partner with IT and Security teams at many... ...This Role Cribl is seeking Technical Support Engineers to ensure customer success by providing...SeniorSoftwareTemporary workWork experience placementRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software Engineer, DGX Cloud AI Infrastructure. Be the first to apply!
- software engineer internship remote Oregon State
- senior software engineer ruby on rails Oregon State
- software developer positions Oregon State
- agile software developer Oregon State
- rust software engineer Oregon State
- entry level software engineer remote Oregon State
- senior software engineer remote Oregon State
- senior software design engineer Oregon State
- senior software engineer Oregon State
- startup software engineer Oregon State


