Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior DGX Cloud AI Infrastructure Software Engineer

$184k - $287.5k

NVIDIA

Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This team focuses on developing tools for optimizing efficiency and resiliency of AI workloads - pre-training, post-training, inference. Our objective is to deliver a stable, scalable environment for AI researchers, providing them with the necessary resources and scale to foster innovation. We are seeking an AI infrastructure software engineer to join our team. You'll be instrumental in designing, building, and maintaining AI infrastructure that enable large-scale AI training and inferencing. The responsibilities include implementing software and systems engineering practices to ensure high efficiency and availability of AI systems.As a senior DGX Cloud AI Infrastructure software engineer at NVIDIA, you will have the opportunity to work on innovative technologies that power the future of AI and data science and be part of a dynamic, diverse, and supportive team that values learning and growth. The role provides the autonomy to work on meaningful projects with the support and mentorship needed to succeed, and contributes to a culture of blameless postmortems, iterative improvement, and risk-taking. If you are seeking an exciting and rewarding career that makes a difference, we invite you to apply now!What you’ll be doing:Develop infrastructure software and tools for large-scale pre-training, post-training, and inference.Develop and optimize tools and libraries to improve infrastructure efficiency and resiliency.Co-design and implement APIs for integration with NVIDIA's resiliency stacks.Enhance infrastructure and products underpinning NVIDIA's AI platforms.Define meaningful and actionable reliability metrics to track and improve system and service reliability.Skilled in problem-solving, root cause analysis, and optimization.Root cause and analyze and triage failures from the application level to the hardware levelWhat we need to see:Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems.Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience).Strong debugging skills and experience in analyzing and triaging AI applications from the application level to the hardware level.Experience with observability platforms for monitoring and logging (e.g., ELK, Prometheus, Loki).Proven track record in building and scaling large-scale distributed systems.Experience with AI training and inferencing infrastructure services.Proficiency in programming languages such as Python, C/C++, script languagesExperience in quality software engineering practices, including test development, defensive programming, version control, and CI.Excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential.Ways to stand out from the crowd:Background in working with the large scale clustersExperience in defining and building observability and telemetry software stackExperience with RDMA software stack (NCCL, IB verbs, ucx, libfabrics)Experience and root cause analysis of failures and datacenter scaleGood understanding on DL frameworks internal PyTorch, TensorFlow, JAX, and RayNVIDIA leads the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until April 6, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR, Remote; US, WA, Remote; US, WA, RedmondType: Full time

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior DGX Cloud AI Infrastructure Software Engineer in Santa Clara, CA vacancy
  • $184k - $287.5k

     ...into the unlimited potential of AI to define the next era of...  ...reliability systems? Join NVIDIA as a Senior Software Engineer - Resilience Engineering, DGX Cloud, and be a pivotal part of a...  ...operating GPU, HPC, or AI training infrastructure with outstanding failure modes.... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    We are looking for a Senior Software Engineer to join our DGX Cloud team and build the foundational systems that drive NVIDIA’s high-performance GPU infrastructure. You will play a critical role in designing...  ...aspects of the NVIDIA AI/ML software stack (e.g., CUDA,... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA DGX Cloud is building and operating large-scale GPU infrastructure for AI research and production workloads. We are looking for Senior Software Engineers to help build the automation, tooling, and operational systems that make GPU clusters reliable, scalable, and... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA is hiring experienced software engineers to help scale up its AI Infrastructure. We expect you to have significant software engineering experience with...  ...learning.What you will be doing:You will be part of an DGX Cloud team responsible for production systems that... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $152k - $241.5k

    NVIDIA is hiring experienced software engineers to help scale up its AI Infrastructure. We expect you to have significant software engineering experience with...  ...learning.What you will be doing:You will be part of an DGX Cloud team responsible for production systems that... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $200k - $322k

     ...NVIDIA’s DGX Cloud team helps some of the most advanced AI builders in the world move from idea to production faster...  ...who enjoy translating complex infrastructure into practical outcomes.Customer...  ...reduce friction.gainsightWork across Engineering, Product, Operations, and... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

    NVIDIA is transforming how the world uses AI, cloud, and accelerated computing, and trust is at the center of that mission. Our Attestation...  ...-native platforms such as Kubernetes and containers, CI/CD, infrastructure as code, observability tools, and at least one major cloud... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means advancing the performance...  ..., networking, storage, and software stacks. We are seeking a Senior Performance Engineer to characterize workloads,...  ...our technically diverse team of infrastructure experts to unlock more efficient... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...the forefront of the generative AI revolution, building the software and systems that power the world...  ...workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking...  ...of large-scale AI clusters, infrastructure, and end-to-end workloads,... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    We are looking for a Senior Software Engineer to become part of our storage management plane team....  ...and supervise our distributed storage infrastructure. Our team is continually dedicated to...  ...recently, GPU deep learning ignited modern AI — the next era of computing — with... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

     ...the unlimited potential of AI to define the next era of computing...  ...workloads worldwide across cloud, neocloud, and on-prem...  ...looking for a hands-on Storage Software Engineer to join the storage team as...  ...file systems on GPU infrastructure, and help operators and internal... 
    Cloud
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    1 day ago
  • $168k - $264.5k

    NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure. The goal is to craft a reliable, scalable...  ...network to support NVIDIA software development workflows and tools,...  ...an existing vacancy. NVIDIA uses AI tools in its recruiting processes... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $168k - $264.5k

     ...seeking a Sr Network Security Engineer to implement and maintain...  ...security across on-premise and cloud environments - enabling business...  ...span Graphics Drivers to AI and Deep Learning. In this role...  ...management of critical security infrastructure, including next-generation... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $176k - $276k

    Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate...  .... We build software and automation to standardize...  ...for a hands-on senior engineer to own the lifecycle...  .... NVIDIA uses AI tools in its recruiting... 
    Cloud
    Senior
    Full time
    Remote work
    Weekend work

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $168k - $264.5k

     ...play. NVIDIA is seeking a Senior Network Deployment Engineer to help build and scale our...  ...procuring and delivering infrastructure and circuits to support Point...  ...plusExperience with multi-cloud networking (AWS VPC, Azure...  ...vacancy. NVIDIA uses AI tools in its recruiting processes... 
    Cloud
    Senior
    Full time
    Contract work
    Work experience placement
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $136k - $224.25k

    NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA, from Graphics...  ...vacancy. NVIDIA uses AI tools in its recruiting... 
    Cloud
    Senior
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $176k - $276k

    Production engineering is a field that involves crafting...  ...various areas, including software and systems...  ...along with open-source cloud-enabling technologies...  ...data access for HPC and AI/ML workloads.Storage Production...  ...production storage infrastructure by supervising availability... 
    Cloud
    Senior
    Full time
    Flexible hours

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...advanced multi-rack, multi-tenant AI/ML datacenters with NVIDIA GB200,...  ...upcoming GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider) Engagements team...  ...(Prometheus, OpenTelemetry), and infrastructure-as-code.Excellent communication-able... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...already made a significant impact in the AI and Software-Defined Networking fields which are...  ...the world’s largest Internet companies, Cloud Service Providers (CSPs) and AI Factories...  ...opportunities. This customer-facing engineering role is for an expert in AI networking... 
    Cloud
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $174k - $252k

     ...networking features to enable AI/ML workloads in GDC...  ...for hybrid/multi cloud, involving data plane...  ...years of experience with software development in Go, C...  ...developing large-scale infrastructure, distributed systems...  ...OSs.Google's software engineers develop the next-generation... 
    Cloud
    Senior

    Google

    Sunnyvale, CA
    9 hours ago
  • $174k - $253k

     ...ML areas, leverage ML infrastructure, and demonstrate expertise...  ..., or launching software products, and 1 year of...  ...technologies.Google's software engineers develop the next-...  ...forward.The AI and Infrastructure team...  ...include Googlers, Google Cloud customers, and billions... 
    Cloud
    Senior
    Worldwide

    Google

    Sunnyvale, CA
    9 hours ago
  • $174k - $253k

     ...years of experience with software development in C++, C,...  ...large-scale infrastructure, distributed systems or...  ...technologies. Google's software engineers develop the next-...  ...software solutions.The AI and Infrastructure team...  ...Googlers, Google Cloud customers, and billions... 
    Cloud
    Senior
    Worldwide

    Google

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

    We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline...  ...supports structural validation, AI-driven runtime behavioral testing,...  ...NVCF (NVIDIA Cloud Functions) or DGX Cloud.Deep familiarity with Isaac... 
    Cloud
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    4 days ago
  • $174k - $253k

     ...maintaining, or launching software products, and 1 year...  ...large-scale infrastructure, distributed systems or...  ...technologies. Google's software engineers develop the next-...  ...software solutions.The AI and Infrastructure team...  ...include Googlers, Googler Cloud customers, and... 
    Cloud
    Senior
    Worldwide

    Google

    Sunnyvale, CA
    9 hours ago
  • $296k - $346k

     ...Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI...  ...home day is currently Tuesday.About the RoleAs a Senior Software Engineer on Lambda’s Core Cloud Platform team, you will build... 
    Cloud
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $155.5k - $315k

    Senior Full Stack Software Engineer (Network)This role has been designed as ‘Hybrid’ with an expectation that...  ...Packard Enterprise is the global edge-to-cloud company advancing the way people live...  ...Linux systems.Prior experience using AI for code writing. Experience in... 
    Cloud
    Senior
    Full time
    Work experience placement
    Work at office
    2 days per week

    Hewlett Packard Enterprise

    San Jose, CA
    9 hours ago
  • $184k - $287.5k

     ...looking for a hands-on Cybersecurity Software Engineer. As part of our team, you will work on...  ...from platform and embedded software to cloud infrastructure, underpinned by safety and performance...  ...for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $132k - $165k

     ...resilient, and secure. As an AI-forward enterprise, we are...  ...data lake to power our cloud-native Zero Trust Exchange...  ....RoleWe are looking for a Senior Software Development Engineer-AI Security to join our team...  ...and implementing core infrastructure components and distributed... 
    Cloud
    Senior
    Full time
    Work at office
    Local area

    Zscaler

    San Jose, CA
    2 days ago
  • $272k - $431.25k

     ...looking for a Principal Software Engineer to join our DGX Cloud team and build the foundational...  ...’s high-performance GPU infrastructure. You will play a...  ...that fuels the future of AI and cloud computing.What...  ...mentoring, and encouraging senior engineers, elevating the... 
    Cloud
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $300k - $425k

     ...experiences is why we’re actively looking for a Senior Software Engineer, Content Platform who can drive further...  ..., and paid time off.How will I use AI at Roku?At Roku, we don’t just use AI,...  ...architecture.Experience with cloud platforms like AWS, Azure, or Google Cloud... 
    Cloud
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior DGX Cloud AI Infrastructure Software Engineer. Be the first to apply!