Principal AI and ML Infra Software Engineer, GPU Clusters
$272k - $431.25kNVIDIA
We are seeking a Principal AI and ML Infra Software Engineer, GPU Clusters at NVIDIA to join our Hardware Infrastructure team. As an Engineer, you will have a pivotal role in enhancing efficiency for our researchers by implementing progressions throughout the entire stack. Your main task will revolve around collaborating closely with customers to pinpoint and address infrastructure deficiencies, facilitating groundbreaking AI and ML research on GPU Clusters. Together, we can craft potent, effective, and scalable solutions as we mold the future of AI/ML technology!
What you will be doing:
Engage closely with our AI and ML research teams to discern their infrastructure requirements and barriers, converting those insights into actionable improvements.
Proactively identify researcher efficiency bottlenecks and lead initiatives to systematically improve it. Drive the direction and long-term roadmaps for such initiatives.
Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization.
Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results.
Work closely with a variety of teams, such as researchers, data engineers, and DevOps professionals, to develop a cohesive AI/ML infrastructure ecosystem.
Keep up to date with the most recent developments in AI/ML technologies, frameworks, and successful strategies, and advocate for their integration within the organization.
What we need to see:
BS or similar background in Computer Science or related area (or equivalent experience).
15+ years of demonstrated expertise in AI/ML and HPC tasks and systems.
Hands-on experience in using or operating High Performance Computing (HPC) grade infrastructure as well as in-depth knowledge of accelerated computing (e.g., GPU, custom silicon), storage (e.g., Lustre, GPFS, BeeGFS), scheduling & orchestration (e.g., Slurm, Kubernetes, LSF), high-speed networking (e.g., Infiniband, RoCE, Amazon EFA), and containers technologies (Docker, Enroot).
Capability in supervising and improving substantial distributed training operations using PyTorch (DDP, FSDP), NeMo, or JAX. Moreover, an in-depth understanding of AI/ML workflows, involving data processing, model training, and inference pipelines.
Proficiency in programming & scripting languages such as Python, Go, Bash, as well as familiarity with cloud computing platforms (e.g., AWS, GCP, Azure) in addition to experience with parallel computing frameworks and paradigms.
Dedication to ongoing learning and staying updated on new technologies and innovative methods in the AI/ML infrastructure sector.
Excellent communication and collaboration skills, with the ability to work effectively with teams and individuals of different backgrounds.
NVIDIA offers competitive salaries and a comprehensive benefits package. Our engineering teams are growing rapidly due to outstanding expansion. If you're a passionate and independent engineer with a love for technology, we want to hear from you.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until May 1, 2026.This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.- ...RoboForce RoboForce is an AI robotics company building Physical... ...We are looking for a Senior Software Engineer to build scalable AI... ...will work across cloud systems, GPU clusters, data pipelines, and robotics... ...proficiency with C++, Python, and ML frameworks (e.g., PyTorch,...SuggestedFull timeWork at officeVisa sponsorship
$275.8k - $340.5k
...About the team: The AV ML Infra team at GM builds ML infrastructure... ...meet the unique demands of AI and ML innovation, supporting... ...the productivity of ML engineers, and drive the adoption of cutting... ...Position Overview: The Principal AI/ML Engineer will lead a growing...PrincipalLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$272k - $425.5k
Principal Software Engineer – Large-Scale LLM Memory and Storage Systems... ...for serving generative AI and reasoning models... ..., Dynamo orchestrates GPU shards, routes requests... ...across heterogeneous clusters so that many accelerators... ...storage, or ML systems infrastructure...PrincipalLocal areaRemote work$144k - $236k
...scaling LinkedIn’s AI model training, feature engineering and serving with hundreds... ...feature engineering infra for all AI use cases... ...data infra, compute software, and hardware to... ...harness the power of our GPU fleet with thousands... ...hundreds of new ML models per quarter...SuggestedFull timeFor contractorsWork experience placementWork at officeFlexible hours$123.24k - $200k
Senior / Principal AI Engineer for Business Intelligence (7063) Overview of Role As a Sr./Principal... ...0+ years of professional experience in software engineering, machine learning engineering... ...(GCP Vertex AI, AWS SageMaker, Azure ML) and containerized workflows (Docker, Kubernetes...PrincipalWork at office$165k - $242k
...is The Essential Cloud for AI™. Built for pioneers by pioneers... ...You'll Do: As a Senior Software Engineer II (IC4) on the AI... ...Improve scheduling latency, cluster utilization, and workload reliability... ...Familiarity with GPU-based workloads, ML training, or inference pipelines...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$182k - $242k
...Essential Cloud for AI™. Built for pioneers... ...for high-performance GPU infrastructure across AI/ML, visual effects,... ...inference. Our stack is engineered for speed, scale,... ...including workload setup, cluster configuration,... ...computing, GPU/accelerator software, or performance-...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$139k - $204k
...The Essential Cloud for AI™. Built for pioneers by... ...role As part of the Cluster Orchestration team, you... ...efficiently across massive GPU clusters. By building... ...You'll Do As a Senior Software Engineer I (IC3), you will own... ...based applications, or ML pipelines. Knowledge...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...We are seeking a Principal Engineer, AI Safety to lead the architecture and technical strategy for safeguarding large language models (LLMs) and multimodal AI systems from misuse. This role focuses on developing scalable defenses against jailbreaks, prompt injection attacks...Principal
$248k - $379.5k
...built on top of our GPU technology,... ...architectures and software optimizations allowing... ..., Prescriptive and AI-augmented Analytics... ...Data Scientists and Engineers working on global deployment... ...deployment of AI/ML based solutions for... ...feedback based clustering and alerting and LLM...Principal- Figure is an AI robotics company developing autonomous... ...for a Helix AI Engineer, Generative AI to build... ...model performance Solid software engineering skills and... ...training (multi-node, GPU clusters, etc.) Familiarity... ...embodied AI, or real-world ML systems Publication...Full timeWork at office
- ...Principal Ai Engineer TENEX is an AI-native, automation-first, built-for-scale Managed Detection... ...real-world attacker behavior into robust ML and rule-based detections. Push... ...Required Skills & Qualifications ~ Software Engineering & Architecture Expertise...PrincipalWork from home
$180k - $240k
...integrates advanced software and hardware powering... ...are seeking a Senior AI Infrastructure Engineer to design, build, and... ...Distributed Training & ML Systems Support... ...Architect and optimize multi-GPU setups, ensuring... ...techniques across H100/A100 clusters. Networking &...Odd jobWork at office- ...aligning with company goals. Identify and capitalize on emerging AI trends in automated shopping, voice commerce, and... ...digital transformation initiatives. Experience working with AI/ML models, personalization engines, and automation tools. Strong leadership, cross-functional...Full time
- ...combining cutting-edge AI with automotive-grade hardware... ...About the Role As a software engineering intern, you will work... ..., Onboard Systems, ML Infrastructure,... ...Infrastructure: The ML Infra team is the accelerator... ...compute modalities (CPU, GPU, FPGA) etc. You have...Internship
- Build and ship AI-powered product features using LLMs, NLP, and agent-based workflows... ...into working features Collaborate with ML engineers to integrate, evaluate, and... ...and scalability 3-5 years of professional software engineering experience Strong experience...Full time
$45k - $121k
...Job Title: AI Infrastructure Engineer City: San Jose State... ...AI, high-speed software-defined storage, and GPU-accelerated nodes. Your... ...workloads from on-prem AHV clusters to the public cloud.... ...datasets required for AI/ML. Kubernetes & Orchestration...Minimum wageLocal area$272k - $431.25k
...Networking Systems & Software Architecture group is solving some of AI’s hardest... ...interconnects. This Principal Architect role leads... ...communication systems—GPU-to-GPU, GPU-to-storage... ...mentoring senior engineers across the... ...~ Understanding of ML systems concepts—transformer...Principal$160k - $225k
...running the world's best data and AI infrastructure platform so our... ...platform for large‑scale GPU training and fine‑tuning. It gives... ...in the world. As a Senior Software Engineer for AI Runtime, you will play... ...high‑performance computing, or ML systems. Experience with distributed...Local area- ...NVIDIA’s Networking Systems & Software Architecture group is solving some of AI’s hardest infrastructure... ...research and production engineering! What you will be doing... ...co‑optimization with GPU, DPU, NIC, and switch teams... .... ~ Understanding of ML systems concepts—transformer...
$272k - $431.25k
...unlimited potential of AI to define the... ...in which our GPU acts as the... ...At NVIDIA, as a Principal Rack Scale... ...Infrastructure Engineer, you will build... ...development of software systems. These... ...with rack‑ or cluster‑scale systems spanning... ...firmware, and infra management as one...PrincipalShift work$269.1k - $307.2k
Distinguished AI Engineer (Agentic AI Platform)... ...applications of AI & ML are bringing... ...model minutiae or infra plumbing. You will... ...mentoring Staff, Principal and Senior engineers... ...training and inference software to improve... ...mastery (multi-region clusters, sericie mesh)...Full timePart timeWork at officeLocal area- ...About the Opportunity We are seeking a Principal Engineer with a deep expertise in autonomous AI agent architecture and deployment, to spearhead the design, development, and optimization of intelligent agent systems on our global crypto exchange platform. This is a senior...Principal
$152k - $241.5k
...the unlimited potential of AI to define the next era of computing... .... An era in which our GPU acts as the brains of... ...We are looking for a Senior Software Engineer to join our mission to continue... ...or control planes for HPC clusters, large‑scale AI/ML platforms, or systems...- What You’ll Do The Applications Engineering team in the Data Infrastructure organization builds AI-native analytics platforms and... ...CoreWeave's internal data-native software products, transforming data... ...Experience developing and deploying AI/ML/LLM-powered applications in...Permanent employmentTemporary workCasual workWork at officeFlexible hours
- ...building the foundation for physical AI — a unified platform that... ...easy to build and deploy as software. Today, robotics is fragmented... ...We are looking for a Senior AI Engineer to design, build, and ship AI-... ...and webhooks; Python for AI/ML integration and scripting ~...Full time
$184.7k - $324.8k
...Software Engineer (Applied AI) We are looking for a Staff level Software Engineer with experience working with the latest LLM's from OpenAI, Anthropic... ...robust evaluations for prompt optimization and tuning ML workflows Experience with RAG and modern model in context...Relocation- ...computing to make AI accessible to everyone... ...a new kind of software stack: a hardware-agnostic... ...like one seamless engine. Developers can... ...evolving needs of ML engineers and drive... ...developing or maintaining GPU compute libraries... ...large compute clusters. Why Join Lemurian...
$165k - $242k
...CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge services... ...innovation. What You’ll Do: Senior engineers are area owners who lead designs, raise engineering... ...scheduling, cost-per-token analytics, GPU resource isolation). Who You Are: ~~5...Permanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work$185.5k - $265k
...resilient, and secure. As an AI-forward enterprise, we... ...We are looking for a Principal DevOps Engineer to join our team. This... ...to the Sr. Manager, Software Engineering in the... ...production Kubernetes clusters, minimizing human error... ...understanding of AI/ML technologies and experience...PrincipalWork at officeLocal area3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal AI and ML Infra Software Engineer, GPU Clusters. Be the first to apply!
- machine learning engineer Santa Clara, CA
- network software engineer Santa Clara, CA
- software engineer travel Santa Clara, CA
- senior robotics software engineer Santa Clara, CA
- entry level software engineer remote Santa Clara, CA
- cybersecurity software engineer Santa Clara, CA
- federal - software developer Santa Clara, CA
- agile software developer Santa Clara, CA
- financial software developer Santa Clara, CA
- software engineer Santa Clara, CA



