Principal AI and ML Infra Software Engineer, GPU Clusters
$272k - $431.25kNVIDIA
Principal Ai And Ml Infra Software Engineer, Gpu Clusters
We are seeking a Principal AI and ML Infra Software Engineer, GPU Clusters at NVIDIA to join our Hardware Infrastructure team. As an Engineer, you will have a pivotal role in enhancing efficiency for our researchers by implementing progressions throughout the entire stack. Your main task will revolve around collaborating closely with customers to pinpoint and address infrastructure deficiencies, facilitating groundbreaking AI and ML research on GPU Clusters. Together, we can craft potent, effective, and scalable solutions as we mold the future of AI/ML technology!
What you will be doing:
- Engage closely with our AI and ML research teams to discern their infrastructure requirements and barriers, converting those insights into actionable improvements.
- Proactively identify researcher efficiency bottlenecks and lead initiatives to systematically improve it. Drive the direction and long-term roadmaps for such initiatives.
- Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization.
- Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results.
- Work closely with a variety of teams, such as researchers, data engineers, and DevOps professionals, to develop a cohesive AI/ML infrastructure ecosystem.
- Keep up to date with the most recent developments in AI/ML technologies, frameworks, and successful strategies, and advocate for their integration within the organization.
What we need to see:
- BS or similar background in Computer Science or related area (or equivalent experience).
- 15+ years of demonstrated expertise in AI/ML and HPC tasks and systems.
- Hands-on experience in using or operating High Performance Computing (HPC) grade infrastructure as well as in-depth knowledge of accelerated computing (e.g., GPU, custom silicon), storage (e.g., Lustre, GPFS, BeeGFS), scheduling & orchestration (e.g., Slurm, Kubernetes, LSF), high-speed networking (e.g., Infiniband, RoCE, Amazon EFA), and containers technologies (Docker, Enroot).
- Capability in supervising and improving substantial distributed training operations using PyTorch (DDP, FSDP), NeMo, or JAX. Moreover, an in-depth understanding of AI/ML workflows, involving data processing, model training, and inference pipelines.
- Proficiency in programming & scripting languages such as Python, Go, Bash, as well as familiarity with cloud computing platforms (e.g., AWS, GCP, Azure) in addition to experience with parallel computing frameworks and paradigms.
- Dedication to ongoing learning and staying updated on new technologies and innovative methods in the AI/ML infrastructure sector.
- Excellent communication and collaboration skills, with the ability to work effectively with teams and individuals of different backgrounds.
NVIDIA offers competitive salaries and a comprehensive benefits package. Our engineering teams are growing rapidly due to outstanding expansion. If you're a passionate and independent engineer with a love for technology, we want to hear from you.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until May 1, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
- ...with enterprise systems. The team is responsible for delivering AI-powered search and agentic workflows that enhance... ...scalable and reusable code by enforcing best practices around software engineering architecture and processes (Code Reviews, Unit testing, etc.)...Principal
$272k - $425.5k
Principal Software Engineer – Large-Scale LLM Memory and Storage Systems... ...for serving generative AI and reasoning models... ..., Dynamo orchestrates GPU shards, routes requests... ...across heterogeneous clusters so that many accelerators... ...storage, or ML systems infrastructure...PrincipalLocal areaRemote work$165k - $242k
...is The Essential Cloud for AI™. Built for pioneers by pioneers... ...You'll Do: As a Senior Software Engineer II (IC4) on the AI... ...Improve scheduling latency, cluster utilization, and workload reliability... ...Familiarity with GPU-based workloads, ML training, or inference pipelines...SuggestedPermanent employmentTemporary workCasual workWork at officeFlexible hours$139k - $204k
...The Essential Cloud for AI™. Built for pioneers by... ...role As part of the Cluster Orchestration team, you... ...efficiently across massive GPU clusters. By building... ...You'll Do As a Senior Software Engineer I (IC3), you will own... ...based applications, or ML pipelines. Knowledge...SuggestedPermanent employmentTemporary workCasual workWork at officeFlexible hours$182k - $242k
...Essential Cloud for AI™. Built for pioneers... ...for high-performance GPU infrastructure across AI/ML, visual effects,... ...inference. Our stack is engineered for speed, scale,... ...including workload setup, cluster configuration,... ...computing, GPU/accelerator software, or performance-...SuggestedPermanent employmentTemporary workCasual workWork at officeFlexible hours$248k - $379.5k
...built on top of our GPU technology,... ...architectures and software optimizations allowing... ..., Prescriptive and AI-augmented Analytics... ...Data Scientists and Engineers working on global deployment... ...deployment of AI/ML based solutions for... ...feedback based clustering and alerting and LLM...Principal$160k - $180k
...You’ll Do As a Senior AI Systems Engineer, you will architect, deploy... ...AI researchers and Software Engineers to... ...dedicated focus on AI/ML systems, high-performance... ...centric bare-metal and GPU clouds (Nebius AI Cloud... ...paired with cloud‑agnostic cluster abstractors like...Local area$100k - $150k
...Generative AI Engineer - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI... ...candidate combines strong ML intuition with production-... ...large-scale training jobs on GPU clusters, diagnosing failures and...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship$166.7k - $283.4k
...Sr. AI Infrastructure Software Engineer C++ Focus KLA is a global leader in diversified electronics for the... ...Performance Computing - HPC (including GPU), Machine Learning, Deep Learning,... ...infrastructure components that support AI/ML workloads across multiple frameworks...Minimum wageWork experience placementFlexible hours$192k - $265k
...Vectra® is the leader in AI-driven threat detection and... ...We’re hiring an AI/ML Engineer to design, build, and deploy... ...fundamentals (classification, clustering, anomaly detection,... ...explainability techniques. Infra skills for ML (Docker, K8s, GPU scheduling, model serving...Work at officeWorldwide3 days per week$100k - $150k
...AI Systems Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and... .... The role focuses on GPU clusters, distributed training frameworks... ...developer experience for ML engineers and researchers...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship$123.24k - $200k
Overview of Role As a Sr./Principal AI Engineer within TSMC's Artificial Intelligence for Business... ...0+ years of professional experience in software engineering, machine learning... ...platform (GCP Vertex AI, AWS SageMaker, Azure ML) and containerized workflows (Docker,...PrincipalFull timeWork at office$160k - $225k
...running the world's best data and AI infrastructure platform so our... ...platform for large‑scale GPU training and fine‑tuning. It gives... ...in the world. As a Senior Software Engineer for AI Runtime, you will play... ...high‑performance computing, or ML systems. Experience with distributed...Local area$249k - $348.5k
...Principal Software Development Engineer Our Technology Team partners with teams across Expedia Group... ...to replace fragile monolithic clusters with isolated, predictable failure... ...direction. Familiarity with AI‑driven systems and applying AI/ML concepts to cloud or platform...PrincipalFlexible hours$293.6k - $335.1k
...Distinguished AI Engineer (Agentic AI Platform)... ...applications of AI & ML are bringing... ...model minutiae or infra plumbing. You will... ...mentoring Staff, Principal and Senior engineers... ...training and inference software to improve... ...mastery (multi‑region clusters, service mesh). Experience...Full timePart timeWork at officeLocal area$272k - $431.25k
...seeking a highly motivated Principal System Software Engineer to drive next-generation innovations... ..., architecture, kernel, AI, middleware, and platform... ...optimization initiatives across CPU, GPU, memory, storage, networking... ...computing and AI/ML software platforms. Contributions...Principal- ...About the Opportunity We are seeking a Principal Engineer with a deep expertise in autonomous AI agent architecture and deployment, to spearhead the design, development, and optimization of intelligent agent systems on our global crypto exchange platform. This is a senior...Principal
$165k - $242k
...CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge services... ...innovation. What You’ll Do: Senior engineers are area owners who lead designs, raise engineering... ...scheduling, cost-per-token analytics, GPU resource isolation). Who You Are: ~~5...Permanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work$204k - $225k
...Principal Software Engineer – Capella Control Plane Platform Location: San Jose, California Couchbase... ...and integration of state-of-the-art AI/ML knowledge and process execution into... ...services that orchestrate Couchbase clusters across cloud providers. Technical Standards...PrincipalWork at officeWork from home$185.5k - $265k
...resilient, and secure. As an AI-forward enterprise, we... ...We are looking for a Principal DevOps Engineer to join our team. This... ...to the Sr. Manager, Software Engineering in the... ...production Kubernetes clusters, minimizing human error... ...understanding of AI/ML technologies and experience...PrincipalWork at officeLocal area3 days per week$165k - $242k
...The Essential Cloud for AI™. Built for pioneers by... ...About the role Senior engineers are area owners who lead... ...hardware teams to evolve our GPU performance testing... ...in Go and/or Python software development. ~ Hands-on... ...Experience Experience with AI/ML infrastructure and...Permanent employmentTemporary workCasual workWork at officeFlexible hours$313.06k
...processes, and foster a friendly, rewarding, and diverse environment for every OK‑er. About the Opportunity We are looking for a Principal AI Engineer to lead the architecture and deployment of large‑scale, LLM‑powered conversational Chatbot systems serving both enterprise...Principal- ...products. As a Senior Lead Software Engineer at JPMorgan Chase... ...optimized for AI and machine learning workloads... ...containerization (Docker), including cluster operations and... ...architecture, ML training, and inference... ...understanding of NVIDIA GPU infrastructure software...For contractors
$92k - $135k
...CoreWeave is The Essential Cloud for AI™. Built for pioneers by... ...cost for model serving on our GPU platform. As an IC1, you'll implement... ...mentorship from experienced engineers. About the role: Implement... ...deployed a microservice or ML inference demo. Coursework/research...Permanent employmentTemporary workCasual workInternshipWork at officeFlexible hours- ...AI Software Engineer Intern San Jose, Hybrid At Nirmata, our mission is to accelerate adoption of cloud native technologies for enterprises... ...governance platform. About the Role: As an AI/ML Software Engineer Intern, you will define and implement the AI...Internship
$216.4k - $331.6k
..., at minimum. The Role As a Principal Software Engineer in the Vehicle AI division, you will be the technical... ...other deeply embedded systems. AI/ML Deployment: Hands-on experience... ...deploying ML models to edge hardware (NPU/GPU/DSP utilization, quantization,...PrincipalFull timeLocal areaRemote workWork from homeRelocation package$180k - $300k
...the potential of generative AI to power the transformation... ...We are at the forefront of software and hardware innovation,... ...AI Software Application Engineer – Technical lead / Principal d-Matrix is seeking an experienced... ...AI inference and AI/ML software support. In this highly...Principal$185k - $275k
...Essential Cloud for AI™. Built for... ...role As part of the Cluster Orchestration team,... ...efficiently across massive GPU clusters. By... ...'ll Do As a Staff Engineer, you will be a technical... ...Are ~8+ years of software engineering... ...based applications, or ML pipelines. Knowledge...Permanent employmentTemporary workCasual workWork at officeFlexible hours$207k - $275k
...Essential Cloud for AI™. Built for... ...role As part of the Cluster Orchestration team,... ...efficiently across massive GPU clusters. By... ...'ll Do As a Staff Engineer (IC5), you will be... ...years of professional software engineering experience... ...applications, or ML pipelines. Knowledge...Permanent employmentTemporary workCasual workWork at officeFlexible hours$227k - $300k
...the transformation to AI-enabled software-defined vehicles. Traditional... ...Senior Staff AI Engineer to join our seasoned... ...will own the end-to-end ML pipeline—from data... ...algorithms to automatically cluster log patterns and... ...for execution on CPU/GPU-bound targets or embedded...Work at officeWorldwideFlexible hoursShift work3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal AI and ML Infra Software Engineer, GPU Clusters. Be the first to apply!
- machine learning engineer Santa Clara, CA
- machine learning software engineer Santa Clara, CA
- senior ml engineer Santa Clara, CA
- computer vision machine learning engineer Santa Clara, CA
- ngo software engineer Santa Clara, CA
- software developer Santa Clara, CA
- software developer internship no experience Santa Clara, CA
- part time software developer remote Santa Clara, CA
- financial software developer Santa Clara, CA
- senior software engineer ruby on rails Santa Clara, CA


