Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal AI and ML Infra Software Engineer, GPU Clusters

$272k - $431.25k

NVIDIA

We are seeking a Principal AI and ML Infra Software Engineer, GPU Clusters at NVIDIA to join our Hardware Infrastructure team. As an Engineer, you will have a pivotal role in enhancing efficiency for our researchers by implementing progressions throughout the entire stack. Your main task will revolve around collaborating closely with customers to pinpoint and address infrastructure deficiencies, facilitating groundbreaking AI and ML research on GPU Clusters. Together, we can craft potent, effective, and scalable solutions as we mold the future of AI/ML technology!

What you will be doing:

  • Engage closely with our AI and ML research teams to discern their infrastructure requirements and barriers, converting those insights into actionable improvements.

  • Proactively identify researcher efficiency bottlenecks and lead initiatives to systematically improve it. Drive the direction and long-term roadmaps for such initiatives.

  • Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization.

  • Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results.

  • Work closely with a variety of teams, such as researchers, data engineers, and DevOps professionals, to develop a cohesive AI/ML infrastructure ecosystem.

  • Keep up to date with the most recent developments in AI/ML technologies, frameworks, and successful strategies, and advocate for their integration within the organization.

What we need to see:

  • BS or similar background in Computer Science or related area (or equivalent experience).

  • 15+ years of demonstrated expertise in AI/ML and HPC tasks and systems.

  • Hands-on experience in using or operating High Performance Computing (HPC) grade infrastructure as well as in-depth knowledge of accelerated computing (e.g., GPU, custom silicon), storage (e.g., Lustre, GPFS, BeeGFS), scheduling & orchestration (e.g., Slurm, Kubernetes, LSF), high-speed networking (e.g., Infiniband, RoCE, Amazon EFA), and containers technologies (Docker, Enroot).

  • Capability in supervising and improving substantial distributed training operations using PyTorch (DDP, FSDP), NeMo, or JAX. Moreover, an in-depth understanding of AI/ML workflows, involving data processing, model training, and inference pipelines.

  • Proficiency in programming & scripting languages such as Python, Go, Bash, as well as familiarity with cloud computing platforms (e.g., AWS, GCP, Azure) in addition to experience with parallel computing frameworks and paradigms.

  • Dedication to ongoing learning and staying updated on new technologies and innovative methods in the AI/ML infrastructure sector.

  • Excellent communication and collaboration skills, with the ability to work effectively with teams and individuals of different backgrounds.

NVIDIA offers competitive salaries and a comprehensive benefits package. Our engineering teams are growing rapidly due to outstanding expansion. If you're a passionate and independent engineer with a love for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until May 1, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Principal AI and ML Infra Software Engineer, GPU Clusters in Santa Clara, CA vacancy
  •  ...RoboForce RoboForce is an AI robotics company building Physical...  ...We are looking for a Senior Software Engineer to build scalable AI...  ...will work across cloud systems, GPU clusters, data pipelines, and robotics...  ...proficiency with C++, Python, and ML frameworks (e.g., PyTorch,... 
    Suggested
    Full time
    Work at office
    Visa sponsorship

    RoboForce

    Milpitas, CA
    1 day ago
  • $275.8k - $340.5k

     ...About the team: The AV ML Infra team at GM builds ML infrastructure...  ...meet the unique demands of AI and ML innovation, supporting...  ...the productivity of ML engineers, and drive the adoption of cutting...  ...Position Overview: The Principal AI/ML Engineer will lead a growing... 
    Principal
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $272k - $425.5k

    Principal Software Engineer – Large-Scale LLM Memory and Storage Systems...  ...for serving generative AI and reasoning models...  ..., Dynamo orchestrates GPU shards, routes requests...  ...across heterogeneous clusters so that many accelerators...  ...storage, or ML systems infrastructure... 
    Principal
    Local area
    Remote work

    NVIDIA Corporation

    Santa Clara, CA
    4 days ago
  • $144k - $236k

     ...scaling LinkedIn’s AI model training, feature engineering and serving with hundreds...  ...feature engineering infra for all AI use cases...  ...data infra, compute software, and hardware to...  ...harness the power of our GPU fleet with thousands...  ...hundreds of new ML models per quarter... 
    Suggested
    Full time
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    LinkedIn

    Mountain View, CA
    2 days ago
  • $123.24k - $200k

    Senior / Principal AI Engineer for Business Intelligence (7063) Overview of Role As a Sr./Principal...  ...0+ years of professional experience in software engineering, machine learning engineering...  ...(GCP Vertex AI, AWS SageMaker, Azure ML) and containerized workflows (Docker, Kubernetes... 
    Principal
    Work at office

    TSMC - Taiwan Semiconductor Manufacturing Company Limited

    San Jose, CA
    4 days ago
  • $165k - $242k

     ...is The Essential Cloud for AI™. Built for pioneers by pioneers...  ...You'll Do: As a Senior Software Engineer II (IC4) on the AI...  ...Improve scheduling latency, cluster utilization, and workload reliability...  ...Familiarity with GPU-based workloads, ML training, or inference pipelines... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    10 days ago
  • $182k - $242k

     ...Essential Cloud for AI™. Built for pioneers...  ...for high-performance GPU infrastructure across AI/ML, visual effects,...  ...inference. Our stack is engineered for speed, scale,...  ...including workload setup, cluster configuration,...  ...computing, GPU/accelerator software, or performance-... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    17 days ago
  • $139k - $204k

     ...The Essential Cloud for AI™. Built for pioneers by...  ...role As part of the Cluster Orchestration team, you...  ...efficiently across massive GPU clusters. By building...  ...You'll Do As a Senior Software Engineer I (IC3), you will own...  ...based applications, or ML pipelines. Knowledge... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    10 days ago
  •  ...We are seeking a Principal Engineer, AI Safety to lead the architecture and technical strategy for safeguarding large language models (LLMs) and multimodal AI systems from misuse. This role focuses on developing scalable defenses against jailbreaks, prompt injection attacks... 
    Principal

    Insight Global

    San Jose, CA
    5 days ago
  • $248k - $379.5k

     ...built on top of our GPU technology,...  ...architectures and software optimizations allowing...  ..., Prescriptive and AI-augmented Analytics...  ...Data Scientists and Engineers working on global deployment...  ...deployment of AI/ML based solutions for...  ...feedback based clustering and alerting and LLM... 
    Principal

    NVIDIA

    Santa Clara, CA
    4 days ago
  • Figure is an AI robotics company developing autonomous...  ...for a Helix AI Engineer, Generative AI to build...  ...model performance Solid software engineering skills and...  ...training (multi-node, GPU clusters, etc.) Familiarity...  ...embodied AI, or real-world ML systems Publication... 
    Full time
    Work at office

    Figure

    San Jose, CA
    1 day ago
  •  ...Principal Ai Engineer TENEX is an AI-native, automation-first, built-for-scale Managed Detection...  ...real-world attacker behavior into robust ML and rule-based detections. Push...  ...Required Skills & Qualifications ~ Software Engineering & Architecture Expertise... 
    Principal
    Work from home

    TenEx

    San Jose, CA
    5 days ago
  • $180k - $240k

     ...integrates advanced software and hardware powering...  ...are seeking a Senior AI Infrastructure Engineer to design, build, and...  ...Distributed Training & ML Systems Support...  ...Architect and optimize multi-GPU setups, ensuring...  ...techniques across H100/A100 clusters. Networking &... 
    Odd job
    Work at office

    Gatik AI

    Santa Clara, CA
    3 days ago
  •  ...aligning with company goals. Identify and capitalize on emerging AI trends in automated shopping, voice commerce, and...  ...digital transformation initiatives. Experience working with AI/ML models, personalization engines, and automation tools. Strong leadership, cross-functional... 
    Full time

    Paypal

    San Jose, CA
    1 day ago
  •  ...combining cutting-edge AI with automotive-grade hardware...  ...About the Role As a software engineering intern, you will work...  ..., Onboard Systems, ML Infrastructure,...  ...Infrastructure: The ML Infra team is the accelerator...  ...compute modalities (CPU, GPU, FPGA) etc. You have... 
    Internship

    Nuro

    Mountain View, CA
    1 day ago
  • Build and ship AI-powered product features using LLMs, NLP, and agent-based workflows...  ...into working features Collaborate with ML engineers to integrate, evaluate, and...  ...and scalability 3-5 years of professional software engineering experience Strong experience... 
    Full time

    Eightfold

    Santa Clara, CA
    1 day ago
  • $45k - $121k

     ...Job Title: AI Infrastructure Engineer City: San Jose State...  ...AI, high-speed software-defined storage, and GPU-accelerated nodes. Your...  ...workloads from on-prem AHV clusters to the public cloud....  ...datasets required for AI/ML. Kubernetes & Orchestration... 
    Minimum wage
    Local area

    Wipro

    San Jose, CA
    11 hours ago
  • $272k - $431.25k

     ...Networking Systems & Software Architecture group is solving some of AI’s hardest...  ...interconnects. This Principal Architect role leads...  ...communication systems—GPU-to-GPU, GPU-to-storage...  ...mentoring senior engineers across the...  ...~ Understanding of ML systems concepts—transformer... 
    Principal

    NVIDIA Gruppe

    Santa Clara, CA
    5 hours ago
  • $160k - $225k

     ...running the world's best data and AI infrastructure platform so our...  ...platform for large‑scale GPU training and fine‑tuning. It gives...  ...in the world. As a Senior Software Engineer for AI Runtime, you will play...  ...high‑performance computing, or ML systems. Experience with distributed... 
    Local area

    United States Digital Space LLC

    Mountain View, CA
    2 days ago
  •  ...NVIDIA’s Networking Systems & Software Architecture group is solving some of AI’s hardest infrastructure...  ...research and production engineering! What you will be doing...  ...co‑optimization with GPU, DPU, NIC, and switch teams...  .... ~ Understanding of ML systems concepts—transformer... 

    NVIDIA Gruppe

    Santa Clara, CA
    5 hours ago
  • $272k - $431.25k

     ...unlimited potential of AI to define the...  ...in which our GPU acts as the...  ...At NVIDIA, as a Principal Rack Scale...  ...Infrastructure Engineer, you will build...  ...development of software systems. These...  ...with rack‑ or cluster‑scale systems spanning...  ...firmware, and infra management as one... 
    Principal
    Shift work

    NVIDIA Corporation

    Santa Clara, CA
    5 hours ago
  • $269.1k - $307.2k

    Distinguished AI Engineer (Agentic AI Platform)...  ...applications of AI & ML are bringing...  ...model minutiae or infra plumbing. You will...  ...mentoring Staff, Principal and Senior engineers...  ...training and inference software to improve...  ...mastery (multi-region clusters, sericie mesh)... 
    Full time
    Part time
    Work at office
    Local area

    Capital One

    San Jose, CA
    11 hours ago
  •  ...About the Opportunity We are seeking a Principal Engineer with a deep expertise in autonomous AI agent architecture and deployment, to spearhead the design, development, and optimization of intelligent agent systems on our global crypto exchange platform. This is a senior... 
    Principal

    United States Digital Space LLC

    San Jose, CA
    2 days ago
  • $152k - $241.5k

     ...the unlimited potential of AI to define the next era of computing...  .... An era in which our GPU acts as the brains of...  ...We are looking for a Senior Software Engineer to join our mission to continue...  ...or control planes for HPC clusters, large‑scale AI/ML platforms, or systems... 

    NVIDIA Gruppe

    Santa Clara, CA
    1 day ago
  • What You’ll Do The Applications Engineering team in the Data Infrastructure organization builds AI-native analytics platforms and...  ...CoreWeave's internal data-native software products, transforming data...  ...Experience developing and deploying AI/ML/LLM-powered applications in... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    1 day ago
  •  ...building the foundation for physical AI — a unified platform that...  ...easy to build and deploy as software. Today, robotics is fragmented...  ...We are looking for a Senior AI Engineer to design, build, and ship AI-...  ...and webhooks; Python for AI/ML integration and scripting ~... 
    Full time

    Dexmate

    Santa Clara, CA
    1 day ago
  • $184.7k - $324.8k

     ...Software Engineer (Applied AI) We are looking for a Staff level Software Engineer with experience working with the latest LLM's from OpenAI, Anthropic...  ...robust evaluations for prompt optimization and tuning ML workflows Experience with RAG and modern model in context... 
    Relocation

    Apple

    Cupertino, CA
    2 days ago
  •  ...computing to make AI accessible to everyone...  ...a new kind of software stack: a hardware-agnostic...  ...like one seamless engine. Developers can...  ...evolving needs of ML engineers and drive...  ...developing or maintaining GPU compute libraries...  ...large compute clusters. Why Join Lemurian... 

    Lemurian Labs Inc.

    Santa Clara, CA
    12 hours ago
  • $165k - $242k

     ...CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge services...  ...innovation.  What You’ll Do: Senior engineers are area owners who lead designs, raise engineering...  ...scheduling, cost-per-token analytics, GPU resource isolation). Who You Are: ~~5... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Sunnyvale, CA
    more than 2 months ago
  • $185.5k - $265k

     ...resilient, and secure. As an AI-forward enterprise, we...  ...We are looking for a Principal DevOps Engineer to join our team. This...  ...to the Sr. Manager, Software Engineering in the...  ...production Kubernetes clusters, minimizing human error...  ...understanding of AI/ML technologies and experience... 
    Principal
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal AI and ML Infra Software Engineer, GPU Clusters. Be the first to apply!