Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal AI and ML Infra Software Engineer, GPU Clusters

$272k - $431.25k

NVIDIA

Principal Ai And Ml Infra Software Engineer, Gpu Clusters

We are seeking a Principal AI and ML Infra Software Engineer, GPU Clusters at NVIDIA to join our Hardware Infrastructure team. As an Engineer, you will have a pivotal role in enhancing efficiency for our researchers by implementing progressions throughout the entire stack. Your main task will revolve around collaborating closely with customers to pinpoint and address infrastructure deficiencies, facilitating groundbreaking AI and ML research on GPU Clusters. Together, we can craft potent, effective, and scalable solutions as we mold the future of AI/ML technology!

What you will be doing:

  • Engage closely with our AI and ML research teams to discern their infrastructure requirements and barriers, converting those insights into actionable improvements.
  • Proactively identify researcher efficiency bottlenecks and lead initiatives to systematically improve it. Drive the direction and long-term roadmaps for such initiatives.
  • Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization.
  • Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results.
  • Work closely with a variety of teams, such as researchers, data engineers, and DevOps professionals, to develop a cohesive AI/ML infrastructure ecosystem.
  • Keep up to date with the most recent developments in AI/ML technologies, frameworks, and successful strategies, and advocate for their integration within the organization.

What we need to see:

  • BS or similar background in Computer Science or related area (or equivalent experience).
  • 15+ years of demonstrated expertise in AI/ML and HPC tasks and systems.
  • Hands-on experience in using or operating High Performance Computing (HPC) grade infrastructure as well as in-depth knowledge of accelerated computing (e.g., GPU, custom silicon), storage (e.g., Lustre, GPFS, BeeGFS), scheduling & orchestration (e.g., Slurm, Kubernetes, LSF), high-speed networking (e.g., Infiniband, RoCE, Amazon EFA), and containers technologies (Docker, Enroot).
  • Capability in supervising and improving substantial distributed training operations using PyTorch (DDP, FSDP), NeMo, or JAX. Moreover, an in-depth understanding of AI/ML workflows, involving data processing, model training, and inference pipelines.
  • Proficiency in programming & scripting languages such as Python, Go, Bash, as well as familiarity with cloud computing platforms (e.g., AWS, GCP, Azure) in addition to experience with parallel computing frameworks and paradigms.
  • Dedication to ongoing learning and staying updated on new technologies and innovative methods in the AI/ML infrastructure sector.
  • Excellent communication and collaboration skills, with the ability to work effectively with teams and individuals of different backgrounds.

NVIDIA offers competitive salaries and a comprehensive benefits package. Our engineering teams are growing rapidly due to outstanding expansion. If you're a passionate and independent engineer with a love for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until May 1, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Principal AI and ML Infra Software Engineer, GPU Clusters in Santa Clara, CA vacancy
  •  ...with enterprise systems. The team is responsible for delivering AI-powered search and agentic workflows that enhance...  ...scalable and reusable code by enforcing best practices around software engineering architecture and processes (Code Reviews, Unit testing, etc.)... 
    Principal

    Gravity Engineering Services Pvt Ltd.

    Santa Clara, CA
    4 days ago
  • $272k - $425.5k

    Principal Software Engineer – Large-Scale LLM Memory and Storage Systems...  ...for serving generative AI and reasoning models...  ..., Dynamo orchestrates GPU shards, routes requests...  ...across heterogeneous clusters so that many accelerators...  ...storage, or ML systems infrastructure... 
    Principal
    Local area
    Remote work

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $165k - $242k

     ...is The Essential Cloud for AI™. Built for pioneers by pioneers...  ...You'll Do: As a Senior Software Engineer II (IC4) on the AI...  ...Improve scheduling latency, cluster utilization, and workload reliability...  ...Familiarity with GPU-based workloads, ML training, or inference pipelines... 
    Suggested
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    29 days ago
  • $139k - $204k

     ...The Essential Cloud for AI™. Built for pioneers by...  ...role As part of the Cluster Orchestration team, you...  ...efficiently across massive GPU clusters. By building...  ...You'll Do As a Senior Software Engineer I (IC3), you will own...  ...based applications, or ML pipelines. Knowledge... 
    Suggested
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    29 days ago
  • $182k - $242k

     ...Essential Cloud for AI™. Built for pioneers...  ...for high-performance GPU infrastructure across AI/ML, visual effects,...  ...inference. Our stack is engineered for speed, scale,...  ...including workload setup, cluster configuration,...  ...computing, GPU/accelerator software, or performance-... 
    Suggested
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    6 days ago
  • $248k - $379.5k

     ...built on top of our GPU technology,...  ...architectures and software optimizations allowing...  ..., Prescriptive and AI-augmented Analytics...  ...Data Scientists and Engineers working on global deployment...  ...deployment of AI/ML based solutions for...  ...feedback based clustering and alerting and LLM... 
    Principal

    NVIDIA

    Santa Clara, CA
    3 days ago
  • $160k - $180k

     ...You’ll Do As a Senior AI Systems Engineer, you will architect, deploy...  ...AI researchers and Software Engineers to...  ...dedicated focus on AI/ML systems, high-performance...  ...centric bare-metal and GPU clouds (Nebius AI Cloud...  ...paired with cloud‑agnostic cluster abstractors like... 
    Local area

    Archer56

    San Jose, CA
    21 hours ago
  • $100k - $150k

     ...Generative AI Engineer - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI...  ...candidate combines strong ML intuition with production-...  ...large-scale training jobs on GPU clusters, diagnosing failures and... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    San Jose, CA
    1 day ago
  • $166.7k - $283.4k

     ...Sr. AI Infrastructure Software Engineer C++ Focus KLA is a global leader in diversified electronics for the...  ...Performance Computing - HPC (including GPU), Machine Learning, Deep Learning,...  ...infrastructure components that support AI/ML workloads across multiple frameworks... 
    Minimum wage
    Work experience placement
    Flexible hours

    KLA

    Milpitas, CA
    4 days ago
  • $192k - $265k

     ...Vectra® is the leader in AI-driven threat detection and...  ...We’re hiring an AI/ML Engineer to design, build, and deploy...  ...fundamentals (classification, clustering, anomaly detection,...  ...explainability techniques. Infra skills for ML (Docker, K8s, GPU scheduling, model serving... 
    Work at office
    Worldwide
    3 days per week

    Vectra AI

    San Jose, CA
    more than 2 months ago
  • $100k - $150k

     ...AI Systems Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and...  .... The role focuses on GPU clusters, distributed training frameworks...  ...developer experience for ML engineers and researchers... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    San Jose, CA
    1 day ago
  • $123.24k - $200k

    Overview of Role As a Sr./Principal AI Engineer within TSMC's Artificial Intelligence for Business...  ...0+ years of professional experience in software engineering, machine learning...  ...platform (GCP Vertex AI, AWS SageMaker, Azure ML) and containerized workflows (Docker,... 
    Principal
    Full time
    Work at office

    TSMC

    San Jose, CA
    14 hours ago
  • $160k - $225k

     ...running the world's best data and AI infrastructure platform so our...  ...platform for large‑scale GPU training and fine‑tuning. It gives...  ...in the world. As a Senior Software Engineer for AI Runtime, you will play...  ...high‑performance computing, or ML systems. Experience with distributed... 
    Local area

    United States Digital Space LLC

    Mountain View, CA
    1 day ago
  • $249k - $348.5k

     ...Principal Software Development Engineer Our Technology Team partners with teams across Expedia Group...  ...to replace fragile monolithic clusters with isolated, predictable failure...  ...direction. Familiarity with AI‑driven systems and applying AI/ML concepts to cloud or platform... 
    Principal
    Flexible hours

    Traveltechessentialist

    San Jose, CA
    2 days ago
  • $293.6k - $335.1k

     ...Distinguished AI Engineer (Agentic AI Platform)...  ...applications of AI & ML are bringing...  ...model minutiae or infra plumbing. You will...  ...mentoring Staff, Principal and Senior engineers...  ...training and inference software to improve...  ...mastery (multi‑region clusters, service mesh). Experience... 
    Full time
    Part time
    Work at office
    Local area

    Capital One

    San Jose, CA
    4 days ago
  • $272k - $431.25k

     ...seeking a highly motivated Principal System Software Engineer to drive next-generation innovations...  ..., architecture, kernel, AI, middleware, and platform...  ...optimization initiatives across CPU, GPU, memory, storage, networking...  ...computing and AI/ML software platforms. Contributions... 
    Principal

    NVIDIA

    Santa Clara, CA
    1 day ago
  •  ...About the Opportunity We are seeking a Principal Engineer with a deep expertise in autonomous AI agent architecture and deployment, to spearhead the design, development, and optimization of intelligent agent systems on our global crypto exchange platform. This is a senior... 
    Principal

    United States Digital Space LLC

    San Jose, CA
    1 day ago
  • $165k - $242k

     ...CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge services...  ...innovation.  What You’ll Do: Senior engineers are area owners who lead designs, raise engineering...  ...scheduling, cost-per-token analytics, GPU resource isolation). Who You Are: ~~5... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Sunnyvale, CA
    more than 2 months ago
  • $204k - $225k

     ...Principal Software Engineer – Capella Control Plane Platform Location: San Jose, California Couchbase...  ...and integration of state-of-the-art AI/ML knowledge and process execution into...  ...services that orchestrate Couchbase clusters across cloud providers. Technical Standards... 
    Principal
    Work at office
    Work from home

    Couchbase

    San Jose, CA
    2 days ago
  • $185.5k - $265k

     ...resilient, and secure. As an AI-forward enterprise, we...  ...We are looking for a Principal DevOps Engineer to join our team. This...  ...to the Sr. Manager, Software Engineering in the...  ...production Kubernetes clusters, minimizing human error...  ...understanding of AI/ML technologies and experience... 
    Principal
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    3 days ago
  • $165k - $242k

     ...The Essential Cloud for AI™. Built for pioneers by...  ...About the role Senior engineers are area owners who lead...  ...hardware teams to evolve our GPU performance testing...  ...in Go and/or Python software development. ~ Hands-on...  ...Experience Experience with AI/ML infrastructure and... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    29 days ago
  • $313.06k

     ...processes, and foster a friendly, rewarding, and diverse environment for every OK‑er. About the Opportunity We are looking for a Principal AI Engineer to lead the architecture and deployment of large‑scale, LLM‑powered conversational Chatbot systems serving both enterprise... 
    Principal

    United States Digital Space LLC

    San Jose, CA
    1 day ago
  •  ...products. As a Senior Lead Software Engineer at JPMorgan Chase...  ...optimized for AI and machine learning workloads...  ...containerization (Docker), including cluster operations and...  ...architecture, ML training, and inference...  ...understanding of NVIDIA GPU infrastructure software... 
    For contractors

    J.P. Morgan

    Palo Alto, CA
    1 day ago
  • $92k - $135k

     ...CoreWeave is The Essential Cloud for AI™. Built for pioneers by...  ...cost for model serving on our GPU platform. As an IC1, you'll implement...  ...mentorship from experienced engineers. About the role: Implement...  ...deployed a microservice or ML inference demo. Coursework/research... 
    Permanent employment
    Temporary work
    Casual work
    Internship
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    29 days ago
  •  ...AI Software Engineer Intern San Jose, Hybrid At Nirmata, our mission is to accelerate adoption of cloud native technologies for enterprises...  ...governance platform. About the Role: As an AI/ML Software Engineer Intern, you will define and implement the AI... 
    Internship

    Nirmata

    San Jose, CA
    4 days ago
  • $216.4k - $331.6k

     ..., at minimum. The Role As a Principal Software Engineer in the Vehicle AI division, you will be the technical...  ...other deeply embedded systems. AI/ML Deployment: Hands-on experience...  ...deploying ML models to edge hardware (NPU/GPU/DSP utilization, quantization,... 
    Principal
    Full time
    Local area
    Remote work
    Work from home
    Relocation package

    General Motors

    Mountain View, CA
    5 days ago
  • $180k - $300k

     ...the potential of generative AI to power the transformation...  ...We are at the forefront of software and hardware innovation,...  ...AI Software Application Engineer – Technical lead / Principal d-Matrix is seeking an experienced...  ...AI inference and AI/ML software support. In this highly... 
    Principal

    d-Matrix

    Santa Clara, CA
    more than 2 months ago
  • $185k - $275k

     ...Essential Cloud for AI™. Built for...  ...role As part of the Cluster Orchestration team,...  ...efficiently across massive GPU clusters. By...  ...'ll Do As a Staff Engineer, you will be a technical...  ...Are ~8+ years of software engineering...  ...based applications, or ML pipelines. Knowledge... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    29 days ago
  • $207k - $275k

     ...Essential Cloud for AI™. Built for...  ...role As part of the Cluster Orchestration team,...  ...efficiently across massive GPU clusters. By...  ...'ll Do As a Staff Engineer (IC5), you will be...  ...years of professional software engineering experience...  ...applications, or ML pipelines. Knowledge... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    25 days ago
  • $227k - $300k

     ...the transformation to AI-enabled software-defined vehicles. Traditional...  ...Senior Staff AI Engineer to join our seasoned...  ...will own the end-to-end ML pipeline—from data...  ...algorithms to automatically cluster log patterns and...  ...for execution on CPU/GPU-bound targets or embedded... 
    Work at office
    Worldwide
    Flexible hours
    Shift work
    3 days per week

    Sonatus

    San Jose, CA
    6 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal AI and ML Infra Software Engineer, GPU Clusters. Be the first to apply!