Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal AI and ML Infra Software Engineer, GPU Clusters

$272k - $431.25k

NVIDIA

What you will be doing: Engage closely with our AI and ML research teams to discern their infrastructure requirements and barriers, converting those insights into actionable improvements. Proactively identify researcher efficiency bottlenecks and lead initiatives to systematically improve it. Drive the direction and long‑term roadmaps for such initiatives. Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization. Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results. Work closely with a variety of teams, such as researchers, data engineers, and DevOps professionals, to develop a cohesive AI/ML infrastructure ecosystem. Keep up to date with the most recent developments in AI/ML technologies, frameworks, and successful strategies, and advocate for their integration within the organization. What we need to see: BS or similar background in Computer Science or related area (or equivalent experience). 15+ years of demonstrated expertise in AI/ML and HPC tasks and systems. Hands‑on experience in using or operating High Performance Computing (HPC) grade infrastructure as well as in‑depth knowledge of accelerated computing (e.g., GPU, custom silicon), storage (e.g., Lustre, GPFS, BeeGFS), scheduling & orchestration (e.g., Slurm, Kubernetes, LSF), high‑speed networking (e.g., Infiniband, RoCE, Amazon EFA), and containers technologies (Docker, Enroot). Capability in supervising and improving substantial distributed training operations using PyTorch (DDP, FSDP), NeMo, or JAX. Moreover, an in‑depth understanding of AI/ML workflows, involving data processing, model training, and inference pipelines. Proficiency in programming & scripting languages such as Python, Go, Bash, as well as familiarity with cloud computing platforms (e.g., AWS, GCP, Azure) in addition to experience with parallel computing frameworks and paradigms. Dedication to ongoing learning and staying updated on new technologies and innovative methods in the AI/ML infrastructure sector. Excellent communication and collaboration skills, with the ability to work effectively with teams and individuals of different backgrounds. NVIDIA offers competitive salaries and a comprehensive benefits package. Our engineering teams are growing rapidly due to outstanding expansion. If you're a passionate and independent engineer with a love for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until May 1, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. #J-18808-Ljbffr

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Principal AI and ML Infra Software Engineer, GPU Clusters in Santa Clara, CA vacancy
  • $224k - $356.5k

     ...unlimited potential of AI to define the next...  ...era in which our GPU acts as the brains...  ...team is building the software stack that makes...  ...efficiency on edge cluster configurationsProduce...  ...Science, Computer Engineering, Electrical Engineering...  ...in GPU computing, ML systems, or high-... 
    Suggested
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...computing, cloud, and AI. Whether you’re designing...  ...of large-scale AI/ML clustered infrastructure. You will...  ...team of multi-disciplined engineers that operates across industry...  ...years of professional software development experience,...  ...Kubernetes for HPC/AI (GPU operators, device... 
    Suggested
    Flexible hours

    AMD

    Santa Clara, CA
    2 days ago
  • $275.8k - $340.5k

     ...About the team: The AV ML Infra team at GM builds ML infrastructure...  ...meet the unique demands of AI and ML innovation, supporting...  ...the productivity of ML engineers, and drive the adoption of cutting...  ...Position Overview: The Principal AI/ML Engineer will lead a growing... 
    Principal
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    4 days ago
  •  ...RoboForce RoboForce is an AI robotics company building Physical...  ...We are looking for a Senior Software Engineer to build scalable AI...  ...will work across cloud systems, GPU clusters, data pipelines, and robotics...  ...proficiency with C++, Python, and ML frameworks (e.g., PyTorch,... 
    Suggested
    Full time
    Work at office
    Visa sponsorship

    RoboForce

    Milpitas, CA
    18 hours ago
  • $250k - $350k

     ...are looking for a Staff/Lead Software Engineer, AI Infrastructure, to play a critical...  ...design and implementation of GPU infrastructure, AI model...  ...environmentsCollaborate with ML, backend, and platform engineering...  ...)Background in building infra for multi-tenant SaaS, enterprise... 
    Suggested
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    1 day ago
  •  ...computing experiences—from AI and data centers, to...  ...: We are seeking a Principal Software Engineer to serve as the senior...  ...qualification on AMD Instinct™ GPU platforms. You will...  ...CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch,...  ...and large-scale cluster softwareExperience validating... 
    Principal
    Contract work
    Shift work

    AMD

    San Jose, CA
    3 days ago
  •  ...computing experiences—from AI and data centers,...  ...in enhancing GPU kernels, deep learning...  ...and advanced engineering principles to drive...  ...integrating graph compilers. Software Engineering Best...  ...compute into ML frameworks (e.g., PyTorch...  ...utilization across clusters. Compiler... 

    AMD

    Santa Clara, CA
    3 days ago
  •  ...generation computing experiences—from AI and data centers, to PCs,...  ...'re looking for a senior software engineer who combines deep systems performance...  ...who can shape software from GPU kernels through distributed...  ...CUTLASS, Thrust, CUB, NCCL), and ML framework cores such as... 
    Shift work

    AMD

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...highly skilled and motivated software engineers to join us and build AI inference systems that...  ...inference stacks, optimize GPU kernels and compilers, drive...  ...deployments on GPU clusters across clouds.Conduct and...  ...frontier for the field of ML Systems; survey recent publications... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...computing experiences—from AI and data centers, to...  ...driving a unified ROCm software stack across AMD’s...  ...silicon, system software, AI/ML frameworks, libraries,...  .... Workload Performance Engineering: Lead the profiling,...  ...EXPERIENCE: Knowledge in GPU architectures, basic knowledge... 
    Principal

    AMD

    San Jose, CA
    1 day ago
  • $152k - $241.5k

     ...computing model focused on visual and AI computing. For two decades,...  ..., with our invention of the GPU. The GPU has also shown to be...  ...is looking for Architects, Software Engineers, and AI application developers...  ...Proficiency in C++, Python and ML frameworks like LangChain, LangSmith... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $170.6k - $261.3k

     ...join us. About the team: The AV ML Infra team at GM builds end-to-end...  ...meet the unique demands of AI and ML innovation, supporting...  ...enhance the productivity of ML engineers, and drive the adoption of...  ...design and build end-to-end software products, owning everything from... 
    Full time
    Local area
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  •  ...generation computing experiences—from AI and data centers, to PCs,...  ...is looking for an influential software engineer who is passionate about...  ...performance from the lowest-level GPU kernels to large-scale distributed...  ..., or the C++/HIP/CUDA core of ML frameworks like PyTorch,... 

    AMD

    Santa Clara, CA
    15 hours ago
  •  ...computing experiences—from AI and data centers,...  ...- Infrastructure Engineer, Reinforcement...  ...APIs across large GPU fleets. You make RL...  ...and postmortems for infra incidents affecting...  ...machine learning (ML) platforms with deep...  ...training, and GPU cluster orchestrationPrior... 

    AMD

    Santa Clara, CA
    1 day ago
  • $127.1k - $185k

     ...re looking for a talented early-career engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. You'll work on software that enables the world's largest AI models to train across massive GPU clusters, developing support for communication... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  •  ...the potential of generative AI to power the transformation...  ...We are at the forefront of software and hardware innovation,...  ...3 days per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat...  ...closely with other software (ML and compilers) and hardware... 
    Principal
    3 days per week

    d-Matrix

    Santa Clara, CA
    3 days ago
  •  ...world's largest AI chip, 56 times larger...  ...faster than GPU-based hyperscale...  ...the Wafer-Scale Engine (WSE). This team...  ...labs.As a Principal SRE, you will define...  ...customers, and cluster stakeholders can...  ...delivering and running software reliably and at...  ...in AI/ML inference systems... 
    Principal
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $152k - $241.5k

     ...decades. Our invention of the GPU in 1999 sparked the growth of...  ...deep learning ignited modern AI - the new era of computing, positioning...  ...AI seeks a Senior Systems Software Engineer interested in solving client-...  ...and applications using ML/DL frameworks, including ONNX... 
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    3 days ago
  • Principal AI EngineerLocation: Hybrid, Santa Clara Industry: Medical Device REQUIRED: proficiency...  ...For7+ years of experience in AI/ML engineering or related fieldsMust be very hands on...  ...and agent-based tools to enhance software development workflowsArchitect and deploy... 
    Principal

    Real Staffing Group

    Santa Clara, CA
    2 days ago
  •  ...experiences—from AI and data centers...  .... THE ROLE:As a Principal AI Infrastructure Solution Engineer, you will partner with AMD’s AI software teams and...  ...strong expertise in GPU‑accelerated computing...  ...AMD GPU clusters for distributed...  ...checkpointing)Implemented ML observability... 
    Principal

    AMD

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...Always-On, low-overhead GPU profiling service...  ..., scales across cluster environments, and delivers...  ...insights for ML workloads. You will...  ...across system software, drivers, and CUDA...  ...integrate with existing ML/AI workflows (e.g.,...  ...technical direction for an engineering team; mentor... 
    Principal
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...seeking highly skilled Applied AI Engineer (Software Engineer) to build...  ...building predictive and deploying ML/AI solutions for complex PCB...  ...regression, classification, clustering, time series, GNNs, reinforcement...  ...or custom RL environments. GPU training, distributed... 

    Advantest

    San Jose, CA
    3 days ago
  • $270k - $340k

     ...and celebrates all of our team members.What You’ll Do:As a Principal AI and ML fundamentalist who is an expert at developing cutting-edge AI...  ...evaluation, and deployment. Collaborate with other researcher engineers to prototype and validate complex solutions from academic... 
    Principal
    Local area

    Archer Aviation

    San Jose, CA
    5 days ago
  • $200k - $220k

     ...passionate, and committed engineers, technologists, and...  ...an experienced AI Network Software Solution Architect to...  ...requires deep expertise in GPU fabric design, high-speed...  ...scale out, internal cluster traffic and external ingress...  ...infrastructure for AI/ML workloads.... 
    Principal
    Worldwide

    Supermicro

    San Jose, CA
    3 days ago
  • $272k - $431.25k

     ...for serving generative AI and reasoning models...  ..., Dynamo orchestrates GPU shards, routes requests...  ...across heterogeneous clusters so that many accelerators...  ....We are seeking a Principal Systems Engineer to define the vision and...  ...performance storage, or ML systems infrastructure... 
    Principal
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    We're looking for outstanding AI systems engineers to develop groundbreaking...  ...technologies in the inference systems software stack! We build innovative AI...  ..., code generators, and GPU kernel technologies for NVIDIA...  .../ industry) experience with ML/DL systems development preferableStrong... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will have a pivotal...  ...to deliver efficiency in our usage of hardware, software, and infrastructure Proactively monitor fleet wide... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $127.1k - $185k

     ...AWS. Our org spans silicon engineering, hardware design and verification, software, and operations. We've...  ...Inferentia and Trainium ML Accelerators, and scalable...  ...analyze ML workloads on custom AI accelerators. You'll work...  ...architectures (CPU, NPU, GPU).Amazon is an equal... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $165.2k - $223.6k

     ...talented scientists and engineers to innovate on...  ...develop generative AI for shopping. On a...  ...reward optimization) at cluster scaleImprove RL...  ...training efficiency (GPU utilization,...  ...internship professional software development experience...  ...- Knowledge of ML frameworks including... 
    Internship
    Local area
    Flexible hours

    Amazon

    Palo Alto, CA
    1 day ago
  • $248k - $379.5k

     ...built on top of our GPU technology,...  ...architectures and software optimizations allowing...  ..., Prescriptive and AI-augmented Analytics...  ...Data Scientists and Engineers working on global deployment...  ...deployment of AI/ML based solutions for...  ...feedback based clustering and alerting and LLM... 
    Principal
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal AI and ML Infra Software Engineer, GPU Clusters. Be the first to apply!