Principal AI and ML Infra Software Engineer, GPU Clusters (Santa Clara)
$272k - $431.25kNVIDIA
What you will be doing:
- Engage closely with our AI and ML research teams to discern their infrastructure requirements and barriers, converting those insights into actionable improvements.
- Proactively identify researcher efficiency bottlenecks and lead initiatives to systematically improve it. Drive the direction and long‑term roadmaps for such initiatives.
- Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization.
- Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results.
- Work closely with a variety of teams, such as researchers, data engineers, and DevOps professionals, to develop a cohesive AI/ML infrastructure ecosystem.
- Keep up to date with the most recent developments in AI/ML technologies, frameworks, and successful strategies, and advocate for their integration within the organization.
What we need to see:
- BS or similar background in Computer Science or related area (or equivalent experience).
- 15+ years of demonstrated expertise in AI/ML and HPC tasks and systems.
- Hands‑on experience in using or operating High Performance Computing (HPC) grade infrastructure as well as in‑depth knowledge of accelerated computing (e.g., GPU, custom silicon), storage (e.g., Lustre, GPFS, BeeGFS), scheduling & orchestration (e.g., Slurm, Kubernetes, LSF), high‑speed networking (e.g., Infiniband, RoCE, Amazon EFA), and containers technologies (Docker, Enroot).
- Capability in supervising and improving substantial distributed training operations using PyTorch (DDP, FSDP), NeMo, or JAX. Moreover, an in‑depth understanding of AI/ML workflows, involving data processing, model training, and inference pipelines.
- Proficiency in programming & scripting languages such as Python, Go, Bash, as well as familiarity with cloud computing platforms (e.g., AWS, GCP, Azure) in addition to experience with parallel computing frameworks and paradigms.
- Dedication to ongoing learning and staying updated on new technologies and innovative methods in the AI/ML infrastructure sector.
- Excellent communication and collaboration skills, with the ability to work effectively with teams and individuals of different backgrounds.
NVIDIA offers competitive salaries and a comprehensive benefits package. Our engineering teams are growing rapidly due to outstanding expansion. If you're a passionate and independent engineer with a love for technology, we want to hear from you.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until May 1, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
#J-18808-Ljbffr- ...Advanced Micro Devices is looking for a Principal Engineer in Santa Clara, CA to lead AI infrastructure development, define GPU architecture specifications, and drive performance gains in ML systems. The role involves leading innovative techniques, collaborating with...PrincipalFull time
$272k - $431.25k
NVIDIA is seeking an experienced professional to engage closely with AI and ML research teams to improve infrastructure and researcher efficiency. Candidates should have over 15 years of expertise in AI/ML and HPC tasks, hands-on experience with HPC infrastructure, and...SuggestedFull time- ...NVIDIA is hiring a Senior Solutions Architect located in Santa Clara, California, to work with its Cloud Partners. This position involves designing next-generation GPU clusters for cutting-edge AI supercomputers and enterprise AI infrastructures. The role requires extensive...SuggestedFull time
$272k - $425.5k
Principal Software Engineer – Large-Scale LLM Memory and Storage... ...: US, CA, Santa Clara: US, WA, Remote:... ...serving generative AI and reasoning models... ...Dynamo orchestrates GPU shards, routes... ...across heterogeneous clusters so that many... ...performance storage, or ML systems...PrincipalFull timeLocal areaRemote work- ...Oracle is seeking a Principal AI Agent / ML Software Engineer in Santa Clara, California, to provide technical leadership in developing next-generation AI systems on Oracle Cloud Infrastructure. The ideal candidate will have extensive experience in building scalable AI...PrincipalFull time
$272k - $431.25k
...NVIDIA is seeking expert engineers to design next generation rack level solutions for scaling AI supercomputing platforms. With a focus on fleet management solutions... ...and API management. The position is based in Santa Clara, California, with a competitive salary between...PrincipalFull time- ...experiences—from AI and data centers,... ...THE ROLE As a Principal Engineer, you will spearhead... ...infrastructure by defining GPU architecture... ...for distributed ML systems, you will... ...passionate about software engineering and possess... ...Austin, Tx or Santa Clara, Ca strongly preferred...PrincipalFull timeRemote work
$99.6k - $234.6k
...Job Description The Principal AI Agent / ML Software Engineer is a Senior Staff-level, hands‑on technical leadership role responsible for defining, building... ...services optimized for low latency, high throughput, GPU efficiency, reliability, cost, operability, and secure...PrincipalFull timeTemporary workFlexible hours$190k - $280k
...MixMode is seeking a Principal Architect in Santa Clara to accelerate AI application performance through innovative hardware and software solutions. The role involves analyzing ML workloads and collaborating with product and hardware teams to enhance inference accelerators...PrincipalFull time$241.8k - $409.2k
GPGPU Software Architect/ Principal Engineer XPENG is a leading smart technology company at... ...innovation, integrating advanced AI and autonomous driving... ...towards General Purpose GPU (GPGPU) architecture. We'... ...regulations. Location Santa Clara, CA Experience & Seniority...PrincipalFull time$272k - $431.25k
...Networking Systems & Software Architecture group is solving some of AI’s hardest... ...interconnects. This Principal Architect role leads... ...communication systems—GPU-to-GPU, GPU-to-storage... ...mentoring senior engineers across the... ...~ Understanding of ML systems concepts—transformer...PrincipalFull time- ...potential of generative AI to power the... ...the forefront of software and hardware innovation... ...: Hybrid (Santa Clara, CA) or Remote... ...Security Architect (Principal) d-Matrix is seeking... ...research in ML, architecture, and... ...working with the engineering teams to incorporate...PrincipalFull timeRemote work
$184k - $287.5k
...Senior Software Engineer, Cloud-Native Stack – CSP Engagements... ...Apply locations US, CA, Santa Clara US, TX, Austin US, WA,... ...-rack, multi-tenant AI/ML datacenters with NVIDIA... ...-rack, multi-tenant clusters: scheduler behavior, container... ...that expose new GPU capabilities. Drive...Full time- ...NVIDIA is seeking a Sr. Principal Systems Software Engineer in Santa Clara to lead development of GPU-accelerated data processing for the Apache Spark ecosystem. You will design and implement Java, Scala, and CUDA/C++ libraries to speed up DataFrames, I/O, and interoperable...PrincipalFull time
$195k - $292k
...Ampere Computing LLC. is seeking a Principal AI Accelerator Software Engineer-Graph Optimization in Santa Clara, California. The role focuses on optimizing computational graphs to unlock the full potential of Ampere's deep learning hardware. Responsibilities include collaborating...PrincipalFull time- NVIDIA is seeking a Senior Software Architect for its GPU Networking Architecture team in Santa Clara, California. In this role, you will define architectural solutions for Software Defined Networking in AI networks, collaborating closely with teams across GPU and Software...Full time
- ...NVIDIA's research team in Santa Clara is seeking a Systems Software Engineer to tackle AI infrastructure challenges. You'll architect communication and memory management... ...movement across systems, and work closely with GPU and networking teams. The ideal candidate has 12+...Full time
$224k - $336k
...is looking for a Principal Application Support Engineer to join our AI Networking Infrastructure... ...hardware and software technologies... ...of large-scale AI clusters. The Person The... ...Familiarity with GPU operations, Collective... ...field Location Santa Clara, CA or Austin, TX...PrincipalFull time- ...seeking a Senior Solutions Architect to join the Cluster Design and Architecture team with a focus on... ...involves designing and optimizing large-scale AI/HPC GPU clusters, advising on topology, and collaborating with engineering and field teams to meet demanding customer...Full time
$272k - $431.25k
...unlimited potential of AI to define the... ...in which our GPU acts as the... ...At NVIDIA, as a Principal Rack Scale... ...Infrastructure Engineer, you will build... ...development of software systems. These... ...with rack‑ or cluster‑scale systems spanning... ...firmware, and infra management as one...PrincipalFull timeShift work$272k - $431.25k
...seeking a highly motivated Principal System Software Engineer to drive next-generation innovations... ..., architecture, kernel, AI, middleware, and platform... ...optimization initiatives across CPU, GPU, memory, storage, networking... ...computing and AI/ML software platforms. Contributions...PrincipalFull time$272k - $431.25k
...NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC... ...are increasingly known as \the AI computing company.\ We are looking... .... We are looking for expert engineers to come and help design rack level... ...Compute. Experience with ML and multi‑variable optimisation...PrincipalFull time- ...and Inclusion. We weave AI into the fabric of... ...seeking a world‑class Principal Engineer (Sr Manager‑equivalent... ...elevate our standards for software quality, and unlock... ...located at our dynamic Santa Clara California headquarters... ...applying AI/ML/GenAI to solve complex...PrincipalFull timeWork at office3 days per week
- ...A leading technology company based in Santa Clara, California, is seeking a Senior Software Engineer to focus on the cloud-native stack for their AI/ML datacenters. This role entails deep technical work including debugging complex systems and gathering customer requirements...Full time
- ...NVIDIA’s Networking Systems & Software Architecture group is solving some of AI’s hardest infrastructure... ...co-optimization with GPU, DPU, NIC, and switch teams... ...projects, mentoring engineers, conducting design reviews... ...debugging. ~ Understanding of ML systems concepts—...Full time
- ...NVIDIA seeks a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group to drive GPU-accelerated data processing in production deployments. You'll develop Java, Scala, and CUDA/C++ libraries to accelerate Spark DataFrames and I/O on common formats...PrincipalFull time
$147k - $237.5k
...Palo Alto Networks, Inc. in Santa Clara, California, is looking for a skilled technical professional to develop and support cloud-based services. The role includes responsibilities such as requirements analysis, technical design, and collaboration with QA teams. Ideal...PrincipalFull time$224k - $356.5k
...platform upon which every new AI-powered application is... ...a deeply technical software manager to lead production... ...optimized inference engines, model profiles/recipes,... ...Deep understanding of AI/ML fundamentals, innovative... ...containers, Kubernetes, GPU, or inference communities...Full time- ...A technology innovation company is looking for an experienced AI Security Architect to enhance security in their AI systems. This... ...secure computing. The position allows for hybrid work based in Santa Clara, CA, reflecting the company's commitment to innovation and inclusivity...PrincipalFull time
- ...NVIDIA in Santa Clara is seeking a Senior Software Architect to enhance communication systems for AI and HPC. You will design new communication technologies and improve performance across GPU clusters. The ideal candidate must hold a M.S./Ph.D. in relevant fields, with...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal AI and ML Infra Software Engineer, GPU Clusters (Santa Clara). Be the first to apply!
- machine learning engineer Santa Clara, CA
- network software engineer Santa Clara, CA
- software engineer travel Santa Clara, CA
- senior robotics software engineer Santa Clara, CA
- entry level software engineer remote Santa Clara, CA
- cybersecurity software engineer Santa Clara, CA
- federal - software developer Santa Clara, CA
- agile software developer Santa Clara, CA
- financial software developer Santa Clara, CA
- software engineer Santa Clara, CA











