Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal AI and ML Infra Software Engineer, GPU Clusters (Santa Clara)

$272k - $431.25k
Full-time

NVIDIA

What you will be doing:

  • Engage closely with our AI and ML research teams to discern their infrastructure requirements and barriers, converting those insights into actionable improvements.
  • Proactively identify researcher efficiency bottlenecks and lead initiatives to systematically improve it. Drive the direction and long‑term roadmaps for such initiatives.
  • Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization.
  • Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results.
  • Work closely with a variety of teams, such as researchers, data engineers, and DevOps professionals, to develop a cohesive AI/ML infrastructure ecosystem.
  • Keep up to date with the most recent developments in AI/ML technologies, frameworks, and successful strategies, and advocate for their integration within the organization.

What we need to see:

  • BS or similar background in Computer Science or related area (or equivalent experience).
  • 15+ years of demonstrated expertise in AI/ML and HPC tasks and systems.
  • Hands‑on experience in using or operating High Performance Computing (HPC) grade infrastructure as well as in‑depth knowledge of accelerated computing (e.g., GPU, custom silicon), storage (e.g., Lustre, GPFS, BeeGFS), scheduling & orchestration (e.g., Slurm, Kubernetes, LSF), high‑speed networking (e.g., Infiniband, RoCE, Amazon EFA), and containers technologies (Docker, Enroot).
  • Capability in supervising and improving substantial distributed training operations using PyTorch (DDP, FSDP), NeMo, or JAX. Moreover, an in‑depth understanding of AI/ML workflows, involving data processing, model training, and inference pipelines.
  • Proficiency in programming & scripting languages such as Python, Go, Bash, as well as familiarity with cloud computing platforms (e.g., AWS, GCP, Azure) in addition to experience with parallel computing frameworks and paradigms.
  • Dedication to ongoing learning and staying updated on new technologies and innovative methods in the AI/ML infrastructure sector.
  • Excellent communication and collaboration skills, with the ability to work effectively with teams and individuals of different backgrounds.

NVIDIA offers competitive salaries and a comprehensive benefits package. Our engineering teams are growing rapidly due to outstanding expansion. If you're a passionate and independent engineer with a love for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until May 1, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

#J-18808-Ljbffr
Vacancy posted 9 hours ago
Similar jobs that could be interesting for youBased on the Principal AI and ML Infra Software Engineer, GPU Clusters (Santa Clara) in Santa Clara, CA vacancy
  •  ...Advanced Micro Devices is looking for a Principal Engineer in Santa Clara, CA to lead AI infrastructure development, define GPU architecture specifications, and drive performance gains in ML systems. The role involves leading innovative techniques, collaborating with... 
    Principal
    Full time

    Advanced Micro Devices

    Santa Clara, CA
    9 hours ago
  • $272k - $431.25k

    NVIDIA is seeking an experienced professional to engage closely with AI and ML research teams to improve infrastructure and researcher efficiency. Candidates should have over 15 years of expertise in AI/ML and HPC tasks, hands-on experience with HPC infrastructure, and... 
    Suggested
    Full time

    NVIDIA

    Santa Clara, CA
    9 hours ago
  •  ...NVIDIA is hiring a Senior Solutions Architect located in Santa Clara, California, to work with its Cloud Partners. This position involves designing next-generation GPU clusters for cutting-edge AI supercomputers and enterprise AI infrastructures. The role requires extensive... 
    Suggested
    Full time

    NVIDIA

    Santa Clara, CA
    9 hours ago
  • $272k - $425.5k

    Principal Software Engineer – Large-Scale LLM Memory and Storage...  ...: US, CA, Santa Clara: US, WA, Remote:...  ...serving generative AI and reasoning models...  ...Dynamo orchestrates GPU shards, routes...  ...across heterogeneous clusters so that many...  ...performance storage, or ML systems... 
    Principal
    Full time
    Local area
    Remote work

    NVIDIA Corporation

    Santa Clara, CA
    9 hours ago
  •  ...Oracle is seeking a Principal AI Agent / ML Software Engineer in Santa Clara, California, to provide technical leadership in developing next-generation AI systems on Oracle Cloud Infrastructure. The ideal candidate will have extensive experience in building scalable AI... 
    Principal
    Full time

    Oracle

    Santa Clara, CA
    9 hours ago
  • $272k - $431.25k

     ...NVIDIA is seeking expert engineers to design next generation rack level solutions for scaling AI supercomputing platforms. With a focus on fleet management solutions...  ...and API management. The position is based in Santa Clara, California, with a competitive salary between... 
    Principal
    Full time

    NVIDIA

    Santa Clara, CA
    9 hours ago
  •  ...experiences—from AI and data centers,...  ...THE ROLE As a Principal Engineer, you will spearhead...  ...infrastructure by defining GPU architecture...  ...for distributed ML systems, you will...  ...passionate about software engineering and possess...  ...Austin, Tx or Santa Clara, Ca strongly preferred... 
    Principal
    Full time
    Remote work

    Advanced Micro Devices

    Santa Clara, CA
    9 hours ago
  • $99.6k - $234.6k

     ...Job Description The Principal AI Agent / ML Software Engineer is a Senior Staff-level, hands‑on technical leadership role responsible for defining, building...  ...services optimized for low latency, high throughput, GPU efficiency, reliability, cost, operability, and secure... 
    Principal
    Full time
    Temporary work
    Flexible hours

    Oracle

    Santa Clara, CA
    9 hours ago
  • $190k - $280k

     ...MixMode is seeking a Principal Architect in Santa Clara to accelerate AI application performance through innovative hardware and software solutions. The role involves analyzing ML workloads and collaborating with product and hardware teams to enhance inference accelerators... 
    Principal
    Full time

    MixMode

    Santa Clara, CA
    9 hours ago
  • $241.8k - $409.2k

    GPGPU Software Architect/ Principal Engineer XPENG is a leading smart technology company at...  ...innovation, integrating advanced AI and autonomous driving...  ...towards General Purpose GPU (GPGPU) architecture. We'...  ...regulations. Location Santa Clara, CA Experience & Seniority... 
    Principal
    Full time

    XPENG

    Santa Clara, CA
    9 hours ago
  • $272k - $431.25k

     ...Networking Systems & Software Architecture group is solving some of AI’s hardest...  ...interconnects. This Principal Architect role leads...  ...communication systems—GPU-to-GPU, GPU-to-storage...  ...mentoring senior engineers across the...  ...~ Understanding of ML systems concepts—transformer... 
    Principal
    Full time

    NVIDIA Gruppe

    Santa Clara, CA
    9 hours ago
  •  ...potential of generative AI to power the...  ...the forefront of software and hardware innovation...  ...: Hybrid (Santa Clara, CA) or Remote...  ...Security Architect (Principal) d-Matrix is seeking...  ...research in ML, architecture, and...  ...working with the engineering teams to incorporate... 
    Principal
    Full time
    Remote work

    d-Matrix

    Santa Clara, CA
    9 hours ago
  • $184k - $287.5k

     ...Senior Software Engineer, Cloud-Native Stack – CSP Engagements...  ...Apply locations US, CA, Santa Clara US, TX, Austin US, WA,...  ...-rack, multi-tenant AI/ML datacenters with NVIDIA...  ...-rack, multi-tenant clusters: scheduler behavior, container...  ...that expose new GPU capabilities. Drive... 
    Full time

    NVIDIA Corporation

    Santa Clara, CA
    9 hours ago
  •  ...NVIDIA is seeking a Sr. Principal Systems Software Engineer in Santa Clara to lead development of GPU-accelerated data processing for the Apache Spark ecosystem. You will design and implement Java, Scala, and CUDA/C++ libraries to speed up DataFrames, I/O, and interoperable... 
    Principal
    Full time

    NVIDIA

    Santa Clara, CA
    9 hours ago
  • $195k - $292k

     ...Ampere Computing LLC. is seeking a Principal AI Accelerator Software Engineer-Graph Optimization in Santa Clara, California. The role focuses on optimizing computational graphs to unlock the full potential of Ampere's deep learning hardware. Responsibilities include collaborating... 
    Principal
    Full time

    Ampere Computing LLC.

    Santa Clara, CA
    9 hours ago
  • NVIDIA is seeking a Senior Software Architect for its GPU Networking Architecture team in Santa Clara, California. In this role, you will define architectural solutions for Software Defined Networking in AI networks, collaborating closely with teams across GPU and Software... 
    Full time

    NVIDIA

    Santa Clara, CA
    9 hours ago
  •  ...NVIDIA's research team in Santa Clara is seeking a Systems Software Engineer to tackle AI infrastructure challenges. You'll architect communication and memory management...  ...movement across systems, and work closely with GPU and networking teams. The ideal candidate has 12+... 
    Full time

    NVIDIA

    Santa Clara, CA
    9 hours ago
  • $224k - $336k

     ...is looking for a Principal Application Support Engineer to join our AI Networking Infrastructure...  ...hardware and software technologies...  ...of large-scale AI clusters. The Person The...  ...Familiarity with GPU operations, Collective...  ...field Location Santa Clara, CA or Austin, TX... 
    Principal
    Full time

    AMD

    Santa Clara, CA
    9 hours ago
  •  ...seeking a Senior Solutions Architect to join the Cluster Design and Architecture team with a focus on...  ...involves designing and optimizing large-scale AI/HPC GPU clusters, advising on topology, and collaborating with engineering and field teams to meet demanding customer... 
    Full time

    NVIDIA

    Santa Clara, CA
    9 hours ago
  • $272k - $431.25k

     ...unlimited potential of AI to define the...  ...in which our GPU acts as the...  ...At NVIDIA, as a Principal Rack Scale...  ...Infrastructure Engineer, you will build...  ...development of software systems. These...  ...with rack‑ or cluster‑scale systems spanning...  ...firmware, and infra management as one... 
    Principal
    Full time
    Shift work

    NVIDIA Corporation

    Santa Clara, CA
    9 hours ago
  • $272k - $431.25k

     ...seeking a highly motivated Principal System Software Engineer to drive next-generation innovations...  ..., architecture, kernel, AI, middleware, and platform...  ...optimization initiatives across CPU, GPU, memory, storage, networking...  ...computing and AI/ML software platforms. Contributions... 
    Principal
    Full time

    NVIDIA

    Santa Clara, CA
    9 hours ago
  • $272k - $431.25k

     ...NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC...  ...are increasingly known as \the AI computing company.\ We are looking...  .... We are looking for expert engineers to come and help design rack level...  ...Compute. Experience with ML and multi‑variable optimisation... 
    Principal
    Full time

    NVIDIA

    Santa Clara, CA
    9 hours ago
  •  ...and Inclusion. We weave AI into the fabric of...  ...seeking a world‑class Principal Engineer (Sr Manager‑equivalent...  ...elevate our standards for software quality, and unlock...  ...located at our dynamic Santa Clara California headquarters...  ...applying AI/ML/GenAI to solve complex... 
    Principal
    Full time
    Work at office
    3 days per week

    Palo Alto Networks, Inc.

    Santa Clara, CA
    9 hours ago
  •  ...A leading technology company based in Santa Clara, California, is seeking a Senior Software Engineer to focus on the cloud-native stack for their AI/ML datacenters. This role entails deep technical work including debugging complex systems and gathering customer requirements... 
    Full time

    NVIDIA Corporation

    Santa Clara, CA
    9 hours ago
  •  ...NVIDIA’s Networking Systems & Software Architecture group is solving some of AI’s hardest infrastructure...  ...co-optimization with GPU, DPU, NIC, and switch teams...  ...projects, mentoring engineers, conducting design reviews...  ...debugging. ~ Understanding of ML systems concepts—... 
    Full time

    NVIDIA

    Santa Clara, CA
    9 hours ago
  •  ...NVIDIA seeks a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group to drive GPU-accelerated data processing in production deployments. You'll develop Java, Scala, and CUDA/C++ libraries to accelerate Spark DataFrames and I/O on common formats... 
    Principal
    Full time

    Nvidia Corporation

    Santa Clara, CA
    9 hours ago
  • $147k - $237.5k

     ...Palo Alto Networks, Inc. in Santa Clara, California, is looking for a skilled technical professional to develop and support cloud-based services. The role includes responsibilities such as requirements analysis, technical design, and collaboration with QA teams. Ideal... 
    Principal
    Full time

    Palo Alto Networks, Inc.

    Santa Clara, CA
    9 hours ago
  • $224k - $356.5k

     ...platform upon which every new AI-powered application is...  ...a deeply technical software manager to lead production...  ...optimized inference engines, model profiles/recipes,...  ...Deep understanding of AI/ML fundamentals, innovative...  ...containers, Kubernetes, GPU, or inference communities... 
    Full time

    NVIDIA

    Santa Clara, CA
    9 hours ago
  •  ...A technology innovation company is looking for an experienced AI Security Architect to enhance security in their AI systems. This...  ...secure computing. The position allows for hybrid work based in Santa Clara, CA, reflecting the company's commitment to innovation and inclusivity... 
    Principal
    Full time

    d-Matrix

    Santa Clara, CA
    9 hours ago
  •  ...NVIDIA in Santa Clara is seeking a Senior Software Architect to enhance communication systems for AI and HPC. You will design new communication technologies and improve performance across GPU clusters. The ideal candidate must hold a M.S./Ph.D. in relevant fields, with... 
    Full time

    NVIDIA

    Santa Clara, CA
    9 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal AI and ML Infra Software Engineer, GPU Clusters (Santa Clara). Be the first to apply!