Principal AI and ML Infra Software Engineer, GPU Clusters
$272k - $431.25kNVIDIA
What you will be doing: Engage closely with our AI and ML research teams to discern their infrastructure requirements and barriers, converting those insights into actionable improvements. Proactively identify researcher efficiency bottlenecks and lead initiatives to systematically improve it. Drive the direction and long‑term roadmaps for such initiatives. Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization. Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results. Work closely with a variety of teams, such as researchers, data engineers, and DevOps professionals, to develop a cohesive AI/ML infrastructure ecosystem. Keep up to date with the most recent developments in AI/ML technologies, frameworks, and successful strategies, and advocate for their integration within the organization. What we need to see: BS or similar background in Computer Science or related area (or equivalent experience). 15+ years of demonstrated expertise in AI/ML and HPC tasks and systems. Hands‑on experience in using or operating High Performance Computing (HPC) grade infrastructure as well as in‑depth knowledge of accelerated computing (e.g., GPU, custom silicon), storage (e.g., Lustre, GPFS, BeeGFS), scheduling & orchestration (e.g., Slurm, Kubernetes, LSF), high‑speed networking (e.g., Infiniband, RoCE, Amazon EFA), and containers technologies (Docker, Enroot). Capability in supervising and improving substantial distributed training operations using PyTorch (DDP, FSDP), NeMo, or JAX. Moreover, an in‑depth understanding of AI/ML workflows, involving data processing, model training, and inference pipelines. Proficiency in programming & scripting languages such as Python, Go, Bash, as well as familiarity with cloud computing platforms (e.g., AWS, GCP, Azure) in addition to experience with parallel computing frameworks and paradigms. Dedication to ongoing learning and staying updated on new technologies and innovative methods in the AI/ML infrastructure sector. Excellent communication and collaboration skills, with the ability to work effectively with teams and individuals of different backgrounds. NVIDIA offers competitive salaries and a comprehensive benefits package. Our engineering teams are growing rapidly due to outstanding expansion. If you're a passionate and independent engineer with a love for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until May 1, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. #J-18808-Ljbffr
$224k - $356.5k
...unlimited potential of AI to define the next... ...era in which our GPU acts as the brains... ...team is building the software stack that makes... ...efficiency on edge cluster configurationsProduce... ...Science, Computer Engineering, Electrical Engineering... ...in GPU computing, ML systems, or high-...SuggestedFull timeLocal area- ...computing, cloud, and AI. Whether you’re designing... ...of large-scale AI/ML clustered infrastructure. You will... ...team of multi-disciplined engineers that operates across industry... ...years of professional software development experience,... ...Kubernetes for HPC/AI (GPU operators, device...SuggestedFlexible hours
$275.8k - $340.5k
...About the team: The AV ML Infra team at GM builds ML infrastructure... ...meet the unique demands of AI and ML innovation, supporting... ...the productivity of ML engineers, and drive the adoption of cutting... ...Position Overview: The Principal AI/ML Engineer will lead a growing...PrincipalLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours- ...RoboForce RoboForce is an AI robotics company building Physical... ...We are looking for a Senior Software Engineer to build scalable AI... ...will work across cloud systems, GPU clusters, data pipelines, and robotics... ...proficiency with C++, Python, and ML frameworks (e.g., PyTorch,...SuggestedFull timeWork at officeVisa sponsorship
$250k - $350k
...are looking for a Staff/Lead Software Engineer, AI Infrastructure, to play a critical... ...design and implementation of GPU infrastructure, AI model... ...environmentsCollaborate with ML, backend, and platform engineering... ...)Background in building infra for multi-tenant SaaS, enterprise...SuggestedWork at office3 days per week- ...computing experiences—from AI and data centers, to... ...: We are seeking a Principal Software Engineer to serve as the senior... ...qualification on AMD Instinct™ GPU platforms. You will... ...CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch,... ...and large-scale cluster softwareExperience validating...PrincipalContract workShift work
- ...computing experiences—from AI and data centers,... ...in enhancing GPU kernels, deep learning... ...and advanced engineering principles to drive... ...integrating graph compilers. Software Engineering Best... ...compute into ML frameworks (e.g., PyTorch... ...utilization across clusters. Compiler...
- ...generation computing experiences—from AI and data centers, to PCs,... ...'re looking for a senior software engineer who combines deep systems performance... ...who can shape software from GPU kernels through distributed... ...CUTLASS, Thrust, CUB, NCCL), and ML framework cores such as...Shift work
$184k - $287.5k
...highly skilled and motivated software engineers to join us and build AI inference systems that... ...inference stacks, optimize GPU kernels and compilers, drive... ...deployments on GPU clusters across clouds.Conduct and... ...frontier for the field of ML Systems; survey recent publications...Full time- ...computing experiences—from AI and data centers, to... ...driving a unified ROCm software stack across AMD’s... ...silicon, system software, AI/ML frameworks, libraries,... .... Workload Performance Engineering: Lead the profiling,... ...EXPERIENCE: Knowledge in GPU architectures, basic knowledge...Principal
$152k - $241.5k
...computing model focused on visual and AI computing. For two decades,... ..., with our invention of the GPU. The GPU has also shown to be... ...is looking for Architects, Software Engineers, and AI application developers... ...Proficiency in C++, Python and ML frameworks like LangChain, LangSmith...Full time$170.6k - $261.3k
...join us. About the team: The AV ML Infra team at GM builds end-to-end... ...meet the unique demands of AI and ML innovation, supporting... ...enhance the productivity of ML engineers, and drive the adoption of... ...design and build end-to-end software products, owning everything from...Full timeLocal areaWork from homeFlexible hours- ...generation computing experiences—from AI and data centers, to PCs,... ...is looking for an influential software engineer who is passionate about... ...performance from the lowest-level GPU kernels to large-scale distributed... ..., or the C++/HIP/CUDA core of ML frameworks like PyTorch,...
- ...computing experiences—from AI and data centers,... ...- Infrastructure Engineer, Reinforcement... ...APIs across large GPU fleets. You make RL... ...and postmortems for infra incidents affecting... ...machine learning (ML) platforms with deep... ...training, and GPU cluster orchestrationPrior...
$127.1k - $185k
...re looking for a talented early-career engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. You'll work on software that enables the world's largest AI models to train across massive GPU clusters, developing support for communication...InternshipLocal areaFlexible hours- ...the potential of generative AI to power the transformation... ...We are at the forefront of software and hardware innovation,... ...3 days per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat... ...closely with other software (ML and compilers) and hardware...Principal3 days per week
- ...world's largest AI chip, 56 times larger... ...faster than GPU-based hyperscale... ...the Wafer-Scale Engine (WSE). This team... ...labs.As a Principal SRE, you will define... ...customers, and cluster stakeholders can... ...delivering and running software reliably and at... ...in AI/ML inference systems...PrincipalShift work
$152k - $241.5k
...decades. Our invention of the GPU in 1999 sparked the growth of... ...deep learning ignited modern AI - the new era of computing, positioning... ...AI seeks a Senior Systems Software Engineer interested in solving client-... ...and applications using ML/DL frameworks, including ONNX...Full timeLocal area- Principal AI EngineerLocation: Hybrid, Santa Clara Industry: Medical Device REQUIRED: proficiency... ...For7+ years of experience in AI/ML engineering or related fieldsMust be very hands on... ...and agent-based tools to enhance software development workflowsArchitect and deploy...Principal
- ...experiences—from AI and data centers... .... THE ROLE:As a Principal AI Infrastructure Solution Engineer, you will partner with AMD’s AI software teams and... ...strong expertise in GPU‑accelerated computing... ...AMD GPU clusters for distributed... ...checkpointing)Implemented ML observability...Principal
$272k - $431.25k
...Always-On, low-overhead GPU profiling service... ..., scales across cluster environments, and delivers... ...insights for ML workloads. You will... ...across system software, drivers, and CUDA... ...integrate with existing ML/AI workflows (e.g.,... ...technical direction for an engineering team; mentor...PrincipalFull time- ...seeking highly skilled Applied AI Engineer (Software Engineer) to build... ...building predictive and deploying ML/AI solutions for complex PCB... ...regression, classification, clustering, time series, GNNs, reinforcement... ...or custom RL environments. GPU training, distributed...
$270k - $340k
...and celebrates all of our team members.What You’ll Do:As a Principal AI and ML fundamentalist who is an expert at developing cutting-edge AI... ...evaluation, and deployment. Collaborate with other researcher engineers to prototype and validate complex solutions from academic...PrincipalLocal area$200k - $220k
...passionate, and committed engineers, technologists, and... ...an experienced AI Network Software Solution Architect to... ...requires deep expertise in GPU fabric design, high-speed... ...scale out, internal cluster traffic and external ingress... ...infrastructure for AI/ML workloads....PrincipalWorldwide$272k - $431.25k
...for serving generative AI and reasoning models... ..., Dynamo orchestrates GPU shards, routes requests... ...across heterogeneous clusters so that many accelerators... ....We are seeking a Principal Systems Engineer to define the vision and... ...performance storage, or ML systems infrastructure...PrincipalFull timeLocal areaRemote work$184k - $287.5k
We're looking for outstanding AI systems engineers to develop groundbreaking... ...technologies in the inference systems software stack! We build innovative AI... ..., code generators, and GPU kernel technologies for NVIDIA... .../ industry) experience with ML/DL systems development preferableStrong...Full time$152k - $241.5k
We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will have a pivotal... ...to deliver efficiency in our usage of hardware, software, and infrastructure Proactively monitor fleet wide...Full timeRemote work$127.1k - $185k
...AWS. Our org spans silicon engineering, hardware design and verification, software, and operations. We've... ...Inferentia and Trainium ML Accelerators, and scalable... ...analyze ML workloads on custom AI accelerators. You'll work... ...architectures (CPU, NPU, GPU).Amazon is an equal...InternshipLocal areaFlexible hours$165.2k - $223.6k
...talented scientists and engineers to innovate on... ...develop generative AI for shopping. On a... ...reward optimization) at cluster scaleImprove RL... ...training efficiency (GPU utilization,... ...internship professional software development experience... ...- Knowledge of ML frameworks including...InternshipLocal areaFlexible hours$248k - $379.5k
...built on top of our GPU technology,... ...architectures and software optimizations allowing... ..., Prescriptive and AI-augmented Analytics... ...Data Scientists and Engineers working on global deployment... ...deployment of AI/ML based solutions for... ...feedback based clustering and alerting and LLM...PrincipalFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal AI and ML Infra Software Engineer, GPU Clusters. Be the first to apply!
- senior ml engineer Santa Clara, CA
- machine learning engineer Santa Clara, CA
- computer vision machine learning engineer Santa Clara, CA
- ngo software engineer Santa Clara, CA
- software data engineer Santa Clara, CA
- graduate software engineer Santa Clara, CA
- software system engineer Santa Clara, CA
- graduate software developer Santa Clara, CA
- software engineer - early career Santa Clara, CA
- entry level software engineer remote Santa Clara, CA


