Principal High-Performance LLM Training Engineer (Santa Clara)
$272k - $431.25kNvidia
NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post-training workloads across NVIDIA’s full hardware and software stack. This role sits at the intersection of distributed training, GPU architecture, systems software, deep learning frameworks, and performance engineering. You will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs, drive improvements across frameworks such as PyTorch, JAX, NeMo, and NeMo RL, and use insights from real workloads to help shape future NVIDIA GPU, system, and software roadmaps.We are looking for a deeply technical leader who can operate across abstraction layers: from application-level training behavior to framework/runtime internals, CUDA libraries, communication collectives, memory systems, networking, and GPU architecture. At this level, success means both directly improving performance directly as well as setting technical direction, raising the bar for the organization, and influencing multi-functional decisions across NVIDIA.What you will be doing:Lead end-to-end performance analysis and optimization of innovative LLM pre-training and post-training workloads on the latest NVIDIA hardware and software platforms.Drive workloads closer to speed-of-light performance by identifying and removing bottlenecks across compute, memory, communication, scheduling, parallelism strategy, kernel efficiency, framework overhead, and system-level scaling.Develop production-quality software, tools, models, benchmarks, and analysis infrastructure that improve training performance, efficiency, and developer velocity across NVIDIA’s AI software stack.Build and refine performance models, workload characterizations, and simulation methodologies to guide future GPU, networking, system, and software architecture decisions.Serve as a technical authority for AI training performance, partnering closely with teams across GPU architecture, systems, CUDA libraries, compilers, networking, frameworks, product management, and applied AI.Translate workload insights into concrete hardware and software recommendations, and advocate for changes that improve performance and efficiency across the AI ecosystem.Mentor and provide technical leadership to engineers across the organization, helping establish best practices for large-scale AI performance analysis and optimization.What we need to see:A MS, or PhD (or equivalent experience) in Computer Science, Electrical Engineering, Computer Engineering, or a related field, with 12+ years of relevant work or research experience.Demonstrated principal-level technical impact in one or more of the following areas: large-scale AI training systems, GPU performance optimization, distributed systems, high-performance computing, ML frameworks, compilers/runtimes, or hardware/software co-design.Deep hands-on experience analyzing and optimizing performance of large-scale deep learning workloads, especially transformer-based models, LLM pre-training, reinforcement learning, fine-tuning, or other post-training workloads.Strong understanding of GPU and AI accelerator architecture from individual accelerators to datacenter-scale systems.Experience with distributed training techniques such as data parallelism, tensor parallelism, pipeline parallelism, expert parallelism, sequence parallelism, activation checkpointing, mixed precision training, and communication/computation overlap.A strong track record of using profiling, tracing, benchmarking, and performance modeling tools to diagnose complex bottlenecks and drive measurable improvements.Excellent communication and technical leadership skills, with the ability to influence architecture and software decisions across multiple teams without relying on direct authority.GPU computing is the most productive and pervasive platform for deep learning and AI. It begins with the most advanced GPUs and the systems and software we build on top of them. We integrate and optimize every deep learning framework. We work with the major systems companies and every major cloud service provider to make GPUs available in data centers and in the cloud. We craft computers and software to bring AI to edge devices, such as self-driving cars and autonomous robots. AI has the potential to spur a wave of social progress unmatched since the industrial revolution.This opportunity offers you the ability to collaborate with some of the most forward-thinking and hard-working people in the world, shaping the future of AI in a creative and autonomous work environment that encourages innovation. If you're passionate about working across the full hardware & software stack—from GPU architecture to application code—to achieve optimal performance, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until May 2, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
$200k - $250k
...Job Title: Principal Design Verification Engineer - PCIe / High-Speed SoCLocation: Santa Clara, CACompensation: $200K - $250K base DOE plus... ...lead verification engineers, perform code/review sessions, set best... ...of PCIe protocol (link training, LTSSM, transaction/completion...PrincipalTrainingPerformancePart time- ...working onsite at our Santa Clara, Ca headquarters... ...compute. As a Principal Hardware Design Engineer, you will be a... ...'s most massive LLM inference... ...of experience in high-complexity hardware... ...education, and training. We also offer incentive... ...and company performance. This is in...PrincipalTrainingPerformancePart time3 days per week
- ...what’s possible with LLM inference on... ...applied research and engineering team that moves fast... ...orchestration to high-level serving APIs... ...how to squeeze performance out of them at inference... ...of modern LLM training and inference.Why... ..., and benefits in Santa Clara, CA.Equal Opportunity...PrincipalTrainingPerformancePart time
$220.92k - $311.89k
...We are seeking a highly experienced and motivated Principal Analog Design Engineer to lead the design... ...-level link performance.• Proven ability... ...US, California, Santa ClaraAdditional Locations... ...education or training. Your recruiter can... ...California, Santa Clara; US, Arizona,...PrincipalTrainingPerformanceFull timePart timeLocal areaImmediate startShift work$272k - $431.25k
...seeking exceptional engineers to join our... ...Doing:Design and train innovative large-scale... ...train, and fine-tune LLM/VLM/VLA systems... ...environments, ensuring performance, safety, and... ...opportunity employer. As we highly value diversity in... ...: US, CA, Santa Clara; US, CA, RemoteType...PrincipalTrainingPerformanceFull timePart timeWork experience placement$272k - $431.25k
...interconnects.This Principal Architect role... ...SGLang, and TensorRT-LLM.Publishing... ...mentoring senior engineers across the organization... ...expertise in high-performance networking (InfiniBand... ..., or distributed training and inference... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin;...PrincipalTrainingPerformanceFull timePart timeRemote work$272k - $431.25k
...requires a senior software engineer. In this exciting... ...Deep Learning LLM training and inference. Your primary... ..., you will build performance analysis tools and strategies... ...with a focus on high-performance networking... ...: US, CA, Santa Clara; US, TX, Remote; US,...PrincipalTrainingPerformanceFull timePart timeRemote work- ...career. THE ROLE:As a Principal AI Infrastructure Solution Engineer, you will partner with... ...to enable large‑scale LLM training and inference on AMD Instinct... ...operating resilient, high‑performance AI workloads at scale.... ...requiredLOCATION:Santa Clara, Ca or open to discuss...PrincipalTrainingPerformancePart time
$272k - $431.25k
...re looking for a Principal Engineer to join our CSP Engagements... ...for end-to-end performance, working directly... ...of distributed training performance... ...(vLLM, TensorRT-LLM, SGLang, continuous... ...Intelligence, High-Performance Computing... ...: US, CA, Santa Clara; US, TX, Austin;...PrincipalTrainingPerformanceFull timePart timeRemote work$272k - $431.25k
...NVIDIA Dynamo is a high-throughput, low-latency... ...environments. Built in Rust for performance and Python for... ...deployment of cutting-edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision... ...: US, CA, Santa Clara; US, WA, Remote; US, MA...PrincipalPerformanceFull timePart timeLocal areaRemote work$177.82k - $266.4k
...shoulder with customer engineering teams from early... ...You Can ExpectAs Senior Principal Engineer for Signal Integrity... ...meet the highest performance and quality threshold... ...contributor role with high customer and ODM visibility... ...and deliver SI/PI training to customers, ODMs, and...PrincipalTrainingPerformancePart timeInternshipWork from home$272k - $431.25k
...workloads, deep learning (DL), high-performance computing (HPC), cloud... ...to see:BS/MS in Electrical Engineering, Computer Science, Computer... ..., enabling faster AI model training, agentic use-cases, efficient... ....SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR, HillsboroType...PrincipalTrainingPerformanceFull timePart time- ...career. THE ROLE:As a Principal Engineer, you will spearhead... ...enable massive model training at scale. Your expertise will drive 2-3x performance gains in both training... ...multi-trillion parameter LLM training/inference... ...: Austin, Tx or Santa Clara, Ca strongly preferred...PrincipalTrainingPerformancePart timeRemote work
- ...of AI. Location:Hybrid (Santa Clara, CA) or Remote The role: Principal AI Systems Architectd-... ...design scalable and secure high-performance AI accelerator systems... ...working with the engineering teams to incorporate them... ...experience, education, and training. We also offer...PrincipalTrainingPerformancePart timeRemote work
$224.97k - $317.6k
...Development and Customer Engineering (MDCE)... ...improvement, performance optimization, and... ...are seeking a Principal Collateral Device... ...in High-Volume Manufacturing... ...US, California, Santa ClaraAdditional... ...relevant education or training. Your recruiter... ..., Santa Clara; US, Arizona, Phoenix...PrincipalTrainingPerformanceFull timePart timeLocal areaImmediate startWorldwideShift work$220.92k - $311.89k
...Role and ImpactAs a Principal Engineer in AMS IP... ...techniques to optimize performance, power, and area... ...diverse segments, from high-performance... ...US, California, Santa ClaraAdditional Locations... ...education or training. Your recruiter... ..., Santa Clara; US, Oregon, Hillsboro...PrincipalTrainingPerformanceFull timePart timeLocal areaImmediate startShift work$220.92k - $311.89k
...'s next in scalable, high-performance silicon.At Intel, our... ....We're looking for a Principal Engineer who thrives on solving... ...: US, California, Santa ClaraAdditional Locations... ...education or training. Your recruiter can share... ...US, California, Santa Clara; US, Oregon, Hillsboro...PrincipalTrainingPerformanceFull timePart timeLocal areaImmediate startShift work$272k - $431.25k
...systems.Drive low-overhead, high-reliability... .../platform layers, and performance counter/trace providers... ...technical direction for an engineering team; mentor engineers... ...experience tuning ML training/inference loops based... ...SummaryLocation: US, CA, Santa Clara; US, TX, AustinType:...PrincipalTrainingPerformanceFull timePart time$272k - $431.25k
...Machine Learning (ML) Engineer to join the GPU... ..., and ML/DL model training and inference... ...learning solutions for performance prediction and... ...and productionizing high-quality ML/DL solutions... ..., including LLM/GenAI, reinforcement... ...SummaryLocation: US, CA, Santa ClaraType: Full...PrincipalTrainingPerformanceFull timePart time$272k - $431.25k
...with GPU architects, driver engineers, SDK teams, and graphics developers... ..., memory systems, performance analysis, or graphics debugging... ...world working here. If you're highly technical and enthusiastic... ...law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR,...PrincipalPerformanceFull timePart time$158.6k - $237.6k
...SoCs encompass best-in-class performance, advanced die-to-die and... ...customers. These chips use highly advanced technology to facilitate... ...· Coach and mentor junior engineers of the team when necessary to... ...employment.#LI-JT2SummaryLocation: Santa Clara, CAType: Full time...PrincipalPerformancePermanent employmentFull timePart timeInternshipWork from home$184k - $287.5k
...seeking an NCX Senior Engineer to join our DSX team,... ...customers realize efficient performance from NVIDIA's AI... ...including distributed training, inference optimization... ...opportunity employer. As we highly value diversity in our... ...: US, CA, Santa Clara; US, Remote; US, WA, SeattleType...TrainingPerformanceFull timePart timeRemote work$195.2k - $361.2k
...optimize inference engines (llama.cpp, vLLM)... ...impact with the Post-Training teamCut CPU... ...and publish honest performance comparisonsUpstream... ...codeExperience with LLM inference. (attention... ...: US, California, Santa ClaraAdditional... ...California, Santa Clara; US, Oregon, Hillsboro...TrainingPerformanceFull timePart timeInternshipLocal areaImmediate startShift work$150.68k - $225.7k
...ImpactThe PCIe Gen 6, Gen 7 and High-speed Serdes product lines... ...best-in-class SerDes performance, ultra-low power dissipation... ...computer science, Electrical Engineering or related fields and 8+ years... ...process prior to employment.#LI-SA1SummaryLocation: Santa Clara, CAType:...PrincipalPerformancePermanent employmentPart timeInternshipWork from home$164.47k - $269.1k
...Analog Circuit Design Engineer, you will be at the forefront... ...role in optimizing performance, power, area, and... ...voltage regulators, and high-speed I/O interfaces.... ...Phoenix, US, California, Santa Clara, US, Texas,... ...relevant education or training. Your recruiter can share...TrainingPerformanceFull timePart timeInternshipLocal areaImmediate startShift work- ...working onsite at our Santa Clara, CA, headquarters 3+... ...per week.The Role: Principal Software Engineer, KernelsWhat you... ...implementing algorithms in high-level languages such... ..., education, and training. We also offer... ...individual and company performance. This is in addition...PrincipalTrainingPerformancePart timeWork experience placement3 days per week
$272k - $431.25k
...world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide... ...plane development.Make high-quality technical... ...silicon, or other high-performance computing systems.Expertise... ...SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US, TX...PrincipalPerformanceFull timePart timeRemote workShift work$272k - $431.25k
...cloud environments. We are looking for Principal Software Engineers to help shape the technical... ...developments in Artificial Intelligence, High-Performance Computing and Visualization. The... ...protected by law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time...PrincipalPerformanceFull timePart time$136.88k - $205k
...ImpactAs a test development principal engineer in the Operations business... ...highest standards. Innovate High-Speed Testing Hardware:... ...Design of complex and high-performance ATE hardware.Effective interpersonal... ....#LI-TT1SummaryLocation: Santa Clara, CAType: Full time...PrincipalPerformancePermanent employmentFull timePart timeInternshipWork from home$208k - $260k
...and educational organizations.As a Principal Software Engineer on the Network Management System team... ....This role is based out of our Santa Clara, CA headquarters, following a hybrid... ...Drive architecture for extensible, high-performance platforms that support large-scale...PrincipalPerformancePart timeLocal areaWorldwide3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal High-Performance LLM Training Engineer (Santa Clara). Be the first to apply!
- principal architect Santa Clara, CA
- senior principal cloud computing engineer Santa Clara, CA
- senior principal scientist Santa Clara, CA
- principal cloud computing engineer Santa Clara, CA
- principal Santa Clara, CA
- system performance engineer Santa Clara, CA
- performance test engineer Santa Clara, CA
- senior performance tester Santa Clara, CA
- IT performance management Santa Clara, CA
- senior performance engineer Santa Clara, CA


