Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal High-Performance LLM Training Engineer (Santa Clara)

$272k - $431.25k
Part-time

Nvidia

NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post-training workloads across NVIDIA’s full hardware and software stack. This role sits at the intersection of distributed training, GPU architecture, systems software, deep learning frameworks, and performance engineering. You will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs, drive improvements across frameworks such as PyTorch, JAX, NeMo, and NeMo RL, and use insights from real workloads to help shape future NVIDIA GPU, system, and software roadmaps.We are looking for a deeply technical leader who can operate across abstraction layers: from application-level training behavior to framework/runtime internals, CUDA libraries, communication collectives, memory systems, networking, and GPU architecture. At this level, success means both directly improving performance directly as well as setting technical direction, raising the bar for the organization, and influencing multi-functional decisions across NVIDIA.What you will be doing:Lead end-to-end performance analysis and optimization of innovative LLM pre-training and post-training workloads on the latest NVIDIA hardware and software platforms.Drive workloads closer to speed-of-light performance by identifying and removing bottlenecks across compute, memory, communication, scheduling, parallelism strategy, kernel efficiency, framework overhead, and system-level scaling.Develop production-quality software, tools, models, benchmarks, and analysis infrastructure that improve training performance, efficiency, and developer velocity across NVIDIA’s AI software stack.Build and refine performance models, workload characterizations, and simulation methodologies to guide future GPU, networking, system, and software architecture decisions.Serve as a technical authority for AI training performance, partnering closely with teams across GPU architecture, systems, CUDA libraries, compilers, networking, frameworks, product management, and applied AI.Translate workload insights into concrete hardware and software recommendations, and advocate for changes that improve performance and efficiency across the AI ecosystem.Mentor and provide technical leadership to engineers across the organization, helping establish best practices for large-scale AI performance analysis and optimization.What we need to see:A MS, or PhD (or equivalent experience) in Computer Science, Electrical Engineering, Computer Engineering, or a related field, with 12+ years of relevant work or research experience.Demonstrated principal-level technical impact in one or more of the following areas: large-scale AI training systems, GPU performance optimization, distributed systems, high-performance computing, ML frameworks, compilers/runtimes, or hardware/software co-design.Deep hands-on experience analyzing and optimizing performance of large-scale deep learning workloads, especially transformer-based models, LLM pre-training, reinforcement learning, fine-tuning, or other post-training workloads.Strong understanding of GPU and AI accelerator architecture from individual accelerators to datacenter-scale systems.Experience with distributed training techniques such as data parallelism, tensor parallelism, pipeline parallelism, expert parallelism, sequence parallelism, activation checkpointing, mixed precision training, and communication/computation overlap.A strong track record of using profiling, tracing, benchmarking, and performance modeling tools to diagnose complex bottlenecks and drive measurable improvements.Excellent communication and technical leadership skills, with the ability to influence architecture and software decisions across multiple teams without relying on direct authority.GPU computing is the most productive and pervasive platform for deep learning and AI. It begins with the most advanced GPUs and the systems and software we build on top of them. We integrate and optimize every deep learning framework. We work with the major systems companies and every major cloud service provider to make GPUs available in data centers and in the cloud. We craft computers and software to bring AI to edge devices, such as self-driving cars and autonomous robots. AI has the potential to spur a wave of social progress unmatched since the industrial revolution.This opportunity offers you the ability to collaborate with some of the most forward-thinking and hard-working people in the world, shaping the future of AI in a creative and autonomous work environment that encourages innovation. If you're passionate about working across the full hardware & software stack—from GPU architecture to application code—to achieve optimal performance, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until May 2, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 5 hours ago
Similar jobs that could be interesting for youBased on the Principal High-Performance LLM Training Engineer (Santa Clara) in Santa Clara, CA vacancy
  • $200k - $250k

     ...Job Title: Principal Design Verification Engineer - PCIe / High-Speed SoCLocation: Santa Clara, CACompensation: $200K - $250K base DOE plus...  ...lead verification engineers, perform code/review sessions, set best...  ...of PCIe protocol (link training, LTSSM, transaction/completion... 
    Principal
    Training
    Performance
    Part time

    CyberCoders

    Santa Clara, CA
    48 minutes ago
  •  ...working onsite at our Santa Clara, Ca headquarters...  ...compute. As a Principal Hardware Design Engineer, you will be a...  ...'s most massive LLM inference...  ...of experience in high-complexity hardware...  ...education, and training. We also offer incentive...  ...and company performance. This is in... 
    Principal
    Training
    Performance
    Part time
    3 days per week

    d-Matrix

    Santa Clara, CA
    48 minutes ago
  •  ...what’s possible with LLM inference on...  ...applied research and engineering team that moves fast...  ...orchestration to high-level serving APIs...  ...how to squeeze performance out of them at inference...  ...of modern LLM training and inference.Why...  ..., and benefits in Santa Clara, CA.Equal Opportunity... 
    Principal
    Training
    Performance
    Part time

    d-Matrix

    Santa Clara, CA
    46 minutes ago
  • $220.92k - $311.89k

     ...We are seeking a highly experienced and motivated Principal Analog Design Engineer to lead the design...  ...-level link performance.• Proven ability...  ...US, California, Santa ClaraAdditional Locations...  ...education or training. Your recruiter can...  ...California, Santa Clara; US, Arizona,... 
    Principal
    Training
    Performance
    Full time
    Part time
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    46 minutes ago
  • $272k - $431.25k

     ...seeking exceptional engineers to join our...  ...Doing:Design and train innovative large-scale...  ...train, and fine-tune LLM/VLM/VLA systems...  ...environments, ensuring performance, safety, and...  ...opportunity employer. As we highly value diversity in...  ...: US, CA, Santa Clara; US, CA, RemoteType... 
    Principal
    Training
    Performance
    Full time
    Part time
    Work experience placement

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $272k - $431.25k

     ...interconnects.This Principal Architect role...  ...SGLang, and TensorRT-LLM.Publishing...  ...mentoring senior engineers across the organization...  ...expertise in high-performance networking (InfiniBand...  ..., or distributed training and inference...  ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin;... 
    Principal
    Training
    Performance
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $272k - $431.25k

     ...requires a senior software engineer. In this exciting...  ...Deep Learning LLM training and inference. Your primary...  ..., you will build performance analysis tools and strategies...  ...with a focus on high-performance networking...  ...: US, CA, Santa Clara; US, TX, Remote; US,... 
    Principal
    Training
    Performance
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    5 hours ago
  •  ...career. THE ROLE:As a Principal AI Infrastructure Solution Engineer, you will partner with...  ...to enable large‑scale LLM training and inference on AMD Instinct...  ...operating resilient, high‑performance AI workloads at scale....  ...requiredLOCATION:Santa Clara, Ca or open to discuss... 
    Principal
    Training
    Performance
    Part time

    AMD

    Santa Clara, CA
    48 minutes ago
  • $272k - $431.25k

     ...re looking for a Principal Engineer to join our CSP Engagements...  ...for end-to-end performance, working directly...  ...of distributed training performance...  ...(vLLM, TensorRT-LLM, SGLang, continuous...  ...Intelligence, High-Performance Computing...  ...: US, CA, Santa Clara; US, TX, Austin;... 
    Principal
    Training
    Performance
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $272k - $431.25k

     ...NVIDIA Dynamo is a high-throughput, low-latency...  ...environments. Built in Rust for performance and Python for...  ...deployment of cutting-edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision...  ...: US, CA, Santa Clara; US, WA, Remote; US, MA... 
    Principal
    Performance
    Full time
    Part time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $177.82k - $266.4k

     ...shoulder with customer engineering teams from early...  ...You Can ExpectAs Senior Principal Engineer for Signal Integrity...  ...meet the highest performance and quality threshold...  ...contributor role with high customer and ODM visibility...  ...and deliver SI/PI training to customers, ODMs, and... 
    Principal
    Training
    Performance
    Part time
    Internship
    Work from home

    Marvell

    Santa Clara, CA
    45 minutes ago
  • $272k - $431.25k

     ...workloads, deep learning (DL), high-performance computing (HPC), cloud...  ...to see:BS/MS in Electrical Engineering, Computer Science, Computer...  ..., enabling faster AI model training, agentic use-cases, efficient...  ....SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR, HillsboroType... 
    Principal
    Training
    Performance
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  •  ...career. THE ROLE:As a Principal Engineer, you will spearhead...  ...enable massive model training at scale. Your expertise will drive 2-3x performance gains in both training...  ...multi-trillion parameter LLM training/inference...  ...: Austin, Tx or Santa Clara, Ca strongly preferred... 
    Principal
    Training
    Performance
    Part time
    Remote work

    AMD

    Santa Clara, CA
    48 minutes ago
  •  ...of AI. Location:Hybrid (Santa Clara, CA) or Remote The role: Principal AI Systems Architectd-...  ...design scalable and secure high-performance AI accelerator systems...  ...working with the engineering teams to incorporate them...  ...experience, education, and training. We also offer... 
    Principal
    Training
    Performance
    Part time
    Remote work

    d-Matrix

    Santa Clara, CA
    45 minutes ago
  • $224.97k - $317.6k

     ...Development and Customer Engineering (MDCE)...  ...improvement, performance optimization, and...  ...are seeking a Principal Collateral Device...  ...in High-Volume Manufacturing...  ...US, California, Santa ClaraAdditional...  ...relevant education or training. Your recruiter...  ..., Santa Clara; US, Arizona, Phoenix... 
    Principal
    Training
    Performance
    Full time
    Part time
    Local area
    Immediate start
    Worldwide
    Shift work

    Intel

    Santa Clara, CA
    46 minutes ago
  • $220.92k - $311.89k

     ...Role and ImpactAs a Principal Engineer in AMS IP...  ...techniques to optimize performance, power, and area...  ...diverse segments, from high-performance...  ...US, California, Santa ClaraAdditional Locations...  ...education or training. Your recruiter...  ..., Santa Clara; US, Oregon, Hillsboro... 
    Principal
    Training
    Performance
    Full time
    Part time
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    47 minutes ago
  • $220.92k - $311.89k

     ...'s next in scalable, high-performance silicon.At Intel, our...  ....We're looking for a Principal Engineer who thrives on solving...  ...: US, California, Santa ClaraAdditional Locations...  ...education or training. Your recruiter can share...  ...US, California, Santa Clara; US, Oregon, Hillsboro... 
    Principal
    Training
    Performance
    Full time
    Part time
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    46 minutes ago
  • $272k - $431.25k

     ...systems.Drive low-overhead, high-reliability...  .../platform layers, and performance counter/trace providers...  ...technical direction for an engineering team; mentor engineers...  ...experience tuning ML training/inference loops based...  ...SummaryLocation: US, CA, Santa Clara; US, TX, AustinType:... 
    Principal
    Training
    Performance
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $272k - $431.25k

     ...Machine Learning (ML) Engineer to join the GPU...  ..., and ML/DL model training and inference...  ...learning solutions for performance prediction and...  ...and productionizing high-quality ML/DL solutions...  ..., including LLM/GenAI, reinforcement...  ...SummaryLocation: US, CA, Santa ClaraType: Full... 
    Principal
    Training
    Performance
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $272k - $431.25k

     ...with GPU architects, driver engineers, SDK teams, and graphics developers...  ..., memory systems, performance analysis, or graphics debugging...  ...world working here. If you're highly technical and enthusiastic...  ...law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR,... 
    Principal
    Performance
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $158.6k - $237.6k

     ...SoCs encompass best-in-class performance, advanced die-to-die and...  ...customers. These chips use highly advanced technology to facilitate...  ...· Coach and mentor junior engineers of the team when necessary to...  ...employment.#LI-JT2SummaryLocation: Santa Clara, CAType: Full time... 
    Principal
    Performance
    Permanent employment
    Full time
    Part time
    Internship
    Work from home

    Marvell

    Santa Clara, CA
    48 minutes ago
  • $184k - $287.5k

     ...seeking an NCX Senior Engineer to join our DSX team,...  ...customers realize efficient performance from NVIDIA's AI...  ...including distributed training, inference optimization...  ...opportunity employer. As we highly value diversity in our...  ...: US, CA, Santa Clara; US, Remote; US, WA, SeattleType... 
    Training
    Performance
    Full time
    Part time
    Remote work

    Nvidia

    Santa Clara, CA
    48 minutes ago
  • $195.2k - $361.2k

     ...optimize inference engines (llama.cpp, vLLM)...  ...impact with the Post-Training teamCut CPU...  ...and publish honest performance comparisonsUpstream...  ...codeExperience with LLM inference. (attention...  ...: US, California, Santa ClaraAdditional...  ...California, Santa Clara; US, Oregon, Hillsboro... 
    Training
    Performance
    Full time
    Part time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    46 minutes ago
  • $150.68k - $225.7k

     ...ImpactThe PCIe Gen 6, Gen 7 and High-speed Serdes product lines...  ...best-in-class SerDes performance, ultra-low power dissipation...  ...computer science, Electrical Engineering or related fields and 8+ years...  ...process prior to employment.#LI-SA1SummaryLocation: Santa Clara, CAType:... 
    Principal
    Performance
    Permanent employment
    Part time
    Internship
    Work from home

    Marvell

    Santa Clara, CA
    46 minutes ago
  • $164.47k - $269.1k

     ...Analog Circuit Design Engineer, you will be at the forefront...  ...role in optimizing performance, power, area, and...  ...voltage regulators, and high-speed I/O interfaces....  ...Phoenix, US, California, Santa Clara, US, Texas,...  ...relevant education or training. Your recruiter can share... 
    Training
    Performance
    Full time
    Part time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    47 minutes ago
  •  ...working onsite at our Santa Clara, CA, headquarters 3+...  ...per week.The Role: Principal Software Engineer, KernelsWhat you...  ...implementing algorithms in high-level languages such...  ..., education, and training. We also offer...  ...individual and company performance. This is in addition... 
    Principal
    Training
    Performance
    Part time
    Work experience placement
    3 days per week

    d-Matrix

    Santa Clara, CA
    46 minutes ago
  • $272k - $431.25k

     ...world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide...  ...plane development.Make high-quality technical...  ...silicon, or other high-performance computing systems.Expertise...  ...SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US, TX... 
    Principal
    Performance
    Full time
    Part time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $272k - $431.25k

     ...cloud environments. We are looking for Principal Software Engineers to help shape the technical...  ...developments in Artificial Intelligence, High-Performance Computing and Visualization. The...  ...protected by law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time... 
    Principal
    Performance
    Full time
    Part time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $136.88k - $205k

     ...ImpactAs a test development principal engineer in the Operations business...  ...highest standards. Innovate High-Speed Testing Hardware:...  ...Design of complex and high-performance ATE hardware.Effective interpersonal...  ....#LI-TT1SummaryLocation: Santa Clara, CAType: Full time... 
    Principal
    Performance
    Permanent employment
    Full time
    Part time
    Internship
    Work from home

    Marvell

    Santa Clara, CA
    47 minutes ago
  • $208k - $260k

     ...and educational organizations.As a Principal Software Engineer on the Network Management System team...  ....This role is based out of our Santa Clara, CA headquarters, following a hybrid...  ...Drive architecture for extensible, high-performance platforms that support large-scale... 
    Principal
    Performance
    Part time
    Local area
    Worldwide
    3 days per week

    Gigamon

    Santa Clara, CA
    45 minutes ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal High-Performance LLM Training Engineer (Santa Clara). Be the first to apply!