Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal High-Performance LLM Training Engineer

$272k - $431.25k

Jobleads-US

NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post-training workloads across NVIDIA’s full hardware and software stack. This role sits at the intersection of distributed training, GPU architecture, systems software, deep learning frameworks, and performance engineering. You will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs, drive improvements across frameworks such as PyTorch, JAX, NeMo, and NeMo RL, and use insights from real workloads to help shape future NVIDIA GPU, system, and software roadmaps. We are looking for a deeply technical leader who can operate across abstraction layers: from application-level training behavior to framework/runtime internals, CUDA libraries, communication collectives, memory systems, networking, and GPU architecture. At this level, success means both directly improving performance directly as well as setting technical direction, raising the bar for the organization, and influencing multi-functional decisions across NVIDIA.

What you will be doing:

Lead end-to-end performance analysis and optimization of innovative LLM pre-training and post-training workloads on the latest NVIDIA hardware and software platforms. Drive workloads closer to speed-of-light performance by identifying and removing bottlenecks across compute, memory, communication, scheduling, parallelism strategy, kernel efficiency, framework overhead, and system-level scaling. Develop production-quality software, tools, models, benchmarks, and analysis infrastructure that improve training performance, efficiency, and developer velocity across NVIDIA’s AI software stack. Build and refine performance models, workload characterizations, and simulation methodologies to guide future GPU, networking, system, and software architecture decisions. Serve as a technical authority for AI training performance, partnering closely with teams across GPU architecture, systems, CUDA libraries, compilers, networking, frameworks, product management, and applied AI. Translate workload insights into concrete hardware and software recommendations, and advocate for changes that improve performance and efficiency across the AI ecosystem. Mentor and provide technical leadership to engineers across the organization, helping establish best practices for large-scale AI performance analysis and optimization.

What we need to see:

A MS, or PhD (or equivalent experience) in Computer Science, Electrical Engineering, Computer Engineering, or a related field, with 12+ years of relevant work or research experience. Demonstrated principal-level technical impact in one or more of the following areas: large-scale AI training systems, GPU performance optimization, distributed systems, high-performance computing, ML frameworks, compilers/runtimes, or hardware/software co-design. Deep hands‑on experience analyzing and optimizing performance of large-scale deep learning workloads, especially transformer-based models, LLM pre‑training, reinforcement learning, fine‑tuning, or other post‑training workloads. Strong understanding of GPU and AI accelerator architecture from individual accelerators to datacenter-scale systems. Experience with distributed training techniques such as data parallelism, tensor parallelism, pipeline parallelism, expert parallelism, sequence parallelism, activation checkpointing, mixed precision training, and communication/computation overlap. A strong track record of using profiling, tracing, benchmarking, and performance modeling tools to diagnose complex bottlenecks and drive measurable improvements. Excellent communication and technical leadership skills, with the ability to influence architecture and software decisions across multiple teams without relying on direct authority.

GPU computing is the most productive and pervasive platform for deep learning and AI. It begins with the most advanced GPUs and the systems and software we build on top of them. We integrate and optimize every deep learning framework. We work with the major systems companies and every major cloud service provider to make GPUs available in data centers and in the cloud. We craft computers and software to bring AI to edge devices, such as self-driving cars and autonomous robots. AI has the potential to spur a wave of social progress unmatched since the industrial revolution. This opportunity offers you the ability to collaborate with some of the most forward-thinking and hard-working people in the world, shaping the future of AI in a creative and autonomous work environment that encourages innovation.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until May 2, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Principal High-Performance LLM Training Engineer in Santa Clara, CA vacancy
  •  ...Google in San Jose, CA, is seeking a Principal M-LLM Post-Training and Execution Software Engineer, XR to lead production-ready features from state-of-the-...  ...platform. You will bridge model alignment and high-performance execution, ensuring safety, reliability, and... 
    Principal
    Training
    Performance

    Jobleads-US

    San Jose, CA
    3 days ago
  • $184k - $287.5k

     ...Senior High-Performance LLM Training Engineer NVIDIA is seeking experienced engineers specializing in performance analysis and optimization to improve the efficiency of LLM training workloads, which are shaping the world's most advanced computing systems. This position... 
    Training
    Performance
    Work experience placement

    Jobleads-US

    Santa Clara, CA
    2 days ago
  •  ...analyze, profile, and optimize AI training workloads on innovative...  ...the big picture of training performance on GPUs, prioritizing and...  ...Computer Science, Electrical Engineering or Computer Engineering and...  ...opportunity employer. As we highly value diversity in our current... 
    Training
    Performance
    Work experience placement

    NVIDIA Gruppe

    Santa Clara, CA
    2 days ago
  • $307k - $427k

     ...architecture for on-device LLM integration and...  ...to optimize model performance for constrained...  ...leads and staff engineers, fostering a...  ...and useful. As a Principal Software Engineer...  ...capabilities and high-level AI services,...  ...relevant education or training. US: $307000 - $42... 
    Principal
    Training
    Performance

    Google

    San Jose, CA
    2 days ago
  • $126.2k - $210.35k

     ...A Day in Your Life at MKS: As a Principal Electrical Engineer – High-Speed Optoelectronics & RF Systems at...  ...optoelectronic products and subsystems. Perform system-level tradeoff analyses...  ...including qualifications, experience and training, operational and business needs and... 
    Principal
    Training
    Performance
    Permanent employment
    Full time
    Work at office
    Visa sponsorship
    Work visa
    Relocation package

    MKS Instruments

    Milpitas, CA
    2 days ago
  •  ...NVIDIA Corporation seeks a Senior High-Performance LLM Training Engineer to optimize AI training workloads and accelerate performance across thousands of GPUs. You will profile, analyze, and implement production-quality software across the deep learning stack, from... 
    Training
    Performance

    Jobleads-US

    Santa Clara, CA
    2 days ago
  •  ...worldwide.We’re a team of engineers, clinicians, and...  ...work helps care teams perform with greater precision...  ...surgery system that uses highly complex mechanics,...  ...of the human hand. The Principal Engineer will strive to...  ...timeRequired Education and Training Bachelor’s or master’s... 
    Principal
    Training
    Performance
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    3 days ago
  • $145k - $236k

     ...Role:We are seeking a Field Application Engineer (NAM) to support and drive customer...  ...architecture flexibilityMulti-threaded, high-performance computeHigh-performance scalable Neural...  ...employees including recruitment, selection, training, utilization, promotion, compensation,... 
    Principal
    Training
    Performance
    Full time
    Work at office
    Local area
    Immediate start

    Globalfoundries

    Santa Clara, CA
    4 days ago
  • $188k - $220k

     ...Title: Principal Clinical Engineer This position is based in our Campbell, California offices...  ...requirements. Participate in the hiring, training and performance management of assigned team aiming...  ...current on information related to high quality pre-clinical testing... 
    Principal
    Training
    Performance
    Full time
    Work experience placement
    Worldwide

    Imperative Care

    Campbell, CA
    5 days ago
  • $153.5k - $310.5k

     ...Senior Principal Mechanical Engineer This role has been designed as ‘Hybrid’ with...  ...decisions, balancing performance, cost, schedule, reliability...  ...storage, telecommunications, or high‑performance computing...  ...work experience, education/training, and/or skill level.... 
    Principal
    Training
    Performance
    Work experience placement
    Work at office
    2 days per week

    Hobbsnews

    Sunnyvale, CA
    1 day ago
  • $200k - $240k

     ...advanced cell architectures to enable high-performance rechargeable batteries for demanding applications. We are seeking a Principal Engineer – Cell Design & Mechanical Architecture...  ...not limited to, internal equity, experience, education, specialty, and training... 
    Principal
    Training
    Performance

    ENSURGE

    San Jose, CA
    a month ago
  •  ...worldwide.Were a team of engineers, clinicians, and...  ...work helps care teams perform with greater precision...  ...Surgical explores novel, high-potential technologies...  ...intervention. We seek a Principal Systems Research...  ...Required Education and Training Typically requires a minimum... 
    Principal
    Training
    Performance
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    3 days ago
  •  ...reproducing customer issues, performing root cause analysis, and verifying...  ...addressed through fixes and engineering accountability.-...  ...across product areas, identifying high-risk impact areas and working...  ...leverages ongoing feedback and training to improve skills.- Coaches and... 
    Principal
    Training
    Performance
    Temporary work
    Local area
    Flexible hours
    Shift work

    hackajob

    Santa Clara, CA
    3 days ago
  •  ...ultra-low power consumption, high radix, compact chip-scale design...  ...data centers enhanced performance, efficiency, and scalability....  ...a Wafer Bonding Integration Engineer to lead the development, integration...  ...relevant experience, skills, training, education, market demands,... 
    Principal
    Training
    Performance

    nEye.ai

    Santa Clara, CA
    a month ago
  • $148.32k - $203.94k

     ...all electronics, ensuring performance, resilience and scalability...  ...00 applications, including high-growth ones in AI datacenters...  ...Job Summary   The Principal Test Engineer is responsible for the development...  ...experience, education, and training. In addition to base... 
    Principal
    Training
    Performance
    Work experience placement

    SiTime Corporation

    Santa Clara, CA
    a month ago
  • $148.32k - $203.94k

     ...of all electronics, ensuring performance, resilience and scalability....  ...400 applications, including high-growth ones in AI datacenters...  ...Collaborate with Digital Design Engineers, CAD, Systems Engineering,...  ..., experience, education, and training. In addition to base salary... 
    Principal
    Training
    Performance

    SiTime Corporation

    Santa Clara, CA
    a month ago
  • $248k - $396.75k

     ...Site Reliability Engineering (SRE) at NVIDIA is an engineering...  ...their reliability and performance objectives while...  ...environments. As a Principal SRE, you will shape the...  .... Architect highly available, resilient,...  ...infrastructure, inference systems, training environments, or high-... 
    Principal
    Training
    Performance

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $200k - $250k

     ...architectures. We are seeking a Principal Engineer – Thin-Film & Process...  ...processes. This is a highly hands-on role for an engineer...  ...to improve chamber performance, uptime, throughput, and repeatability...  ...equity, experience, education, specialty, and training.... 
    Principal
    Training
    Performance

    ENSURGE

    San Jose, CA
    a month ago
  •  ...ultra-low power consumption, high radix, compact chip-scale design...  ...data centers enhanced performance, efficiency, and scalability....  ...Systems Packaging & Assembly Engineer to lead the development and integration...  ...relevant experience, skills, training, education, market demands,... 
    Principal
    Training
    Performance
    Contract work

    nEye.ai

    Santa Clara, CA
    11 days ago
  •  ...implementation strategies, migration plans, and engineering bills of materials (BOMs) for...  .... ~ Experience supporting AI, high-performance computing (HPC), cloud, and large-scale...  ...location, work experience, education/training, and/or skill level. – United States... 
    Principal
    Training
    Performance
    Work experience placement
    Remote work
    Work from home

    Jobleads-US

    Sunnyvale, CA
    4 days ago
  • $272k - $431.25k

     ...architect with deep experience designing high-performance, high-frequency server-class...  ...need to see: ~ BS/MS in Electrical Engineering, Computer Science, Computer Engineering...  ...technology stack, enabling faster AI model training, agentic use-cases, efficient data processing... 
    Principal
    Training
    Performance

    Jobleads-US

    Santa Clara, CA
    6 days ago
  • $174k - $352.5k

     ...Sr Principal Thermal EngineerThis role has been designed as Onsite...  ...a Senior Principal Thermal Engineer who is a recognized authority...  ...advanced liquid cooling for high-performance networking, compute, and data...  ...work experience, education/training, and/or skill level. – United... 
    Principal
    Training
    Performance
    Work experience placement
    Work at office

    Jobleads-US

    Sunnyvale, CA
    2 days ago
  • $156.8k - $219.52k

     ...experienced Electrical Power Engineer to join the TeraWave team at...  ...hardware that meets the demanding performance, mass, reliability, and...  ...Design and characterize sub-1V, high-current point-of-load (POL)...  ...materials transportation/shipping training. Required for certain Job... 
    Principal
    Training
    Performance
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide

    BLUE ORIGIN

    Cupertino, CA
    4 days ago
  •  ...NVIDIA is seeking a senior or principal engineer specializing in physics simulation within the Generalist Embodied Agent Research...  ...MuJoCo/Isaac Sim-based environments, optimizing GPU performance for large-scale training, and deploying learned models to physical robots.... 
    Principal
    Training
    Performance

    Jobleads-US

    Santa Clara, CA
    2 days ago
  •  ...Advanced Micro Devices, Inc. is seeking a Principal or Fellow level software engineer to advance AI infrastructure, focusing on performance, reliability and scalable workloads for model training and inference. You will collaborate with customers and cross‑functional... 
    Principal
    Training
    Performance

    Jobleads-US

    San Jose, CA
    2 days ago
  •  ...discover, and create.We’re seeking a Principal Machine Learning Systems Engineer (P60) to lead technical...  ...cutting-edge algorithms with reliable, high-performance infrastructure.Working at AtlassianAtlassians...  ...implement scalable systems for training, fine-tuning, and serving large... 
    Principal
    Training
    Performance
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    1 day ago
  • $260k - $275k

     ...Influence architecture and engineering culture at a company...  ...annual security training What You Will Be Doing...  ...implement, and manage highly available and scalable...  ...fault tolerance, and performance Manage and optimize...  ...years of experience as a Principal SRE with a strong... 
    Principal
    Training
    Performance

    Saviynt

    Milpitas, CA
    2 days ago
  •  ...GlobalFoundries, we use the same high-volume processes that...  ...mission of PsiQuantum's Principal Hardware Design Engineer role is to support the Electronics...  ...data and extract key performance parameters. Maintain...  ...qualifications, relevant education and training, competencies, experience,... 
    Principal
    Training
    Performance
    Full time
    Shift work

    PSI Quantum

    Milpitas, CA
    2 days ago
  • $272k - $431.25k

     ...interconnects. This Principal Architect role leads...  ...SGLang, and TensorRT-LLM. Publishing findings...  ...and mentoring senior engineers across the organization...  ...with deep expertise in high-performance networking (InfiniBand...  ...parallelism, or distributed training and inference patterns... 
    Principal
    Training
    Performance

    NVIDIA Gruppe

    Santa Clara, CA
    4 days ago
  • $180k - $260k

     ...capabilities to customers with high demand for artificial...  ...We are seeking a Principal Kubernetes Control Plane Engineer to architect the foundational...  ..." that supports our high-performance AI workloads, ensuring our...  ...before they impact customer training or inference jobs.... 
    Principal
    Training
    Performance
    Remote job
    Full time
    Local area

    Bitdeer

    San Jose, CA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal High-Performance LLM Training Engineer. Be the first to apply!