Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior AI Inference Performance Engineer (CUDA/LLM/VLM)

NVIDIA AI

Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput. Requires over 6 years of experience in full-stack AI inference performance with strong programming skills in Python, C++, or Rust and expertise in CUDA. #J-18808-Ljbffr NVIDIA AI

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior AI Inference Performance Engineer (CUDA/LLM/VLM) in Santa Clara, CA vacancy
  • $184k - $287.5k

     ...skilled and motivated software engineers to join us and build AI inference systems that serve large-...  ...and implement high-performance inference stacks, optimize...  ...and performance: CUDA, memory hierarchy, streams...  ...building and optimizing LLM inference engines (e.g.,... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • CoreWeave is hiring a Senior Engineer for its Benchmarking & Performance team to write, profile, and optimize GPU kernels on the LLM inference path. You will improve latency and throughput and collaborate with product, orchestration, and hardware teams to achieve strict... 
    Senior
    Performance

    CoreWeave

    Sunnyvale, CA
    2 hours ago
  • $184k - $287.5k

     ...projects at the intersection of CUDA and Deep Learning Systems....  ...unlock maximum hardware performance for emerging AI workloads. You will be a...  ...in both training and inference pipelines.Collaborate closely...  ...Computer Science, Computer Engineering, Electrical Engineering,... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • NVIDIA seeks a Senior Software Engineer - AI Inference Performance to push LLM/VLM workloads toward practical performance limits on NVIDIA GPUs. You will lead end-to-...  ...tuning, and developing high-performance kernels with CUDA, Triton, and CUTLASS while collaborating with... 
    Senior
    Performance

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...now looking for a Senior DL Algorithms Engineer! NVIDIA is...  ...who are mindful of performance analysis and optimization...  ...that leads the AI revolution.What...  ...multimodal model inference as part of NVIDIA...  ...production code to TRT-LLM, NVIDIA’s open-...  ...experience (CUDA or OpenCL) is a plusYour... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...re now looking for a Sr. Inference Engineer, for GPU Kernel Optimization...  ...it take to push every LLM inference operation to its performance ceiling? Our LLM Inference...  ...optimization: applying AI-driven analysis to diagnose...  ...kernel optimization — CUDA, CUTLASS, Triton, or equivalent... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...computing experiences—from AI and data centers, to PCs, gaming...  ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end...  ...rocProfiler, Omniperf) and NVIDIA (CUDA, Nsight Systems/Compute,...  ...vs. NVIDIA) on standardized LLM workloads. Produce clear, data... 
    Senior
    Performance

    AMD

    Santa Clara, CA
    1 day ago
  • Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated...  ...open-source inference engines to reduce latency and increase...  ...in full-stack AI inference performance with...  ...Rust and expertise in CUDA. A degree in Computer Science... 
    Senior
    Performance

    NVIDIA AI

    Santa Clara, CA
    1 day ago
  •  ...evolve a state-of-the-art inference framework in modern C++...  ...with teams across CUDA, kernel libraries, compilers...  ...to deliver high-performance, production-ready solutions...  ...Skills C++, TensorRT, LLM, VLM, GEMM, CUDA, Attention,...  ...Infrastructure, Robotics, Embedded AI Benefits Equity, Health... 
    Senior
    Performance

    NVIDIA AI

    Santa Clara, CA
    12 hours ago
  • $184k - $287.5k

     ...the platform upon which every new AI-powered application is built. We are seeking a Senior Software Engineer - AI Inference Performance to advance innovative LLM and VLM inference. You will push...  ...distributed runtimes, communication, CUDA kernels, and GPU architecture. Deliver... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...driven Software Engineer to bring ground...  ...of Physical AI!What you'll be...  ...containerized inference execution for the...  ...models accuracy and performance (latency,...  ...patches)Contribute VLM-related...  ...Torch, TRT, TRT-LLM).Proficiency with...  ...of ML models (CUDA kernels)Strong... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $193.3k - $261.5k

    We are looking for a Senior Inference Engineer to own inference for real...  ...AI. This is a full-stack...  ...Develop and tune high-performance kernels for critical operations...  ...fall outside standard LLM serving patterns — sustained...  ...(CUTLASS, Triton, raw CUDA/PTX), fused attention... 
    Senior
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    2 days ago
  • $184k - $287.5k

     ...globally. We seek a Senior Engineer to lead technical...  ...advanced AI agent frameworks and...  ...combining powerful local inference (Nemotron models)...  ...understanding of LLM inference pipelines...  ...accelerated computing (CUDA, TensorRT), and...  ...C++ (for performance-critical systems/OS... 
    Senior
    Performance
    Full time
    Local area
    Shift work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...unlimited potential of AI to define the next...  ...Technology Engineer, you will be at the...  ...suboptimal runtime performance.Conduct hands-on trainings...  ....Improve LLM & GenAI user experience...  ....Experience with CUDA and NVIDIA's...  ...GPU-accelerated AI inference driven by NVIDIA APIs... 
    Senior
    Performance
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    3 days ago
  • NVIDIA in Santa Clara, CA is seeking an outstanding High Performance AI Engineer to build groundbreaking multi-agent systems for the CUDA ecosystem, accelerating agent planning, tool-use, code generation, and other AI workloads. You will collaborate with software and hardware... 
    Senior
    Performance

    Thomas To

    Santa Clara, CA
    1 day ago
  • NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and...  ...drives experimental agents, optimize performance, and contribute to cutting-edge AI research... 
    Senior
    Performance

    Nvidia Corporation in

    Santa Clara, CA
    2 days ago
  •  ...headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize...  ...proficiency in C/C++/Python on Linux, and experience with distributed, high-performance software. #J-18808-Ljbffr Jobleads-US
    Senior
    Performance

    Jobleads-US

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    NVIDIA in Santa Clara, CA is seeking a senior systems/C++/CUDA engineer to develop first-of-its-kind storage and IO acceleration libraries. You will optimize performance, work with research teams, and contribute to CUDA and C++ code, with base salary ranging from $184,... 
    Senior
    Performance

    Thomas To

    Santa Clara, CA
    2 days ago
  • NVIDIA is seeking a Senior Software Engineer to advance AI inference performance on GPU-accelerated systems. You will optimize LLM/VLM workloads, profile with Nsight tools, and contribute to open-source inference engines while collaborating with model, kernel, and networking... 
    Senior
    Performance

    Nvidia Corporation in

    Santa Clara, CA
    12 hours ago
  • $184k - $287.5k

     ...the generative AI revolution! The...  ...language models (LLM) and diffusion...  ...for maximal inference efficiency using...  ...by research and engineering teams alike developing...  ...looking for a Senior Deep Learning...  ...improving high-performance kernel implementations in CUDA, TRT-LLM, and... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and...  ..., build new abstractions for LLM serving engines, and contribute to...  ...collaborate across teams, work on CUDA C/C++, Triton, and cutting-edge MLIR... 
    Senior

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...We are now looking for a Senior High-Performance LLM Training Engineer!NVIDIA is seeking experienced engineers specializing...  ...generation of GPUs powering the AI revolution.What you will be doing:...  ...Programming skills in C++, Python, and CUDA.GPU computing is the most productive... 
    Senior
    Performance
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    3 days ago
  • $148k - $235.75k

     ...looking for a Senior Technical Product...  ...pivotal in our inference marketing. You...  ...on working with engineering to understand...  ...CPUs, networking, CUDA libraries,...  ...leadership position in AI inference.Want...  ...and high performance computing. Come...  ...experience in LLM, AI/ML development... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...potential of generative AI to power the...  ...what’s possible with LLM inference on heterogeneous...  ...research and engineering team that moves fast...  ...and how to squeeze performance out of them at inference...  ...programming (CUDA/Triton) and performance...  ...hardware.• Small, senior team with high... 
    Senior
    Performance

    d-Matrix

    Santa Clara, CA
    4 days ago
  • NVIDIA is seeking a Senior Validation Engineer in the DGX Server Product Engineering Team....  ...role focuses on system validation, performance testing, and model enablement in...  ...systems, running real-world ML/LLM workloads, and leveraging CUDA, TensorRT, Slurm and BCM tools.... 
    Senior
    Performance

    NVIDIA

    Santa Clara, CA
    2 hours ago
  • $152k - $241.5k

     ...Intelligence, High Performance Computing and Visualization...  ...Deep Learning engineer to bring advanced...  ...technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX,...  ...up to 100K GPUs to inference down at microsecond...  ...with Python, C++, CUDA or related DSLs (Triton... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...experiences—from AI and data...  ...AMD is seeking a Senior Product Manager...  ...large-scale model inference on AMD Instinct...  ...community, high-performance computing, and...  ...will influence engineering roadmaps, represent...  ...of LLM inference, including...  ..., whether HIP, CUDA, or Triton kernels... 
    Senior
    Performance
    Remote work

    AMD

    Santa Clara, CA
    4 days ago
  •  ...centers, delivering low-latency, high-throughput AI for multi-node GPU workloads. As a Senior Engineer, you will shape core infrastructure and architecture decisions, lead performance optimizations, and own the inference engine to scale research and production workloads... 
    Senior
    Performance

    Sanas

    Palo Alto, CA
    4 days ago
  • $184k - $287.5k

    We are now looking for a Senior High-Performance LLM Training Engineer! NVIDIA is seeking experienced engineers specializing...  ...generation of GPUs powering the AI revolution. What you will be doing...  ...skills in C++, Python, and CUDA. #LI-Hybrid Your base salary will... 
    Senior
    Performance
    Work experience placement

    NVIDIA

    Santa Clara, CA
    12 hours ago
  • $193.3k - $261.5k

     ...unparalleled ML inference and training performance.The Inference Enablement...  ...boundary, our engineers build systematic...  ...'s possible in AI acceleration.As part...  ...wide variety of LLM model families,...  ...mentorship. Our senior members enjoy one...  ...Familiarity with CUDA kernels or equivalent... 
    Senior
    Performance
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior AI Inference Performance Engineer (CUDA/LLM/VLM). Be the first to apply!