Senior AI Inference Performance Engineer (CUDA/LLM/VLM)
NVIDIA AI
Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput. Requires over 6 years of experience in full-stack AI inference performance with strong programming skills in Python, C++, or Rust and expertise in CUDA. #J-18808-Ljbffr NVIDIA AI
$184k - $287.5k
...skilled and motivated software engineers to join us and build AI inference systems that serve large-... ...and implement high-performance inference stacks, optimize... ...and performance: CUDA, memory hierarchy, streams... ...building and optimizing LLM inference engines (e.g.,...SeniorPerformanceFull time- CoreWeave is hiring a Senior Engineer for its Benchmarking & Performance team to write, profile, and optimize GPU kernels on the LLM inference path. You will improve latency and throughput and collaborate with product, orchestration, and hardware teams to achieve strict...SeniorPerformance
$184k - $287.5k
...projects at the intersection of CUDA and Deep Learning Systems.... ...unlock maximum hardware performance for emerging AI workloads. You will be a... ...in both training and inference pipelines.Collaborate closely... ...Computer Science, Computer Engineering, Electrical Engineering,...SeniorPerformanceFull time- NVIDIA seeks a Senior Software Engineer - AI Inference Performance to push LLM/VLM workloads toward practical performance limits on NVIDIA GPUs. You will lead end-to-... ...tuning, and developing high-performance kernels with CUDA, Triton, and CUTLASS while collaborating with...SeniorPerformance
$184k - $287.5k
...now looking for a Senior DL Algorithms Engineer! NVIDIA is... ...who are mindful of performance analysis and optimization... ...that leads the AI revolution.What... ...multimodal model inference as part of NVIDIA... ...production code to TRT-LLM, NVIDIA’s open-... ...experience (CUDA or OpenCL) is a plusYour...SeniorPerformanceFull time$184k - $287.5k
...re now looking for a Sr. Inference Engineer, for GPU Kernel Optimization... ...it take to push every LLM inference operation to its performance ceiling? Our LLM Inference... ...optimization: applying AI-driven analysis to diagnose... ...kernel optimization — CUDA, CUTLASS, Triton, or equivalent...SeniorPerformanceFull time- ...computing experiences—from AI and data centers, to PCs, gaming... ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end... ...rocProfiler, Omniperf) and NVIDIA (CUDA, Nsight Systems/Compute,... ...vs. NVIDIA) on standardized LLM workloads. Produce clear, data...SeniorPerformance
- Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated... ...open-source inference engines to reduce latency and increase... ...in full-stack AI inference performance with... ...Rust and expertise in CUDA. A degree in Computer Science...SeniorPerformance
- ...evolve a state-of-the-art inference framework in modern C++... ...with teams across CUDA, kernel libraries, compilers... ...to deliver high-performance, production-ready solutions... ...Skills C++, TensorRT, LLM, VLM, GEMM, CUDA, Attention,... ...Infrastructure, Robotics, Embedded AI Benefits Equity, Health...SeniorPerformance
$184k - $287.5k
...the platform upon which every new AI-powered application is built. We are seeking a Senior Software Engineer - AI Inference Performance to advance innovative LLM and VLM inference. You will push... ...distributed runtimes, communication, CUDA kernels, and GPU architecture. Deliver...SeniorPerformanceFull time$152k - $241.5k
...driven Software Engineer to bring ground... ...of Physical AI!What you'll be... ...containerized inference execution for the... ...models accuracy and performance (latency,... ...patches)Contribute VLM-related... ...Torch, TRT, TRT-LLM).Proficiency with... ...of ML models (CUDA kernels)Strong...SeniorPerformanceFull time$193.3k - $261.5k
We are looking for a Senior Inference Engineer to own inference for real... ...AI. This is a full-stack... ...Develop and tune high-performance kernels for critical operations... ...fall outside standard LLM serving patterns — sustained... ...(CUTLASS, Triton, raw CUDA/PTX), fused attention...SeniorPerformanceInternshipLocal areaFlexible hours$184k - $287.5k
...globally. We seek a Senior Engineer to lead technical... ...advanced AI agent frameworks and... ...combining powerful local inference (Nemotron models)... ...understanding of LLM inference pipelines... ...accelerated computing (CUDA, TensorRT), and... ...C++ (for performance-critical systems/OS...SeniorPerformanceFull timeLocal areaShift work$152k - $241.5k
...unlimited potential of AI to define the next... ...Technology Engineer, you will be at the... ...suboptimal runtime performance.Conduct hands-on trainings... ....Improve LLM & GenAI user experience... ....Experience with CUDA and NVIDIA's... ...GPU-accelerated AI inference driven by NVIDIA APIs...SeniorPerformanceFull timeLocal area- NVIDIA in Santa Clara, CA is seeking an outstanding High Performance AI Engineer to build groundbreaking multi-agent systems for the CUDA ecosystem, accelerating agent planning, tool-use, code generation, and other AI workloads. You will collaborate with software and hardware...SeniorPerformance
- NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and... ...drives experimental agents, optimize performance, and contribute to cutting-edge AI research...SeniorPerformance
- ...headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize... ...proficiency in C/C++/Python on Linux, and experience with distributed, high-performance software. #J-18808-Ljbffr Jobleads-USSeniorPerformance
$184k - $287.5k
NVIDIA in Santa Clara, CA is seeking a senior systems/C++/CUDA engineer to develop first-of-its-kind storage and IO acceleration libraries. You will optimize performance, work with research teams, and contribute to CUDA and C++ code, with base salary ranging from $184,...SeniorPerformance- NVIDIA is seeking a Senior Software Engineer to advance AI inference performance on GPU-accelerated systems. You will optimize LLM/VLM workloads, profile with Nsight tools, and contribute to open-source inference engines while collaborating with model, kernel, and networking...SeniorPerformance
$184k - $287.5k
...the generative AI revolution! The... ...language models (LLM) and diffusion... ...for maximal inference efficiency using... ...by research and engineering teams alike developing... ...looking for a Senior Deep Learning... ...improving high-performance kernel implementations in CUDA, TRT-LLM, and...SeniorPerformanceFull time- ...Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and... ..., build new abstractions for LLM serving engines, and contribute to... ...collaborate across teams, work on CUDA C/C++, Triton, and cutting-edge MLIR...Senior
$184k - $287.5k
...We are now looking for a Senior High-Performance LLM Training Engineer!NVIDIA is seeking experienced engineers specializing... ...generation of GPUs powering the AI revolution.What you will be doing:... ...Programming skills in C++, Python, and CUDA.GPU computing is the most productive...SeniorPerformanceFull timeWork experience placement$148k - $235.75k
...looking for a Senior Technical Product... ...pivotal in our inference marketing. You... ...on working with engineering to understand... ...CPUs, networking, CUDA libraries,... ...leadership position in AI inference.Want... ...and high performance computing. Come... ...experience in LLM, AI/ML development...SeniorPerformanceFull time- ...potential of generative AI to power the... ...what’s possible with LLM inference on heterogeneous... ...research and engineering team that moves fast... ...and how to squeeze performance out of them at inference... ...programming (CUDA/Triton) and performance... ...hardware.• Small, senior team with high...SeniorPerformance
- NVIDIA is seeking a Senior Validation Engineer in the DGX Server Product Engineering Team.... ...role focuses on system validation, performance testing, and model enablement in... ...systems, running real-world ML/LLM workloads, and leveraging CUDA, TensorRT, Slurm and BCM tools....SeniorPerformance
$152k - $241.5k
...Intelligence, High Performance Computing and Visualization... ...Deep Learning engineer to bring advanced... ...technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX,... ...up to 100K GPUs to inference down at microsecond... ...with Python, C++, CUDA or related DSLs (Triton...SeniorPerformanceFull timeRemote work- ...experiences—from AI and data... ...AMD is seeking a Senior Product Manager... ...large-scale model inference on AMD Instinct... ...community, high-performance computing, and... ...will influence engineering roadmaps, represent... ...of LLM inference, including... ..., whether HIP, CUDA, or Triton kernels...SeniorPerformanceRemote work
- ...centers, delivering low-latency, high-throughput AI for multi-node GPU workloads. As a Senior Engineer, you will shape core infrastructure and architecture decisions, lead performance optimizations, and own the inference engine to scale research and production workloads...SeniorPerformance
$184k - $287.5k
We are now looking for a Senior High-Performance LLM Training Engineer! NVIDIA is seeking experienced engineers specializing... ...generation of GPUs powering the AI revolution. What you will be doing... ...skills in C++, Python, and CUDA. #LI-Hybrid Your base salary will...SeniorPerformanceWork experience placement$193.3k - $261.5k
...unparalleled ML inference and training performance.The Inference Enablement... ...boundary, our engineers build systematic... ...'s possible in AI acceleration.As part... ...wide variety of LLM model families,... ...mentorship. Our senior members enjoy one... ...Familiarity with CUDA kernels or equivalent...SeniorPerformanceWork experience placementInternshipLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Inference Performance Engineer (CUDA/LLM/VLM). Be the first to apply!
- ai prompt engineer Santa Clara, CA
- senior ai engineer Santa Clara, CA
- ai developer Santa Clara, CA
- ai engineer Santa Clara, CA
- ai engineer remote Santa Clara, CA
- senior network engineer remote Santa Clara, CA
- senior app developer Santa Clara, CA
- senior manager legal Santa Clara, CA
- sr project manager Santa Clara, CA
- senior account executive Santa Clara, CA
