Senior AI Inference Systems Engineer | GPU Kernels & Runtime
NVIDIA
NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models and AI workloads. You will collaborate across teams, work on CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and participate in open source projects. #J-18808-Ljbffr NVIDIA
- ...Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You will create libraries, code generators, and GPU kernel innovations for LLM workloads. Join a team that...Senior
- NVIDIA is seeking outstanding AI systems engineers in Santa Clara to advance the inference software stack. You will build libraries, code generators, and GPU kernels for NVIDIA hardware, designing abstractions for LLM serving engines and JIT compilers to accelerate large...Suggested
- NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs with limited resources. You will collaborate... ...while optimizing AI models, data pipelines, and inference runtimes for performance on next-generation GPUs. The role...SeniorLocal area
$184k - $287.5k
NVIDIA is seeking an experienced AI systems engineer to innovate and develop cutting-edge technologies in AI inference systems. You will design and optimize kernel technologies to accelerate workloads for NVIDIA's hardware architecture. The ideal candidate holds a Master...Senior$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it... ..., and agentic optimization systems that improve GPU kernels at... ...kernel optimization: applying AI-driven analysis to diagnose... ..., compiler decisions, and runtime scheduling.Direct...SeniorFull time$184k - $287.5k
...are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency... ...high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale...SeniorFull time$184k - $356.5k
NVIDIA is hiring a Senior Engineer to join its GPU Software team in Santa Clara, California. This role focuses on designing and developing GPU kernel drivers and embedded software, impacting both... ...and kernel expertise in various systems. Attractive salary range from 18...Senior$160k - $253k
...transforming into AI factories, and NVIDIA... ...computing is the engine of artificial intelligence... ...are looking for a Senior Technical Marketing... ...NVIDIA's GPU architecture, server... ...efficiency for AI inference & training.What you... ...GPU and rack-scale systems. This role bridges...SeniorFull time$182k - $242k
...Essential Cloud for AI™. Built for... ...high-performance GPU infrastructure across... ..., and real-time inference. Our stack is engineered for speed, scale,... ...re looking for a Senior Engineer for CoreWeave... ...team, focused on kernel authoring and... ...-critical systems. ~ Hands-on CUDA...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel-level optimization for LLM inference and end-to-end model serving, focusing on CUDA kernels and throughput/latency improvements. You will lead kernel design reviews, mentor engineers...Senior
- CoreWeave is seeking a Senior Software Engineer for the Systems Engineering team to own kernel tracing and patching across Kubernetes and container runtimes. You will debug complex failures, trace root causes in the Linux kernel, and upstream fixes where appropriate. This...Senior
- Accellor is seeking a Technical Architect — AI Systems, Inference & Platform Internals to design, scale, and optimize internal AI systems... ...OpenAI API workloads. The role focuses on inference runtime, model serving, GPU infrastructure, and distributed systems for production...Senior
$184k - $287.5k
...Deep Learning Systems. As the complexity... ...for emerging AI workloads. You... ...optimization, custom kernel development,... ...from a single GPU to... ...both training and inference pipelines.Collaborate... ...exploratory tools and runtime systems to... ...Science, Computer Engineering, Electrical...SeniorFull time- ...computing experiences—from AI and data centers, to... ...gaming and embedded systems. Grounded in a... ...are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance... ...the software runtime, and drive... ...compute utilization, kernel scheduling, memory allocation...Senior
$184k - $287.5k
...'re looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build... ..., code generators, and GPU kernel technologies for NVIDIA's hardware... ..., new LLM inference runtimes components, and kernel code...SeniorFull time- NVIDIA in Santa Clara, CA seeks a senior software engineer focusing on GPU computing and ML inference to optimize LLM workloads on edge AI hardware. You will track open-source inference frameworks, map architectures to NVIDIA GPUs, and report performance metrics. With...Senior
$224k - $356.5k
...unlimited potential of AI to define the... ...in which our GPU acts as the... ...open-source LLM inference frameworks —... ...Science, Computer Engineering, Electrical... ...computing, ML systems, or high-performance... ...with GPU kernel development or... ...optimization, runtime configuration,...SeniorFull timeLocal area- ...computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture... ...critical in enhancing GPU kernels, deep learning models, and training/inference performance across... ...technologies and advanced engineering principles to drive continuous...Senior
$152k - $241.5k
...looking for a Senior Software Engineer for Deep Learning Inference! Would you like... ...learning experts, GPU architects and... ...developing System Software.Proficiency... ...in GPU kernel programming using... ...TensorFlow, ONNX Runtime or other ML frameworks... .... NVIDIA uses AI tools in its...SeniorFull time- ...experiences—from AI and data... ...and embedded systems. Grounded in... ...is seeking a Senior Product Manager... ...open-source GPU software stack... ...-scale model inference on AMD Instinct... ...influence engineering roadmaps, represent... ...JAX, serving runtimes like vLLM and... ...libraries, kernel, runtime, and...SeniorRemote work
$169.78k - $338.69k
...and mission-driven Senior/Staff AI Infrastructure Engineer, Inference & Optimization to lead... ...pipelines, build system-level stability... ...performance bottlenecks and runtime anomalies.... ...utilization across target GPU architectures.... ...quantization (INT8/FP8/AWQ), kernel fusion, or graph...SeniorFull time- ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware...Senior
- ...the potential of generative AI to power the... ...3 days per week.The role: Senior toStaff Runtime Systems EngineerWhat You Will Do:d... ...in-memory compute for AI inference in datacenters.This position... ...position is for runtime software engineering, working on the architecture...Senior3 days per week
- ...computing experiences—from AI and data centers, to... ...gaming and embedded systems. Grounded in a... ...a Principal GenAI Inference Optimization Engineer to join our Models and... ...workloads on AMD GPU platforms. You will... ...multiple layers—from kernels and runtimes to frameworks and serving...
- d-Matrix is seeking a Senior to Staff Runtime Systems Engineer to design and implement runtime firmware and software for our AI compute platform. You will work on multi-core SoC, drivers, and systems software, ensuring performance and reliability while collaborating with...Senior
$224k - $356.5k
...globally. We seek a Senior Engineer to lead technical... ...deploying advanced AI agent frameworks and local runtimes on Windows and NVIDIA... ...powerful local inference (Nemotron models) with... ...desktop AI operating system.What You Will Be... ...Llama.cpp, vLLM), GPU-accelerated computing...SeniorFull timeLocal areaShift work$262k - $364k
Senior Staff Software Engineer, GPU System Software link Copy link Google Sunnyvale, CA, USA Advanced Experience... ...launching software products. Experience in Kernel programming and firmware. Coding... ...enhance software solutions. The AI and Infrastructure team is...SeniorWorldwide- Senior AI Systems Performance Engineer Palo Alto, California, United States The... ...across compiler, runtime, and hardware layers... ...for large‑scale AI inference. Responsibilities... ...Compiler, runtime, or kernel‑level optimization... ...or TensorRT. Strong GPU programming skills (...SeniorFull timeTemporary workLocal areaFlexible hours
- A cutting-edge AI company in California is looking... ...of Technical Staff for Kernel/Compiler/Communication.... ...expertise in CUDA and GPU optimization, along with... ...in performance engineering. The ideal candidate will... ...performance kernels and optimize systems for large GPU clusters,...Senior
$150k - $230k
...About Clockwork Systems Clockwork.io –... ...to increase GPU cluster utilization... ...systems engineers who share a vision... ...computing. As AI workloads grow... ...PyTorch, NCCL, CUDA runtime—not as a user,... .... Examples: Kernel subsystems, device... ...stack. Senior Expectations...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Inference Systems Engineer | GPU Kernels & Runtime. Be the first to apply!
- ai engineer Santa Clara, CA
- senior ai engineer Santa Clara, CA
- ai prompt engineer Santa Clara, CA
- ai engineer remote Santa Clara, CA
- ai developer Santa Clara, CA
- distributed systems engineer Santa Clara, CA
- operations support system engineer Santa Clara, CA
- system performance engineer Santa Clara, CA
- unix linux systems engineer Santa Clara, CA
- mission system engineer Santa Clara, CA

