Senior AI Inference Systems Engineer | GPU Kernels & Runtime
NVIDIA
NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models and AI workloads. You will collaborate across teams, work on CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and participate in open source projects. #J-18808-Ljbffr NVIDIA
- CoreWeave is hiring a Senior Engineer for its Benchmarking & Performance team to write, profile, and optimize GPU kernels on the LLM inference path. You will improve latency and throughput and collaborate with product, orchestration, and hardware teams to achieve strict...Senior
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it... ..., and agentic optimization systems that improve GPU kernels at... ...kernel optimization: applying AI-driven analysis to diagnose... ..., compiler decisions, and runtime scheduling.Direct...SeniorFull time$184k - $287.5k
...are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency... ...high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale...SeniorFull time$180k - $320k
...technology company for AI and Bitcoin mining... ...We are seeking a Senior GPU Systems & Fabric Engineer to serve as the critical... ...expertise in Linux kernel internals, GPU... ...efficient multi-tenant inference workloads and maximize... ...device drivers, and runtime libraries (CUDA, NCCL...SeniorRemote jobFull timeLocal area- Accellor is seeking a Technical Architect — AI Systems, Inference & Platform Internals to design, scale, and optimize internal AI systems... ...OpenAI API workloads. The role focuses on inference runtime, model serving, GPU infrastructure, and distributed systems for production...Senior
$160k - $253k
...transforming into AI factories, and NVIDIA... ...computing is the engine of artificial intelligence... ...are looking for a Senior Technical Marketing... ...NVIDIA's GPU architecture, server... ...efficiency for AI inference & training.What you... ...GPU and rack-scale systems. This role bridges...SeniorFull time- NVIDIA seeks a Senior Software Engineer - AI Inference Performance to push LLM/VLM workloads... ...software, and distributed runtimes. The role spans profiling,... ...developing high-performance kernels with CUDA, Triton, and... ..., kernel, networking, and GPU architecture #J-18808-Ljbffr...Senior
- MatX is building custom silicon for LLM inference and training. You will develop the host-side interface library, manage... ...format to enable safe evolution of compiler-runtime contracts. You will design the kernel ABI and Python bindings (PyO3) to move tensors to the...
$184k - $287.5k
...Deep Learning Systems. As the complexity... ...for emerging AI workloads. You... ...optimization, custom kernel development,... ...from a single GPU to... ...both training and inference pipelines.Collaborate... ...exploratory tools and runtime systems to... ...Science, Computer Engineering, Electrical...SeniorFull time$182k - $242k
...Essential Cloud for AI. Built for... ...high-performance GPU infrastructure across... ...rendering, and real-time inference. Our stack is engineered for speed, scale,... ...re looking for a Senior Engineer for... ...team, focused on kernel authoring and... ...performance-critical systems. Hands-on CUDA...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...computing experiences—from AI and data centers, to... ...gaming and embedded systems. Grounded in a... ...are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance... ...the software runtime, and drive... ...compute utilization, kernel scheduling, memory allocation...Senior
$224k - $356.5k
...unlimited potential of AI to define the... ...in which our GPU acts as the... ...open-source LLM inference frameworks —... ...Science, Computer Engineering, Electrical... ...computing, ML systems, or high-performance... ...with GPU kernel development or... ...optimization, runtime configuration,...SeniorFull timeLocal area$152k - $241.5k
NVIDIA’s invention of the GPU in 1999 sparked the growth of... ...deep learning ignited modern AI — the next era of computing... ...advancement.Are you a motivated system software engineer with a deep understanding of... ...modelsBackground with kernel mode developmentExperience with...SeniorFull time$184k - $287.5k
...'re looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build... ..., code generators, and GPU kernel technologies for NVIDIA's hardware... ..., new LLM inference runtimes components, and kernel code...SeniorFull time$184k - $287.5k
...upon which every new AI-powered application... ...built. We are seeking a Senior Software Engineer - AI Inference Performance to... ...performance limits on NVIDIA GPU-accelerated systems. Your work will span... ..., distributed runtimes, communication, CUDA kernels, and GPU...SeniorFull time- ...computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture... ...critical in enhancing GPU kernels, deep learning models, and training/inference performance across... ...technologies and advanced engineering principles to drive continuous...Senior
$150k - $230k
...About Clockwork Systems Clockwork.io –... ...to increase GPU cluster utilization... ...systems engineers who share a vision... ...computing. As AI workloads grow... ...PyTorch, NCCL, CUDA runtime—not as a user,... .... Examples: Kernel subsystems, device... ...stack. Senior Expectations...Senior- ...experiences—from AI and data... ...and embedded systems. Grounded in... ...is seeking a Senior Product Manager... ...open-source GPU software stack... ...-scale model inference on AMD Instinct... ...influence engineering roadmaps, represent... ...JAX, serving runtimes like vLLM and... ...libraries, kernel, runtime, and...SeniorRemote work
- ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware...Senior
- ...the potential of generative AI to power the... ...3 days per week.The role: Senior toStaff Runtime Systems EngineerWhat You Will Do:d... ...in-memory compute for AI inference in datacenters.This position... ...position is for runtime software engineering, working on the architecture...Senior3 days per week
- d-Matrix is seeking a Senior to Staff Runtime Systems Engineer to design and implement runtime firmware and software for our AI compute platform. You will work on multi-core SoC, drivers, and systems software, ensuring performance and reliability while collaborating with...Senior
- ...computing experiences—from AI and data centers, to... ...gaming and embedded systems. Grounded in a... ...a Principal GenAI Inference Optimization Engineer to join our Models and... ...workloads on AMD GPU platforms. You will... ...multiple layers—from kernels and runtimes to frameworks and serving...
$262k - $364k
Senior Staff Software Engineer, GPU System Software Google | Sunnyvale, CA, USA Advanced Experience owning outcomes... ...software products. Experience in Kernel programming and firmware. Coding experience... ...to push technology forward. The AI and Infrastructure team is...SeniorWorldwide- Senior AI Systems Performance Engineer Palo Alto, California, United States The... ...across compiler, runtime, and hardware layers... ...for large‑scale AI inference. Responsibilities... ...Compiler, runtime, or kernel‑level optimization... ...or TensorRT. Strong GPU programming skills (...SeniorFull timeTemporary workLocal areaFlexible hours
- A cutting-edge AI company in California is looking... ...of Technical Staff for Kernel/Compiler/Communication.... ...expertise in CUDA and GPU optimization, along with... ...in performance engineering. The ideal candidate will... ...performance kernels and optimize systems for large GPU clusters,...Senior
$184k - $287.5k
...wants to change how the AI inference industry works. With... ...disaggregation down to kernel optimization. You’ll... ...architectures adopted by 10k+ gpu clusters. You’ll teach... ..., distinguished engineers and datacenter... ...scale their businesses, systems and infrastructure. This...SeniorFull time- ...and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase... ...years of experience in full-stack AI inference performance with...Senior
$148.7k - $201.2k
...available to Generative AI customers? Do you... ...scale?AWS Hardware Engineering is looking for a Systems Development Engineer... ...resolve Linux boot and runtime issues across... ...Power, NIC, NVMe, and GPU subsystems- Build automation... ...for AI training, inference, and compute...InternshipLocal areaWorldwideFlexible hours$184k - $287.5k
A leading technology company is seeking a Senior Software Engineer for AI and DL Kernel Libraries in Santa Clara, CA. The role involves designing and optimizing... ...6+ years of experience preferably in deep learning systems. The position offers a salary range of $184,000 - $287...SeniorRemote job- NVIDIA is seeking a Senior Software Engineer to advance AI inference performance on GPU-accelerated systems. You will optimize LLM/VLM workloads, profile with Nsight tools, and... ...engines while collaborating with model, kernel, and networking teams. The role emphasizes...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Inference Systems Engineer | GPU Kernels & Runtime. Be the first to apply!
- ai prompt engineer Santa Clara, CA
- senior ai engineer Santa Clara, CA
- ai developer Santa Clara, CA
- ai engineer Santa Clara, CA
- ai engineer remote Santa Clara, CA
- senior staff systems engineer Santa Clara, CA
- application system engineer Santa Clara, CA
- system engineer contract Santa Clara, CA
- operations support system engineer Santa Clara, CA
- sr systems engineer Santa Clara, CA

