Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA

NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models and AI workloads. You will collaborate across teams, work on CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and participate in open source projects. #J-18808-Ljbffr NVIDIA

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior AI Inference Systems Engineer | GPU Kernels & Runtime in Santa Clara, CA vacancy
  •  ...Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You will create libraries, code generators, and GPU kernel innovations for LLM workloads. Join a team that... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    3 days ago
  • NVIDIA is seeking outstanding AI systems engineers in Santa Clara to advance the inference software stack. You will build libraries, code generators, and GPU kernels for NVIDIA hardware, designing abstractions for LLM serving engines and JIT compilers to accelerate large... 
    Suggested

    Segment (Twilio)

    Santa Clara, CA
    3 days ago
  • NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs with limited resources. You will collaborate...  ...while optimizing AI models, data pipelines, and inference runtimes for performance on next-generation GPUs. The role... 
    Senior
    Local area

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA is seeking an experienced AI systems engineer to innovate and develop cutting-edge technologies in AI inference systems. You will design and optimize kernel technologies to accelerate workloads for NVIDIA's hardware architecture. The ideal candidate holds a Master... 
    Senior

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it...  ..., and agentic optimization systems that improve GPU kernels at...  ...kernel optimization: applying AI-driven analysis to diagnose...  ..., compiler decisions, and runtime scheduling.Direct... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency...  ...high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $356.5k

    NVIDIA is hiring a Senior Engineer to join its GPU Software team in Santa Clara, California. This role focuses on designing and developing GPU kernel drivers and embedded software, impacting both...  ...and kernel expertise in various systems. Attractive salary range from 18... 
    Senior

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $160k - $253k

     ...transforming into AI factories, and NVIDIA...  ...computing is the engine of artificial intelligence...  ...are looking for a Senior Technical Marketing...  ...NVIDIA's GPU architecture, server...  ...efficiency for AI inference & training.What you...  ...GPU and rack-scale systems. This role bridges... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $182k - $242k

     ...Essential Cloud for AI™. Built for...  ...high-performance GPU infrastructure across...  ..., and real-time inference. Our stack is engineered for speed, scale,...  ...re looking for a Senior Engineer for CoreWeave...  ...team, focused on kernel authoring and...  ...-critical systems. ~ Hands-on CUDA... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    18 days ago
  • CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel-level optimization for LLM inference and end-to-end model serving, focusing on CUDA kernels and throughput/latency improvements. You will lead kernel design reviews, mentor engineers... 
    Senior

    Neura Market

    Sunnyvale, CA
    1 day ago
  • CoreWeave is seeking a Senior Software Engineer for the Systems Engineering team to own kernel tracing and patching across Kubernetes and container runtimes. You will debug complex failures, trace root causes in the Linux kernel, and upstream fixes where appropriate. This... 
    Senior

    CoreWeave

    Sunnyvale, CA
    5 hours ago
  • Accellor is seeking a Technical Architect — AI Systems, Inference & Platform Internals to design, scale, and optimize internal AI systems...  ...OpenAI API workloads. The role focuses on inference runtime, model serving, GPU infrastructure, and distributed systems for production... 
    Senior

    Accellor

    Mountain View, CA
    2 days ago
  • $184k - $287.5k

     ...Deep Learning Systems. As the complexity...  ...for emerging AI workloads. You...  ...optimization, custom kernel development,...  ...from a single GPU to...  ...both training and inference pipelines.Collaborate...  ...exploratory tools and runtime systems to...  ...Science, Computer Engineering, Electrical... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    15 hours ago
  •  ...computing experiences—from AI and data centers, to...  ...gaming and embedded systems. Grounded in a...  ...are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance...  ...the software runtime, and drive...  ...compute utilization, kernel scheduling, memory allocation... 
    Senior

    AMD

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...'re looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build...  ..., code generators, and GPU kernel technologies for NVIDIA's hardware...  ..., new LLM inference runtimes components, and kernel code... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    15 hours ago
  • NVIDIA in Santa Clara, CA seeks a senior software engineer focusing on GPU computing and ML inference to optimize LLM workloads on edge AI hardware. You will track open-source inference frameworks, map architectures to NVIDIA GPUs, and report performance metrics. With... 
    Senior

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...unlimited potential of AI to define the...  ...in which our GPU acts as the...  ...open-source LLM inference frameworks —...  ...Science, Computer Engineering, Electrical...  ...computing, ML systems, or high-performance...  ...with GPU kernel development or...  ...optimization, runtime configuration,... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture...  ...critical in enhancing GPU kernels, deep learning models, and training/inference performance across...  ...technologies and advanced engineering principles to drive continuous... 
    Senior

    AMD

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...looking for a Senior Software Engineer for Deep Learning Inference! Would you like...  ...learning experts, GPU architects and...  ...developing System Software.Proficiency...  ...in GPU kernel programming using...  ...TensorFlow, ONNX Runtime or other ML frameworks...  .... NVIDIA uses AI tools in its... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...experiences—from AI and data...  ...and embedded systems. Grounded in...  ...is seeking a Senior Product Manager...  ...open-source GPU software stack...  ...-scale model inference on AMD Instinct...  ...influence engineering roadmaps, represent...  ...JAX, serving runtimes like vLLM and...  ...libraries, kernel, runtime, and... 
    Senior
    Remote work

    AMD

    Santa Clara, CA
    4 days ago
  • $169.78k - $338.69k

     ...and mission-driven Senior/Staff AI Infrastructure Engineer, Inference & Optimization to lead...  ...pipelines, build system-level stability...  ...performance bottlenecks and runtime anomalies....  ...utilization across target GPU architectures....  ...quantization (INT8/FP8/AWQ), kernel fusion, or graph... 
    Senior
    Full time

    DiDi Labs

    San Jose, CA
    2 days ago
  •  ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware... 
    Senior

    Jobleads-US

    Santa Clara, CA
    3 days ago
  •  ...the potential of generative AI to power the...  ...3 days per week.The role: Senior toStaff Runtime Systems EngineerWhat You Will Do:d...  ...in-memory compute for AI inference in datacenters.This position...  ...position is for runtime software engineering, working on the architecture... 
    Senior
    3 days per week

    d-Matrix

    Santa Clara, CA
    2 days ago
  •  ...computing experiences—from AI and data centers, to...  ...gaming and embedded systems. Grounded in a...  ...a Principal GenAI Inference Optimization Engineer to join our Models and...  ...workloads on AMD GPU platforms. You will...  ...multiple layers—from kernels and runtimes to frameworks and serving... 

    AMD

    San Jose, CA
    15 hours ago
  • d-Matrix is seeking a Senior to Staff Runtime Systems Engineer to design and implement runtime firmware and software for our AI compute platform. You will work on multi-core SoC, drivers, and systems software, ensuring performance and reliability while collaborating with... 
    Senior

    Entrada Ventures

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...globally. We seek a Senior Engineer to lead technical...  ...deploying advanced AI agent frameworks and local runtimes on Windows and NVIDIA...  ...powerful local inference (Nemotron models) with...  ...desktop AI operating system.What You Will Be...  ...Llama.cpp, vLLM), GPU-accelerated computing... 
    Senior
    Full time
    Local area
    Shift work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $262k - $364k

    Senior Staff Software Engineer, GPU System Software link Copy link Google Sunnyvale, CA, USA Advanced Experience...  ...launching software products. Experience in Kernel programming and firmware. Coding...  ...enhance software solutions. The AI and Infrastructure team is... 
    Senior
    Worldwide

    Google Inc.

    Sunnyvale, CA
    1 day ago
  • Senior AI Systems Performance Engineer Palo Alto, California, United States The...  ...across compiler, runtime, and hardware layers...  ...for large‑scale AI inference. Responsibilities...  ...Compiler, runtime, or kernel‑level optimization...  ...or TensorRT. Strong GPU programming skills (... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    SambaNova

    Palo Alto, CA
    5 hours ago
  • A cutting-edge AI company in California is looking...  ...of Technical Staff for Kernel/Compiler/Communication....  ...expertise in CUDA and GPU optimization, along with...  ...in performance engineering. The ideal candidate will...  ...performance kernels and optimize systems for large GPU clusters,... 
    Senior

    RadixArk

    Palo Alto, CA
    2 days ago
  • $150k - $230k

     ...About Clockwork Systems Clockwork.io –...  ...to increase GPU cluster utilization...  ...systems engineers who share a vision...  ...computing. As AI workloads grow...  ...PyTorch, NCCL, CUDA runtime—not as a user,...  .... Examples: Kernel subsystems, device...  ...stack. Senior Expectations... 
    Senior

    Clockwork.io

    Palo Alto, CA
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior AI Inference Systems Engineer | GPU Kernels & Runtime. Be the first to apply!