Inference Systems Engineer — GPU Kernels & LLMs
Theory Ventures
Sail is building cutting-edge software to run AI inference and host agents at scale. You will own token processing at the kernel level, optimize perf, and design parallelism across heterogeneous hardware to maximize throughput in production. You will work with state-of-the-art engines, profiling tools, and advanced GPU techniques while collaborating with a hands-on team in a fast-paced SF office environment. #J-18808-Ljbffr Theory Ventures
- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This...Suggested
- Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will leverage your HPC background to push the bleeding edge of AI, working with CUDA/C++ and a modern tech stack. Ideal candidates have...SuggestedFull time
- A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal...Suggested
- ...the da Vinci surgical system and Ion—have transformed... ...worldwide.We’re a team of engineers, clinicians, and... ...PositionAs a Senior Systems GPU Engineer - AI &... ...: Development of Linux kernel internals, device drivers... ...models to real-time onboard inference—while serving as a core...SuggestedLocal areaWorldwideFlexible hours
- ...shipping excellence. We seek engineers with strong intrinsic... ...We’re looking for a systems engineer with HPC or... ...experience to help scale AI inference. You’ll leverage your... ...systems to optimize GPU performance at the bleeding... ...and optimize GPU kernels and tensor libraries Translate...SuggestedFull timeWork at office
- ...mission-critical inference for the world's most... ...the platform engineers turn to to ship AI... ...global operating system for distributed, heterogeneous... ...to lead our GPU Networking efforts... ...Speeds for LLMs: You will work deeply... ...behaviors. Optimize Kernels: You will work...Full timeFlexible hours
- ...Baseten powers mission-critical inference for the world's most dynamic... ...and help build the platform engineers turn to to ship AI products.... ...ROLE We’re seeking a GPU Kernel Engineer to join our team... ...directly influence production systems serving millions of users...Full timeFlexible hours
- ...About the Team OpenAI’s Inference team ensures that our... ...build and optimize the systems that power our... ...inference stack - including kernels, communication libraries... ...the Role We’re hiring engineers to scale and optimize... ...infrastructure across emerging GPU platforms. You’ll work...Full time
- ...scale. As part of the inference team, you’ll be responsible... ...our GPUs by designing kernels, tuning memory layouts,... ...for a kernel-focused engineer to lead efforts in writing... ..., and optimizing GPU kernels used in inference... ...code used in production systems. Understand GPU memory...Full time
- ...Francisco is seeking a Staff ML Systems Engineer to design and prototype... ...low-latency, high-throughput inference. You will implement changes in... ...inference engines, including kernel backends and ATLAS-style systems, while profiling across GPU, networking, and memory to improve...
$167.2k - $209k
...DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean... ...at the inference engine and GPU kernel layers, ensuring our infrastructure... ...open-source software projects. Systems Design: Excellent system design...Local areaRemote workWorldwideFlexible hours$100k - $120k
...foundation models. As training and inference workloads grow, we need kernel‑level innovations to reduce... ...Responsibilities Lead a team of kernel and system engineers focused on performance-critical... ...compute kernels for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware...$180k - $280k
...build reliable and general AI systems to power economically... ...Since mid-2024, we've been engineering the foundation for what comes... ...the role We're looking for a GPU kernel engineer with deep, low-level... ...expertise to make our training and inference faster and more efficient....Work at officeVisa sponsorshipShift work- Mirai Labs in San Francisco seeks engineers to join a senior team building the full... ...You will primarily work on uzu, our inference engine, and focus on supporting new modalities... ...in writing high-performance GPU kernels or Rust systems programming. We welcome applications...Local area
$100k - $150k
...vertically integrated AI cloud engineered for AI. We own and... ...energy, data centres, GPU superclusters,... ...stacks on AI training and inference clusters. Confident with... ...failures. ~ Linux systems engineering at scale.... ...modern Linux distributions, kernel modules, systemd,...Full timeRemote workFlexible hours- Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-... .... You will work on high-throughput systems, latency-optimized serving, and... ...building efficient inference stacks, GPU-aware optimization, and deep learning...
$220k
Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience...$142.2k - $204.6k
...This RoleAs a software engineer for GenAI inference, you will help design... ...model (LLM) serving systems are fast, scalable,... ...inference stack — from kernels and runtimes to... ...optimized for large-scale LLMs inferenceCollaborate... ...experience with CUDA, GPU programming, and key...Local areaWorldwide- ...the world’s most efficient software for inference and agent hosting. In this role, you’ll own... ...the lowest layers of the stack, optimize kernel performance, develop new request... ...exotic parallelism schemes, write custom GPU kernels for regimes like cascade attention...
- ...AI Systems Engineer - Codex Core Agents Location San Francisco Employment Type Full... ...harness issues, model behavior, inference/runtime issues, and product... ...loops. Background in compilers, kernels, runtimes, inference optimization, GPU systems, benchmarking, profiling...Full timeWork at officeLocal areaRelocation packageFlexible hours
$220k - $320k
inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques...$160k - $230k
...efficient and scalable inference for large language models (LLMs). Our mission is to... ...and Optimization Engineer to design, develop,... ...inference, GPU/accelerator optimizations... ...frameworks, distributed systems, or high‑... ...compiled, efficient kernels. Soft Skills: Strong...Full time- San Francisco Tensor Company is seeking a Founding GPU Kernel Engineer to enhance GPU performance for AI applications. You will optimize and write kernels while collaborating with compiler teams to improve efficiencies across architectures. The ideal candidate has deep...Work at officeRelocation package
- MakerMaker.AI in San Francisco is seeking a skilled Software Engineer to write and optimize GPU kernels. You will work on deep low-level tasks that directly... ...over 4 years of experience with GPU kernels, strong systems expertise, and a proven track record in kernel...
$285k - $315k
About The Role We're looking for a Founding GPU Kernel Engineer who lives right at the boundary between hardware and software. Someone who thinks... ...with low-level profiling tools: Nsight Compute, Nsight Systems, rocprof, or equivalents Experience reading and reasoning about...Full timeWork at officeRelocation package$285k - $315k
SF Tensor is looking for a Founding GPU Kernel Engineer in San Francisco, specializing in GPU architecture and kernel optimization for machine learning workloads. The ideal candidate has deep expertise, proven capabilities in hand-optimizing performance-critical kernels...Full timeRelocation package$135.2k - $306.4k
Job Overview Oracle hardware platform development engineering is seeking a highly driven GPU/CPU Platform System Engineer at the Principal Engineer level. The GPU System Engineer will work within development engineering with a small team of talented engineers who lead...Temporary workWork experience placementRemote workFlexible hours- nineDots.io is hiring a Senior Software Engineer to build secure GPU sandbox environments and scalable GPU compute platforms. You will help define... ...AI agents. You’ll work across GPU virtualization, Linux kernel, and graphics drivers, with opportunities to influence software...
- Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize...
- ...a Member of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training... ...training loops. You will own the infrastructure enabling fast inference and scalable RL iteration, balancing KV-cache strategies,...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Inference Systems Engineer — GPU Kernels & LLMs. Be the first to apply!
- lead system engineer San Francisco, CA
- distributed systems engineer San Francisco, CA
- operations support system engineer San Francisco, CA
- computer systems engineer San Francisco, CA
- system performance engineer San Francisco, CA
- unix linux systems engineer San Francisco, CA
- microsoft systems engineer San Francisco, CA
- mission system engineer San Francisco, CA
- system engineer remote San Francisco, CA
- application system engineer San Francisco, CA


