Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Kernel Engineer: Build Fast AI Inference at Scale

Baseten

A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal candidates have 1–5 years of CUDA development experience and a strong understanding of GPU architecture. This position offers competitive compensation, including equity, and comprehensive benefits including medical coverage and generous PTO. #J-18808-Ljbffr Baseten

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the GPU Kernel Engineer: Build Fast AI Inference at Scale in San Francisco, CA vacancy
  •  ...the Team Our Inference team brings...  ...state-of-the-art AI models, allowing...  ...We’re hiring engineers to scale and optimize OpenAI...  ...emerging GPU platforms. You’...  ...from low-level kernel performance to...  ...partner teams to build, integrate and...  ...part of a small, fast-moving team building... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    more than 2 months ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor...  .... Join us and help build the platform engineers turn to to ship AI products...  ...We’re seeking a GPU Kernel Engineer to join our...  .... You'll work in a fast-paced, intellectually... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  • $190k - $250k

     ...Sciforium is an AI infrastructure company developing...  ...-on support from AMD engineers the team is scaling rapidly to build the full stack powering...  ...seeking a highly skilled GPU Kernel Engineer who is passionate...  ...large-scale training and inference. This role is ideal... 
    Suggested
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    3 days ago
  • $315k

    A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams... 
    Suggested

    Anthropic

    San Francisco, CA
    22 hours ago
  •  ...mission-critical inference for the world's most dynamic AI companies, like Cursor...  ...Join us and help build the platform engineers turn to to ship AI...  ..., and improve GPU efficiency via profiling...  .... Build large-scale, real-time...  ...customization - enable fast evaluation, safe rollout... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  •  ...the Team OpenAI’s Inference team powers the...  .... We're a small, fast-moving team of engineers focused on delivering...  ...of what AI can do. We’re...  ...multimodal inference, building the infrastructure...  ...multimodal models at scale. You’ll be part...  ...including GPU utilization, tensor... 
    Full time

    OpenAI

    San Francisco, CA
    more than 2 months ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like...  .... Join us and help build the platform engineers turn to to ship AI...  ...is building its own GPU infrastructure for large-scale inference. As we...  ...RNIC issues, host kernel stalls, GPU driver... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  •  ..., is a leader in AI cloud infrastructure...  ...One person, one GPU. If you'd like to build the world's best...  ...to protect large-scale AI/ML workloads...  ...groups. Develop kernel integrity...  ...level threats. Engineer security capabilities...  ...employees, and growing fast ~ Our investors... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    more than 2 months ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference at scale. The company is...  ...Senior / Staff Site Reliability Engineer to support and scale large-scale... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • Crusoe is seeking a senior Linux kernel engineer to set the technical direction for its kernel team...  ...production-ready infrastructure at scale. You will mentor engineers, backport CVEs...  ...management, scheduling, networking, storage, and GPU subsystems. #J-18808-Ljbffr Socket.dev

    Socket.dev

    San Francisco, CA
    3 days ago
  •  ...COMPANY We're building autonomous...  ...operate the inference systems that...  ...This is an engineering role, not a research...  ...their work fast and reliable...  ...quantization, custom kernels, scheduling...  ...grade, large‑scale serving...  ...Experience with GPU‑accelerated...  ...#J-18808-Ljbffr MakerMaker.AI

    MakerMaker.AI

    San Francisco, CA
    4 days ago
  • Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack for real-...  ...You will primarily work on uzu, our inference engine, and focus on supporting...  ...experience in writing high-performance GPU kernels or Rust systems programming. We... 
    Local area

    Mirai Labs

    San Francisco, CA
    22 hours ago
  • $100k - $120k

    Coda Robotics is scaling the compute infrastructure that powers...  ...models. As training and inference workloads grow, we need kernel‑level innovations to...  ...team of kernel and system engineers focused on performance-critical...  ...for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware... 

    Coda Robotics

    San Francisco, CA
    4 days ago
  • SF Tensor is building the fastest GPU compiler and a versatile kernel optimizer. We seek a Member of Technical Staff for GPU Kernel Engineering to push the envelope on what hardware can do before...  ...pre-training, post-training and inference across NVIDIA, AMD, TPU and Trainium... 

    SF Tensor

    San Francisco, CA
    4 days ago
  • $161.3k - $241.9k

     ...frontier agentic AI, an enterprise-grade...  ...chance to help build a generational company...  ...support. We’re scaling fast and defining a new...  ...interaction, every model inference, and every...  ...for a Production Engineer to help build and...  ...Experience operating GPU fleets, high-performance... 
    Full time

    Harvey, Inc.

    San Francisco, CA
    a month ago
  • $175k - $300k

     ...had to do. Powerful AI will be the biggest...  .... There are groups building AI who don't share...  ...software. Speed and scale are our key...  ...everything forward as fast as possible. First...  ...The Production Engineering Team Examples of...  ...: at our scale, a GPU failure isn't a ticket... 
    Full time
    Local area

    Fluidstack

    San Francisco, CA
    more than 2 months ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor,...  .... Join us and help build the platform engineers turn to to ship AI products...  ...that powers large-scale LLM inference across our...  ...runtimes, networking, and GPU workloads Make thoughtful... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  •  ...fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the...  ...horizon. Our customers build fast-growing businesses around...  ...growth. The Fraud Engineering team works within our Applied...  ...About OpenAI OpenAI is an AI research and deployment... 
    Full time
    Immediate start

    OpenAI

    San Francisco, CA
    more than 2 months ago
  •  ...mission-critical inference for the world'...  ...most dynamic AI companies,...  ...Join us and help build the platform engineers turn to to...  ...platform are fast, reliable, and...  ...AI models at scale. RESPONSIBILITIES...  ...TensorRT-LLM kernels, analyze CUDA...  ...across multi-GPU setups Productionize... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  •  ...About the Team Our Inference team brings OpenAI’...  ...start-of-the-art AI models, allowing them...  ...are looking for an engineer who wants to take...  ...inference stack. Build tools to give us visibility...  ...and every GB of GPU RAM of our hardware...  ...rapidly increasing scale. Are self-... 
    Full time

    OpenAI

    San Francisco, CA
    more than 2 months ago
  • $320k

     ...interpretable, and steerable AI systems. We want AI...  ...researchers, engineers, policy experts,...  ...together to build beneficial AI systems...  ...mandate is to make inference deployment boring and...  ...deployment systems at scale and gravitate...  ...production across GPU, TPU, and Trainium... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    more than 2 months ago
  •  ...Tokens-as-a-Service (TaaS) Engineer to help build the systems that convert large-scale infrastructure capacity...  ...stack, ensuring GPU capacity can be onboarded...  ...Experience with GPU clusters, AI infrastructure,...  ...Familiarity with model porting, inference/training workloads, token... 
    Full time

    OpenAI

    San Francisco, CA
    more than 2 months ago
  •  ...enabling data and AI teams to solve the...  ...breakthroughs. We do this by building and running the...  .... Founded by engineers — and customer-...  ...interfacing with data to scaling our services and...  ...Foundation Model Inference team is the...  ...infrastructure, or GPU orchestration ~... 
    Full time

    Databricks

    San Francisco, CA
    more than 2 months ago
  • $175k - $250k

     ...We're a well-funded AI infrastructure startup...  ...We're looking for an engineer to help build and maintain a high-performance inference library designed to support...  ...ROCm, Triton, or similar GPU/accelerator programming...  ...performance-critical compute kernels Understanding of... 
    Local area

    Jobot

    San Francisco, CA
    4 days ago
  • $188k - $275k

     ...Essential Cloud for AI™. Built for...  ...enables innovators to build and scale AI with confidence...  ...Do: The Field Engineering organization at CoreWeave...  ...can train and inference on at scale,...  ...lifecycle: leading new GPU cluster bring-up...  ...have fun, and move fast!  We're in an... 
    Permanent employment
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    19 days ago
  •  ...efficiency through AI-driven computer vision...  .... Our Solutions Engineering team ensures customers...  ...adopt, deploy, and scale Voxel across...  ...iteration Design and build perception pipeline...  ..., evaluation, and inference optimization (Triton...  ...validate it with a fast proof of concept... 
    Full time
    Work at office

    Voxel

    San Francisco, CA
    3 days ago
  • $180k - $250k

     ...the next generation of AI products. We build the infrastructure,...  ...production, and do it at scale without compromise....  ...high-performance inference, orchestration, and observability...  ...experienced software engineer who thrives on...  ..., scheduling, GPU autoscaling, large scale... 
    Full time
    Currently hiring
    Remote work
    Relocation package

    Falò

    San Francisco, CA
    a month ago
  •  ...in San Francisco is hiring a Member of Technical Staff for GPU Kernel Engineering to push the limits of what the hardware can do before any search...  ...microarchitectural level, reason about PTX and SASS, and build models that feed the compiler search space. Relocation assistance... 
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    2 days ago
  • $220k - $320k

    inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over...  ...in ML systems and GPU programming. Key responsibilities...  ..., and values curiosity and fast learning. You will join a team... 

    inference.net

    San Francisco, CA
    2 days ago
  •  ...alternative chips. Inference is fragmenting:...  ...faster than existing GPU-based competitors....  ...frontier labs, fast-growing AI application companies...  ...it. You'll build and own the inference...  ...traffic and real scale. This is a founding...  ...backlog. As a founding engineer, you'll help... 

    General Compute

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Kernel Engineer: Build Fast AI Inference at Scale. Be the first to apply!