Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Kernel Engineer: Build Fast AI Inference at Scale

Baseten

A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal candidates have 1–5 years of CUDA development experience and a strong understanding of GPU architecture. This position offers competitive compensation, including equity, and comprehensive benefits including medical coverage and generous PTO. #J-18808-Ljbffr Baseten

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the GPU Kernel Engineer: Build Fast AI Inference at Scale in San Francisco, CA vacancy
  • Sciforium is seeking a GPU Kernel Engineer to push performance on modern accelerators. You will design and optimize custom GPU kernels...  ...to integrating ops in ML frameworks used for large-scale training and inference. Ideal candidates have 5+ years in GPU kernel development... 
    Suggested

    Sciforium

    San Francisco, CA
    1 day ago
  •  ...the Team Our Inference team brings...  ...state-of-the-art AI models, allowing...  ...We’re hiring engineers to scale and optimize OpenAI...  ...emerging GPU platforms. You’...  ...from low-level kernel performance to...  ...partner teams to build, integrate and...  ...part of a small, fast-moving team building... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    16 hours ago
  •  ...seeking a talented software engineer to join their dynamic Inference team. This role involves...  ...infrastructure for large-scale multimodal models, focusing...  ...to push the boundaries of AI technology, ensuring reliable...  .... If you thrive in fast-paced environments and enjoy... 
    Suggested

    Jobleads-US

    San Francisco, CA
    1 day ago
  • $180k - $280k

     ...frontier model lab. We build reliable and general AI systems to power...  ...shift on the scale of the...  ...We're a small, fast-moving team from...  ...024, we've been engineering the foundation for...  ...re looking for a GPU kernel engineer with deep...  ...our training and inference faster and more... 
    Suggested
    Work at office
    Visa sponsorship
    Shift work

    TypeSafe AI

    San Francisco, CA
    3 days ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor...  .... Join us and help build the platform engineers turn to to ship AI products...  ...We’re seeking a GPU Kernel Engineer to join our...  .... You'll work in a fast-paced, intellectually... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    16 hours ago
  •  ...mission-critical inference for the world's most dynamic AI companies, like Cursor...  ...Join us and help build the platform engineers turn to to ship AI...  ...-modal workloads scale, the network is...  ...engineers to lead our GPU Networking efforts...  .... Optimize Kernels: You will work with... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    16 hours ago
  • $285k - $315k

     ...looking for a Founding GPU Kernel Engineer who lives right at the...  ...different GPU architectures Build tools and methods for...  ..., AMD, and emerging AI accelerators - understand...  ...experience with large-scale scientific computing,...  ...to know why things are fast or slow on the hardware... 
    Full time
    Work at office
    Relocation package

    SF Tensor

    San Francisco, CA
    3 days ago
  • Sciforium is an AI infrastructure company developing...  ...-on support from AMD engineers the team is scaling rapidly to build the full stack powering...  ...seeking a highly skilled GPU Kernel Engineer who is passionate...  ...large-scale training and inference. This role is ideal for... 
    Flexible hours

    Sciforium

    San Francisco, CA
    1 day ago
  • $167.2k - $209k

     ...relentless in their drive to build the simplest scalable...  ...are energized by the fast-paced environment of a...  ...is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean...  ...inference engine and GPU kernel layers, ensuring our... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    DigitalOcean

    San Francisco, CA
    3 days ago
  • $315k

    A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams... 

    Anthropic

    San Francisco, CA
    16 hours ago
  • Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building... 

    Linuxcareers

    San Francisco, CA
    2 days ago
  •  ...the Team OpenAI’s Inference team powers the...  .... We're a small, fast-moving team of engineers focused on delivering...  ...of what AI can do. We’re...  ...multimodal inference, building the infrastructure...  ...multimodal models at scale. You’ll be part...  ...including GPU utilization, tensor... 
    Full time

    OpenAI

    San Francisco, CA
    16 hours ago
  •  ...mission-critical inference for the world's most dynamic AI companies, like Cursor...  ...Join us and help build the platform engineers turn to to ship AI...  ..., and improve GPU efficiency via profiling...  .... Build large-scale, real-time...  ...customization - enable fast evaluation, safe rollout... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    16 hours ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like...  .... Join us and help build the platform engineers turn to to ship AI...  ...is building its own GPU infrastructure for large-scale inference. As we...  ...RNIC issues, host kernel stalls, GPU driver... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    16 hours ago
  • LeoForce is seeking a Global Inference Library Engineer to design and optimize a...  ...library for modern AI models. You will work across...  ...of AI infrastructure, GPU programming, and low-level kernels, with exposure to...  ...technically focused startup building cutting-edge AI software... 

    Leoforce

    San Francisco, CA
    1 day ago
  • $220k

    Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience... 

    Perplexity

    San Francisco, CA
    16 hours ago
  • Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing...  ...lowest layers of the stack, optimize kernel performance, develop new request...  ...schemes, write custom GPU kernels for regimes like cascade... 

    Sail

    San Francisco, CA
    16 hours ago
  • $190.9k - $232.8k

     ...staff software engineer for GenAI Performance and Kernel, you will own the...  ...high-performance GPU kernels powering our GenAI inference stack. You will...  ...performance at scale.What You Will DoLead...  ...building high-performance...  ...is the data and AI company. More than... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    4 days ago
  • $190.9k - $232.8k

     ...a staff software engineer for GenAI inference, you will lead the...  ..., and robust scaling. Your work will encompass...  ...inference stack: kernels, runtimes,...  ...guide standards to build and maintain instrumentation...  ...with CUDA, GPU programming, and...  ...is the data and AI company. More... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    4 days ago
  • $220k

    We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack...  ...in API Gateway. GPU kernels migration to CuTe DSL....  ...directed. You do well in fast-moving environments where... 

    Perplexity

    San Francisco, CA
    1 day ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference at scale. The company is...  ...Senior / Staff Site Reliability Engineer to support and scale large-scale... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • Fluidstack is building civilization-scale infrastructure for AI, delivering specialized data center capacity with a focus on extreme ownership and velocity. This role focuses on mechanical engineering for assigned sites, reviewing cooling and piping scope, and providing... 

    Fluidstack

    San Francisco, CA
    4 days ago
  • $188k - $275k

     ...Essential Cloud for AI™. Built for...  ...enables innovators to build and scale AI with confidence...  ...Do: The Field Engineering organization at CoreWeave...  ...can train and inference on at scale,...  ...lifecycle: leading new GPU cluster bring-up...  ...have fun, and move fast!  We're in an... 
    Permanent employment
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    28 days ago
  •  ...worldwide.We’re a team of engineers, clinicians, and...  ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will...  ...Development of Linux kernel internals, device...  ...to real-time onboard inference—while serving as a...  ...C++, Python, Bash) and building real-time, multi-... 
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    16 hours ago
  • $215k - $285k

    Senior Software Engineer, GPU Sandboxes Location: San...  ...company building the compute layer for AI agents. Its platform...  ...virtualization, GPU drivers, kernels, scheduling, and...  ...problems in a small, fast-moving team....  ...model training, or inference workloads. Compensation... 
    Full time
    Work at office

    Recruiting From Scratch

    San Francisco, CA
    3 days ago
  •  ...COMPANY We're building autonomous...  ...operate the inference systems that...  ...This is an engineering role, not a research...  ...their work fast and reliable...  ...quantization, custom kernels, scheduling...  ...grade, large‑scale serving...  ...Experience with GPU‑accelerated...  ...#J-18808-Ljbffr MakerMaker.AI

    MakerMaker.AI

    San Francisco, CA
    3 days ago
  • Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack for real-...  ...You will primarily work on uzu, our inference engine, and focus on supporting...  ...experience in writing high-performance GPU kernels or Rust systems programming. We... 
    Local area

    Mirai Labs

    San Francisco, CA
    16 hours ago
  •  ...out into multiple AI inference requests running...  ...sits a large GPU fleet spread across...  ..., our inference engineers and researchers build models while...  ...GPU clusters at scale: NVIDIA hardware...  ...that makes them fast (InfiniBand or RoCE...  .... GPU kernel work in CUDA or... 
    Shift work

    Neura Market

    San Francisco, CA
    16 hours ago
  • $185k

    About the RoleThe Engineering Acceleration team builds and operates the foundational systems...  ...integration systems for a fast-growing engineering organization...  ...bottlenecks.Use modern AI tools to rethink CI failure...  ...or operated CI systems at scale, especially in environments... 
    Work at office
    Local area
    Remote work
    Flexible hours

    OpenAI

    San Francisco, CA
    2 days ago
  • Plaud Inc. is seeking senior AI researchers to join our...  ...San Francisco. You will help build and train large-scale audio/speech models and push...  ...optimization for real-time inference. You will work across the stack...  ...training, collaborating with a fast-growing team. #J-18808-... 

    Plaud

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Kernel Engineer: Build Fast AI Inference at Scale. Be the first to apply!