Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Kernel Engineer: Build Fast AI Inference at Scale

Baseten

A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal candidates have 1–5 years of CUDA development experience and a strong understanding of GPU architecture. This position offers competitive compensation, including equity, and comprehensive benefits including medical coverage and generous PTO. #J-18808-Ljbffr Baseten

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the GPU Kernel Engineer: Build Fast AI Inference at Scale in San Francisco, CA vacancy
  • Sail is building cutting-edge software to run AI inference and host agents at scale. You will own token processing at the kernel level, optimize perf, and...  ...-of-the-art engines, profiling tools, and advanced GPU techniques while...  ...hands-on team in a fast-paced SF office... 
    Suggested
    Work at office

    Theory Ventures

    San Francisco, CA
    3 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 
    Suggested

    Vast.ai Inc.

    San Francisco, CA
    3 days ago
  •  ...Team OpenAI’s Inference team ensures...  ...reliably, and at scale. We build and optimize...  ...stack - including kernels, communication...  ...We’re hiring engineers to scale and optimize...  ...emerging GPU platforms. You’...  ...part of a small, fast-moving team building...  ...OpenAI is an AI research and... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...seeking a talented software engineer to join their dynamic Inference team. This role involves...  ...infrastructure for large-scale multimodal models, focusing...  ...to push the boundaries of AI technology, ensuring reliable...  .... If you thrive in fast-paced environments and enjoy... 
    Suggested

    Jobleads-US

    San Francisco, CA
    5 days ago
  • $180k - $280k

     ...frontier model lab. We build reliable and general AI systems to power...  ...shift on the scale of the...  ...We're a small, fast-moving team from...  ...024, we've been engineering the foundation for...  ...re looking for a GPU kernel engineer with deep...  ...our training and inference faster and more... 
    Suggested
    Work at office
    Visa sponsorship
    Shift work

    TypeSafe AI

    San Francisco, CA
    1 day ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor...  .... Join us and help build the platform engineers turn to to ship AI products...  ...We’re seeking a GPU Kernel Engineer to join our...  .... You'll work in a fast-paced, intellectually... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...the Team We’re building high-performance...  ...models at massive scale. As part of the inference team, you’ll be...  ...GPUs by designing kernels, tuning memory layouts...  ...a kernel-focused engineer to lead efforts...  ..., and optimizing GPU kernels used in...  ...OpenAI is an AI research and deployment... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...mission-critical inference for the world's most dynamic AI companies, like Cursor...  ...Join us and help build the platform engineers turn to to ship AI...  ...-modal workloads scale, the network is...  ...engineers to lead our GPU Networking efforts...  .... Optimize Kernels: You will work with... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will leverage your HPC background...  ...is full-time and on-site, offering equity and a fast-paced startup environment. #J-18808-Ljbffr Vast... 
    Full time

    Vast.ai

    San Francisco, CA
    3 days ago
  • $255k - $405k

     ...Lambda is the #1 GPU Cloud for ML/AI teams training...  ...models, where engineers can easily,...  ...and affordably build, test and...  ...AI products at scale. Lambda’s product...  ...clouds and managed inference services –...  ...hardening, kernel integrity monitoring...  ...Enjoy moving fast and making a... 
    Full time
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    1 day ago
  • $285k - $315k

     ...believe the future of AI and high-...  ...portable. We are building a Kernel Optimizer that automatically...  ...with researchers, engineers, and organizations...  ...for a Founding GPU Kernel Engineer...  ...experience with large‑scale scientific...  ...know why things are fast or slow on the hardware... 
    Full time
    Work at office
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    1 day ago
  • $167.2k - $209k

     ...relentless in their drive to build the simplest scalable...  ...are energized by the fast-paced environment of a...  ...is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean...  ...inference engine and GPU kernel layers, ensuring our... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    DigitalOcean

    San Francisco, CA
    1 day ago
  • $315k

    A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams... 

    Anthropic

    San Francisco, CA
    4 days ago
  • $200k - $300k

     ...Labs: We build foundational...  ...— unlocking AI's full potential...  ...research to systems engineering to product...  ...and serve as fast as the...  ...world models at scale is a novel systems...  ...— in kernels, in the serving...  ...Optimize inference and serving end...  ...Write and tune GPU kernels (CUDA... 
    Full time

    World Labs

    San Francisco, CA
    3 days ago
  • Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building... 

    Linuxcareers

    San Francisco, CA
    1 day ago
  •  ...state-of-the-art AI models -...  ...performance model inference and accelerating...  ...optimization, and scaling of our...  ...role, you’ll lead engineering efforts to ensure...  ...performance at the kernel level, and...  ...pipelines. Build tooling and observability...  ...engineers on GPU performance,... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...the Team OpenAI’s Inference team powers the...  .... We're a small, fast-moving team of engineers focused on delivering...  ...of what AI can do. We’re...  ...multimodal inference, building the infrastructure...  ...multimodal models at scale. You’ll be part...  ...including GPU utilization, tensor... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...We're building the company which will de-risk...  ...When people finance GPU clusters, the datacenters...  ...for computer and inference, but sell to...  ...Otherwise, as AI scales, compute only becomes...  ...the same artifact, fast incremental feedback for every engineer, and a credible roadmap... 
    Long term contract
    Full time
    Contract work
    Fixed term contract
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Shift work

    San Francisco Compute Company

    San Francisco, CA
    1 day ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like...  .... Join us and help build the platform engineers turn to to ship AI...  ...is building its own GPU infrastructure for large-scale inference. As we...  ...RNIC issues, host kernel stalls, GPU driver... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...mission-critical inference for the world's most dynamic AI companies, like Cursor...  ...Join us and help build the platform engineers turn to to ship AI...  ..., and improve GPU efficiency via profiling...  .... Build large-scale, real-time...  ...customization - enable fast evaluation, safe rollout... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $220k

    Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience... 

    Perplexity

    San Francisco, CA
    4 days ago
  • Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing...  ...lowest layers of the stack, optimize kernel performance, develop new request...  ...schemes, write custom GPU kernels for regimes like cascade... 

    SAIL

    San Francisco, CA
    4 days ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference at scale. The company is...  ...Senior / Staff Site Reliability Engineer to support and scale large-scale... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $190.9k - $232.8k

     ...a staff software engineer for GenAI inference, you will lead the...  ..., and robust scaling. Your work will encompass...  ...inference stack: kernels, runtimes,...  ...guide standards to build and maintain instrumentation...  ...with CUDA, GPU programming, and...  ...is the data and AI company. More... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  • $190.9k - $232.8k

     ...staff software engineer for GenAI Performance and Kernel, you will own the...  ...high-performance GPU kernels powering our GenAI inference stack. You will...  ...performance at scale.What You Will DoLead...  ...building high-performance...  ...is the data and AI company. More than... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  •  ...Our mission is to scale intelligence to serve...  ...who are building AI systems to power magical...  ...work hard and move fast to do what’s best...  ...team of researchers, engineers, designers, and...  ...with Kubernetes, and GPU workloads on those...  ...and throughput of inference. ~ Strong... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    1 day ago
  • $300k

     ...GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its...  ...workloads? This team is building low-latency AI systems where...  ...memory hierarchy, kernel launch overhead, occupancy...  ..., profiling large-scale speech and...  ...model ideas into fast, production-ready... 
    Relocation
    Visa sponsorship
    Free visa

    techire ai

    San Francisco, CA
    5 days ago
  • $220k

    We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack...  ...in API Gateway. GPU kernels migration to CuTe DSL....  ...directed. You do well in fast-moving environments where... 

    Perplexity

    San Francisco, CA
    5 days ago
  • Fluidstack is building civilization-scale infrastructure for AI, delivering specialized data center capacity with a focus on extreme ownership and velocity. This role focuses on mechanical engineering for assigned sites, reviewing cooling and piping scope, and providing... 

    Fluidstack

    San Francisco, CA
    2 days ago
  •  ...worldwide.We’re a team of engineers, clinicians, and...  ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will...  ...Development of Linux kernel internals, device...  ...to real-time onboard inference—while serving as a...  ...C++, Python, Bash) and building real-time, multi-... 
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Kernel Engineer: Build Fast AI Inference at Scale. Be the first to apply!