Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Kernel Engineer: Build Fast AI Inference at Scale

Baseten

A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal candidates have 1–5 years of CUDA development experience and a strong understanding of GPU architecture. This position offers competitive compensation, including equity, and comprehensive benefits including medical coverage and generous PTO. #J-18808-Ljbffr Baseten

Vacancy posted 15 hours ago
Similar jobs that could be interesting for youBased on the GPU Kernel Engineer: Build Fast AI Inference at Scale in San Francisco, CA vacancy
  •  ...the Team Our Inference team brings...  ...state-of-the-art AI models, allowing...  ...We’re hiring engineers to scale and optimize OpenAI...  ...emerging GPU platforms. You’...  ...from low-level kernel performance to...  ...partner teams to build, integrate and...  ...part of a small, fast-moving team building... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    15 hours ago
  •  ...seeking a talented software engineer to join their dynamic Inference team. This role involves...  ...infrastructure for large-scale multimodal models, focusing...  ...to push the boundaries of AI technology, ensuring reliable...  .... If you thrive in fast-paced environments and enjoy... 
    Suggested

    Jobleads-US

    San Francisco, CA
    15 hours ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor...  .... Join us and help build the platform engineers turn to to ship AI products...  ...We’re seeking a GPU Kernel Engineer to join our...  .... You'll work in a fast-paced, intellectually... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    15 hours ago
  •  ...the Team We’re building high-performance...  ...models at massive scale. As part of the inference team, you’ll be...  ...GPUs by designing kernels, tuning memory layouts...  ...a kernel-focused engineer to lead efforts...  ..., and optimizing GPU kernels used in...  ...OpenAI is an AI research and deployment... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    15 hours ago
  •  ...mission-critical inference for the world's most dynamic AI companies, like Cursor...  ...Join us and help build the platform engineers turn to to ship AI...  ...-modal workloads scale, the network is...  ...engineers to lead our GPU Networking efforts...  .... Optimize Kernels: You will work with... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    15 hours ago
  • $255k - $405k

     ...Lambda is the #1 GPU Cloud for ML/AI teams training...  ...models, where engineers can easily,...  ...and affordably build, test and...  ...AI products at scale. Lambda’s product...  ...clouds and managed inference services –...  ...hardening, kernel integrity monitoring...  ...Enjoy moving fast and making a... 
    Full time
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    15 hours ago
  • $315k

    A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams... 

    Anthropic

    San Francisco, CA
    4 days ago
  • $200k - $300k

     ...Labs: We build foundational...  ...— unlocking AI's full potential...  ...research to systems engineering to product...  ...and serve as fast as the...  ...world models at scale is a novel systems...  ...— in kernels, in the serving...  ...Optimize inference and serving end...  ...Write and tune GPU kernels (CUDA... 
    Full time

    World Labs

    San Francisco, CA
    3 days ago
  •  ...state-of-the-art AI models -...  ...performance model inference and accelerating...  ...optimization, and scaling of our...  ...role, you’ll lead engineering efforts to ensure...  ...performance at the kernel level, and...  ...pipelines. Build tooling and observability...  ...engineers on GPU performance,... 
    Full time

    OpenAI

    San Francisco, CA
    15 hours ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like...  .... Join us and help build the platform engineers turn to to ship AI...  ...is building its own GPU infrastructure for large-scale inference. As we...  ...RNIC issues, host kernel stalls, GPU driver... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    15 hours ago
  •  ...mission-critical inference for the world's most dynamic AI companies, like Cursor...  ...Join us and help build the platform engineers turn to to ship AI...  ..., and improve GPU efficiency via profiling...  .... Build large-scale, real-time...  ...customization - enable fast evaluation, safe rollout... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    15 hours ago
  •  ...the Team OpenAI’s Inference team powers the...  .... We're a small, fast-moving team of engineers focused on delivering...  ...of what AI can do. We’re...  ...multimodal inference, building the infrastructure...  ...multimodal models at scale. You’ll be part...  ...including GPU utilization, tensor... 
    Full time

    OpenAI

    San Francisco, CA
    15 hours ago
  •  ...We're building the company which will de-risk...  ...When people finance GPU clusters, the datacenters...  ...for computer and inference, but sell to...  ...Otherwise, as AI scales, compute only becomes...  ...the same artifact, fast incremental feedback for every engineer, and a credible roadmap... 
    Long term contract
    Full time
    Contract work
    Fixed term contract
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Shift work

    San Francisco Compute Company

    San Francisco, CA
    15 hours ago
  • $220k

    Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience... 

    Perplexity

    San Francisco, CA
    4 days ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference at scale. The company is...  ...Senior / Staff Site Reliability Engineer to support and scale large-scale... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $190.9k - $232.8k

     ...a staff software engineer for GenAI inference, you will lead the...  ..., and robust scaling. Your work will encompass...  ...inference stack: kernels, runtimes,...  ...guide standards to build and maintain instrumentation...  ...with CUDA, GPU programming, and...  ...is the data and AI company. More... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  • $190.9k - $232.8k

     ...staff software engineer for GenAI Performance and Kernel, you will own the...  ...high-performance GPU kernels powering our GenAI inference stack. You will...  ...performance at scale.What You Will DoLead...  ...building high-performance...  ...is the data and AI company. More than... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  •  ...Our mission is to scale intelligence to serve...  ...who are building AI systems to power magical...  ...work hard and move fast to do what’s best...  ...team of researchers, engineers, designers, and...  ...with Kubernetes, and GPU workloads on those...  ...and throughput of inference. ~ Strong... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    15 hours ago
  • $300k

     ...GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its...  ...workloads? This team is building low-latency AI systems where...  ...memory hierarchy, kernel launch overhead, occupancy...  ..., profiling large-scale speech and...  ...model ideas into fast, production-ready... 
    Relocation
    Visa sponsorship
    Free visa

    techire ai

    San Francisco, CA
    5 days ago
  •  ...worldwide.We’re a team of engineers, clinicians, and...  ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will...  ...Development of Linux kernel internals, device...  ...to real-time onboard inference—while serving as a...  ...C++, Python, Bash) and building real-time, multi-... 
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    3 days ago
  • $100k - $120k

    Coda Robotics is scaling the compute infrastructure that powers...  ...models. As training and inference workloads grow, we need kernel‑level innovations to...  ...team of kernel and system engineers focused on performance-critical...  ...for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware... 

    Coda Robotics

    San Francisco, CA
    2 days ago
  • $185k

    About the RoleThe Engineering Acceleration team builds and operates the foundational systems...  ...integration systems for a fast-growing engineering organization...  ...bottlenecks.Use modern AI tools to rethink CI failure...  ...or operated CI systems at scale, especially in environments... 
    Work at office
    Local area
    Remote work
    Flexible hours

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...COMPANY We're building autonomous...  ...operate the inference systems that...  ...This is an engineering role, not a research...  ...their work fast and reliable...  ...quantization, custom kernels, scheduling...  ...grade, large‑scale serving...  ...Experience with GPU‑accelerated...  ...#J-18808-Ljbffr MakerMaker.AI

    MakerMaker.AI

    San Francisco, CA
    2 days ago
  • Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack for real-...  ...You will primarily work on uzu, our inference engine, and focus on supporting...  ...experience in writing high-performance GPU kernels or Rust systems programming. We... 
    Local area

    Mirai Labs

    San Francisco, CA
    4 days ago
  • $250k - $300k

     ...vertically integrated AI infrastructure...  ...a team that is building the future,...  ...believe in the scale of our ambition...  ...owning the inference stack end to end...  ...with customer engineering teams to tailor...  ...SGLang to the CUDA kernels underneath,...  ...of concept, run fast experiments to... 
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  • $206.3k - $388k

     ...for a Principal ML Engineer to architect and scale the multimodal data...  ...engineering and applied ML building distributed, GPU-accelerated systems...  ...determine how fast and how well Adobe...  ...data Scale up inference throughput across the...  ...impact, powered by AI and driven by human... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    15 hours ago
  • $342k

     ...demands of advanced AI workloads. The...  ...for building the next generation...  ...the RoleAs an Engineer on our hardware...  ...work with our kernel, compiler and machine...  ...training and inference on our models....  ...drive decisions on scale up, scale out,...  ...of GPU and/or other AI... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    4 days ago
  •  ...Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These...  ...Engineer to help scale the engineering systems...  ...work. The systems you build directly impact OpenAI...  ...instability, GPU scheduling, or test environment...  ...OpenAI is an AI research and deployment... 
    Full time

    OpenAI

    San Francisco, CA
    15 hours ago
  • $320k

     ...interpretable, and steerable AI systems. We want AI...  ...researchers, engineers, policy experts,...  ...together to build beneficial AI systems...  ...mandate is to make inference deployment boring and...  ...deployment systems at scale and gravitate...  ...production across GPU, TPU, and Trainium... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    15 hours ago
  •  ...About the Team The scaling team team builds foundational components...  ...Role As a software engineer on the Scaling team, you...  ..., building custom kernels, contributing to compiler...  ...excited to work in a fast-paced, highly collaborative...  ...OpenAI is an AI research and deployment... 
    Full time
    Work at office
    Local area
    Relocation package
    3 days per week

    OpenAI

    San Francisco, CA
    15 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Kernel Engineer: Build Fast AI Inference at Scale. Be the first to apply!