Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff AI Inference Kernel Engineer

Sail

Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing down to the lowest layers of the stack, optimize kernel performance, develop new request scheduling and parallelism strategies, and help us use a heterogeneous mix of hardware at max efficiency. You’ll design and implement exotic parallelism schemes, write custom GPU kernels for regimes like cascade attention, and understand every microsecond of GPU time spent during a #J-18808-Ljbffr Sail

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Staff AI Inference Kernel Engineer in San Francisco, CA vacancy
  • A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal... 
    Suggested

    Baseten

    San Francisco, CA
    1 day ago
  • OpenAI in San Francisco seeks an experienced Software Engineer to help bring inference workloads to AWS Trainium and build the software stack to run...  ...platform. This deeply technical, cross‑stack role covers kernels, compilers, and model execution. You will develop high‑... 
    Suggested

    Slope

    San Francisco, CA
    3 days ago
  • Magic is hiring a Kernel Engineer in San Francisco to design, implement, and optimize high-performance...  ...kernels for long-context training and inference. You will tackle memory usage, data...  ...extensive testing, and quality. Experience with AI accelerators and GPU kernel frameworks... 
    Suggested
    Visa sponsorship

    Magic

    San Francisco, CA
    2 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 
    Suggested

    Vast.ai Inc.

    San Francisco, CA
    4 days ago
  • An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure...  ...researchers and product teams to push the boundaries of AI technology, ensuring reliable production services. If... 
    Suggested

    Jobleads-US

    San Francisco, CA
    1 day ago
  • Crusoe in San Francisco is seeking a Staff Technical Program Manager to lead the Managed Inference platform team. You will ensure end-to-end program delivery for LLM workloads, driving innovation in AI infrastructure powered by clean energy. The ideal candidate has extensive... 

    Crusoe

    San Francisco, CA
    5 days ago
  •  ...is seeking a Member of Technical Staff to own coverage of the serverless inference landscape. You will benchmark endpoints...  ...and keep us ahead in frontier AI benchmarking. You will analyze metrics...  ...frontiers, collaborating with engineers and industry leaders. #J-18808-Ljbffr... 

    Artificial Analysis, Inc.

    San Francisco, CA
    5 days ago
  •  ...is seeking a Product Manager to own LiveKit Inference—the gateway for developers to access the best models for voice AI through a single integration. The PM team is...  ...will set the vision and roadmap, partner with engineering, manage model providers and deployment, and aim... 
    Remote job

    LiveKit

    San Francisco, CA
    3 days ago
  • $180k - $280k

     ...We build reliable and general AI systems to power economically...  .... Since mid-2024, we've been engineering the foundation for what comes...  ...role We're looking for a GPU kernel engineer with deep, low-level...  ...to make our training and inference faster and more efficient. You... 
    Work at office
    Visa sponsorship
    Shift work

    TypeSafe AI

    San Francisco, CA
    3 days ago
  • Jobot is seeking an engineer to build and maintain a high-performance inference library for modern AI models across diverse compute architectures. The role targets someone who understands how LLM inference works under the hood and aims to squeeze maximum performance from... 

    Jobot

    San Francisco, CA
    1 day ago
  • Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open... 

    Anyscale

    San Francisco, CA
    1 day ago
  • $175k - $250k

     ...bit about us We're a well-funded AI infrastructure startup...  ...Job Details We're looking for an engineer to help build and maintain a high-performance inference library designed to support modern...  ...optimizing performance-critical compute kernels Understanding of modern... 
    Local area

    Jobot

    San Francisco, CA
    1 day ago
  • $175k - $250k

    Global Inference Library Engineer Experience: Senior Level Salary: $175,000 - $250,000 per year Job Details...  ...library designed to support modern AI models across a variety of compute...  ...optimizing performance-critical compute kernels Understanding of modern transformer... 

    LeoForce

    San Francisco, CA
    1 day ago
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 

    Causal Labs

    San Francisco, CA
    4 days ago
  • Together AI in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms, architectures, and...  ...for low-latency, high-throughput inference. You will implement changes in production...  ...inference engines, including kernel backends and ATLAS-style systems... 

    Together

    San Francisco, CA
    5 days ago
  • Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control...  ..., and explore KV caching for memory/compute trade-offs in LLM inference stacks. You will contribute to deep observability, tracing... 

    Sail

    San Francisco, CA
    5 days ago
  • Sail Research in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be responsible for building high-performance schedulers and optimizing global routing while focusing... 

    Sail Research

    San Francisco, CA
    1 day ago
  • Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-based LPM initiative. You will work on high-throughput systems, latency-optimized serving, and distributed orchestration across Kubernetes, Ray, and Slurm... 

    Kindredventures

    San Francisco, CA
    1 day ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability. The ideal candidate will have 3+ years of experience in production... 

    MakerMaker.AI

    San Francisco, CA
    3 days ago
  • A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and...  ...Ideal candidates should have strong software engineering skills and experience with ML inference... 

    Gimlet Labs

    San Francisco, CA
    4 days ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ...Senior Systems GPU Engineer - AI & Robotics, you will be...  ...Virtualization: Development of Linux kernel internals, device drivers,...  ...foundational models to real-time onboard inference—while serving as a core... 
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    4 days ago
  •  ...ll write and optimize the GPU kernels and supporting systems...  ...that makes our training and inference workloads fast. This is deep,...  ...actually use. We hire kernel engineers because the gap between "this...  ...end‑to‑end impact frustrates you #J-18808-Ljbffr MakerMaker.AI
    Shift work

    MakerMaker.AI

    San Francisco, CA
    3 days ago
  • Inception is seeking engineers and scientists to design, optimize, and maintain compute foundations for large‑scale language model training and inference. You will develop high‑performance ML kernels, enable efficient low‑precision arithmetic, and improve the distributed... 

    Inception

    San Francisco, CA
    1 day ago
  • $167.2k - $209k

     ...builders in the world. DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean aims to be the...  ...optimizations at the inference engine and GPU kernel layers, ensuring our infrastructure extracts maximum... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    DigitalOcean

    San Francisco, CA
    3 days ago
  • $315k

    A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams... 

    Anthropic

    San Francisco, CA
    5 days ago
  •  ...team of YC and unicorn founders and senior engineers with deep expertise in 3D, generative...  ...'re looking for a Founding Engineer, ML Inference with deep expertise in high-performance...  ...optimizations using torch.compile, custom CUDA kernels, and specialized inference frameworks... 
    Relocation
    Visa sponsorship
    Relocation package

    Reactor

    San Francisco, CA
    2 days ago
  • $315k

     ...interpretable, and steerable AI systems. We want AI to...  ...committed researchers, engineers, policy experts, and...  ...the Role As a TPU Kernel Engineer, you'll be responsible...  ..., training, and inference. A significant portion...  ...Currently, we expect all staff to be in one of our offices... 
    Contract work
    For contractors
    For subcontractor
    Work at office
    Relocation
    Visa sponsorship
    Work visa
    Flexible hours

    Anthropic

    San Francisco, CA
    3 days ago
  • $100k - $120k

     ...generation robotic foundation models. As training and inference workloads grow, we need kernel‑level innovations to reduce latency, memory usage, and...  ...faster. Responsibilities Lead a team of kernel and system engineers focused on performance-critical code Design, implement... 

    Coda Robotics

    San Francisco, CA
    3 days ago
  • $160k - $230k

    About the Role At Together.ai, we are building state-of-the-art...  ...enable efficient and scalable inference for large language models (LLMs...  ...Frameworks and Optimization Engineer to design, develop, and optimize...  ...graph, compiled, efficient kernels. Soft Skills: Strong analytical... 
    Full time

    Togetherai

    San Francisco, CA
    1 day ago
  •  ...ROLE You build and operate the inference systems that serve our models...  ...real workloads. This is an engineering role, not a research role....  ...research team (quantization, custom kernels, scheduling improvements,...  ...systems) isn't appealing #J-18808-Ljbffr MakerMaker.AI

    MakerMaker.AI

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff AI Inference Kernel Engineer. Be the first to apply!