Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer - GPU Kernels

Baseten

GPU Kernel Engineer

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products.

We're seeking a GPU Kernel Engineer to join our team at the cutting edge of AI acceleration, where your code directly impacts the performance of state-of-the-art machine learning models. As a GPU Kernel Engineer, you'll craft the foundation that powers modern AI workloads, optimizing every microsecond of computation to enable breakthrough applications.

You'll work in a fast-paced, intellectually stimulating environment where technical excellence is paramount and your contributions directly influence production systems serving millions of users across numerous products. This role offers exceptional growth potential for engineers passionate about low-level optimization and high-impact systems work.

You'll get to work on these types of projects as part of our Model Performance team:

  • Baseten Embeddings Inference: The fastest embeddings solution available
  • The Baseten Inference Stack
  • Driving model performance optimization

Core Engineering Responsibilities:

  • Design and implement high-performance GPU kernels for key ML operations, including matrix multiplications, attention mechanisms, and mixture-of-experts routing
  • Write and optimize code using CUDA, PTX assembly, and architecture-specific techniques
  • Apply advanced performance optimization methods such as memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlap

Performance & Innovation:

  • Implement cutting-edge features like quantization (FP8/FP4), sparsity, and compute/communication overlap
  • Identify and resolve performance bottlenecks using tools like Nsight Systems, Nsight Compute, and Torch Profiler
  • Collaborate with research teams to productionize theoretical advancements

Impact & Collaboration:

  • Contribute to internal and open-source GPU libraries
  • Present technical contributions at industry conferences (e.g., NVIDIA GTC, AWS re:Invent)

Requirements:

  • Strong understanding of GPU architecture and programming paradigms: Memory hierarchy (global, shared, registers, L1/L2 cache), Thread/block/grid organization, Synchronization techniques and race condition mitigation
  • Proficient in C++ and GPU performance profiling tools
  • Knowledge of: CUDA C++ API, Memory access patterns and bandwidth optimization, Numerical precision and quantization strategies, Modern GPU features (e.g., tensor cores, async operations)

Nice to Have:

  • Experience with Transformer models and attention optimization (e.g., Flash Attention)
  • Familiarity with GPU kernel libraries: Cutlass, Triton, Thrust, CUB
  • Background in GEMM tuning and distributed/multi-GPU compute
  • Contributions to open-source GPU projects
  • Research publications or conference presentations on GPU performance

Benefits:

  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k)
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Software Engineer - GPU Kernels in San Francisco, CA vacancy
  •  ...including BOND, IVP, Spark Capital, Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a GPU Kernel Engineer to join our team at the cutting edge of AI acceleration, where your code directly... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    14 hours ago
  • $100k - $120k

    Coda Robotics is looking for an experienced engineer to join their founding team, focusing on low-level compute kernels to enhance robotic foundation models. The ideal candidate...  ...programming (C/C++, assembly), expertise in GPU optimizations, and familiarity with ML... 
    Suggested

    Coda Robotics

    San Francisco, CA
    3 days ago
  •  ...us and help build the platform engineers turn to to ship AI products....  ...Baseten is building its own GPU infrastructure for large-scale...  ...bad optics, RNIC issues, host kernel stalls, GPU driver problems, and...  ...not. We are hiring a Lead Software Engineer to build a first-class... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    14 hours ago
  •  ...progression via model inference. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across emerging GPU platforms. You’ll work across the stack - from low-level kernel performance to high-level distributed execution - and... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    14 hours ago
  •  ...and help build the platform engineers turn to to ship AI products....  ...foundational engineers to lead our GPU Networking efforts, making...  ...to architect the software fabric that unifies thousands...  ...system behaviors. Optimize Kernels: You will work with communication... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    14 hours ago
  • $215k - $285k

    Senior Software Engineer, GPU Sandboxes Location: San Francisco, CA Company Stage of Funding: Series A Office Type: In Person Salary: $21...  ..., spanning hypervisors, Linux virtualization, GPU drivers, kernels, scheduling, and the orchestration infrastructure required... 
    Full time
    Work at office

    Recruiting From Scratch

    San Francisco, CA
    3 days ago
  • $175k - $300k

     ...them - with teams spanning hardware and software. Speed and scale are our key differentiators...  ...frontier forward.   The Production Engineering Team Examples of key exciting...  ...pace with a 10 GW fleet: at our scale, a GPU failure isn't a ticket. It's a throughput... 
    Full time
    Local area

    Fluidstack

    San Francisco, CA
    14 hours ago
  •  ...generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly...  ...About the Role We are looking for a systems-minded engineer to help advance our kernel development, performance engineering, and hardware-software... 
    Full time

    OpenAI

    San Francisco, CA
    14 hours ago
  • $190.9k - $232.8k

    P-1285About This RoleAs a staff software engineer for GenAI Performance and Kernel, you will own the design, implementation, optimization, and correctness of the high-performance GPU kernels powering our GenAI inference stack. You will lead development of highly-tuned,... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    4 days ago
  • $230k

     ...responsible AI deployment over unchecked growth.About the roleAs a software engineer on the Fleet High Performance Computing (HPC) team, you will...  ...(e.g., PCIe, Infiniband, networking, power management, kernel perf tuning)Knowledge of hardware management protocols (e.g.,... 
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    3 days ago
  •  ...accelerate the abundance of energy and intelligence. We are seeking a Software Engineer to join Crusoe’s Data Center Infrastructure Engineering team, focusing on software for managing a fleet of GPU servers and the data centers that house them. You will build advanced... 

    Crusoe

    San Francisco, CA
    1 day ago
  • Give every AI agent its own GPU Senior Software Engineer – GPU Sandboxes San Francisco onsite | Relocation considered | Equity offered AI agents...  ...tracing difficult failures across the hypervisor, Linux kernel, drivers and physical hardware. Go, Rust or C/C++ can all... 
    Full time
    Relocation

    nineDots Technology Recruitment

    San Francisco, CA
    11 days ago
  •  ...Job Description Job Description Senior Neural Network Kernel Software Development Engineer Our client is making substantial investments in software to enhance the seamless deployment of neural networks on their hardware, streamlining the experience for researchers... 

    Targeted Talent

    San Francisco, CA
    12 days ago
  • Sciforium is seeking a GPU Kernel Engineer to push performance on modern accelerators. You will design and optimize custom GPU kernels, from low-level development to integrating ops in ML frameworks used for large-scale training and inference. Ideal candidates have 5+ years... 

    Sciforium

    San Francisco, CA
    1 day ago
  • $167.2k - $209k

     ...world. DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI...  ...at the inference engine and GPU kernel layers, ensuring our infrastructure extracts...  ...modern GPU families (NVIDIA/AMD) and their software stacks (CUDA, ROCm, TensorRT, OpenAI Triton... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    DigitalOcean

    San Francisco, CA
    3 days ago
  • A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal... 

    Baseten

    San Francisco, CA
    1 day ago
  • Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This contract-based opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using... 
    Contract work
    Freelance

    Obsidian

    San Francisco, CA
    4 days ago
  • $180k - $280k

     ...FAIR, backed by top-tier investors. Since mid-2024, we've been engineering the foundation for what comes after the current "state-of-the...  ...gets things done. About the role We're looking for a GPU kernel engineer with deep, low-level CUDA expertise to make our training... 
    Work at office
    Visa sponsorship
    Shift work

    TypeSafe AI

    San Francisco, CA
    3 days ago
  •  ...from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the...  ...the role We are seeking a highly skilled GPU Kernel Engineer who is passionate about pushing...  .... You will work across the hardware-software stack, from low-level kernel development... 
    Flexible hours

    Sciforium

    San Francisco, CA
    1 day ago
  • $285k - $315k

    SF Tensor is looking for a Founding GPU Kernel Engineer in San Francisco, specializing in GPU architecture and kernel optimization for machine learning workloads. The ideal candidate has deep expertise, proven capabilities in hand-optimizing performance-critical kernels... 
    Full time
    Relocation package

    SF Tensor

    San Francisco, CA
    3 days ago
  • SF Tensor in San Francisco is hiring a Member of Technical Staff for GPU Kernel Engineering to push the limits of what the hardware can do before any search begins. You will hand-write and optimize kernels to achieve unprecedented throughput, profile at the microarchitectural... 
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    1 day ago
  • $285k - $315k

    About The Role We're looking for a Founding GPU Kernel Engineer who lives right at the boundary between hardware and software. Someone who thinks in warps, occupancy, and memory hierarchies, and can squeeze every last FLOP out of a GPU. Your job is to go deeper than anyone... 
    Full time
    Work at office
    Relocation package

    SF Tensor

    San Francisco, CA
    3 days ago
  • $100k - $120k

     ...training and inference workloads grow, we need kernel‑level innovations to reduce latency,...  ...Lead a team of kernel and system engineers focused on performance-critical code Design...  ...compute kernels for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware accelerators Find... 

    Coda Robotics

    San Francisco, CA
    3 days ago
  • $115k - $140k

     ...MITHRIL   is the critical hardware & software necessary to augment any off-the-...  ...About the Job As a   Software Engineer - Perception   at Lodestar, you’ll be...  ...and geometric algorithms using CUDA kernels, TensorRT, or other GPU acceleration frameworks Experience... 
    Permanent employment
    Full time
    Flexible hours

    Lodestar Corporation

    San Francisco, CA
    14 hours ago
  •  ...About the role We're looking for a Software Engineer to help move data from AI models to actuators...  ...You'll work close to the hardware, from kernel and firmware up through the controls and...  ...: camera drivers, image transport, GPU inference, or sensor synchronization... 
    Full time
    Internship
    Immediate start

    Gradient Robotics

    San Francisco, CA
    14 hours ago
  •  ...responsible for the architectural and engineering backbone of OpenAI’s...  ...models. Our work spans system software, networking, platform...  ...of a system, including CPU, GPU, memory subsystem, frontend,...  ...Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth... 
    Full time

    OpenAI

    San Francisco, CA
    14 hours ago
  •  ...don’t believe culture can be engineered – but when it falls into place...  ...We're looking for a software engineer to optimize and deploy...  ...—using TensorRT, custom CUDA kernels, and low-level systems engineering...  ...maximize throughput on embedded GPU platforms Collaborate with... 
    Full time
    Local area

    Humble Robotics

    San Francisco, CA
    14 hours ago
  • $139k - $257.55k

     ...experience and stronger capabilitiesBuild out next-generation GPU-accelerated features and modernize existing featuresCollaborate...  ...usersCollaborate with technology and platform teamsWrite and review engineering documents and design specsEnsure quality in all phases of... 
    Full time
    Temporary work
    Local area
    Worldwide
    Shift work

    Adobe Systems

    San Francisco, CA
    2 days ago
  • $180k - $250k

     ...Fal Engineer Position Fal is the generative media ecosystem powering the...  ...hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive. You...  ...Linux systems for AI workloads: kernel parameters, NUMA topology, CPU... 
    Local area
    Relocation package

    fal

    San Francisco, CA
    2 days ago
  •  ...Opportunity Gradient is looking for a Founding Software Engineer to own how the robots move data from AI...  ...will work close to the hardware — from kernel and firmware up through the controls and...  ...: camera drivers, image transport, GPU inference, or sensor synchronization •... 
    Full time

    Rethink recruit

    San Francisco, CA
    14 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer - GPU Kernels. Be the first to apply!