Software Engineer - GPU Kernels
Baseten
GPU Kernel EngineerBaseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products.We're seeking a GPU Kernel Engineer to join our team at the cutting edge of AI acceleration, where your code directly impacts the performance of state-of-the-art machine learning models. As a GPU Kernel Engineer, you'll craft the foundation that powers modern AI workloads, optimizing every microsecond of computation to enable breakthrough applications.You'll work in a fast-paced, intellectually stimulating environment where technical excellence is paramount and your contributions directly influence production systems serving millions of users across numerous products. This role offers exceptional growth potential for engineers passionate about low-level optimization and high-impact systems work.You'll get to work on these types of projects as part of our Model Performance team:Baseten Embeddings Inference: The fastest embeddings solution availableThe Baseten Inference StackDriving model performance optimizationCore Engineering Responsibilities:Design and implement high-performance GPU kernels for key ML operations, including matrix multiplications, attention mechanisms, and mixture-of-experts routingWrite and optimize code using CUDA, PTX assembly, and architecture-specific techniquesApply advanced performance optimization methods such as memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlapPerformance & Innovation:Implement cutting-edge features like quantization (FP8/FP4), sparsity, and compute/communication overlapIdentify and resolve performance bottlenecks using tools like Nsight Systems, Nsight Compute, and Torch ProfilerCollaborate with research teams to productionize theoretical advancementsImpact & Collaboration:Contribute to internal and open-source GPU librariesPresent technical contributions at industry conferences (e.g., NVIDIA GTC, AWS re:Invent)Requirements:Strong understanding of GPU architecture and programming paradigms: Memory hierarchy (global, shared, registers, L1/L2 cache), Thread/block/grid organization, Synchronization techniques and race condition mitigationProficient in C++ and GPU performance profiling toolsKnowledge of: CUDA C++ API, Memory access patterns and bandwidth optimization, Numerical precision and quantization strategies, Modern GPU features (e.g., tensor cores, async operations)Nice to Have:Experience with Transformer models and attention optimization (e.g., Flash Attention)Familiarity with GPU kernel libraries: Cutlass, Triton, Thrust, CUBBackground in GEMM tuning and distributed/multi-GPU computeContributions to open-source GPU projectsResearch publications or conference presentations on GPU performanceBenefits:Competitive compensation, including meaningful equity.100% coverage of medical, dental, and vision insurance for employee and dependentsFlexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)Paid parental leaveFertility and family-building stipend through CarrotCompany-facilitated 401(k)Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).
$100k - $120k
Coda Robotics is looking for an experienced engineer to join their founding team, focusing on low-level compute kernels to enhance robotic foundation models. The ideal candidate... ...programming (C/C++, assembly), expertise in GPU optimizations, and familiarity with ML...Suggested- Baseten is seeking an Engineering Manager to lead our GPU Kernel Engineering team, directing the low-level CUDA work that accelerates Baseten's inference stack. This player-coach role combines hands-on kernel work with team leadership to maximize impact. You will steer...Suggested
- Give every AI agent its own GPU Senior Software Engineer - GPU Sandboxes San Francisco onsite | Relocation considered | Equity offered AI agents... ...comfortable tracing difficult failures across the hypervisor, Linux kernel, drivers and physical hardware. Go, Rust or C/C++ can all...SuggestedRelocation
$215k - $285k
Senior Software Engineer, GPU Sandboxes Location: San Francisco, CA Company Stage of Funding: Series A Office Type: In Person Salary: $21... ..., spanning hypervisors, Linux virtualization, GPU drivers, kernels, scheduling, and the orchestration infrastructure required...SuggestedFull timeWork at office- ...them - with teams spanning hardware and software. Speed and scale are our key differentiators... ...the frontier forward.The Production Engineering TeamExamples of key exciting problems the... ...of 10s to 100s of GWs: at our scale, a GPU failure isn't a ticket. It's a throughput...SuggestedLocal area
- nineDots.io is hiring a Senior Software Engineer to build secure GPU sandbox environments and scalable GPU compute platforms. You will help define architecture... ...AI agents. You’ll work across GPU virtualization, Linux kernel, and graphics drivers, with opportunities to influence...
$190.9k - $232.8k
P-1285About This RoleAs a staff software engineer for GenAI Performance and Kernel, you will own the design, implementation, optimization, and correctness of the high-performance GPU kernels powering our GenAI inference stack. You will lead development of highly-tuned,...Local areaWorldwide- ...accelerate the abundance of energy and intelligence. We are seeking a Software Engineer to join Crusoe’s Data Center Infrastructure Engineering team, focusing on software for managing a fleet of GPU servers and the data centers that house them. You will build advanced...
- MakerMaker.AI in San Francisco is seeking a skilled Software Engineer to write and optimize GPU kernels. You will work on deep low-level tasks that directly impact the performance of machine learning models. The ideal candidate has over 4 years of experience with GPU kernels...
$100k - $120k
...training and inference workloads grow, we need kernel‑level innovations to reduce latency,... ...Lead a team of kernel and system engineers focused on performance-critical code Design... ...compute kernels for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware accelerators Find...- San Francisco Tensor Company is seeking a Founding GPU Kernel Engineer to enhance GPU performance for AI applications. You will optimize and write kernels while collaborating with compiler teams to improve efficiencies across architectures. The ideal candidate has deep...Work at officeRelocation package
$285k - $315k
SF Tensor is looking for a Founding GPU Kernel Engineer in San Francisco, specializing in GPU architecture and kernel optimization for machine learning workloads. The ideal candidate has deep expertise, proven capabilities in hand-optimizing performance-critical kernels...Full timeRelocation package- A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal...
$167.2k - $209k
...world. DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI... ...at the inference engine and GPU kernel layers, ensuring our infrastructure extracts... ...modern GPU families (NVIDIA/AMD) and their software stacks (CUDA, ROCm, TensorRT, OpenAI Triton...Local areaRemote workWorldwideFlexible hours$285k - $315k
About The Role We're looking for a Founding GPU Kernel Engineer who lives right at the boundary between hardware and software. Someone who thinks in warps, occupancy, and memory hierarchies, and can squeeze every last FLOP out of a GPU. Your job is to go deeper than anyone...Full timeWork at officeRelocation package$180k - $280k
...FAIR, backed by top-tier investors. Since mid-2024, we've been engineering the foundation for what comes after the current "state-of-the... ...gets things done. About the role We're looking for a GPU kernel engineer with deep, low-level CUDA expertise to make our training...Work at officeVisa sponsorshipShift work$180k - $250k
...Fal Engineer PositionFal is the generative media ecosystem powering the next... ...hands-on engineer who builds the software and processes that keep a large fleet of GPU servers healthy and productive.... ...Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning...Local areaRelocation package$293k - $385k
...responsible for the architectural and engineering backbone of OpenAI’s... ...models. Our work spans system software, networking, platform... ...of a system, including CPU, GPU, memory subsystem, frontend,... ...Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth...Work at officeLocal areaFlexible hours- ...Job Description Job Description GPU Programmer / Software Engineer Job Type: Contractor Location: Remote Job Overview We are seeking... ...using CUDA, WebGPU, or GLSL . Profile and optimize GPU kernels and shaders for performance and efficiency. Develop...Remote jobFor contractors
- ...TO SEE *ALL* OF OUR JOB OPENINGS!Senior Software Engineer — Backend PerformanceAs a Senior Software... ...runs out of room, and reach for the GPU when it earns its keep. This is engineers... ...algorithms fundamentals.Nice to haveCUDA / kernel-level optimization or hardware-aware...
- ...patients worldwide.We’re a team of engineers, clinicians, and innovators... ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you... ...Integration: Design and optimize software interfaces between robotic... ...Virtualization: Development of Linux kernel internals, device drivers,...Local areaWorldwideFlexible hours
$220k
Perplexity is looking for an engineer to join their team in San Francisco. You will work... ...engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime... ...candidate has 3+ years of experience in software engineering with a focus on ML inference...- ...seeking a Global Inference Library Engineer to design and optimize a high‑performance... ...knowledge of AI infrastructure, GPU programming, and low-level kernels, with exposure to frameworks like... ...focused startup building cutting-edge AI software. #J-18808-Ljbffr LeoForce
- A leading consulting firm is seeking a Software Engineer (C++ Systems) in San Francisco to optimize microsecond-level performance in GPU virtualization software. Ideal candidates will have elite C++ expertise, with at least 2 years of experience in low-level systems engineering...
$115k - $140k
...Software Engineer: PerceptionLos Angeles, USAbout LodestarLodestar's mission is to develop the first "Protect and Defend" capability... ...neural networks and geometric algorithms using CUDA kernels, TensorRT, or other GPU acceleration frameworksExperience with distributed...Permanent employmentFull time- ...Senior Software EngineerLambda, the superintelligence cloud, is a... ...superintelligence. One person, one GPU.If you'd like to build the... ....The Lambda Infrastructure Engineering organization forges the foundation... ...Python.Experience with Linux kernel internals and system-level...Local areaFlexible hours
- 1. Role Overview Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance...Contract workFreelance
- ...Senior Software EngineerLambda, the superintelligence cloud, is a... ...superintelligence. One person, one GPU.If you'd like to build the... ....The Lambda Infrastructure Engineering organization forges the foundation... ...Python.Experience with Linux kernel internals and system-level...Local areaFlexible hours
- B Capital is seeking a Systems Engineer to join its Compute Platform team in San Francisco. This role involves maintaining a K8s-based platform and solving complex systems challenges, focusing on GPU infrastructures and multi-cloud environments. The ideal candidate has...
- Sail builds the world’s most efficient software for inference and agent hosting. In this role... ...the lowest layers of the stack, optimize kernel performance, develop new request... ...exotic parallelism schemes, write custom GPU kernels for regimes like cascade attention...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer - GPU Kernels. Be the first to apply!
- cybersecurity software engineer San Francisco, CA
- graduate software engineer San Francisco, CA
- software developer fintech San Francisco, CA
- new graduate software engineer San Francisco, CA
- senior robotics software engineer San Francisco, CA
- software engineer visa sponsorship San Francisco, CA
- network software engineer San Francisco, CA
- software engineer remote San Francisco, CA
- part time software developer San Francisco, CA
- part time software developer remote San Francisco, CA

