GPU Kernel Engineer: Build Fast AI Inference at Scale
Baseten
A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal candidates have 1–5 years of CUDA development experience and a strong understanding of GPU architecture. This position offers competitive compensation, including equity, and comprehensive benefits including medical coverage and generous PTO. #J-18808-Ljbffr Baseten
- Sciforium is seeking a GPU Kernel Engineer to push performance on modern accelerators. You will design and optimize custom GPU kernels... ...to integrating ops in ML frameworks used for large-scale training and inference. Ideal candidates have 5+ years in GPU kernel development...Suggested
- ...the Team Our Inference team brings... ...state-of-the-art AI models, allowing... ...We’re hiring engineers to scale and optimize OpenAI... ...emerging GPU platforms. You’... ...from low-level kernel performance to... ...partner teams to build, integrate and... ...part of a small, fast-moving team building...SuggestedFull time
- ...seeking a talented software engineer to join their dynamic Inference team. This role involves... ...infrastructure for large-scale multimodal models, focusing... ...to push the boundaries of AI technology, ensuring reliable... .... If you thrive in fast-paced environments and enjoy...Suggested
$180k - $280k
...frontier model lab. We build reliable and general AI systems to power... ...shift on the scale of the... ...We're a small, fast-moving team from... ...024, we've been engineering the foundation for... ...re looking for a GPU kernel engineer with deep... ...our training and inference faster and more...SuggestedWork at officeVisa sponsorshipShift work- ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor... .... Join us and help build the platform engineers turn to to ship AI products... ...We’re seeking a GPU Kernel Engineer to join our... .... You'll work in a fast-paced, intellectually...SuggestedFull timeFlexible hours
- ...mission-critical inference for the world's most dynamic AI companies, like Cursor... ...Join us and help build the platform engineers turn to to ship AI... ...-modal workloads scale, the network is... ...engineers to lead our GPU Networking efforts... .... Optimize Kernels: You will work with...Full timeFlexible hours
$285k - $315k
...looking for a Founding GPU Kernel Engineer who lives right at the... ...different GPU architectures Build tools and methods for... ..., AMD, and emerging AI accelerators - understand... ...experience with large-scale scientific computing,... ...to know why things are fast or slow on the hardware...Full timeWork at officeRelocation package- Sciforium is an AI infrastructure company developing... ...-on support from AMD engineers the team is scaling rapidly to build the full stack powering... ...seeking a highly skilled GPU Kernel Engineer who is passionate... ...large-scale training and inference. This role is ideal for...Flexible hours
$167.2k - $209k
...relentless in their drive to build the simplest scalable... ...are energized by the fast-paced environment of a... ...is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean... ...inference engine and GPU kernel layers, ensuring our...Local areaRemote workWorldwideFlexible hours$315k
A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams...- Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building...
- ...the Team OpenAI’s Inference team powers the... .... We're a small, fast-moving team of engineers focused on delivering... ...of what AI can do. We’re... ...multimodal inference, building the infrastructure... ...multimodal models at scale. You’ll be part... ...including GPU utilization, tensor...Full time
- ...mission-critical inference for the world's most dynamic AI companies, like Cursor... ...Join us and help build the platform engineers turn to to ship AI... ..., and improve GPU efficiency via profiling... .... Build large-scale, real-time... ...customization - enable fast evaluation, safe rollout...Full timeFlexible hours
- ...powers mission-critical inference for the world's most dynamic AI companies, like... .... Join us and help build the platform engineers turn to to ship AI... ...is building its own GPU infrastructure for large-scale inference. As we... ...RNIC issues, host kernel stalls, GPU driver...Full timeFlexible hours
- LeoForce is seeking a Global Inference Library Engineer to design and optimize a... ...library for modern AI models. You will work across... ...of AI infrastructure, GPU programming, and low-level kernels, with exposure to... ...technically focused startup building cutting-edge AI software...
$220k
Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience...- Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing... ...lowest layers of the stack, optimize kernel performance, develop new request... ...schemes, write custom GPU kernels for regimes like cascade...
$190.9k - $232.8k
...staff software engineer for GenAI Performance and Kernel, you will own the... ...high-performance GPU kernels powering our GenAI inference stack. You will... ...performance at scale.What You Will DoLead... ...building high-performance... ...is the data and AI company. More than...Local areaWorldwide$190.9k - $232.8k
...a staff software engineer for GenAI inference, you will lead the... ..., and robust scaling. Your work will encompass... ...inference stack: kernels, runtimes,... ...guide standards to build and maintain instrumentation... ...with CUDA, GPU programming, and... ...is the data and AI company. More...Local areaWorldwide$220k
We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack... ...in API Gateway. GPU kernels migration to CuTe DSL.... ...directed. You do well in fast-moving environments where...$250k
...opportunities? Join a rapidly scaling AI cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference at scale. The company is... ...Senior / Staff Site Reliability Engineer to support and scale large-scale...Full timeRemote work- Fluidstack is building civilization-scale infrastructure for AI, delivering specialized data center capacity with a focus on extreme ownership and velocity. This role focuses on mechanical engineering for assigned sites, reviewing cooling and piping scope, and providing...
$188k - $275k
...Essential Cloud for AI™. Built for... ...enables innovators to build and scale AI with confidence... ...Do: The Field Engineering organization at CoreWeave... ...can train and inference on at scale,... ...lifecycle: leading new GPU cluster bring-up... ...have fun, and move fast! We're in an...Permanent employmentFull timeContract workTemporary workCasual workWork at officeFlexible hours- ...worldwide.We’re a team of engineers, clinicians, and... ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will... ...Development of Linux kernel internals, device... ...to real-time onboard inference—while serving as a... ...C++, Python, Bash) and building real-time, multi-...Local areaWorldwideFlexible hours
$215k - $285k
Senior Software Engineer, GPU Sandboxes Location: San... ...company building the compute layer for AI agents. Its platform... ...virtualization, GPU drivers, kernels, scheduling, and... ...problems in a small, fast-moving team.... ...model training, or inference workloads. Compensation...Full timeWork at office- ...COMPANY We're building autonomous... ...operate the inference systems that... ...This is an engineering role, not a research... ...their work fast and reliable... ...quantization, custom kernels, scheduling... ...grade, large‑scale serving... ...Experience with GPU‑accelerated... ...#J-18808-Ljbffr MakerMaker.AI
- Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack for real-... ...You will primarily work on uzu, our inference engine, and focus on supporting... ...experience in writing high-performance GPU kernels or Rust systems programming. We...Local area
- ...out into multiple AI inference requests running... ...sits a large GPU fleet spread across... ..., our inference engineers and researchers build models while... ...GPU clusters at scale: NVIDIA hardware... ...that makes them fast (InfiniBand or RoCE... .... GPU kernel work in CUDA or...Shift work
$185k
About the RoleThe Engineering Acceleration team builds and operates the foundational systems... ...integration systems for a fast-growing engineering organization... ...bottlenecks.Use modern AI tools to rethink CI failure... ...or operated CI systems at scale, especially in environments...Work at officeLocal areaRemote workFlexible hours- Plaud Inc. is seeking senior AI researchers to join our... ...San Francisco. You will help build and train large-scale audio/speech models and push... ...optimization for real-time inference. You will work across the stack... ...training, collaborating with a fast-growing team. #J-18808-...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU Kernel Engineer: Build Fast AI Inference at Scale. Be the first to apply!



