GPU Kernel Engineer: Build Fast AI Inference at Scale
Baseten
A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal candidates have 1–5 years of CUDA development experience and a strong understanding of GPU architecture. This position offers competitive compensation, including equity, and comprehensive benefits including medical coverage and generous PTO. #J-18808-Ljbffr Baseten
- Sail is building cutting-edge software to run AI inference and host agents at scale. You will own token processing at the kernel level, optimize perf, and... ...-of-the-art engines, profiling tools, and advanced GPU techniques while... ...hands-on team in a fast-paced SF office...SuggestedWork at office
- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This...Suggested
- ...Team OpenAI’s Inference team ensures... ...reliably, and at scale. We build and optimize... ...stack - including kernels, communication... ...We’re hiring engineers to scale and optimize... ...emerging GPU platforms. You’... ...part of a small, fast-moving team building... ...OpenAI is an AI research and...SuggestedFull time
- ...seeking a talented software engineer to join their dynamic Inference team. This role involves... ...infrastructure for large-scale multimodal models, focusing... ...to push the boundaries of AI technology, ensuring reliable... .... If you thrive in fast-paced environments and enjoy...Suggested
$180k - $280k
...frontier model lab. We build reliable and general AI systems to power... ...shift on the scale of the... ...We're a small, fast-moving team from... ...024, we've been engineering the foundation for... ...re looking for a GPU kernel engineer with deep... ...our training and inference faster and more...SuggestedWork at officeVisa sponsorshipShift work- ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor... .... Join us and help build the platform engineers turn to to ship AI products... ...We’re seeking a GPU Kernel Engineer to join our... .... You'll work in a fast-paced, intellectually...Full timeFlexible hours
- ...the Team We’re building high-performance... ...models at massive scale. As part of the inference team, you’ll be... ...GPUs by designing kernels, tuning memory layouts... ...a kernel-focused engineer to lead efforts... ..., and optimizing GPU kernels used in... ...OpenAI is an AI research and deployment...Full time
- ...mission-critical inference for the world's most dynamic AI companies, like Cursor... ...Join us and help build the platform engineers turn to to ship AI... ...-modal workloads scale, the network is... ...engineers to lead our GPU Networking efforts... .... Optimize Kernels: You will work with...Full timeFlexible hours
- Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will leverage your HPC background... ...is full-time and on-site, offering equity and a fast-paced startup environment. #J-18808-Ljbffr Vast...Full time
$255k - $405k
...Lambda is the #1 GPU Cloud for ML/AI teams training... ...models, where engineers can easily,... ...and affordably build, test and... ...AI products at scale. Lambda’s product... ...clouds and managed inference services –... ...hardening, kernel integrity monitoring... ...Enjoy moving fast and making a...Full timeWork at officeLocal areaWork from homeFlexible hours$285k - $315k
...believe the future of AI and high-... ...portable. We are building a Kernel Optimizer that automatically... ...with researchers, engineers, and organizations... ...for a Founding GPU Kernel Engineer... ...experience with large‑scale scientific... ...know why things are fast or slow on the hardware...Full timeWork at officeRelocation package$167.2k - $209k
...relentless in their drive to build the simplest scalable... ...are energized by the fast-paced environment of a... ...is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean... ...inference engine and GPU kernel layers, ensuring our...Local areaRemote workWorldwideFlexible hours$315k
A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams...$200k - $300k
...Labs: We build foundational... ...— unlocking AI's full potential... ...research to systems engineering to product... ...and serve as fast as the... ...world models at scale is a novel systems... ...— in kernels, in the serving... ...Optimize inference and serving end... ...Write and tune GPU kernels (CUDA...Full time- Linuxcareers in San Francisco is building AI research infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the research team. The role emphasizes extending orchestration with Kubernetes/Slurm, building...
- ...state-of-the-art AI models -... ...performance model inference and accelerating... ...optimization, and scaling of our... ...role, you’ll lead engineering efforts to ensure... ...performance at the kernel level, and... ...pipelines. Build tooling and observability... ...engineers on GPU performance,...Full time
- ...the Team OpenAI’s Inference team powers the... .... We're a small, fast-moving team of engineers focused on delivering... ...of what AI can do. We’re... ...multimodal inference, building the infrastructure... ...multimodal models at scale. You’ll be part... ...including GPU utilization, tensor...Full time
- ...We're building the company which will de-risk... ...When people finance GPU clusters, the datacenters... ...for computer and inference, but sell to... ...Otherwise, as AI scales, compute only becomes... ...the same artifact, fast incremental feedback for every engineer, and a credible roadmap...Long term contractFull timeContract workFixed term contractWork at officeLocal areaRemote workVisa sponsorshipShift work
- ...powers mission-critical inference for the world's most dynamic AI companies, like... .... Join us and help build the platform engineers turn to to ship AI... ...is building its own GPU infrastructure for large-scale inference. As we... ...RNIC issues, host kernel stalls, GPU driver...Full timeFlexible hours
- ...mission-critical inference for the world's most dynamic AI companies, like Cursor... ...Join us and help build the platform engineers turn to to ship AI... ..., and improve GPU efficiency via profiling... .... Build large-scale, real-time... ...customization - enable fast evaluation, safe rollout...Full timeFlexible hours
$220k
Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience...- Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing... ...lowest layers of the stack, optimize kernel performance, develop new request... ...schemes, write custom GPU kernels for regimes like cascade...
$250k
...opportunities? Join a rapidly scaling AI cloud infrastructure provider building a next-generation GPU platform designed for AI training, experimentation, and inference at scale. The company is... ...Senior / Staff Site Reliability Engineer to support and scale large-scale...Full timeRemote work$190.9k - $232.8k
...a staff software engineer for GenAI inference, you will lead the... ..., and robust scaling. Your work will encompass... ...inference stack: kernels, runtimes,... ...guide standards to build and maintain instrumentation... ...with CUDA, GPU programming, and... ...is the data and AI company. More...Local areaWorldwide$190.9k - $232.8k
...staff software engineer for GenAI Performance and Kernel, you will own the... ...high-performance GPU kernels powering our GenAI inference stack. You will... ...performance at scale.What You Will DoLead... ...building high-performance... ...is the data and AI company. More than...Local areaWorldwide- ...Our mission is to scale intelligence to serve... ...who are building AI systems to power magical... ...work hard and move fast to do what’s best... ...team of researchers, engineers, designers, and... ...with Kubernetes, and GPU workloads on those... ...and throughput of inference. ~ Strong...Full timeWork experience placementWork at officeRemote workFlexible hours
$300k
...GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its... ...workloads? This team is building low-latency AI systems where... ...memory hierarchy, kernel launch overhead, occupancy... ..., profiling large-scale speech and... ...model ideas into fast, production-ready...RelocationVisa sponsorshipFree visa$220k
We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack... ...in API Gateway. GPU kernels migration to CuTe DSL.... ...directed. You do well in fast-moving environments where...- Fluidstack is building civilization-scale infrastructure for AI, delivering specialized data center capacity with a focus on extreme ownership and velocity. This role focuses on mechanical engineering for assigned sites, reviewing cooling and piping scope, and providing...
- ...worldwide.We’re a team of engineers, clinicians, and... ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will... ...Development of Linux kernel internals, device... ...to real-time onboard inference—while serving as a... ...C++, Python, Bash) and building real-time, multi-...Local areaWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU Kernel Engineer: Build Fast AI Inference at Scale. Be the first to apply!



