CUDA Engineer - Kernel Optimization
Mercor
1. Role Overview Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using profiler-guided analysis. You’ll help evaluate, optimize, and reason about GPU kernels across modern hardware environments. This is a contract-based opportunity for specialists who enjoy squeezing performance out of modern GPU architectures. 2. Key Responsibilities Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization Use profiler metrics such as L2 cache hit rate, L2 throughput, occupancy, and related signals to guide kernel improvements Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms Write, modify, and reason about C++17, Python, and GPU programming code Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes Document optimization decisions clearly, including when specific profiler metrics are or are not useful 3. Ideal Qualifications Available to work at least 20 hrs/wk Fluent in core C++ features through C++17 Working knowledge of Python and Git Fluent in at least one GPU programming model, such as CUDA, HIP, Slang, HLSL, GLSL, or related kernel programming At least 1 year of professional or graduate-level research experience working with GPUs Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels Ability to optimize GPU kernels without needing deep prior context on every algorithm Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization is a plus Experience optimizing kernels for NVIDIA Blackwell hardware is a plus Familiarity with NSight Compute is a plus Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm is a plus Open-source contributions related to GPU kernel optimization are a plus #J-18808-Ljbffr Mercor
- Baseten is seeking an Engineering Manager to lead our GPU Kernel Engineering team, directing the low-level CUDA work that accelerates Baseten's inference stack. This player-coach role combines hands-on kernel work with team leadership to maximize impact. You will steer...Suggested
- Inception is seeking engineers and scientists to design, optimize, and maintain compute foundations for large‑scale... ...You will develop high‑performance ML kernels, enable efficient low‑precision... ...large models. The role emphasizes CUDA/CuTe/Triton kernel design, memory bandwidth...Suggested
- Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This contract-based opportunity is designed... ...analysis. Expect to write C++17 and Python code, apply CUDA or HIP, and document decisions clearly. #J-18808-Ljbffr MercorSuggestedContract workFreelance
- ...last FLOP from our GPUs by designing kernels, tuning memory layouts, and optimizing model execution at the lowest... ...are looking for a kernel-focused engineer to lead efforts in writing, porting... ...role requires deep familiarity with CUDA or equivalent kernel programming environments...SuggestedFull time
- ...site ABOUT THE ROLE You’ll write and optimize the GPU kernels and supporting systems software that makes... ...models actually use. We hire kernel engineers because the gap between "this works"... ...DO Write and optimize GPU kernels (CUDA, ROCm, Triton, or similar) for training...SuggestedShift work
$167.2k - $209k
...DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean aims to be... ...the inference engine and GPU kernel layers, ensuring our... ...AMD) and their software stacks (CUDA, ROCm, TensorRT, OpenAI Triton...Local areaRemote workWorldwideFlexible hours- ...AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible... ...computation efficiency. Ideal candidates have 1-5 years of CUDA development experience and a strong understanding of GPU...
- ...accelerate model deployment. This is a founding engineer role in downtown San Francisco, full-... ...help design the core compiler, write CUDA kernels, and drive performance improvements for... ...required; strong experience with GPU optimization, CUDA, and Rust is preferred. #J-18808...Full time
$285k - $315k
...Role We're looking for a Founding GPU Kernel Engineer who lives right at the boundary between... ...then turn that knowledge into compiler optimization passes that help every model we compile... ...Solid systems programming in C++ and CUDA (or ROCm/HIP) Good understanding of how...Full timeWork at officeRelocation package$285k - $315k
SF Tensor is looking for a Founding GPU Kernel Engineer in San Francisco, specializing in GPU architecture and kernel optimization for machine learning workloads. The ideal candidate... ..., and strong programming skills in C++ and CUDA. This full-time position offers a...Full timeRelocation package- Luminal is hiring a Founding Compiler Engineer for an on-site role in downtown San Francisco... ...shape the core compiler, implement CUDA kernels, and review model performance to... ...models, with a focus on high-performance optimization and practical code paths for real-world...
$180k - $280k
...investors. Since mid-2024, we've been engineering the foundation for what comes... ...the role We're looking for a GPU kernel engineer with deep, low-level CUDA expertise to make our training and... ...more efficient. You'll write and optimize custom kernels, profile and eliminate...Work at officeVisa sponsorshipShift work$100k - $120k
...inference workloads grow, we need kernel‑level innovations to reduce... ...team to architect and optimize low‑level compute kernels, drivers... ...a team of kernel and system engineers focused on performance-critical... ...for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware...- ...and help build the platform engineers turn to to ship AI products.... ...ROLE We’re seeking a GPU Kernel Engineer to join our team at... ...powers modern AI workloads, optimizing every microsecond of computation... ...Write and optimize code using CUDA, PTX assembly, and...Full timeFlexible hours
$160k - $230k
...language models (LLMs). Our mission is to optimize inference frameworks, algorithms, and... ...anInference Frameworks and Optimization Engineer to design, develop, and optimize distributed... ...parallelism for high-performance serving.Apply CUDA graph optimizations, TensorRT/TRT-LLM...Full time$166k - $244k
...following: Machine Learning Optimization (e.g., quantization, distillation... .../TPU hardware architecture, Kernel programming, or... ...in Computer Science, Computer Engineering, or a related technical field... ...with kernel programming (e.g., CUDA, OpenCL, Vulkan, Triton), compiler...Full timeTemporary work- ...(More Big More Better). You will own optimizations on both the training and on-robot inference... ...ML optimizations anywhere: From the CUDA kernels, to ML architecture, to frontend or... ...GPT-4 to hundreds of millions of users, engineered the foundations of autonomous driving,...Full time
$342k
...accelerate innovation and enable hardware optimized specifically for AI.About the RoleAs an Engineer on our hardware optimization and... .... You will work with our kernel, compiler and machine learning engineers... ...AI acceleratorsExperience with CUDA, Triton or a related accelerator...Work at officeLocal areaRelocation packageFlexible hours$350k
Mirendil is seeking an engineer to design and optimize custom ML kernels to enhance our model development stack. This role involves working at the intersection of hardware and frontier AI research, focusing on performance optimization. The ideal candidate will have experience...- ...AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components... .... You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization across...
$250k - $300k
...direction for Crusoe's Linux kernel team, owning the roadmap and... ...while mentoring and growing the engineers around you to deliver... ...and HPC workloads. You will optimize the stack for low latency and... ...and related technologies like CUDA or ROCm.Experience with high-...Temporary work$190.9k - $232.8k
P-1285About This RoleAs a staff software engineer for GenAI Performance and Kernel, you will own the design, implementation, optimization, and correctness of the high-performance GPU... ...writing and tuning compute kernels (CUDA, Triton, OpenCL, LLVM IR, assembly or similar...Local areaWorldwide- ...Member of Technical Staff focused on kernels and GPU performance. This role involves optimizing GPU and accelerator kernels for... ...candidates have strong software engineering foundations and experience with... .... Familiarity with tools like CUDA and performance profiling is...
- ...patients worldwide.We’re a team of engineers, clinicians, and innovators... ...algorithms into performance optimized, robust, validated and... ...Virtualization: Development of Linux kernel internals, device drivers,... ...Expert in GPU Compute API - CUDA, OpenCL• Proficiency in multiple...Local areaWorldwideFlexible hours
$280k
...growing group of committed researchers, engineers, policy experts, and business leaders working... ...AI systems. About the Role As a TPU Kernel Engineer, you'll be responsible for... ...of this work will involve designing and optimizing kernels for the TPU. You will also provide...Work at officeVisa sponsorshipFlexible hours- ...corporation headquartered in San Francisco, is seeking a TPU Kernel Engineer to identify and address performance issues across ML systems,... ...including research, training, and inference. You will design and optimize kernels for the TPU and provide feedback to researchers on...
$315k
...growing group of committed researchers, engineers, policy experts, and business leaders working... ...AI systems. About the Role As a TPU Kernel Engineer, you'll be responsible for... ...of this work will involve designing and optimizing kernels for the TPU. You will also provide...Contract workFor contractorsFor subcontractorWork at officeRelocationVisa sponsorshipWork visaFlexible hours- MakerMaker.AI in San Francisco is seeking a skilled Software Engineer to write and optimize GPU kernels. You will work on deep low-level tasks that directly impact the performance of machine learning models. The ideal candidate has over 4 years of experience with GPU kernels...
- Magic is hiring a Kernel Engineer in San Francisco to design, implement, and optimize high-performance kernels for long-context training and inference. You will tackle memory usage, data movement, and throughput challenges in real-time workloads. You’ll work across training...Visa sponsorship
- San Francisco Tensor Company is seeking a Founding GPU Kernel Engineer to enhance GPU performance for AI applications. You will optimize and write kernels while collaborating with compiler teams to improve efficiencies across architectures. The ideal candidate has deep...Work at officeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to CUDA Engineer - Kernel Optimization. Be the first to apply!

