Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

CUDA Engineer - Kernel Optimization

Mercor

1. Role Overview Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using profiler-guided analysis. You’ll help evaluate, optimize, and reason about GPU kernels across modern hardware environments. This is a contract-based opportunity for specialists who enjoy squeezing performance out of modern GPU architectures. 2. Key Responsibilities Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization Use profiler metrics such as L2 cache hit rate, L2 throughput, occupancy, and related signals to guide kernel improvements Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms Write, modify, and reason about C++17, Python, and GPU programming code Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes Document optimization decisions clearly, including when specific profiler metrics are or are not useful 3. Ideal Qualifications Available to work at least 20 hrs/wk Fluent in core C++ features through C++17 Working knowledge of Python and Git Fluent in at least one GPU programming model, such as CUDA, HIP, Slang, HLSL, GLSL, or related kernel programming At least 1 year of professional or graduate-level research experience working with GPUs Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels Ability to optimize GPU kernels without needing deep prior context on every algorithm Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization is a plus Experience optimizing kernels for NVIDIA Blackwell hardware is a plus Familiarity with NSight Compute is a plus Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm is a plus Open-source contributions related to GPU kernel optimization are a plus #J-18808-Ljbffr Mercor

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the CUDA Engineer - Kernel Optimization in San Francisco, CA vacancy
  • Baseten is seeking an Engineering Manager to lead our GPU Kernel Engineering team, directing the low-level CUDA work that accelerates Baseten's inference stack. This player-coach role combines hands-on kernel work with team leadership to maximize impact. You will steer... 
    Suggested

    Baseten

    San Francisco, CA
    5 days ago
  • Inception is seeking engineers and scientists to design, optimize, and maintain compute foundations for large‑scale...  ...You will develop high‑performance ML kernels, enable efficient low‑precision...  ...large models. The role emphasizes CUDA/CuTe/Triton kernel design, memory bandwidth... 
    Suggested

    Inception

    San Francisco, CA
    5 days ago
  • Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This contract-based opportunity is designed...  ...analysis. Expect to write C++17 and Python code, apply CUDA or HIP, and document decisions clearly. #J-18808-Ljbffr Mercor
    Suggested
    Contract work
    Freelance

    Mercor

    San Francisco, CA
    2 days ago
  •  ...last FLOP from our GPUs by designing kernels, tuning memory layouts, and optimizing model execution at the lowest...  ...are looking for a kernel-focused engineer to lead efforts in writing, porting...  ...role requires deep familiarity with CUDA or equivalent kernel programming environments... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...site ABOUT THE ROLE You’ll write and optimize the GPU kernels and supporting systems software that makes...  ...models actually use. We hire kernel engineers because the gap between "this works"...  ...DO Write and optimize GPU kernels (CUDA, ROCm, Triton, or similar) for training... 
    Suggested
    Shift work

    MakerMaker.AI

    San Francisco, CA
    1 day ago
  • $167.2k - $209k

     ...DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean aims to be...  ...the inference engine and GPU kernel layers, ensuring our...  ...AMD) and their software stacks (CUDA, ROCm, TensorRT, OpenAI Triton... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    DigitalOcean

    San Francisco, CA
    1 day ago
  •  ...AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible...  ...computation efficiency. Ideal candidates have 1-5 years of CUDA development experience and a strong understanding of GPU... 

    Baseten

    San Francisco, CA
    5 days ago
  •  ...accelerate model deployment. This is a founding engineer role in downtown San Francisco, full-...  ...help design the core compiler, write CUDA kernels, and drive performance improvements for...  ...required; strong experience with GPU optimization, CUDA, and Rust is preferred. #J-18808... 
    Full time

    Luminal

    San Francisco, CA
    2 days ago
  • $285k - $315k

     ...Role We're looking for a Founding GPU Kernel Engineer who lives right at the boundary between...  ...then turn that knowledge into compiler optimization passes that help every model we compile...  ...Solid systems programming in C++ and CUDA (or ROCm/HIP) Good understanding of how... 
    Full time
    Work at office
    Relocation package

    SF Tensor

    San Francisco, CA
    1 day ago
  • $285k - $315k

    SF Tensor is looking for a Founding GPU Kernel Engineer in San Francisco, specializing in GPU architecture and kernel optimization for machine learning workloads. The ideal candidate...  ..., and strong programming skills in C++ and CUDA. This full-time position offers a... 
    Full time
    Relocation package

    SF Tensor

    San Francisco, CA
    1 day ago
  • Luminal is hiring a Founding Compiler Engineer for an on-site role in downtown San Francisco...  ...shape the core compiler, implement CUDA kernels, and review model performance to...  ...models, with a focus on high-performance optimization and practical code paths for real-world... 

    Slope

    San Francisco, CA
    1 day ago
  • $180k - $280k

     ...investors. Since mid-2024, we've been engineering the foundation for what comes...  ...the role We're looking for a GPU kernel engineer with deep, low-level CUDA expertise to make our training and...  ...more efficient. You'll write and optimize custom kernels, profile and eliminate... 
    Work at office
    Visa sponsorship
    Shift work

    TypeSafe AI

    San Francisco, CA
    2 days ago
  • $100k - $120k

     ...inference workloads grow, we need kernel‑level innovations to reduce...  ...team to architect and optimize low‑level compute kernels, drivers...  ...a team of kernel and system engineers focused on performance-critical...  ...for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware... 

    Coda Robotics

    San Francisco, CA
    2 days ago
  •  ...and help build the platform engineers turn to to ship AI products....  ...ROLE We’re seeking a GPU Kernel Engineer to join our team at...  ...powers modern AI workloads, optimizing every microsecond of computation...  ...Write and optimize code using CUDA, PTX assembly, and... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $160k - $230k

     ...language models (LLMs). Our mission is to optimize inference frameworks, algorithms, and...  ...anInference Frameworks and Optimization Engineer to design, develop, and optimize distributed...  ...parallelism for high-performance serving.Apply CUDA graph optimizations, TensorRT/TRT-LLM... 
    Full time

    Together AI

    San Francisco, CA
    3 days ago
  • $166k - $244k

     ...following: Machine Learning Optimization (e.g., quantization, distillation...  .../TPU hardware architecture, Kernel programming, or...  ...in Computer Science, Computer Engineering, or a related technical field...  ...with kernel programming (e.g., CUDA, OpenCL, Vulkan, Triton), compiler... 
    Full time
    Temporary work

    Google

    San Francisco, CA
    3 days ago
  •  ...(More Big More Better). You will own optimizations on both the training and on-robot inference...  ...ML optimizations anywhere: From the CUDA kernels, to ML architecture, to frontend or...  ...GPT-4 to hundreds of millions of users, engineered the foundations of autonomous driving,... 
    Full time

    The Generalist

    San Francisco, CA
    1 day ago
  • $342k

     ...accelerate innovation and enable hardware optimized specifically for AI.About the RoleAs an Engineer on our hardware optimization and...  .... You will work with our kernel, compiler and machine learning engineers...  ...AI acceleratorsExperience with CUDA, Triton or a related accelerator... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    4 days ago
  • $350k

    Mirendil is seeking an engineer to design and optimize custom ML kernels to enhance our model development stack. This role involves working at the intersection of hardware and frontier AI research, focusing on performance optimization. The ideal candidate will have experience... 

    Mirendil

    San Francisco, CA
    1 day ago
  •  ...AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components...  .... You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization across... 

    Genesis AI

    San Francisco, CA
    1 day ago
  • $250k - $300k

     ...direction for Crusoe's Linux kernel team, owning the roadmap and...  ...while mentoring and growing the engineers around you to deliver...  ...and HPC workloads. You will optimize the stack for low latency and...  ...and related technologies like CUDA or ROCm.Experience with high-... 
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  • $190.9k - $232.8k

    P-1285About This RoleAs a staff software engineer for GenAI Performance and Kernel, you will own the design, implementation, optimization, and correctness of the high-performance GPU...  ...writing and tuning compute kernels (CUDA, Triton, OpenCL, LLVM IR, assembly or similar... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  •  ...Member of Technical Staff focused on kernels and GPU performance. This role involves optimizing GPU and accelerator kernels for...  ...candidates have strong software engineering foundations and experience with...  .... Familiarity with tools like CUDA and performance profiling is... 

    Gimlet Labs

    San Francisco, CA
    3 days ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ...algorithms into performance optimized, robust, validated and...  ...Virtualization: Development of Linux kernel internals, device drivers,...  ...Expert in GPU Compute API - CUDA, OpenCL• Proficiency in multiple... 
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    3 days ago
  • $280k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  ...AI systems. About the Role As a TPU Kernel Engineer, you'll be responsible for...  ...of this work will involve designing and optimizing kernels for the TPU. You will also provide... 
    Work at office
    Visa sponsorship
    Flexible hours

    Neura Market

    San Francisco, CA
    5 days ago
  •  ...corporation headquartered in San Francisco, is seeking a TPU Kernel Engineer to identify and address performance issues across ML systems,...  ...including research, training, and inference. You will design and optimize kernels for the TPU and provide feedback to researchers on... 

    Neura Market

    San Francisco, CA
    5 days ago
  • $315k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  ...AI systems. About the Role As a TPU Kernel Engineer, you'll be responsible for...  ...of this work will involve designing and optimizing kernels for the TPU. You will also provide... 
    Contract work
    For contractors
    For subcontractor
    Work at office
    Relocation
    Visa sponsorship
    Work visa
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  • MakerMaker.AI in San Francisco is seeking a skilled Software Engineer to write and optimize GPU kernels. You will work on deep low-level tasks that directly impact the performance of machine learning models. The ideal candidate has over 4 years of experience with GPU kernels... 

    MakerMaker.AI

    San Francisco, CA
    2 days ago
  • Magic is hiring a Kernel Engineer in San Francisco to design, implement, and optimize high-performance kernels for long-context training and inference. You will tackle memory usage, data movement, and throughput challenges in real-time workloads. You’ll work across training... 
    Visa sponsorship

    Magic

    San Francisco, CA
    1 day ago
  • San Francisco Tensor Company is seeking a Founding GPU Kernel Engineer to enhance GPU performance for AI applications. You will optimize and write kernels while collaborating with compiler teams to improve efficiencies across architectures. The ideal candidate has deep... 
    Work at office
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to CUDA Engineer - Kernel Optimization. Be the first to apply!