Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

CUDA Kernel Performance Engineer

Mercor

Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This contract-based opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to squeeze performance out of modern GPU architectures. You will analyze, optimize, and reason about GPU kernels across modern hardware, using profiler-guided analysis. Expect to write C++17 and Python code, apply CUDA or HIP, and document decisions clearly. #J-18808-Ljbffr Mercor

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the CUDA Kernel Performance Engineer in San Francisco, CA vacancy
  • Inception is seeking engineers and scientists to design, optimize, and maintain compute...  ...and inference. You will develop high‑performance ML kernels, enable efficient low‑precision arithmetic...  ...of large models. The role emphasizes CUDA/CuTe/Triton kernel design, memory... 
    Performance

    Inception

    San Francisco, CA
    5 days ago
  • 1. Role Overview Mercor is seeking GPU kernel optimization experts to contribute to a project...  ..., and the ability to improve kernel performance using profiler-guided analysis. You’ll help...  ...Python, and GPU programming code Apply CUDA, HIP, shader programming, or related kernel... 
    Performance
    Contract work
    Freelance

    Mercor

    San Francisco, CA
    2 days ago
  •  ...acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for...  ...computation efficiency. Ideal candidates have 1-5 years of CUDA development experience and a strong understanding of GPU... 
    Performance

    Baseten

    San Francisco, CA
    5 days ago
  • Luminal is hiring a Founding Compiler Engineer for an on-site role in downtown San Francisco. You will help shape the core compiler, implement CUDA kernels, and review model performance to accelerate AI pipelines. As a founding team member, you will contribute to production... 
    Performance

    Slope

    San Francisco, CA
    1 day ago
  •  ...AI compiler to accelerate model deployment. This is a founding engineer role in downtown San Francisco, full-time on-site. You will help design the core compiler, write CUDA kernels, and drive performance improvements for production AI models. The team values practical... 
    Performance
    Full time

    Luminal

    San Francisco, CA
    2 days ago
  • $285k - $315k

    SF Tensor is looking for a Founding GPU Kernel Engineer in San Francisco, specializing in GPU architecture...  ..., proven capabilities in hand-optimizing performance-critical kernels, and strong programming skills in C++ and CUDA. This full-time position offers a competitive... 
    Performance
    Full time
    Relocation package

    SF Tensor

    San Francisco, CA
    1 day ago
  • $100k - $120k

     ...and inference workloads grow, we need kernel‑level innovations to reduce latency,...  ...Responsibilities Lead a team of kernel and system engineers focused on performance-critical code Design, implement, and...  ...kernels for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware accelerators Find... 
    Performance

    Coda Robotics

    San Francisco, CA
    2 days ago
  •  ...ll write and optimize the GPU kernels and supporting systems...  ...This is deep, low-level work (performance counters, memory bandwidth, warp...  ...actually use. We hire kernel engineers because the gap between "this...  ...Write and optimize GPU kernels (CUDA, ROCm, Triton, or similar)... 
    Performance
    Shift work

    MakerMaker.AI

    San Francisco, CA
    1 day ago
  • $180k - $280k

     ...investors. Since mid-2024, we've been engineering the foundation for what comes...  ...the role We're looking for a GPU kernel engineer with deep, low-level CUDA expertise to make our training and...  ...Write, optimize, and maintain high-performance GPU kernels (e.g., in CUDA / CuTe... 
    Performance
    Work at office
    Visa sponsorship
    Shift work

    TypeSafe AI

    San Francisco, CA
    2 days ago
  • $285k - $315k

     ...Role We're looking for a Founding GPU Kernel Engineer who lives right at the boundary between...  ...attention, normalization, etc.) to set the performance ceilings Profile at the...  ...Solid systems programming in C++ and CUDA (or ROCm/HIP) Good understanding of how... 
    Performance
    Full time
    Work at office
    Relocation package

    SF Tensor

    San Francisco, CA
    1 day ago
  • $167.2k - $209k

     ...DigitalOcean is seeking a Senior Engineer 2 to play a key technical...  ...can offer the industry-leading performance for our inference services....  ...the inference engine and GPU kernel layers, ensuring our infrastructure...  ...) and their software stacks (CUDA, ROCm, TensorRT, OpenAI... 
    Performance
    Local area
    Remote work
    Worldwide
    Flexible hours

    DigitalOcean

    San Francisco, CA
    1 day ago
  • Anthropic, a public benefit corporation headquartered in San Francisco, is seeking a TPU Kernel Engineer to identify and address performance issues across ML systems, including research, training, and inference. You will design and optimize kernels for the TPU and provide... 
    Performance

    Neura Market

    San Francisco, CA
    5 days ago
  • MakerMaker.AI in San Francisco is seeking a skilled Software Engineer to write and optimize GPU kernels. You will work on deep low-level tasks that directly impact the performance of machine learning models. The ideal candidate has over 4 years of experience with GPU kernels... 
    Performance

    MakerMaker.AI

    San Francisco, CA
    2 days ago
  • Baseten is seeking an Engineering Manager to lead our GPU Kernel Engineering team, directing the low-level CUDA work that accelerates Baseten's inference stack. This player-coach...  ...building processes and culture for a high-performing, distributed kernel team. #J-18808-Ljbffr... 

    Baseten

    San Francisco, CA
    5 days ago
  • Magic is hiring a Kernel Engineer in San Francisco to design, implement, and optimize high-performance kernels for long-context training and inference. You will tackle memory usage, data movement, and throughput challenges in real-time workloads. You’ll work across training... 
    Performance
    Visa sponsorship

    Magic

    San Francisco, CA
    1 day ago
  • San Francisco Tensor Company is seeking a Founding GPU Kernel Engineer to enhance GPU performance for AI applications. You will optimize and write kernels while collaborating with compiler teams to improve efficiencies across architectures. The ideal candidate has deep... 
    Performance
    Work at office
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    2 days ago
  • $280k

     ...quickly growing group of committed researchers, engineers, policy experts, and business leaders...  ...beneficial AI systems. About the Role As a TPU Kernel Engineer, you'll be responsible for identifying and addressing performance issues across many different ML systems,... 
    Performance
    Work at office
    Visa sponsorship
    Flexible hours

    Neura Market

    San Francisco, CA
    5 days ago
  • $315k

     ...quickly growing group of committed researchers, engineers, policy experts, and business leaders...  ...beneficial AI systems. About the Role As a TPU Kernel Engineer, you'll be responsible for identifying and addressing performance issues across many different ML systems, including... 
    Performance
    Contract work
    For contractors
    For subcontractor
    Work at office
    Relocation
    Visa sponsorship
    Work visa
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  •  ...Team We’re building high-performance infrastructure to serve OpenAI...  ...FLOP from our GPUs by designing kernels, tuning memory layouts, and...  ...looking for a kernel-focused engineer to lead efforts in writing,...  ...requires deep familiarity with CUDA or equivalent kernel... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...of what's possible in video generation.We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and...  ...solutions that achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll ensure our infrastructure... 
    Performance

    Genmo

    San Francisco, CA
    3 days ago
  •  ...in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization across... 
    Performance

    Genesis AI

    San Francisco, CA
    1 day ago
  • $350k

    Mirendil is seeking an engineer to design and optimize custom ML kernels to enhance our model development stack. This role involves working at the intersection...  ...of hardware and frontier AI research, focusing on performance optimization. The ideal candidate will have... 
    Performance

    Mirendil

    San Francisco, CA
    1 day ago
  • TypeSafe AI in San Francisco is seeking a GPU kernel engineer with deep CUDA expertise to accelerate our training and inference workloads. You will write, optimize, and maintain high-performance kernels close to the metal. Work alongside research and platform engineers... 
    Performance
    Work at office

    TypeSafe AI

    San Francisco, CA
    2 days ago
  • $172.5k - $210k

     ...strategies, and be part of a high-performing team that believes in each...  ...: As an Automated Testing Engineer, you will be responsible for...  ...Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL)...  ...: Knowledge of Linux kernel internals, specifically PCIe... 
    Performance
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ..., our work helps care teams perform with greater precision and patients...  ...: Development of Linux kernel internals, device drivers, memory...  ...• Expert in GPU Compute API - CUDA, OpenCL• Proficiency in... 
    Performance
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    3 days ago
  •  ...optimization Highly experienced with CUDA Highly experienced with Rust...  ...on-site role for a Founding Compiler Engineer located in downtown San Francisco. You...  ...Day-to-day tasks will include writing CUDA kernels, conducting model performance reviews. #J-18808-Ljbffr Slope
    Performance
    Full time

    Slope

    San Francisco, CA
    1 day ago
  • $220k - $320k

    A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over 2 years of experience in ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation...
    Performance
    Local area

    Inference

    San Francisco, CA
    4 days ago
  • $220k - $320k

    inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques... 
    Performance

    inference.net

    San Francisco, CA
    5 days ago
  •  .... They are looking for a Runtime Engineer to build the execution layer at the...  ...spanning parallel execution, kernel scheduling, runtime architecture and performance analysis. You will work close to...  ...development experience GPU programming, CUDA/ROCm, HPC, large clusters,... 
    Performance

    Oho Group

    San Francisco, CA
    4 days ago
  •  ...real workloads. This is an engineering role, not a research role. You...  ...at high throughput Own the performance characteristics of those systems...  ...team (quantization, custom kernels, scheduling improvements, memory...  ...following languages: C++, CUDA, ROCm or Triton Track record... 
    Performance

    MakerMaker.AI

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to CUDA Kernel Performance Engineer. Be the first to apply!