Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

CUDA Engineer - Kernel Optimization

Obsidian

1. Role Overview Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using profiler-guided analysis. You'll help evaluate, optimize, and reason about GPU kernels across modern hardware environments. This is a contract-based opportunity for specialists who enjoy squeezing performance out of modern GPU architectures. 2. Key Responsibilities Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization Use profiler metrics such as L2 cache hit rate, L2 throughput, occupancy, and related signals to guide kernel improvements Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms Write, modify, and reason about C++17, Python, and GPU programming code Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes Document optimization decisions clearly, including when specific profiler metrics are or are not useful 3. Ideal Qualifications Available to work at least 20 hrs/wk Fluent in core C++ features through C++17 Working knowledge of Python and Git Fluent in at least one GPU programming model, such as CUDA, HIP, Slang, HLSL, GLSL, or related kernel programming At least 1 year of professional or graduate-level research experience working with GPUs Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels Ability to optimize GPU kernels without needing deep prior context on every algorithm Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization is a plus Experience optimizing kernels for NVIDIA Blackwell hardware is a plus Familiarity with NSight Compute is a plus Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm is a plus Open-source contributions related to GPU kernel optimization are a plus #J-18808-Ljbffr Obsidian

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the CUDA Engineer - Kernel Optimization in San Francisco, CA vacancy
  • Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This contract-based opportunity is designed...  ...analysis. Expect to write C++17 and Python code, apply CUDA or HIP, and document decisions clearly. #J-18808-Ljbffr Mercor
    Suggested
    Contract work
    Freelance

    Mercor

    San Francisco, CA
    2 days ago
  • Inception is seeking engineers and scientists to design, optimize, and maintain compute foundations for large‑scale...  ...You will develop high‑performance ML kernels, enable efficient low‑precision...  ...large models. The role emphasizes CUDA/CuTe/Triton kernel design, memory bandwidth... 
    Suggested

    Inception

    San Francisco, CA
    5 days ago
  • $190k - $250k

     ...with hands-on support from AMD engineers the team is scaling rapidly...  ...seeking a highly skilled GPU Kernel Engineer who is passionate about...  ...role, you will design and optimize custom GPU kernels that power...  ...GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas... 
    Suggested
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    4 days ago
  •  ..., and hardware-appropriateness of Neuron Kernel Interface (NKI) development tasks used to...  ...frontier AI lab's models. You will assess CUDA→NKI migration fidelity, Trainium-specific...  ...systems group, you will work with kernel engineers and ML developers to ensure robust... 
    Suggested

    Mercor

    San Francisco, CA
    4 days ago
  • $160k - $230k

     ...models (LLMs). Our mission is to optimize inference frameworks,...  ...Frameworks and Optimization Engineer to design, develop, and optimize...  ...‑performance serving. Apply CUDA graph optimizations, TensorRT...  ...CUDA graph, compiled, efficient kernels. Soft Skills: Strong... 
    Suggested
    Full time

    Togetherai

    San Francisco, CA
    5 days ago
  • Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This contract-based opportunity is designed...  ...about GPU kernels across modern hardware, writing and reviewing C++17, Python, and CUDA/HIP code. #J-18808-Ljbffr Obsidian
    Contract work
    Freelance

    Obsidian

    San Francisco, CA
    3 days ago
  • $166k - $244k

     ...following: Machine Learning Optimization (e.g., quantization, distillation...  .../TPU hardware architecture, Kernel programming, or...  ...in Computer Science, Computer Engineering, or a related technical field...  ...with kernel programming (e.g., CUDA, OpenCL, Vulkan, Triton), compiler... 
    Full time
    Temporary work

    Google

    San Francisco, CA
    2 days ago
  • $140k - $210k

     ...and deliver with high agency. The Role As a Software Engineer on the Multi-Agent Systems team , you will design the...  ...robot planning and coordination capabilities, and build algorithm optimization that enables system robustness, reliability, and scale. This role... 
    Full time
    Local area
    Flexible hours

    Nimble Robotics

    San Francisco, CA
    1 day ago
  •  ...AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components...  .... You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization across... 

    Genesis AI

    San Francisco, CA
    1 day ago
  • $342k

     ...accelerate innovation and enable hardware optimized specifically for AI.About the RoleAs an Engineer on our hardware optimization and...  .... You will work with our kernel, compiler and machine learning engineers...  ...AI acceleratorsExperience with CUDA, Triton or a related accelerator... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    4 days ago
  •  ...Member of Technical Staff focused on kernels and GPU performance. This role involves optimizing GPU and accelerator kernels for...  ...candidates have strong software engineering foundations and experience with...  .... Familiarity with tools like CUDA and performance profiling is... 

    Gimlet Labs

    San Francisco, CA
    3 days ago
  • Magic is hiring a Kernel Engineer in San Francisco to design, implement, and optimize high-performance kernels for long-context training and inference. You will tackle memory usage, data movement, and throughput challenges in real-time workloads. You’ll work across training... 
    Visa sponsorship

    Magic

    San Francisco, CA
    1 day ago
  • $190.9k - $232.8k

    P-1285About This RoleAs a staff software engineer for GenAI Performance and Kernel, you will own the design, implementation, optimization, and correctness of the high-performance GPU...  ...writing and tuning compute kernels (CUDA, Triton, OpenCL, LLVM IR, assembly or similar... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  • Bot Auto is hiring an experienced GPU-focused engineer to advance autonomous driving workloads. You will optimize end-to-end GPU performance, including sensor processing...  ..., and control subsystems. The role emphasizes CUDA-based development, profiling, and deployment on embedded... 
    Relocation

    Bot-Auto

    San Francisco, CA
    4 days ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ...algorithms into performance optimized, robust, validated and...  ...Virtualization: Development of Linux kernel internals, device drivers,...  ...Expert in GPU Compute API - CUDA, OpenCL• Proficiency in multiple... 
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    4 days ago
  •  ...shipping excellence. We seek engineers with strong intrinsic drive,...  ...high-performance systems to optimize GPU performance at the bleeding...  ...SF or LA offices Tech Stack CUDA/C++, GPGPU, Python, Linux...  ...Responsibilities Design and optimize GPU kernels and tensor libraries... 
    Full time
    Work at office

    Vast.ai Inc.

    San Francisco, CA
    3 days ago
  • $93.6k - $106.08k

     ...assembling a diverse, world-class team—engineers, designers, researchers, and product minds...  ..., investigate performance bottlenecks, optimize system behavior, and build software that...  ...apps or AOSP Experience with Linux Kernel driver development Experience porting... 
    Hourly pay
    Full time
    Temporary work
    Summer work
    Internship
    Local area
    Flexible hours

    HP IQ

    San Francisco, CA
    1 day ago
  •  ...We’re looking for an experienced Research Scientist/Engineer with a focus on model optimization to join our core AI team. Our ideal partner-in-crime...  ...TVM, Triton, or XLA Experience writing custom Triton/CUDA kernels or low‑level performance tuning Experience with experiment... 
    Remote work
    Relocation package
    Flexible hours

    Tavus

    San Francisco, CA
    4 days ago
  •  ...and help build the platform engineers turn to to ship AI products....  ...THE ROLE We’re seeking a GPU Kernel Engineer to join our team at...  ...powers modern AI workloads, optimizing every microsecond of computation...  ...Write and optimize code using CUDA, PTX assembly, and architecture... 
    Flexible hours

    Baseten

    San Francisco, CA
    4 days ago
  • $350k

     ...for a Staff Site Reliability Engineer to lead the reliability of large...  ..., including PyTorch, NCCL, CUDA, drivers, networking fabrics,...  ...systems expertise, including kernel tuning, CUDA lifecycle...  ...Lustre, or GPFS Experience optimizing distributed training efficiency... 
    Full time
    San Francisco, CA
    a month ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 

    Vast.ai Inc.

    San Francisco, CA
    3 days ago
  •  ...Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference...  .... You will push performance from kernel code to distributed engines, profiling...  ...improvements that scale. Responsibilities include CUDA / Triton optimizations, designing... 

    TensorScale AI

    San Francisco, CA
    5 days ago
  • $130k - $200k

    Join us to apply for the Founding Engineer (Systems + ML) role at Partcl . Get...  ...00/yr Responsibilities Develop and optimize GPU‑accelerated engines (C++/CUDA) for timing analysis, gate‑sizing,...  ...to‑end pipelines: high‑performance kernels, efficient file IO, training models... 
    Full time

    Partcl

    San Francisco, CA
    4 days ago
  •  ...000 patients worldwide. As a Staff Systems Engineer with expertise in CT imaging, you will be the...  ...and validation (V&V) strategy for optimizing Heartflow product outputs across scanners, detectors, reconstruction kernels, and imaging techniques, so Heartflow produces... 
    Local area
    Worldwide
    Relocation

    HeartFlow

    San Francisco, CA
    3 days ago
  • $250k - $300k

     ...Senior Staff/Principal Deployment Automation Engineer for the Compute Team, you will be...  ...scenarios.System Internals: Knowledge of Linux kernel internals, specifically PCIe topology,...  ...GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  •  ...generation.We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits...  ...achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll... 

    Genmo

    San Francisco, CA
    3 days ago
  • $172.5k - $210k

     ....About the Role: As an Automated Testing Engineer, you will be responsible for the end-to-end...  ...GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a...  ....System Internals: Knowledge of Linux kernel internals, specifically PCIe topology, VFIO... 
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission...  ...libraries, or custom kernels/fused ops. Experience with multi...  ...debugging performance issues across CUDA/NCCL, networking, IO, and...  ...with data pipeline optimization, sharded datasets, or caching... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    3 days ago
  • $190.9k - $232.8k

    A leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong... 

    Jobleads-US

    San Francisco, CA
    3 days ago
  • OpenAI in San Francisco seeks an experienced Software Engineer to help bring inference workloads to AWS Trainium and build the software...  ...on the platform. This deeply technical, cross‑stack role covers kernels, compilers, and model execution. You will develop high‑... 

    Slope

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to CUDA Engineer - Kernel Optimization. Be the first to apply!