CUDA Engineer - Kernel Optimization
Obsidian
1. Role Overview Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for freelancers with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using profiler-guided analysis. You'll help evaluate, optimize, and reason about GPU kernels across modern hardware environments. This is a contract-based opportunity for specialists who enjoy squeezing performance out of modern GPU architectures. 2. Key Responsibilities Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization Use profiler metrics such as L2 cache hit rate, L2 throughput, occupancy, and related signals to guide kernel improvements Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms Write, modify, and reason about C++17, Python, and GPU programming code Apply CUDA, HIP, shader programming, or related kernel programming expertise to improve performance outcomes Document optimization decisions clearly, including when specific profiler metrics are or are not useful 3. Ideal Qualifications Available to work at least 20 hrs/wk Fluent in core C++ features through C++17 Working knowledge of Python and Git Fluent in at least one GPU programming model, such as CUDA, HIP, Slang, HLSL, GLSL, or related kernel programming At least 1 year of professional or graduate-level research experience working with GPUs Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels Ability to optimize GPU kernels without needing deep prior context on every algorithm Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization is a plus Experience optimizing kernels for NVIDIA Blackwell hardware is a plus Familiarity with NSight Compute is a plus Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm is a plus Open-source contributions related to GPU kernel optimization are a plus #J-18808-Ljbffr Obsidian
- Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This contract-based opportunity is designed... ...analysis. Expect to write C++17 and Python code, apply CUDA or HIP, and document decisions clearly. #J-18808-Ljbffr MercorSuggestedContract workFreelance
- Inception is seeking engineers and scientists to design, optimize, and maintain compute foundations for large‑scale... ...You will develop high‑performance ML kernels, enable efficient low‑precision... ...large models. The role emphasizes CUDA/CuTe/Triton kernel design, memory bandwidth...Suggested
$190k - $250k
...with hands-on support from AMD engineers the team is scaling rapidly... ...seeking a highly skilled GPU Kernel Engineer who is passionate about... ...role, you will design and optimize custom GPU kernels that power... ...GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas...SuggestedFull timeFlexible hours- ..., and hardware-appropriateness of Neuron Kernel Interface (NKI) development tasks used to... ...frontier AI lab's models. You will assess CUDA→NKI migration fidelity, Trainium-specific... ...systems group, you will work with kernel engineers and ML developers to ensure robust...Suggested
$160k - $230k
...models (LLMs). Our mission is to optimize inference frameworks,... ...Frameworks and Optimization Engineer to design, develop, and optimize... ...‑performance serving. Apply CUDA graph optimizations, TensorRT... ...CUDA graph, compiled, efficient kernels. Soft Skills: Strong...SuggestedFull time- Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This contract-based opportunity is designed... ...about GPU kernels across modern hardware, writing and reviewing C++17, Python, and CUDA/HIP code. #J-18808-Ljbffr ObsidianContract workFreelance
$166k - $244k
...following: Machine Learning Optimization (e.g., quantization, distillation... .../TPU hardware architecture, Kernel programming, or... ...in Computer Science, Computer Engineering, or a related technical field... ...with kernel programming (e.g., CUDA, OpenCL, Vulkan, Triton), compiler...Full timeTemporary work$140k - $210k
...and deliver with high agency. The Role As a Software Engineer on the Multi-Agent Systems team , you will design the... ...robot planning and coordination capabilities, and build algorithm optimization that enables system robustness, reliability, and scale. This role...Full timeLocal areaFlexible hours- ...AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components... .... You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization across...
$342k
...accelerate innovation and enable hardware optimized specifically for AI.About the RoleAs an Engineer on our hardware optimization and... .... You will work with our kernel, compiler and machine learning engineers... ...AI acceleratorsExperience with CUDA, Triton or a related accelerator...Work at officeLocal areaRelocation packageFlexible hours- ...Member of Technical Staff focused on kernels and GPU performance. This role involves optimizing GPU and accelerator kernels for... ...candidates have strong software engineering foundations and experience with... .... Familiarity with tools like CUDA and performance profiling is...
- Magic is hiring a Kernel Engineer in San Francisco to design, implement, and optimize high-performance kernels for long-context training and inference. You will tackle memory usage, data movement, and throughput challenges in real-time workloads. You’ll work across training...Visa sponsorship
$190.9k - $232.8k
P-1285About This RoleAs a staff software engineer for GenAI Performance and Kernel, you will own the design, implementation, optimization, and correctness of the high-performance GPU... ...writing and tuning compute kernels (CUDA, Triton, OpenCL, LLVM IR, assembly or similar...Local areaWorldwide- Bot Auto is hiring an experienced GPU-focused engineer to advance autonomous driving workloads. You will optimize end-to-end GPU performance, including sensor processing... ..., and control subsystems. The role emphasizes CUDA-based development, profiling, and deployment on embedded...Relocation
- ...patients worldwide.We’re a team of engineers, clinicians, and innovators... ...algorithms into performance optimized, robust, validated and... ...Virtualization: Development of Linux kernel internals, device drivers,... ...Expert in GPU Compute API - CUDA, OpenCL• Proficiency in multiple...Local areaWorldwideFlexible hours
- ...shipping excellence. We seek engineers with strong intrinsic drive,... ...high-performance systems to optimize GPU performance at the bleeding... ...SF or LA offices Tech Stack CUDA/C++, GPGPU, Python, Linux... ...Responsibilities Design and optimize GPU kernels and tensor libraries...Full timeWork at office
$93.6k - $106.08k
...assembling a diverse, world-class team—engineers, designers, researchers, and product minds... ..., investigate performance bottlenecks, optimize system behavior, and build software that... ...apps or AOSP Experience with Linux Kernel driver development Experience porting...Hourly payFull timeTemporary workSummer workInternshipLocal areaFlexible hours- ...We’re looking for an experienced Research Scientist/Engineer with a focus on model optimization to join our core AI team. Our ideal partner-in-crime... ...TVM, Triton, or XLA Experience writing custom Triton/CUDA kernels or low‑level performance tuning Experience with experiment...Remote workRelocation packageFlexible hours
- ...and help build the platform engineers turn to to ship AI products.... ...THE ROLE We’re seeking a GPU Kernel Engineer to join our team at... ...powers modern AI workloads, optimizing every microsecond of computation... ...Write and optimize code using CUDA, PTX assembly, and architecture...Flexible hours
$350k
...for a Staff Site Reliability Engineer to lead the reliability of large... ..., including PyTorch, NCCL, CUDA, drivers, networking fabrics,... ...systems expertise, including kernel tuning, CUDA lifecycle... ...Lustre, or GPFS Experience optimizing distributed training efficiency...Full time- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This...
- ...Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference... .... You will push performance from kernel code to distributed engines, profiling... ...improvements that scale. Responsibilities include CUDA / Triton optimizations, designing...
$130k - $200k
Join us to apply for the Founding Engineer (Systems + ML) role at Partcl . Get... ...00/yr Responsibilities Develop and optimize GPU‑accelerated engines (C++/CUDA) for timing analysis, gate‑sizing,... ...to‑end pipelines: high‑performance kernels, efficient file IO, training models...Full time- ...000 patients worldwide. As a Staff Systems Engineer with expertise in CT imaging, you will be the... ...and validation (V&V) strategy for optimizing Heartflow product outputs across scanners, detectors, reconstruction kernels, and imaging techniques, so Heartflow produces...Local areaWorldwideRelocation
$250k - $300k
...Senior Staff/Principal Deployment Automation Engineer for the Compute Team, you will be... ...scenarios.System Internals: Knowledge of Linux kernel internals, specifically PCIe topology,... ...GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a...Temporary work- ...generation.We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits... ...achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll...
$172.5k - $210k
....About the Role: As an Automated Testing Engineer, you will be responsible for the end-to-end... ...GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a... ....System Internals: Knowledge of Linux kernel internals, specifically PCIe topology, VFIO...Temporary work- Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission... ...libraries, or custom kernels/fused ops. Experience with multi... ...debugging performance issues across CUDA/NCCL, networking, IO, and... ...with data pipeline optimization, sharded datasets, or caching...Full timeWork at officeRemote workFlexible hours
$190.9k - $232.8k
A leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong...- OpenAI in San Francisco seeks an experienced Software Engineer to help bring inference workloads to AWS Trainium and build the software... ...on the platform. This deeply technical, cross‑stack role covers kernels, compilers, and model execution. You will develop high‑...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to CUDA Engineer - Kernel Optimization. Be the first to apply!

