GPU Performance Engineer
Genmo
We are Genmo, a research lab developing the world’s most sophisticated video world models to understand, simulate, and interact with the physical world. Our mission is to unlock the right brain of AGI. Join us in advancing physical intelligence and enabling robots to learn and act in a changing world.We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits.The RoleYou'll be our performance optimization expert, using advanced profiling tools to identify bottlenecks and implementing solutions that achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll ensure our infrastructure delivers world-class performance. This role is perfect for someone who gets excited about microsecond optimizations and pushing hardware to its theoretical limits.Key ResponsibilitiesProfile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentationWrite high-performance CUDA and Triton kernels for critical model operationsOptimize cold start latency from seconds to milliseconds for our serving infrastructureTune memory access patterns, kernel fusion, and GPU utilizationCollaborate with ML engineers to optimize model implementationsDebug performance issues across the full stack from application to hardwareImplement custom memory pooling and allocation strategiesShare optimization techniques and build performance culture across teamsQualificationsBachelor's or Master's degree in Computer Science, Electrical Engineering, or related field5+ years systems programming experience with 3+ years focused on GPU optimizationExpert proficiency with GPU profiling tools (Nsight Systems, nvprof)Strong CUDA programming skills with production kernel developmentDeep understanding of GPU architecture (memory hierarchy, SMs, warps)Track record of achieving significant performance improvements (5-10x)Experience with Python and C++ in production environmentsWe ValueExperience with Triton kernel developmentKnowledge of CUTLASS or similar high-performance librariesBackground in ML-specific optimizations (attention, transformers)RDMA/InfiniBand optimization experienceContributions to GPU libraries or frameworksLow-level debugging skills (PTX/SASS reading)Genmo is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law. Genmo, Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish.LocationSan Francisco HQEmployment TypeFull timeDepartmentEngineering
$300k
...GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its limits — not in theory, but in production systems handling real-time speech and multimodal workloads? This team is building low-latency AI systems where milliseconds actually matter....PerformanceRelocationVisa sponsorshipFree visa- ...Sciforium is seeking a GPU Kernel Engineer to push performance on modern accelerators, designing and optimizing custom GPU kernels powering large-scale AI systems across low-level kernel development and high-level ML frameworks. You will collaborate with ML researchers...Performance
- ...Hamilton Barnes is seeking a Senior Storage Engineer to own the high-performance storage layer for large-scale GPU clusters and AI workloads. You will design, deploy, and operate storage platforms, collaborating with infrastructure, compute, and networking teams to scale...PerformanceRemote job
- ...sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full... ...the role We are seeking a highly skilled GPU Kernel Engineer who is passionate about pushing the limits of performance on modern accelerators. In this role, you...PerformanceFlexible hours
$170k - $250k
...over $500K in revenue within six months and is scaling rapidly with a small, high-performing team. This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The role offers hands-on involvement with a production...PerformanceFull timeVisa sponsorshipFlexible hours$250k
...infrastructure provider building a next-generation GPU platform designed for AI training,... ...for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and... ...the reliability, scalability, and performance of HPC and cloud infrastructure environments...PerformanceFull timeRemote work- ...Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton... ..., identify bottlenecks, and improve performance for large-scale LLM training and inference... ...with ML researchers, distributed systems engineers, model-serving teams, and NVIDIA/AMD hardware...PerformanceFull timeFlexible hours
- ...Design, implement, and optimize custom GPU kernels for large-scale AI systems. Profile ML operations, develop performance models, identify bottlenecks, and improve training... ...with ML researchers, distributed systems engineers, model-serving teams, and NVIDIA/AMD...PerformanceFull timeFlexible hours
$190k - $235k
San Francisco, CaliforniaSoftware Engineering /Full time Exempt /HybridRETHINK MANUFACTURING... ...deep learningOptimize model inference for GPU deployment, leveraging CUDA, TensorRT,... ...GenICam, CoaXPress)C/C++ experience for performance-critical componentsExperience with MLOps...PerformanceFull time$220k - $260k
..., what belongs in the inference graph, and where the performance ceiling really is.Owning perception quality end to end... ...edge compute, reproducibility, and honest evaluation.GPU inference in practice: TensorRT engines, batching, fp16, and reasoning about where latency...PerformanceLocal areaNight shift$347k
...unfettered growth.About the RoleOpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency... ...infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond.Develop tooling and metrics that...PerformanceWork at officeLocal areaFlexible hours- ...Vision Systems Engineer Exploration Technology Group San Francisco, California, United... ...‑in‑the‑loop validation, and own performance verification against program requirements... ...track‑before‑detect — on embedded FPGA and GPU platforms. Performance Modeling & Test...PerformanceFlexible hours
- ...The role You'll own the cost and performance of our inference stack. Your work will determine... .... You'll work closely with the engineers operating the serving fleet while owning... ...systems language. ~ Experience with GPU performance, including CUDA, NCCL, mixed...PerformanceFlexible hours
$125k - $200k
...Founding RL Engineer (San Francisco, on-site, full-time) About the role We are looking... ...frameworks to measure model and agent performance Run experiments fast, read results, and... ...what to try next Scale training on GPU clusters and keep pipelines reliable...PerformanceFull time$220k - $320k
...Customer engineering at iframe.ai is not solutions architecture and is not customer support... ...time, owning their training and inference performance end-to-end. The same engineers who tune... ...or AI-native scale-ups running 256–2048-GPU jobs. 02 Profile distributed training...PerformanceVisa sponsorshipFlexible hours$120k - $190k
...Houston; New York; San Francisco; Seattle About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior...PerformanceFlexible hours- ...Responsibilities As a ML System Engineer on the AI & ML Platform’s Inference team... ...scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding... ...required if you have:1+ years of system performance optimization experienceLow-level inference...PerformanceWork at officeLocal area
- ...Site Reliability Engineer - AI Infrastructure Location: Global Remote / San Francisco · Full... ...Have Exposure to ML/AI infrastructure or GPU-based systems (CUDA, Slurm, Triton, etc.). Familiarity with high-performance networking (InfiniBand, NVLink) or distributed...PerformanceFull timeRemote work
$182k - $242k
...CoreWeave combines superior infrastructure performance with deep technical expertise to... ...Do: We're looking for a Senior Storage Engineer, File & Block to help build and operate... ...CoreWeave's storage from a follower of GPU growth into a multiplier of it: owning the...PerformancePermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...About This RoleAs a Senior ML System Engineer on the AI & ML Platform’s Inference team... ...scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding... ...experience, 2+ years of system performance optimization experienceDeep low-level systems...PerformanceWork at officeLocal area
- ...centers to deliver Faster AI. Today’s AI performance is frequently limited by communication... ...improvement in performance and unlock greater GPU utilization to speed training job... ...in execution mode and has a world‑class engineering team with decades of experience in state...Performance
$170k - $205k
...their AI strategies, and be part of a high-performing team that believes in each other, come... ...the Role:As a Senior Deployment Automation Engineer for the Compute Team, you will be... ...testing automation of large-scale, multi-node GPU clusters. You will develop the CI/CD infrastructure...PerformanceTemporary work$209k - $253k
...their AI strategies, and be part of a high-performing team that believes in each other, come... ..., and our Compute-focused Production Engineers are the backbone of that mission. This role... ...for AI and HPC workloads across CPU, GPU, and DPU/NIC resources. You will participate...PerformanceTemporary work- # Staff Mechanical Engineer· San Francisco Office (Fremont St)Critical FacilitiesFullTimeremotePosted... ...of superintelligence. One person, one GPU. If you'd like to build the world's best... ...to-chip liquid-cooling infrastructure. - Perform cooling-load, heat-transfer, hydraulic,...PerformanceWork at officeLocal areaWork from homeFlexible hours
$215k - $260k
...their AI strategies, and be part of a high-performing team that believes in each other, come... ...seeking a Hardware Production / Sustaining Engineer to strengthen Crusoe’s Hardware Systems... ...resolution, and reliability across Crusoe Cloud’s GPU- and CPU-based infrastructure.You will...Performance- ...stack used to run inference and training workloads across our fleet. We work at the boundary of Linux, containers, filesystems, storage, GPU drivers, and distributed systems. Our goal is to make demanding ML workloads start quickly, run efficiently, and remain securely...Performance
- ...Modal is seeking a systems engineer to advance the container runtime stack used for inference... ...close to Linux, containers, storage, GPU drivers, and distributed systems. You... ...in Rust and Go, diagnose tough Linux and performance issues, and influence the architecture of...Performance
- ...seeking a Principal Machine Learning Systems Engineer (P60) to lead technical directions of... ...-edge algorithms with reliable, high-performance infrastructure.Working at AtlassianAtlassians... ...systems, high-performance computing, or GPU optimization.Familiarity with search/...PerformanceWork at officeLocal area
$165k - $200k
...their AI strategies, and be part of a high-performing team that believes in each other, come... ...detail-oriented Senior Network Production Engineer to support the physical and logical implementation... ...of high-performance compute (HPC) and GPU-based AI infrastructure, this role plays...PerformanceFull timeTemporary workRemote work$172.5k - $210k
...their AI strategies, and be part of a high-performing team that believes in each other, come... ...About the Role: As an Automated Testing Engineer, you will be responsible for the end-to-end validation of large-scale, multi-node GPU clusters. You will help own the automated...PerformanceFull timeTemporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU Performance Engineer. Be the first to apply!
- sports performance coach San Francisco, CA
- performance food service San Francisco, CA
- performance food group San Francisco, CA
- performance shop San Francisco, CA
- performance specialist San Francisco, CA
- performance windows San Francisco, CA
- performance nutrition San Francisco, CA
- high performance computing engineer San Francisco, CA
- sports performance San Francisco, CA
- performance improvement consultant San Francisco, CA




