Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Performance Engineer

Genmo

We are Genmo, a research lab developing the world’s most sophisticated video world models to understand, simulate, and interact with the physical world. Our mission is to unlock the right brain of AGI. Join us in advancing physical intelligence and enabling robots to learn and act in a changing world.We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits.The RoleYou'll be our performance optimization expert, using advanced profiling tools to identify bottlenecks and implementing solutions that achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll ensure our infrastructure delivers world-class performance. This role is perfect for someone who gets excited about microsecond optimizations and pushing hardware to its theoretical limits.Key ResponsibilitiesProfile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentationWrite high-performance CUDA and Triton kernels for critical model operationsOptimize cold start latency from seconds to milliseconds for our serving infrastructureTune memory access patterns, kernel fusion, and GPU utilizationCollaborate with ML engineers to optimize model implementationsDebug performance issues across the full stack from application to hardwareImplement custom memory pooling and allocation strategiesShare optimization techniques and build performance culture across teamsQualificationsBachelor's or Master's degree in Computer Science, Electrical Engineering, or related field5+ years systems programming experience with 3+ years focused on GPU optimizationExpert proficiency with GPU profiling tools (Nsight Systems, nvprof)Strong CUDA programming skills with production kernel developmentDeep understanding of GPU architecture (memory hierarchy, SMs, warps)Track record of achieving significant performance improvements (5-10x)Experience with Python and C++ in production environmentsWe ValueExperience with Triton kernel developmentKnowledge of CUTLASS or similar high-performance librariesBackground in ML-specific optimizations (attention, transformers)RDMA/InfiniBand optimization experienceContributions to GPU libraries or frameworksLow-level debugging skills (PTX/SASS reading)Genmo is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law. Genmo, Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish.LocationSan Francisco HQEmployment TypeFull timeDepartmentEngineering

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the GPU Performance Engineer in San Francisco, CA vacancy
  • $300k

     ...GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its limits — not in theory, but in production systems handling real-time speech and multimodal workloads? This team is building low-latency AI systems where milliseconds actually matter.... 
    Performance
    Relocation
    Visa sponsorship
    Free visa

    Vincere Ltd

    San Francisco, CA
    3 days ago
  •  ...Sciforium is seeking a GPU Kernel Engineer to push performance on modern accelerators, designing and optimizing custom GPU kernels powering large-scale AI systems across low-level kernel development and high-level ML frameworks. You will collaborate with ML researchers... 
    Performance

    Jobleads-US

    San Francisco, CA
    1 day ago
  •  ...Hamilton Barnes is seeking a Senior Storage Engineer to own the high-performance storage layer for large-scale GPU clusters and AI workloads. You will design, deploy, and operate storage platforms, collaborating with infrastructure, compute, and networking teams to scale... 
    Performance
    Remote job

    Jobleads-US

    San Francisco, CA
    2 days ago
  •  ...sponsorship from AMD with hands-on support from AMD engineers the team is scaling rapidly to build the full...  ...the role We are seeking a highly skilled GPU Kernel Engineer who is passionate about pushing the limits of performance on modern accelerators. In this role, you... 
    Performance
    Flexible hours

    Jobleads-US

    San Francisco, CA
    1 day ago
  • $170k - $250k

     ...over $500K in revenue within six months and is scaling rapidly with a small, high-performing team. This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The role offers hands-on involvement with a production... 
    Performance
    Full time
    Visa sponsorship
    Flexible hours
    San Francisco, CA
    a month ago
  • $250k

     ...infrastructure provider building a next-generation GPU platform designed for AI training,...  ...for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and...  ...the reliability, scalability, and performance of HPC and cloud infrastructure environments... 
    Performance
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton...  ..., identify bottlenecks, and improve performance for large-scale LLM training and inference...  ...with ML researchers, distributed systems engineers, model-serving teams, and NVIDIA/AMD hardware... 
    Performance
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    7 days ago
  •  ...Design, implement, and optimize custom GPU kernels for large-scale AI systems. Profile ML operations, develop performance models, identify bottlenecks, and improve training...  ...with ML researchers, distributed systems engineers, model-serving teams, and NVIDIA/AMD... 
    Performance
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    7 days ago
  • $190k - $235k

    San Francisco, CaliforniaSoftware Engineering /Full time Exempt /HybridRETHINK MANUFACTURING...  ...deep learningOptimize model inference for GPU deployment, leveraging CUDA, TensorRT,...  ...GenICam, CoaXPress)C/C++ experience for performance-critical componentsExperience with MLOps... 
    Performance
    Full time

    Bright Machines

    San Francisco, CA
    12 hours ago
  • $220k - $260k

     ..., what belongs in the inference graph, and where the performance ceiling really is.Owning perception quality end to end...  ...edge compute, reproducibility, and honest evaluation.GPU inference in practice: TensorRT engines, batching, fp16, and reasoning about where latency... 
    Performance
    Local area
    Night shift

    Atomic

    San Francisco, CA
    2 days ago
  • $347k

     ...unfettered growth.About the RoleOpenAI is looking for an experienced Performance Engineer to help us scale the performance, reliability, and efficiency...  ...infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond.Develop tooling and metrics that... 
    Performance
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    3 days ago
  •  ...Vision Systems Engineer Exploration Technology Group San Francisco, California, United...  ...‑in‑the‑loop validation, and own performance verification against program requirements...  ...track‑before‑detect — on embedded FPGA and GPU platforms. Performance Modeling & Test... 
    Performance
    Flexible hours

    Exploration Technology Group

    San Francisco, CA
    1 day ago
  •  ...The role You'll own the cost and performance of our inference stack. Your work will determine...  .... You'll work closely with the engineers operating the serving fleet while owning...  ...systems language. ~ Experience with GPU performance, including CUDA, NCCL, mixed... 
    Performance
    Flexible hours

    Jobleads-US

    San Francisco, CA
    2 days ago
  • $125k - $200k

     ...Founding RL Engineer (San Francisco, on-site, full-time) About the role We are looking...  ...frameworks to measure model and agent performance Run experiments fast, read results, and...  ...what to try next Scale training on GPU clusters and keep pipelines reliable... 
    Performance
    Full time

    Jobleads-US

    San Francisco, CA
    1 day ago
  • $220k - $320k

     ...Customer engineering at iframe.ai is not solutions architecture and is not customer support...  ...time, owning their training and inference performance end-to-end. The same engineers who tune...  ...or AI-native scale-ups running 256–2048-GPU jobs. 02 Profile distributed training... 
    Performance
    Visa sponsorship
    Flexible hours

    Jobleads-US

    San Francisco, CA
    3 days ago
  • $120k - $190k

     ...Houston; New York; San Francisco; Seattle About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior... 
    Performance
    Flexible hours

    Jobleads-US

    San Francisco, CA
    1 day ago
  •  ...Responsibilities As a ML System Engineer on the AI & ML Platform’s Inference team...  ...scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding...  ...required if you have:1+ years of system performance optimization experienceLow-level inference... 
    Performance
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    4 days ago
  •  ...Site Reliability Engineer - AI Infrastructure Location: Global Remote / San Francisco · Full...  ...Have Exposure to ML/AI infrastructure or GPU-based systems (CUDA, Slurm, Triton, etc.). Familiarity with high-performance networking (InfiniBand, NVLink) or distributed... 
    Performance
    Full time
    Remote work

    Andromeda Cluster

    San Francisco, CA
    1 day ago
  • $182k - $242k

     ...CoreWeave combines superior infrastructure performance with deep technical expertise to...  ...Do: We're looking for a Senior Storage Engineer, File & Block to help build and operate...  ...CoreWeave's storage from a follower of GPU growth into a multiplier of it: owning the... 
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    20 days ago
  •  ...About This RoleAs a Senior ML System Engineer on the AI & ML Platform’s Inference team...  ...scaling) to deep low-level optimizations (GPU kernels, quantization, speculative decoding...  ...experience, 2+ years of system performance optimization experienceDeep low-level systems... 
    Performance
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    2 days ago
  •  ...centers to deliver Faster AI. Today’s AI performance is frequently limited by communication...  ...improvement in performance and unlock greater GPU utilization to speed training job...  ...in execution mode and has a world‑class engineering team with decades of experience in state... 
    Performance

    Eridu

    San Francisco, CA
    3 days ago
  • $170k - $205k

     ...their AI strategies, and be part of a high-performing team that believes in each other, come...  ...the Role:As a Senior Deployment Automation Engineer for the Compute Team, you will be...  ...testing automation of large-scale, multi-node GPU clusters. You will develop the CI/CD infrastructure... 
    Performance
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  • $209k - $253k

     ...their AI strategies, and be part of a high-performing team that believes in each other, come...  ..., and our Compute-focused Production Engineers are the backbone of that mission. This role...  ...for AI and HPC workloads across CPU, GPU, and DPU/NIC resources. You will participate... 
    Performance
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • # Staff Mechanical Engineer· San Francisco Office (Fremont St)Critical FacilitiesFullTimeremotePosted...  ...of superintelligence. One person, one GPU. If you'd like to build the world's best...  ...to-chip liquid-cooling infrastructure. - Perform cooling-load, heat-transfer, hydraulic,... 
    Performance
    Work at office
    Local area
    Work from home
    Flexible hours

    DataCenterGuidelines

    San Francisco, CA
    2 days ago
  • $215k - $260k

     ...their AI strategies, and be part of a high-performing team that believes in each other, come...  ...seeking a Hardware Production / Sustaining Engineer to strengthen Crusoe’s Hardware Systems...  ...resolution, and reliability across Crusoe Cloud’s GPU- and CPU-based infrastructure.You will... 
    Performance

    Crusoe

    San Francisco, CA
    3 days ago
  •  ...stack used to run inference and training workloads across our fleet. We work at the boundary of Linux, containers, filesystems, storage, GPU drivers, and distributed systems. Our goal is to make demanding ML workloads start quickly, run efficiently, and remain securely... 
    Performance

    Jobleads-US

    San Francisco, CA
    2 days ago
  •  ...Modal is seeking a systems engineer to advance the container runtime stack used for inference...  ...close to Linux, containers, storage, GPU drivers, and distributed systems. You...  ...in Rust and Go, diagnose tough Linux and performance issues, and influence the architecture of... 
    Performance

    Jobleads-US

    San Francisco, CA
    3 days ago
  •  ...seeking a Principal Machine Learning Systems Engineer (P60) to lead technical directions of...  ...-edge algorithms with reliable, high-performance infrastructure.Working at AtlassianAtlassians...  ...systems, high-performance computing, or GPU optimization.Familiarity with search/... 
    Performance
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    1 day ago
  • $165k - $200k

     ...their AI strategies, and be part of a high-performing team that believes in each other, come...  ...detail-oriented Senior Network Production Engineer to support the physical and logical implementation...  ...of high-performance compute (HPC) and GPU-based AI infrastructure, this role plays... 
    Performance
    Full time
    Temporary work
    Remote work

    Crusoe

    San Francisco, CA
    2 days ago
  • $172.5k - $210k

     ...their AI strategies, and be part of a high-performing team that believes in each other, come...  ...About the Role: As an Automated Testing Engineer, you will be responsible for the end-to-end validation of large-scale, multi-node GPU clusters. You will help own the automated... 
    Performance
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    13 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Performance Engineer. Be the first to apply!