Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Performance Engineer

Genmo

We are Genmo, a research lab dedicated to building open, state-of-the-art models for video generation towards unlocking the right brain of AGI. Join us in shaping the future of AI and pushing the boundaries of what's possible in video generation.We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits.The RoleYou'll be our performance optimization expert, using advanced profiling tools to identify bottlenecks and implementing solutions that achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll ensure our infrastructure delivers world-class performance. This role is perfect for someone who gets excited about microsecond optimizations and pushing hardware to its theoretical limits.Key ResponsibilitiesProfile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentationWrite high-performance CUDA and Triton kernels for critical model operationsOptimize cold start latency from seconds to milliseconds for our serving infrastructureTune memory access patterns, kernel fusion, and GPU utilizationCollaborate with ML engineers to optimize model implementationsDebug performance issues across the full stack from application to hardwareImplement custom memory pooling and allocation strategiesShare optimization techniques and build performance culture across teamsQualificationsBachelor's or Master's degree in Computer Science, Electrical Engineering, or related field5+ years systems programming experience with 3+ years focused on GPU optimizationExpert proficiency with GPU profiling tools (Nsight Systems, nvprof)Strong CUDA programming skills with production kernel developmentDeep understanding of GPU architecture (memory hierarchy, SMs, warps)Track record of achieving significant performance improvements (5-10x)Experience with Python and C++ in production environmentsWe ValueExperience with Triton kernel developmentKnowledge of CUTLASS or similar high-performance librariesBackground in ML-specific optimizations (attention, transformers)RDMA/InfiniBand optimization experienceContributions to GPU libraries or frameworksLow-level debugging skills (PTX/SASS reading)Genmo is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law. Genmo, Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish.LocationSan Francisco HQEmployment TypeFull timeDepartmentEngineering

Vacancy posted 16 hours ago
Similar jobs that could be interesting for youBased on the GPU Performance Engineer in San Francisco, CA vacancy
  • $315k

    A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams... 
    Performance

    Anthropic

    San Francisco, CA
    1 day ago
  • $220k - $320k

    inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques... 
    Performance

    inference.net

    San Francisco, CA
    2 days ago
  • A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal... 
    Performance

    Baseten

    San Francisco, CA
    2 days ago
  • Inferact is seeking a hands-on cluster administration engineer to own and operate its high-performance GPU compute infrastructure. You will ensure health, availability, observability, and usability around the clock for HP compute clusters across providers. You will collaborate... 
    Performance

    Inferact

    San Francisco, CA
    1 day ago
  • MakerMaker.AI in San Francisco is seeking a skilled Software Engineer to write and optimize GPU kernels. You will work on deep low-level tasks that directly impact the performance of machine learning models. The ideal candidate has over 4 years of experience with GPU kernels... 
    Performance

    MakerMaker.AI

    San Francisco, CA
    4 days ago
  • $285k - $315k

    SF Tensor is looking for a Founding GPU Kernel Engineer in San Francisco, specializing in GPU architecture and kernel optimization for machine...  ...has deep expertise, proven capabilities in hand-optimizing performance-critical kernels, and strong programming skills in C++ and... 
    Performance
    Full time
    Relocation package

    SF Tensor

    San Francisco, CA
    4 days ago
  •  ...infrastructure company in San Francisco hire a Senior Software Engineer to build and operate GPU-backed sandboxes for AI agents. You will own the runtime...  ...from architecture to deployment, with focus on performance and security. The role demands deep Linux virtualization... 
    Performance

    Recruiting from Scratch

    San Francisco, CA
    4 days ago
  • San Francisco Tensor Company is seeking a Founding GPU Kernel Engineer to enhance GPU performance for AI applications. You will optimize and write kernels while collaborating with compiler teams to improve efficiencies across architectures. The ideal candidate has deep... 
    Performance
    Work at office
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    4 days ago
  • Adobe in San Francisco seeks engineers to rebuild Photoshop's core rendering engine with a GPU-first architecture. You will design, implement, and optimize GPU-based...  ...across Photoshop Desktop. You will profile performance, write tests, and participate in reviews while collaborating... 
    Performance

    Adobe

    San Francisco, CA
    2 days ago
  • Bot Auto is hiring an experienced GPU-focused engineer to advance autonomous driving workloads. You will optimize end-to-end GPU performance, including sensor processing (camera, LiDAR) and neural network inference, and collaborate with software and hardware teams to optimize... 
    Performance
    Relocation

    Bot-Auto

    San Francisco, CA
    1 day ago
  • Adobe is seeking a Graphics/Engine Software Engineer to help rebuild Photoshop's GPU-first core engine (NGE) in California. You will design, build and optimize...  ...workflows. You will also write shaders, profile performance, and contribute to testing, reviews and documentation... 
    Performance

    Adobe Inc.

    San Francisco, CA
    2 days ago
  • Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack for real-time local intelligence...  ...modern language models work, and experience in writing high-performance GPU kernels or Rust systems programming. We welcome applications... 
    Performance
    Local area

    Mirai Labs

    San Francisco, CA
    1 day ago
  •  ...millions of patients worldwide.We’re a team of engineers, clinicians, and innovators united by one...  .... Every day, our work helps care teams perform with greater precision and patients...  ...Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be responsible... 
    Performance
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    1 day ago
  • $100k - $120k

     ...cheaper and faster. Responsibilities Lead a team of kernel and system engineers focused on performance-critical code Design, implement, and optimize custom compute kernels for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware accelerators Find bottlenecks in memory hierarchy... 
    Performance

    Coda Robotics

    San Francisco, CA
    4 days ago
  • $180k - $280k

     ...tier investors. Since mid-2024, we've been engineering the foundation for what comes after the...  .... About the role We're looking for a GPU kernel engineer with deep, low-level CUDA...  ...include: Write, optimize, and maintain high-performance GPU kernels (e.g., in CUDA / CuTe DSL)... 
    Performance
    Work at office
    Visa sponsorship
    Shift work

    TypeSafe AI

    San Francisco, CA
    4 days ago
  •  ...reality. You would collaborate with software engineers, AI researchers, and hardware specialists to develop high-performance solutions that meet the stringent requirements...  ...mobility. Key Responsibilities Optimize end-to-end GPU performance for real-time autonomous driving... 
    Performance

    Bot Auto

    San Francisco, CA
    1 day ago
  • $285k - $315k

    About The Role We're looking for a Founding GPU Kernel Engineer who lives right at the boundary between hardware and software. Someone who...  ...workloads (matmuls, attention, normalization, etc.) to set the performance ceilings Profile at the microarchitectural level: look... 
    Performance
    Full time
    Work at office
    Relocation package

    SF Tensor

    San Francisco, CA
    4 days ago
  • $285k - $315k

     ...Company, we believe the future of AI and high-performance computing depends on rethinking the...  .... We are partnering with researchers, engineers, and organizations who share our belief...  ...About the Role We're hiring a Founding GPU Compiler Engineer to build the core compilation... 
    Performance
    Full time
    Work at office
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    4 days ago
  • $315k

    Performance Engineer, GPU Join to apply for the Performance Engineer, GPU role at Anthropic . About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  • $167.2k - $209k

     ...world. DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI...  ...we can offer the industry-leading performance for our inference services. You will be...  ...optimizations at the inference engine and GPU kernel layers, ensuring our infrastructure... 
    Performance
    Local area
    Remote work
    Worldwide
    Flexible hours

    DigitalOcean

    San Francisco, CA
    4 days ago
  •  ...infrastructure. You will design, deploy, and operate large-scale GPU clusters powering training, evaluation, and serving for the...  ...ensuring reliability with observability. You'll work with researchers to optimize performance and placement. #J-18808-Ljbffr Linuxcareers
    Performance

    Linuxcareers

    San Francisco, CA
    3 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to...  ...inference. You will design and optimize GPU kernels and tensor libraries, leveraging...  ...frameworks to push the bleeding edge of AI performance. This role is based on-site in San... 
    Performance

    Vast.ai Inc.

    San Francisco, CA
    16 hours ago
  • TensorScale AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and implementing... 
    Performance

    TensorScale AI

    San Francisco, CA
    2 days ago
  •  ...and leadership is earned by shipping excellence. We seek engineers with strong intrinsic drive, a true passion for...  ...scale AI inference. You’ll leverage your knowledge of high-performance systems to optimize GPU performance at the bleeding edge of AI. Full-Time On-site... 
    Performance
    Full time
    Work at office

    Vast.ai Inc.

    San Francisco, CA
    16 hours ago
  • CoreWeave is seeking a Bare Metal Support Engineer in San Francisco, CA, to ensure high performance and reliability of our GPU infrastructure. You will engage directly with customers and collaborate with engineering teams to resolve issues and improve our cloud services... 
    Performance

    CoreWeave

    San Francisco, CA
    2 days ago
  • $135.2k - $306.4k

    Job Overview Oracle hardware platform development engineering is seeking a highly driven GPU/CPU Platform System Engineer at the Principal Engineer level...  ...development, design reviews, system integration, performance testing and characterization. You will interact closely... 
    Performance
    Temporary work
    Work experience placement
    Remote work
    Flexible hours

    Oracle

    San Francisco, CA
    3 days ago
  • A technology startup is seeking a Founding Engineer (Systems + ML) to develop GPU-accelerated engines and build end-to-end pipelines. The ideal candidate...  ...GPU code, alongside a deep understanding of systems performance. You'll join as the first technical hire, expecting... 
    Performance
    Full time

    Partcl

    San Francisco, CA
    16 hours ago
  • A leading consulting firm is seeking a Software Engineer (C++ Systems) in San Francisco to optimize microsecond-level performance in GPU virtualization software. Ideal candidates will have elite C++ expertise, with at least 2 years of experience in low-level systems engineering... 
    Performance

    SK HR Consultants.com

    San Francisco, CA
    16 hours ago
  • $170k - $250k

     ...over $500K in revenue within six months and is scaling rapidly with a small, high-performing team. This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The role offers hands-on involvement with a production... 
    Performance
    Full time
    Visa sponsorship
    Flexible hours
    San Francisco, CA
    15 days ago
  • An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and...  ...infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate closely... 
    Performance

    Jobleads-US

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Performance Engineer. Be the first to apply!