Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Performance Engineer: Scale AI Inference

$315k

Anthropic

A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams, and troubleshooting performance issues. The position offers a competitive salary range of $315,000—$560,000, equity opportunities, and a supportive work environment that values communication and collaboration. #J-18808-Ljbffr Anthropic

Vacancy posted 10 hours ago
Similar jobs that could be interesting for youBased on the GPU Performance Engineer: Scale AI Inference in San Francisco, CA vacancy
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure...  ...a next-generation GPU platform designed for...  ...experimentation, and inference at scale. The company...  ...Site Reliability Engineer to support and scale large...  ...reliability, scalability, and performance of HPC and cloud... 
    Performance
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $175k - $250k

     ...We're a well-funded AI infrastructure startup...  ...artificial intelligence, high-performance computing, and...  ...We're looking for an engineer to help build and maintain a high-performance inference library designed to support...  ...ROCm, Triton, or similar GPU/accelerator programming... 
    Performance
    Local area

    Jobot

    San Francisco, CA
    3 days ago
  • $190k - $250k

     ...Sciforium is an AI infrastructure company developing...  ...-on support from AMD engineers the team is scaling rapidly to build the...  ...a highly skilled GPU Kernel Engineer who...  ...pushing the limits of performance on modern...  ...large-scale training and inference. This role is ideal... 
    Performance
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    2 days ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like...  ...build the platform engineers turn to to ship AI...  ...multi-modal workloads scale, the network is the...  ...to lead our GPU Networking efforts,...  ...validate networking performance on bleeding-edge clusters... 
    Performance
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  • $140k - $200k

     ...pioneering the future of Physical AI. Our advanced vision...  ...Data Infrastructure / Quality Engineering role will play a crucial...  ...validation, and production-scale release Experience architecting...  ...and optimizing high-performance GPU cloud inference services, with specific expertise... 
    Performance
    Full time
    Work experience placement
    Local area

    Ouster

    San Francisco, CA
    a month ago
  •  ...the world’s largest AI infrastructure networks...  ...highly available GPU infrastructure for AI training and inference workloads.About the...  ...Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics...  ...infrastructure.Perform root-cause analysis... 
    Performance
    Permanent employment

    OpenAI

    San Francisco, CA
    12 days ago
  • $188k - $275k

     ...Essential Cloud for AI™. Built for...  ...innovators to build and scale AI with confidence...  ...infrastructure performance with deep technical...  ...Do: The Field Engineering organization at CoreWeave...  ...can train and inference on at scale,...  ...lifecycle: leading new GPU cluster bring-up... 
    Performance
    Permanent employment
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    18 days ago
  • $106.9k - $200.6k

     ...opportunity We are seeking an AI Systems Engineer to own the delivery, model-...  ...pipelines, operating high-performance inference (GPUs, model servers,...  ...sink), DCGM exporter for GPU telemetry, processor batching...  ...NIM) on GPUs at production scale. Deep observability... 
    Performance
    Full time
    Summer holiday
    Flexible hours

    EY

    San Francisco, CA
    2 days ago
  •  ...About the Team Our Inference team brings OpenAI’s most...  ...our state-of-the-art AI models, allowing them...  ...before. We focus on performant and efficient model inference...  ...Role We’re hiring engineers to scale and optimize OpenAI’s...  ...across emerging GPU platforms. You’ll work... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    more than 2 months ago
  • $350k

     ...Join a rapidly growing AI infrastructure...  ...provider delivering large-scale compute solutions for AI training and inference across global cloud and GPU environments. The...  ...build reliable, high-performance platforms supporting...  ...Site Reliability Engineer to lead the reliability... 
    Performance
    Full time
    San Francisco, CA
    more than 2 months ago
  • $185k - $260k

     ...integrate with advanced AI, and are ultimately...  ...Working with scientists and engineers, we combine simulation,...  .... You will define performance objectives and develop,...  ...efficient enough for large-scale offline processing....  ...profiling and improving GPU performance and memory... 
    Performance
    Full time

    Merge Labs

    San Francisco, CA
    2 days ago
  • $229.5k - $255k

     ...the future of physical AI. Powered by a proprietary...  ...autonomy teams and engineers overcome is the sim-to-...  ...geometrically coherent, correctly scaled and aligned, and usable...  ...implicit. Improve Performance, Throughput, and Cost - Profile and optimize GPU and distributed... 
    Performance
    Full time
    Work at office
    3 days per week

    Niantic Spatial

    San Francisco, CA
    4 days ago
  • $125.5k - $261.6k

     ...We are seeking AI Systems Engineers to build and operate...  ...from bare-metal and GPU infrastructure through...  ...fabric, and multi-tenant scaling. You will be responsible...  ...secure execution and inference: Ray Serve, vLLM/NIM...  ...rewarded based on your performance and recognized for... 
    Performance
    Full time
    Contract work
    Summer holiday
    Flexible hours

    EY

    San Francisco, CA
    2 days ago
  • $125k - $165k

     ..., delivering power to AI data centers in months...  ...proprietary energy architecture engineered entirely in-house. The...  ..., and the operational scale to lead it....  ...characteristics, and performance parameters — and serve...  ...employment information, and inferences drawn from your PI. We... 
    Performance
    Full time

    Redwood Materials

    San Francisco, CA
    1 day ago
  •  ...Head of GPU Cloud About the Company Developing...  ...platform for large-scale AI infrastructure....  ...Cloud to spearhead the engineering team dedicated to the...  ...model for large-scale AI inference services, with a direct...  ...platform's reliability and performance. Close collaboration... 
    Performance

    Confidential

    San Francisco, CA
    1 day ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco...  ...with generative AI. They are the team behind...  ...and maintaining large-scale distributed systems that...  ...ensuring scalability, performance, and reliability across...  ...scale Manage and automate GPU compute clusters using... 
    Performance
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    19 hours ago
  •  ...compute infrastructure scales efficiently to...  ...increasingly sophisticated AI models.We’re...  ...with Capacity Systems Engineering, Infrastructure,...  ...Research to optimize inference capacity across our global GPU fleet. This role...  ...infrastructure investments, performance-efficiency trade-... 
    Performance

    OpenAI

    San Francisco, CA
    11 days ago
  • $180k - $250k

     ...the next generation of AI products. We build the...  ..., and do it at scale without compromise. For...  ...unified platform where high-performance inference, orchestration, and...  ...experienced software engineer who thrives on building...  ...orchestration, scheduling, GPU autoscaling, large... 
    Performance
    Full time
    Currently hiring
    Remote work
    Relocation package

    Falò

    San Francisco, CA
    a month ago
  • $200k

     ...high-speed networks powering the AI era? Join a trailblazing leader in GPU-accelerated computing, designing and...  ...real-time visibility into the performance of thousands of interconnected nodes...  ...Terraform, or SaltStack. High-Scale Networking: A strong foundation in... 
    Performance
    Full time
    San Francisco, CA
    more than 2 months ago
  •  ...deeply in Generative AI — pioneering...  ...Learning Systems Engineer (P60) to lead technical...  ...on building and scaling the systems that power...  ...reliable, high-performance infrastructure.Working...  ...model training, inference pipelines, or...  ...performance computing, or GPU optimization.... 
    Performance
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    1 day ago
  •  ...-a-Service (TaaS) Engineer to help build the...  ...that convert large-scale infrastructure capacity...  ...will work across performance benchmarking,...  ...infrastructure stack, ensuring GPU capacity can be...  ...GPU clusters, AI infrastructure,...  ...model porting, inference/training workloads... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    more than 2 months ago
  • $165k - $200k

     ...vertically integrated AI infrastructure company...  ...urgency, who believe in the scale of our ambition and...  ...and be part of a high-performing team that believes in each...  ...Network Production Engineer to support the physical...  ...performance compute (HPC) and GPU-based AI infrastructure... 
    Performance
    Full time
    Temporary work
    Remote work

    Crusoe

    San Francisco, CA
    4 days ago
  • $231.6k - $289.5k

     ...delivering modular AI infrastructure...  ...factory with speed, scale and sovereignty. Named...  ...a Distinguished Engineer / Technical Fellow...  ...in high-density GPU orchestration and...  ...such as in-orbit inference, adaptive learning...  ...a culture of high-performance engineering and architectural... 
    Performance
    Remote work
    Flexible hours

    Armada

    San Francisco, CA
    2 days ago
  • $161.3k - $241.9k

     ...frontier agentic AI, an enterprise-grade...  ...support. We’re scaling fast and defining...  ...interaction, every model inference, and every...  ...for a Production Engineer to help build and...  ...reliability, and performance. Improve compute...  ...Experience operating GPU fleets, high-performance... 
    Performance
    Full time

    Harvey, Inc.

    San Francisco, CA
    a month ago
  • $155k - $269k

     ...Description Waabi, founded by AI visionary Raquel...  ...Scientists and Engineers building the content backbone...  ...and curate assets at the scale of tens of thousands of...  ..., and debugging GPU jobs in the cloud (AWS,...  ...incentive awards and an annual performance bonus. Perks/... 
    Performance
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    29 days ago
  • $300k

     ...startup building out their AI and cloud platform, powered...  ...ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability...  ...you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring... 
    Performance
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $342k

     ...demands of advanced AI workloads. The...  ...About the RoleAs an Engineer on our hardware optimization...  ...and performance. You will work with...  ...efficient training and inference on our models. If...  ...decisions on scale up, scale out, front...  ...understanding of GPU and/or other AI acceleratorsExperience... 
    Performance
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    19 hours ago
  • $106.9k - $200.6k

     ...opportunity We are seeking AI Systems Engineers to own the security and...  ...lifecycle at production scale across multiple environments...  ...challenges of confidential AI inference (models/secrets inside...  ...be rewarded based on your performance and recognized for the value... 
    Performance
    Full time
    Work experience placement
    Summer holiday
    Remote work
    Flexible hours

    EY

    San Francisco, CA
    2 days ago
  • $250k

     ...Ready to architect AI infrastructure...  ...building a serverless inference platform,...  ...Inference Platform Engineer at an early stage...  ...systems to maximise GPU utilisation and minimise...  ...best practices in performance and efficiency....  ...experience building large-scale, fault-tolerant... 
    Performance
    Full time
    San Francisco, CA
    more than 2 months ago
  •  ...Senior Systems Engineer San Francisco, California Onsite or Remote...  ...posture, database performance, AI inference infrastructure: you can cover...  ...platforms (vLLM, Cloud Run, GPU provisioning), latency SLOs, cost optimization, auto-scaling characteristics, and capacity... 
    Performance
    Remote work
    Work from home

    Evidently

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Performance Engineer: Scale AI Inference. Be the first to apply!