Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Performance Engineer: Scale AI Inference

$315k

Anthropic

A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams, and troubleshooting performance issues. The position offers a competitive salary range of $315,000—$560,000, equity opportunities, and a supportive work environment that values communication and collaboration. #J-18808-Ljbffr Anthropic

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the GPU Performance Engineer: Scale AI Inference in San Francisco, CA vacancy
  •  ...seeking a talented software engineer to join their dynamic Inference team. This role involves designing...  ...infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image...  ...to push the boundaries of AI technology, ensuring reliable... 
    Performance

    Jobleads-US

    San Francisco, CA
    4 days ago
  • A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal... 
    Performance

    Baseten

    San Francisco, CA
    3 days ago
  • Sciforium is seeking a GPU Kernel Engineer to push performance on modern accelerators. You will design and optimize custom GPU kernels, from...  ...to integrating ops in ML frameworks used for large-scale training and inference. Ideal candidates have 5+ years in GPU kernel... 
    Performance

    Sciforium

    San Francisco, CA
    3 days ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like...  ...build the platform engineers turn to to ship AI...  ...multi-modal workloads scale, the network is the...  ...to lead our GPU Networking efforts,...  ...validate networking performance on bleeding-edge clusters... 
    Performance
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    12 hours ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure...  ...a next-generation GPU platform designed for...  ...experimentation, and inference at scale. The company...  ...Site Reliability Engineer to support and scale large...  ...reliability, scalability, and performance of HPC and cloud... 
    Performance
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ..., our work helps care teams perform with greater precision and...  ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be...  ...models to real-time onboard inference—while serving as a core contributor... 
    Performance
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    3 days ago
  • $325k

    A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate...  ...collaboration with researchers and focus on performance optimization. Compensation ranges from $325K to $4... 
    Performance

    Jobleads-US

    San Francisco, CA
    4 days ago
  •  ...About the Team Our Inference team brings OpenAI’s most...  ...our state-of-the-art AI models, allowing them...  ...before. We focus on performant and efficient model inference...  ...Role We’re hiring engineers to scale and optimize OpenAI’s...  ...across emerging GPU platforms. You’ll work... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    12 hours ago
  • $160k - $230k

     ...the RoleAt Together.ai, we are building...  ...efficient and scalable inference for large language...  ...the boundaries of performance, scalability, and...  ...and Optimization Engineer to design, develop...  ...models at scale. This role will focus...  ...throughput inference, GPU/accelerator... 
    Performance
    Full time

    Together AI

    San Francisco, CA
    2 days ago
  • $170k - $250k

     ...infrastructure company operating in the AI space, backed by a leading...  ...within six months and is scaling rapidly with a small, high-performing team. This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The... 
    Performance
    Full time
    Visa sponsorship
    Flexible hours
    San Francisco, CA
    22 days ago
  • $188k - $275k

     ...Essential Cloud for AI™. Built for...  ...innovators to build and scale AI with confidence...  ...infrastructure performance with deep technical...  ...Do: The Field Engineering organization at CoreWeave...  ...can train and inference on at scale,...  ...lifecycle: leading new GPU cluster bring-up... 
    Performance
    Permanent employment
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    26 days ago
  • $140k - $200k

     ...pioneering the future of Physical AI. Our advanced vision...  ...Data Infrastructure / Quality Engineering role will play a crucial...  ...validation, and production-scale release Experience architecting...  ...and optimizing high-performance GPU cloud inference services, with specific expertise... 
    Performance
    Work experience placement
    Local area

    Ouster

    San Francisco, CA
    7 days ago
  •  ...build and operate the inference systems that serve...  .... This is an engineering role, not a research...  ...Own the performance characteristics of...  ...production‑grade, large‑scale serving infrastructure...  ...Experience with GPU‑accelerated inference...  ...#J-18808-Ljbffr MakerMaker.AI
    Performance

    MakerMaker.AI

    San Francisco, CA
    1 day ago
  • $100k - $120k

    Coda Robotics is scaling the compute infrastructure that powers...  ...models. As training and inference workloads grow, we need...  ...team of kernel and system engineers focused on performance-critical code Design, implement...  ...for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware... 
    Performance

    Coda Robotics

    San Francisco, CA
    1 day ago
  • Sciforium is an AI infrastructure company developing...  ...-on support from AMD engineers the team is scaling rapidly to build the...  ...a highly skilled GPU Kernel Engineer who...  ...pushing the limits of performance on modern...  ...large-scale training and inference. This role is ideal... 
    Performance
    Flexible hours

    Sciforium

    San Francisco, CA
    4 days ago
  •  ...-a-Service (TaaS) Engineer to help build the...  ...that convert large-scale infrastructure capacity...  ...will work across performance benchmarking,...  ...infrastructure stack, ensuring GPU capacity can be...  ...GPU clusters, AI infrastructure,...  ...model porting, inference/training workloads... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    12 hours ago
  • $342k

     ...demands of advanced AI workloads. The...  ...About the RoleAs an Engineer on our hardware optimization...  ...and performance. You will work with...  ...efficient training and inference on our models. If...  ...decisions on scale up, scale out, front...  ...understanding of GPU and/or other AI acceleratorsExperience... 
    Performance
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    3 days ago
  • $170k - $245k

     ...accelerate the progress of AI applications out into the real...  ...or data scientist can scale an ML application from their...  ...the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations...  ...that push the boundaries of performance for inference at large scale... 
    Performance
    Work at office

    Anyscale

    San Francisco, CA
    5 days ago
  • $264.8k - $331k

     ...Systems Research Engineer, Agent Post-training...  ...our society. At Scale, our mission is to...  ...the development of AI applications. For...  ...algorithms to reach the performance necessary for...  ...our training and inference framework.Post-train...  ...of the modern GPU clusterExperience... 
    Performance
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • $155k - $269k

     ...Description Waabi, founded by AI visionary Raquel...  ...Scientists and Engineers building the content backbone...  ...and curate assets at the scale of tens of thousands of...  ..., and debugging GPU jobs in the cloud (AWS,...  ...incentive awards and an annual performance bonus. Perks/... 
    Performance
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    7 days ago
  • $300 per month

     ...vertically integrated AI infrastructure company...  ...urgency, who believe in the scale of our ambition and...  ...and be part of a high-performing team that believes in...  ..., our Production Engineering team ensures the reliability...  ...AI pipelines and inference servicesDefine, measure... 
    Performance
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  • $350k

     ...Join a rapidly growing AI infrastructure...  ...provider delivering large-scale compute solutions for AI training and inference across global cloud and GPU environments. The...  ...build reliable, high-performance platforms supporting...  ...Site Reliability Engineer to lead the reliability... 
    Performance
    Full time
    San Francisco, CA
    a month ago
  • $264.8k - $331k

     ...Systems Research Engineer, Agent Post-training...  ...Enterprise GenAI AI is becoming vitally...  ...our society. At Scale, our mission is to...  ...algorithms to reach the performance necessary for...  ...our training and inference framework. Post-train...  ...of the modern GPU cluster Experience... 
    Performance
    Full time
    Contract work
    For contractors
    For subcontractor
    Work at office

    Scale LLP

    San Francisco, CA
    3 days ago
  • $315k

     ...interpretable, and steerable AI systems. We want AI to...  ...researchers, engineers, policy experts, and business...  ...innovations in GPU performance and systems engineering...  ...performance at unprecedented scale, developing cutting-...  ...dramatically improve inference efficiency. Working... 
    Performance
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    more than 2 months ago
  • $350k

     ...knowledge and tools to make AI work for their unique...  ...We are scientists, engineers, and builders who’ve created...  ...for architecting and scaling the core infrastructure...  ...clusters with GPU workloads, or building...  ...controllers/operators, or performance profiling. Familiarity... 
    Performance
    Full time
    Local area
    Immediate start
    Visa sponsorship
    Work visa
    Relocation package
    Flexible hours

    Thinking Machines Lab

    San Francisco, CA
    12 hours ago
  • $286.2k - $326.7k

     ...Senior Distinguished Engineer, AI Compute (Remote Eligible...  ...and scalable, high-performance AI infrastructure. At...  ...system delivering the high-scale developer and runtime...  ...on top of CPU and GPU substrates. Your contributions...  ...model training, model inference and feature generation... 
    Performance
    Full time
    Part time
    Local area
    Remote work

    Capital One

    San Francisco, CA
    25 days ago
  •  ...build, and operate the platform that schedules AI workloads across thousands of nodes. This...  ...of architecture, reliability, and performance in a fast-growing environment. You will collaborate with founders and engineers from Nvidia, Google AI, Intel, and Pixie Labs... 
    Performance

    Acceler8 Talent

    San Francisco, CA
    3 days ago
  • $189.6k - $237k

    Scale’s ML platform (RLXF) team builds our internal...  ...model training and inference. The platform has been...  ...heart of the field of AI as an indispensable provider...  ...software engineering skills, proficient in...  ...qualifications, interview performance, and relevant education... 
    Performance
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  •  ...About the Team Our Inference team brings OpenAI...  ...start-of-the-art AI models, allowing...  ...before. We focus on performant and efficient...  ...are looking for an engineer who wants to take...  ...FLOP and every GB of GPU RAM of our hardware...  ...rapidly increasing scale. Are self-... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    12 hours ago
  • $193k - $234k

     ...vertically integrated AI infrastructure company...  ...urgency, who believe in the scale of our ambition and...  ...and be part of a high-performing team that believes in each...  ...Network Production Engineer to lead the physical and...  ...performance compute (HPC) and GPU-based AI infrastructure... 
    Performance
    Temporary work
    Remote work

    Crusoe

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Performance Engineer: Scale AI Inference. Be the first to apply!