Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Performance Engineer: Scale AI Inference

$315k

Anthropic

A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams, and troubleshooting performance issues. The position offers a competitive salary range of $315,000—$560,000, equity opportunities, and a supportive work environment that values communication and collaboration. #J-18808-Ljbffr Anthropic

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the GPU Performance Engineer: Scale AI Inference in San Francisco, CA vacancy
  •  ...seeking a talented software engineer to join their dynamic Inference team. This role involves designing...  ...infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image...  ...to push the boundaries of AI technology, ensuring reliable... 
    Performance

    Jobleads-US

    San Francisco, CA
    3 days ago
  • A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal... 
    Performance

    Baseten

    San Francisco, CA
    2 days ago
  • Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open... 
    Performance

    Anyscale

    San Francisco, CA
    3 days ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like...  ...build the platform engineers turn to to ship AI...  ...multi-modal workloads scale, the network is the...  ...to lead our GPU Networking efforts,...  ...validate networking performance on bleeding-edge clusters... 
    Performance
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $250k

     ...opportunities? Join a rapidly scaling AI cloud infrastructure...  ...a next-generation GPU platform designed for...  ...experimentation, and inference at scale. The company...  ...Site Reliability Engineer to support and scale large...  ...reliability, scalability, and performance of HPC and cloud... 
    Performance
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ..., our work helps care teams perform with greater precision and...  ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be...  ...models to real-time onboard inference—while serving as a core contributor... 
    Performance
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    1 day ago
  • $220k - $320k

    inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques... 
    Performance

    inference.net

    San Francisco, CA
    3 days ago
  • Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device...  .... You will primarily work on uzu, our inference engine, and focus on supporting new...  ...models work, and experience in writing high-performance GPU kernels or Rust systems programming. We... 
    Performance
    Local area

    Mirai Labs

    San Francisco, CA
    2 days ago
  • $325k

    A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate...  ...collaboration with researchers and focus on performance optimization. Compensation ranges from $325K to $4... 
    Performance

    Jobleads-US

    San Francisco, CA
    3 days ago
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 
    Performance

    Causal Labs

    San Francisco, CA
    20 hours ago
  •  ...Team OpenAI’s Inference team ensures that...  ...reliably, and at scale. We build and optimize...  ...- to increase performance, flexibility, and...  ...Role We’re hiring engineers to scale and optimize...  ...across emerging GPU platforms. You’ll...  ...OpenAI OpenAI is an AI research and... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • Sail Research in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be responsible for building high-performance schedulers and optimizing global routing while... 
    Performance

    Sail Research

    San Francisco, CA
    3 days ago
  • $170k - $250k

     ...infrastructure company operating in the AI space, backed by a leading...  ...within six months and is scaling rapidly with a small, high-performing team. This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The... 
    Performance
    Full time
    Visa sponsorship
    Flexible hours
    San Francisco, CA
    1 day ago
  • $160k - $230k

     ...the RoleAt Together.ai, we are building...  ...efficient and scalable inference for large language...  ...the boundaries of performance, scalability, and...  ...and Optimization Engineer to design, develop...  ...models at scale. This role will focus...  ...throughput inference, GPU/accelerator... 
    Performance
    Full time

    Together AI

    San Francisco, CA
    1 day ago
  •  ...foundation Model and seeks infrastructure engineers for high-throughput, low-latency inference at scale in San Francisco. You will work on...  ...collaborating with researchers to push model performance. We value deep learning expertise, GPU-aware optimization, and production-... 
    Performance

    causal

    San Francisco, CA
    3 days ago
  • A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying...  ...GPU systems for high-throughput inference and model performance optimization. The ideal... 
    Performance

    Reflection AI

    San Francisco, CA
    3 days ago
  •  ...NEAR AI was started by Illia Polosukhin, co...  ...source AI at a global scale. We are specifically...  ...an expert in high‑performance LLM serving systems and inference optimization. In this...  ...major inference engines such as SGLang, vLLM...  ...of state‑of‑the‑art GPU architectures, and effectively... 
    Performance

    NEAR.AI

    San Francisco, CA
    1 day ago
  • $300k

     ...GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its limits — not in theory, but in production systems handling...  ...team is building low-latency AI systems where milliseconds...  ...runtime layers, profiling large-scale speech and multimodal models... 
    Performance
    Relocation
    Visa sponsorship
    Free visa

    techire ai

    San Francisco, CA
    3 days ago
  • $350k

    Mirendil is looking for engineers to build infrastructure for frontier reasoning models at...  ...Francisco location. This role focuses on large-scale reinforcement learning (RL) model...  ...reliable training infrastructure, implement performance optimizations, and develop evaluation... 
    Performance

    Mirendil

    San Francisco, CA
    3 days ago
  • $100k - $120k

    Coda Robotics is scaling the compute infrastructure that powers...  ...models. As training and inference workloads grow, we need...  ...team of kernel and system engineers focused on performance-critical code Design, implement...  ...for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware... 
    Performance

    Coda Robotics

    San Francisco, CA
    5 days ago
  •  ...build and operate the inference systems that serve...  .... This is an engineering role, not a research...  ...Own the performance characteristics of...  ...production‑grade, large‑scale serving infrastructure...  ...Experience with GPU‑accelerated inference...  ...#J-18808-Ljbffr MakerMaker.AI
    Performance

    MakerMaker.AI

    San Francisco, CA
    5 days ago
  •  ...-a-Service (TaaS) Engineer to help build the...  ...that convert large-scale infrastructure capacity...  ...will work across performance benchmarking,...  ...infrastructure stack, ensuring GPU capacity can be...  ...GPU clusters, AI infrastructure,...  ...model porting, inference/training workloads... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • About Us Vast.ai’s cloud powers AI projects and businesses...  ...excellence. We seek engineers with strong intrinsic drive...  ...programming experience to help scale AI inference. You’ll leverage your knowledge of high-performance systems to optimize GPU performance at the bleeding... 
    Performance
    Full time
    Work at office

    Vast.ai Inc.

    San Francisco, CA
    20 hours ago
  • CoreWeave is seeking a Bare Metal Support Engineer in San Francisco, CA, to ensure high performance and reliability of our GPU infrastructure. You will engage directly with customers...  ...support processes. Join us to be part of the AI infrastructure revolution! #J-18808-Ljbffr... 
    Performance

    CoreWeave

    San Francisco, CA
    3 days ago
  • $342k

     ...demands of advanced AI workloads. The...  ...About the RoleAs an Engineer on our hardware optimization...  ...and performance. You will work with...  ...efficient training and inference on our models. If...  ...decisions on scale up, scale out, front...  ...understanding of GPU and/or other AI acceleratorsExperience... 
    Performance
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    2 days ago
  • $180k - $250k

     ...the next generation of AI products. We build the...  ..., and do it at scale without compromise. For...  ...unified platform where high-performance inference, orchestration, and...  ...experienced software engineer who thrives on building...  ...orchestration, scheduling, GPU autoscaling, large... 
    Performance
    Currently hiring
    Relocation package

    features and labels

    San Francisco, CA
    1 day ago
  • A tech startup focused on AI workloads is seeking a Member of...  ...Staff to design and optimize inference systems. The role involves managing...  ...and improving execution performance across various components....  ...should have strong software engineering skills and experience with ML... 
    Performance

    Gimlet Labs

    San Francisco, CA
    20 hours ago
  • $170k - $245k

     ...accelerate the progress of AI applications out into the real...  ...or data scientist can scale an ML application from their...  ...the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations...  ...that push the boundaries of performance for inference at large scale... 
    Performance
    Work at office

    Anyscale

    San Francisco, CA
    4 days ago
  • $227.2k - $284k

     ...Scale's Physical AI business unit is dedicated to solving the...  ...As an ML Systems Engineer on the Physical AI team...  ...fault-tolerant, high-performance systems for serving robotics...  ...tracking of model inference. Lead: Own...  ...environments, including GPU-level algorithm optimizations... 
    Performance
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • $140k - $200k

     ...pioneering the future of Physical AI. Our advanced vision...  ...Data Infrastructure / Quality Engineering role will play a crucial...  ...validation, and production-scale release Experience architecting...  ...and optimizing high-performance GPU cloud inference services, with specific expertise... 
    Performance
    Full time
    Work experience placement
    Local area

    Ouster

    San Francisco, CA
    15 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Performance Engineer: Scale AI Inference. Be the first to apply!