Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Performance Engineer, Inference Systems

$350k
Full-time

Anthropic

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the Role

Anthropic's inference fleet serves Claude to millions of users across our own products and the world's largest cloud platforms. The stack that makes this possible is deep and tightly coupled: accelerator kernels, model servers, distributed routing, autoscaling, capacity management. Every layer affects the others, often in ways that are hard to see in isolation.

The Inference System Dynamics team is responsible for understanding that whole system and holding it to a high bar across four dimensions: throughput, latency, reliability, and correctness . We measure how the fleet performs against its theoretical performance frontier, run cross-layer investigations to explain the gaps, and own the correctness checks that make sure Claude's outputs are right, not just fast, across hardware platforms and serving configurations. We don't own the individual components. We instrument and model them, find the highest-leverage opportunities across them, and partner with the owning teams to land the wins.

You'll work across all four areas. One week that might mean tracing a tail-latency regression from request timing down through routing and batching into a kernel overhead; the next it might mean tightening a correctness eval so it catches an output regression introduced by a quantization change. We're looking for performance engineers who treat correctness as part of performance.

Key Responsibilities

  • Run cross-layer performance investigations across throughput, latency, and reliability, sizing the gap between actual fleet performance and theoretical rooflines, identifying root causes, and quantifying the value of closing them
  • Own and improve the correctness evaluation pipeline that validates model output quality across hardware platforms, numerics, and serving configurations, and lead the investigation when it catches a regression
  • Build the observability, dashboards, and modeling tools that make throughput, latency, cost, reliability, correctness, and their interactions legible across the stack
  • Partner with kernel, serving, routing, autoscaling, and capacity teams to prioritize and land the highest-impact optimizations your analysis surfaces
  • Ruthlessly stack-rank a large surface area of opportunities by impact and effort, and say no to the ones that don't make the cut

Minimum Qualifications

  • Hands-on performance engineering experience: profiling, roofline analysis, latency/throughput optimization, and root-cause investigation in complex production systems
  • Proficiency in Python, with the ability to read, instrument, and contribute to large production codebases you didn’t write
  • Solid data analysis skills (e.g. SQL, pandas, or similar) sufficient to turn raw telemetry into clear findings
  • Ability to communicate quantitative results clearly in writing to influence priorities on teams you don't manage
  • Genuine interest in correctness as an engineering discipline: numerics, evaluation design, regression detection

Preferred Qualifications

  • Experience with ML systems, especially training or inference infrastructure or general LLM serving stacks. Direct large-scale inference experience is a strong plus
  • Familiarity with GPU/TPU/accelerator performance concepts (memory bandwidth, kernel overheads, quantization, collective communication). Reasoning about these matters more than having written kernels yourself
  • Experience with reliability engineering for high-throughput services: autoscaling, load balancing, request routing, tail latency
  • Experience with model evaluation or numerical regression-detection pipelines
  • Experience building observability or telemetry for distributed systems
  • Comfortable having impact through influence and evidence rather than direct ownership

Representative Projects

  • Trace a 350ms latency gap on a new accelerator platform from end-to-end request timing down to a server scheduling overhead, quantify the win, and land the fix directly or with the owning team
  • Redesign the correctness eval gate: determine which signals reliably catch real model-output regressions versus noise, and make it the trusted release criterion across hardware backends
  • Build a FLOPs funnel that breaks down where compute actually goes across the fleet, exposing the gap between achieved throughput and kernel rooflines
  • Root-cause a numerical divergence between two hardware platforms to a specific kernel change, and define the acceptance threshold going forward
  • Model the latency–cost impact of changing batch-sizing and utilization targets, and turn the result into the signal the autoscaler uses in production

Deadline to apply: None. Applications will be reviewed on a rolling basis.

The annual compensation range for this role is listed below.

For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.

Annual Salary:

$350,000—$850,000 USD

Logistics

Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience

Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience

Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position

Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.

We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.

Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic.com/careers directly for confirmed position openings.

How we're different

We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills.

The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.

Come work with us!

Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Performance Engineer, Inference Systems in San Francisco, CA vacancy
  • Sail Research in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be responsible for building high-performance schedulers and optimizing global routing while... 
    Performance

    Sail Research

    San Francisco, CA
    1 day ago
  • Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control...  ...will also build LV routing systems to dispatch workloads with...  .../compute trade-offs in LLM inference stacks. You will contribute to... 
    Performance

    Sail

    San Francisco, CA
    16 hours ago
  • $225k

    Dormont Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role involves designing distributed systems, optimizing performance, and ensuring high reliability for RL and post-training workflows. The ideal candidate... 
    Performance

    Dormont Manufacturing Co

    San Francisco, CA
    16 hours ago
  • MakerMaker in San Francisco is seeking a Senior ML systems engineer to build and operate production inference systems for large models. You will own performance, profiling, and optimizations to ensure high throughput and low latency in production. You will collaborate with... 
    Performance

    MakerMaker

    San Francisco, CA
    4 days ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability. The ideal candidate will have 3+ years of experience in production... 
    Performance

    MakerMaker.AI

    San Francisco, CA
    3 days ago
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 
    Performance

    Causal Labs

    San Francisco, CA
    4 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries...  ...to push the bleeding edge of AI performance. This role is based on-site in San Francisco... 
    Performance

    Vast.ai Inc.

    San Francisco, CA
    4 days ago
  • TensorScale AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and implementing... 
    Performance

    TensorScale AI

    San Francisco, CA
    1 day ago
  •  ...Member of Technical Staff for distributed systems to design, build, and operate the platform...  ...ownership of architecture, reliability, and performance in a fast-growing environment. You will collaborate with founders and engineers from Nvidia, Google AI, Intel, and Pixie... 
    Performance

    Acceler8 Talent

    San Francisco, CA
    16 hours ago
  •  ...of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving execution performance across various components. Ideal candidates should have strong software engineering skills and experience with ML inference systems... 
    Performance

    Gimlet Labs

    San Francisco, CA
    4 days ago
  •  ...Member of Technical Staff focused on ML systems and inference in San Francisco. You will design and...  ...fast, predictable, and scalable performance. Key responsibilities include optimizing...  ...have strong foundations in software engineering, experience with ML inference systems... 
    Performance

    Gimlet Labs, Inc.

    San Francisco, CA
    1 day ago
  • $249.5k - $273.5k

     ...competitive advantage.Unlike legacy systems built to route and answer,...  ...Research, Design, and Engineering leadership, you will lead a...  ...feedback, career development, and performance management. Create clarity...  ...emerging agent frameworks, inference optimization techniques,... 
    Performance
    Work at office

    Dialpad

    San Francisco, CA
    3 days ago
  •  ...in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms,...  ...scheduling for low-latency, high-throughput inference. You will implement changes in...  ...RL and post-training pipelines, drive performance improvements, and mentor engineers on... 
    Performance

    Together

    San Francisco, CA
    16 hours ago
  • OpenAI in San Francisco is seeking an experienced systems generalist to build an automated inference optimization platform across hardware, compiler, and...  ...focusing on reliable long-running workflows, reproducible performance results, and secure handling of model and hardware... 
    Performance

    Slope

    San Francisco, CA
    3 days ago
  • $180k - $250k

     ...maintain its frontier position on model performance for generative media models. Design...  ...serving architecture on top of our in-house inference engine, focusing on maximizing throughput...  ...: Strong foundation in systems programming with expertise in identifying... 
    Performance
    Full time
    Currently hiring
    Relocation
    Visa sponsorship

    Fal

    San Francisco, CA
    16 hours ago
  • $130k - $200k

    Join us to apply for the Founding Engineer (Systems + ML) role at Partcl . Get AI-powered advice on...  ...Architect and build end‑to‑end pipelines: high‑performance kernels, efficient file IO, training models, latency-sensitive inference, CLI/UI layers. Communicate with... 
    Performance
    Full time

    Partcl

    San Francisco, CA
    16 hours ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic...  ...and help build the platform engineers turn to to ship AI products....  ...the global operating system for distributed, heterogeneous...  ...characterize and validate networking performance on bleeding-edge clusters (H... 
    Performance
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    16 hours ago
  •  ...AI Engineer, Search & Knowledge Systems About Pi Pi is building an agentic product security platform...  ...resolution, ontology design, relationship inference, and semantic enrichment....  ...improve latency, and optimize cost/performance tradeoffs. Partner with product,... 
    Performance
    Full time

    Pi Security

    San Francisco, CA
    16 hours ago
  • $189.6k - $237k

     ...language model training and inference. The platform has been powering...  ...:Strong excitement about system optimizationExperience with...  ...distributed ML systemsStrong software engineering skills, proficient in...  ..., qualifications, interview performance, and relevant education or... 
    Performance
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  • $150k - $250k

     ...adaptive intelligence. Our systems evolve through use, reflecting...  ...to experts across fields: engineering, research, operations, and beyond...  ...billing, or IAM flows with performance in mind Write maintainable,...  ...at Stealth AI Startup by 2x Inferred from the description for... 
    Performance
    Full time
    Summer work
    Internship
    Remote work
    Flexible hours

    Stealth AI Startup

    San Francisco, CA
    4 days ago
  • $285.45k

     ...we use AI in our recruiting process here.As a principal engineer on the Online Systems team, you’ll join a team that powers Pinterest’s most business...  ...strong technical judgment in reliability, scalability, performance, and infrastructure efficiency; proficiency in at least... 
    Performance
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    2 days ago
  •  ...technologies—like the da Vinci surgical system and Ion—have transformed how...  ...worldwide.We’re a team of engineers, clinicians, and innovators...  ..., our work helps care teams perform with greater precision and...  ...models to real-time onboard inference—while serving as a core contributor... 
    Performance
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    12 hours ago
  •  ...AI/ML Engineer (RL & Physical Systems) FLUIX is building the AI Operating System for data centers. We...  ...based experiments to validate model performance under noisy, dynamic, and safety-critical...  ...relevant (knowledge distillation, inference orchestration, etc.). Lead or... 
    Performance
    Weekend work

    Fluix AI

    San Francisco, CA
    3 days ago
  •  ...the Role We’re hiring a hands‑on Vision Systems Engineer to own the detection, tracking, and...  ...imagery, optimized for real‑time embedded inference via quantization, pruning, or...  ...hardware‑in‑the‑loop validation, and own performance verification against program requirements... 
    Performance

    Exploration Technology Group

    San Francisco, CA
    5 days ago
  •  ...format. Our users are AI/ML researchers and AI infra engineers developing models in complex domains, such as...  ...engineers with deep experience in low-level, high-performance software and cloud-scale storage systems. ~5+ years of experience writing high-quality production... 
    Performance
    Full time
    Work at office

    Spiral

    San Francisco, CA
    16 hours ago
  •  ...AI Systems Engineer - Codex Core Agents Location San Francisco Employment Type Full time Department...  ...also includes generous equity, performance-related bonus(es) for eligible employees...  ...harness issues, model behavior, inference/runtime issues, and product failures.... 
    Performance
    Full time
    Work at office
    Local area
    Relocation package
    Flexible hours

    Slope

    San Francisco, CA
    4 days ago
  •  ...About the Team The Platform Systems team operates at the intersection of cutting-edge...  ...AI and distributed systems. We do the engineering and research required to train our flagship...  ...getting into low level details about performance. Are passionate about building... 
    Performance
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    16 hours ago
  •  ...swing at the most ambitious engineering challenge of our careers. Everyone...  ...competitive market. Inference is just one piece of an effective...  ...and build the rest of the system, that turns billions of...  ...model (Deepseek, Qwen, Llama) performance. Claude Code is all about agent... 
    Performance
    Work at office
    Immediate start

    Sail Research

    San Francisco, CA
    4 days ago
  •  ...About the Team: The Database Systems team specializes in high-performance distributed databases. Our team built Rockset, the real-time search, analytics...  ...cases. About the Role : We are looking for engineers passionate about distributed systems, close-to-the-metal... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    16 hours ago
  • Inception in San Francisco is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in production. You will build and operate infrastructure that handles billions of inference requests, optimizing for latency, throughput, cost, and reliability... 

    Inception

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Performance Engineer, Inference Systems. Be the first to apply!