Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Engineer, Performance RL (Reinforcement Learning)

$350k
Full-time

Anthropic

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the RL Teams

Our Reinforcement Learning teams lead Anthropic's reinforcement learning research and development, playing a critical role in advancing our AI systems. We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.6 and Opus 4.6. Our work spans several key areas:

  • Developing systems that enable models to use computers effectively

  • Advancing code generation through reinforcement learning

  • Pioneering fundamental RL research for large language models

  • Building scalable RL infrastructure and training methodologies

  • Enhancing model reasoning capabilities

We collaborate closely with Anthropic's alignment and frontier red teams to ensure our systems are both capable and safe. We partner with the applied production training team to bring research innovations into deployed models, and are dedicated to implement our research at scale. Our Reinforcement Learning teams sit at the intersection of cutting-edge research and engineering excellence, with a deep commitment to building high-quality, scalable systems that push the boundaries of what AI can accomplish.

About the Role

We're hiring for the Code RL team within the RL organization. As a Research Engineer, you'll advance our models' ability to safely write correct, fast code for accelerators.

You'll need to know accelerator performance well to turn it into tasks and signals models can learn from. Specifically, you will:

  • Invent, design and implement RL environments and evaluations.

  • Conduct experiments and shape our research roadmap.

  • Deliver your work into training runs.

  • Collaborate with other researchers, engineers, and performance engineering specialists across and outside Anthropic.

You may be a good fit if you:

  • Have expertise with accelerators (CUDA, ROCm, Triton, Pallas), ML framework programming (JAX or PyTorch).

  • Have worked across the stack – kernels, model code, distributed systems.

  • Know how to balance research exploration with engineering implementation.

  • Are passionate about AI's potential and committed to developing safe and beneficial systems.

Strong candidates may also have:

  • Experience with reinforcement learning.

  • Experience porting ML workloads between different types of accelerators.

  • Familiarity with LLM training methodologies.

The annual compensation range for this role is listed below.

For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.

Annual Salary:

$350,000—$850,000 USD

Logistics

Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience

Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience

Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position

Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.

We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.

Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic.com/careers directly for confirmed position openings.

How we're different

We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills.

The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.

Come work with us!

Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Research Engineer, Performance RL (Reinforcement Learning) in San Francisco, CA vacancy
  •  ...quickly growing group of committed researchers, engineers, policy experts, and business...  ...beneficial AI systems. About the RL Teams Our Reinforcement Learning teams play a critical role in...  ...autonomous engineering, to high-performance code for accelerators — and we'... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    2 days ago
  •  ...growing group of committed researchers, engineers, policy experts, and business...  .... About the role The RL Velocity team owns the efficiency...  ...Own the reliability and performance of research runs end-to-end...  ...Problems in AI Safety, and Learning from Human Preferences. Come... 
    Performance
    Remote job
    Work at office
    Visa sponsorship
    Flexible hours

    Neura Market

    San Francisco, CA
    4 days ago
  •  ...specialist to own end-to-end creation of reinforcement learning environments for new capabilities. You...  ...relationships, and measure impact on model performance. You will collaborate with domain experts to design data pipelines, run RL experiments, and translate capability... 
    Performance

    Neura Market

    San Francisco, CA
    4 days ago
  • $227.2k - $284k

     ...bring AI into production that performs when it matters most,...  ...production, paired with applied ML research, design, and evaluation to...  ...the RoleAs a Staff Machine Learning Research Engineer, you will operate across...  ...traces, online or offline RL — and validate them with rigorous... 
    Performance
    Full time

    Scale AI

    San Francisco, CA
    3 days ago
  •  ...growing group of committed researchers, engineers, policy experts, and...  ...at the frontier of machine learning, implementing and improving...  ...ML Systems Engineer on our Reinforcement Learning Engineering team,...  ...obsessively on improving the performance, robustness, and usability... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  • $264.8k - $331k

     ...training algorithms to reach the performance necessary for complex agents...  ...the world. The Enterprise ML Research Lab works on the front lines...  ...on applying our Agent RL Training + Building algorithms...  ...coverage, retirement benefits, a learning and development stipend, and... 
    Performance
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  • $264.8k - $331k

     ...RoleAs a Senior/Staff Machine Learning Engineer (MLE) on the General Agents...  ...metrics to measure agent performance, reliability, and business...  ...problem spaces, balancing research-driven approaches with pragmatic...  ...fine-tuning (SFT), reinforcement learning with verifiable rewards... 
    Performance
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  • $250k - $350k

    Research Engineer / Scientist (Robot Learning) About World Labs: We build foundational world models that can perceive...  ...on sim-to-real transfer, robust performance, scalable training and inference...  ..., including imitation learning, reinforcement learning for manipulation.... 
    Performance

    World Labs Inc.

    San Francisco, CA
    1 day ago
  •  ...organizations across IT, Engineering, Financial Services...  ...of models learn and improve. As an RL Environment Engineer...  ...work hands‑on with research teams at top AI labs...  ...to build and refine reinforcement learning environments...  ...and evaluate model performance beyond static tests... 
    Performance
    H1b
    Work at office
    Visa sponsorship
    Flexible hours
    Night shift

    AI Talent Now

    San Francisco, CA
    3 days ago
  • $110.7k - $379.2k

    Position Summary Research Engineer — Post-Training & Small Language...  ...optimization, and reinforcement learning / alignment workflows. • Build...  ...decisioning using verifiable-reward RL — designing reward signals...  ...adherence, and task performance. • Curate, clean,... 
    Performance
    Local area
    Visa sponsorship

    Deloitte

    San Francisco, CA
    3 days ago
  • $264.8k - $331k

     ...Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI AI is becoming vitally...  ...post-training algorithms to reach the performance necessary for complex agents in enterprises...  ...algorithms for our next-gen Agent RL training platform, support large... 
    Performance
    Full time
    Contract work
    For contractors
    For subcontractor
    Work at office

    Scale LLP

    San Francisco, CA
    17 hours ago
  • HeyMilo AI is hiring a Research Engineer to join our Applied AI team in San Francisco. You’ll design and build reinforcement learning environments with verifiable rewards for real-world...  ...harnesses to measure model performance. The role blends ML research with systems... 
    Performance

    HeyMilo AI

    San Francisco, CA
    4 days ago
  • $350k

     ...growing group of committed researchers, engineers, policy experts, and...  ...to-end process of creating RL environments for new capabilities...  ...measuring impact on model performance. Responsibilities Own...  ...cases Have experience with reinforcement learning, reward design, or... 
    Performance
    Work at office
    Visa sponsorship
    Flexible hours

    Neura Market

    San Francisco, CA
    4 days ago
  • $190k - $270k

     ...the TeamThe Databricks AI Research organization is pushing the...  ..., harness design, agentic reinforcement learning (RL), and the construction of...  ...exploration with product and engineering rigor.Clear communication...  ...eligibility for annual performance bonus, equity, and the benefits... 
    Performance
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $180k - $340k

     ...Research EngineerYou'll own the quality of AI across everything...  ...Gamma creates. As our Research Engineer, you'll design evaluation...  ...to improve baseline performance, preventing issues like poor...  ...techniques for LLMs including reinforcement learning and supervised fine-tuningExceptional... 
    Performance
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    17 hours ago
  • $176k - $255k

     ...accelerate progress in GenAI research. We are looking for...  ...and Research Engineers with expertise in LLM...  ...Computer Science, Machine Learning, AI, or a related...  ...understanding of deep learning, reinforcement learning, and large-...  ..., interview performance, and relevant education... 
    Performance
    Full time
    Shift work

    Scale AI

    San Francisco, CA
    more than 2 months ago
  • $220k - $300k

     ...Research Engineer, Post-TrainingVizcom is where design teams at companies...  ...that judgment — and learning how to model it — is the challenge...  ...optimization, and reinforcement learning.Build rigorous evaluations...  ...HaveExperience building high-performance training or inference... 
    Performance

    Vizcom

    San Francisco, CA
    3 days ago
  •  ...the intersection of machine learning infrastructure, applied AI,...  ...As a Staff Machine Learning Engineer, you will lead the technical...  ...frameworks to measure model performance across diverse use cases and...  ...repeatable pipelines that translate research into production impact. You... 
    Performance
    Full time
    Work experience placement
    Local area
    Immediate start

    Plaid Inc.

    San Francisco, CA
    17 hours ago
  •  ...applications. As an Applied AI Research Engineer, you’ll focus on human‑...  ...‑training, distillation, reinforcement learning, and rigorous evaluation...  ...evaluate model and agent performance, reliability, traceability...  ...ambitious ideas in GenAI, RL, agentic workflows, evaluation... 
    Performance
    Remote work
    Relocation package
    Flexible hours

    Code Metal

    San Francisco, CA
    3 days ago
  • Preference Model is seeking Research Engineers or Research Scientists to advance self-directed learning in AI. The role involves training and evaluating models within proprietary RL environments and optimizing ML infrastructure. Candidates will benefit from competitive... 

    Preference Model

    San Francisco, CA
    3 days ago
  • $220k - $290k

     ...Sciences LLC) is an Alphabet-founded research and development company whose...  ...Position Description:Calico seeks Machine Learning Research Engineers to join our rapidly growing ML team....  ...to optimize training and inference performance, implementing advanced strategies such... 
    Performance

    Calico Life Sciences

    South San Francisco, CA
    1 day ago
  • Research Engineer, Post-Training (All Industry Levels) Join to apply for the Research...  ...tuning AI models, optimizing their performance, and ensuring they meet the highest...  ...understanding of modern machine learning techniques (reinforcement learning, transformers, etc) Track... 
    Performance

    Character.AI

    San Francisco, CA
    17 hours ago
  •  ...AI models at leading research labs and enterprises....  ...requires continuous learning and evolution. You'll...  ...an Applied Research Engineer, you will be at the forefront...  ...processes, such as Reinforcement Learning from Human...  ...) impact model performance and alignment. Optimize... 
    Performance
    Flexible hours

    HRB

    San Francisco, CA
    1 day ago
  •  ...AI that empowers software engineers by automating production engineering...  ...end‑to‑end, balancing research and engineering to create...  ...AI models, improve performance, and reduce compute costs...  ...novel techniques, including reinforcement learning, retrieval‑augmented generation... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Resolve AI

    San Francisco, CA
    17 hours ago
  • $165k - $310k

    Senior Research Engineer, LLM Training & Post-Training New York, New York...  ...action over perfection and learn by shipping. Take...  ..., preference optimization, reinforcement learning, evaluation, and experimentation...  ...overhead, and performance bottlenecks. Design evaluation... 
    Performance
    For contractors
    For subcontractor
    Work at office
    Remote work
    Work from home
    Flexible hours
    2 days per week

    Lightning AI

    San Francisco, CA
    2 days ago
  • $365k

     ...growing group of committed researchers, engineers, policy experts, and...  ...About the Team Our Reinforcement Learning teams are central to advancing...  ..., code generation through RL, fundamental RL research for...  ...in RL research, covering performance against baselines, experiment... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    1 day ago
  • $350k

     ...growing group of committed researchers, engineers, policy experts, and...  ...creating training data and RL environments targeting visual...  ...Have experience with reinforcement learning, reward design, or training...  ...Finetuning Claude to maximize its performance using a particular set of... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    2 days ago
  • $159k - $296k

     ...world in a positive way. To learn more visit: The Motion...  ...self-driving trucks. As a research engineer for Learnable Planner you will...  ...(e.g., imitation and reinforcement learning, optimization-based...  ...incentive awards and an annual performance bonus. Perks/Benefits:... 
    Performance
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    9 days ago
  •  ...Overview We are seeking a ML/RL Engineer to join our Algo team and...  ...of Multi‑Agent Reinforcement Learning (MARL) and safety‑critical...  ...Constrained Learning: Lead the research and implementation of advanced...  ...experience, with opportunities for performance bonuses and equity.... 
    Performance
    Shift work

    Bot Auto

    San Francisco, CA
    3 days ago
  • $150k - $250k

     ...Join a lean, high-caliber engineering team building a synthetic...  ...Olympiad medalists and published researchers, with direct ownership...  ...Analyze model and agent performance on synthetic tasks to understand...  ...yet. ~ Familiarity with reinforcement learning, agentic AI workflows, or... 
    Performance
    Full time
    Visa sponsorship

    Clera

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Engineer, Performance RL (Reinforcement Learning). Be the first to apply!