Research Engineer, Performance RL (Reinforcement Learning)
$350kAnthropic
About Anthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
About the RL Teams
Our Reinforcement Learning teams lead Anthropic's reinforcement learning research and development, playing a critical role in advancing our AI systems. We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.6 and Opus 4.6. Our work spans several key areas:
-
Developing systems that enable models to use computers effectively
-
Advancing code generation through reinforcement learning
-
Pioneering fundamental RL research for large language models
-
Building scalable RL infrastructure and training methodologies
-
Enhancing model reasoning capabilities
We collaborate closely with Anthropic's alignment and frontier red teams to ensure our systems are both capable and safe. We partner with the applied production training team to bring research innovations into deployed models, and are dedicated to implement our research at scale. Our Reinforcement Learning teams sit at the intersection of cutting-edge research and engineering excellence, with a deep commitment to building high-quality, scalable systems that push the boundaries of what AI can accomplish.
About the Role
We're hiring for the Code RL team within the RL organization. As a Research Engineer, you'll advance our models' ability to safely write correct, fast code for accelerators.
You'll need to know accelerator performance well to turn it into tasks and signals models can learn from. Specifically, you will:
-
Invent, design and implement RL environments and evaluations.
-
Conduct experiments and shape our research roadmap.
-
Deliver your work into training runs.
-
Collaborate with other researchers, engineers, and performance engineering specialists across and outside Anthropic.
You may be a good fit if you:
-
Have expertise with accelerators (CUDA, ROCm, Triton, Pallas), ML framework programming (JAX or PyTorch).
-
Have worked across the stack – kernels, model code, distributed systems.
-
Know how to balance research exploration with engineering implementation.
-
Are passionate about AI's potential and committed to developing safe and beneficial systems.
Strong candidates may also have:
-
Experience with reinforcement learning.
-
Experience porting ML workloads between different types of accelerators.
-
Familiarity with LLM training methodologies.
The annual compensation range for this role is listed below.
For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.
Annual Salary:
$350,000—$850,000 USD
Logistics
Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.
Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.
Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic.com/careers directly for confirmed position openings.How we're different
We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills.
The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.
Come work with us!
Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.
- ...quickly growing group of committed researchers, engineers, policy experts, and business... ...beneficial AI systems. About the RL Teams Our Reinforcement Learning teams play a critical role in... ...autonomous engineering, to high-performance code for accelerators — and we'...PerformanceFull timeWork at officeVisa sponsorshipFlexible hours
- ...growing group of committed researchers, engineers, policy experts, and business... .... About the role The RL Velocity team owns the efficiency... ...Own the reliability and performance of research runs end-to-end... ...Problems in AI Safety, and Learning from Human Preferences. Come...PerformanceRemote jobWork at officeVisa sponsorshipFlexible hours
- ...specialist to own end-to-end creation of reinforcement learning environments for new capabilities. You... ...relationships, and measure impact on model performance. You will collaborate with domain experts to design data pipelines, run RL experiments, and translate capability...Performance
$227.2k - $284k
...bring AI into production that performs when it matters most,... ...production, paired with applied ML research, design, and evaluation to... ...the RoleAs a Staff Machine Learning Research Engineer, you will operate across... ...traces, online or offline RL — and validate them with rigorous...PerformanceFull time- ...growing group of committed researchers, engineers, policy experts, and... ...at the frontier of machine learning, implementing and improving... ...ML Systems Engineer on our Reinforcement Learning Engineering team,... ...obsessively on improving the performance, robustness, and usability...PerformanceFull timeWork at officeVisa sponsorshipFlexible hours
$264.8k - $331k
...training algorithms to reach the performance necessary for complex agents... ...the world. The Enterprise ML Research Lab works on the front lines... ...on applying our Agent RL Training + Building algorithms... ...coverage, retirement benefits, a learning and development stipend, and...PerformanceFull time$264.8k - $331k
...RoleAs a Senior/Staff Machine Learning Engineer (MLE) on the General Agents... ...metrics to measure agent performance, reliability, and business... ...problem spaces, balancing research-driven approaches with pragmatic... ...fine-tuning (SFT), reinforcement learning with verifiable rewards...PerformanceFull time$250k - $350k
Research Engineer / Scientist (Robot Learning) About World Labs: We build foundational world models that can perceive... ...on sim-to-real transfer, robust performance, scalable training and inference... ..., including imitation learning, reinforcement learning for manipulation....Performance- ...organizations across IT, Engineering, Financial Services... ...of models learn and improve. As an RL Environment Engineer... ...work hands‑on with research teams at top AI labs... ...to build and refine reinforcement learning environments... ...and evaluate model performance beyond static tests...PerformanceH1bWork at officeVisa sponsorshipFlexible hoursNight shift
$110.7k - $379.2k
Position Summary Research Engineer — Post-Training & Small Language... ...optimization, and reinforcement learning / alignment workflows. • Build... ...decisioning using verifiable-reward RL — designing reward signals... ...adherence, and task performance. • Curate, clean,...PerformanceLocal areaVisa sponsorship$264.8k - $331k
...Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI AI is becoming vitally... ...post-training algorithms to reach the performance necessary for complex agents in enterprises... ...algorithms for our next-gen Agent RL training platform, support large...PerformanceFull timeContract workFor contractorsFor subcontractorWork at office- HeyMilo AI is hiring a Research Engineer to join our Applied AI team in San Francisco. You’ll design and build reinforcement learning environments with verifiable rewards for real-world... ...harnesses to measure model performance. The role blends ML research with systems...Performance
$350k
...growing group of committed researchers, engineers, policy experts, and... ...to-end process of creating RL environments for new capabilities... ...measuring impact on model performance. Responsibilities Own... ...cases Have experience with reinforcement learning, reward design, or...PerformanceWork at officeVisa sponsorshipFlexible hours$190k - $270k
...the TeamThe Databricks AI Research organization is pushing the... ..., harness design, agentic reinforcement learning (RL), and the construction of... ...exploration with product and engineering rigor.Clear communication... ...eligibility for annual performance bonus, equity, and the benefits...PerformanceLocal areaWorldwide$180k - $340k
...Research EngineerYou'll own the quality of AI across everything... ...Gamma creates. As our Research Engineer, you'll design evaluation... ...to improve baseline performance, preventing issues like poor... ...techniques for LLMs including reinforcement learning and supervised fine-tuningExceptional...PerformanceFull timeWork at officeWork from home$176k - $255k
...accelerate progress in GenAI research. We are looking for... ...and Research Engineers with expertise in LLM... ...Computer Science, Machine Learning, AI, or a related... ...understanding of deep learning, reinforcement learning, and large-... ..., interview performance, and relevant education...PerformanceFull timeShift work$220k - $300k
...Research Engineer, Post-TrainingVizcom is where design teams at companies... ...that judgment — and learning how to model it — is the challenge... ...optimization, and reinforcement learning.Build rigorous evaluations... ...HaveExperience building high-performance training or inference...Performance- ...the intersection of machine learning infrastructure, applied AI,... ...As a Staff Machine Learning Engineer, you will lead the technical... ...frameworks to measure model performance across diverse use cases and... ...repeatable pipelines that translate research into production impact. You...PerformanceFull timeWork experience placementLocal areaImmediate start
- ...applications. As an Applied AI Research Engineer, you’ll focus on human‑... ...‑training, distillation, reinforcement learning, and rigorous evaluation... ...evaluate model and agent performance, reliability, traceability... ...ambitious ideas in GenAI, RL, agentic workflows, evaluation...PerformanceRemote workRelocation packageFlexible hours
- Preference Model is seeking Research Engineers or Research Scientists to advance self-directed learning in AI. The role involves training and evaluating models within proprietary RL environments and optimizing ML infrastructure. Candidates will benefit from competitive...
$220k - $290k
...Sciences LLC) is an Alphabet-founded research and development company whose... ...Position Description:Calico seeks Machine Learning Research Engineers to join our rapidly growing ML team.... ...to optimize training and inference performance, implementing advanced strategies such...Performance- Research Engineer, Post-Training (All Industry Levels) Join to apply for the Research... ...tuning AI models, optimizing their performance, and ensuring they meet the highest... ...understanding of modern machine learning techniques (reinforcement learning, transformers, etc) Track...Performance
- ...AI models at leading research labs and enterprises.... ...requires continuous learning and evolution. You'll... ...an Applied Research Engineer, you will be at the forefront... ...processes, such as Reinforcement Learning from Human... ...) impact model performance and alignment. Optimize...PerformanceFlexible hours
- ...AI that empowers software engineers by automating production engineering... ...end‑to‑end, balancing research and engineering to create... ...AI models, improve performance, and reduce compute costs... ...novel techniques, including reinforcement learning, retrieval‑augmented generation...PerformanceFull timeWork at officeVisa sponsorshipFlexible hours
$165k - $310k
Senior Research Engineer, LLM Training & Post-Training New York, New York... ...action over perfection and learn by shipping. Take... ..., preference optimization, reinforcement learning, evaluation, and experimentation... ...overhead, and performance bottlenecks. Design evaluation...PerformanceFor contractorsFor subcontractorWork at officeRemote workWork from homeFlexible hours2 days per week$365k
...growing group of committed researchers, engineers, policy experts, and... ...About the Team Our Reinforcement Learning teams are central to advancing... ..., code generation through RL, fundamental RL research for... ...in RL research, covering performance against baselines, experiment...PerformanceFull timeWork at officeVisa sponsorshipFlexible hoursShift work$350k
...growing group of committed researchers, engineers, policy experts, and... ...creating training data and RL environments targeting visual... ...Have experience with reinforcement learning, reward design, or training... ...Finetuning Claude to maximize its performance using a particular set of...PerformanceFull timeWork at officeVisa sponsorshipFlexible hours$159k - $296k
...world in a positive way. To learn more visit: The Motion... ...self-driving trucks. As a research engineer for Learnable Planner you will... ...(e.g., imitation and reinforcement learning, optimization-based... ...incentive awards and an annual performance bonus. Perks/Benefits:...PerformanceFull timeWork at officeWork from homeFlexible hours- ...Overview We are seeking a ML/RL Engineer to join our Algo team and... ...of Multi‑Agent Reinforcement Learning (MARL) and safety‑critical... ...Constrained Learning: Lead the research and implementation of advanced... ...experience, with opportunities for performance bonuses and equity....PerformanceShift work
$150k - $250k
...Join a lean, high-caliber engineering team building a synthetic... ...Olympiad medalists and published researchers, with direct ownership... ...Analyze model and agent performance on synthetic tasks to understand... ...yet. ~ Familiarity with reinforcement learning, agentic AI workflows, or...PerformanceFull timeVisa sponsorship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Engineer, Performance RL (Reinforcement Learning). Be the first to apply!
- research programmer San Francisco, CA
- research engineer San Francisco, CA
- senior research engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- deep learning research engineer San Francisco, CA
- research software engineer San Francisco, CA
- research assistant engineering San Francisco, CA
- ai research engineer San Francisco, CA
- human performance consultant San Francisco, CA
- performance test architect San Francisco, CA




