Research Engineer / Research Scientist, RL Frontiers
Anthropic
About Anthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
About the role
Reinforcement learning is how Claude learns to reason, write code, and act autonomously over long horizons. The RL Scaling team works on how RL scales: what happens to throughput, stability, and learning efficiency as models get larger, episodes get longer, and compute grows by orders of magnitude, and what has to change in our algorithms and systems to keep getting returns from that scale.
This role sits squarely across research and engineering. You'll develop next-generation architectures and RL algorithms, take them from a small-scale result to a frontier-scale run, and understand every place they behave differently along the way. You'll build the systems that set how fast the team can iterate: how many experiments, at what scale, and how quickly we can trust the results. And you'll work on Anthropic's largest and fastest RL runs, where the gap between a good idea and a working one is often a problem no one has solved yet.
Key responsibilities
- Study how RL training and sampling scale with model size, context length, and compute, and find the algorithmic and systems changes that keep scaling efficient
- Develop next-generation model architectures and RL algorithms, and make them run efficiently at frontier scale
- Take promising small-scale results to frontier-scale runs, and diagnose why they behave differently when they get there, whether the cause is numerical, algorithmic, or systemic
- Build the experimental infrastructure that sets research velocity: fast, reproducible comparisons of architecture and algorithm variants at meaningful scale
- Own end-to-end performance of our largest RL runs, from research code down to the hardware
- Build performance and cost models for proposed architecture and algorithm changes, and use them to decide which ideas get scaled
- Investigate training dynamics at scale, including instabilities, divergence, and throughput regressions, and trace them to root cause
Minimum qualifications
- Deep familiarity with modern transformer language models, including their architecture, training dynamics, and the behavior of large-scale optimization
- Hands-on experience training large models in a distributed setting, including the tradeoffs between data, tensor, and pipeline parallelism
- A track record of original technical work in ML training or systems, such as new methods, architectures, or optimizations, demonstrated through research, open-source, or production impact
- Ability to design rigorous experiments at scale, including baselines, ablations, and enough statistical care to trust a result that costs real compute
- Ability to reason quantitatively about the compute, memory, and communication costs of a model or algorithm
- Strong programming skills in Python and JAX or PyTorch, and comfort reading and changing code at every layer of the stack
Preferred qualifications
- Research experience in reinforcement learning, optimization, or large-scale training, published or otherwise
- Experience developing RL algorithms for language models
- Experience with scaling laws or other quantitative models of training efficiency
- Experience designing or modifying transformer architectures beyond standard configurations
- Experience scaling training to large fleets of accelerators and debugging the problems that only appear at scale
- Deep understanding of numerics in large-scale training, including low-precision formats and sources of instability
- Familiarity with how GPU or TPU performance characteristics shape architecture and algorithm choices
- Experience with C++ or Rust
Representative projects
- Characterize how a new RL algorithm's throughput and learning efficiency change from small models to frontier scale, and fix what breaks
- Develop a new attention variant, get it working at full scale, and measure how its quality and throughput compare to the baseline
- Prepare our next largest-ever RL run: find what breaks when model size, context length, and compute all grow at once, and fix it before launch
- Trace a loss instability that only appears past a certain scale to its root cause, and work out whether the fix belongs in the algorithm, the numerics, or the system
- Build a model that predicts the throughput and cost of a proposed architecture change before anyone writes the kernel
The annual compensation range for this role is listed below.
For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.
Annual Salary:
$500,000—$850,000 USD
Logistics
Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.
Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.
We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.
Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic.com/careers directly for confirmed position openings.How we're different
We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills.
The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.
Come work with us!
Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.
- ...whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to... ...write code, and act autonomously over long horizons. At frontier scale, an RL run is an unusually demanding distributed system....SuggestedFull timeWork at officeVisa sponsorshipFlexible hoursShift work
- ...Description Job Description We are looking for a hybrid Systems Engineer and AI Researcher to lead the development of our agent evaluation framework... ...+ step action trajectories safely. Scale Post-Training & RL Pipelines: Implement high-throughput post-training...SuggestedWork at office
- ...Job Description Job Description San Francisco, California | Primarily On-site We are seeking an Research Engineer – RL Infrastructure & Agent Environments to build the environments, evaluation systems, and supporting infrastructure used to train and assess long...Suggested
$180k - $280k
...The Impact You’ll Make Our research team is expanding to keep pace with a wave of frontier‑facing work: internal research streams... ...of the field. As a Research Engineer, you’ll take a research direction... ...Own projects (for example, an RL/agentic environment build for a...SuggestedFull time$275k
...path to safe AGI lies in automating research and code generation to improve... ...can alone. Our approach combines frontier-scale pre-training, domain-specific RL, ultra-long context, and inference... ...About the role As a Research Engineer, you'll work on training, evaluating...SuggestedRelocationVisa sponsorship- ...AfterQuery is an applied research lab curating data solutions for... ...development. We serve every frontier AI lab with the mission of delivering... ...the opportunity to shape the engineering organization and lead major... .... Through controlled SFT and RL - base post-training...Local area
- ...Research Engineer New York - Hybrid; San Francisco Bay Area - Hybrid About Invisible... ...methodology, benchmarks, and RL environments for frontier labs and enterprise clients. As a Research... ...Partner with Research Scientists on methodology review, with Solutions...Full timeWork at officeLocal areaRemote work
$200k - $350k
...second-time technical founders, engineers that made 100+ games for... ...3D environments. Our current research spans: Distributed multi... ...generative vision, world models, or RL systems. Strong... ...mission-driven team working at the frontier of AI world generation. We're...Visa sponsorshipRelocation package- ...infrastructure / Reinforcement Learning (RL) training data & evaluations... ...learning (RL) training data and evaluation for frontier AI agents. Their platform is used by advanced... ...Opportunity Our partner is hiring a Research Engineer to help scale the quality assurance (QA...Remote work
- ...Ando Research Team Member Ando is a messaging platform where AI agents take on work alongside... ..., vs where do we genuinely need RL and continual learning? What you'll... ...in-house. Part of the job is evaluating frontier vendors and research teams (eval infrastructure...Work from home
$150k - $200k
...About Kovari Kovari develops frontier technology to transform physical businesses. We partner... ...a small team of builders, operators, and researchers. We deploy on site, alongside the... ...robot policies on hardware (model-based, RL, or imitation learning) Sim-to-real or...Full time- ...is a quickly growing group of committed researchers, engineers, policy experts, and business leaders... ...partner with research teams across every RL domain, bringing their priorities into... ...impacts of your work and about shipping frontier models responsibly Strong...Full timeWork at officeVisa sponsorshipFlexible hours
- ...The Role As a Research Engineer, you'll build the systems that let us post-train models continuously: the training pipelines, environment... ...trace collection, data curation, environment generation, and the RL stack that improves agents over time. When a training run...
$175k - $250k
...Research Engineer About Scorecard We’re a small, nimble team backed by top-tier investors. We have multi-billion dollar customers and a... ...design, grader calibration, or human annotation pipelines. RL experience: building RL environments for LLMs, reward design,...Work at office$100k - $300k
...contribute to our innovative projects. Position Overview We are hiring Research Engineers to develop scalable robotic systems aimed at achieving general... ...research across multiple disciplines (Perception, Robotics, RL/IL, Machine Learning, etc.). Collaborate with a...Full time- ...Agentic AI that empowers software engineers by automating production... ...Unusual Ventures, Jeff Dean (Chief Scientist, Google DeepMind), Thomas... ...end‑to‑end, balancing research and engineering to create production... ...Future: Help build the next frontier in enterprise software and...Full timeWork at officeVisa sponsorshipFlexible hours
- ...Research Engineer On Physical AI Team Hedra is a pioneering generative modeling company — first models to market — now building a Physical... ...and your impact will be direct. If you want to work at the frontier of generative modeling and physical AI, this is the team....Work at office
- ...suite for molecules. We train frontier models that learn the... ...structure and interaction, so scientists can move faster and pursue targets... ...way it reinvented software engineering, and Chai is at the... ...Role We are seeking an AI Research Engineer to help design, train...Shift work
- ...About Phonic Phonic is a product and research lab focused on powering the most realistic... ...and perform agentic tasks with frontier intelligence. Our team includes top-tier... ...office. About The Role As a Research Engineer at Phonic, you'll sit at the intersection...Work at office
$200k - $400k
...Research Engineer Decagon is the leading conversational AI platform empowering every brand to deliver concierge customer experiences.... ...capability, and efficiency forward, and you'll design and implement frontier approaches for training, evaluation, and orchestration...Full timeWork at officeLocal area- ...well known in the AI community for seminal research accomplishments at top AI labs, have run... ...Chai Discovery is seeking an AI Engineer to play a crucial role in developing our... ...We Offer Highly engaging work at the frontier of AI-driven drug discovery that will fundamentally...Full time
- ...Goliath Partners has exclusively partnered with a fast-growing, venture-backed frontier GenAI Research Lab that is operating at serious scale & expanding. They’re looking for a Research Engineer to help build the data behind their next-generation world models. You'll own...
- ...About Human Archive Human Archive is a research lab backed by Y Combinator focused on... .... The Opportunity As a Research Engineer, you'll work on multimodal sensing... ...datasets. Your work will help shape how frontier labs and leading robotics companies train...Shift work
- ...enables innovation across Plaid. As a Staff Machine Learning Engineer, you will lead the technical strategy and development of Plaid’... ...cases and build scalable, repeatable pipelines that translate research into production impact. You will also partner closely with teams...Full timeWork experience placementLocal areaImmediate start
$350k
...a quickly growing group of committed researchers, engineers, policy experts, and business leaders... ...You could describe yourself as both a scientist and an engineer. As a Research Engineer... ...Interpretability, Fine-Tuning, and the Frontier Red Team. Our blog provides an overview...Work at officeVisa sponsorshipFlexible hours- ...Self-Improvement (RSI) team works across research, engineering, product, and infrastructure to build... ...About the Role We're hiring research scientists , research engineers , and AI... ...through agent harnesses, synthetic data, RL environments, and model training. Build...
$150k - $180k
...reinforcement learning environments for frontier model training. You will own... ...data quality, shaping research culture around what makes... ...evaluate thousands of tasks across RL environments, synthetic data,... .... Partner with research engineers, domain experts, and data...Full timeVisa sponsorship$365k
...is a quickly growing group of committed researchers, engineers, policy experts, and business leaders... ...interpretability, and safety, each operating at the frontier of AI development. As a Technical... ...research areas like compute, evals, RL environments, and emerging research...Work at officeVisa sponsorshipFlexible hoursShift work$200k - $350k
We’re partnering with a frontier AI startup We’re working with an early-stage AI company that’... ...deeply curious—building at the intersection of research, product, and creativity . The Role As a Machine Learning Research Engineer , you’ll own end-to-end research cycles—...$190k - $310k
...Intelligence legible and intuitive to our audience of engineers, researchers, and heads of AI/ML at some of the largest companies... ...public-facing work. Technical fluency. We publish frontier work on post-training, RL, and async training, and would be excited to work with...Full timeWork at officeVisa sponsorshipRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Engineer / Research Scientist, RL Frontiers. Be the first to apply!
- deep learning research engineer San Francisco, CA
- research software engineer San Francisco, CA
- research engineer San Francisco, CA
- research programmer San Francisco, CA
- ai research engineer San Francisco, CA
- safety scientist San Francisco, CA
- graduate scientist San Francisco, CA
- remote scientist San Francisco, CA
- research scientist - biology San Francisco, CA
- scientist 1 San Francisco, CA



