Research Scientist - RL Training
Snorkel AI
Role Description
We're looking for a Research Scientist to work on reinforcement learning for training and aligning large language models. This is a foundational research role focused on one of the most consequential open data problems in AI:
- How to generate the data, reward signals, and training procedures that steer LLM behavior in reliable and generalizable directions.
- A core capability that directly differentiates Snorkel's data-as-a-service offering.
You'll work closely with Snorkel's research, engineering, and delivery teams to advance our RL data capabilities:
- Translating research ideas into the preference datasets, reward models, and RL-ready corpora we produce for frontier AI labs.
- Contributing to a research agenda that is central to Snorkel's long-term differentiation as a provider of bespoke training data.
Qualifications
- Deep expertise in reinforcement learning from human or AI feedback, reward modeling, and credit attribution.
- Experience training or fine-tuning 30B+ large language models at scale.
- Strong proficiency in Python and ML frameworks, especially PyTorch and HuggingFace.
- Hands-on experience with RL frameworks such as Verl and SkyRL.
- Solid software engineering fundamentals.
- Familiarity with ML infrastructure and cloud platforms and tools (AWS, GCP, Kubernetes, Slurm, etc.).
- Comfort operating in a high-iteration environment with open-ended research questions.
- Ph.D. in machine learning, reinforcement learning, or a related field strongly preferred.
Requirements
- Research and implement reinforcement learning techniques — including GRPO, RLHF, RLAIF, DPO, and reward modeling.
- Design and build data pipelines that generate high-quality training signals for RL workflows.
- Prototype and iterate on end-to-end RL training recipes.
- Work closely with research scientists, ML engineers, and delivery teams.
- Stay current with the latest developments in large-scale multi-node LLM training, alignment research, and scalable RL methods.
- Contribute to Snorkel's research publications and internal knowledge base in RL and model training.
Benefits
- Become part of a company with market-proven solutions and robust funding.
- Opportunities to shape priorities and initiatives.
- Influence key strategic decisions and directly impact ongoing success.
- Support in building your career in an environment designed for growth, learning, and shared success.
$204k - $259k
...foster collaborations with other research teams in Alphabet. AI... ...you will report to a Principal Scientist. Responsibilities Participate in... ...’s Foundation World Model post‑training and evaluation Research and develop cutting edge RL and Distillation techniques for...TrainingTemporary workRemote work$175k - $220k
...passionate about using reinforcement learning to solve dexterous manipulation tasks? As an RL Research Scientist, you'll lead research projects on visual sim-to-real transfer and post-training of Vision‑Language‑Action (VLA) models. Your job is to turn unlabeled data from...Training£87.68k - £120.56k per year
Role Description Research Scientists at Phaidra lead our efforts in developing novel algorithmic architecture... ...across systems, including the training pipelines — pretraining, curriculum... ...Research and implement methods for safe RL, constrained control, scenario planning...TrainingFull timeRemote workFlexible hours$150k - $173k
...will join a small but world-class Applied Research and AI team and work on genuinely hard, open... ...security operations — and can be used to train, evaluate, and iterate on decision-making... ...implementing and evaluating deep RL algorithms; fluency in policy gradient methods...TrainingBi-weekly payFull timeShift work£120.28k - £165.38k per year
...Australia, and India. Who You Are Research Scientists at Phaidra lead our efforts in... ...generalize across systems, including the training pipelines — pretraining, curriculum learning... ...Research and implement methods for e.g. safe RL, constrained control, scenario planning...TrainingRemote jobFull timeTemporary workWork at officeLocal areaFlexible hours$300k
...Research Scientist — Frontier World Models & RL A stealth, exceptionally well-backed applied AI lab is hiring Research Scientists to solve open problems at... ...current limits, with real users and a data flywheel to train against the moment beta launches. The open problems you...TrainingRemote workVisa sponsorshipRelocation package- ...operating at the frontier of machine learning research. Founded by leading researchers, repeat... ...AI, agentic systems, and advanced training methodologies. These teams are tackling some... ...Reinforcement Learning, RLHF, RLAIF, online RL, offline RL, and scalable RL systems. Mid...TrainingCurrently hiring
- ...alignment through entertainment. We need a Research Scientist to help design and execute research... ...improves language models through game-based training. You'll work on improving a 260B+... ...model we have under contract by EOY, using RL environments and game-generated data to...TrainingContract workRemote workVisa sponsorshipFlexible hours
- ...Member of Technical Staff, Reinforcement Learning Research Our client is a well-funded, early-stage AI lab building a real-time,... ...on problems that are genuinely unsolved. About the Role Own RL and post-training for large-scale multimodal models at a frontier lab, from building...TrainingInternshipRelocation packageShift work
£87.68k - £165.38k per year
...Our partner is looking for a Senior AI Research Scientist (Model-based RL) based in Spain. This role offers the opportunity to shape the future of... ...capable of generalizing across different systems, including training approaches such as pretraining, curriculum learning,...TrainingRemote workFlexible hours$126k - $423k
...team We are looking for multiple passionate Research Scientists to join the Research Group at Applied Intuition... ...Conduct research on reinforcement learning (RL) related topics including large‑scale self‑play RL, VLA post‑training, large‑scale closed‑loop RL based on neural...TrainingFull timeFor contractorsFor subcontractorCasual workWork at officeImmediate startRemote workDay shift- ...About the role We’re looking for a top‑tier Research Scientist to join our tech team. Your core... ...ranking systems for AI agents Prototype, train, and evaluate new models for factual search... ...progress of our search engine Design a RL framework to integrate AI preferences into...TrainingRemote workFlexible hours
- ...Role We’re looking for Applied Scientists to join Wayve Labs and help... ...Wayve, we are a high‑conviction research team with the strategic... ...transformers, MoE, large‑scale training). Generative world modeling (... ...Reinforcement learning (e.g., offline RL, RLHF, reward modeling)....TrainingFull timeWork at officeWork from homeVisa sponsorshipRelocation packageFlexible hours
$200k - $250k
...enterprise workflows ~Post-train LLM agents using RLHF, DPO,... ...engineering standards ~Mentor researchers and engineers; drive... ...equivalent) ~5+ years hands-on RL — environment design, reward... ...Subject: Senior Staff Research Scientist – RL Company Description...TrainingFull time$66k - $94k
Associate Research Scientist – Evidence Generation (Primary Data Collection) Are you looking to work with a dynamic team of industry leading... ...including but not limited to: skill sets, experience and training, licensure and certifications, and other business and organizational...TrainingFull timeInternshipLocal areaRemote workWorldwide$225k - $300k
...Research Scientist About Latent Health Healthcare today is only truly personalized for two groups: those with wealth and access,... ...work on: Verifiable reinforcement learning at scale Mid-training and post-training of foundation models Novel objectives...TrainingFull timeWork at officeImmediate start$168k - $264.5k
Role Description We're looking for a Research Scientist who shares our vision for transforming healthcare. This is a unique opportunity to... ...clinical applications. ~5 plus years of experience building and training machine learning models at scale, along with the compute...TrainingFull timeShift work$175k - $225k
Role Description Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation...TrainingFull time$100k - $150k
Role Description We are seeking an AI Research Engineer to bridge cutting-edge applied research and production engineering, designing and... ...ML frameworks such as PyTorch or JAX. ~Hands-on experience training, fine-tuning, and evaluating deep learning models at non-...TrainingFull timeLocal areaImmediate start$133.19k - $190.28k
Role Description We are seeking Research Scientists (across all levels of seniority) to join our Artist-First AI Music Lab. Our team pioneers... ...ML-based audio processing and signal processing. ~Post-Training: Research in post-training techniques for music generation,...TrainingFull timeFlexible hours$126.89k - $166.13k
Role Description We are looking for a Senior Microwave Research Scientist. As a Senior Microwave Research Scientist, you’ll be part of a cross... ...field with 5+ years of experience ~Experience hiring and training new teams for research and development work ~Knowledge of...TrainingFull timeWork at office$86.01 - $147.94 per hour
...candidate will be dependent on a variety of factors including, but not limited to, the candidate’s experience, qualifications, location, training, licenses, shifts worked and compensation model. Carle Health offers a comprehensive benefits package for team members and...TrainingFull timeLocal areaRemote workRelocation packageShift work$40 per hour
...DataAnnotation is seeking a Research Scientist (Biology) to contribute to AI model training and evaluation. In this role, you will assess AI chatbots with complex biology questions and evaluate their outputs for quality and performance. The ideal candidate will have expertise...TrainingHourly payRemote workFlexible hours$40 per hour
A leading AI training company is seeking a Biology Research Scientist for a remote role in the United States. This position involves training AI models by measuring their progress, evaluating their logic, and solving complex biology problems. The ideal candidate should...TrainingHourly payRemote work- ...A leading AI healthcare organization in the United States seeks a Research Scientist Intern to advance their AI models. The role involves architecting algorithms, training models on high-performance clusters, and collaborating with healthcare professionals. Ideal candidates...TrainingInternshipRemote work
- ...solve real‑world business problems. We’re training and deploying frontier models for... ...for our customers. Cohere is a team of researchers, engineers, designers, and more, who are... ...measure LLM progress. As a Senior Research Scientist, Model Evaluation, you will: Create...TrainingFull timeWork at officeLocal areaRemote workHome office
$30 per hour
A leading AI training firm is seeking an Applied Physics Research Scientist to evaluate AI chatbots on physics-related challenges. This is a remote position ideal for experts with a strong grasp of classical mechanics and related fields. Applicants must possess fluency...TrainingHourly payRemote workFlexible hours- A leading AI development firm is seeking a Mathematician (PhD) to join their remote AI training project. In this role, you will evaluate AI-generated mathematical responses, ensuring accuracy and clarity. To qualify, you must have a PhD in Mathematics/Statistics, significant...TrainingWeekly payRemote workFlexible hours
$40 per hour
A leading AI training company is looking for a Math Expert to enhance AI models by solving complex mathematical problems and evaluating their outputs. This REMOTE position requires fluency in English and expertise in arithmetic, algebra, geometry, calculus, and statistics...TrainingHourly payContract workRemote work$40 per hour
A leading AI training company is seeking a Research Scientist (Biology) to evaluate and train AI chatbots by measuring their logic and solving complex biological questions. This role requires an expert understanding of biology and related fields. Candidates should be fluent...TrainingHourly payRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist - RL Training. Be the first to apply!









