Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Scientist - RL Training

Full-time

Snorkel AI

Role Description

We're looking for a Research Scientist to work on reinforcement learning for training and aligning large language models. This is a foundational research role focused on one of the most consequential open data problems in AI:

  • How to generate the data, reward signals, and training procedures that steer LLM behavior in reliable and generalizable directions.
  • A core capability that directly differentiates Snorkel's data-as-a-service offering.

You'll work closely with Snorkel's research, engineering, and delivery teams to advance our RL data capabilities:

  • Translating research ideas into the preference datasets, reward models, and RL-ready corpora we produce for frontier AI labs.
  • Contributing to a research agenda that is central to Snorkel's long-term differentiation as a provider of bespoke training data.

Qualifications

  • Deep expertise in reinforcement learning from human or AI feedback, reward modeling, and credit attribution.
  • Experience training or fine-tuning 30B+ large language models at scale.
  • Strong proficiency in Python and ML frameworks, especially PyTorch and HuggingFace.
  • Hands-on experience with RL frameworks such as Verl and SkyRL.
  • Solid software engineering fundamentals.
  • Familiarity with ML infrastructure and cloud platforms and tools (AWS, GCP, Kubernetes, Slurm, etc.).
  • Comfort operating in a high-iteration environment with open-ended research questions.
  • Ph.D. in machine learning, reinforcement learning, or a related field strongly preferred.

Requirements

  • Research and implement reinforcement learning techniques — including GRPO, RLHF, RLAIF, DPO, and reward modeling.
  • Design and build data pipelines that generate high-quality training signals for RL workflows.
  • Prototype and iterate on end-to-end RL training recipes.
  • Work closely with research scientists, ML engineers, and delivery teams.
  • Stay current with the latest developments in large-scale multi-node LLM training, alignment research, and scalable RL methods.
  • Contribute to Snorkel's research publications and internal knowledge base in RL and model training.

Benefits

  • Become part of a company with market-proven solutions and robust funding.
  • Opportunities to shape priorities and initiatives.
  • Influence key strategic decisions and directly impact ongoing success.
  • Support in building your career in an environment designed for growth, learning, and shared success.
Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Research Scientist - RL Training in Remote vacancy
  • $204k - $259k

     ...foster collaborations with other research teams in Alphabet. AI...  ...you will report to a Principal Scientist. Responsibilities Participate in...  ...’s Foundation World Model post‑training and evaluation Research and develop cutting edge RL and Distillation techniques for... 
    Training
    Temporary work
    Remote work

    Waymo

    San Francisco, CA
    3 days ago
  • $175k - $220k

     ...passionate about using reinforcement learning to solve dexterous manipulation tasks? As an RL Research Scientist, you'll lead research projects on visual sim-to-real transfer and post-training of Vision‑Language‑Action (VLA) models. Your job is to turn unlabeled data from... 
    Training

    Boston Dynamics

    Waltham, MA
    2 days ago
  • £87.68k - £120.56k per year

    Role Description Research Scientists at Phaidra lead our efforts in developing novel algorithmic architecture...  ...across systems, including the training pipelines — pretraining, curriculum...  ...Research and implement methods for safe RL, constrained control, scenario planning... 
    Training
    Full time
    Remote work
    Flexible hours

    Phaidra

    Remote
    9 hours ago
  • $150k - $173k

     ...will join a small but world-class Applied Research and AI team and work on genuinely hard, open...  ...security operations — and can be used to train, evaluate, and iterate on decision-making...  ...implementing and evaluating deep RL algorithms; fluency in policy gradient methods... 
    Training
    Bi-weekly pay
    Full time
    Shift work

    The Nuclear Company

    Washington DC
    2 hours ago
  • £120.28k - £165.38k per year

     ...Australia, and India. Who You Are Research Scientists at Phaidra lead our efforts in...  ...generalize across systems, including the training pipelines — pretraining, curriculum learning...  ...Research and implement methods for e.g. safe RL, constrained control, scenario planning... 
    Training
    Remote job
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours

    Phaidra

    Remote
    6 days ago
  • $300k

     ...Research Scientist — Frontier World Models & RL A stealth, exceptionally well-backed applied AI lab is hiring Research Scientists to solve open problems at...  ...current limits, with real users and a data flywheel to train against the moment beta launches. The open problems you... 
    Training
    Remote work
    Visa sponsorship
    Relocation package

    Harnham

    San Francisco, CA
    4 days ago
  •  ...operating at the frontier of machine learning research. Founded by leading researchers, repeat...  ...AI, agentic systems, and advanced training methodologies. These teams are tackling some...  ...Reinforcement Learning, RLHF, RLAIF, online RL, offline RL, and scalable RL systems. Mid... 
    Training
    Currently hiring

    Alexander Chapman

    California, MO
    5 days ago
  •  ...alignment through entertainment. We need a Research Scientist to help design and execute research...  ...improves language models through game-based training. You'll work on improving a 260B+...  ...model we have under contract by EOY, using RL environments and game-generated data to... 
    Training
    Contract work
    Remote work
    Visa sponsorship
    Flexible hours

    Good Start Labs

    New York, NY
    2 days ago
  •  ...Member of Technical Staff, Reinforcement Learning Research Our client is a well-funded, early-stage AI lab building a real-time,...  ...on problems that are genuinely unsolved. About the Role Own RL and post-training for large-scale multimodal models at a frontier lab, from building... 
    Training
    Internship
    Relocation package
    Shift work

    RecruitSeq

    Seattle, WA
    3 days ago
  • £87.68k - £165.38k per year

     ...Our partner is looking for a Senior AI Research Scientist (Model-based RL) based in Spain. This role offers the opportunity to shape the future of...  ...capable of generalizing across different systems, including training approaches such as pretraining, curriculum learning,... 
    Training
    Remote work
    Flexible hours

    Jobgether

    San Antonio, TX
    10 hours ago
  • $126k - $423k

     ...team We are looking for multiple passionate Research Scientists to join the Research Group at Applied Intuition...  ...Conduct research on reinforcement learning (RL) related topics including large‑scale self‑play RL, VLA post‑training, large‑scale closed‑loop RL based on neural... 
    Training
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Immediate start
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    3 days ago
  •  ...About the role We’re looking for a top‑tier Research Scientist to join our tech team. Your core...  ...ranking systems for AI agents Prototype, train, and evaluate new models for factual search...  ...progress of our search engine Design a RL framework to integrate AI preferences into... 
    Training
    Remote work
    Flexible hours

    Linkup Inc

    San Francisco, CA
    1 day ago
  •  ...Role We’re looking for Applied Scientists to join Wayve Labs and help...  ...Wayve, we are a high‑conviction research team with the strategic...  ...transformers, MoE, large‑scale training). Generative world modeling (...  ...Reinforcement learning (e.g., offline RL, RLHF, reward modeling).... 
    Training
    Full time
    Work at office
    Work from home
    Visa sponsorship
    Relocation package
    Flexible hours

    Icehouseventures

    Sunnyvale, CA
    3 days ago
  • $200k - $250k

     ...enterprise workflows ~Post-train LLM agents using RLHF, DPO,...  ...engineering standards ~Mentor researchers and engineers; drive...  ...equivalent) ~5+ years hands-on RL — environment design, reward...  ...Subject: Senior Staff Research Scientist – RL Company Description... 
    Training
    Full time

    Centific

    Remote
    1 day ago
  • $66k - $94k

    Associate Research Scientist – Evidence Generation (Primary Data Collection) Are you looking to work with a dynamic team of industry leading...  ...including but not limited to: skill sets, experience and training, licensure and certifications, and other business and organizational... 
    Training
    Full time
    Internship
    Local area
    Remote work
    Worldwide

    Precision AQ

    United States
    3 days ago
  • $225k - $300k

     ...Research Scientist About Latent Health Healthcare today is only truly personalized for two groups: those with wealth and access,...  ...work on: Verifiable reinforcement learning at scale Mid-training and post-training of foundation models Novel objectives... 
    Training
    Full time
    Work at office
    Immediate start

    Latent

    Remote
    9 hours ago
  • $168k - $264.5k

    Role Description We're looking for a Research Scientist who shares our vision for transforming healthcare. This is a unique opportunity to...  ...clinical applications. ~5 plus years of experience building and training machine learning models at scale, along with the compute... 
    Training
    Full time
    Shift work

    NVIDIA

    Remote
    14 hours ago
  • $175k - $225k

    Role Description Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation... 
    Training
    Full time

    Innodata Inc.

    Remote
    6 days ago
  • $100k - $150k

    Role Description We are seeking an AI Research Engineer to bridge cutting-edge applied research and production engineering, designing and...  ...ML frameworks such as PyTorch or JAX. ~Hands-on experience training, fine-tuning, and evaluating deep learning models at non-... 
    Training
    Full time
    Local area
    Immediate start

    Bright Vision Technologies

    Remote
    9 hours ago
  • $133.19k - $190.28k

    Role Description We are seeking Research Scientists (across all levels of seniority) to join our Artist-First AI Music Lab. Our team pioneers...  ...ML-based audio processing and signal processing. ~Post-Training: Research in post-training techniques for music generation,... 
    Training
    Full time
    Flexible hours

    Spotify

    Remote
    21 days ago
  • $126.89k - $166.13k

    Role Description We are looking for a Senior Microwave Research Scientist. As a Senior Microwave Research Scientist, you’ll be part of a cross...  ...field with 5+ years of experience ~Experience hiring and training new teams for research and development work ~Knowledge of... 
    Training
    Full time
    Work at office

    IonQ

    Remote
    9 hours ago
  • $86.01 - $147.94 per hour

     ...candidate will be dependent on a variety of factors including, but not limited to, the candidate’s experience, qualifications, location, training, licenses, shifts worked and compensation model. Carle Health offers a comprehensive benefits package for team members and... 
    Training
    Full time
    Local area
    Remote work
    Relocation package
    Shift work

    Carle Health

    Peoria, IL
    5 days ago
  • $40 per hour

     ...DataAnnotation is seeking a Research Scientist (Biology) to contribute to AI model training and evaluation. In this role, you will assess AI chatbots with complex biology questions and evaluate their outputs for quality and performance. The ideal candidate will have expertise... 
    Training
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Wausau, WI
    1 day ago
  • $40 per hour

    A leading AI training company is seeking a Biology Research Scientist for a remote role in the United States. This position involves training AI models by measuring their progress, evaluating their logic, and solving complex biology problems. The ideal candidate should... 
    Training
    Hourly pay
    Remote work

    DataAnnotation

    Topeka, KS
    1 day ago
  •  ...A leading AI healthcare organization in the United States seeks a Research Scientist Intern to advance their AI models. The role involves architecting algorithms, training models on high-performance clusters, and collaborating with healthcare professionals. Ideal candidates... 
    Training
    Internship
    Remote work

    White Rabbit

    New York, NY
    1 day ago
  •  ...solve real‑world business problems. We’re training and deploying frontier models for...  ...for our customers. Cohere is a team of researchers, engineers, designers, and more, who are...  ...measure LLM progress. As a Senior Research Scientist, Model Evaluation, you will: Create... 
    Training
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    3 days ago
  • $30 per hour

    A leading AI training firm is seeking an Applied Physics Research Scientist to evaluate AI chatbots on physics-related challenges. This is a remote position ideal for experts with a strong grasp of classical mechanics and related fields. Applicants must possess fluency... 
    Training
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Bismarck, ND
    1 day ago
  • A leading AI development firm is seeking a Mathematician (PhD) to join their remote AI training project. In this role, you will evaluate AI-generated mathematical responses, ensuring accuracy and clarity. To qualify, you must have a PhD in Mathematics/Statistics, significant... 
    Training
    Weekly pay
    Remote work
    Flexible hours

    SME Careers

    Cambridge, MA
    1 day ago
  • $40 per hour

    A leading AI training company is looking for a Math Expert to enhance AI models by solving complex mathematical problems and evaluating their outputs. This REMOTE position requires fluency in English and expertise in arithmetic, algebra, geometry, calculus, and statistics... 
    Training
    Hourly pay
    Contract work
    Remote work

    DataAnnotation

    Washington DC
    1 day ago
  • $40 per hour

    A leading AI training company is seeking a Research Scientist (Biology) to evaluate and train AI chatbots by measuring their logic and solving complex biological questions. This role requires an expert understanding of biology and related fields. Candidates should be fluent... 
    Training
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Santa Fe, NM
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Scientist - RL Training. Be the first to apply!