Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Scientist - Post-training / RL

Epsilon Labs, Inc.

About Us We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting‑edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product‑market fit with a substantial customer pipeline already in place. Role Overview We're seeking a Research Scientist with deep expertise in post‑training and reinforcement learning to join our ML Research team . You'll be at the forefront of developing and deploying state‑of‑the‑art multimodal models for clinical use in radiology settings. This role owns every stage after pretraining: supervised fine‑tuning, reward modeling, reinforcement learning against verifiable and learned reward signals, reasoning and tool‑use training, and inference‑time strategy. You'll work with one of the largest and most diverse medical imaging datasets in the industry, advancing the state‑of‑the‑art in grounded report generation, reward design, and inference‑time reasoning while maintaining the clinical rigor required for healthcare deployment. Key Responsibilities Design reinforcement learning with verifiable rewards for report generation, including clinical label and entity‑relation matching, grounding IoU, measurement accuracy, and reporting schema compliance. Extend reinforcement learning to unverifiable and noisy objectives such as report quality and clinical usefulness, using learned reward models and radiologist feedback pipelines ( RLHF ) built on expert preferences and report edits. Run GRPO‑family algorithms with complex multi‑reward objectives , tuning reward composition and diagnosing reward hacking, entropy collapse, and diversity loss. Train explicit reward models , including multimodal reward models conditioned on the image, with both outcome and process supervision. Train chain‑of‑thought reasoning over image regions, including evidence localization and verification loops that keep reasoning grounded in the image rather than in language priors. Train multimodal tool use — windowing, zoom and crop, detector and segmentation calls, prior study retrieval — with credit assignment across multi‑turn trajectories. Develop inference‑time methods including best‑of‑N sampling against reward models and grounding‑aware decoding, and distill the resulting gains back into the policy. Tune output stylization to institutional reporting conventions, keeping style rewards separated from clinical content rewards. Stay current with cutting‑edge research in reinforcement learning, reward modeling, and multimodal post‑training. Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for post‑training medical VLMs at scale. Qualifications 6+ years of academia/industry experience in reinforcement learning, post‑training, or multimodal machine learning Deep expertise in post‑training large language or vision‑language models (e.g., Qwen‑VL, InternVL, LLaVA, or similar architectures) Strong foundation in modern post‑training and reinforcement learning techniques including: Group‑relative policy optimization and its successors (GRPO, DAPO, GSPO, CISPO) with multi‑reward objectives Reinforcement learning with verifiable rewards, and with noisy, sparse, or learned reward signals Reward model training: pairwise and generative reward models, outcome and process supervision Preference optimization methods (DPO, IPO, ORPO, KTO) and RLHF Inference‑time compute scaling, including best‑of‑N sampling and verifier‑guided decoding Practical experience diagnosing and mitigating reward hacking and reward over‑optimization Track record of implementing complex models from research papers and adapting them to new domains Proficiency in PyTorch or JAX, with experience training large models on multi‑GPU/distributed systems Experience with reinforcement learning infrastructure at scale, including rollout generation (vLLM, SGLang) and frameworks such as verl, TRL, or OpenRLHF Experience with autoregressive language modeling and instruction tuning Strong software engineering skills and ability to write production‑quality code Preferred Qualifications Publications at top‑tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, MICCAI) Hands‑on experience with medical imaging applications, particularly radiology report generation Experience with agentic or multi‑turn reinforcement learning, including credit assignment over tool‑use trajectories Experience with grounded generation tasks (visual grounding, referring expression comprehension) Knowledge of evaluation methodologies for long‑form generation, including factuality assessment and hallucination detection Experience mitigating catastrophic forgetting of supervised capabilities during reinforcement learning Familiarity with clinical NLP and medical knowledge representation Experience with model interpretability, explainability, and uncertainty quantification in safety‑critical applications #J-18808-Ljbffr Epsilon Labs, Inc.

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Research Scientist - Post-training / RL in San Francisco, CA vacancy
  • $204k - $259k

     ...Waymo AI Foundations Team Scientist Waymo is an autonomous driving...  ...foster collaborations with other research teams in Alphabet. AI...  ...Waymo's Foundation World Model post-training and evaluation Research and develop cutting edge RL and Distillation techniques for... 
    Training
    Temporary work

    Latent Logic

    San Francisco, CA
    5 days ago
  • Traverse is a research data lab building reinforcement learning environments for frontier AI labs. As a Research Scientist, you will design and build RL environments that teach models to do work that...  ...knowledge, turning that into training signals that actually work. We... 
    Training

    Traverse

    San Francisco, CA
    3 days ago
  •  ...invites applications for a Senior Scientist role in the AI Foundations team. In...  ...Scientist and collaborate across research and engineering to advance RL and foundation models for autonomous...  ...will contribute to world model post‑training, publish high‑impact work, and help... 
    Training

    Waymo

    San Francisco, CA
    2 days ago
  • cursor is hiring a Research Scientist to drive research in reinforcement learning at their New York...  ...should have a deep background in RL, be excellent programmers, and be comfortable...  ...RL, improving data quality for model training, and executing realtime RL for coding agents... 
    Training
    Work at office

    cursor

    San Francisco, CA
    2 days ago
  • Epsilon Labs, Inc. is seeking a Research Scientist with deep expertise in post-training and reinforcement learning to advance multimodal models for clinical radiology use. You will own stages after pretraining, including supervised fine-tuning, reward modeling, and inference... 
    Training

    Epsilon Labs, Inc.

    San Francisco, CA
    3 days ago
  • Xterraai, based in San Francisco, is seeking research scientists to develop innovative AI systems that reason about complex scientific problems...  ...on reinforcement learning, evaluation, and building robust training infrastructure. Ideal candidates possess strong machine... 
    Training
    Remote work

    Xterraai

    San Francisco, CA
    2 days ago
  • $216k - $270k

    Scale Labs, Research Scientist — Safety Post TrainingAs the leading data and evaluation partner for frontier...  ...Scientist working on Safety Post-Training you will develop and apply post-training...  ...advance.Experience with post-training and RL techniques such as RLHF, DPO, GRPO,... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    3 days ago
  • $117.2k - $313.7k

     ...ExperienceSalesforce AI Research is a global leader in...  ...breakthrough in multimodal AI; trained state-of-the-art large...  ...Research Scientists who want to build, ship...  ...reinforcement learning (RL), reasoning and planning...  ...processing.Core Modeling and Post-Training: Pre-training... 
    Training
    Full time
    Worldwide

    Salesforce

    San Francisco, CA
    2 days ago
  •  ...Research Scientist Engineering · Full-time · San Francisco; New York Our mission is to automate coding. The first step...  ...Cursor is building the future of coding. We train frontier coding agents and scale RL on real user data to make them increasingly effective... 
    Training
    Full time

    Anysphere

    San Francisco, CA
    1 day ago
  •  ...Pantograph is training general models that start by watching internet-scale video and...  ..., durable robots. We're looking for research scientists who want to scale simple methods across...  ...supervised, goal-conditioned, or unsupervised RL Robotics models, especially those... 
    Training

    Pantograph

    San Francisco, CA
    2 days ago
  •  ...AfterQuery AfterQuery is an applied research lab curating data solutions for foundation...  ...data works. You will design and run training experiments that isolate the impact of...  ...model behavior. This includes SFT and RL-based post-training, where you'll measure how different... 
    Training
    Local area
    Shift work

    AfterQuery

    San Francisco, CA
    3 days ago
  • $300k - $320k

     ...quickly growing group of committed researchers, engineers, policy experts, and...  ...re seeking an exceptional Research Scientist to join our Life Sciences team at...  ...capabilities on scientific tasks through post-training, evaluation design, and RL environment development. As a... 
    Training
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  • $290.4k - $363k

     ...intersection of cutting-edge research, large-scale engineering, and...  ...evaluation methodologies, and agent/RL infrastructure that will...  ...in large-language models, post-training, evaluation, and agentic/RL environments...  ..., and deployed.As a Research Scientist Manager, you will lead a... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $200k

     ...idler is a frontier data research lab. We build the evals...  ...use to measure and train their models. After...  ...role As a Research Scientist at idler, you'll own measuring...  ...labs — to design novel post-training recipes and...  ...years of experience doing RL in production at a... 
    Training
    Work at office
    Relocation package

    Idler

    San Francisco, CA
    8 days ago
  • Traverse is a research data lab building reinforcement learning...  ...has figured out how to train models on. We work directly...  ...About the Role As a Research Scientist, you will design and build RL environments that teach...  ...integrate environments into their post-training pipelines Your... 
    Training

    Traverse

    San Francisco, CA
    3 days ago
  • $250k - $400k

     ...on genuinely novel AI for Science research, combining frontier reasoning, post-training and reinforcement learning with a...  ...experience in LLM post-training, RL for reasoning, reasoning datasets...  ...or experienced Research Scientist. What matters most is hands-on ownership... 
    Training

    techire ai

    San Francisco, CA
    2 days ago
  •  ...Role Pretraining gives us a general model. Post-training makes it useful, controllable, safe, and...  ...the places in between. This is where research meets reality. You’ll be responsible for...  ...techniques such as imitation learning, RL, distillation, synthetic data, and curriculum... 
    Training

    Generalist

    San Francisco, CA
    5 days ago
  •  ...place. Role Overview We're seeking a Research Scientist with deep expertise in large-scale...  ...task mixtures, and the large-scale training runs that build grounded visual understanding...  ...required for healthcare deployment. Post-training and RL are owned by a partner role you'll... 
    Training

    Epsilon Labs, Inc.

    San Francisco, CA
    5 days ago
  • $166k - $225k

    As a Research Scientist on the GenAI Team at Databricks, you will be responsible for keeping up with...  ...with diverse backgrounds and technical training. And most importantly, you will love...  ...billions of parameters. Have solid ML and RL software engineering and scientific... 
    Training
    Work at office
    Local area

    Databricks, Inc.

    San Francisco, CA
    2 days ago
  •  ...via reinforcement learning: Designing and training reasoning systems using RLHF, RLAIF, and...  ...Contributing to alignment and oversight research - figuring out how to reliably supervise...  ...ideally applied to language models, but strong RL backgrounds from other domains (robotics,... 
    Training
    Full time
    Internship

    Xterraai

    San Francisco, CA
    2 days ago
  • Research San Francisco or New York · SF preferred Research Scientist You will define the research agenda for models that can launch...  ...You’ll work across reasoning, post-training, continual learning, evals,...  ...models in house: fine-tuning, SFT, RL, reward modeling, preference... 
    Training

    Economic Intelligence Co.

    San Francisco, CA
    20 hours ago
  • $120k - $250k

     ...behaviors in unstructured environments Research and implement state-of-the-art robot learning...  ...Optimize robot policies for distributed training at scale and real-time edge deployment...  ...manipulation tech stacks with imitation learning or RL-based methods Background in real‑time ML... 
    Training

    Rainfall Ventures

    San Francisco, CA
    5 days ago
  • $350k

     ...quickly growing group of committed researchers, engineers, policy experts,...  ...We're looking for a Research Scientist who has done hands-on...  ...models (pretraining, fine-tuning, RL, evals, or agents scaffolds)...  ...Strong candidates may also have Trained or RL'd frontier models hands... 
    Training
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic Limited

    San Francisco, CA
    2 days ago
  •  ...Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer...  ...open problems at the intersection of post‑training methodology and performant inference,...  ...areas (e.g., both interpretability and RL, or both systems and training... 
    Training
    Immediate start
    Flexible hours

    The Consensus

    San Francisco, CA
    5 days ago
  • $117.2k - $313.7k

    Salesforce AI Research is looking for outstanding AI Research Scientists and Research Engineers to discover new research problems...  ...agents, reinforcement learning (RL), reasoning and planning,...  ...processing. Core Modeling and Post‑Training: Machine learning methodology, pre... 
    Training

    salesforce.com, inc.

    San Francisco, CA
    1 day ago
  •  ...deployment. We are founded by leading scientists in robot reinforcement...  ...of the world's best robotics researchers are already building the...  ...Python and PyTorch, including training large-scale neural networks...  ...Language Model Fine-Tuning (SFT and RL-based) Transformer-based 3D... 
    Training

    Flexion Robotics

    San Francisco, CA
    13 days ago
  • $216k - $270k

    Research Scientist, AI Controls and Monitoring Scale Labs, Research Scientist - AI Controls and Monitoring As the leading data and evaluation...  ...oversight, interpretability, debate). Experience with post-training and RL techniques such as RLHF, DPO, GRPO, and similar... 
    Training
    Full time

    Scale

    San Francisco, CA
    3 days ago
  • $218.4k - $273k

     ...Environments (ACE) team, part of Scale’s Research organization, brings together...  ...on agent environments and RL reward signals, benchmarking...  ...range displayed on each job posting reflects the minimum and...  ...performance, and relevant education or training. Scale employees in eligible... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  •  ...Research Scientist, Research Engineer, AI Systems EngineerThe Recursive Self...  ...designing evaluations, and training models to develop missing capabilities...  ...harnesses, synthetic data, RL environments, and model...  ...that you believe this job posting is non-compliant, please... 
    Training

    OpenAI

    San Francisco, CA
    2 days ago
  •  ...pretraining, midtraining, reinforcement learning, post‑training, evaluations, harnessing, and deployment—and connect that research to the patients, clinicians, and real‑world...  ...in pretraining, reinforcement learning (RL) / post‑training, or evals; and researchers with... 
    Training
    Work at office
    Relocation package

    Triwill Group

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Scientist - Post-training / RL. Be the first to apply!