Research Scientist - Post-training / RL
Epsilon Labs, Inc.
About Us We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting‑edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product‑market fit with a substantial customer pipeline already in place. Role Overview We're seeking a Research Scientist with deep expertise in post‑training and reinforcement learning to join our ML Research team . You'll be at the forefront of developing and deploying state‑of‑the‑art multimodal models for clinical use in radiology settings. This role owns every stage after pretraining: supervised fine‑tuning, reward modeling, reinforcement learning against verifiable and learned reward signals, reasoning and tool‑use training, and inference‑time strategy. You'll work with one of the largest and most diverse medical imaging datasets in the industry, advancing the state‑of‑the‑art in grounded report generation, reward design, and inference‑time reasoning while maintaining the clinical rigor required for healthcare deployment. Key Responsibilities Design reinforcement learning with verifiable rewards for report generation, including clinical label and entity‑relation matching, grounding IoU, measurement accuracy, and reporting schema compliance. Extend reinforcement learning to unverifiable and noisy objectives such as report quality and clinical usefulness, using learned reward models and radiologist feedback pipelines ( RLHF ) built on expert preferences and report edits. Run GRPO‑family algorithms with complex multi‑reward objectives , tuning reward composition and diagnosing reward hacking, entropy collapse, and diversity loss. Train explicit reward models , including multimodal reward models conditioned on the image, with both outcome and process supervision. Train chain‑of‑thought reasoning over image regions, including evidence localization and verification loops that keep reasoning grounded in the image rather than in language priors. Train multimodal tool use — windowing, zoom and crop, detector and segmentation calls, prior study retrieval — with credit assignment across multi‑turn trajectories. Develop inference‑time methods including best‑of‑N sampling against reward models and grounding‑aware decoding, and distill the resulting gains back into the policy. Tune output stylization to institutional reporting conventions, keeping style rewards separated from clinical content rewards. Stay current with cutting‑edge research in reinforcement learning, reward modeling, and multimodal post‑training. Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for post‑training medical VLMs at scale. Qualifications 6+ years of academia/industry experience in reinforcement learning, post‑training, or multimodal machine learning Deep expertise in post‑training large language or vision‑language models (e.g., Qwen‑VL, InternVL, LLaVA, or similar architectures) Strong foundation in modern post‑training and reinforcement learning techniques including: Group‑relative policy optimization and its successors (GRPO, DAPO, GSPO, CISPO) with multi‑reward objectives Reinforcement learning with verifiable rewards, and with noisy, sparse, or learned reward signals Reward model training: pairwise and generative reward models, outcome and process supervision Preference optimization methods (DPO, IPO, ORPO, KTO) and RLHF Inference‑time compute scaling, including best‑of‑N sampling and verifier‑guided decoding Practical experience diagnosing and mitigating reward hacking and reward over‑optimization Track record of implementing complex models from research papers and adapting them to new domains Proficiency in PyTorch or JAX, with experience training large models on multi‑GPU/distributed systems Experience with reinforcement learning infrastructure at scale, including rollout generation (vLLM, SGLang) and frameworks such as verl, TRL, or OpenRLHF Experience with autoregressive language modeling and instruction tuning Strong software engineering skills and ability to write production‑quality code Preferred Qualifications Publications at top‑tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, MICCAI) Hands‑on experience with medical imaging applications, particularly radiology report generation Experience with agentic or multi‑turn reinforcement learning, including credit assignment over tool‑use trajectories Experience with grounded generation tasks (visual grounding, referring expression comprehension) Knowledge of evaluation methodologies for long‑form generation, including factuality assessment and hallucination detection Experience mitigating catastrophic forgetting of supervised capabilities during reinforcement learning Familiarity with clinical NLP and medical knowledge representation Experience with model interpretability, explainability, and uncertainty quantification in safety‑critical applications #J-18808-Ljbffr Epsilon Labs, Inc.
$204k - $259k
...Waymo AI Foundations Team Scientist Waymo is an autonomous driving... ...foster collaborations with other research teams in Alphabet. AI... ...Waymo's Foundation World Model post-training and evaluation Research and develop cutting edge RL and Distillation techniques for...TrainingTemporary work- Traverse is a research data lab building reinforcement learning environments for frontier AI labs. As a Research Scientist, you will design and build RL environments that teach models to do work that... ...knowledge, turning that into training signals that actually work. We...Training
- ...invites applications for a Senior Scientist role in the AI Foundations team. In... ...Scientist and collaborate across research and engineering to advance RL and foundation models for autonomous... ...will contribute to world model post‑training, publish high‑impact work, and help...Training
- cursor is hiring a Research Scientist to drive research in reinforcement learning at their New York... ...should have a deep background in RL, be excellent programmers, and be comfortable... ...RL, improving data quality for model training, and executing realtime RL for coding agents...TrainingWork at office
- Epsilon Labs, Inc. is seeking a Research Scientist with deep expertise in post-training and reinforcement learning to advance multimodal models for clinical radiology use. You will own stages after pretraining, including supervised fine-tuning, reward modeling, and inference...Training
- Xterraai, based in San Francisco, is seeking research scientists to develop innovative AI systems that reason about complex scientific problems... ...on reinforcement learning, evaluation, and building robust training infrastructure. Ideal candidates possess strong machine...TrainingRemote work
$216k - $270k
Scale Labs, Research Scientist — Safety Post TrainingAs the leading data and evaluation partner for frontier... ...Scientist working on Safety Post-Training you will develop and apply post-training... ...advance.Experience with post-training and RL techniques such as RLHF, DPO, GRPO,...TrainingFull time$117.2k - $313.7k
...ExperienceSalesforce AI Research is a global leader in... ...breakthrough in multimodal AI; trained state-of-the-art large... ...Research Scientists who want to build, ship... ...reinforcement learning (RL), reasoning and planning... ...processing.Core Modeling and Post-Training: Pre-training...TrainingFull timeWorldwide- ...Research Scientist Engineering · Full-time · San Francisco; New York Our mission is to automate coding. The first step... ...Cursor is building the future of coding. We train frontier coding agents and scale RL on real user data to make them increasingly effective...TrainingFull time
- ...Pantograph is training general models that start by watching internet-scale video and... ..., durable robots. We're looking for research scientists who want to scale simple methods across... ...supervised, goal-conditioned, or unsupervised RL Robotics models, especially those...Training
- ...AfterQuery AfterQuery is an applied research lab curating data solutions for foundation... ...data works. You will design and run training experiments that isolate the impact of... ...model behavior. This includes SFT and RL-based post-training, where you'll measure how different...TrainingLocal areaShift work
$300k - $320k
...quickly growing group of committed researchers, engineers, policy experts, and... ...re seeking an exceptional Research Scientist to join our Life Sciences team at... ...capabilities on scientific tasks through post-training, evaluation design, and RL environment development. As a...TrainingWork at officeVisa sponsorshipFlexible hours$290.4k - $363k
...intersection of cutting-edge research, large-scale engineering, and... ...evaluation methodologies, and agent/RL infrastructure that will... ...in large-language models, post-training, evaluation, and agentic/RL environments... ..., and deployed.As a Research Scientist Manager, you will lead a...TrainingFull time$200k
...idler is a frontier data research lab. We build the evals... ...use to measure and train their models. After... ...role As a Research Scientist at idler, you'll own measuring... ...labs — to design novel post-training recipes and... ...years of experience doing RL in production at a...TrainingWork at officeRelocation package- Traverse is a research data lab building reinforcement learning... ...has figured out how to train models on. We work directly... ...About the Role As a Research Scientist, you will design and build RL environments that teach... ...integrate environments into their post-training pipelines Your...Training
$250k - $400k
...on genuinely novel AI for Science research, combining frontier reasoning, post-training and reinforcement learning with a... ...experience in LLM post-training, RL for reasoning, reasoning datasets... ...or experienced Research Scientist. What matters most is hands-on ownership...Training- ...Role Pretraining gives us a general model. Post-training makes it useful, controllable, safe, and... ...the places in between. This is where research meets reality. You’ll be responsible for... ...techniques such as imitation learning, RL, distillation, synthetic data, and curriculum...Training
- ...place. Role Overview We're seeking a Research Scientist with deep expertise in large-scale... ...task mixtures, and the large-scale training runs that build grounded visual understanding... ...required for healthcare deployment. Post-training and RL are owned by a partner role you'll...Training
$166k - $225k
As a Research Scientist on the GenAI Team at Databricks, you will be responsible for keeping up with... ...with diverse backgrounds and technical training. And most importantly, you will love... ...billions of parameters. Have solid ML and RL software engineering and scientific...TrainingWork at officeLocal area- ...via reinforcement learning: Designing and training reasoning systems using RLHF, RLAIF, and... ...Contributing to alignment and oversight research - figuring out how to reliably supervise... ...ideally applied to language models, but strong RL backgrounds from other domains (robotics,...TrainingFull timeInternship
- Research San Francisco or New York · SF preferred Research Scientist You will define the research agenda for models that can launch... ...You’ll work across reasoning, post-training, continual learning, evals,... ...models in house: fine-tuning, SFT, RL, reward modeling, preference...Training
$120k - $250k
...behaviors in unstructured environments Research and implement state-of-the-art robot learning... ...Optimize robot policies for distributed training at scale and real-time edge deployment... ...manipulation tech stacks with imitation learning or RL-based methods Background in real‑time ML...Training$350k
...quickly growing group of committed researchers, engineers, policy experts,... ...We're looking for a Research Scientist who has done hands-on... ...models (pretraining, fine-tuning, RL, evals, or agents scaffolds)... ...Strong candidates may also have Trained or RL'd frontier models hands...TrainingWork at officeVisa sponsorshipFlexible hours- ...Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer... ...open problems at the intersection of post‑training methodology and performant inference,... ...areas (e.g., both interpretability and RL, or both systems and training...TrainingImmediate startFlexible hours
$117.2k - $313.7k
Salesforce AI Research is looking for outstanding AI Research Scientists and Research Engineers to discover new research problems... ...agents, reinforcement learning (RL), reasoning and planning,... ...processing. Core Modeling and Post‑Training: Machine learning methodology, pre...Training- ...deployment. We are founded by leading scientists in robot reinforcement... ...of the world's best robotics researchers are already building the... ...Python and PyTorch, including training large-scale neural networks... ...Language Model Fine-Tuning (SFT and RL-based) Transformer-based 3D...Training
$216k - $270k
Research Scientist, AI Controls and Monitoring Scale Labs, Research Scientist - AI Controls and Monitoring As the leading data and evaluation... ...oversight, interpretability, debate). Experience with post-training and RL techniques such as RLHF, DPO, GRPO, and similar...TrainingFull time$218.4k - $273k
...Environments (ACE) team, part of Scale’s Research organization, brings together... ...on agent environments and RL reward signals, benchmarking... ...range displayed on each job posting reflects the minimum and... ...performance, and relevant education or training. Scale employees in eligible...TrainingFull time- ...Research Scientist, Research Engineer, AI Systems EngineerThe Recursive Self... ...designing evaluations, and training models to develop missing capabilities... ...harnesses, synthetic data, RL environments, and model... ...that you believe this job posting is non-compliant, please...Training
- ...pretraining, midtraining, reinforcement learning, post‑training, evaluations, harnessing, and deployment—and connect that research to the patients, clinicians, and real‑world... ...in pretraining, reinforcement learning (RL) / post‑training, or evals; and researchers with...TrainingWork at officeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist - Post-training / RL. Be the first to apply!
- materials scientist San Francisco, CA
- scientist assay development San Francisco, CA
- entry level research scientist San Francisco, CA
- health scientist San Francisco, CA
- quality control scientist San Francisco, CA
- deep learning scientist San Francisco, CA
- research associate scientist San Francisco, CA
- application scientist San Francisco, CA
- scientist antibody discovery San Francisco, CA
- senior analytical scientist San Francisco, CA


