Research Scientist - Post-training / RL
Epsilon Labs, Inc.
About Us We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting‑edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product‑market fit with a substantial customer pipeline already in place. Role Overview We're seeking a Research Scientist with deep expertise in post‑training and reinforcement learning to join our ML Research team . You'll be at the forefront of developing and deploying state‑of‑the‑art multimodal models for clinical use in radiology settings. This role owns every stage after pretraining: supervised fine‑tuning, reward modeling, reinforcement learning against verifiable and learned reward signals, reasoning and tool‑use training, and inference‑time strategy. You'll work with one of the largest and most diverse medical imaging datasets in the industry, advancing the state‑of‑the‑art in grounded report generation, reward design, and inference‑time reasoning while maintaining the clinical rigor required for healthcare deployment. Key Responsibilities Design reinforcement learning with verifiable rewards for report generation, including clinical label and entity‑relation matching, grounding IoU, measurement accuracy, and reporting schema compliance. Extend reinforcement learning to unverifiable and noisy objectives such as report quality and clinical usefulness, using learned reward models and radiologist feedback pipelines ( RLHF ) built on expert preferences and report edits. Run GRPO‑family algorithms with complex multi‑reward objectives , tuning reward composition and diagnosing reward hacking, entropy collapse, and diversity loss. Train explicit reward models , including multimodal reward models conditioned on the image, with both outcome and process supervision. Train chain‑of‑thought reasoning over image regions, including evidence localization and verification loops that keep reasoning grounded in the image rather than in language priors. Train multimodal tool use — windowing, zoom and crop, detector and segmentation calls, prior study retrieval — with credit assignment across multi‑turn trajectories. Develop inference‑time methods including best‑of‑N sampling against reward models and grounding‑aware decoding, and distill the resulting gains back into the policy. Tune output stylization to institutional reporting conventions, keeping style rewards separated from clinical content rewards. Stay current with cutting‑edge research in reinforcement learning, reward modeling, and multimodal post‑training. Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for post‑training medical VLMs at scale. Qualifications 6+ years of academia/industry experience in reinforcement learning, post‑training, or multimodal machine learning Deep expertise in post‑training large language or vision‑language models (e.g., Qwen‑VL, InternVL, LLaVA, or similar architectures) Strong foundation in modern post‑training and reinforcement learning techniques including: Group‑relative policy optimization and its successors (GRPO, DAPO, GSPO, CISPO) with multi‑reward objectives Reinforcement learning with verifiable rewards, and with noisy, sparse, or learned reward signals Reward model training: pairwise and generative reward models, outcome and process supervision Preference optimization methods (DPO, IPO, ORPO, KTO) and RLHF Inference‑time compute scaling, including best‑of‑N sampling and verifier‑guided decoding Practical experience diagnosing and mitigating reward hacking and reward over‑optimization Track record of implementing complex models from research papers and adapting them to new domains Proficiency in PyTorch or JAX, with experience training large models on multi‑GPU/distributed systems Experience with reinforcement learning infrastructure at scale, including rollout generation (vLLM, SGLang) and frameworks such as verl, TRL, or OpenRLHF Experience with autoregressive language modeling and instruction tuning Strong software engineering skills and ability to write production‑quality code Preferred Qualifications Publications at top‑tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, MICCAI) Hands‑on experience with medical imaging applications, particularly radiology report generation Experience with agentic or multi‑turn reinforcement learning, including credit assignment over tool‑use trajectories Experience with grounded generation tasks (visual grounding, referring expression comprehension) Knowledge of evaluation methodologies for long‑form generation, including factuality assessment and hallucination detection Experience mitigating catastrophic forgetting of supervised capabilities during reinforcement learning Familiarity with clinical NLP and medical knowledge representation Experience with model interpretability, explainability, and uncertainty quantification in safety‑critical applications #J-18808-Ljbffr Epsilon Labs, Inc.
$204k - $259k
...foster collaborations with other research teams in Alphabet. AI... ...you will report to a Principal Scientist. Responsibilities Participate... ...Waymo’s Foundation World Model post‑training and evaluation Research and develop cutting edge RL and Distillation techniques for...TrainingTemporary workRemote work- Product Pulse is building post-training research to prove our data works. We’re hiring 2-3 Research Scientists to design and run training experiments, isolate the impact of our... ...and partnerships. You’ll run SFT and RL post-training studies, quantify lift across reasoning...Training
- Epsilon Labs, Inc. is seeking a Research Scientist with deep expertise in post-training and reinforcement learning to advance multimodal models for clinical radiology use. You will own stages after pretraining, including supervised fine-tuning, reward modeling, and inference...Training
$150k - $250k
...behaviors in unstructured environments Research and implement state-of-the-art robot learning... ...Optimize robot policies for distributed training at scale and real-time edge deployment... ...manipulation tech stacks with imitation learning or RL‑based methods Background in real‑time ML...Training- Xterraai, based in San Francisco, is seeking research scientists to develop innovative AI systems that reason about complex scientific problems... ...on reinforcement learning, evaluation, and building robust training infrastructure. Ideal candidates possess strong machine...TrainingRemote work
$216k - $270k
Scale Labs, Research Scientist — Safety Post TrainingAs the leading data and evaluation partner for frontier... ...Scientist working on Safety Post-Training you will develop and apply post-training... ...advance.Experience with post-training and RL techniques such as RLHF, DPO, GRPO,...TrainingFull time$117.2k - $313.7k
...Salesforce.The ExperienceSalesforce AI Research is looking for outstanding AI Research Scientists and Research Engineers. Our team... ...agents, reinforcement learning (RL), reasoning and planning,... ...audio processing.Core Modeling and Post-Training: Machine learning methodology,...TrainingFull time$300k - $320k
...quickly growing group of committed researchers, engineers, policy experts, and... ...are seeking an exceptional Research Scientist to join our Life Sciences team at... ...capabilities on scientific tasks through post-training, evaluation design, and RL environment development. As a core...TrainingVisa sponsorship- ...About the role We’re looking for a top‑tier Research Scientist to join our tech team. Your core... ...ranking systems for AI agents Prototype, train, and evaluate new models for factual search... ...progress of our search engine Design a RL framework to integrate AI preferences into...TrainingRemote workFlexible hours
$290.4k - $363k
...intersection of cutting-edge research, large-scale engineering, and... ...evaluation methodologies, and agent/RL infrastructure that will... ...in large-language models, post-training, evaluation, and agentic/RL environments... ..., and deployed.As a Research Scientist Manager, you will lead a...TrainingFull time- ...deployment. We are founded by leading scientists in robot reinforcement... ...of the world's best robotics researchers are already building the... ...Python and PyTorch, including training large-scale neural networks... ...Language Model Fine-Tuning (SFT and RL-based) Transformer-based 3D...Training
- ...well-funded, early-stage AI lab building research systems aimed at one of the hardest open problems... .... What you'll be working on: Build and train reinforcement learning agents capable of... ...distributed training. Experience applying RL to mathematics, coding, simulation, or...Training
- ...programmers, using a combination of inventive research, design, and engineering. Our... ...crazy ideas, and shipping code. Research Scientist Cursor is building the future of coding. We train frontier coding agents and scale RL on real user data to make them increasingly...Training
- Traverse is a research data lab building reinforcement learning... ...has figured out how to train models on. We work directly... ...About the Role As a Research Scientist, you will design and build RL environments that teach... ...integrate environments into their post-training pipelines Your...Training
$250k - $400k
...on genuinely novel AI for Science research, combining frontier reasoning, post-training and reinforcement learning with a... ...experience in LLM post-training, RL for reasoning, reasoning datasets... ...or experienced Research Scientist. What matters most is hands‑on ownership...Training$300k
Research Scientist — Frontier World Models & RL A stealth, exceptionally well-backed applied AI lab is hiring Research... ...real users and a data flywheel to train against the moment beta launches. The... ...in an autoregressive setting. RL / Post-Training Advance RL post-training...TrainingRemote workVisa sponsorshipRelocation package- We build training data and evaluation infrastructure that frontier... ...'re a small, early team (post-Series A) where... ...building out our post-training research team and hiring 2-3 Research Scientists to work together on this... ..., including SFT and RL-based post-training, to measure...TrainingInternshipShift work
$250k
...AfterQuery AfterQuery builds the training data and evaluation... .... We are a small, early team (post Series A) where individual contributors... .... Working directly with research teams at top AI labs, you’ll experiment... ...worked for/interned for any RL environment companies in the...Training- ...retains scraps at best. We train models to study your world and... ...join a small, focused team of researchers and engineers working at the... ...learning and memory. As a Research Scientist, you’ll design experiments,... ...and agentic retrieval. RL and online training — exploring...TrainingWork at office
- ...via reinforcement learning: Designing and training reasoning systems using RLHF, RLAIF, and... ...Contributing to alignment and oversight research - figuring out how to reliably supervise... ...ideally applied to language models, but strong RL backgrounds from other domains (robotics,...TrainingFull timeInternship
- The role As a research scientist, you will design, implement, and optimize the large-scale training infrastructure that powers our frontier reinforcement... ...with researchers to make our RL stack reliable, fast, and... ...to bring frontier post-training capabilities into production...TrainingWork at officeVisa sponsorshipRelocation package
$120k - $250k
...behaviors in unstructured environments Research and implement state-of-the-art robot learning... ...Optimize robot policies for distributed training at scale and real-time edge deployment... ...manipulation tech stacks with imitation learning or RL-based methods Background in real‑time ML...Training$350k
...quickly growing group of committed researchers, engineers, policy experts,... ...We're looking for a Research Scientist who has done hands-on... ...models (pretraining, fine-tuning, RL, evals, or agents scaffolds)... ...Strong candidates may also have Trained or RL\'d frontier models...TrainingWork at officeVisa sponsorshipFlexible hours$218.4k - $273k
...Environments (ACE) team, part of Scale’s Research organization, brings together... ...on agent environments and RL reward signals, benchmarking... ...range displayed on each job posting reflects the minimum and... ...performance, and relevant education or training. Scale employees in eligible...TrainingFull time$216k - $270k
Scale Labs, Research Scientist - AI Controls and Monitoring As the leading data and evaluation partner for frontier AI companies, Scale... ...oversight, interpretability, debate). Experience with post‑training and RL techniques such as RLHF, DPO, GRPO, and similar approaches...TrainingFull time$192.6k - $344.85k
...3Research Lead / Principal Scientist & ManagerPost-Training · Alignment · Reinforcement... ...capable systems is still an open research problem.Autodesk touches... ...are evolving — and post-training is the layer that... ...world-class talent across ML, RL, alignment, and foundation...TrainingFull timeFor contractorsRemote work- We are Genmo, a research lab dedicated to building open, state-of-the-art models for video generation towards... ...overview:We are seeking an exceptional Research Scientist to join our team, focusing on alignment and post-training techniques for large-scale video generation...TrainingRelocation
$120.7k - $238.6k
The OpportunityAdobe Research is looking for research scientists in Generative AI to join a world-class research... ...Experience on large-scale generative model training· Experience of working with large-... ...in Colorado (as listed on the job posting), the application window will...TrainingFull timeTemporary workLocal areaWorldwide$290.4k - $363k
...intersection of cutting‑edge research, large‑scale engineering, and... ...evaluation methodologies, and agent/RL infrastructure that will... ...in large‑language models, post‑training, evaluation, and agentic/RL environments... ...and deployed. As a Research Scientist Manager, you will lead a...TrainingFull time$192.6k - $344.85k
## AI Research Manager/Scientist, Reinforcement LearningApplylocations: San Francisco... ...type: Full timeposted on: Posted 15 Days Agojob requisition id... ...Manager to lead our post-training and *model* alignment efforts... ...with strengths across *ML*, RL, evaluation, and human-centered...TrainingRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist - Post-training / RL. Be the first to apply!
- scientist ii San Francisco, CA
- scientist 1 San Francisco, CA
- image scientist San Francisco, CA
- downstream processing scientist San Francisco, CA
- qc scientist San Francisco, CA
- research scientist San Francisco, CA
- analytical scientist San Francisco, CA
- research scientist - biology San Francisco, CA
- genomics scientist San Francisco, CA
- research associate scientist San Francisco, CA


