RL Research Scientist: Data-Driven LLM Alignment
Snorkel AI
Snorkel AI in San Francisco seeks a Research Scientist to work on reinforcement learning for training and aligning large language models. This foundational role focuses on data, reward signals, and training procedures that steer LLM behavior in reliable directions – a core capability that differentiates Snorkel’s data‑as‑a‑service offering. You will collaborate with Snorkel’s research, engineering, and delivery teams to advance RL data capabilities, translating research ideas into preference #J-18808-Ljbffr Snorkel AI
- Product Pulse is building post-training research to prove our data works. We’re hiring 2-3 Research Scientists to design and run training experiments, isolate the impact... ...development and partnerships. You’ll run SFT and RL post-training studies, quantify lift across reasoning...Data
$204k - $259k
...collaborations with other research teams in Alphabet. AI... ...report to a Principal Scientist. You will: Participate... ...develop cutting edge RL and Distillation... ...manner such as Data parallel, FSDP and other... ...on-policy learning or alignment with human preferences...DataFull timeTemporary workRemote work$150k - $250k
...building high-signal training data and evaluation infrastructure... ...on model behavior (SFT and RL post-training), and turn the... ...Work closely with the other Research Scientists to build shared experimental... ...and benchmarks. Tech stack: LLM post-training — SFT, RL. Requirements...DataFull timeVisa sponsorshipShift work$192.6k - $344.85k
...Research Lead / Principal Scientist & ManagerPost-Training · Alignment · Reinforcement LearningAutodesk... ...systems is still an open research problem.Autodesk touches... ...: rich structured data, long-horizon reasoning tasks... ...world-class talent across ML, RL, alignment, and foundation...DataFull timeFor contractorsRemote work$117.2k - $313.7k
...ExperienceSalesforce AI Research is looking for... ...outstanding AI Research Scientists and Research... ...Large language model (LLM)-powered agents,... ...reinforcement learning (RL), reasoning and... ...including techniques for alignment, safety, and... ...use your personal data and your rights, including...DataFull time- Distyl AI is seeking researchers to join our Post-Training team, translating foundation models... ...optimization, and continual adaptation to align models with Distyl’s workflows. You will... ...conducting experiments that demonstrate impact using real-world data. #J-18808-Ljbffr DistylData
$234.3k - $349k
...in their company's data and fueled by... ...About the roleAI research at WRITER isn't just... ...As an AI research scientist, you'll be at the... ...DPO, and emerging alignment techniques — with... ...improvement — including LLM-as-judge... ...immersive purpose-driven culture, career development...DataFull timeWork at officeLocal area- ...small, focused team of researchers and engineers... ...memory. As a Research Scientist, you'll design... ...methods). Synthetic data and self-study - understanding... ...retrieval. RL and online training... ...and problem-driven, collaborative to our... ...Familiarity with LLM training infrastructure...DataWork at office
- ...AfterQuery is an applied research lab curating data solutions for foundation model... ...behavior. This includes SFT and RL-based post-training, where... ..., generalization, and alignment. Working closely with partner... ...Strong familiarity with LLM training and evaluation methodologies...DataShift work
$300k - $320k
...Research Scientist Anthropic's mission is to create reliable,... ...evaluation design, and RL environment development... ...surface — tool scaffolding, data pipelines, eval... ...experience ~ Experience with LLM post-training: RLHF, RL... ...data curation, or eval-driven development ~ Direct...DataWork at officeVisa sponsorshipFlexible hours- Traverse is a research data lab building reinforcement learning environments for frontier AI labs. As a Research Scientist, you will design and build RL environments that teach models to do work that has historically required years of human expertise. You’ll work at the...Data
- Harnham is seeking a Research Scientist to tackle frontier problems in world models and RL at an on-site SF lab. This role emphasizes shipping research into real products with a data-driven approach and close collaboration with a stealth AI team. You will own end-to-end...DataRelocation package
- We build training data and evaluation infrastructure... ...our post-training research team and hiring 2-3 Research Scientists to work together on... ...including SFT and RL-based post-training... ..., and alignment. Working closely with... ...Strong familiarity with LLM training and...DataInternshipShift work
$250k
...AfterQuery builds the training data and evaluation... ...Working directly with research teams at top AI labs, you... ...downstream impact on model alignment and capability Partner... ...for/interned for any RL environment companies... ...instincts and familiarity with LLM training pipelines,...Data- ...job is to prove that our data works . You will design... ...This includes SFT and RL-based post-training,... ...capability, generalization, and alignment. Working closely with... ...external-facing research, blog posts, and technical... ...familiarity with LLM training and evaluation...DataShift work
- ...sleep into a personalized, data‑driven recovery experience. We are... ...Eight Sleep is looking for research scientists with a passion for using AI... ...ML, model optimization, NLP, LLM) Strong interest in applying... ...succeeds, so do you - perfectly aligning your achievements with the...DataFull timeImmediate startWorldwideNight shift
$218.4k - $273k
...has been the leading AI data foundry, helping fuel... ...team, part of Scale’s Research organization, brings together... ...agent environments and RL reward signals,... ...not only expertise in LLM agents and planning algorithms... ...concepts, and their alignment with our organizational...DataFull time- ...at the intersection of research, product, and partnerships... ...across modeling, data, systems, and evaluation... ...Features : You will use SFT, RL, distillation, and... ...metrics, building user‑aligned evaluations, and iterating... ..., or human‑feedback‑driven refinement. You are a product...Data
$357k
...helping enterprises unify data, applications,... ...of their roles . We are driven by innovation and are looking... ...Lead AI Research Scientist in Workato’s AI Research... ...automated post‑training, RL, evaluation techniques,... ...PyTorch or JAX and modern LLM frameworks. Proven track...DataWork at officeFlexible hours$216k - $270k
Scale Labs, Research Scientist — Frontier Risk EvaluationsAs the leading data and evaluation partner for frontier AI companies,... ...data science roles related to LLM technologies.Experience with red... ...of research concepts, and their alignment with our organizational culture...DataFull time- ...and helping enterprises unify data, applications, processes, and... ...of their roles. We are driven by innovation and looking for... ...looking for an exceptional AI Research Scientist to join our growing team. In... ...with PyTorch/JAX and modern LLM frameworks.Strong publication...DataRemote workFlexible hours
- ...Research Scientist ThirdLayer is solving one of the hardest problems in deploying... ...a continuous stream of data across your browser and work... ...develop new post-training and RL methods, and set the... ...touching post-training, RL, or alignment. A visible research track...Data
- ...auditing methods, and improving alignment. We were the first to show... ...used on the company’s largest RL training runs to detect misbehavior... ...monitoring, and frontier‑risk research. We care most about... ...Investigate how pre‑training, synthetic data, mid‑training, post‑training,...DataWork at officeRelocation package3 days per week
$200k - $350k
...Research Scientist At Snorkel, we believe meaningful AI doesn't start with... ...the model, it starts with the data. We're on a mission to help... ...utilizing techniques such as LLM as a Judge Prototype and build... ...implementation of research-driven innovations. Move fast and...DataLocal areaRemote work- ...Research Scientist / Machine Learning Scientist Location: SF Bay Area/Hybrid... ..., analyzing preference data, and disentangling performance... ...model reliability and alignment. Fluent in the full experimental... ..., and community-driven research What we offer:...DataFull timeRemote work
$259.2k - $324k
...has been the leading AI data foundry, helping fuel... ...team, part of Scale’s Research organization, brings together... ...agent environments and RL reward signals,... ...not only expertise in LLM agents and planning algorithms... ...concepts, and their alignment with our organizational...DataFull time$200k - $320k
...strengthen our expansive research team. We are looking... ...physicists, and computer scientists-who are passionate... ...systematic analysis of large data, and more. As a Senior... ...conversational AI, and LLM safety; turn them into... ...do. If you're hungry, driven, and ready to build something...DataWork at officeLocal areaRelocation- ...community for seminal research accomplishments at top... ...experienced AI Research Scientist to play a crucial role... ...projects, ensuring they align with business goals. Scale... ...AI to biomolecular data is a plus. Experience managing... ...at the frontier of AI-driven drug discovery that...Data
- ...We’re looking for a top‑tier Research Scientist to join our tech team. Your core... ...our search engine Design a RL framework to integrate AI... .... An AI can absorb oceans of data, while human attention barely... ...coming agentic web, Google’s ad‑driven model has no natural home. Search...DataRemote workFlexible hours
- ...Research Scientist Position 100% on-site Position Summary The CV Translational... ...to understand hypothesis-driven disease pathophysiology to... ...translational projects which are aligned with CV & FIB Translational... ...laboratory work, literature/data review, and/or computational...Data
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to RL Research Scientist: Data-Driven LLM Alignment. Be the first to apply!
- scientist ii San Francisco, CA
- scientist 1 San Francisco, CA
- image scientist San Francisco, CA
- downstream processing scientist San Francisco, CA
- qc scientist San Francisco, CA
- research scientist San Francisco, CA
- analytical scientist San Francisco, CA
- research scientist - biology San Francisco, CA
- genomics scientist San Francisco, CA
- research associate scientist San Francisco, CA

