Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

RL Research Scientist: Data-Driven LLM Alignment

Snorkel AI

Snorkel AI in San Francisco seeks a Research Scientist to work on reinforcement learning for training and aligning large language models. This foundational role focuses on data, reward signals, and training procedures that steer LLM behavior in reliable directions – a core capability that differentiates Snorkel’s data‑as‑a‑service offering. You will collaborate with Snorkel’s research, engineering, and delivery teams to advance RL data capabilities, translating research ideas into preference #J-18808-Ljbffr Snorkel AI

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the RL Research Scientist: Data-Driven LLM Alignment in San Francisco, CA vacancy
  • Product Pulse is building post-training research to prove our data works. We’re hiring 2-3 Research Scientists to design and run training experiments, isolate the impact...  ...development and partnerships. You’ll run SFT and RL post-training studies, quantify lift across reasoning... 
    Data

    Product Pulse

    San Francisco, CA
    1 day ago
  • $204k - $259k

     ...collaborations with other research teams in Alphabet. AI...  ...report to a Principal Scientist. You will: Participate...  ...develop cutting edge RL and Distillation...  ...manner such as Data parallel, FSDP and other...  ...on-policy learning or alignment with human preferences... 
    Data
    Full time
    Temporary work
    Remote work

    Latent Logic

    San Francisco, CA
    4 days ago
  • $150k - $250k

     ...building high-signal training data and evaluation infrastructure...  ...on model behavior (SFT and RL post-training), and turn the...  ...Work closely with the other Research Scientists to build shared experimental...  ...and benchmarks. Tech stack: LLM post-training — SFT, RL. Requirements... 
    Data
    Full time
    Visa sponsorship
    Shift work

    David Joseph & Company

    San Francisco, CA
    3 days ago
  • $192.6k - $344.85k

     ...Research Lead / Principal Scientist & ManagerPost-Training · Alignment · Reinforcement LearningAutodesk...  ...systems is still an open research problem.Autodesk touches...  ...: rich structured data, long-horizon reasoning tasks...  ...world-class talent across ML, RL, alignment, and foundation... 
    Data
    Full time
    For contractors
    Remote work

    Autodesk

    San Francisco, CA
    1 day ago
  • $117.2k - $313.7k

     ...ExperienceSalesforce AI Research is looking for...  ...outstanding AI Research Scientists and Research...  ...Large language model (LLM)-powered agents,...  ...reinforcement learning (RL), reasoning and...  ...including techniques for alignment, safety, and...  ...use your personal data and your rights, including... 
    Data
    Full time

    Salesforce

    San Francisco, CA
    1 day ago
  • Distyl AI is seeking researchers to join our Post-Training team, translating foundation models...  ...optimization, and continual adaptation to align models with Distyl’s workflows. You will...  ...conducting experiments that demonstrate impact using real-world data. #J-18808-Ljbffr Distyl
    Data

    Distyl

    San Francisco, CA
    4 days ago
  • $234.3k - $349k

     ...in their company's data and fueled by...  ...About the roleAI research at WRITER isn't just...  ...As an AI research scientist, you'll be at the...  ...DPO, and emerging alignment techniques — with...  ...improvement — including LLM-as-judge...  ...immersive purpose-driven culture, career development... 
    Data
    Full time
    Work at office
    Local area

    Writer

    San Francisco, CA
    3 days ago
  •  ...small, focused team of researchers and engineers...  ...memory. As a Research Scientist, you'll design...  ...methods). Synthetic data and self-study - understanding...  ...retrieval. RL and online training...  ...and problem-driven, collaborative to our...  ...Familiarity with LLM training infrastructure... 
    Data
    Work at office

    Engram

    San Francisco, CA
    1 day ago
  •  ...AfterQuery is an applied research lab curating data solutions for foundation model...  ...behavior. This includes SFT and RL-based post-training, where...  ..., generalization, and alignment. Working closely with partner...  ...Strong familiarity with LLM training and evaluation methodologies... 
    Data
    Shift work

    AfterQuery

    San Francisco, CA
    2 days ago
  • $300k - $320k

     ...Research Scientist Anthropic's mission is to create reliable,...  ...evaluation design, and RL environment development...  ...surface — tool scaffolding, data pipelines, eval...  ...experience ~ Experience with LLM post-training: RLHF, RL...  ...data curation, or eval-driven development ~ Direct... 
    Data
    Work at office
    Visa sponsorship
    Flexible hours

    Colorwave Inc

    San Francisco, CA
    2 days ago
  • Traverse is a research data lab building reinforcement learning environments for frontier AI labs. As a Research Scientist, you will design and build RL environments that teach models to do work that has historically required years of human expertise. You’ll work at the... 
    Data

    Traverse

    San Francisco, CA
    2 days ago
  • Harnham is seeking a Research Scientist to tackle frontier problems in world models and RL at an on-site SF lab. This role emphasizes shipping research into real products with a data-driven approach and close collaboration with a stealth AI team. You will own end-to-end... 
    Data
    Relocation package

    Harnham

    San Francisco, CA
    9 hours ago
  • We build training data and evaluation infrastructure...  ...our post-training research team and hiring 2-3 Research Scientists to work together on...  ...including SFT and RL-based post-training...  ..., and alignment. Working closely with...  ...Strong familiarity with LLM training and... 
    Data
    Internship
    Shift work

    Product Pulse

    San Francisco, CA
    1 day ago
  • $250k

     ...AfterQuery builds the training data and evaluation...  ...Working directly with research teams at top AI labs, you...  ...downstream impact on model alignment and capability Partner...  ...for/interned for any RL environment companies...  ...instincts and familiarity with LLM training pipelines,... 
    Data

    AfterQuery

    San Francisco, CA
    2 days ago
  •  ...job is to prove that our data works . You will design...  ...This includes SFT and RL-based post-training,...  ...capability, generalization, and alignment. Working closely with...  ...external-facing research, blog posts, and technical...  ...familiarity with LLM training and evaluation... 
    Data
    Shift work

    AfterQuery

    San Francisco, CA
    1 day ago
  •  ...sleep into a personalized, data‑driven recovery experience. We are...  ...Eight Sleep is looking for research scientists with a passion for using AI...  ...ML, model optimization, NLP, LLM) Strong interest in applying...  ...succeeds, so do you - perfectly aligning your achievements with the... 
    Data
    Full time
    Immediate start
    Worldwide
    Night shift

    eightsleep

    San Francisco, CA
    4 days ago
  • $218.4k - $273k

     ...has been the leading AI data foundry, helping fuel...  ...team, part of Scale’s Research organization, brings together...  ...agent environments and RL reward signals,...  ...not only expertise in LLM agents and planning algorithms...  ...concepts, and their alignment with our organizational... 
    Data
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  •  ...at the intersection of research, product, and partnerships...  ...across modeling, data, systems, and evaluation...  ...Features : You will use SFT, RL, distillation, and...  ...metrics, building user‑aligned evaluations, and iterating...  ..., or human‑feedback‑driven refinement. You are a product... 
    Data

    Luma AI

    San Francisco, CA
    3 days ago
  • $357k

     ...helping enterprises unify data, applications,...  ...of their roles . We are driven by innovation and are looking...  ...Lead AI Research Scientist in Workato’s AI Research...  ...automated post‑training, RL, evaluation techniques,...  ...PyTorch or JAX and modern LLM frameworks. Proven track... 
    Data
    Work at office
    Flexible hours

    Workato

    San Francisco, CA
    1 day ago
  • $216k - $270k

    Scale Labs, Research Scientist — Frontier Risk EvaluationsAs the leading data and evaluation partner for frontier AI companies,...  ...data science roles related to LLM technologies.Experience with red...  ...of research concepts, and their alignment with our organizational culture... 
    Data
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  •  ...and helping enterprises unify data, applications, processes, and...  ...of their roles. We are driven by innovation and looking for...  ...looking for an exceptional AI Research Scientist to join our growing team. In...  ...with PyTorch/JAX and modern LLM frameworks.Strong publication... 
    Data
    Remote work
    Flexible hours

    Workato

    San Francisco, CA
    2 days ago
  •  ...Research Scientist ThirdLayer is solving one of the hardest problems in deploying...  ...a continuous stream of data across your browser and work...  ...develop new post-training and RL methods, and set the...  ...touching post-training, RL, or alignment. A visible research track... 
    Data

    Thirdlayer (yc W25)

    San Francisco, CA
    4 days ago
  •  ...auditing methods, and improving alignment. We were the first to show...  ...used on the company’s largest RL training runs to detect misbehavior...  ...monitoring, and frontier‑risk research. We care most about...  ...Investigate how pre‑training, synthetic data, mid‑training, post‑training,... 
    Data
    Work at office
    Relocation package
    3 days per week

    United States Digital Space LLC

    San Francisco, CA
    1 day ago
  • $200k - $350k

     ...Research Scientist At Snorkel, we believe meaningful AI doesn't start with...  ...the model, it starts with the data. We're on a mission to help...  ...utilizing techniques such as LLM as a Judge Prototype and build...  ...implementation of research-driven innovations. Move fast and... 
    Data
    Local area
    Remote work

    Snorkel AI

    San Francisco, CA
    6 hours ago
  •  ...Research Scientist / Machine Learning Scientist Location: SF Bay Area/Hybrid...  ..., analyzing preference data, and disentangling performance...  ...model reliability and alignment. Fluent in the full experimental...  ..., and community-driven research What we offer:... 
    Data
    Full time
    Remote work

    Lead Allies Inc.

    San Francisco, CA
    3 days ago
  • $259.2k - $324k

     ...has been the leading AI data foundry, helping fuel...  ...team, part of Scale’s Research organization, brings together...  ...agent environments and RL reward signals,...  ...not only expertise in LLM agents and planning algorithms...  ...concepts, and their alignment with our organizational... 
    Data
    Full time

    Scale AI, Inc.

    San Francisco, CA
    1 day ago
  • $200k - $320k

     ...strengthen our expansive research team. We are looking...  ...physicists, and computer scientists-who are passionate...  ...systematic analysis of large data, and more. As a Senior...  ...conversational AI, and LLM safety; turn them into...  ...do. If you're hungry, driven, and ready to build something... 
    Data
    Work at office
    Local area
    Relocation

    EliseAI

    San Francisco, CA
    4 days ago
  •  ...community for seminal research accomplishments at top...  ...experienced AI Research Scientist to play a crucial role...  ...projects, ensuring they align with business goals. Scale...  ...AI to biomolecular data is a plus. Experience managing...  ...at the frontier of AI-driven drug discovery that... 
    Data

    Menlo Ventures

    San Francisco, CA
    2 days ago
  •  ...We’re looking for a top‑tier Research Scientist to join our tech team. Your core...  ...our search engine Design a RL framework to integrate AI...  .... An AI can absorb oceans of data, while human attention barely...  ...coming agentic web, Google’s ad‑driven model has no natural home. Search... 
    Data
    Remote work
    Flexible hours

    Linkup Inc

    San Francisco, CA
    1 day ago
  •  ...Research Scientist Position 100% on-site Position Summary The CV Translational...  ...to understand hypothesis-driven disease pathophysiology to...  ...translational projects which are aligned with CV & FIB Translational...  ...laboratory work, literature/data review, and/or computational... 
    Data

    Omni Inclusive

    Brisbane, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to RL Research Scientist: Data-Driven LLM Alignment. Be the first to apply!