Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Benchmark Lead Research Scientist

Vals AI

Vals AI in San Francisco is seeking exceptional researchers and research engineers to design and build the next generation of AI benchmarks. You will create high-impact, challenging evaluations that push the boundaries of what we can measure in foundation models. This role is perfect for someone with deep research expertise who wants to see their work directly influence how the world evaluates AI systems. We are building the standard for evaluating the ability of LLMs to perform real-world #J-18808-Ljbffr Vals AI

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the LLM Benchmark Lead Research Scientist in San Francisco, CA vacancy
  •  ...Inc. in San Francisco is seeking strong engineers to own the leaderboards that evaluate real-world tasks by LLMs. You will test and benchmark new models across domains like law, tax, coding, finance and more, analyzing error modes and collaborating with the communications... 
    Suggested

    Vals AI, Inc.

    San Francisco, CA
    5 days ago
  • $200k - $325k

     ...when Snorkel started as a research project in the Stanford...  ...to empower scientists, engineers, financial experts...  ...collaborate with partners and lead the development of the next frontier benchmarks and datasets. This is a...  ...Stay at the frontier of LLM evaluation research and... 
    Suggested
    Local area

    Snorkel AI

    San Francisco, CA
    1 day ago
  • Camb in San Francisco is searching for a Lead LLM R&D Researcher. This full-time position involves leading innovations in natural language processing and colloquial translation as well as collaborating with R&D teams. The ideal candidate will have a PhD or equivalent experience... 
    Suggested
    Full time

    Camb

    San Francisco, CA
    3 days ago
  • Enam, Inc. is seeking a senior ML researcher to lead exploration of new neural network foundations and architectures, with a strong emphasis on LLM applications. You will design experiments, validate novel hypotheses, and translate research ideas into scalable systems in... 
    Suggested

    Enam, Inc.

    San Francisco, CA
    3 days ago
  • $357k

     ...enhance life inside and outside of work. Responsibilities Lead AI Research Scientist in Workato’s AI Research Lab. This research‑first role focuses...  ...teams. Hands‑on experience with PyTorch or JAX and modern LLM frameworks. Proven track record of transitioning research... 
    Suggested
    Work at office
    Flexible hours

    Workato

    San Francisco, CA
    2 days ago
  • Bay Area Full time Lead LLM R&D, innovate in zero-shot performance and colloquial translation, collaborate, and publish research in top conferences. Join Camb.ai's pioneering AI team...  ...just innovating; we're setting SOTA benchmarks. Backed by elite global VCs, partnering... 
    Full time

    Camb

    San Francisco, CA
    3 days ago
  • Snorkel AI in San Francisco seeks a Research Scientist to work on reinforcement learning for training and aligning large language models. This...  ...on data, reward signals, and training procedures that steer LLM behavior in reliable directions - a core capability that differentiates... 

    Snorkel AI

    San Francisco, CA
    2 days ago
  • Cincinnatus LLC is recruiting researchers to design and author multi-step evaluation tasks for frontier AI benchmarks. The role emphasizes translating scientific method into...  ...full-time W-2 engagement, integrated with leading AI lab teams through Cincinnatus’ extended... 
    Remote job
    Full time
    Part time

    Obsidian

    San Francisco, CA
    2 days ago
  • Infinity is building ResearchBench, a high-quality research-replication benchmark. A decomposition-and-testing agent converts papers into roughly 400 leaves across datasets, compute, and function implementation. A separate harness then runs backends like Claude Code, Codex... 

    Infinity Artificial Intelligence Institute

    San Francisco, CA
    1 day ago
  • $150k - $250k

     ...goods, and global social organizations.We research and deploy technologies that power AI-...  ...of patient journeys.Distyl is backed by leading investors including Lightspeed Venture Partners...  ...to drive incremental improvements on benchmarks or optimize an existing process but... 
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    2 days ago
  • $216k - $270k

    Scale Labs, Research Scientist — Frontier Risk EvaluationsAs the leading data and evaluation partner for frontier AI companies, Scale...  ...Experience in crafting evaluations and benchmarks, or a background in data science roles related to LLM technologies.Experience with red-... 
    Full time

    Scale AI

    San Francisco, CA
    3 days ago
  • $262.5k - $299.6k

     ...Applied Researcher II (AI Foundations, LLM Core and Agentic AI) Overview At Capital One, we are creating trustworthy...  .... For years, Capital One has been leading the industry in using machine...  ...with a cross‑functional team of data scientists, software engineers, machine... 
    Full time
    Part time
    Local area
    Flexible hours

    Capital One

    San Francisco, CA
    1 day ago
  • A leading investment firm in San Francisco is seeking a Chief Scientist to advance state-of-the-art technologies in Machine Learning for cybersecurity. Ideal candidates...  ...background and experience with reputed research groups. While prior startup experience is a plus... 

    Greylock Partners

    San Francisco, CA
    2 days ago
  • Scale AI in New York seeks a Research Scientist Manager to lead a world-class team, define the research roadmap, and drive from prototyping to deployment. You will mentor researchers, collaborate with engineering and product, publish results, and influence strategic directions... 

    Scale AI

    San Francisco, CA
    3 days ago
  • $141.1k - $281.7k

     ...image, and videos, and creates an industry‑leading technical platform to improve creative...  ...marketers and designers. Responsibilities Research and implement ranking ML models using...  ...models Conduct innovative research on novel LLM architectures and training methods,... 

    Adobe

    San Francisco, CA
    2 days ago
  • Scale AI, Inc. seeks a Research Scientist Manager to lead a world-class team at the forefront of GenAI research, evaluation, and agent/RL infrastructure. You will define the research roadmap and drive execution from prototyping to deployment, balancing technical leadership... 

    Scale AI, Inc.

    San Francisco, CA
    5 days ago
  •  ...AI at scale in San Francisco. We are seeking a research-leaning engineer to own end-to-end inference research bets for LLM serving, including speculative decoding, quantization...  ...management. You will work with the research lead and Forward Deployed Engineers to deploy models... 

    Modal

    San Francisco, CA
    2 days ago
  •  ...workersResponsibilitiesWe are looking for an exceptional AI Research Scientist to join our growing team. In this role,...  ...‑agent collaboration.Prototype and benchmark models; present findings internally and...  ....Hands‑on with PyTorch/JAX and modern LLM frameworks.Strong publication record in... 
    Remote work
    Flexible hours

    Workato

    San Francisco, CA
    3 days ago
  • $234.3k - $349k

     ...is where the world's leading enterprises orchestrate...  ...AI. About the roleAI research at WRITER isn't just about...  .... As an AI research scientist, you'll be at the...  ...workflowsBuild novel evaluation benchmarks and methodologies that...  ...— including LLM-as-judge frameworks, synthetic... 
    Full time
    Work at office
    Local area

    Writer

    San Francisco, CA
    4 days ago
  • A biotech company focusing on AI in drug discovery seeks a highly skilled AI Research Scientist to drive the development of next-generation AI models used in drug discovery workflows. The ideal candidate possesses a Master's or PhD in a relevant field along with at least... 

    Menlo Ventures

    San Francisco, CA
    4 days ago
  • WRITER seeks a staff AI research scientist to advance a high-impact research program on large language models, agentic reasoning, and system-level capabilities for enterprise-scale AI. This role blends independent research with direct product impact. Based in San Francisco... 

    Neura Market

    San Francisco, CA
    4 days ago
  • $300k - $320k

     ...Research Scientist Anthropic's mission is to create reliable, interpretable...  ...model training objectives, benchmarks, and agentic workflows. You'...  ...~ Experience with LLM post-training: RLHF, RL from...  ...pipelines — target identification, lead optimization, ADMET modeling... 
    Work at office
    Visa sponsorship
    Flexible hours

    Colorwave Inc

    San Francisco, CA
    3 days ago
  •  ...a technical leader to head its robotics coverage, developing benchmarking methodologies, leaderboards, and datasets that set industry standards...  ...concepts into actionable insights. This role blends product, research, and client-facing work with broad scope to shape our robotics... 

    Artificial Analysis

    San Francisco, CA
    4 days ago
  • Artificial Analysis, Inc. in San Francisco is seeking a Member of Technical Staff to lead its new robotics vertical. This role entails defining benchmarking methodologies and producing leaderboards to evaluate robotics capabilities. Ideal candidates will possess a strong... 

    Artificial Analysis, Inc.

    San Francisco, CA
    2 days ago
  • The Token Company is seeking an ML Researcher to own a slice of open problems in applied AI, focusing on what information in an LLM context matters and how to represent it efficiently. This high-autonomy role requires running many experiments, reproducing papers, and shipping... 

    davidjoseph-co

    San Francisco, CA
    1 day ago
  • A pioneering research lab is seeking a Senior Software Engineer focused on building next-generation conversational AI interfaces. The ideal...  ...will have deep expertise in Python and experience with modern LLM frameworks. This position involves ownership of end-to-end development... 

    Jack & Jill

    San Francisco, CA
    2 days ago
  • $125k - $225k

     ...FutureSearch is looking for exceptional Research Scientists to evaluate and improve state-of-the-art forecasting and agentic LLM web research. We are an elite team of engineers...  ...our ICLR Workshop paper in Jan 2026, and benchmarks like Deep Research Bench and Bench to the... 
    Remote work
    Flexible hours

    Future Research Corp

    San Francisco, CA
    2 days ago
  • Carnaby Fox is seeking a Member of Technical Staff (AI Research) in San Francisco to help shape the research direction for frontier AI...  ..., evaluate LLMs, and improve data quality for high-stakes AI benchmarks. The role emphasizes independent ownership, pioneering benchmark... 

    Carnaby Fox

    San Francisco, CA
    5 days ago
  •  ...also happen to be the world's leading generative AI studio—we're...  ...We’ve developed an in-house LLM storytelling system that blends...  ...Project Astra, and top-tier AI researchers. As an early member of...  ...saving techniques, scalable benchmarking and checkpoint selection, are... 
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    18 days ago
  • $172.5k - $260.1k

     ...through AI, driving innovation, and keeping Salesforce's core values at the heart of it all.Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of... 
    Full time

    Salesforce

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Benchmark Lead Research Scientist. Be the first to apply!