Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Benchmark Lead Research Scientist

Vals AI

Vals AI in San Francisco is seeking exceptional researchers and research engineers to design and build the next generation of AI benchmarks. You will create high-impact, challenging evaluations that push the boundaries of what we can measure in foundation models. This role is perfect for someone with deep research expertise who wants to see their work directly influence how the world evaluates AI systems. We are building the standard for evaluating the ability of LLMs to perform real-world #J-18808-Ljbffr Vals AI

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the LLM Benchmark Lead Research Scientist in San Francisco, CA vacancy
  • Vibehackers in San Francisco is seeking a Senior Staff Researcher to design and implement novel benchmarks that evaluate real-world LLM capabilities. You will lead research and engineering efforts, publish results, and collaborate with model labs and enterprise partners... 
    Suggested
    Relocation package

    Vibehackers

    San Francisco, CA
    2 days ago
  • A leading AI CRM company is seeking an AI Research Scientist/Research Engineer in San Francisco. You will focus on advancing AI techniques, developing models, and solving real-world enterprise problems. Ideal candidates hold a Ph.D. or have strong AI product experience... 
    Suggested

    Niebles

    San Francisco, CA
    16 hours ago
  •  ...Inc. in San Francisco is seeking strong engineers to own the leaderboards that evaluate real-world tasks by LLMs. You will test and benchmark new models across domains like law, tax, coding, finance and more, analyzing error modes and collaborating with the communications... 
    Suggested

    Vals AI, Inc.

    San Francisco, CA
    1 day ago
  • $200k - $325k

     ...when Snorkel started as a research project in the Stanford...  ...to empower scientists, engineers, financial experts...  ...collaborate with partners and lead the development of the next frontier benchmarks and datasets. This is a...  ...Stay at the frontier of LLM evaluation research and... 
    Suggested
    Local area

    Snorkel AI

    San Francisco, CA
    3 days ago
  • Innovaccer in San Francisco seeks a Principal AI Researcher to lead production-grade AI systems, solving complex problems at scale. You will...  ...design, develop, and deploy AI-powered applications, including LLM-based solutions, AI agents, RAG, and automation workflows.... 
    Suggested

    Socket.dev

    San Francisco, CA
    4 days ago
  • Camb in San Francisco is searching for a Lead LLM R&D Researcher. This full-time position involves leading innovations in natural language processing and colloquial translation as well as collaborating with R&D teams. The ideal candidate will have a PhD or equivalent experience... 
    Full time

    Camb

    San Francisco, CA
    4 days ago
  • Enam, Inc. is seeking a senior ML researcher to lead exploration of new neural network foundations and architectures, with a strong emphasis on LLM applications. You will design experiments, validate novel hypotheses, and translate research ideas into scalable systems in... 

    Enam, Inc.

    San Francisco, CA
    4 days ago
  • Bay Area Full time Lead LLM R&D, innovate in zero-shot performance and colloquial translation, collaborate, and publish research in top conferences. Join Camb.ai's pioneering AI team...  ...just innovating; we're setting SOTA benchmarks. Backed by elite global VCs, partnering... 
    Full time

    Camb

    San Francisco, CA
    4 days ago
  • Mercor is seeking a researcher to join its GenAI team to translate the scientific method into complex, multi-step tasks and to design studies...  ...with researchers to ensure consistent, high-quality benchmark design and evaluation of frontier models. #J-18808-Ljbffr Mercor
    Remote job

    Mercor

    San Francisco, CA
    3 days ago
  • Cincinnatus LLC is recruiting researchers to design and author multi-step evaluation tasks for frontier AI benchmarks. The role emphasizes translating scientific method into...  ...full-time W-2 engagement, integrated with leading AI lab teams through Cincinnatus’ extended... 
    Remote job
    Full time
    Part time

    Obsidian

    San Francisco, CA
    3 days ago
  • Scale AI, Inc. seeks Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation. The role focuses on building benchmarks and diagnosing model failure modes in text and multimodal modalities within the GenAI... 

    Scale AI, Inc.

    San Francisco, CA
    3 days ago
  • $150k - $250k

     ...goods, and global social organizations.We research and deploy technologies that power AI-...  ...of patient journeys.Distyl is backed by leading investors including Lightspeed Venture Partners...  ...to drive incremental improvements on benchmarks or optimize an existing process but... 
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    3 days ago
  •  ...standard for how AI is measured and to work with leading AI labs and enterprises. You’ll develop benchmarking methodologies, manage relationships with top AI labs...  ...drive product direction. The role blends product, research, and client-facing work for those who enjoy... 

    Artificial Analysis

    San Francisco, CA
    1 day ago
  • Pocket FM is seeking a Research Scientist to advance generative AI for long-form storytelling. You will work at the intersection of LLMs, multimodal AI, agentic systems, and scalable machine learning, translating research into production systems that delight millions of... 
    Worldwide

    Pocket FM

    San Francisco, CA
    3 days ago
  • $141.1k - $281.7k

     ...image, and videos, and creates an industry‑leading technical platform to improve creative...  ...marketers and designers. Responsibilities Research and implement ranking ML models using...  ...models Conduct innovative research on novel LLM architectures and training methods,... 

    Adobe

    San Francisco, CA
    3 days ago
  • A leading investment firm in San Francisco is seeking a Chief Scientist to advance state-of-the-art technologies in Machine Learning for cybersecurity. Ideal candidates...  ...background and experience with reputed research groups. While prior startup experience is a plus... 

    Greylock Partners

    San Francisco, CA
    3 days ago
  • $180.6k - $225.75k

     ...works with the industry's leading AI labs to provide high quality...  ...progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF,...  ...and will focus on building benchmarks and diagnosing model failure... 
    Full time

    Scale AI

    San Francisco, CA
    3 days ago
  • $216k - $270k

    Scale Labs, Research Scientist — Frontier Risk EvaluationsAs the leading data and evaluation partner for frontier AI companies, Scale...  ...Experience in crafting evaluations and benchmarks, or a background in data science roles related to LLM technologies.Experience with red-... 
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  • $262.5k - $299.6k

    Applied Researcher II (AI Foundations, LLM Core and Agentic AI) Overview: At Capital One, we are creating...  ...good. For years, Capital One has been leading the industry in using machine...  ...with a cross-functional team of data scientists, software engineers, machine learning... 
    Full time
    Part time
    Local area
    Flexible hours

    Capital One

    San Francisco, CA
    4 days ago
  • A biotech company focusing on AI in drug discovery seeks a highly skilled AI Research Scientist to drive the development of next-generation AI models used in drug discovery workflows. The ideal candidate possesses a Master's or PhD in a relevant field along with at least... 

    Menlo Ventures

    San Francisco, CA
    16 hours ago
  • A pioneering research lab is seeking a Senior Software Engineer focused on building next-generation conversational AI interfaces. The ideal...  ...will have deep expertise in Python and experience with modern LLM frameworks. This position involves ownership of end-to-end development... 

    Jack & Jill

    San Francisco, CA
    3 days ago
  • Build AI is hiring an Evals Lead to benchmark video and world-model capabilities with and without our data. You will deeply understand evaluation validity, leakage controls, and whether numbers reflect real capability. You should have previously pushed evals that shifted... 
    Shift work

    Build AI

    San Francisco, CA
    4 days ago
  • Distyl AI is a frontier AI company seeking researchers to redefine how software is used in...  ...environments. You will design evaluation benchmarks and explore novel paradigms for...  ...equity, a hybrid work model, and access to leading AI models for enterprise-scale problems... 

    Distyl

    San Francisco, CA
    4 days ago
  • Artificial Analysis, Inc. in San Francisco is seeking a Member of Technical Staff to lead its new robotics vertical. This role entails defining benchmarking methodologies and producing leaderboards to evaluate robotics capabilities. Ideal candidates will possess a strong... 

    Artificial Analysis, Inc.

    San Francisco, CA
    3 days ago
  •  ...a technical leader to head its robotics coverage, developing benchmarking methodologies, leaderboards, and datasets that set industry standards...  ...concepts into actionable insights. This role blends product, research, and client-facing work with broad scope to shape our robotics... 

    Artificial Analysis

    San Francisco, CA
    16 hours ago
  • Innovaccer is seeking a Senior AI Researcher to design, develop, and deploy production-grade AI systems...  ...product, engineering, and data teams to build LLM-based solutions, AI agents, and RAG-driven workflows. You will lead experiments from hypothesis to production, read... 

    Socket

    San Francisco, CA
    4 days ago
  • $234.3k - $349k

     ...is where the world's leading enterprises orchestrate...  ...AI. About the roleAI research at WRITER isn't just about...  .... As an AI research scientist, you'll be at the...  ...workflowsBuild novel evaluation benchmarks and methodologies that...  ...— including LLM-as-judge frameworks, synthetic... 
    Full time
    Work at office
    Local area

    Writer

    San Francisco, CA
    5 days ago
  •  ...workersResponsibilitiesWe are looking for an exceptional AI Research Scientist to join our growing team. In this role,...  ...‑agent collaboration.Prototype and benchmark models; present findings internally and...  ....Hands‑on with PyTorch/JAX and modern LLM frameworks.Strong publication record in... 
    Remote work
    Flexible hours

    Workato

    San Francisco, CA
    4 days ago
  • $300k - $320k

     ...growing group of committed researchers, engineers, policy...  ...exceptional Research Scientist to join our Life Sciences...  ...training objectives, benchmarks, and agentic workflows...  ...~ Experience with LLM post-training: RLHF, RL...  ...target identification, lead optimization, ADMET modeling... 
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    2 days ago
  • Innovaccer in San Francisco is seeking an AI Researcher to join the AI team and help build production-grade AI systems, including LLM-based solutions, AI agents, RAG, and intelligent automation workflows. You will work closely with product, engineering, and data teams... 

    Uncover

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Benchmark Lead Research Scientist. Be the first to apply!