LLM Benchmark Lead Research Scientist
Vals AI
Vals AI in San Francisco is seeking exceptional researchers and research engineers to design and build the next generation of AI benchmarks. You will create high-impact, challenging evaluations that push the boundaries of what we can measure in foundation models. This role is perfect for someone with deep research expertise who wants to see their work directly influence how the world evaluates AI systems. We are building the standard for evaluating the ability of LLMs to perform real-world #J-18808-Ljbffr Vals AI
- ...Inc. in San Francisco is seeking strong engineers to own the leaderboards that evaluate real-world tasks by LLMs. You will test and benchmark new models across domains like law, tax, coding, finance and more, analyzing error modes and collaborating with the communications...Suggested
$200k - $325k
...when Snorkel started as a research project in the Stanford... ...to empower scientists, engineers, financial experts... ...collaborate with partners and lead the development of the next frontier benchmarks and datasets. This is a... ...Stay at the frontier of LLM evaluation research and...SuggestedLocal area- Camb in San Francisco is searching for a Lead LLM R&D Researcher. This full-time position involves leading innovations in natural language processing and colloquial translation as well as collaborating with R&D teams. The ideal candidate will have a PhD or equivalent experience...SuggestedFull time
- Enam, Inc. is seeking a senior ML researcher to lead exploration of new neural network foundations and architectures, with a strong emphasis on LLM applications. You will design experiments, validate novel hypotheses, and translate research ideas into scalable systems in...Suggested
$357k
...enhance life inside and outside of work. Responsibilities Lead AI Research Scientist in Workato’s AI Research Lab. This research‑first role focuses... ...teams. Hands‑on experience with PyTorch or JAX and modern LLM frameworks. Proven track record of transitioning research...SuggestedWork at officeFlexible hours- Bay Area Full time Lead LLM R&D, innovate in zero-shot performance and colloquial translation, collaborate, and publish research in top conferences. Join Camb.ai's pioneering AI team... ...just innovating; we're setting SOTA benchmarks. Backed by elite global VCs, partnering...Full time
- Snorkel AI in San Francisco seeks a Research Scientist to work on reinforcement learning for training and aligning large language models. This... ...on data, reward signals, and training procedures that steer LLM behavior in reliable directions - a core capability that differentiates...
- Cincinnatus LLC is recruiting researchers to design and author multi-step evaluation tasks for frontier AI benchmarks. The role emphasizes translating scientific method into... ...full-time W-2 engagement, integrated with leading AI lab teams through Cincinnatus’ extended...Remote jobFull timePart time
- Infinity is building ResearchBench, a high-quality research-replication benchmark. A decomposition-and-testing agent converts papers into roughly 400 leaves across datasets, compute, and function implementation. A separate harness then runs backends like Claude Code, Codex...
$150k - $250k
...goods, and global social organizations.We research and deploy technologies that power AI-... ...of patient journeys.Distyl is backed by leading investors including Lightspeed Venture Partners... ...to drive incremental improvements on benchmarks or optimize an existing process but...Work at office3 days per week$216k - $270k
Scale Labs, Research Scientist — Frontier Risk EvaluationsAs the leading data and evaluation partner for frontier AI companies, Scale... ...Experience in crafting evaluations and benchmarks, or a background in data science roles related to LLM technologies.Experience with red-...Full time$262.5k - $299.6k
...Applied Researcher II (AI Foundations, LLM Core and Agentic AI) Overview At Capital One, we are creating trustworthy... .... For years, Capital One has been leading the industry in using machine... ...with a cross‑functional team of data scientists, software engineers, machine...Full timePart timeLocal areaFlexible hours- A leading investment firm in San Francisco is seeking a Chief Scientist to advance state-of-the-art technologies in Machine Learning for cybersecurity. Ideal candidates... ...background and experience with reputed research groups. While prior startup experience is a plus...
- Scale AI in New York seeks a Research Scientist Manager to lead a world-class team, define the research roadmap, and drive from prototyping to deployment. You will mentor researchers, collaborate with engineering and product, publish results, and influence strategic directions...
$141.1k - $281.7k
...image, and videos, and creates an industry‑leading technical platform to improve creative... ...marketers and designers. Responsibilities Research and implement ranking ML models using... ...models Conduct innovative research on novel LLM architectures and training methods,...- Scale AI, Inc. seeks a Research Scientist Manager to lead a world-class team at the forefront of GenAI research, evaluation, and agent/RL infrastructure. You will define the research roadmap and drive execution from prototyping to deployment, balancing technical leadership...
- ...AI at scale in San Francisco. We are seeking a research-leaning engineer to own end-to-end inference research bets for LLM serving, including speculative decoding, quantization... ...management. You will work with the research lead and Forward Deployed Engineers to deploy models...
- ...workersResponsibilitiesWe are looking for an exceptional AI Research Scientist to join our growing team. In this role,... ...‑agent collaboration.Prototype and benchmark models; present findings internally and... ....Hands‑on with PyTorch/JAX and modern LLM frameworks.Strong publication record in...Remote workFlexible hours
$234.3k - $349k
...is where the world's leading enterprises orchestrate... ...AI. About the roleAI research at WRITER isn't just about... .... As an AI research scientist, you'll be at the... ...workflowsBuild novel evaluation benchmarks and methodologies that... ...— including LLM-as-judge frameworks, synthetic...Full timeWork at officeLocal area- A biotech company focusing on AI in drug discovery seeks a highly skilled AI Research Scientist to drive the development of next-generation AI models used in drug discovery workflows. The ideal candidate possesses a Master's or PhD in a relevant field along with at least...
- WRITER seeks a staff AI research scientist to advance a high-impact research program on large language models, agentic reasoning, and system-level capabilities for enterprise-scale AI. This role blends independent research with direct product impact. Based in San Francisco...
$300k - $320k
...Research Scientist Anthropic's mission is to create reliable, interpretable... ...model training objectives, benchmarks, and agentic workflows. You'... ...~ Experience with LLM post-training: RLHF, RL from... ...pipelines — target identification, lead optimization, ADMET modeling...Work at officeVisa sponsorshipFlexible hours- ...a technical leader to head its robotics coverage, developing benchmarking methodologies, leaderboards, and datasets that set industry standards... ...concepts into actionable insights. This role blends product, research, and client-facing work with broad scope to shape our robotics...
- Artificial Analysis, Inc. in San Francisco is seeking a Member of Technical Staff to lead its new robotics vertical. This role entails defining benchmarking methodologies and producing leaderboards to evaluate robotics capabilities. Ideal candidates will possess a strong...
- The Token Company is seeking an ML Researcher to own a slice of open problems in applied AI, focusing on what information in an LLM context matters and how to represent it efficiently. This high-autonomy role requires running many experiments, reproducing papers, and shipping...
- A pioneering research lab is seeking a Senior Software Engineer focused on building next-generation conversational AI interfaces. The ideal... ...will have deep expertise in Python and experience with modern LLM frameworks. This position involves ownership of end-to-end development...
$125k - $225k
...FutureSearch is looking for exceptional Research Scientists to evaluate and improve state-of-the-art forecasting and agentic LLM web research. We are an elite team of engineers... ...our ICLR Workshop paper in Jan 2026, and benchmarks like Deep Research Bench and Bench to the...Remote workFlexible hours- Carnaby Fox is seeking a Member of Technical Staff (AI Research) in San Francisco to help shape the research direction for frontier AI... ..., evaluate LLMs, and improve data quality for high-stakes AI benchmarks. The role emphasizes independent ownership, pioneering benchmark...
- ...also happen to be the world's leading generative AI studio—we're... ...We’ve developed an in-house LLM storytelling system that blends... ...Project Astra, and top-tier AI researchers. As an early member of... ...saving techniques, scalable benchmarking and checkpoint selection, are...Work at officeVisa sponsorship
$172.5k - $260.1k
...through AI, driving innovation, and keeping Salesforce's core values at the heart of it all.Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Benchmark Lead Research Scientist. Be the first to apply!
- scientist ii San Francisco, CA
- scientist 1 San Francisco, CA
- image scientist San Francisco, CA
- downstream processing scientist San Francisco, CA
- entry level research scientist San Francisco, CA
- qc scientist San Francisco, CA
- research scientist San Francisco, CA
- analytical scientist San Francisco, CA
- research scientist - biology San Francisco, CA
- genomics scientist San Francisco, CA

