LLM Benchmark Lead Research Scientist
Vals AI
Vals AI in San Francisco is seeking exceptional researchers and research engineers to design and build the next generation of AI benchmarks. You will create high-impact, challenging evaluations that push the boundaries of what we can measure in foundation models. This role is perfect for someone with deep research expertise who wants to see their work directly influence how the world evaluates AI systems. We are building the standard for evaluating the ability of LLMs to perform real-world #J-18808-Ljbffr Vals AI
- Vibehackers in San Francisco is seeking a Senior Staff Researcher to design and implement novel benchmarks that evaluate real-world LLM capabilities. You will lead research and engineering efforts, publish results, and collaborate with model labs and enterprise partners...SuggestedRelocation package
- A leading AI CRM company is seeking an AI Research Scientist/Research Engineer in San Francisco. You will focus on advancing AI techniques, developing models, and solving real-world enterprise problems. Ideal candidates hold a Ph.D. or have strong AI product experience...Suggested
- ...Inc. in San Francisco is seeking strong engineers to own the leaderboards that evaluate real-world tasks by LLMs. You will test and benchmark new models across domains like law, tax, coding, finance and more, analyzing error modes and collaborating with the communications...Suggested
$200k - $325k
...when Snorkel started as a research project in the Stanford... ...to empower scientists, engineers, financial experts... ...collaborate with partners and lead the development of the next frontier benchmarks and datasets. This is a... ...Stay at the frontier of LLM evaluation research and...SuggestedLocal area- Innovaccer in San Francisco seeks a Principal AI Researcher to lead production-grade AI systems, solving complex problems at scale. You will... ...design, develop, and deploy AI-powered applications, including LLM-based solutions, AI agents, RAG, and automation workflows....Suggested
- Camb in San Francisco is searching for a Lead LLM R&D Researcher. This full-time position involves leading innovations in natural language processing and colloquial translation as well as collaborating with R&D teams. The ideal candidate will have a PhD or equivalent experience...Full time
- Enam, Inc. is seeking a senior ML researcher to lead exploration of new neural network foundations and architectures, with a strong emphasis on LLM applications. You will design experiments, validate novel hypotheses, and translate research ideas into scalable systems in...
- Bay Area Full time Lead LLM R&D, innovate in zero-shot performance and colloquial translation, collaborate, and publish research in top conferences. Join Camb.ai's pioneering AI team... ...just innovating; we're setting SOTA benchmarks. Backed by elite global VCs, partnering...Full time
- Mercor is seeking a researcher to join its GenAI team to translate the scientific method into complex, multi-step tasks and to design studies... ...with researchers to ensure consistent, high-quality benchmark design and evaluation of frontier models. #J-18808-Ljbffr MercorRemote job
- Cincinnatus LLC is recruiting researchers to design and author multi-step evaluation tasks for frontier AI benchmarks. The role emphasizes translating scientific method into... ...full-time W-2 engagement, integrated with leading AI lab teams through Cincinnatus’ extended...Remote jobFull timePart time
- Scale AI, Inc. seeks Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation. The role focuses on building benchmarks and diagnosing model failure modes in text and multimodal modalities within the GenAI...
$150k - $250k
...goods, and global social organizations.We research and deploy technologies that power AI-... ...of patient journeys.Distyl is backed by leading investors including Lightspeed Venture Partners... ...to drive incremental improvements on benchmarks or optimize an existing process but...Work at office3 days per week- ...standard for how AI is measured and to work with leading AI labs and enterprises. You’ll develop benchmarking methodologies, manage relationships with top AI labs... ...drive product direction. The role blends product, research, and client-facing work for those who enjoy...
- Pocket FM is seeking a Research Scientist to advance generative AI for long-form storytelling. You will work at the intersection of LLMs, multimodal AI, agentic systems, and scalable machine learning, translating research into production systems that delight millions of...Worldwide
$141.1k - $281.7k
...image, and videos, and creates an industry‑leading technical platform to improve creative... ...marketers and designers. Responsibilities Research and implement ranking ML models using... ...models Conduct innovative research on novel LLM architectures and training methods,...- A leading investment firm in San Francisco is seeking a Chief Scientist to advance state-of-the-art technologies in Machine Learning for cybersecurity. Ideal candidates... ...background and experience with reputed research groups. While prior startup experience is a plus...
$180.6k - $225.75k
...works with the industry's leading AI labs to provide high quality... ...progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise in LLM post-training (SFT, RLHF,... ...and will focus on building benchmarks and diagnosing model failure...Full time$216k - $270k
Scale Labs, Research Scientist — Frontier Risk EvaluationsAs the leading data and evaluation partner for frontier AI companies, Scale... ...Experience in crafting evaluations and benchmarks, or a background in data science roles related to LLM technologies.Experience with red-...Full time$262.5k - $299.6k
Applied Researcher II (AI Foundations, LLM Core and Agentic AI) Overview: At Capital One, we are creating... ...good. For years, Capital One has been leading the industry in using machine... ...with a cross-functional team of data scientists, software engineers, machine learning...Full timePart timeLocal areaFlexible hours- A biotech company focusing on AI in drug discovery seeks a highly skilled AI Research Scientist to drive the development of next-generation AI models used in drug discovery workflows. The ideal candidate possesses a Master's or PhD in a relevant field along with at least...
- A pioneering research lab is seeking a Senior Software Engineer focused on building next-generation conversational AI interfaces. The ideal... ...will have deep expertise in Python and experience with modern LLM frameworks. This position involves ownership of end-to-end development...
- Build AI is hiring an Evals Lead to benchmark video and world-model capabilities with and without our data. You will deeply understand evaluation validity, leakage controls, and whether numbers reflect real capability. You should have previously pushed evals that shifted...Shift work
- Distyl AI is a frontier AI company seeking researchers to redefine how software is used in... ...environments. You will design evaluation benchmarks and explore novel paradigms for... ...equity, a hybrid work model, and access to leading AI models for enterprise-scale problems...
- Artificial Analysis, Inc. in San Francisco is seeking a Member of Technical Staff to lead its new robotics vertical. This role entails defining benchmarking methodologies and producing leaderboards to evaluate robotics capabilities. Ideal candidates will possess a strong...
- ...a technical leader to head its robotics coverage, developing benchmarking methodologies, leaderboards, and datasets that set industry standards... ...concepts into actionable insights. This role blends product, research, and client-facing work with broad scope to shape our robotics...
- Innovaccer is seeking a Senior AI Researcher to design, develop, and deploy production-grade AI systems... ...product, engineering, and data teams to build LLM-based solutions, AI agents, and RAG-driven workflows. You will lead experiments from hypothesis to production, read...
$234.3k - $349k
...is where the world's leading enterprises orchestrate... ...AI. About the roleAI research at WRITER isn't just about... .... As an AI research scientist, you'll be at the... ...workflowsBuild novel evaluation benchmarks and methodologies that... ...— including LLM-as-judge frameworks, synthetic...Full timeWork at officeLocal area- ...workersResponsibilitiesWe are looking for an exceptional AI Research Scientist to join our growing team. In this role,... ...‑agent collaboration.Prototype and benchmark models; present findings internally and... ....Hands‑on with PyTorch/JAX and modern LLM frameworks.Strong publication record in...Remote workFlexible hours
$300k - $320k
...growing group of committed researchers, engineers, policy... ...exceptional Research Scientist to join our Life Sciences... ...training objectives, benchmarks, and agentic workflows... ...~ Experience with LLM post-training: RLHF, RL... ...target identification, lead optimization, ADMET modeling...Work at officeVisa sponsorshipFlexible hours- Innovaccer in San Francisco is seeking an AI Researcher to join the AI team and help build production-grade AI systems, including LLM-based solutions, AI agents, RAG, and intelligent automation workflows. You will work closely with product, engineering, and data teams...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Benchmark Lead Research Scientist. Be the first to apply!
- materials scientist San Francisco, CA
- scientist assay development San Francisco, CA
- entry level research scientist San Francisco, CA
- health scientist San Francisco, CA
- quality control scientist San Francisco, CA
- deep learning scientist San Francisco, CA
- application scientist San Francisco, CA
- scientist antibody discovery San Francisco, CA
- senior analytical scientist San Francisco, CA
- decision scientist San Francisco, CA

