Research Scientist - Frontier Benchmarks
Snorkel AI
About Snorkel
At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data.
We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler!
ABOUT THE ROLE
We're looking for a Research Scientist to collaborate with partners and lead the development of the next frontier benchmarks and datasets. This is a highly visible, customer-facing role at the intersection of research, company strategy, and go-to-market. You'll design datasets taking into account frontier model performance and work with our academic partners, and then partner with delivery, product and go-to-market to scale out production. You will also serve as a credible technical partner for our customers, prospects, and drive results that impact the broader research community.
This role reports directly to the Head of Research and is ideal for someone who is energized by cross-functional work and wants to understand how startups operate across research, data operations, and commercial teams.
MAIN RESPONSIBILITIES
- Design state of the art datasets that drive frontier model training and evaluation based on current model performance and academic partnerships
- Translate benchmark insights into clear, compelling narratives that articulate the ROI of expert-curated data for customer-facing presentations, technical reports, and go-to-market materials.
- Work cross-functionally with data operations, product, engineering, and strategy to surface research findings that inform the company roadmap.
- Stay at the frontier of LLM evaluation research and bring best practices into Snorkel's workflows
- Represent Snorkel's research externally through publications, blog posts, conference talks, and customer engagements that advance the conversation around data-centric AI
PREFERRED QUALIFICATIONS
- Strong research background in AI/ML evaluation, NLP, or related fields, with a track record of rigorous experimental design — especially around measuring the impact of training and evaluation data on model behavior.
- Exceptional communication skills — able to present complex technical findings clearly to both technical and non-technical audiences
- Comfort operating in a fast-moving, cross-functional environment with ambiguous problem spaces
- Genuine interest in GTM strategy, startup dynamics, and the commercial side of AI data services.
- Ph.D. in machine learning, NLP, or a related field preferred; equivalent industry or research lab experience considered.
Be Your Best at Snorkel
Joining Snorkel AI means becoming part of a company that has market proven solutions, robust funding, and is scaling rapidly—offering a unique combination of stability and the excitement of high growth. As a member of our team, you’ll have meaningful opportunities to shape priorities and initiatives, influence key strategic decisions, and directly impact our ongoing success. Whether you’re looking to deepen your technical expertise, explore leadership opportunities, or learn new skills across multiple functions, you’re fully supported in building your career in an environment designed for growth, learning, and shared success.
Snorkel AI is proud to be an Equal Employment Opportunity employer and is committed to building a team that represents a variety of backgrounds, perspectives, and skills. Snorkel AI embraces diversity and provides equal employment opportunities to all employees and applicants for employment. Snorkel AI prohibits discrimination and harassment of any type on the basis of race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local law. All employment is decided on the basis of qualifications, performance, merit, and business need.
We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.
$176k - $304k
...Research Scientist, Frontier Capabilities Cambridge, MA USA; San Francisco, CA USA Your impact at LILA We're building a talent-dense... ...hypothesis space is vast and reward signal is thin. Existing benchmarks do not capture these nuances. The goal is to build...SuggestedFull timeWork at officeLocal areaFlexible hoursShift work$216k - $270k
Scale Labs, Research Scientist — Frontier Risk EvaluationsAs the leading data and evaluation partner for frontier AI companies, Scale plays an integral... ...team.Nice to have:Experience in crafting evaluations and benchmarks, or a background in data science roles related to LLM...SuggestedFull time- ...computing offers a path through these bottlenecks. As an ML Research Scientist, you'll work at the frontier of generative modeling and quantum acceleration,... ..., denoising, and likelihood estimation Develop and benchmark novel solver methods for diffusion ODEs/SDEs Quantum...SuggestedFull timeCasual workVisa sponsorship
- Cincinnatus LLC is recruiting researchers to design and author multi-step evaluation tasks for frontier AI benchmarks. The role emphasizes translating scientific method into practical tasks, with a focus on Python-based analysis, rigorous evaluation, and clear written conclusions...SuggestedRemote jobFull timePart time
- ...profound global impact. About the Role Frontier AI is moving toward scientific... ...for this next era. We are looking for a Research Scientist who can help define Quantum AI: not just... ...propose, test, and refine hypotheses. Benchmarking frameworks that reveal when a new computational...SuggestedCasual workVisa sponsorship
$250k
...data and evaluation infrastructure that frontier AI labs use to make their models better.... ...rigorous evaluations that go beyond static benchmarks. We are a small, early team (post Series... ...and measured. Working directly with research teams at top AI labs, you’ll experiment...$150k - $250k
...rearchitect critical operations for the frontier of AI. Our customers include the largest... ...goods, and global social organizations.We research and deploy technologies that power AI-... ...want to drive incremental improvements on benchmarks or optimize an existing process but...Work at office3 days per week- Traverse is a research data lab building reinforcement learning environments for frontier AI labs. As a Research Scientist, you will design and build RL environments that teach models to do work that has historically required years of human expertise. You’ll work at the...
- Infinity is building ResearchBench, a high-quality research-replication benchmark. A decomposition-and-testing agent converts papers into roughly 400 leaves across datasets, compute, and function implementation. A separate harness then runs backends like Claude Code, Codex...
- United States Digital Space LLC is seeking energetic researchers and engineers to join their Secure Intelligence Institute (SII). You'll... ...conduct impactful research to improve the security and privacy of frontier intelligence systems. Responsibilities include developing...
- Vals AI in San Francisco is seeking exceptional researchers and research engineers to design and build the next generation of AI benchmarks. You will create high-impact, challenging evaluations that push the boundaries of what we can measure in foundation models. This role...
- ...Research Scientist / Machine Learning Scientist Location: SF Bay Area/Hybrid / Remote Type... ...protocols that go beyond traditional benchmarks Analyze large-scale human voting and... ...by industry leaders pushing the frontier of safe and reliable AI. Sundar Pichai...Full timeRemote work
$200k - $335k
...The role As a research scientist, you will design, implement, and optimize the large-scale... ...training infrastructure that powers our frontier reinforcement learning stack. This is... ...by partners like Kleiner Perkins, Benchmark, Sequoia, Lux, and Greenoaks. Who Thrives...Full timeWork at officeVisa sponsorshipRelocation package- ...Researcher Position at Hedra Hedra is building a world-class Physical AI research team... ...researchers who are excited to go beyond benchmarks and build models that operate in the... ...research into production Stay at the frontier of the field — synthesizing relevant literature...Work at office
$125k - $225k
...FutureSearch is looking for exceptional Research Scientists to evaluate and improve state-of-the-... ...of engineers and researchers at the frontier of AI epistemics. We have the best publicly... ...ICLR Workshop paper in Jan 2026, and benchmarks like Deep Research Bench and Bench to...Remote workFlexible hours$320k
...is a quickly growing group of committed researchers, engineers, policy experts, and... ...beneficial AI systems. About the Team The Frontier Red Team (FRT) is a small, focused technical... ...massive threat surface. As a Research Scientist on FRT focusing on cyber, you'll build...Work at officeRelocationVisa sponsorshipFlexible hours$160k - $220k
...year runway.About the RoleWe’re looking for an AI Research Scientist to advance the methodological frontier of AI in healthcare. This role is ideal for someone... ...healthcare validation standards are higher than benchmark culture alone, and you are energized by the opportunity...Temporary workWork at officeMonday to FridayMonday to Thursday$140k - $200k
...breakthrough AI models at leading research labs and enterprises. Since... ...integrated solutions for frontier AI development: Enterprise... ...not a traditional research scientist role. You will not spend months... ...with evaluation and benchmarking of LLMs — designing metrics,...Work at officeFlexible hours2 days per week- Velvet is a data research company building the datasets that power the next... ...quality audiovisual training data for frontier labs. We’re hiring a Research Scientist to develop and fine‑tune models... ...Build evaluation frameworks and benchmarks to rigorously measure enhancement...Immediate startShift work
$234.3k - $349k
...work with AI. About the roleAI research at WRITER isn't just about... ...world. As an AI research scientist, you'll be at the center of... ...workflowsBuild novel evaluation benchmarks and methodologies that push... ...representing WRITER at the frontier of the field and contributing...Full timeWork at officeLocal area- ...applying those ideas to one of the next frontiers in machine learning. At Granica, we're... ...from structured information at scale. Our research is led by Prof. Andrea Montanari (... ...approaches. Design rigorous experiments and benchmarks to measure model quality and efficiency....Work at office
$250k - $400k
...stealth AI start-up building frontier reasoning models for... ...genuinely novel AI for Science research, combining frontier reasoning... ...Building evaluation frameworks and benchmarks for complex, multi-step... ...Researcher or experienced Research Scientist. What matters most is hands‑...$300k - $320k
...Research Scientist Anthropic's mission is to create reliable, interpretable, and steerable... ...knowledge into model training objectives, benchmarks, and agentic workflows. You'll help... ...accelerated biology while shaping how frontier models reason about and execute computational...Work at officeVisa sponsorshipFlexible hours$250k
...and evaluation infrastructure that frontier AI labs use to improve their... ...evaluations that go beyond static benchmarks. We're a small, early team (post–... ...re building out our post-training research team and hiring 2–3 Research Scientists to work together on this mission....Full timeInternshipShift work$150k - $250k
...data and evaluation infrastructure that frontier AI labs use to improve their models, partnering... ...evaluations that go beyond static benchmarks. It's a small, early team where... ...top quant firms, big tech, and leading AI research labs. Founded 2025 · 11–50 people · Industry...Full timeShift work$295k
...at OpenAI, and is guided by OpenAI's Preparedness Framework. Frontier AI models have the potential to benefit all of humanity, but also... ...productivity can also accelerate exploitation. As a Researcher for cybersecurity risks, you will help design and implement an...- OpenAI in San Francisco seeks exceptional researchers to push the frontier of safety mitigations, helping derisk frontier models and advance techniques from interpretability, robustness and alignment to ensure safe deployments. This role requires deep technical experience...
- ...enterprises. We aim to push the frontier of AI that understands real,... ...role is for an experienced scientist who thrives both in... ...and deep content extraction. Research, evaluate, and integrate the... ...product impact. Develop new benchmarks, datasets, and evaluation methodologies...
- OpenAI is seeking a Researcher for Frontier Cybersecurity Risks to design and implement an end-to-end mitigation stack that reduces severe cyber misuse across OpenAI products. The role requires deep technical expertise in deep learning, transformer models, and safeguarding...
- Crucibl is building judgment at scale for enterprise AI. You will push the frontier of how models reason on high-stakes business decisions, and translate findings into testable experiments that inform product direction. You’ll design evaluation frameworks for uncertainty...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist - Frontier Benchmarks. Be the first to apply!
- scientist ii San Francisco, CA
- scientist 1 San Francisco, CA
- image scientist San Francisco, CA
- downstream processing scientist San Francisco, CA
- entry level research scientist San Francisco, CA
- qc scientist San Francisco, CA
- research scientist San Francisco, CA
- analytical scientist San Francisco, CA
- research scientist - biology San Francisco, CA
- genomics scientist San Francisco, CA


