Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Scientist Benchmark Fellow

$250k
Part-time

Openrefinery.ai

What’s the next AI frontier?

Scientific discovery, automated. The question is, “how do you judge if an AI model is doing ‘good research’”. The answer comes from researchers who use AI models on real research. So, we're seeking one Fellow per STEM field who wants to build their field's first benchmark and put their name on it.

Core Responsibilities 

  • Validate results of the work produced by AI models’ paper reproduction and simulations.
  • Map out your field, including the most important research areas and the taxonomy within.
  • Set the standards for refining AI traces into benchmark samples; create benchmarks to evaluate frontier AI models against.

Authorship

  • Your benchmark will be published as a citable public release, including a paper and an open eval set.  You’ll be named a lead author. 
  • Bonus: Get more citations for your papers.


Requirements

  • PhD students or postdocs in biology, medicine, neuroscience, materials science, chemistry, physics, mechanical engineering, or chemical engineering.
  • Have published a paper in your field; can judge a paper in your field and catch author’s errors.
  • Currently doing AI for Science research in their fields. (For example, you are using AI agents like Claude Code or Codex for 20+ hours per week in your research).

How it works

  • Can be remote or in-person.
  • You can set your hours. You can take a break if there’s an academic deadline.
  • $50-100/hr (up to $250K annualized). If you’re in the US, F1/visa-friendly structures are available.
  • You get to work directly with our founders and with researchers at frontier AI labs.

About OpenRefinery

Our mission is to advance scientific research via AI advancement. We scaled Turing from $0 to 2.2B (Series E). The cofounders Kai, Alex, and Char led ML programs Facebook AI Research (FAIR) and Zoox, exited two companies before, and built a $300M ARR data business together.

Backed by Zetta Ventures (invested in Kaggle and Domino Data Lab), Chair of the MIT Board of Trustees, Audacious Ventures (investors in Reflection AI and Decagon), and AI researchers from OpenAI, Anthropic, Google DeepMind, Thinking Machines, Nvidia, and the United Arab Emirates’ AI institute. 

We partner with researchers at Stanford, MIT, Caltech, ETH Zürich, UPenn, Berkeley, NUS, and more.

How to Apply

Email View email address on nature.com with your field, how you use AI in your research, and one thing models consistently fails in your field. Résumé/CV is optional; you can just share your academic profile.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI Scientist Benchmark Fellow in United States vacancy
  • Snorkel AI in San Francisco is searching for a Research Scientist to lead the development of datasets and benchmarks for AI models. This customer-facing role involves working with academic partners and cross-functional teams to ensure impactful research outcomes. The ideal... 
    Suggested

    Snorkel AI

    San Francisco, CA
    3 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will author original, executable research problems that today’s frontier models cannot solve, sourced from published papers, datasets,... 
    Suggested

    Obsidian

    San Francisco, CA
    22 hours ago
  • $50 per hour

     ...in Biology or related fields for remote roles. You will design biology challenges for AI systems, evaluate their reasoning, and collaborate with top researchers to improve benchmarks. The position offers flexible hours at a competitive rate of $50+/hour, allowing access... 
    Suggested
    Remote job
    Hourly pay
    Flexible hours

    Turing

    Seattle, WA
    3 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks and participate in a new benchmark for scientific computing. You will author original, executable research problems that frontier models cannot solve, focusing on materials science and related subdomains... 
    Suggested
    Part time
    Immediate start

    Obsidian

    New York, NY
    22 hours ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, collaborating with leading AI labs. You will author original, executable research problems that frontier models cannot solve. Domains require... 
    Suggested
    Part time

    Mercor

    New York, NY
    3 days ago
  • Mercor is seeking PhD and Master’s scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will source material, craft executable problems, and build the grading criteria to measure model performance on challenging scenarios. You will... 

    Obsidian

    Miami, FL
    1 day ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering with leading AI labs. You will author original, executable research problems that frontier models cannot solve. Engagement lasts... 
    Part time
    Immediate start

    Mercor

    San Francisco, CA
    2 days ago
  • $234.3k - $349k

     ...leading enterprises orchestrate AI-powered work. Our vision is...  ...the world. As an AI research scientist, you'll be at the center of...  ...workflowsBuild novel evaluation benchmarks and methodologies that push...  ...communityMentor and uplevel fellow researchers and engineers on... 
    Full time
    Work at office
    Local area

    Writer

    San Francisco, CA
    5 days ago
  • Mercor is seeking a remote contractor for a benchmark dataset project focused on evaluating AI models in the Education domain. The role requires 15-20 hours per week and offers fully remote collaboration across US/Canada. You will work on tasks with clear ground-truth... 
    Remote job
    Weekly pay
    Part time
    For contractors

    Obsidian

    Washington DC
    1 day ago
  • $50 per hour

    A leading AI research accelerator is seeking remote PhD holders in Biology or related fields. As a contractor, you will design advanced biology questions, evaluate AI outputs, and collaborate with researchers on cutting-edge projects with top AI labs. This role offers... 
    Remote job
    For contractors
    Flexible hours

    Turing

    Chicago, IL
    3 days ago
  • A leading AI research organization is looking for remote candidates with PhDs in Biology or related fields to fine-tune AI language models. In this role, you will design biology questions and assess AI's problem-solving capabilities while collaborating with top researchers... 
    Remote job
    Contract work

    Turing

    New York, NY
    3 days ago
  • $233.3k - $385k

     ...together. Because at UKG, your work matters—and so do you.   About the Role: We are seeking a distinguished and customer-facing Fellow AI Engineer to define and lead UKG’s enterprise-wide Agentic AI strategy across the full UKG Product Suite. This role extends beyond... 
    Full time

    Ukg

    United States
    22 hours ago
  • $77.6k - $176k

    AI and ML Data ScientistThe Opportunity: As an Agentic AI Engineer and Data Scientist for military intelligence, you’re excited by the opportunity to design, develop, and...  ...prompt engineering, prompt evaluation, model benchmarking, fine-tuning, synthetic data generation,... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    McLean, VA
    5 days ago
  •  ...the RoleLexisNexis Legal & Professional is hiring an Applied AI Data Scientist to help shape the next generation of AI-powered legal...  ...Claude APIs to prototype and refine AI capabilities.Design the Benchmark & Standard Design rigorous, domain-aware evaluation methodologies... 
    Full time
    Remote work
    Worldwide

    RELX Group

    New York, NY
    3 days ago
  • Mercor is seeking a talented AI/education domain specialist to support a benchmark dataset project evaluating AI models on visual document understanding and instruction-following in the Education domain. The role is fully remote, roughly 15-20 hours per week, and open to... 
    Remote job
    For contractors

    Mercor Inc

    New York, NY
    22 hours ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. The role emphasizes material sourcing, prompt design, and robust grading criteria against... 
    Part time
    Immediate start

    Obsidian

    San Francisco, CA
    4 days ago
  • Mercor is seeking AI/ML specialists for a benchmark dataset project evaluating AI models on visual document understanding and instruction-following in Education. This is a remote, independent contractor role requiring ~15-20 hours per week and can be performed from the... 
    Remote job
    For contractors

    Obsidian

    San Francisco, CA
    1 day ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will join a collaboration with leading AI labs to publish original, executable research problems that frontier models cannot solve. Candidates should have PhD in chemistry-related... 
    Part time

    Mercor

    New York, NY
    4 days ago
  • $65 - $85 per hour

     ...is seeking a Computational Biologist for a remote role with a pay range of $65 to $85 per hour. In this position, you will design benchmark problems and analyze biological datasets. Applicants must possess an advanced degree in a relevant field and experience with bioinformatics... 
    Remote job
    Hourly pay
    Contract work
    10 hours per week
    Flexible hours

    Crossing Hurdles

    New York, NY
    3 days ago
  • $130k - $162.5k

     ...everyone. From advanced data analytics and AI to cybersecurity, we use innovative...  ...Group's enterprise AI team. We are AI/ML scientists and engineers with deep expertise in AI/ML...  ...LLMsCollaborating with the Optum AI team and fellow scientist and engineersDriving adoption... 
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Bellevue, WA
    5 days ago
  • $150k - $200k

     ...Beach, FLThe Role: We are seeking an exceptionally talented AI Scientist to join the Artificial Intelligence team at WorldQuant. The successful...  ...pay ranges for all roles based on job function and level, benchmarked against similar stage organizations. When finalizing an offer... 
    Work experience placement
    Casual work
    Flexible hours

    WorldQuant

    New York, NY
    3 days ago
  • $105k - $145k

    OverviewWe are looking for an AI Evaluation Scientistto design and execute evaluation...  ...across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data...  ...detect performance drift over time. Develop benchmark datasets, challenge sets, and scenario-... 
    Local area

    Steampunk

    McLean, VA
    5 days ago
  • $126k - $234k

     ...oncology drug discovery, computational biology, AI/ML, and data engineering. We are seeking an enthusiastic AI/ML scientist with deep expertise in digital pathology and...  ....Contribute to evaluation, validation, and benchmarking of image analysis algorithms and workflows.Perform... 
    Full time
    Local area
    Relocation
    Flexible hours

    Novartis

    Cambridge, MA
    4 days ago
  • $184k - $287.5k

     ...amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU...  ...searching for a highly motivated engineer to lead performance benchmarking and optimization efforts for our data center products. You will... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $60 - $90 per hour

     ...About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position... 
    Hourly pay
    Weekly pay
    Full time
    Contract work
    For contractors
    Summer work
    Remote work

    Mercor

    Remote
    22 hours ago
  • $225k - $280k

     ...Wizard Wizard is the top-performing AI Shopping Agent, delivering the best products...  ...Role We're looking for an Applied Scientist to own how we measure, understand, and...  ...Build and maintain evaluation datasets, benchmarks, and scoring frameworks Improve the LLM... 
    Remote work
    Flexible hours

    Wizard

    United States
    2 days ago
  • $179k - $217k

     ...Chief Science Officer, we are looking for a talented Senior AI Data Scientist to join our team and help shape the future of AI. ~Design...  ...designing A/B tests, offline evaluations, and ML benchmarking frameworks. ~Familiarity with vector databases, embedding... 
    Full time
    Work at office
    Remote work

    Chaos Labs

    Remote
    a month ago
  • $77.02k - $117.58k

     ...named one of the Tampa Bay Times’ Top Workplaces.SummaryThe AI Data Scientist - Precision Oncology develops and applies advanced analytics...  ...healthcare outcomes.ResponsibilitiesDevelop, evaluate, and benchmark machine learning models for prediction, prognosis, treatment... 

    Moffitt Cancer Center

    Tampa, FL
    1 day ago
  • $192k - $304.75k

     ...approach to accelerated computing. We're looking for a passionate AI research scientist with deep quantum computing expertise to path-find the...  ...and develop open AI models, curated datasets, and rigorous benchmarks that advance the state of the art and empower the broader... 
    Full time

    Nvidia

    Santa Clara, CA
    7 days ago
  • $69.4k - $158k

    AI and ML Research ScientistThe Opportunity:As an analytics professional, you’re excited...  ...Army applications.As an advanced data scientist or researcher on our Maritime and Ground...  ...establishing evaluation protocols such as benchmarks, ablations, and error analysesKnowledge of... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Adelphi, MD
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Scientist Benchmark Fellow. Be the first to apply!