Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Scientist Benchmark Fellow

$50 - $100 per hour
Part-time

Openrefinery.ai

What’s the next AI frontier?

Scientific discovery, automated. The question is, “how do you judge if an AI model is doing ‘good research’”. The answer comes from researchers who use AI models on real research. So, we're seeking one Fellow per STEM field who wants to build their field's first benchmark and put their name on it.

Core Responsibilities 

  • Validate results of the work produced by AI models’ paper reproduction and simulations.
  • Map out your field, including the most important research areas and the taxonomy within.
  • Set the standards for refining AI traces into benchmark samples; create benchmarks to evaluate frontier AI models against.

Authorship

  • Your benchmark will be published as a citable public release, including a paper and an open eval set. You’ll be named a lead author. 
  • Bonus: Get more citations for your papers.


Requirements

  • PhD students or postdocs in biology, medicine, neuroscience, materials science, chemistry, physics, mechanical engineering, or chemical engineering.
  • Have published a paper in your field; can judge a paper in your field and catch author’s errors.
  • Currently doing AI for Science research in their fields. (For example, you are using AI agents like Claude Code or Codex for 20+ hours per week in your research).

How it works

  • Can be remote or in-person.
  • You can set your hours. You can take a break if there’s an academic deadline.
  • $50-100/hr (up to $250K annualized). If you’re in the US, F1/visa-friendly structures are available.
  • You get to work directly with our founders and with researchers at frontier AI labs.

About OpenRefinery

Our mission is to advance scientific research via AI advancement. We scaled Turing from $0 to 2.2B (Series E). The cofounders Kai, Alex, and Char led ML programs Facebook AI Research (FAIR) and Zoox, exited two companies before, and built a $300M ARR data business together.

Backed by Zetta Ventures (invested in Kaggle and Domino Data Lab), Chair of the MIT Board of Trustees, Audacious Ventures (investors in Reflection AI and Decagon), and AI researchers from OpenAI, Anthropic, Google DeepMind, Thinking Machines, Nvidia, and the United Arab Emirates’ AI institute. 

We partner with researchers at Stanford, MIT, Caltech, ETH Zürich, UPenn, Berkeley, NUS, and more.

How to Apply

Email View email address on us.fitly.work with your field, how you use AI in your research, and one thing models consistently fails in your field. Résumé/CV is optional; you can just share your academic profile.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the AI Scientist Benchmark Fellow in United States vacancy
  • Snorkel AI in San Francisco is searching for a Research Scientist to lead the development of datasets and benchmarks for AI models. This customer-facing role involves working with academic partners and cross-functional teams to ensure impactful research outcomes. The ideal... 
    Suggested

    Snorkel AI

    San Francisco, CA
    15 hours ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will author original, executable research problems that today’s frontier models cannot solve, sourced from published papers, datasets,... 
    Suggested

    Obsidian

    San Francisco, CA
    2 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering with leading AI labs. You will author original, executable research problems that frontier models cannot solve. Engagement lasts... 
    Suggested
    Part time
    Immediate start

    Mercor

    San Francisco, CA
    4 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks and participate in a new benchmark for scientific computing. You will author original, executable research problems that frontier models cannot solve, focusing on materials science and related subdomains... 
    Suggested
    Part time
    Immediate start

    Obsidian

    New York, NY
    2 days ago
  • $50 per hour

     ...in Biology or related fields for remote roles. You will design biology challenges for AI systems, evaluate their reasoning, and collaborate with top researchers to improve benchmarks. The position offers flexible hours at a competitive rate of $50+/hour, allowing access... 
    Suggested
    Remote job
    Hourly pay
    Flexible hours

    Turing

    Seattle, WA
    15 hours ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, collaborating with leading AI labs. You will author original, executable research problems that frontier models cannot solve. Domains require... 
    Part time

    Mercor

    New York, NY
    15 hours ago
  • Mercor is seeking PhD and Master’s scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will source material, craft executable problems, and build the grading criteria to measure model performance on challenging scenarios. You will... 

    Obsidian

    Miami, FL
    3 days ago
  • $234.3k - $349k

     ...leading enterprises orchestrate AI-powered work. Our vision is...  ...the world. As an AI research scientist, you'll be at the center of...  ...workflowsBuild novel evaluation benchmarks and methodologies that push...  ...communityMentor and uplevel fellow researchers and engineers on... 
    Full time
    Work at office
    Local area

    Writer

    San Francisco, CA
    2 days ago
  • $50 per hour

    A leading AI research accelerator is seeking remote PhD holders in Biology or related fields. As a contractor, you will design advanced biology questions, evaluate AI outputs, and collaborate with researchers on cutting-edge projects with top AI labs. This role offers... 
    Remote job
    For contractors
    Flexible hours

    Turing

    Chicago, IL
    15 hours ago
  • $233.3k - $385k

     ...together. Because at UKG, your work matters—and so do you.   About the Role: We are seeking a distinguished and customer-facing Fellow AI Engineer to define and lead UKG’s enterprise-wide Agentic AI strategy across the full UKG Product Suite. This role extends beyond... 
    Full time

    Ukg

    United States
    15 hours ago
  • $119k - $286k

     ..., Data Science and IT to prototype, build, and scale practical AI-driven solutions that improve quality, cost, cycle time, and engineering...  ...skills, retrieval-augmented generation (RAG and Graph-RAG), benchmarking, and quality evaluation. Machine Learning Production: Develop... 
    Full time
    Local area
    Immediate start

    Micron

    San Jose, CA
    4 days ago
  •  ...the RoleLexisNexis Legal & Professional is hiring an Applied AI Data Scientist to help shape the next generation of AI-powered legal...  ...Claude APIs to prototype and refine AI capabilities.Design the Benchmark & Standard Design rigorous, domain-aware evaluation methodologies... 
    Full time
    Remote work
    Worldwide

    RELX Group

    New York, NY
    15 hours ago
  • $77.6k - $176k

    AI and ML Data ScientistThe Opportunity: As an Agentic AI Engineer and Data Scientist for military intelligence, you’re excited by the opportunity to design, develop, and...  ...prompt engineering, prompt evaluation, model benchmarking, fine-tuning, synthetic data generation,... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    McLean, VA
    2 days ago
  • Mercor is seeking AI/ML specialists for a benchmark dataset project evaluating AI models on visual document understanding and instruction-following in Education. This is a remote, independent contractor role requiring ~15-20 hours per week and can be performed from the... 
    Remote job
    For contractors

    Obsidian

    San Francisco, CA
    3 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. The role emphasizes material sourcing, prompt design, and robust grading criteria against... 
    Part time
    Immediate start

    Obsidian

    San Francisco, CA
    1 day ago
  • Mercor is seeking a talented AI/education domain specialist to support a benchmark dataset project evaluating AI models on visual document understanding and instruction-following in the Education domain. The role is fully remote, roughly 15-20 hours per week, and open to... 
    Remote job
    For contractors

    Mercor Inc

    New York, NY
    2 days ago
  • $130k - $162.5k

     ...everyone. From advanced data analytics and AI to cybersecurity, we use innovative...  ...Group's enterprise AI team. We are AI/ML scientists and engineers with deep expertise in AI/ML...  ...LLMsCollaborating with the Optum AI team and fellow scientist and engineersDriving adoption... 
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Bellevue, WA
    2 days ago
  • $150k - $200k

     ...Beach, FLThe Role: We are seeking an exceptionally talented AI Scientist to join the Artificial Intelligence team at WorldQuant. The successful...  ...pay ranges for all roles based on job function and level, benchmarked against similar stage organizations. When finalizing an offer... 
    Work experience placement
    Casual work
    Flexible hours

    WorldQuant

    New York, NY
    15 hours ago
  • $105k - $145k

    OverviewWe are looking for an AI Evaluation Scientistto design and execute evaluation...  ...across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data...  ...detect performance drift over time. Develop benchmark datasets, challenge sets, and scenario-... 
    Local area

    Steampunk

    McLean, VA
    2 days ago
  • $86.32k - $154.96k

     ...ecosystem where experimental, clinical, and data scientists collaborate to accelerate discoveries,...  ...a skilled and collaborative Agentic AI Scientist to lead the design, development...  ...of scientific contribution.)Publication benchmark: 1 to 2 first-author papers with a journal... 
    Full time
    Work experience placement
    Work at office

    St.Jude Children´s Research Hospital

    Memphis, TN
    1 day ago
  • $126k - $234k

     ...oncology drug discovery, computational biology, AI/ML, and data engineering. We are seeking an enthusiastic AI/ML scientist with deep expertise in digital pathology and...  ....Contribute to evaluation, validation, and benchmarking of image analysis algorithms and workflows.Perform... 
    Full time
    Local area
    Relocation
    Flexible hours

    Novartis

    Cambridge, MA
    1 day ago
  • $65 - $85 per hour

     ...is seeking a Computational Biologist for a remote role with a pay range of $65 to $85 per hour. In this position, you will design benchmark problems and analyze biological datasets. Applicants must possess an advanced degree in a relevant field and experience with bioinformatics... 
    Remote job
    Hourly pay
    Contract work
    10 hours per week
    Flexible hours

    Crossing Hurdles

    New York, NY
    15 hours ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will join a collaboration with leading AI labs to publish original, executable research problems that frontier models cannot solve. Candidates should have PhD in chemistry-related... 
    Part time

    Mercor

    New York, NY
    1 day ago
  • $184k - $287.5k

     ...amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU...  ...searching for a highly motivated engineer to lead performance benchmarking and optimization efforts for our data center products. You will... 
    Full time

    Nvidia

    Santa Clara, CA
    15 hours ago
  • $60 - $90 per hour

     ...About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position... 
    Hourly pay
    Weekly pay
    Full time
    Contract work
    For contractors
    Summer work
    Remote work

    Mercor

    Remote
    15 hours ago
  • $225k - $280k

     ...Wizard Wizard is the top-performing AI Shopping Agent, delivering the best products...  ...Role We're looking for an Applied Scientist to own how we measure, understand, and...  ...Build and maintain evaluation datasets, benchmarks, and scoring frameworks Improve the LLM... 
    Remote work
    Flexible hours

    Wizard

    United States
    4 days ago
  • $179k - $217k

     ...Chief Science Officer, we are looking for a talented Senior AI Data Scientist to join our team and help shape the future of AI. ~Design...  ...designing A/B tests, offline evaluations, and ML benchmarking frameworks. ~Familiarity with vector databases, embedding... 
    Full time
    Work at office
    Remote work

    Chaos Labs

    Remote
    a month ago
  • $69.4k - $158k

    AI and ML Research ScientistThe Opportunity:As an analytics professional, you’re excited...  ...Army applications.As an advanced data scientist or researcher on our Maritime and Ground...  ...establishing evaluation protocols such as benchmarks, ablations, and error analysesKnowledge of... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Adelphi, MD
    15 hours ago
  •  ...intelligence by establishing the first distributed AI infrastructure dedicated to personalized...  ...of AI.About the Role:The AI Research Scientist will contribute to webAI’s development of...  ...for distributed AIDevelop and evaluate benchmarks, datasets, and experimental frameworks to... 
    Live out
    Work at office
    Local area

    webAI

    Austin, TX
    15 hours ago
  • $160k - $220k

     ...enjoy multi-year runway.About the RoleWe’re looking for an AI Research Scientist to advance the methodological frontier of AI in healthcare....  ...understand that healthcare validation standards are higher than benchmark culture alone, and you are energized by the opportunity to... 
    Temporary work
    Work at office
    Monday to Friday
    Monday to Thursday

    Sprinter Health

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Scientist Benchmark Fellow. Be the first to apply!