Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Scientist Benchmark Fellow

$50 - $100 per hour
Part-time

Openrefinery.ai

What’s the next AI frontier?

Scientific discovery, automated. The question is, “how do you judge if an AI model is doing ‘good research’”. The answer comes from researchers who use AI models on real research. So, we're seeking one Fellow per STEM field who wants to build their field's first benchmark and put their name on it.

Core Responsibilities 

  • Validate results of the work produced by AI models’ paper reproduction and simulations.
  • Map out your field, including the most important research areas and the taxonomy within.
  • Set the standards for refining AI traces into benchmark samples; create benchmarks to evaluate frontier AI models against.

Authorship

  • Your benchmark will be published as a citable public release, including a paper and an open eval set. You’ll be named a lead author. 
  • Bonus: Get more citations for your papers.


Requirements

  • PhD students or postdocs in biology, medicine, neuroscience, materials science, chemistry, physics, mechanical engineering, or chemical engineering.
  • Have published a paper in your field; can judge a paper in your field and catch author’s errors.
  • Currently doing AI for Science research in their fields. (For example, you are using AI agents like Claude Code or Codex for 20+ hours per week in your research).

How it works

  • Can be remote or in-person.
  • You can set your hours. You can take a break if there’s an academic deadline.
  • $50-100/hr (up to $250K annualized). If you’re in the US, F1/visa-friendly structures are available.
  • You get to work directly with our founders and with researchers at frontier AI labs.

About OpenRefinery

Our mission is to advance scientific research via AI advancement. We scaled Turing from $0 to 2.2B (Series E). The cofounders Kai, Alex, and Char led ML programs Facebook AI Research (FAIR) and Zoox, exited two companies before, and built a $300M ARR data business together.

Backed by Zetta Ventures (invested in Kaggle and Domino Data Lab), Chair of the MIT Board of Trustees, Audacious Ventures (investors in Reflection AI and Decagon), and AI researchers from OpenAI, Anthropic, Google DeepMind, Thinking Machines, Nvidia, and the United Arab Emirates’ AI institute. 

We partner with researchers at Stanford, MIT, Caltech, ETH Zürich, UPenn, Berkeley, NUS, and more.

How to Apply

Email View email address on us.fitly.work with your field, how you use AI in your research, and one thing models consistently fails in your field. Résumé/CV is optional; you can just share your academic profile.

Vacancy posted 24 days ago
Similar jobs that could be interesting for youBased on the AI Scientist Benchmark Fellow in United States vacancy
  • Mercor is seeking PhD and Master’s scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will source material, craft executable problems, and build the grading criteria to measure model performance on challenging scenarios. You will... 
    Suggested

    Obsidian

    Miami, FL
    3 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks and participate in a new benchmark for scientific computing. You will author original, executable research problems that frontier models cannot solve, focusing on materials science and related subdomains... 
    Suggested
    Part time
    Immediate start

    Obsidian

    New York, NY
    3 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will author original, executable research problems that today's frontier models cannot solve. Researchers with depth in quantum and computational... 
    Suggested
    Part time
    Immediate start

    Obsidian

    Los Angeles, CA
    2 days ago
  • Mercor is seeking PhD and Master's scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will craft original, executable research problems that current frontier models cannot solve. Responsibilities include sourcing materials, writing... 
    Suggested
    Part time

    Obsidian

    New York, NY
    4 days ago
  •  ...leading enterprises orchestrate AI-powered work. Our vision is...  ...the world. As an AI research scientist, you'll be at the center of...  ...workflows Build novel evaluation benchmarks and methodologies that push...  ...Mentor and uplevel fellow researchers and engineers on... 
    Suggested
    Full time
    Work at office
    Local area
    Flexible hours

    Writer Corporation

    New York, NY
    4 days ago
  • SupportFinity™ is seeking a Mathematical Modeling Specialist to train AI models and measure their progress. This role includes working on mathematics problems to evaluate AI chatbot outputs for quality and performance. The position is remote and offers flexibility in choosing... 
    Remote job
    Hourly pay

    SupportFinity™

    New York, NY
    3 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. In collaboration with leading AI labs, you will source your own material, write scientific... 
    Part time
    Immediate start

    Obsidian

    Dallas, TX
    2 days ago
  • $233.3k - $385k

     ...together. Because at UKG, your work matters—and so do you.   About the Role: We are seeking a distinguished and customer-facing Fellow AI Engineer to define and lead UKG’s enterprise-wide Agentic AI strategy across the full UKG Product Suite. This role extends beyond... 
    Full time

    Ukg

    United States
    a month ago
  • $77.6k - $176k

    AI and ML Data ScientistThe Opportunity: As an Agentic AI Engineer and Data Scientist for military intelligence, you’re excited by the opportunity to design, develop, and...  ...prompt engineering, prompt evaluation, model benchmarking, fine-tuning, synthetic data generation,... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    McLean, VA
    5 days ago
  •  ...Role LexisNexis Legal & Professional is hiring an Applied AI Data Scientist to help shape the next generation of AI-powered legal...  ...APIs to prototype and refine AI capabilities. Design the Benchmark & Standard Design rigorous, domain-aware evaluation methodologies... 
    Full time
    Remote work
    Worldwide
    New York, NY
    19 days ago
  • $65 - $85 per hour

     ...is seeking a Computational Biologist for a remote role with a pay range of $65 to $85 per hour. In this position, you will design benchmark problems and analyze biological datasets. Applicants must possess an advanced degree in a relevant field and experience with bioinformatics... 
    Remote job
    Hourly pay
    Contract work
    10 hours per week
    Flexible hours

    Crossing Hurdles

    New York, NY
    6 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will join a collaboration with leading AI labs to publish original, executable research problems that frontier models cannot solve. Candidates should have PhD in chemistry-related... 
    Part time

    Mercor

    New York, NY
    2 days ago
  • $111k - $151k

     ...looking for 6+ years of experience in a Data Scientist, Applied Scientist, Machine Learning Engineer...  ...rigorous evaluation frameworks and metrics for AI/ML systems, covering offline and online testing, error analysis, benchmark development, and quality measurement. We... 
    Full time
    Worldwide

    LexisNexis

    New York, NY
    20 days ago
  • $60 - $90 per hour

     ...About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position... 
    Hourly pay
    Weekly pay
    Full time
    Contract work
    For contractors
    Summer work
    Remote work

    Mercor

    Remote
    a month ago
  • $219k - $246k

     ...technology company purpose-built to power AI-enabled precision health solutions that...  .... Description As an AI Applied Scientist, you will occupy a unique, highly impactful...  ...-judge\" grading rubrics to continuously benchmark agent safety, accuracy, factuality, and... 
    Full time
    Remote work

    Verily Life Sciences

    Eastern, KY
    2 days ago
  • $130k - $162.5k

     ...everyone. From advanced data analytics and AI to cybersecurity, we use innovative...  ...Group's enterprise AI team. We are AI/ML scientists and engineers with deep expertise in AI/ML...  ...LLMsCollaborating with the Optum AI team and fellow scientist and engineersDriving adoption... 
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    Bellevue, WA
    a month ago
  • $100k - $120k

     ...requirements.   What you’ll need to succeed as an Applied AI Scientist at XPO: Minimum qualifications: ~ Bachelor's degree or...  ...experience ~1 year of experience designing evaluation harnesses or benchmarks to rigorously assess model or agent performance against... 

    XPO Logistics

    Boston, MA
    a month ago
  • $86.32k - $154.96k

     ...ecosystem where experimental, clinical, and data scientists collaborate to accelerate discoveries,...  ...a skilled and collaborative Agentic AI Scientist to lead the design, development...  ...scientific contribution.) Publication benchmark: 1 to 2 first-author papers with a journal... 
    Full time
    Work experience placement
    Work at office

    St Jude Children's Research Hospital

    Memphis, TN
    24 days ago
  • $138.8k - $208.2k

     ...impact every day at Zebra.What We're Looking For:The Advanced AI Scientist is a highly experienced technical individual contributor...  ...model selection, prompt and context engineering, evaluation, benchmarking, guardrails, robust tool execution, and continuous improvement... 
    Full time
    Summer work
    Local area
    Flexible hours

    Zebra Technologies Corporation

    Austin, TX
    3 days ago
  • $150k - $200k

     ...Beach, FLThe Role: We are seeking an exceptionally talented AI Scientist to join the Artificial Intelligence team at WorldQuant. The successful...  ...pay ranges for all roles based on job function and level, benchmarked against similar stage organizations. When finalizing an offer... 
    Work experience placement
    Casual work
    Flexible hours

    WorldQuant

    New York, NY
    12 days ago
  • $219k - $246k

     ...technology company purpose-built to power AI-enabled precision health solutions that...  ...precise. Description As an AI Applied Scientist, you will occupy a unique, highly...  ...a-judge" grading rubrics to continuously benchmark agent safety, accuracy, factuality, and compliance... 
    Full time
    Remote work

    Verily Health

    San Bruno, CA
    1 day ago
  • $105k - $145k

     ...that ensure our predictive and generative AI systems areaccurate, reliable, safe, and...  ...across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data...  ...detect performance drift over time. Develop benchmark datasets, challenge sets, and scenario-... 
    Local area

    Steampunk.com

    McLean, VA
    5 days ago
  • $77.6k - $176k

    Applied AI Health ScientistThe Opportunity:To achieve an organization’s mission, leaders...  ...we need you, an experienced Applied AI Scientist for Health who can contribute expertise...  ...and curation, model training and benchmarking, and deployment and validation in real-world... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Immediate start
    Remote work
    Flexible hours
    Shift work

    Booz Allen Hamilton

    Washington DC
    5 days ago
  • $184k - $287.5k

     ...amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU...  ...searching for a highly motivated engineer to lead performance benchmarking and optimization efforts for our data center products. You will... 
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $50 per hour

    A leading AI research organization is seeking PhDs in Chemistry or related fields for a remote contract. The role involves designing...  ...developing solutions, evaluating AI outputs, and refining benchmarks. The pay rate is $50+/hour, depending on expertise. This flexibly... 
    Remote job
    Contract work

    Turing

    San Francisco, CA
    2 days ago
  • $109k - $174.8k

     ...step of the way. Learn more at Position Summary Senior Scientist - Data (AI) Scientist, Translational Safety to join the Data, Data Science...  ...scientific validation : assess biological plausibility, benchmark model performance, produce explainability and validation... 
    Work at office
    Local area
    Immediate start

    J&J Family of Companies

    Spring House, PA
    6 days ago
  • $160k - $220k

     ...enjoy multi-year runway.About the RoleWe’re looking for an AI Research Scientist to advance the methodological frontier of AI in healthcare....  ...understand that healthcare validation standards are higher than benchmark culture alone, and you are energized by the opportunity to... 
    Temporary work
    Work at office
    Monday to Friday
    Monday to Thursday

    Sprinter Health

    San Francisco, CA
    a month ago
  • $185.8k - $243.8k

     ...AI Research Scientist - AI BioDesign The Allen Institute accelerates science for a healthier world through large-scale research designed to...  ...to the broader scientific community through publications, benchmarks, and reusable methods. The role operates in high-ambiguity,... 
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Work visa
    Relocation package

    Allen Institute

    Seattle, WA
    more than 2 months ago
  • $170.5k - $315.49k

     ...that will power the next generation of physical AI systems. We are seeking a Neuromorphic AI Research Scientist to advance state-of-the-art neuromorphic...  ...software artifacts including APIs, kernels, and benchmarking tools that enable collaborative innovation Prototype... 
    Internship
    Work at office
    Local area
    Immediate start
    Shift work

    Intel

    Hillsboro, OR
    1 day ago
  • $192k - $304.75k

     ...approach to accelerated computing. We're looking for a passionate AI research scientist with deep quantum computing expertise to path-find the...  ...and develop open AI models, curated datasets, and rigorous benchmarks that advance the state of the art and empower the broader... 
    Full time

    NVIDIA

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Scientist Benchmark Fellow. Be the first to apply!