AI Scientist Benchmark Fellow
$250kOpenrefinery.ai
What’s the next AI frontier?
Scientific discovery, automated. The question is, “how do you judge if an AI model is doing ‘good research’”. The answer comes from researchers who use AI models on real research. So, we're seeking one Fellow per STEM field who wants to build their field's first benchmark and put their name on it.
Core Responsibilities
- Validate results of the work produced by AI models’ paper reproduction and simulations.
- Map out your field, including the most important research areas and the taxonomy within.
- Set the standards for refining AI traces into benchmark samples; create benchmarks to evaluate frontier AI models against.
Authorship
- Your benchmark will be published as a citable public release, including a paper and an open eval set. You’ll be named a lead author.
- Bonus: Get more citations for your papers.
Requirements
- PhD students or postdocs in biology, medicine, neuroscience, materials science, chemistry, physics, mechanical engineering, or chemical engineering.
- Have published a paper in your field; can judge a paper in your field and catch author’s errors.
- Currently doing AI for Science research in their fields. (For example, you are using AI agents like Claude Code or Codex for 20+ hours per week in your research).
How it works
- Can be remote or in-person.
- You can set your hours. You can take a break if there’s an academic deadline.
- $50-100/hr (up to $250K annualized). If you’re in the US, F1/visa-friendly structures are available.
- You get to work directly with our founders and with researchers at frontier AI labs.
About OpenRefinery
Our mission is to advance scientific research via AI advancement. We scaled Turing from $0 to 2.2B (Series E). The cofounders Kai, Alex, and Char led ML programs Facebook AI Research (FAIR) and Zoox, exited two companies before, and built a $300M ARR data business together.
Backed by Zetta Ventures (invested in Kaggle and Domino Data Lab), Chair of the MIT Board of Trustees, Audacious Ventures (investors in Reflection AI and Decagon), and AI researchers from OpenAI, Anthropic, Google DeepMind, Thinking Machines, Nvidia, and the United Arab Emirates’ AI institute.
We partner with researchers at Stanford, MIT, Caltech, ETH Zürich, UPenn, Berkeley, NUS, and more.
How to Apply
Email View email address on nature.com with your field, how you use AI in your research, and one thing models consistently fails in your field. Résumé/CV is optional; you can just share your academic profile.
- Snorkel AI in San Francisco is searching for a Research Scientist to lead the development of datasets and benchmarks for AI models. This customer-facing role involves working with academic partners and cross-functional teams to ensure impactful research outcomes. The ideal...Suggested
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will author original, executable research problems that today’s frontier models cannot solve, sourced from published papers, datasets,...Suggested
$50 per hour
...in Biology or related fields for remote roles. You will design biology challenges for AI systems, evaluate their reasoning, and collaborate with top researchers to improve benchmarks. The position offers flexible hours at a competitive rate of $50+/hour, allowing access...SuggestedRemote jobHourly payFlexible hours- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks and participate in a new benchmark for scientific computing. You will author original, executable research problems that frontier models cannot solve, focusing on materials science and related subdomains...SuggestedPart timeImmediate start
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, collaborating with leading AI labs. You will author original, executable research problems that frontier models cannot solve. Domains require...SuggestedPart time
- Mercor is seeking PhD and Master’s scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will source material, craft executable problems, and build the grading criteria to measure model performance on challenging scenarios. You will...
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering with leading AI labs. You will author original, executable research problems that frontier models cannot solve. Engagement lasts...Part timeImmediate start
$234.3k - $349k
...leading enterprises orchestrate AI-powered work. Our vision is... ...the world. As an AI research scientist, you'll be at the center of... ...workflowsBuild novel evaluation benchmarks and methodologies that push... ...communityMentor and uplevel fellow researchers and engineers on...Full timeWork at officeLocal area- Mercor is seeking a remote contractor for a benchmark dataset project focused on evaluating AI models in the Education domain. The role requires 15-20 hours per week and offers fully remote collaboration across US/Canada. You will work on tasks with clear ground-truth...Remote jobWeekly payPart timeFor contractors
$50 per hour
A leading AI research accelerator is seeking remote PhD holders in Biology or related fields. As a contractor, you will design advanced biology questions, evaluate AI outputs, and collaborate with researchers on cutting-edge projects with top AI labs. This role offers...Remote jobFor contractorsFlexible hours- A leading AI research organization is looking for remote candidates with PhDs in Biology or related fields to fine-tune AI language models. In this role, you will design biology questions and assess AI's problem-solving capabilities while collaborating with top researchers...Remote jobContract work
$233.3k - $385k
...together. Because at UKG, your work matters—and so do you. About the Role: We are seeking a distinguished and customer-facing Fellow AI Engineer to define and lead UKG’s enterprise-wide Agentic AI strategy across the full UKG Product Suite. This role extends beyond...Full time$77.6k - $176k
AI and ML Data ScientistThe Opportunity: As an Agentic AI Engineer and Data Scientist for military intelligence, you’re excited by the opportunity to design, develop, and... ...prompt engineering, prompt evaluation, model benchmarking, fine-tuning, synthetic data generation,...Full timeContract workPart timeWork at officeLocal areaRemote work- ...the RoleLexisNexis Legal & Professional is hiring an Applied AI Data Scientist to help shape the next generation of AI-powered legal... ...Claude APIs to prototype and refine AI capabilities.Design the Benchmark & Standard Design rigorous, domain-aware evaluation methodologies...Full timeRemote workWorldwide
- Mercor is seeking a talented AI/education domain specialist to support a benchmark dataset project evaluating AI models on visual document understanding and instruction-following in the Education domain. The role is fully remote, roughly 15-20 hours per week, and open to...Remote jobFor contractors
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. The role emphasizes material sourcing, prompt design, and robust grading criteria against...Part timeImmediate start
- Mercor is seeking AI/ML specialists for a benchmark dataset project evaluating AI models on visual document understanding and instruction-following in Education. This is a remote, independent contractor role requiring ~15-20 hours per week and can be performed from the...Remote jobFor contractors
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will join a collaboration with leading AI labs to publish original, executable research problems that frontier models cannot solve. Candidates should have PhD in chemistry-related...Part time
$65 - $85 per hour
...is seeking a Computational Biologist for a remote role with a pay range of $65 to $85 per hour. In this position, you will design benchmark problems and analyze biological datasets. Applicants must possess an advanced degree in a relevant field and experience with bioinformatics...Remote jobHourly payContract work10 hours per weekFlexible hours$130k - $162.5k
...everyone. From advanced data analytics and AI to cybersecurity, we use innovative... ...Group's enterprise AI team. We are AI/ML scientists and engineers with deep expertise in AI/ML... ...LLMsCollaborating with the Optum AI team and fellow scientist and engineersDriving adoption...Minimum wageFull timeWork experience placementWork at officeLocal areaRemote work$150k - $200k
...Beach, FLThe Role: We are seeking an exceptionally talented AI Scientist to join the Artificial Intelligence team at WorldQuant. The successful... ...pay ranges for all roles based on job function and level, benchmarked against similar stage organizations. When finalizing an offer...Work experience placementCasual workFlexible hours$105k - $145k
OverviewWe are looking for an AI Evaluation Scientistto design and execute evaluation... ...across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data... ...detect performance drift over time. Develop benchmark datasets, challenge sets, and scenario-...Local area$126k - $234k
...oncology drug discovery, computational biology, AI/ML, and data engineering. We are seeking an enthusiastic AI/ML scientist with deep expertise in digital pathology and... ....Contribute to evaluation, validation, and benchmarking of image analysis algorithms and workflows.Perform...Full timeLocal areaRelocationFlexible hours$184k - $287.5k
...amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU... ...searching for a highly motivated engineer to lead performance benchmarking and optimization efforts for our data center products. You will...Full time$60 - $90 per hour
...About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position...Hourly payWeekly payFull timeContract workFor contractorsSummer workRemote work$225k - $280k
...Wizard Wizard is the top-performing AI Shopping Agent, delivering the best products... ...Role We're looking for an Applied Scientist to own how we measure, understand, and... ...Build and maintain evaluation datasets, benchmarks, and scoring frameworks Improve the LLM...Remote workFlexible hours$179k - $217k
...Chief Science Officer, we are looking for a talented Senior AI Data Scientist to join our team and help shape the future of AI. ~Design... ...designing A/B tests, offline evaluations, and ML benchmarking frameworks. ~Familiarity with vector databases, embedding...Full timeWork at officeRemote work$77.02k - $117.58k
...named one of the Tampa Bay Times’ Top Workplaces.SummaryThe AI Data Scientist - Precision Oncology develops and applies advanced analytics... ...healthcare outcomes.ResponsibilitiesDevelop, evaluate, and benchmark machine learning models for prediction, prognosis, treatment...$192k - $304.75k
...approach to accelerated computing. We're looking for a passionate AI research scientist with deep quantum computing expertise to path-find the... ...and develop open AI models, curated datasets, and rigorous benchmarks that advance the state of the art and empower the broader...Full time$69.4k - $158k
AI and ML Research ScientistThe Opportunity:As an analytics professional, you’re excited... ...Army applications.As an advanced data scientist or researcher on our Maritime and Ground... ...establishing evaluation protocols such as benchmarks, ablations, and error analysesKnowledge of...Full timeContract workPart timeWork at officeLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Scientist Benchmark Fellow. Be the first to apply!



