AI Scientist Benchmark Fellow
$50 - $100 per hourOpenrefinery.ai
What’s the next AI frontier?
Scientific discovery, automated. The question is, “how do you judge if an AI model is doing ‘good research’”. The answer comes from researchers who use AI models on real research. So, we're seeking one Fellow per STEM field who wants to build their field's first benchmark and put their name on it.
Core Responsibilities
- Validate results of the work produced by AI models’ paper reproduction and simulations.
- Map out your field, including the most important research areas and the taxonomy within.
- Set the standards for refining AI traces into benchmark samples; create benchmarks to evaluate frontier AI models against.
Authorship
- Your benchmark will be published as a citable public release, including a paper and an open eval set. You’ll be named a lead author.
- Bonus: Get more citations for your papers.
Requirements
- PhD students or postdocs in biology, medicine, neuroscience, materials science, chemistry, physics, mechanical engineering, or chemical engineering.
- Have published a paper in your field; can judge a paper in your field and catch author’s errors.
- Currently doing AI for Science research in their fields. (For example, you are using AI agents like Claude Code or Codex for 20+ hours per week in your research).
How it works
- Can be remote or in-person.
- You can set your hours. You can take a break if there’s an academic deadline.
- $50-100/hr (up to $250K annualized). If you’re in the US, F1/visa-friendly structures are available.
- You get to work directly with our founders and with researchers at frontier AI labs.
About OpenRefinery
Our mission is to advance scientific research via AI advancement. We scaled Turing from $0 to 2.2B (Series E). The cofounders Kai, Alex, and Char led ML programs Facebook AI Research (FAIR) and Zoox, exited two companies before, and built a $300M ARR data business together.
Backed by Zetta Ventures (invested in Kaggle and Domino Data Lab), Chair of the MIT Board of Trustees, Audacious Ventures (investors in Reflection AI and Decagon), and AI researchers from OpenAI, Anthropic, Google DeepMind, Thinking Machines, Nvidia, and the United Arab Emirates’ AI institute.
We partner with researchers at Stanford, MIT, Caltech, ETH Zürich, UPenn, Berkeley, NUS, and more.
How to Apply
Email View email address on us.fitly.work with your field, how you use AI in your research, and one thing models consistently fails in your field. Résumé/CV is optional; you can just share your academic profile.
- Snorkel AI in San Francisco is searching for a Research Scientist to lead the development of datasets and benchmarks for AI models. This customer-facing role involves working with academic partners and cross-functional teams to ensure impactful research outcomes. The ideal...Suggested
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will author original, executable research problems that today’s frontier models cannot solve, sourced from published papers, datasets,...Suggested
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering with leading AI labs. You will author original, executable research problems that frontier models cannot solve. Engagement lasts...SuggestedPart timeImmediate start
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks and participate in a new benchmark for scientific computing. You will author original, executable research problems that frontier models cannot solve, focusing on materials science and related subdomains...SuggestedPart timeImmediate start
$50 per hour
...in Biology or related fields for remote roles. You will design biology challenges for AI systems, evaluate their reasoning, and collaborate with top researchers to improve benchmarks. The position offers flexible hours at a competitive rate of $50+/hour, allowing access...SuggestedRemote jobHourly payFlexible hours- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, collaborating with leading AI labs. You will author original, executable research problems that frontier models cannot solve. Domains require...Part time
- Mercor is seeking PhD and Master’s scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will source material, craft executable problems, and build the grading criteria to measure model performance on challenging scenarios. You will...
$234.3k - $349k
...leading enterprises orchestrate AI-powered work. Our vision is... ...the world. As an AI research scientist, you'll be at the center of... ...workflowsBuild novel evaluation benchmarks and methodologies that push... ...communityMentor and uplevel fellow researchers and engineers on...Full timeWork at officeLocal area$50 per hour
A leading AI research accelerator is seeking remote PhD holders in Biology or related fields. As a contractor, you will design advanced biology questions, evaluate AI outputs, and collaborate with researchers on cutting-edge projects with top AI labs. This role offers...Remote jobFor contractorsFlexible hours$233.3k - $385k
...together. Because at UKG, your work matters—and so do you. About the Role: We are seeking a distinguished and customer-facing Fellow AI Engineer to define and lead UKG’s enterprise-wide Agentic AI strategy across the full UKG Product Suite. This role extends beyond...Full time$119k - $286k
..., Data Science and IT to prototype, build, and scale practical AI-driven solutions that improve quality, cost, cycle time, and engineering... ...skills, retrieval-augmented generation (RAG and Graph-RAG), benchmarking, and quality evaluation. Machine Learning Production: Develop...Full timeLocal areaImmediate start- ...the RoleLexisNexis Legal & Professional is hiring an Applied AI Data Scientist to help shape the next generation of AI-powered legal... ...Claude APIs to prototype and refine AI capabilities.Design the Benchmark & Standard Design rigorous, domain-aware evaluation methodologies...Full timeRemote workWorldwide
$77.6k - $176k
AI and ML Data ScientistThe Opportunity: As an Agentic AI Engineer and Data Scientist for military intelligence, you’re excited by the opportunity to design, develop, and... ...prompt engineering, prompt evaluation, model benchmarking, fine-tuning, synthetic data generation,...Full timeContract workPart timeWork at officeLocal areaRemote work- Mercor is seeking AI/ML specialists for a benchmark dataset project evaluating AI models on visual document understanding and instruction-following in Education. This is a remote, independent contractor role requiring ~15-20 hours per week and can be performed from the...Remote jobFor contractors
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. The role emphasizes material sourcing, prompt design, and robust grading criteria against...Part timeImmediate start
- Mercor is seeking a talented AI/education domain specialist to support a benchmark dataset project evaluating AI models on visual document understanding and instruction-following in the Education domain. The role is fully remote, roughly 15-20 hours per week, and open to...Remote jobFor contractors
$130k - $162.5k
...everyone. From advanced data analytics and AI to cybersecurity, we use innovative... ...Group's enterprise AI team. We are AI/ML scientists and engineers with deep expertise in AI/ML... ...LLMsCollaborating with the Optum AI team and fellow scientist and engineersDriving adoption...Minimum wageFull timeWork experience placementWork at officeLocal areaRemote work$150k - $200k
...Beach, FLThe Role: We are seeking an exceptionally talented AI Scientist to join the Artificial Intelligence team at WorldQuant. The successful... ...pay ranges for all roles based on job function and level, benchmarked against similar stage organizations. When finalizing an offer...Work experience placementCasual workFlexible hours$105k - $145k
OverviewWe are looking for an AI Evaluation Scientistto design and execute evaluation... ...across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data... ...detect performance drift over time. Develop benchmark datasets, challenge sets, and scenario-...Local area$86.32k - $154.96k
...ecosystem where experimental, clinical, and data scientists collaborate to accelerate discoveries,... ...a skilled and collaborative Agentic AI Scientist to lead the design, development... ...of scientific contribution.)Publication benchmark: 1 to 2 first-author papers with a journal...Full timeWork experience placementWork at office$126k - $234k
...oncology drug discovery, computational biology, AI/ML, and data engineering. We are seeking an enthusiastic AI/ML scientist with deep expertise in digital pathology and... ....Contribute to evaluation, validation, and benchmarking of image analysis algorithms and workflows.Perform...Full timeLocal areaRelocationFlexible hours$65 - $85 per hour
...is seeking a Computational Biologist for a remote role with a pay range of $65 to $85 per hour. In this position, you will design benchmark problems and analyze biological datasets. Applicants must possess an advanced degree in a relevant field and experience with bioinformatics...Remote jobHourly payContract work10 hours per weekFlexible hours- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will join a collaboration with leading AI labs to publish original, executable research problems that frontier models cannot solve. Candidates should have PhD in chemistry-related...Part time
$184k - $287.5k
...amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU... ...searching for a highly motivated engineer to lead performance benchmarking and optimization efforts for our data center products. You will...Full time$60 - $90 per hour
...About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position...Hourly payWeekly payFull timeContract workFor contractorsSummer workRemote work$225k - $280k
...Wizard Wizard is the top-performing AI Shopping Agent, delivering the best products... ...Role We're looking for an Applied Scientist to own how we measure, understand, and... ...Build and maintain evaluation datasets, benchmarks, and scoring frameworks Improve the LLM...Remote workFlexible hours$179k - $217k
...Chief Science Officer, we are looking for a talented Senior AI Data Scientist to join our team and help shape the future of AI. ~Design... ...designing A/B tests, offline evaluations, and ML benchmarking frameworks. ~Familiarity with vector databases, embedding...Full timeWork at officeRemote work$69.4k - $158k
AI and ML Research ScientistThe Opportunity:As an analytics professional, you’re excited... ...Army applications.As an advanced data scientist or researcher on our Maritime and Ground... ...establishing evaluation protocols such as benchmarks, ablations, and error analysesKnowledge of...Full timeContract workPart timeWork at officeLocal areaRemote work- ...intelligence by establishing the first distributed AI infrastructure dedicated to personalized... ...of AI.About the Role:The AI Research Scientist will contribute to webAI’s development of... ...for distributed AIDevelop and evaluate benchmarks, datasets, and experimental frameworks to...Live outWork at officeLocal area
$160k - $220k
...enjoy multi-year runway.About the RoleWe’re looking for an AI Research Scientist to advance the methodological frontier of AI in healthcare.... ...understand that healthcare validation standards are higher than benchmark culture alone, and you are energized by the opportunity to...Temporary workWork at officeMonday to FridayMonday to Thursday
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Scientist Benchmark Fellow. Be the first to apply!



