AI Scientist Benchmark Fellow
$50 - $100 per hourOpenrefinery.ai
What’s the next AI frontier?
Scientific discovery, automated. The question is, “how do you judge if an AI model is doing ‘good research’”. The answer comes from researchers who use AI models on real research. So, we're seeking one Fellow per STEM field who wants to build their field's first benchmark and put their name on it.
Core Responsibilities
- Validate results of the work produced by AI models’ paper reproduction and simulations.
- Map out your field, including the most important research areas and the taxonomy within.
- Set the standards for refining AI traces into benchmark samples; create benchmarks to evaluate frontier AI models against.
Authorship
- Your benchmark will be published as a citable public release, including a paper and an open eval set. You’ll be named a lead author.
- Bonus: Get more citations for your papers.
Requirements
- PhD students or postdocs in biology, medicine, neuroscience, materials science, chemistry, physics, mechanical engineering, or chemical engineering.
- Have published a paper in your field; can judge a paper in your field and catch author’s errors.
- Currently doing AI for Science research in their fields. (For example, you are using AI agents like Claude Code or Codex for 20+ hours per week in your research).
How it works
- Can be remote or in-person.
- You can set your hours. You can take a break if there’s an academic deadline.
- $50-100/hr (up to $250K annualized). If you’re in the US, F1/visa-friendly structures are available.
- You get to work directly with our founders and with researchers at frontier AI labs.
About OpenRefinery
Our mission is to advance scientific research via AI advancement. We scaled Turing from $0 to 2.2B (Series E). The cofounders Kai, Alex, and Char led ML programs Facebook AI Research (FAIR) and Zoox, exited two companies before, and built a $300M ARR data business together.
Backed by Zetta Ventures (invested in Kaggle and Domino Data Lab), Chair of the MIT Board of Trustees, Audacious Ventures (investors in Reflection AI and Decagon), and AI researchers from OpenAI, Anthropic, Google DeepMind, Thinking Machines, Nvidia, and the United Arab Emirates’ AI institute.
We partner with researchers at Stanford, MIT, Caltech, ETH Zürich, UPenn, Berkeley, NUS, and more.
How to Apply
Email View email address on us.fitly.work with your field, how you use AI in your research, and one thing models consistently fails in your field. Résumé/CV is optional; you can just share your academic profile.
- Mercor is seeking PhD and Master’s scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will source material, craft executable problems, and build the grading criteria to measure model performance on challenging scenarios. You will...Suggested
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks and participate in a new benchmark for scientific computing. You will author original, executable research problems that frontier models cannot solve, focusing on materials science and related subdomains...SuggestedPart timeImmediate start
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will author original, executable research problems that today's frontier models cannot solve. Researchers with depth in quantum and computational...SuggestedPart timeImmediate start
- Mercor is seeking PhD and Master's scientists to author AI evaluation tasks for a new benchmark in scientific computing. You will craft original, executable research problems that current frontier models cannot solve. Responsibilities include sourcing materials, writing...SuggestedPart time
- ...leading enterprises orchestrate AI-powered work. Our vision is... ...the world. As an AI research scientist, you'll be at the center of... ...workflows Build novel evaluation benchmarks and methodologies that push... ...Mentor and uplevel fellow researchers and engineers on...SuggestedFull timeWork at officeLocal areaFlexible hours
- SupportFinity™ is seeking a Mathematical Modeling Specialist to train AI models and measure their progress. This role includes working on mathematics problems to evaluate AI chatbot outputs for quality and performance. The position is remote and offers flexibility in choosing...Remote jobHourly pay
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. In collaboration with leading AI labs, you will source your own material, write scientific...Part timeImmediate start
$233.3k - $385k
...together. Because at UKG, your work matters—and so do you. About the Role: We are seeking a distinguished and customer-facing Fellow AI Engineer to define and lead UKG’s enterprise-wide Agentic AI strategy across the full UKG Product Suite. This role extends beyond...Full time$77.6k - $176k
AI and ML Data ScientistThe Opportunity: As an Agentic AI Engineer and Data Scientist for military intelligence, you’re excited by the opportunity to design, develop, and... ...prompt engineering, prompt evaluation, model benchmarking, fine-tuning, synthetic data generation,...Full timeContract workPart timeWork at officeLocal areaRemote work- ...Role LexisNexis Legal & Professional is hiring an Applied AI Data Scientist to help shape the next generation of AI-powered legal... ...APIs to prototype and refine AI capabilities. Design the Benchmark & Standard Design rigorous, domain-aware evaluation methodologies...Full timeRemote workWorldwide
$65 - $85 per hour
...is seeking a Computational Biologist for a remote role with a pay range of $65 to $85 per hour. In this position, you will design benchmark problems and analyze biological datasets. Applicants must possess an advanced degree in a relevant field and experience with bioinformatics...Remote jobHourly payContract work10 hours per weekFlexible hours- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will join a collaboration with leading AI labs to publish original, executable research problems that frontier models cannot solve. Candidates should have PhD in chemistry-related...Part time
$111k - $151k
...looking for 6+ years of experience in a Data Scientist, Applied Scientist, Machine Learning Engineer... ...rigorous evaluation frameworks and metrics for AI/ML systems, covering offline and online testing, error analysis, benchmark development, and quality measurement. We...Full timeWorldwide$60 - $90 per hour
...About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position...Hourly payWeekly payFull timeContract workFor contractorsSummer workRemote work$219k - $246k
...technology company purpose-built to power AI-enabled precision health solutions that... .... Description As an AI Applied Scientist, you will occupy a unique, highly impactful... ...-judge\" grading rubrics to continuously benchmark agent safety, accuracy, factuality, and...Full timeRemote work$130k - $162.5k
...everyone. From advanced data analytics and AI to cybersecurity, we use innovative... ...Group's enterprise AI team. We are AI/ML scientists and engineers with deep expertise in AI/ML... ...LLMsCollaborating with the Optum AI team and fellow scientist and engineersDriving adoption...Minimum wageFull timeWork experience placementWork at officeLocal areaRemote work$100k - $120k
...requirements. What you’ll need to succeed as an Applied AI Scientist at XPO: Minimum qualifications: ~ Bachelor's degree or... ...experience ~1 year of experience designing evaluation harnesses or benchmarks to rigorously assess model or agent performance against...$86.32k - $154.96k
...ecosystem where experimental, clinical, and data scientists collaborate to accelerate discoveries,... ...a skilled and collaborative Agentic AI Scientist to lead the design, development... ...scientific contribution.) Publication benchmark: 1 to 2 first-author papers with a journal...Full timeWork experience placementWork at office$138.8k - $208.2k
...impact every day at Zebra.What We're Looking For:The Advanced AI Scientist is a highly experienced technical individual contributor... ...model selection, prompt and context engineering, evaluation, benchmarking, guardrails, robust tool execution, and continuous improvement...Full timeSummer workLocal areaFlexible hours$150k - $200k
...Beach, FLThe Role: We are seeking an exceptionally talented AI Scientist to join the Artificial Intelligence team at WorldQuant. The successful... ...pay ranges for all roles based on job function and level, benchmarked against similar stage organizations. When finalizing an offer...Work experience placementCasual workFlexible hours$219k - $246k
...technology company purpose-built to power AI-enabled precision health solutions that... ...precise. Description As an AI Applied Scientist, you will occupy a unique, highly... ...a-judge" grading rubrics to continuously benchmark agent safety, accuracy, factuality, and compliance...Full timeRemote work$105k - $145k
...that ensure our predictive and generative AI systems areaccurate, reliable, safe, and... ...across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data... ...detect performance drift over time. Develop benchmark datasets, challenge sets, and scenario-...Local area$77.6k - $176k
Applied AI Health ScientistThe Opportunity:To achieve an organization’s mission, leaders... ...we need you, an experienced Applied AI Scientist for Health who can contribute expertise... ...and curation, model training and benchmarking, and deployment and validation in real-world...Full timeContract workPart timeWork at officeLocal areaImmediate startRemote workFlexible hoursShift work$184k - $287.5k
...amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU... ...searching for a highly motivated engineer to lead performance benchmarking and optimization efforts for our data center products. You will...Full time$50 per hour
A leading AI research organization is seeking PhDs in Chemistry or related fields for a remote contract. The role involves designing... ...developing solutions, evaluating AI outputs, and refining benchmarks. The pay rate is $50+/hour, depending on expertise. This flexibly...Remote jobContract work$109k - $174.8k
...step of the way. Learn more at Position Summary Senior Scientist - Data (AI) Scientist, Translational Safety to join the Data, Data Science... ...scientific validation : assess biological plausibility, benchmark model performance, produce explainability and validation...Work at officeLocal areaImmediate start$160k - $220k
...enjoy multi-year runway.About the RoleWe’re looking for an AI Research Scientist to advance the methodological frontier of AI in healthcare.... ...understand that healthcare validation standards are higher than benchmark culture alone, and you are energized by the opportunity to...Temporary workWork at officeMonday to FridayMonday to Thursday$185.8k - $243.8k
...AI Research Scientist - AI BioDesign The Allen Institute accelerates science for a healthier world through large-scale research designed to... ...to the broader scientific community through publications, benchmarks, and reusable methods. The role operates in high-ambiguity,...Work at officeLocal areaRemote workVisa sponsorshipWork visaRelocation package$170.5k - $315.49k
...that will power the next generation of physical AI systems. We are seeking a Neuromorphic AI Research Scientist to advance state-of-the-art neuromorphic... ...software artifacts including APIs, kernels, and benchmarking tools that enable collaborative innovation Prototype...InternshipWork at officeLocal areaImmediate startShift work$192k - $304.75k
...approach to accelerated computing. We're looking for a passionate AI research scientist with deep quantum computing expertise to path-find the... ...and develop open AI models, curated datasets, and rigorous benchmarks that advance the state of the art and empower the broader...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Scientist Benchmark Fellow. Be the first to apply!



