STEM PhD Researcher for AI Model Evaluation
$70 - $100 per hourSaidGig
Join a leading AI lab''s cutting-edge research team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced LLMs. Overview
Advanced STEM Researchers and PhD-level subject-matter experts (SMEs) are needed to contribute to a project supporting a frontier-model evaluation effort focused on rigorous scientific and technical reasoning. The AI lab is building next-generation models capable of solving complex, research-grade problems across the sciences and requires deep domain expertise to design, solve, and evaluate the challenging tasks that train and benchmark these systems.
This is a W-2 employment position with Cincinnatus LLC, requiring a commitment of 40 hours per week during weekdays. This position will be placed at a leading AI Lab as part of their extended workforce.
Key Responsibilities- Guide research teams to close knowledge gaps in STEM domains by surfacing edge cases, ambiguities, and frontier problems where current models underperform.
- Design challenging, rigorous domain tasks and write accurate, well-reasoned solutions that demonstrate expert-level scientific and technical reasoning.
- Evaluate tasks and solutions produced by AI agents and other contributors, providing clear written technical feedback grounded in domain expertise.
- Develop evaluation frameworks and rubrics for assessing scientific reasoning quality across STEM domains.
- Collaborate with other subject matter experts to ensure consistency and accuracy in training data.
- PhD (completed, enrolled, or equivalent research track) in Physics, Chemistry, Biology, Mathematics, Statistics, Computer Science, Electrical Engineering, Mechanical Engineering, Civil Engineering, Materials Science, or another STEM discipline.
- 3+ years of research, academic, or industry experience in their primary STEM domain.
- Demonstrated technical expertise in at least one domain: computational modeling, laboratory methods, data analysis, statistical inference, programming, or equivalent scientific methods.
- Ability to commit to 40 hours per week during weekdays for the duration of the engagement.
- Prior experience with data annotation, labeling, evaluation, or human feedback collection is a strong plus.
- Experience with LLMs, AI systems, or agentic workflows; familiarity with agentic frameworks is a plus.
- Strong written communication skills; ability to explain complex scientific or technical concepts clearly in writing.
Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives.
Equal Employment Opportunity:Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.
$60 - $90 per hour
...Help build next-generation agentic evaluation benchmarks for frontier AI models by turning rigorous scientific... ...results at the level of a careful researcher. Define scoring and success criteria... .... Qualifications MSc or PhD in a STEM field, or in a computational...For phdHourly payFreelanceImmediate startRemote work$20 - $55 per hour
...conduct literature and document research, prepare and review complex... ...that help train and evaluate next-generation AI systems. You will contribute... ...healthcare, policy, finance, and STEM engineering. Key Responsibilities... .... Master''s degree, PhD, or JD from a recognized...For phdHourly payFor contractorsRemote work$400k
...Define the frontier of AI-powered legal reasoning by building rigorous evaluation frameworks and benchmarks... ...combines applied legal research, dataset curation, and... ...in large language models and agentic workflows.... ...such as JD, LLM, SJD, PhD in Law, or equivalent....For phdFull timeRemote work$60 - $90 per hour
...next-generation agentic evaluation benchmarks for frontier AI models by acting as a ground-truth... ...and execute realistic, research-style analysis tasks that... ...Qualifications MSc or PhD in statistics, data science... ...or another quantitative STEM field, or equivalent practical...For phdHourly payFull timeFreelanceRemote work$80 - $110 per hour
...Overview Contribute expert, research-level chemistry judgment to train and evaluate next-generation AI systems that reason... ..., and identify model failure modes at the level... ...an experienced PhD chemist would spot immediately... ...or a related STEM field. Receipt of a...For phdHourly payPart timeImmediate startRemote work$75 per hour
...Role Overview AI and Machine Learning Researchers apply deep expertise in machine... ...science research to evaluate AI-generated outputs... ...data that improves model understanding of advanced... ...Qualifications PhD in Computer Science,... ...that requirement. STEM OPT is not supported...For phdHourly payFull timeContract workPart timeRemote workFlexible hours$400k
...Join a dynamic research team as a Member of Technical Staff (MTS) focused... ...a pivotal role in advancing AI systems designed to enhance... ...emphasizes the development of robust evaluation frameworks that assess medical... ...healthcare field (MD, DO, PhD, MPH, PharmD, etc.). Deep...For phdFull timeRemote work$40 per hour
...Chemist to join their team remotely. In this role, you will train AI models by evaluating their performance on complex chemistry questions. The... ...knowledge, with native or bilingual proficiency in English. A Master's or PhD is preferred but not mandatory. #J-18808-LjbffrFor phdHourly payRemote work$80 - $135 per hour
...). The role involves solving frontier research-level physics problems end-to-end, auditing... ...solutions so that large language models can be evaluated on rigorous physics reasoning. Physics... ...cases Qualifications ~ Solver: PhD or postdoc in the relevant subfield,...For phdHourly payRemote work10 hours per week$40 per hour
A data annotation company seeks an R&D Biologist to enhance AI models by evaluating chatbot outputs related to complex biology queries. This remote... ...to detail, and fluent English skills. A Master's or PhD in Biology or related fields is preferred, although not mandatory...For phdHourly payRemote work$40 per hour
A leading AI training firm in the United States is seeking an R&D Biologist... .... In this remote position, you will evaluate AI chatbots and enhance their models while ensuring the biological accuracy... ..., with a preference for Master's or PhD qualifications. You can choose your...For phdHourly payRemote work$70 - $90 per hour
...to a cutting-edge project with a leading AI research lab, producing high-quality, hard problems... ...of state-of-the-art large language models. The work emphasizes rigorous subject-matter... ...clearly. Qualifications Completed PhD in Computer Science. Undergraduate degree...For phdHourly payRemote work$40 per hour
A research organization is seeking a Biology Research Scientist to evaluate AI chatbots' performance and improve their quality. The role requires an expert level of biology... ...at $40+. Candidates with a Master's or PhD in a related field are preferred. Only applicants...For phdHourly payContract workRemote work- ...Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities include fact-checking AI technical...For phdHourly payRemote workFlexible hours
- ...Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content... ...and validation, requiring candidates with a BS, MS, or PhD in relevant sciences. Responsibilities include fact-checking...For phdRemote workFlexible hours
- A leading AI training company in the United States is seeking a Biology Research Scientist. In this remote role, you will evaluate AI chatbots by posing complex biology questions and assessing... ..., or related fields and master's or PhD qualifications are preferred. Payment...For phdHourly payRemote work
$60 per hour
...Professionals to join their Expert Network to evaluate AI-generated science. This role allows you... ...rate of up to $60 per hour for reviewing model responses, validating technical claims,... ...educational background with a BS, MS, or PhD in life sciences and the ability to...For phdHourly payRemote workWork from homeFlexible hours$60 per hour
...developing cutting-edge AI systems, while enjoying... ...AI development. AI models are increasingly capable... ...AI models on tasks like evaluating AI-generated quantitative... ..., operations research, or any other quantitative... ...similar); a master's or PhD is a plus. ~ Relevant...For phdHourly payFull timeRemote workFlexible hours$40 per hour
A leading AI training company is seeking a Biotechnology R&D Scientist to evaluate and improve AI models by posing complex biological questions. This remote position allows candidates... ...and fluency in English, and a Master's or PhD is preferred but not required. #J-18808-...For phdRemote jobHourly pay$40 per hour
A leading AI training company is seeking a Postdoctoral Physics Associate to evaluate AI chatbot outputs and enhance model quality. This remote position requires an expert level of understanding... ...in physics concepts, along with a PhD. Candidates will assess the...For phdHourly payRemote workFlexible hours$40 per hour
A leading AI research firm is seeking a Postdoctoral Researcher in Chemistry to evaluate AI chatbot outputs and improve their performance. This remote role offers flexibility... ...in English, and a current or completed PhD. Applicants from the United States are encouraged...For phdHourly payRemote work$40 per hour
...Physics Associate to join their team. In this remote role, you will train AI models, measure their progress, and evaluate their performance in solving complex physics problems. Applicants should have a PhD in Physics or a related field, along with a strong grasp of classical...For phdHourly payRemote workFlexible hours$40 per hour
...Process Development Chemist to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the... ..., in progress, or completed Master's and/or PhD is preferred but not required Notes Payment...For phdHourly payFull timeContract workPart timeRemote work$40 per hour
...data annotation company is seeking a Microbiologist to train AI models. You will evaluate AI chatbot outputs on complex biology questions and assess... ...biology, with a preference for those holding a Master's or PhD. This remote position offers flexible scheduling, hourly...For phdHourly payRemote workFlexible hours$40 per hour
...leading data annotation company seeks a Statistician to enhance AI models through evaluating chatbot logic and solving complex mathematical problems.... ...English and strong expertise in mathematics. A Master's or PhD is preferable, but not mandatory. Only candidates from the...For phdHourly payContract workRemote work$40 per hour
...a Microbiologist to join their team remotely. You will train AI models, evaluate outputs, and ensure the quality of AI chatbots in biological... ...scheduling, with hourly pay starting at $40+ USD. A Master's or PhD in Biology or related fields is preferred, though not...For phdHourly payRemote work$40 per hour
A cutting-edge AI training company in the United States is looking for a Microbiologist to evaluate and improve AI chatbot models. This role allows you to work remotely and choose projects... ...biology and genetics. A Master's or PhD is preferred but not required. Only applicants...For phdHourly payContract workRemote work- ...data technology company is seeking a Statistician to enhance AI models by evaluating their logic and progress. The role demands expert-level... ...AI models and assessing outputs for quality. A Master's or PhD is preferred, and hourly compensation starts at $40+, with bonuses...For phdHourly payFull timePart timeRemote workFlexible hours
$40 per hour
A tech company specializing in AI training is looking for a Statistician... ...this remote role, you'll train AI models by providing complex math problems and evaluating their outputs for quality and... ...fluency in English. A Master's or PhD is preferred but not required. This...For phdHourly payRemote workFlexible hours$70 - $75 per hour
...Role Overview Apply your STEM PhD expertise to design challenging domain problems and create high-quality data used to train and evaluate state-of-the-art large language models for a leading AI research lab. This role focuses on generating difficult, domain-specific problems...For phdHourly payRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to STEM PhD Researcher for AI Model Evaluation. Be the first to apply!
- survey researcher United States
- lead researcher United States
- blockchain researcher United States
- senior design researcher United States
- machine learning researcher United States
- academic researcher United States
- vulnerability researcher United States
- work from home court researcher United States
- music researcher United States
- visiting researcher United States


