STEM Researcher for AI Model Evaluation
$60 - $90 per hourSaidGig
Help build next-generation agentic evaluation benchmarks for frontier AI models by turning rigorous scientific practice into challenging, multi-step evaluation tasks. You will design experiments, implement and run them in code, analyze results, and write careful conclusions that expose where state-of-the-art models fail. Tasks are designed to require one to two days of focused work and combine study design, computational implementation, and clear written criteria for success.
Key Responsibilities- Design tasks that operationalize experimental design, hypothesis testing, and rigorous evaluation into multi-step challenges that current models do not reliably solve.
- Author and execute solutions in Python and notebook environments, documenting methods and expected results at the level of a careful researcher.
- Define scoring and success criteria, specifying what distinguishes sound scientific reasoning from superficially plausible answers.
- Evaluate model outputs against your tasks, identify methodological and reasoning errors, and flag mistakes a working researcher would notice immediately.
- Collaborate with researchers and subject-matter experts to align evaluations, share findings, and keep assessments consistent.
- Produce reproducible code, analyses, and written summaries for each task you create.
- MSc or PhD in a STEM field, or in a computational social science or humanities discipline, or equivalent hands-on experience in a research-heavy role involving coding and data analysis.
- At least 1 year of experience in an active research position, in academia, industry, or national labs.
- Significant computational research experience, including Python-based analysis, simulation, modeling, or building data pipelines.
- Strong grounding in experimental design, hypothesis testing, and rigorous evaluation.
- Familiarity with Git, development IDEs, and notebook environments such as Jupyter or Colab.
- Preferred: prior experience with AI model training, model evaluation, or authoring benchmarks and tasks.
- Highly detail oriented, creative in task design, excellent written communication, and able to work independently on ambiguous, open-ended problems.
- Ability to commit approximately 35 hours per week.
- This is a W-2 employment role with Cincinnatus LLC, intended as a structured, role-based position rather than a freelance or project-only engagement.
- Placement will be with a leading AI research team as part of their extended workforce, working closely within client teams and standard enterprise workflows.
- The position is fully remote within the United States.
- Typical time commitment is approximately 35 hours per week. Individual tasks generally take one to two days of continuous focused effort.
- Hourly pay range: $60.00 to $90.00 per hour.
- Compensation is paid through Cincinnatus LLC as the employer of record.
- Because this is W-2 employment with a U.S. employer, candidates must be able to be employed by a United States-based company while working remotely within the United States.
- Cincinnatus LLC handles employment administration, including offers, onboarding, payroll, benefits, and compliance for this role.
- Opportunities may be listed or discovered via partner platforms, however hiring, onboarding, and employment administration are managed by Cincinnatus LLC.
- Cincinnatus is an Equal Employment Opportunity employer and will provide reasonable accommodations for qualified individuals with disabilities during the application process.
$20 - $55 per hour
...expertise to conduct literature and document research, prepare and review complex Word and... ...and presentations that help train and evaluate next-generation AI systems. You will contribute evidence-... ..., healthcare, policy, finance, and STEM engineering. Key Responsibilities...SuggestedHourly payFor contractorsRemote work$50 per hour
...and solve challenging STEM problems to help fine-tune large language models, producing clear, step-... ...on helping define new evaluation benchmarks based on physics... ...help you learn to use AI tools to become a... ...reasoning. Collaborate with researchers to align problems and...SuggestedContract workFor contractorsFreelanceRemote work$400k
...Overview Define the frontier of AI-powered legal reasoning by building rigorous evaluation frameworks and benchmarks for... ...This role combines applied legal research, dataset curation, and collaboration... ...intelligence in large language models and agentic workflows. About...SuggestedFull timeRemote work$80 - $110 per hour
...Overview Contribute expert, research-level chemistry judgment to train and evaluate next-generation AI systems that reason about organic... ...literature, and identify model failure modes at the level of... ...institution in chemistry or a related STEM field. Receipt of a...SuggestedHourly payPart timeImmediate startRemote work- ...clinical and workplace expertise to evaluate AI-generated content in their... ...feedback that improves model performance on medical tasks... ...language. This hourly, temporary AI research engagement runs year-round,... ...course requirements. STEM OPT is not supported. Refer...SuggestedHourly payTemporary workPart timeRemote workFlexible hours
$75 per hour
...Role Overview AI and Machine Learning Researchers apply deep expertise in machine learning... ...science research to evaluate AI-generated outputs and create... ...data that improves model understanding of advanced... ...meet that requirement. STEM OPT is not supported for this...Hourly payFull timeContract workPart timeRemote workFlexible hours$60 - $90 per hour
...Help build next-generation agentic evaluation benchmarks for frontier AI models by acting as a ground-truth expert... ...will design and execute realistic, research-style analysis tasks that test model... ...data science, or another quantitative STEM field, or equivalent practical...Hourly payFull timeFreelanceRemote work- ...clinical imaging and diagnostic expertise to evaluate AI model outputs, assess field-specific content,... ...Work asynchronously with external AI research teams to complete assigned tasks and... ...may not satisfy that requirement. STEM OPT is not supported for this program....Hourly payTemporary workPart timeRemote workFlexible hours
$40 per hour
...DataAnnotation is seeking a Biotechnology R&D Scientist to train AI models. In this role, you will evaluate the outputs of AI chatbots and assess their logic to improve model quality. The ideal candidate should have a deep understanding of cell biology, genetics, biochemistry...Hourly payFor contractorsRemote work$40 - $65 per hour
...Overview Contribute domain expertise to evaluate and harden frontier large language models by crafting adversarial multi-turn... ...on improving how next-generation AI systems learn, reason, and behave,... ...analysis-centric fields such as research, editorial, technical writing, or...Remote jobHourly payFor contractors$80 - $135 per hour
...benchmark (arXiv:2509.26574v3). The role involves solving frontier research-level physics problems end-to-end, auditing expert... ...adjudicating between competing solutions so that large language models can be evaluated on rigorous physics reasoning. Physics Subdomains...Hourly payRemote work10 hours per week$40 per hour
A technology company in Massachusetts is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal candidates will hold advanced qualifications in biology or biochemistry. This position allows full...Hourly payFull timePart timeRemote work$400k
...Join a dynamic research team as a Member of Technical Staff (MTS) focused on Medical & Health... ...will play a pivotal role in advancing AI systems designed to enhance healthcare, clinical... ...emphasizes the development of robust evaluation frameworks that assess medical reasoning,...Full timeRemote work$40 per hour
A leading data annotation company is seeking an R&D Biologist to improve AI models by evaluating their performance and logic. In this role, you will utilize your expertise in biology and related fields to enhance the quality of chatbots through detailed evaluations. Candidates...Hourly payFor contractorsRemote workFlexible hours$40 per hour
...data solutions company is seeking a Process Development Chemist to join their team remotely. In this role, you will train AI models by evaluating their performance on complex chemistry questions. The position offers flexibility to work on chosen projects at an hourly rate...Hourly payRemote work$60 - $80 per hour
...marketing subject-matter expertise to a GenAI team building foundational AI models. You will create realistic marketing tasks, evaluate model outputs against structured rubrics, and advise research and engineering teams on brand strategy, growth marketing, and campaign-...Hourly payWeekday work$40 per hour
...focused company is seeking a Process Development Chemist to evaluate and improve AI chatbots through complex chemistry questions. Candidates should... ...from the United States only. Join us and use your chemistry expertise to drive AI model quality improvement! #J-18808-LjbffrHourly payFull timePart timeRemote work$40 per hour
A research organization is seeking a Biology Research Scientist to evaluate AI chatbots' performance and improve their quality. The role requires an expert level of biology, with a focus on cell biology and genetics. Applicants can work remotely and choose their projects...Hourly payContract workRemote work$40 per hour
A data annotation company seeks an R&D Biologist to enhance AI models by evaluating chatbot outputs related to complex biology queries. This remote position allows flexibility in project selection and scheduling, with hourly rates starting at $40+ USD. Ideal candidates...Hourly payRemote work$40 per hour
A data annotation company is seeking a Biology Research Scientist to evaluate and improve AI models by testing their logic with complex biology questions. The ideal candidate should have a strong background in biology or biochemistry and be detail-oriented. This position...Hourly payRemote workFlexible hours$40 per hour
A technology company focused on AI model training is seeking an R&D Biologist to evaluate AI chatbots' performance using complex biological queries. The position is flexible, allowing you to choose projects and work on your own schedule. Candidates should have an expert...Hourly payRemote workFlexible hours$40 per hour
A leading AI training firm in the United States is seeking an R&D Biologist to join their team. In this remote position, you will evaluate AI chatbots and enhance their models while ensuring the biological accuracy of their outputs. The ideal candidate should have an expert...Hourly payRemote work$40 per hour
...company in the United States is seeking an R&D Biologist to train AI models and improve their quality. This position offers remote... ...selected projects at your own schedule. Responsibilities include evaluating the performance of AI chatbots on complex biology topics. Candidates...Hourly payRemote work- A technology firm is seeking a Process Development Chemist to train AI models remotely. The role requires expertise in chemistry to evaluate AI chatbot responses to complex problems. Candidates should have solid knowledge of chemistry concepts with fluency in English....Hourly payFor contractorsRemote workFlexible hours
- A leading AI training company in the United States is seeking a Biology Research Scientist. In this remote role, you will evaluate AI chatbots by posing complex biology questions and assessing their logic and performance. An expert-level understanding of biology, along...Hourly payRemote work
$40 per hour
...We are looking for a Process Development Chemist to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the quality of each model. In this role you will need to hold an expert level...Hourly payFull timeContract workPart timeRemote work$40 per hour
A leading technology company in the United States is looking for a Biology Research Scientist to train AI models by measuring and evaluating the models. This role involves a deep understanding of biology and related fields. You will choose your projects and work at your...Hourly payRemote workFlexible hours$70 - $90 per hour
...Contribute domain expertise to a cutting-edge project with a leading AI research lab, producing high-quality, hard problems and data that guide development of state-of-the-art large language models. The work emphasizes rigorous subject-matter accuracy and careful adherence...Hourly payRemote work$13.01 - $23.43 per hour
...Work directly with generative music systems by listening to and evaluating AI-produced songs in Thai and English, producing detailed annotations and technical ratings that guide model improvements. This role combines musical judgment, audio engineering vocabulary, and...Hourly payPart timeImmediate startRemote work10 hours per week$40 per hour
A leading AI training company is seeking a Biotechnology R&D Scientist to evaluate and improve AI models by posing complex biological questions. This remote position allows candidates to choose projects and work on their own schedule, with rates starting at $40+ per hour...Remote jobHourly pay
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to STEM Researcher for AI Model Evaluation. Be the first to apply!
- survey researcher United States
- lead researcher United States
- blockchain researcher United States
- senior design researcher United States
- machine learning researcher United States
- freelance researcher United States
- academic researcher United States
- vulnerability researcher United States
- work from home court researcher United States
- music researcher United States


