Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

STEM Expert for AI Model Evaluation

SaidGig

Role Overview

Drive the creation and evaluation of challenging STEM problems used to fine-tune and benchmark large language models. You will design multi-step physics and math problems, produce clear step-by-step solutions with rigorous reasoning, and collaborate with researchers to build evaluation benchmarks that probe model limitations. This role is fully remote, contract-based, and ideal for candidates currently engaged in advanced STEM study or research.

Company overview

Based in San Francisco, the company accelerates frontier AI research and helps enterprises deploy reliable, high-impact AI systems. Work here means contributing high-quality data, advanced training pipelines, and domain expertise to support cutting-edge LLM research and production deployments.

Key Responsibilities
  • Design and solve challenging STEM problems to probe the limitations of large language models, with an emphasis on areas where models struggle, such as abstraction, multi-step reasoning, and symbolic manipulation.
  • Create clear, high-quality, step-by-step solutions with well-articulated reasoning suitable for evaluation and training use.
  • Collaborate with LLM researchers to align problem sets and solutions with evaluation goals and to define success criteria.
  • Help develop new evaluation benchmarks based on Physics curricula spanning early undergraduate through PhD-level topics.
  • Provide constructive feedback and detailed annotations on model outputs and dataset items.
Qualifications
  • Education and experience, preferred: currently pursuing or holding a Master’s, PhD, or Postdoctoral degree in STEM, Applied Physics, or a closely related field.
  • Analytical skills: strong research aptitude and the ability to analyze and solve complex physics and STEM problems using a structured, logical approach.
  • Communication: excellent structured written communication, ability to explain STEM concepts clearly in simple language, and to use visuals and physics reasoning where appropriate.
  • Creative thinking: capacity for creative and lateral thinking when designing problem prompts and solutions.
  • Feedback and annotation: experience or aptitude for providing detailed, constructive annotations and review notes.
  • Remote work skills: self-motivated, able to work independently, and effective at collaborating in a distributed environment.
  • Technical setup: access to a desktop or laptop with a reliable internet connection.
Work Terms
  • Location: Remote.
  • Engagement: Contract, contractor assignment or freelancer status.
  • Benefits: This engagement does not include medical or paid leave.
  • Duration and extension: Contracts may be extended based on performance and project needs.
Perks
  • Work fully remotely on cutting-edge AI projects.
  • Opportunity to contribute to leading LLM research and enterprise AI deployments.
  • Gain experience leveraging AI tools to strengthen analytical skills and future-proof your career.
Eligibility

Candidates currently pursuing or holding a Master’s, PhD, or Postdoctoral degree in STEM, Applied Physics, or a related field are eligible and encouraged to apply.

How to Apply

If you meet the eligibility above and are interested in contributing to LLM evaluation and benchmark creation, please submit an application. Eligible applicants will be considered based on their qualifications and fit for current project needs.

Vacancy posted more than 2 months ago
Similar jobs that could be interesting for youBased on the STEM Expert for AI Model Evaluation in United States vacancy
  • $50 per hour

     ...Design and author challenging STEM problems and clear, step-by-step...  ...help fine-tune large language models such as ChatGPT. You will...  ...model limitations, contribute evaluation benchmarks across physics curricula...  ...provide experience applying AI to improve analytical workflows... 
    Suggested
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    2 days ago
  • $55 per hour

     ...Biology experts contribute their scientific knowledge to AI research projects by helping improve how large language models understand and explain specialized biological...  ...for AI systems. Evaluate large language model...  ...school''s requirements. STEM OPT is not supported.... 
    Suggested
    Hourly pay
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $65 per hour

     ...security expertise to design domain-specific prompts and evaluate large language model outputs for AI research projects, improving model behavior, safety,...  ...course, this program may not satisfy that requirement. STEM OPT is not supported. Application Process Create... 
    Suggested
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Suggested
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    5 days ago
  • $60 - $80 per hour

     ...building foundational large language models by applying deep insurance domain expertise to create, evaluate, and refine training data. You...  ...claims practice. Evaluate AI model outputs against...  ...Collaborate with other subject-matter experts to ensure consistency and accuracy... 
    Suggested
    Hourly pay
    Weekday work

    SaidGig

    United States
    a month ago
  •  ...improve and validate large language models, by creating realistic retail...  ...LLC and placed with a leading AI lab. Key Responsibilities...  ...in real retail practice. Evaluate AI model outputs against...  ...Collaborate with other subject matter experts to ensure consistency and... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    5 days ago
  • $60 per hour

     ...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Hourly pay
    Remote work
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    5 days ago
  • Prolific is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities... 
    Remote job
    Hourly pay
    Flexible hours

    Prolific

    Jacksonville, FL
    3 days ago
  • $60 per hour

    Prolific is seeking Biology Experts and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with...  ...pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    Dallas, TX
    3 days ago
  • $60 per hour

    Prolific, located in Arizona, is seeking Biology Experts and Life Science Professionals to join their Expert Network. This role involves evaluating and training AI models with real scientific expertise. Successful candidates will review AI-generated content for accuracy... 
    Hourly pay

    Prolific

    Arizona City, AZ
    2 days ago
  • $55 per hour

     ...Role Overview Biology experts apply their domain knowledge to design domain-specific prompts, evaluate large language model outputs, and guide AI research across biological subfields. This role...  ...may not meet that requirement. STEM OPT is not supported. Application... 
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    2 days ago
  •  ...Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems,...  ...electromagnetism, use adversarial prompting to surface errors, and provide expert critique of AI responses while working with project #J-1880... 
    Remote job

    Dorado

    New York, NY
    3 days ago
  • $60 per hour

    Prolific is seeking Chemistry Experts and Chemical Engineers to join their Expert Network. Participants will evaluate AI-generated chemistry through tasks that assess factual accuracy...  ..., enabling cutting-edge advancements in AI models. The position requires a strong educational... 
    Hourly pay

    Prolific

    Arizona City, AZ
    2 days ago
  • $80 per hour

     ...and distribution experience to craft expert-level training content and evaluate AI-generated responses against real-...  ...written feedback that helps improve AI model performance and response quality....  ...not satisfy that requirement. STEM OPT is not supported for this work.... 
    Part time
    Remote work

    SaidGig

    United States
    20 days ago
  • $80 - $100 per hour

     ...improve next-generation AI systems through practical technical input, evaluations, and high-quality training...  ...coding agents in complex STEM work. Key Responsibilities Provide expert analysis, feedback, and practical...  ...data challenges for AI model development. Document... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Remote
    9 days ago
  • $50 per hour

     ...to create and assess domain-specific prompts and to evaluate large language model responses, helping improve AI performance on chemistry problems and explanations....  ...before applying. Candidates who require a new STEM OPT I-983 are not eligible at this time. Candidates... 
    Part time
    H1b
    Remote work
    Visa sponsorship
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $60 - $90 per hour

     ...author rigorous, multi-step evaluation tasks that translate...  ...state-of-the-art models cannot yet solve reliably...  ...researchers and subject-matter experts to ensure consistent,...  ...MSc or PhD in a STEM field, or in a computational...  ...Prior experience in AI training, model evaluation... 
    Hourly pay
    Full time
    Freelance
    Remote work

    SaidGig

    United States
    1 day ago
  • $17 - $54 per hour

     ...Music & Lyrics Expert - French | Remote AI Model Evaluation is a remote evaluation track for reviewing french generalist evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured... 
    Remote job
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    3 days ago
  • $60 - $80 per hour

     ...Apply deep retail domain expertise to help build and evaluate foundational generative AI models. You will design realistic retail tasks, produce authoritative...  ...tasks. Collaborate with other subject matter experts to ensure consistency and accuracy in training and... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    1 day ago
  • $150 per hour

     ...professionals apply their domain expertise to evaluate AI-generated outputs, assess technical...  ...clear, structured feedback that improves models'' understanding of aerospace tasks, terminology...  ...may not meet that requirement. STEM OPT is not supported for participation in... 
    Hourly pay
    Temporary work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $85 per hour

     ...Role Overview Psychology experts apply clinical and research...  ...domain-specific prompts, and evaluate large language models to improve their...  ...This role supports year-round AI research projects that vary...  ...fulfill that requirement. STEM OPT is not supported for participation... 
    Hourly pay
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $75 per hour

     ...Finance professionals apply their financial analysis, modeling, and advisory expertise to evaluate AI-generated financial content, create job-relevant prompts...  ...course, the program may not meet that requirement. STEM OPT is not supported. Application Process Create... 
    Hourly pay
    Full time
    Contract work
    Part time
    For contractors
    Bank staff
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $80 per hour

     ...analysis, forecasting, and REC markets to evaluate AI-generated content and produce expert-level analytical outputs that...  ..., actionable feedback to improve model accuracy for renewable energy and environmental...  ...not satisfy that requirement. STEM OPT is not supported.... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $75 per hour

     ...survey, and photogrammetric expertise to evaluate AI-generated maps and geospatial content, verify...  ...clear, structured feedback that improves model outputs. No prior AI experience is...  ...CPT course, this program may not qualify. STEM OPT is not supported. Application process... 
    Hourly pay
    Temporary work
    Part time
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $80 per hour

     ...inventory planning, and supply chain operations to create expert training data and evaluate AI-generated responses for accuracy and relevance. This is...  ...course, these projects may not meet that requirement. STEM OPT is not supported. Refer to the program help resources... 
    Hourly pay
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $20 - $40 per hour

     ...expertise to improve how next-generation AI systems learn and reason by analyzing...  ...becomes high-quality training data and evaluations for AI models. The role supports a unique customer project...  ..., technical clarifications, and expert-level commentary on complex engineering... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Indiana
    20 days ago
  • $105 per hour

     ...Role Overview Web Development Experts apply their web engineering skills to help refine and evaluate Large Language Models for high-stakes business...  ...intersection of web development and AI research, supporting model...  ...meet that requirement. STEM OPT is not supported.... 
    Part time
    Work experience placement
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $20 - $75 per hour

     ...Role Overview Provide expert telecommunications guidance to help train and evaluate next-generation AI systems. You will convert real-world telecom knowledge into high-quality...  ...data, evaluations, and feedback that improve model learning, reasoning, and performance. micro1... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $50 - $101 per hour

     ...knowledge to help train next-generation AI systems. In this remote contractor role supporting...  ...programs and nutrition guidance so models learn accurate, practical, safe fitness...  ...general wellness guidance. Create and evaluate sample fitness programs and nutrition plans... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Major County, OK
    20 days ago
  •  ...definition of excellent enterprise selling for a cutting-edge generative AI team by auditing multi-step sales workflows, producing end-to-end expert examples, and shaping the evaluation standards the model learns from. You will apply deep, practical selling experience to... 
    Hourly pay
    Full time
    Work at office
    Remote work

    SaidGig

    Remote
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to STEM Expert for AI Model Evaluation. Be the first to apply!