STEM Expert for AI Model Evaluation
SaidGig
Drive the creation and evaluation of challenging STEM problems used to fine-tune and benchmark large language models. You will design multi-step physics and math problems, produce clear step-by-step solutions with rigorous reasoning, and collaborate with researchers to build evaluation benchmarks that probe model limitations. This role is fully remote, contract-based, and ideal for candidates currently engaged in advanced STEM study or research.
Company overviewBased in San Francisco, the company accelerates frontier AI research and helps enterprises deploy reliable, high-impact AI systems. Work here means contributing high-quality data, advanced training pipelines, and domain expertise to support cutting-edge LLM research and production deployments.
Key Responsibilities- Design and solve challenging STEM problems to probe the limitations of large language models, with an emphasis on areas where models struggle, such as abstraction, multi-step reasoning, and symbolic manipulation.
- Create clear, high-quality, step-by-step solutions with well-articulated reasoning suitable for evaluation and training use.
- Collaborate with LLM researchers to align problem sets and solutions with evaluation goals and to define success criteria.
- Help develop new evaluation benchmarks based on Physics curricula spanning early undergraduate through PhD-level topics.
- Provide constructive feedback and detailed annotations on model outputs and dataset items.
- Education and experience, preferred: currently pursuing or holding a Master’s, PhD, or Postdoctoral degree in STEM, Applied Physics, or a closely related field.
- Analytical skills: strong research aptitude and the ability to analyze and solve complex physics and STEM problems using a structured, logical approach.
- Communication: excellent structured written communication, ability to explain STEM concepts clearly in simple language, and to use visuals and physics reasoning where appropriate.
- Creative thinking: capacity for creative and lateral thinking when designing problem prompts and solutions.
- Feedback and annotation: experience or aptitude for providing detailed, constructive annotations and review notes.
- Remote work skills: self-motivated, able to work independently, and effective at collaborating in a distributed environment.
- Technical setup: access to a desktop or laptop with a reliable internet connection.
- Location: Remote.
- Engagement: Contract, contractor assignment or freelancer status.
- Benefits: This engagement does not include medical or paid leave.
- Duration and extension: Contracts may be extended based on performance and project needs.
- Work fully remotely on cutting-edge AI projects.
- Opportunity to contribute to leading LLM research and enterprise AI deployments.
- Gain experience leveraging AI tools to strengthen analytical skills and future-proof your career.
Candidates currently pursuing or holding a Master’s, PhD, or Postdoctoral degree in STEM, Applied Physics, or a related field are eligible and encouraged to apply.
How to ApplyIf you meet the eligibility above and are interested in contributing to LLM evaluation and benchmark creation, please submit an application. Eligible applicants will be considered based on their qualifications and fit for current project needs.
$50 per hour
...Design and author challenging STEM problems and clear, step-by-step... ...help fine-tune large language models such as ChatGPT. You will... ...model limitations, contribute evaluation benchmarks across physics curricula... ...provide experience applying AI to improve analytical workflows...SuggestedContract workFor contractorsFreelanceRemote work$55 per hour
...Biology experts contribute their scientific knowledge to AI research projects by helping improve how large language models understand and explain specialized biological... ...for AI systems. Evaluate large language model... ...school''s requirements. STEM OPT is not supported....SuggestedHourly payPart timeRemote workFlexible hours$65 per hour
...security expertise to design domain-specific prompts and evaluate large language model outputs for AI research projects, improving model behavior, safety,... ...course, this program may not satisfy that requirement. STEM OPT is not supported. Application Process Create...SuggestedPart timeRemote workFlexible hours- ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...SuggestedRemote workFlexible hours
$60 - $80 per hour
...building foundational large language models by applying deep insurance domain expertise to create, evaluate, and refine training data. You... ...claims practice. Evaluate AI model outputs against... ...Collaborate with other subject-matter experts to ensure consistency and accuracy...SuggestedHourly payWeekday work- ...improve and validate large language models, by creating realistic retail... ...LLC and placed with a leading AI lab. Key Responsibilities... ...in real retail practice. Evaluate AI model outputs against... ...Collaborate with other subject matter experts to ensure consistency and...Hourly payWeekday work
$60 per hour
...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Hourly payRemote workWork from homeFlexible hours- Prolific is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities...Remote jobHourly payFlexible hours
$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with... ...pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing...Remote jobHourly payWork from homeFlexible hours$60 per hour
Prolific, located in Arizona, is seeking Biology Experts and Life Science Professionals to join their Expert Network. This role involves evaluating and training AI models with real scientific expertise. Successful candidates will review AI-generated content for accuracy...Hourly pay$55 per hour
...Role Overview Biology experts apply their domain knowledge to design domain-specific prompts, evaluate large language model outputs, and guide AI research across biological subfields. This role... ...may not meet that requirement. STEM OPT is not supported. Application...Part timeRemote workFlexible hours- ...Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems,... ...electromagnetism, use adversarial prompting to surface errors, and provide expert critique of AI responses while working with project #J-1880...Remote job
$60 per hour
Prolific is seeking Chemistry Experts and Chemical Engineers to join their Expert Network. Participants will evaluate AI-generated chemistry through tasks that assess factual accuracy... ..., enabling cutting-edge advancements in AI models. The position requires a strong educational...Hourly pay$80 per hour
...and distribution experience to craft expert-level training content and evaluate AI-generated responses against real-... ...written feedback that helps improve AI model performance and response quality.... ...not satisfy that requirement. STEM OPT is not supported for this work....Part timeRemote work$80 - $100 per hour
...improve next-generation AI systems through practical technical input, evaluations, and high-quality training... ...coding agents in complex STEM work. Key Responsibilities Provide expert analysis, feedback, and practical... ...data challenges for AI model development. Document...Hourly payFor contractorsRemote work$50 per hour
...to create and assess domain-specific prompts and to evaluate large language model responses, helping improve AI performance on chemistry problems and explanations.... ...before applying. Candidates who require a new STEM OPT I-983 are not eligible at this time. Candidates...Part timeH1bRemote workVisa sponsorship10 hours per weekFlexible hours$60 - $90 per hour
...author rigorous, multi-step evaluation tasks that translate... ...state-of-the-art models cannot yet solve reliably... ...researchers and subject-matter experts to ensure consistent,... ...MSc or PhD in a STEM field, or in a computational... ...Prior experience in AI training, model evaluation...Hourly payFull timeFreelanceRemote work$17 - $54 per hour
...Music & Lyrics Expert - French | Remote AI Model Evaluation is a remote evaluation track for reviewing french generalist evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured...Remote jobFor contractors10 hours per week$60 - $80 per hour
...Apply deep retail domain expertise to help build and evaluate foundational generative AI models. You will design realistic retail tasks, produce authoritative... ...tasks. Collaborate with other subject matter experts to ensure consistency and accuracy in training and...Hourly payWeekday work$150 per hour
...professionals apply their domain expertise to evaluate AI-generated outputs, assess technical... ...clear, structured feedback that improves models'' understanding of aerospace tasks, terminology... ...may not meet that requirement. STEM OPT is not supported for participation in...Hourly payTemporary workPart timeRemote workFlexible hours$85 per hour
...Role Overview Psychology experts apply clinical and research... ...domain-specific prompts, and evaluate large language models to improve their... ...This role supports year-round AI research projects that vary... ...fulfill that requirement. STEM OPT is not supported for participation...Hourly payPart timeRemote workFlexible hours$75 per hour
...Finance professionals apply their financial analysis, modeling, and advisory expertise to evaluate AI-generated financial content, create job-relevant prompts... ...course, the program may not meet that requirement. STEM OPT is not supported. Application Process Create...Hourly payFull timeContract workPart timeFor contractorsBank staffRemote workFlexible hours$80 per hour
...analysis, forecasting, and REC markets to evaluate AI-generated content and produce expert-level analytical outputs that... ..., actionable feedback to improve model accuracy for renewable energy and environmental... ...not satisfy that requirement. STEM OPT is not supported....Hourly payFull timeContract workPart timeRemote workFlexible hours$75 per hour
...survey, and photogrammetric expertise to evaluate AI-generated maps and geospatial content, verify... ...clear, structured feedback that improves model outputs. No prior AI experience is... ...CPT course, this program may not qualify. STEM OPT is not supported. Application process...Hourly payTemporary workPart timeRemote work$80 per hour
...inventory planning, and supply chain operations to create expert training data and evaluate AI-generated responses for accuracy and relevance. This is... ...course, these projects may not meet that requirement. STEM OPT is not supported. Refer to the program help resources...Hourly payContract workPart timeRemote workFlexible hours$20 - $40 per hour
...expertise to improve how next-generation AI systems learn and reason by analyzing... ...becomes high-quality training data and evaluations for AI models. The role supports a unique customer project... ..., technical clarifications, and expert-level commentary on complex engineering...Hourly payFor contractorsRemote work$105 per hour
...Role Overview Web Development Experts apply their web engineering skills to help refine and evaluate Large Language Models for high-stakes business... ...intersection of web development and AI research, supporting model... ...meet that requirement. STEM OPT is not supported....Part timeWork experience placementRemote workFlexible hours$20 - $75 per hour
...Role Overview Provide expert telecommunications guidance to help train and evaluate next-generation AI systems. You will convert real-world telecom knowledge into high-quality... ...data, evaluations, and feedback that improve model learning, reasoning, and performance. micro1...Hourly payFor contractorsRemote work$50 - $101 per hour
...knowledge to help train next-generation AI systems. In this remote contractor role supporting... ...programs and nutrition guidance so models learn accurate, practical, safe fitness... ...general wellness guidance. Create and evaluate sample fitness programs and nutrition plans...Hourly payFor contractorsRemote work- ...definition of excellent enterprise selling for a cutting-edge generative AI team by auditing multi-step sales workflows, producing end-to-end expert examples, and shaping the evaluation standards the model learns from. You will apply deep, practical selling experience to...Hourly payFull timeWork at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to STEM Expert for AI Model Evaluation. Be the first to apply!
- fruit expert United States
- subject matter expert United States
- expert data analyst United States
- guest service support expert United States
- expert systems engineer United States
- technology expert United States
- fulfillment expert United States
- subject matter expert senior United States
- subject matter expert work from home United States
- sql expert United States



