Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Physics Researcher for AI Model Evaluation

$80 - $135 per hour

SaidGig

Role Overview

Create and verify human-quality reference solutions for the CritPt benchmark (arXiv:2509.26574v3), a frontier research-level physics benchmark. The role produces fully human-verified reference data used to evaluate large language model performance on frontier physics reasoning. Work includes solving CritPt research-level problems end-to-end, auditing other experts'' solutions, and adjudicating between parallel solution attempts to determine the golden reference.

Physics subdomains covered

  • High Energy Physics and Mathematical Physics
  • Biophysics and Statistical Physics
  • Condensed Matter and AMO
  • Gravitation, Cosmology, and Astrophysics
  • Quantum Information
  • Optical Properties of Materials
  • Magnetic Materials
  • Measurements in Quantum Mechanics
Key Responsibilities
  • Solve research-level physics challenges end-to-end, with verifiable derivations, runnable code, and peer-reviewed references.
  • Decompose challenges into standalone checkpoint sub-problems that require genuine physical reasoning and can be independently verified.
  • Author Python answer templates that include automated grading functions for symbolic and numerical answers.
  • Audit submitted solutions for correctness, scope, and soundness of method, providing actionable feedback across iterations.
  • Adjudicate between parallel solver attempts and decide which solution becomes the golden reference for a problem.
  • Document detailed chain-of-thought reasoning, specify error tolerances, present equivalent symbolic forms, and supply verification test cases.
Qualifications
  • Solver track: PhD or postdoc in the relevant subfield, senior PhD student minimum.
  • Auditor track: Postdoc or junior professor in the relevant subfield, PhD minimum.
  • Adjudicator track: Full professor or industry research principal investigator in the relevant subfield, senior postdoc or junior professor minimum.
  • Hands-on familiarity with at least two canonical methods of the target subfield, demonstrated through publications, broader coverage preferred.
  • Provide 3 to 5 representative publications, with arXiv ID or DOI, ideally within the last approximately 5 years and in the target subfield.
  • Working proficiency with LaTeX, Python, Jupyter, and SymPy.
  • Strong written English, B2, C1, or C2 level minimum; native or near-native preferred.
Work Terms
  • Location: Remote.
  • Employment type: hourly.
  • Expected commitment: approximately 10 hours per week, sustained across an 8 to 10 week window per task pool.
  • Work is asynchronous.
Compensation
  • Pay range: $80 to $135 per hour, based on role and demonstrated expertise.
Eligibility
  • Candidates must hold the academic or research standing specified under Qualifications for the track they apply to.
  • Applicants must be able to provide the requested publications (arXiv ID or DOI) and demonstrate working proficiency with the listed tools.
  • Strong written English is required to prepare the human-verified reference solutions and feedback.
Vacancy posted 22 days ago
Similar jobs that could be interesting for youBased on the Physics Researcher for AI Model Evaluation in United States vacancy
  •  ...Overview Design graduate-level, research-focused computational problems that test whether advanced AI systems can perform real...  ...will be exercised against top AI models and iteratively refined until...  ...validators in Python. Run and evaluate problems against state-of-the-... 
    Suggested
    Hourly pay
    Remote work

    SaidGig

    United States
    6 days ago
  • $80 - $150 per hour

     ...Role Overview Apply your physics expertise to evaluate and improve scientific reasoning...  ...will train next generation AI systems supporting a...  ...theoretical arguments produced by researchers or AI systems. Detect...  ...Jupyter for theoretical modeling and computational validation... 
    Suggested
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Remote
    4 days ago
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public...  ...team at Scale deploys advanced AI systems—including LLMs,...  ...systems. Ability to convert research insights into measurable...  ...accommodations to applicants with physical and mental disabilities. If... 
    Suggested
    Full time

    Scale Ai

    Washington DC
    1 day ago
  • $40 per hour

    A biotechnology company is seeking a Biotechnology R&D Scientist to train AI models and evaluate their outputs. This role involves measuring the progress of AI chatbots with complex biology questions and ensuring their performance and correctness. Candidates should have... 
    Suggested
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Vermont
    5 days ago
  • $224k - $356.5k

     ...tapping into the unlimited potential of AI to define the next era of computing. An...  ...Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful...  ...and communicate effectively across research, engineering, and product teams.Ways to... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    5 days ago
  • $40 per hour

    A forward-thinking AI development firm seeks experienced quantitative professionals to evaluate AI-generated work, applying their skills in statistical analysis, predictive modeling, and technical writing. This fully remote opportunity offers a flexible schedule and projects... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Sioux Falls, SD
    5 days ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting... 
    Hourly pay
    Remote work

    DataAnnotation

    Juneau, AK
    5 days ago
  • $40 per hour

     ...A leading AI development firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully remote position allows for a flexible schedule, offering competitive pay starting at $40+ per hour. Ideal candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Bismarck, ND
    5 days ago
  • $40 per hour

    A leading AI training company is seeking a Biotechnology R&D Scientist to evaluate and improve AI models by posing complex biological questions. This remote position allows candidates to choose projects and work on their own schedule, with rates starting at $40+ per hour... 
    Hourly pay
    Remote work

    DataAnnotation

    Sioux Falls, SD
    3 days ago
  • $150 per hour

     ...Overview Aerospace engineering professionals apply their domain expertise to evaluate AI-generated outputs, assess technical content, and provide clear, structured feedback that improves models'' understanding of aerospace tasks, terminology, and practices. You will... 
    Hourly pay
    Temporary work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...Role Overview Play a central role on a GenAI research team by applying hands-on legal practice experience to improve how frontier AI models perform real legal work. In this position you will evaluate model outputs, create high-quality instruction specifications and authoritative... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    4 days ago
  • $40 per hour

    A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative work and provide critical feedback. This role offers the flexibility of remote work, allowing you to set your own schedule while focusing on impactful projects... 
    Hourly pay
    Remote work

    DataAnnotation

    Wisconsin
    5 days ago
  • $85 per hour

     ...environmental assessment, GIS, and renewable energy siting expertise to evaluate AI-generated outputs and to create expert-level training data...  ...use in your daily work. Collaborate asynchronously with AI research teams, documenting edge cases, common errors, and context... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $30 - $90 per hour

     ...maintain backend services in Go while evaluating and training alpha-stage AI coding tools. This contract role...  ...optimizations. Test and evaluate alpha AI models using Cursor, conducted over...  .... Collaborate with the research team via Slack, providing real-time... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to work remotely. In this role, you'll evaluate AI-generated quantitative work and solve technical problems while providing feedback to shape AI systems. Qualifications include 2+ years... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Salt Lake City, UT
    5 days ago
  • A leading AI development company is seeking experienced quantitative professionals for remote work evaluating AI-generated quantitative analysis. Ideal candidates will have a robust background in fields like data science, economics, or biostatistics, with at least 2 years... 
    Remote work

    DataAnnotation

    New York, NY
    5 days ago
  •  ...Drive the creation and evaluation of challenging STEM problems...  ...benchmark large language models. You will design multi-step physics and math problems,...  ...reasoning, and collaborate with researchers to build evaluation...  ...company accelerates frontier AI research and helps enterprises... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $100 - $150 per hour

     ...Role Overview Provide senior legal subject-matter expertise to a GenAI research team, creating authoritative instruction specifications, golden solutions, and evaluation benchmarks so frontier AI models reason correctly about real legal work. This role centers on hands-on... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Local area
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    4 days ago
  • $40 per hour

     ...analytics company seeks experienced quantitative professionals to evaluate AI-generated analysis and help advance AI development. This...  ...particularly those with experience in statistical methods and predictive modeling. Join to impact the next generation of AI systems dedicated to... 
    Hourly pay
    Remote work

    DataAnnotation

    El Paso, TX
    5 days ago
  • $85 per hour

     ...Overview Psychology experts apply clinical and research knowledge to design tasks, create domain-specific prompts, and evaluate large language models to improve their understanding and...  ...research. This role supports year-round AI research projects that vary by domain and... 
    Hourly pay
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $60 per hour

     ...developing cutting-edge AI systems, while enjoying...  ...AI development. AI models are increasingly capable...  ...AI models on tasks like evaluating AI-generated quantitative...  ..., operations research, or any other quantitative...  ...statistics, economics, finance, physics, biology, epidemiology,... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Charleston, WV
    5 days ago
  •  ...that develops large language models, shaping training data by designing...  ...practice with rigorous evaluation to improve model behavior for...  ...Responsibilities Work with research and engineering teams to close...  ...marketing practice. Evaluate AI model outputs using structured... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    17 days ago
  • $85 per hour

     ...renewable energy generation, REC trading, and portfolio management to evaluate AI-generated content and create expert training material. This...  ...and communicate effectively in writing with a distributed AI research team. Application Process Create a contributor profile... 
    Hourly pay
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $65 - $90 per hour

     ...Role Overview Apply your real-world architecture expertise to evaluate and improve how AI systems understand and reason about architecture. In this flexible, part-time, remote role you will review content for technical accuracy, answer domain-specific questions, and provide... 
    Hourly pay
    Part time
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Remote
    15 days ago
  • $60 - $80 per hour

     ...building foundational large language models by applying deep insurance domain expertise to create, evaluate, and refine training data. You...  ...Responsibilities Work with research and engineering teams to close...  ...claims practice. Evaluate AI model outputs against... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    a month ago
  • $40 per hour

    A leading AI development company seeks experienced quantitative professionals to evaluate AI-generated work and solve quantitative problems. This fully remote role offers a flexible schedule with competitive pay starting at $40+ per hour. Candidates should have 2+ years... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Little Rock, AR
    5 days ago
  • $40 per hour

     ...A leading AI company in the United States is seeking experienced quantitative professionals to evaluate and validate AI-generated analytical work. This fully remote position allows you to set your own schedule, with competitive hourly pay starting at $40 USD. Responsibilities... 
    Hourly pay
    Remote work

    DataAnnotation

    Jackson, MS
    5 days ago
  • $40 per hour

    A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated analytics and provide technical feedback for model improvement. The role offers fully remote work from multiple countries and a flexible schedule to choose your projects... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Lansing, MI
    5 days ago
  • $85 per hour

     ...geospatial and environmental expertise to evaluate AI-generated outputs used in environmental...  ..., and provide clear feedback to improve model behavior. Work with familiar tools...  .... Collaborate asynchronously with AI research teams while working independently to meet... 
    Hourly pay
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Physics Researcher for AI Model Evaluation. Be the first to apply!