Physics Researcher for AI Model Evaluation
$80 - $135 per hourSaidGig
Create and verify human-quality reference solutions for the CritPt benchmark (arXiv:2509.26574v3), a frontier research-level physics benchmark. The role produces fully human-verified reference data used to evaluate large language model performance on frontier physics reasoning. Work includes solving CritPt research-level problems end-to-end, auditing other experts'' solutions, and adjudicating between parallel solution attempts to determine the golden reference.
Physics subdomains covered
- High Energy Physics and Mathematical Physics
- Biophysics and Statistical Physics
- Condensed Matter and AMO
- Gravitation, Cosmology, and Astrophysics
- Quantum Information
- Optical Properties of Materials
- Magnetic Materials
- Measurements in Quantum Mechanics
- Solve research-level physics challenges end-to-end, with verifiable derivations, runnable code, and peer-reviewed references.
- Decompose challenges into standalone checkpoint sub-problems that require genuine physical reasoning and can be independently verified.
- Author Python answer templates that include automated grading functions for symbolic and numerical answers.
- Audit submitted solutions for correctness, scope, and soundness of method, providing actionable feedback across iterations.
- Adjudicate between parallel solver attempts and decide which solution becomes the golden reference for a problem.
- Document detailed chain-of-thought reasoning, specify error tolerances, present equivalent symbolic forms, and supply verification test cases.
- Solver track: PhD or postdoc in the relevant subfield, senior PhD student minimum.
- Auditor track: Postdoc or junior professor in the relevant subfield, PhD minimum.
- Adjudicator track: Full professor or industry research principal investigator in the relevant subfield, senior postdoc or junior professor minimum.
- Hands-on familiarity with at least two canonical methods of the target subfield, demonstrated through publications, broader coverage preferred.
- Provide 3 to 5 representative publications, with arXiv ID or DOI, ideally within the last approximately 5 years and in the target subfield.
- Working proficiency with LaTeX, Python, Jupyter, and SymPy.
- Strong written English, B2, C1, or C2 level minimum; native or near-native preferred.
- Location: Remote.
- Employment type: hourly.
- Expected commitment: approximately 10 hours per week, sustained across an 8 to 10 week window per task pool.
- Work is asynchronous.
- Pay range: $80 to $135 per hour, based on role and demonstrated expertise.
- Candidates must hold the academic or research standing specified under Qualifications for the track they apply to.
- Applicants must be able to provide the requested publications (arXiv ID or DOI) and demonstrate working proficiency with the listed tools.
- Strong written English is required to prepare the human-verified reference solutions and feedback.
- ...Overview Design graduate-level, research-focused computational problems that test whether advanced AI systems can perform real... ...will be exercised against top AI models and iteratively refined until... ...validators in Python. Run and evaluate problems against state-of-the-...SuggestedHourly payRemote work
$80 - $150 per hour
...Role Overview Apply your physics expertise to evaluate and improve scientific reasoning... ...will train next generation AI systems supporting a... ...theoretical arguments produced by researchers or AI systems. Detect... ...Jupyter for theoretical modeling and computational validation...SuggestedRemote jobHourly payFor contractors$208k - $300k
...Machine Learning Engineer - Model Evaluations, Public Sector The Public... ...team at Scale deploys advanced AI systems—including LLMs,... ...systems. Ability to convert research insights into measurable... ...accommodations to applicants with physical and mental disabilities. If...SuggestedFull time$40 per hour
A biotechnology company is seeking a Biotechnology R&D Scientist to train AI models and evaluate their outputs. This role involves measuring the progress of AI chatbots with complex biology questions and ensuring their performance and correctness. Candidates should have...SuggestedHourly payRemote workFlexible hours$224k - $356.5k
...tapping into the unlimited potential of AI to define the next era of computing. An... ...Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful... ...and communicate effectively across research, engineering, and product teams.Ways to...SuggestedFull time- ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote workFlexible hours
$40 per hour
A forward-thinking AI development firm seeks experienced quantitative professionals to evaluate AI-generated work, applying their skills in statistical analysis, predictive modeling, and technical writing. This fully remote opportunity offers a flexible schedule and projects...Hourly payRemote workFlexible hours$40 per hour
A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting...Hourly payRemote work$40 per hour
...A leading AI development firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully remote position allows for a flexible schedule, offering competitive pay starting at $40+ per hour. Ideal candidates...Hourly payRemote workFlexible hours$40 per hour
A leading AI training company is seeking a Biotechnology R&D Scientist to evaluate and improve AI models by posing complex biological questions. This remote position allows candidates to choose projects and work on their own schedule, with rates starting at $40+ per hour...Hourly payRemote work$150 per hour
...Overview Aerospace engineering professionals apply their domain expertise to evaluate AI-generated outputs, assess technical content, and provide clear, structured feedback that improves models'' understanding of aerospace tasks, terminology, and practices. You will...Hourly payTemporary workPart timeRemote workFlexible hours- ...Role Overview Play a central role on a GenAI research team by applying hands-on legal practice experience to improve how frontier AI models perform real legal work. In this position you will evaluate model outputs, create high-quality instruction specifications and authoritative...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package
$40 per hour
A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative work and provide critical feedback. This role offers the flexibility of remote work, allowing you to set your own schedule while focusing on impactful projects...Hourly payRemote work$85 per hour
...environmental assessment, GIS, and renewable energy siting expertise to evaluate AI-generated outputs and to create expert-level training data... ...use in your daily work. Collaborate asynchronously with AI research teams, documenting edge cases, common errors, and context...Hourly payFull timeContract workPart timeRemote workFlexible hours$30 - $90 per hour
...maintain backend services in Go while evaluating and training alpha-stage AI coding tools. This contract role... ...optimizations. Test and evaluate alpha AI models using Cursor, conducted over... .... Collaborate with the research team via Slack, providing real-time...Hourly payContract workPart timeFor contractorsRemote work$40 per hour
A leading AI development company is seeking experienced quantitative professionals to work remotely. In this role, you'll evaluate AI-generated quantitative work and solve technical problems while providing feedback to shape AI systems. Qualifications include 2+ years...Hourly payRemote workFlexible hours- A leading AI development company is seeking experienced quantitative professionals for remote work evaluating AI-generated quantitative analysis. Ideal candidates will have a robust background in fields like data science, economics, or biostatistics, with at least 2 years...Remote work
- ...Drive the creation and evaluation of challenging STEM problems... ...benchmark large language models. You will design multi-step physics and math problems,... ...reasoning, and collaborate with researchers to build evaluation... ...company accelerates frontier AI research and helps enterprises...Contract workFor contractorsFreelanceRemote work
$100 - $150 per hour
...Role Overview Provide senior legal subject-matter expertise to a GenAI research team, creating authoritative instruction specifications, golden solutions, and evaluation benchmarks so frontier AI models reason correctly about real legal work. This role centers on hands-on...Hourly payFull timeFreelanceInternshipLive inLocal areaRelocationRelocation package$40 per hour
...analytics company seeks experienced quantitative professionals to evaluate AI-generated analysis and help advance AI development. This... ...particularly those with experience in statistical methods and predictive modeling. Join to impact the next generation of AI systems dedicated to...Hourly payRemote work$85 per hour
...Overview Psychology experts apply clinical and research knowledge to design tasks, create domain-specific prompts, and evaluate large language models to improve their understanding and... ...research. This role supports year-round AI research projects that vary by domain and...Hourly payPart timeRemote workFlexible hours$60 per hour
...developing cutting-edge AI systems, while enjoying... ...AI development. AI models are increasingly capable... ...AI models on tasks like evaluating AI-generated quantitative... ..., operations research, or any other quantitative... ...statistics, economics, finance, physics, biology, epidemiology,...Hourly payFull timeRemote workFlexible hours- ...that develops large language models, shaping training data by designing... ...practice with rigorous evaluation to improve model behavior for... ...Responsibilities Work with research and engineering teams to close... ...marketing practice. Evaluate AI model outputs using structured...Hourly payWeekday work
$85 per hour
...renewable energy generation, REC trading, and portfolio management to evaluate AI-generated content and create expert training material. This... ...and communicate effectively in writing with a distributed AI research team. Application Process Create a contributor profile...Hourly payContract workPart timeRemote workFlexible hours$65 - $90 per hour
...Role Overview Apply your real-world architecture expertise to evaluate and improve how AI systems understand and reason about architecture. In this flexible, part-time, remote role you will review content for technical accuracy, answer domain-specific questions, and provide...Hourly payPart timeRemote work10 hours per weekFlexible hours$60 - $80 per hour
...building foundational large language models by applying deep insurance domain expertise to create, evaluate, and refine training data. You... ...Responsibilities Work with research and engineering teams to close... ...claims practice. Evaluate AI model outputs against...Hourly payWeekday work$40 per hour
A leading AI development company seeks experienced quantitative professionals to evaluate AI-generated work and solve quantitative problems. This fully remote role offers a flexible schedule with competitive pay starting at $40+ per hour. Candidates should have 2+ years...Hourly payRemote workFlexible hours$40 per hour
...A leading AI company in the United States is seeking experienced quantitative professionals to evaluate and validate AI-generated analytical work. This fully remote position allows you to set your own schedule, with competitive hourly pay starting at $40 USD. Responsibilities...Hourly payRemote work$40 per hour
A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated analytics and provide technical feedback for model improvement. The role offers fully remote work from multiple countries and a flexible schedule to choose your projects...Hourly payRemote workFlexible hours$85 per hour
...geospatial and environmental expertise to evaluate AI-generated outputs used in environmental... ..., and provide clear feedback to improve model behavior. Work with familiar tools... .... Collaborate asynchronously with AI research teams while working independently to meet...Hourly payContract workPart timeRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Physics Researcher for AI Model Evaluation. Be the first to apply!




