Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Benchmark Regression Model Evaluation Specialist [Remote]

AuraOne Human Data

Remote
  • Remote job

Benchmark Regression Model Evaluation Specialist is a remote evaluation track for reviewing benchmark regression model evaluation evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.

Why this role matters

AI data reviewers help turn benchmark regression model evaluation evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.

Responsibilities

  • Evaluate benchmark regression model evaluation evaluation model outputs against a versioned rubric and assign severity tags for Benchmark Regression Model Evaluation Specialist assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
  • Document recurring failure modes so the modeling team can target them in the next training run.

Qualifications

  • Prior evaluation, annotation, or human-rater experience on benchmark regression model evaluation evaluation or adjacent content for Benchmark Regression Model Evaluation Specialist work.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.

Example tasks

  • Compare two benchmark regression model evaluation evaluation model responses to the same prompt and pick the stronger one with rationale.
  • Tag an unsafe response with the correct policy category and severity.
  • Audit a 50-row batch for rubric consistency and report drift to the program lead.
  • Propose a rubric clarification after spotting a recurring failure mode.

Nice to have

  • Background in linguistics, content moderation, or trust & safety review.
  • Experience with inter-rater agreement metrics and calibration cycles.
  • Domain expertise that lets you spot subject-matter errors automated checks miss.

Skills

  • Model output evaluation
  • Rubric-based annotation
  • Severity tagging
  • Inter-rater calibration
  • Benchmark Regression Model Evaluation evaluation
  • Frontier evaluation
  • Rubric calibration
  • Failure analysis
  • Benchmark
  • Regression

Work model

Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.

Compensation

Hourly rate confirmed after the interview process.

Application process

Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.

Vacancy posted 19 days ago
Similar jobs that could be interesting for youBased on the Benchmark Regression Model Evaluation Specialist [Remote] in Remote vacancy
  • Dorado is seeking a Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems, probe model reasoning at the frontier, and help identify where models fail under rigorous scientific scrutiny.... 
    Suggested
    Remote job

    Dorado

    New York, NY
    2 days ago
  •  ...company. We build cutting-edge foundation AI models and end-to-end products that are...  ...Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling...  ...you will:Create ambitious new evaluation benchmarks that push the limits of what our models... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    2 days ago
  • $60 - $90 per hour

     ...labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo ,...  ...Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation... 
    Suggested
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    7 hours ago
  • $180k - $240k

     ...looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team...  ...pass/fail criteria grounded in Research benchmarks. ~Build monitoring systems that keep...  ...monitoring systems that detect quality regressions in production. ~Partner with DevOps... 
    Suggested
    Full time

    Deepgram

    Remote
    2 days ago
  • $40 - $50 per hour

     ...Role Overview Apply your linguistics expertise to evaluate large language model outputs and help train next-generation AI systems. You will analyze model-human conversations for linguistic precision, assess adherence to instructions, and deliver clear, actionable feedback... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    Immediate start

    SaidGig

    Europe
    4 days ago
  •  ...Causal Reasoning Model Evaluation Specialist is a remote review track for evaluating AI outputs across causal reasoning model evaluation research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results,... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    24 days ago
  • $15 - $20 per hour

     ...tools. ~Generate high-quality human evaluation data by identifying response strengths,...  ...and completeness of responses. ~Ensure model responses align with expected conversational...  ...in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'... 
    Part time
    Summer work

    Mercor

    Remote
    a month ago
  • $100 per hour

     ...improve the performance of large language models on finance tasks. You will work with AI...  ...AI systems. Key Responsibilities Evaluate LLM performance in finance areas where models...  ...approaches, evaluation strategies, and benchmark development. Qualifications... 
    Hourly pay
    Contract work
    For contractors
    Freelance
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear,...  ...corrections and feedback. Help define new evaluation benchmarks based on mathematics curricula spanning early undergraduate... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    24 days ago
  • Overview Lynker Corporation is seeking a Sea Ice Model Evaluation and Applications Scientist to support the Ocean and Tsunami Center (OTC...  ...compensation plan regularly calibrated against industry and location benchmarks 401(k) retirement plan with company-matching Employee Stock... 
    Temporary work
    Seasonal work
    Local area
    Remote work
    Flexible hours

    Lynker Corporation

    Suitland, MD
    4 days ago
  • Frontier Model Misuse Risk Evaluator is a remote red-team track for stress-testing AI systems against...  ...modeling team can reproduce, fix, and regress-test them. Responsibilities Design adversarial...  ...— US-eligible. Remote · Independent specialist contractor. Employment type:... 
    Hourly pay
    For contractors
    Remote work
    10 hours per week

    AuraOne

    New York, NY
    4 days ago
  • $80 - $135 per hour

     ...verify human-quality reference solutions for the CritPt benchmark (arXiv:2509.26574v3), a frontier research-level physics benchmark...  ...produces fully human-verified reference data used to evaluate large language model performance on frontier physics reasoning. Work includes... 
    Hourly pay
    Remote work
    10 hours per week

    SaidGig

    United States
    26 days ago
  • $60 - $90 per hour

     ...research by creating rigorous, real-world data analysis evaluations for generative AI models. You will design and complete complex analytical tasks...  .... Experience with AI training, model evaluation, or benchmark or task authoring is preferred. Exceptional attention... 
    Hourly pay
    Full time
    Freelance
    Remote work

    SaidGig

    United States
    27 days ago
  • $60 - $90 per hour

     ...run rigorous, multi-step machine learning evaluation tasks for a leading generative AI...  ...analyze results to determine where frontier models succeed or fail. Typical tasks require one...  ...in AI training, model evaluation, or benchmark or task authoring is preferred. High... 
    Hourly pay
    Full time
    Freelance
    Remote work

    SaidGig

    United States
    19 days ago
  • $15 - $20 per hour

     ...in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel ,...  ...tools . Generate high-quality human evaluation data by identifying response strengths,...  ...and completeness of responses. Ensure model responses align with expected conversational... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    21 days ago
  •  ...company. We build cutting‑edge foundation AI models and end‑to‑end products that are...  ...Germany and Paris. Join us! Why this role? Evaluation is critical to making progress in...  ...Evaluation, you will: Develop evaluation benchmarks, datasets, and environments for measuring... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    cohere

    New York, NY
    1 day ago
  • $228.7k - $343.1k

     ...enormous scale, and one bad model can mean millions in credit...  ...at scale, so you critically evaluate what it produces and own the...  ...conceptual-soundness review, benchmarking, stress testing, and outcomes...  ...across model families, from regression and tree ensembles to deep... 
    Remote job
    Full time
    Local area
    Shift work

    Block

    New York, NY
    15 hours ago
  •  ...Frontier Model Misuse Red Team Specialist is a remote red-team track for stress-testing AI systems against...  ...this role matters Adversarial evaluation is how AuraOne hardens AI models before...  ...team can reproduce, fix, and regress-test them. Responsibilities Design... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    15 days ago
  •  ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities... 
    Hourly pay
    Remote work
    Flexible hours

    Prolific

    Jacksonville, FL
    3 days ago
  • $195.2k - $262.2k

     ...and enterprises from data and model training through to...  ...verifiers, task environments, and evaluation sets for reasoning, coding, tool...  ...reasoning, tool use, safety, and regression risk. Write clear experiment plans, design docs, benchmark reports, and runbooks, and... 
    Full time
    Temporary work
    Immediate start
    Remote work

    Nebius

    Palo Alto, CA
    1 day ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    4 days ago
  • $85 per hour

     ...Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D...  ...frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving cloud... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    a month ago
  • Role Description We are sourcing independent Audio Evaluation Specialists for an AI benchmark evaluation project assessing advanced agentic audio models. The objective of this project is to autonomously produce high-quality evaluation tasks through simulated interactions... 
    Hourly pay
    Contract work
    For contractors
    Freelance
    Remote work

    Meridial

    Remote
    5 days ago
  • $238k - $302k

     ...billions in simulation across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements...  ...-scale simulations. Conduct data analysis to diagnose regressions in ML models. Collaborate with world-class engineering... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    15 hours ago
  • $89 - $95 per hour

    Role Description ORAU is seeking a fully remote Senior Advisor – Payment Model Evaluation to support the Centers for Medicare and Medicaid Services Innovation Center (CMMI) as an ORAU employee. This is a part-time, temporary role expected to last 8 months or longer.... 
    Hourly pay
    Temporary work
    Part time
    Remote work

    ORAU

    Remote
    1 day ago
  •  ...looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team...  ...canaries, and test frameworks that catch regressions, hallucinations, and quality issues before...  .../fail criteria grounded in Research benchmarks, and build the monitoring that keeps... 
    Full time

    Deepgram

    Remote
    16 days ago
  • $40 per hour

    A technology company in Massachusetts is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal candidates will hold advanced qualifications in biology or biochemistry. This position allows... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Oklahoma City, OK
    4 days ago
  •  ...leading AI lab network. The role involves probing frontier models, designing robust evaluation tasks, and documenting findings for reproducibility....  ...and close collaboration with researchers to strengthen benchmark tasks and defenses against failure modes in AI systems.... 
    Full time
    Remote work

    Mercor

    New York, NY
    2 days ago
  • $85 per hour

     ...Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D...  ...frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving cloud... 
    Contract work
    Summer work
    Remote work

    Mercor

    Miami, FL
    11 days ago
  •  ...initiative. You will craft and verify rigorous multiple-choice questions across core chemistry domains, evaluate solution quality, and help establish gold-standard benchmarks for advancing AI capabilities. You will be assigned one of two task types: question authoring or... 
    Remote job

    Mercor

    New York, NY
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Benchmark Regression Model Evaluation Specialist [Remote]. Be the first to apply!