Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Benchmark Regression Model Evaluation Specialist [Remote]

AuraOne Human Data

Remote
  • Remote job

Benchmark Regression Model Evaluation Specialist is a remote evaluation track for reviewing benchmark regression model evaluation evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.

Why this role matters

AI data reviewers help turn benchmark regression model evaluation evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.

Responsibilities

  • Evaluate benchmark regression model evaluation evaluation model outputs against a versioned rubric and assign severity tags for Benchmark Regression Model Evaluation Specialist assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
  • Document recurring failure modes so the modeling team can target them in the next training run.

Qualifications

  • Prior evaluation, annotation, or human-rater experience on benchmark regression model evaluation evaluation or adjacent content for Benchmark Regression Model Evaluation Specialist work.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.

Example tasks

  • Compare two benchmark regression model evaluation evaluation model responses to the same prompt and pick the stronger one with rationale.
  • Tag an unsafe response with the correct policy category and severity.
  • Audit a 50-row batch for rubric consistency and report drift to the program lead.
  • Propose a rubric clarification after spotting a recurring failure mode.

Nice to have

  • Background in linguistics, content moderation, or trust & safety review.
  • Experience with inter-rater agreement metrics and calibration cycles.
  • Domain expertise that lets you spot subject-matter errors automated checks miss.

Skills

  • Model output evaluation
  • Rubric-based annotation
  • Severity tagging
  • Inter-rater calibration
  • Benchmark Regression Model Evaluation evaluation
  • Frontier evaluation
  • Rubric calibration
  • Failure analysis
  • Benchmark
  • Regression

Work model

Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.

Compensation

Hourly rate confirmed after the interview process.

Application process

Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.

Vacancy posted 18 hours ago
Similar jobs that could be interesting for youBased on the Benchmark Regression Model Evaluation Specialist [Remote] in Remote vacancy
  •  ...As the Manager of Model Validation & Verification (VnV) for Behavior Autonomy,...  ...and data science team responsible for evaluating, benchmarking, and validating the machine learning...  ...performance. Statistical uncertainties, and regressions into clear, data-driven... 
    Suggested
    Full time
    Temporary work
    Relocation package

    Zoox

    California
    9 days ago
  • $20 - $60 per hour

     ...Help train next-generation AI systems by creating rigorous, real-world evaluations that test how well advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates, advanced-degree holders, and professionals from any background... 
    Suggested
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    1 day ago
  •  ...Causal Reasoning Model Evaluation Specialist is a remote review track for evaluating AI outputs across causal reasoning model evaluation research review reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results,... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    18 hours ago
  • $295k

     ...company. We build cutting-edge foundation AI models and end-to-end products that are...  ...Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling...  ...you will:Create ambitious new evaluation benchmarks that push the limits of what our models... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    1 day ago
  • $60 - $90 per hour

     ...labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo ,...  ...Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation... 
    Suggested
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    1 day ago
  • $20 per hour

     ...in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo...  ...tools. Generate high-quality human evaluation data by identifying response strengths,...  ..., and completeness of responses. Ensure model responses align with expected conversational... 
    Remote job
    Contract work
    Part time
    Summer work

    Mercor

    New York, NY
    4 days ago
  • $100 per hour

     ...your finance expertise to help improve AI models across complex financial problem-solving...  ...is required. Key Responsibilities Evaluate language models in finance domains where...  ...approaches, evaluation strategies, and benchmarks. Qualifications At least 2 years... 
    Hourly pay
    Contract work
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $70 - $90 per hour

     ...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness...  ...and INT8 data types. Experience benchmarking machine-learning training workloads... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    29 days ago
  • $70 - $90 per hour

     ...accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical correctness, completeness, fair performance benchmarking, appropriate scope, and whether kernels compile and run... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    29 days ago
  •  ...that happens: every product built on a model is bounded by what it costs to run, so the...  ...for: ~ This role builds the evaluation and decision systems that make agentic inference...  ...dataset construction, automated scoring, regression testing, and failure analysis In addition... 
    Full time

    Bitdeer Technologies Group

    Austin, TX
    23 days ago
  •  ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear,...  ...corrections and feedback. Help define new evaluation benchmarks based on mathematics curricula spanning early undergraduate... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    a month ago
  • $60 - $90 per hour

     ...to improve how advanced AI models reason through real-world engineering...  ..., and create rigorous evaluations grounded in industry...  ...workflows. Design challenging benchmarks and evaluation sets spanning...  ...researchers and adjacent-domain specialists to calibrate consistent... 
    Hourly pay
    Full time
    Remote work

    SaidGig

    United States
    6 days ago
  • $90.58k - $114.59k

     ...YesAgency: Mental Health, Office ofTitle: Mental Hygiene Program Evaluation Specialist 3 (NY HELPS), Kingsboro Psychiatric Center, P27987...  ...identify trends and proactively mitigate risks.• Establishing benchmarks for quality assurance.• Coordinating and monitoring survey... 
    Permanent employment
    Full time
    Work at office
    Remote work

    State of New York, USA

    New York, NY
    4 days ago
  • $65 - $105 per hour

     ...judgment to help frontier AI models reason more accurately about...  ...-quality engineering work, evaluate model performance, and turn...  ...practice into clear standards and benchmarks. Key Responsibilities...  ...Collaborate with researchers and specialists in adjacent disciplines to... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    a month ago
  • $65 - $105 per hour

     ...Help advance frontier AI models by bringing rigorous life sciences...  ...judgment into task design, evaluation, and model improvement. You...  ..., reference solutions, benchmarks, and evaluation standards. You...  ...researchers and adjacent-field specialists, translating expert... 
    Hourly pay
    Full time
    Freelance
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    a month ago
  • $70 - $110 per hour

     ...partner with an AI research team to evaluate medical knowledge tasks, define...  ...quality clinical standards, and develop benchmarks that measure meaningful model improvement. Key...  ...with researchers and adjacent-domain specialists to calibrate standards and translate... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    a month ago
  • $228.7k - $343.1k

     ...enormous scale, and one bad model can mean millions in credit...  ...at scale, so you critically evaluate what it produces and own the...  ...conceptual-soundness review, benchmarking, stress testing, and outcomes...  ...across model families, from regression and tree ensembles to deep... 
    Remote job
    Full time
    Local area
    Shift work

    Block

    New York, NY
    1 day ago
  • $147.5k - $211k

     ...President, Data Scientist - Model Validation and AI Governance...  ...challenging model methodologies, evaluation approaches, controls, and...  ...challenging evaluation frameworks, benchmark and golden datasets, rubric-...  ..., human review protocols, regression suites, offline and online testing... 
    Local area
    3 days per week

    New York Life Insurance Company

    New York, NY
    3 days ago
  •  ...leading AI lab network. The role involves probing frontier models, designing robust evaluation tasks, and documenting findings for reproducibility....  ...and close collaboration with researchers to strengthen benchmark tasks and defenses against failure modes in AI systems.... 
    Full time
    Remote work

    Mercor

    New York, NY
    6 days ago
  •  ...write and verify rigorous multiple-choice questions across core history and political science domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. You will be assigned two task types: Question Authoring and... 
    Remote job

    Mercor Inc

    New York, NY
    6 days ago
  • $20 per hour

    SupportFinity™ in Maine is seeking an Editorial Proofreader to join their team focused on training AI models. The position involves evaluating AI chatbot outputs and improving model quality through expert editing and writing skills. This flexible role allows you to work... 
    Remote job
    Hourly pay
    Flexible hours

    SupportFinity™

    Montgomery, AL
    2 days ago
  • $70 - $80 per hour

     ...Role Overview Apply advanced drug safety expertise to help improve AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data. This remote contract role focuses on the quality, accuracy, and regulatory alignment of complex safety reports... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    a month ago
  •  ...Refusal Preference Reward Model Evaluator is a remote red-team track for stress-testing AI systems...  ...modeling team can reproduce, fix, and regress-test them. Responsibilities...  ...Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    13 days ago
  • $20 per hour

    SupportFinity™ is looking for an Editorial Proofreader to join our team for AI model training. In this remote role, you'll evaluate AI chatbots and enhance model quality. Candidates should have fluency in English and strong editing skills. This position can be full‑time... 
    Remote job
    Hourly pay
    Full time
    Contract work
    Part time

    SupportFinity™

    New York, NY
    6 days ago
  •  ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities... 
    Remote job
    Hourly pay
    Flexible hours

    Prolific

    Jacksonville, FL
    6 days ago
  • Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote job
    Flexible hours

    Prolific

    Charlotte, NC
    6 days ago
  • $60 - $70 per hour

     ...alignment, and overall quality of frontier AI model outputs on complex, policy sensitive,...  ...area" topics. Work through structured evaluations to identify unsafe behavior, reasoning...  ...used for RLHF, SFT, and AI safety benchmarking. Identify unsafe outputs, hallucinations... 
    Hourly pay
    Remote work

    SaidGig

    United States
    1 day ago
  • $20 per hour

    SupportFinity™ is seeking an Editorial Proofreader to evaluate AI models and improve their quality through expert writing and editing skills. This role can be part‑time or full‑time, allowing for a flexible schedule and project selection. Applicants must be fluent in English... 
    Remote job
    Hourly pay
    Full time
    Part time
    Flexible hours

    SupportFinity™

    Sioux Falls, SD
    2 days ago
  • SupportFinity™ is looking for an Editorial Proofreader to join our team to train AI models. In this role, you will measure AI chatbot progress, evaluate logic, and solve problems to enhance model quality. Applicants should have a strong command of English and experience... 
    Remote job
    Hourly pay
    Full time
    Part time
    Flexible hours

    SupportFinity™

    Columbia, SC
    2 days ago
  • $60 per hour

    Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Benchmark Regression Model Evaluation Specialist [Remote]. Be the first to apply!