Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Rubric Calibration Model Evaluation Specialist [Remote]

AuraOne Human Data

Remote
  • Remote job

Rubric Calibration Model Evaluation Specialist is a remote evaluation track for reviewing rubric calibration model evaluation evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.

Why this role matters

AI data reviewers help turn rubric calibration model evaluation evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.

Responsibilities

  • Evaluate rubric calibration model evaluation evaluation model outputs against a versioned rubric and assign severity tags for Rubric Calibration Model Evaluation Specialist assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
  • Document recurring failure modes so the modeling team can target them in the next training run.

Qualifications

  • Prior evaluation, annotation, or human-rater experience on rubric calibration model evaluation evaluation or adjacent content for Rubric Calibration Model Evaluation Specialist work.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.

Example tasks

  • Compare two rubric calibration model evaluation evaluation model responses to the same prompt and pick the stronger one with rationale.
  • Tag an unsafe response with the correct policy category and severity.
  • Audit a 50-row batch for rubric consistency and report drift to the program lead.
  • Propose a rubric clarification after spotting a recurring failure mode.

Nice to have

  • Background in linguistics, content moderation, or trust & safety review.
  • Experience with inter-rater agreement metrics and calibration cycles.
  • Domain expertise that lets you spot subject-matter errors automated checks miss.

Skills

  • Model output evaluation
  • Rubric-based annotation
  • Severity tagging
  • Inter-rater calibration
  • Rubric Calibration Model Evaluation evaluation
  • Frontier evaluation
  • Rubric calibration
  • Failure analysis
  • Rubric
  • Calibration

Work model

Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.

Compensation

Hourly rate confirmed after the interview process.

Application process

Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Rubric Calibration Model Evaluation Specialist [Remote] in Remote vacancy
  •  ...Causal Reasoning Model Evaluation Specialist is a remote review track for evaluating AI outputs across...  ...reviewer-quality scores in inter-rater calibration cycles. Qualifications Graduate...  .... Comfort applying multi-page rubrics consistently across long batches.... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    3 days ago
  • Dorado is seeking a Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems, probe model reasoning at the frontier, and help identify where models fail under rigorous scientific scrutiny.... 
    Suggested
    Remote job

    Dorado

    New York, NY
    1 day ago
  •  ...Calibration Technology Evaluation Specialist is a remote review track for evaluating AI outputs across calibration...  ...the right next step so the modeling team can train on it. Why this role...  ...and stakeholder fit on a structured rubric. Flag operational risk, missed... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    a month ago
  •  ...Frontier Model Misuse Red Team Specialist is a remote red-team track for stress-testing...  ...jailbreak with the rubric clause it violated so the...  ...role matters Adversarial evaluation is how AuraOne hardens AI...  ...new red-team rubrics. Calibrate against the broader red-team... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    29 days ago
  • $70 - $90 per hour

     ...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness...  ...Trainium environments. Provide clear, rubric-based written feedback.... 
    Suggested
    Hourly pay
    Remote work

    SaidGig

    Remote
    9 days ago
  • $70 - $90 per hour

     ...accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical...  ...provide clear written feedback using defined evaluation rubrics. Key Responsibilities Evaluate GPU and accelerator... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    9 days ago
  •  ...Quant Finance Reasoning Model Evaluator is a remote review track for...  ...quality scores in inter-rater calibration cycles. Qualifications...  ...Comfort applying multi-page rubrics consistently across long batches...  .... Remote · Independent specialist contractor. Employment type... 
    Remote job
    Hourly pay
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    a month ago
  •  ...Refusal Preference Reward Model Evaluator is a remote red-team track for...  ...successful jailbreak with the rubric clause it violated so the...  ...propose new red-team rubrics. Calibrate against the broader red-team...  .... Remote · Independent specialist contractor. Employment type:... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    a month ago
  •  ...Policy Preference Reward Model Evaluator is a remote review track for...  ...quality scores in inter-rater calibration cycles. Qualifications...  ...Comfort applying multi-page rubrics consistently across long...  ...eligible. Remote · Independent specialist contractor. Employment type... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  •  ...Medical Document OCR Model Evaluator is a remote clinical-review track...  ...flag handling on a structured rubric. Flag patient-safety issues...  ...in weekly inter-rater calibration cycles. Qualifications...  ...eligible. Remote · Independent specialist contractor. Employment type:... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    6 days ago
  • $100 - $150 per hour

     ...will be considered for future projects evaluating how well AI systems perform real-world data...  ...AI or human-produced analyses and models, document decisions in writing, and iterate...  ...work Comfort receiving feedback and calibrating judgment against established standards... 
    Hourly pay
    Immediate start
    Remote work

    SaidGig

    United States
    a month ago
  •  ...that happens: every product built on a model is bounded by what it costs to run, so the...  ...for: ~ This role builds the evaluation and decision systems that make agentic inference...  ...and routing, uncertainty estimation and calibration, cascading and escalation, or stage-aware... 
    Full time

    Bitdeer Technologies Group

    Austin, TX
    3 days ago
  • $100 per hour

     ...-generated content for a customer-facing evaluation and optimization project. You will help shape...  .... Develop and refine large language model prompts using structured problem-solving...  ...content reviews, quality assurance, and rubric-based assessments of AI outputs. Annotate... 
    Hourly pay
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    13 days ago
  • $65 - $105 per hour

     ...engineering judgment to help frontier AI models reason more accurately about real-...  ...high-quality engineering work, evaluate model performance, and turn expert...  ...Collaborate with researchers and specialists in adjacent disciplines to calibrate standards and translate tacit... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    13 days ago
  • $65 - $105 per hour

     ...Help advance frontier AI models by bringing rigorous life sciences...  ...judgment into task design, evaluation, and model improvement. You will...  ...with the research team. Calibrate standards with client researchers and adjacent-field specialists, translating expert scientific... 
    Hourly pay
    Full time
    Freelance
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    13 days ago
  •  ...help train next-generation AI systems. Your work will shape how models learn, reason, and perform through high-quality, real-world...  ...patterns and problem-solving strategies in backend engineering. Evaluate real-world Node.js scenarios to help models learn robust, scalable... 
    Remote job
    Contract work

    YO IT Consulting

    Austin, CO
    3 days ago
  •  ...first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world...  ...City, Montreal, Seoul, Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling intelligence. As... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    1 day ago
  • $100 per hour

     ...improve the performance of large language models on finance tasks. You will work with AI...  ...advanced AI systems. Key Responsibilities Evaluate LLM performance in finance areas where...  ...errors. Design clear evaluation rubrics and assessment criteria for finance-specific... 
    Hourly pay
    Contract work
    For contractors
    Freelance
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $70 - $110 per hour

     ...clinical judgment to improve how frontier AI models reason about real medical work. In...  ...standards for correct answers, and evaluate whether models are making meaningful progress...  ...with researchers and adjacent-domain specialists to calibrate consistent standards and translate... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    13 days ago
  • $100 per hour

    A leading technology firm is seeking finance experts to enhance AI models. Responsibilities include evaluating performance in capital markets and creating assessment rubrics. Candidates should have 2+ years in finance fields like investment banking and possess strong financial... 
    Remote job
    Hourly pay
    10 hours per week

    Turing

    Seattle, WA
    4 days ago
  • $60 - $90 per hour

     ...Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation: $60–$90/hour Location: Remote Commitment:... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    9 hours ago
  •  ...across film, television, digital, or live events to support an AI evaluation project in a fully remote capacity. You will craft realistic...  ..., contracts, crew, and vendor coordination skills to the table, delivering clear evaluation rubrics #J-18808-Ljbffr YO AI Labs
    Remote work

    YO AI Labs

    Austin, TX
    3 days ago
  • $91k - $169k

     ...the company’s success. As a Portfolio Analytics & Strategy Specialist (Fraud Model Analyst) within PNC's Technology organization, you will be...  ...monitoring, QC findings, and model controls.• Leads in evaluating identified model risks, defects, and issues, and reaches conclusions... 
    Full time
    Temporary work
    Part time
    Work experience placement
    Work at office
    Remote work

    The PNC Financial Services Group

    Denver, CO
    5 days ago
  • $11 - $19 per hour

     ...Evaluate AI-generated music and lyrics across a broad range of genres, applying your Bengali music expertise to help assess quality, originality, and natural expression. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    22 days ago
  •  ...Role Overview Apply research-grade expertise to help evaluate and improve AI reasoning across technical and humanities disciplines. This remote contractor role supports AI-model training through rigorous analysis, high-quality feedback, and clearly articulated academic... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    United States
    8 days ago
  • $15 per hour

     ...Evaluate AI-generated music and lyrics in Malayalam and English, helping assess outputs across a broad range of genres against detailed quality standards. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities. Rate lyrics... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    11 days ago
  • $28 - $60 per hour

     ...Evaluate AI-generated music and lyrics across a wide range of genres, applying your knowledge of the Dutch music scene and strong editorial judgment to detailed quality standards. Key Responsibilities Assess AI-generated music and rate it against detailed quality... 
    Hourly pay
    For contractors
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    11 days ago
  • $20 per hour

     ...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths, areas for improvement,...  ...quality, clarity, tone, and completeness of responses. Ensure model responses align with expected conversational behavior and system... 
    Remote job
    Contract work
    Part time
    Summer work

    Mercor

    New York, NY
    4 days ago
  • $20 - $60 per hour

     ...Role Overview Help train and evaluate next-generation AI systems by creating rigorous, real-world assessments that test how advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates and other researchers and writers with... 
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    11 days ago
  • Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote job
    Flexible hours

    Prolific

    Charlotte, NC
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Rubric Calibration Model Evaluation Specialist [Remote]. Be the first to apply!