Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Model Grading Model Evaluation Specialist [Remote]

AuraOne Human Data

Remote
  • Remote job

Model Grading Model Evaluation Specialist is a remote evaluation track for reviewing model grading model evaluation evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.

Why this role matters

AI data reviewers help turn model grading model evaluation evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.

Responsibilities

  • Evaluate model grading model evaluation evaluation model outputs against a versioned rubric and assign severity tags for Model Grading Model Evaluation Specialist assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
  • Document recurring failure modes so the modeling team can target them in the next training run.

Qualifications

  • Prior evaluation, annotation, or human-rater experience on model grading model evaluation evaluation or adjacent content for Model Grading Model Evaluation Specialist work.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.

Example tasks

  • Compare two model grading model evaluation evaluation model responses to the same prompt and pick the stronger one with rationale.
  • Tag an unsafe response with the correct policy category and severity.
  • Audit a 50-row batch for rubric consistency and report drift to the program lead.
  • Propose a rubric clarification after spotting a recurring failure mode.

Nice to have

  • Background in linguistics, content moderation, or trust & safety review.
  • Experience with inter-rater agreement metrics and calibration cycles.
  • Domain expertise that lets you spot subject-matter errors automated checks miss.

Skills

  • Model output evaluation
  • Rubric-based annotation
  • Severity tagging
  • Inter-rater calibration
  • Model Grading Model Evaluation evaluation
  • Frontier evaluation
  • Rubric calibration
  • Failure analysis
  • Model
  • Grading

Work model

Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.

Compensation

Hourly rate confirmed after the interview process.

Application process

Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Model Grading Model Evaluation Specialist [Remote] in Remote vacancy
  • $100 - $150 per hour

     ...will be considered for future projects evaluating how well AI systems perform real-world data...  ...AI or human-produced analyses and models, document decisions in writing, and iterate...  ...Responsibilities Design precise, task-specific grading criteria for data science deliverables,... 
    Suggested
    Hourly pay
    Immediate start
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $60 - $90 per hour

     ...Overview Design rigorous, real-world materials science and engineering tasks that evaluate an AI model’s expert reasoning. You will create prompts, supporting data rooms, and grading criteria, then test and refine each task until it clearly distinguishes strong from... 
    Suggested
    Hourly pay
    Work at office
    Remote work

    SaidGig

    United States
    27 days ago
  • $20 - $60 per hour

     ...Help train next-generation AI systems by creating rigorous, real-world evaluations that test how well advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates, advanced-degree holders, and professionals from any background... 
    Suggested
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    2 days ago
  •  ...Causal Reasoning Model Evaluation Specialist is a remote review track for evaluating AI outputs across causal reasoning model evaluation research...  ...reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    4 days ago
  •  ...Research Engineer - Code Generation & Model Evaluation is a remote engineering review track for...  ...experienced engineers with the modeling team to grade outputs the way a code reviewer would....  ...— US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    4 days ago
  • $60 - $90 per hour

     ...Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation: $60–$90/hour Location: Remote Commitment:... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    a month ago
  • $100 per hour

     ...Role Overview Apply your finance expertise to help improve AI models across complex financial problem-solving areas, including capital...  ...No prior AI experience is required. Key Responsibilities Evaluate language models in finance domains where performance needs... 
    Hourly pay
    Contract work
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $70 - $90 per hour

     ...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness, CUDA-to-NKI migration fidelity, and whether implementations are well suited to... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    a month ago
  • $70 - $90 per hour

     ...Review GPU and accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical correctness, completeness, fair performance benchmarking, appropriate scope, and whether kernels... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    a month ago
  •  ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets... 
    For contractors
    Remote work

    SaidGig

    United States
    a month ago
  •  ...As the Manager of Model Validation & Verification (VnV) for Behavior Autonomy, you will lead an engineering and data science team responsible for evaluating, benchmarking, and validating the machine learning models and behavioral algorithms that drive our autonomous vehicle... 
    Full time
    Temporary work
    Relocation package

    Zoox

    California
    13 days ago
  •  ...Apply your data science and quantitative expertise to improve how AI models reason through statistics, machine learning, experimentation, and analytical problems. Key Responsibilities Evaluate AI model outputs for data science, statistics, machine learning, and quantitative... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    a month ago
  • $100 - $150 per hour

     ...Role Overview Apply your data science expertise to evaluate AI-generated slides, spreadsheets, and documents for real-world quality and...  ...data science contexts. Use deep subject-matter expertise to grade output quality. Identify factual, aesthetic, and... 
    Hourly pay
    Work at office
    Remote work

    SaidGig

    United States
    5 days ago
  • $20 per hour

    SupportFinity™ in Maine is seeking an Editorial Proofreader to join their team focused on training AI models. The position involves evaluating AI chatbot outputs and improving model quality through expert editing and writing skills. This flexible role allows you to work... 
    Remote job
    Hourly pay
    Flexible hours

    SupportFinity™

    Montgomery, AL
    1 day ago
  • $20 per hour

    SupportFinity™ is seeking an Editorial Proofreader to evaluate AI models and improve their quality through expert writing and editing skills. This role can be part‑time or full‑time, allowing for a flexible schedule and project selection. Applicants must be fluent in English... 
    Remote job
    Hourly pay
    Full time
    Part time
    Flexible hours

    SupportFinity™

    Sioux Falls, SD
    1 day ago
  • SupportFinity™ is looking for an Editorial Proofreader to join our team to train AI models. In this role, you will measure AI chatbot progress, evaluate logic, and solve problems to enhance model quality. Applicants should have a strong command of English and experience... 
    Remote job
    Hourly pay
    Full time
    Part time
    Flexible hours

    SupportFinity™

    Columbia, SC
    1 day ago
  • $20 per hour

     ...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths, areas for improvement,...  ...quality, clarity, tone, and completeness of responses. Ensure model responses align with expected conversational behavior and system... 
    Remote job
    Contract work
    Part time
    Summer work

    Mercor

    New York, NY
    3 days ago
  • $20 per hour

    SupportFinity™ is looking for an Editorial Proofreader to join our team for AI model training. In this remote role, you'll evaluate AI chatbots and enhance model quality. Candidates should have fluency in English and strong editing skills. This position can be full‑time... 
    Remote job
    Hourly pay
    Full time
    Contract work
    Part time

    SupportFinity™

    New York, NY
    5 days ago
  •  ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities... 
    Remote job
    Hourly pay
    Flexible hours

    Prolific

    Jacksonville, FL
    5 days ago
  • Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote job
    Flexible hours

    Prolific

    Charlotte, NC
    5 days ago
  • $40 - $80 per hour

     ...operational updates. Your work will help models learn to produce effective, compliant government...  ...responses for AI training. Create grading rubrics and assessment guidelines based...  ...and military writing best practices. Evaluate AI-generated reports and provide detailed... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Remote
    a month ago
  • $60 per hour

    Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    1 day ago
  •  ...0 per hour. In this role, you will help a leading AI lab improve frontier AI systems by evaluating AI-generated code, creating challenging engineering problems, and supporting model performance on real-world software tasks. Scope of Work Evaluate AI-generated code... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    United States
    a month ago
  •  ...Role Title: Research Engineer - Code Generation & Model Evaluation Role Type: Contractor Location: Remote micro1 is engaging Research Engineers to participate in a project focused on code generation and model evaluation for a customer's initiative. In this role... 
    Temporary work
    For contractors
    Remote work

    micro1

    Remote
    9 days ago
  •  ...to work with a leading AI lab's GenAI team, shaping how frontier models reason about real design and manufacturing tasks. You will review task quality, write instruction specs, and help define evaluation criteria in a hands-on, outcomes-focused role. This remote, US-based... 
    Remote job

    Mercor

    New York, NY
    5 days ago
  •  ...use. Inference is where that happens: every product built on a model is bounded by what it costs to run, so the economics of serving...  ...What you will be responsible for: ~ This role builds the evaluation and decision systems that make agentic inference measurable, reliable... 
    Full time

    Bitdeer Technologies Group

    Austin, TX
    27 days ago
  •  ...Role Overview Apply advanced mathematical reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically structured solutions, verify numerical results with code, and review... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $60 - $90 per hour

     ...engineering judgment to improve how advanced AI models reason through real-world engineering...  ...correct solutions, and create rigorous evaluations grounded in industry practice. Key...  ...with researchers and adjacent-domain specialists to calibrate consistent standards and... 
    Hourly pay
    Full time
    Remote work

    SaidGig

    United States
    9 days ago
  • $50 - $100 per hour

     ...Apply your software engineering expertise to help train and evaluate next-generation AI systems through real-world coding tasks, technical...  ...remote contract role focuses on code generation workflows and model evaluation; prior AI experience is not required. Key... 
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    13 days ago
  • $60 - $80 per hour

     ...specific role. Role Overview Qualified experts may support AI research by applying mathematics knowledge to real-world tasks, model evaluation, and domain-specific feedback. Key Responsibilities Train and evaluate AI models in mathematics. Create tasks and... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Model Grading Model Evaluation Specialist [Remote]. Be the first to apply!