Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Multi-Step Reasoning Model Evaluation Specialist [Remote]

AuraOne Human Data

Remote
  • Remote job

Multi-Step Reasoning Model Evaluation Specialist is a remote evaluation track for reviewing multi step reasoning model evaluation evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.

Why this role matters

AI data reviewers help turn multi step reasoning model evaluation evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.

Responsibilities

  • Evaluate multi step reasoning model evaluation evaluation model outputs against a versioned rubric and assign severity tags for Multi-Step Reasoning Model Evaluation Specialist assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
  • Document recurring failure modes so the modeling team can target them in the next training run.

Qualifications

  • Prior evaluation, annotation, or human-rater experience on multi step reasoning model evaluation evaluation or adjacent content for Multi-Step Reasoning Model Evaluation Specialist work.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.

Example tasks

  • Compare two multi step reasoning model evaluation evaluation model responses to the same prompt and pick the stronger one with rationale.
  • Tag an unsafe response with the correct policy category and severity.
  • Audit a 50-row batch for rubric consistency and report drift to the program lead.
  • Propose a rubric clarification after spotting a recurring failure mode.

Nice to have

  • Background in linguistics, content moderation, or trust & safety review.
  • Experience with inter-rater agreement metrics and calibration cycles.
  • Domain expertise that lets you spot subject-matter errors automated checks miss.

Skills

  • Model output evaluation
  • Rubric-based annotation
  • Severity tagging
  • Inter-rater calibration
  • Multi Step Reasoning Model Evaluation evaluation
  • Frontier evaluation
  • Rubric calibration
  • Failure analysis
  • Multi
  • Step
  • Reasoning

Work model

Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.

Compensation

Hourly rate confirmed after the interview process.

Application process

Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Multi-Step Reasoning Model Evaluation Specialist [Remote] in Remote vacancy
  •  ...Causal Reasoning Model Evaluation Specialist is a remote review track for evaluating AI outputs across causal...  ...checking assumptions, reproducing key steps, and capturing the right method...  ...professional level. Comfort applying multi-page rubrics consistently across long... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  •  ...Pull Request Reasoning Evaluation Specialist is a remote review track for evaluating AI outputs across...  ...risk; and document the right next step so the modeling team can train on it. Why this role...  ...Specialist work. Comfort applying multi-page rubrics consistently across... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  •  ...Quant Finance Reasoning Model Evaluator is a remote review track for evaluating AI outputs across finance...  ...experience. Comfort applying multi-page rubrics consistently across long...  ...— US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  • $80 - $150 per hour

     ...Professor), you will play a critical role in evaluating and enhancing the training of next-generation AI models, ensuring they learn and reason effectively through your domain knowledge...  ...platforms. Detect errors, unjustified steps, missing assumptions, dimensional... 
    Suggested
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Indiana
    14 days ago
  • $17 per hour

     ...AI Evaluation Specialists contribute to the advancement of Large Language Models (LLMs) by testing and providing feedback in collaboration with leading AI labs. This role...  ...labs. A belief that your knowledge and reasoning skills can challenge today’s most advanced AI... 
    Suggested
    Temporary work
    Part time
    Remote work

    SaidGig

    United States
    9 days ago
  • $20 per hour

     ...Generate high-quality human evaluation data by identifying response...  ...factual inaccuracies. Assess reasoning quality, clarity, tone, and...  ...completeness of responses. Ensure model responses align with expected...  ...AI interview and application steps to be considered for this... 
    Remote job
    Contract work
    Part time
    Summer work

    Mercor

    San Francisco, CA
    a month ago
  • $60 - $90 per hour

     ...driven sales work. You will audit multi-step sales workflows, produce end-to-end expert examples the model learns from, and shape the criteria used to evaluate AI outputs. This role is remote...  ...Opportunity employer and will provide reasonable accommodations for qualified... 
    Remote job
    Hourly pay
    Full time
    Part time
    Freelance
    Work at office

    SaidGig

    Remote
    1 day ago
  • $50 per hour

     ...fine-tune large language models, enhancing their...  ...Develop high-quality, step-by-step solutions with clear and rigorous reasoning. Collaborate with LLM...  ...align problem types with evaluation goals, particularly in...  ...conceptual abstraction, multi-step reasoning, and data... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    a month ago
  • $50 per hour

     ...opportunity to engage with advanced mathematical reasoning and problem-solving in the context of improving and evaluating large language models. You will apply your expertise to design...  ...large language models, particularly in multi-step, abstract, and proof-based settings.... 
    Contract work
    Freelance
    Remote work

    SaidGig

    United States
    9 days ago
  • $350k

     ...pivotal role in shaping the future of AI-powered legal reasoning. This position focuses on the intersection of large language models, agentic systems, and legal workflows, emphasizing the development of rigorous evaluation frameworks to measure and enhance AI performance in... 
    Remote job
    Full time

    SaidGig

    United States
    3 days ago
  •  ...Frontier Model Misuse Red Team Specialist is a remote red-team track for stress-testing...  ...matters Adversarial evaluation is how AuraOne hardens AI...  ...attack with reproduction steps and the policy clause it violated...  ...across single-turn and multi-turn conversations.... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  • $70 - $90 per hour

     ...and low-level programming experts to apply their knowledge in systems programming and security concepts to enhance AI models'' ability to detect and reason about potential threats. The opportunity begins with a work trial and may extend into a two-month project based on... 
    Hourly pay
    Remote work

    SaidGig

    United Kingdom
    26 days ago
  • $15 - $20 per hour

     ...external tools. ~Generate high-quality human evaluation data by identifying response strengths,...  ...improvement, and factual inaccuracies. ~Assess reasoning quality, clarity, tone, and completeness of responses. ~Ensure model responses align with expected conversational... 
    Part time
    Summer work

    Mercor

    Remote
    19 days ago
  • $105 per hour

     ...collaborating with a community of experts to refine and evaluate the capabilities of Large Language Models (LLMs) in creating impactful presentations and...  ...AI to enhance project outcomes. Collaborate with multi-disciplinary teams to translate complex engineering concepts... 
    Work experience placement
    Remote work
    Flexible hours

    SaidGig

    United States
    3 days ago
  •  ...judgment and frontline medical experience to evaluate AI-generated medical content, ensuring clinical accuracy, sound diagnostic reasoning, appropriate treatment considerations, and...  ...Assess clarity, relevance, and safety of model outputs in realistic care scenarios.... 
    Full time
    For contractors
    Private practice
    Remote work
    Flexible hours

    SaidGig

    United States
    3 days ago
  • $60 per hour

     ...professionals to help advance AI development. AI models are increasingly capable of performing complex analytical and scientific reasoning — but these systems still need...  ...state-of-the-art AI models on tasks like evaluating AI-generated quantitative analysis, solving... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Little Rock, AR
    1 day ago
  • $40 per hour

     ...firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully...  ...AI systems' development. Join us to shape the future of AI reasoning while working from anywhere in the US, Canada, UK, Ireland, Australia... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Bismarck, ND
    15 hours ago
  • $40 per hour

     ...seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the evolution of AI models. Candidates should have a background in data...  ...problems and provide technical feedback on AI reasoning systems, shaping the future of data analytics.... 
    Hourly pay
    Remote work

    DataAnnotation

    Lansing, MI
    15 hours ago
  • $40 per hour

     ...company is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the development of cutting...  ...or statistics and strong coding skills. Join us to directly impact the future of AI analytics and model reasoning. #J-18808-Ljbffr... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Lincoln, NE
    15 hours ago
  • $40 - $65 per hour

     ...high-impact project focused on the evaluation and enhancement of frontier language models as an LLM Red-Teamer. In this...  ...systems, shaping how models learn, reason, and perform through high-quality...  ...Develop complex, adversarial multi-turn conversations and task-based... 
    Remote job
    Hourly pay
    For contractors
    Immediate start

    SaidGig

    Frontier County, NE
    13 days ago
  • $85 per hour

     ...well planning, and energy data systems to evaluate AI-generated content and create expert...  ...knowledge to judge accuracy and relevance of model outputs and deliver clear feedback that...  ...verification as part of onboarding. Application steps include creating an account, completing... 
    Hourly pay
    Contract work
    Part time
    Work at office
    Remote work
    Flexible hours

    SaidGig

    United States
    6 days ago
  • A data technology company is seeking a Statistician to enhance AI models by evaluating their logic and progress. The role demands expert-level mathematical reasoning and offers a flexible working schedule, allowing for full-time or part-time remote work. Responsibilities... 
    Hourly pay
    Full time
    Part time
    Remote work
    Flexible hours

    DataAnnotation

    Raleigh, NC
    15 hours ago
  •  ...data services company is seeking an Applied Mathematician to evaluate AI models by providing complex mathematical problems to chatbots and...  ...schedule. Applicants should possess expert-level mathematical reasoning and fluency in English. Payment is hourly, starting at $40,... 
    Hourly pay
    Remote work

    DataAnnotation

    New York, NY
    2 days ago
  •  ...Role Overview Psychiatrists evaluate AI-generated psychiatric content, use clinical judgment to assess model responses, and deliver clear, structured feedback that improves...  ..., treatment recommendations, diagnostic reasoning, and patient-facing language. Develop prompts... 
    Hourly pay
    Temporary work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    5 days ago
  • $40 per hour

     ...tech firm is seeking experienced quantitative professionals to evaluate AI-generated analyses and enhance AI development. The role...  ...This opportunity impacts the evolution of AI systems focused on reasoning and data analytics. Applicants can work alongside other roles... 
    Remote job
    Hourly pay

    DataAnnotation

    Brooklyn, NY
    1 day ago
  • $40 per hour

     ...development company seeks experienced quantitative professionals to evaluate AI-generated work and solve quantitative problems. This fully...  ...us to contribute to shaping the future of AI systems focused on reasoning about data and analytics. #J-18808-Ljbffr DataAnnotation
    Remote job
    Hourly pay
    Flexible hours

    DataAnnotation

    Little Rock, AR
    4 days ago
  • $40 per hour

     ...for experienced quantitative professionals to evaluate AI-generated quantitative work and contribute...  ...familiarity with statistical methods and predictive modeling. Join us and help shape the future of AI systems built to reason about data and analytics. #J-18808-Ljbffr... 
    Remote job
    Hourly pay
    Flexible hours

    DataAnnotation

    Boston, MA
    4 days ago
  • $40 per hour

     ...company is seeking experienced quantitative professionals to evaluate AI-generated quantitative work, solve quantitative problems, and...  ...quantitative field. Join a team that helps shape the future of AI systems tailored for data reasoning. #J-18808-Ljbffr DataAnnotation
    Remote job
    Hourly pay
    Flexible hours

    DataAnnotation

    Juneau, AK
    4 days ago
  • $40 per hour

     ...seeking experienced quantitative professionals to evaluate AI-generated work and help advance AI development...  ...the development of AI systems while influencing their ability to reason about data. Join us to help shape the future of AI models! #J-18808-Ljbffr DataAnnotation
    Remote job
    Hourly pay

    DataAnnotation

    New York, NY
    1 day ago
  •  ...contribute to AI research projects that enhance the understanding of workplace tasks and language in their field. This role involves evaluating AI model outputs, assessing content related to your profession, and providing structured feedback to improve AI performance. The... 
    Remote work
    Flexible hours

    SaidGig

    United States
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Multi-Step Reasoning Model Evaluation Specialist [Remote]. Be the first to apply!