Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Tool-Use Reasoning Model Evaluation Specialist [Remote]

AuraOne Human Data

Remote
  • Remote job

Tool-Use Reasoning Model Evaluation Specialist is a remote evaluation track for reviewing tool use reasoning model evaluation evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.

Why this role matters

AI data reviewers help turn tool use reasoning model evaluation evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.

Responsibilities

  • Evaluate tool use reasoning model evaluation evaluation model outputs against a versioned rubric and assign severity tags for Tool-Use Reasoning Model Evaluation Specialist assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
  • Document recurring failure modes so the modeling team can target them in the next training run.

Qualifications

  • Prior evaluation, annotation, or human-rater experience on tool use reasoning model evaluation evaluation or adjacent content for Tool-Use Reasoning Model Evaluation Specialist work.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.

Example tasks

  • Compare two tool use reasoning model evaluation evaluation model responses to the same prompt and pick the stronger one with rationale.
  • Tag an unsafe response with the correct policy category and severity.
  • Audit a 50-row batch for rubric consistency and report drift to the program lead.
  • Propose a rubric clarification after spotting a recurring failure mode.

Nice to have

  • Background in linguistics, content moderation, or trust & safety review.
  • Experience with inter-rater agreement metrics and calibration cycles.
  • Domain expertise that lets you spot subject-matter errors automated checks miss.

Skills

  • Model output evaluation
  • Rubric-based annotation
  • Severity tagging
  • Inter-rater calibration
  • Tool Use Reasoning Model Evaluation evaluation
  • Frontier evaluation
  • Rubric calibration
  • Failure analysis
  • Tool
  • Reasoning

Work model

Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.

Compensation

Hourly rate confirmed after the interview process.

Application process

Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Tool-Use Reasoning Model Evaluation Specialist [Remote] in Remote vacancy
  •  ...Causal Reasoning Model Evaluation Specialist is a remote review track for evaluating AI outputs across causal...  ...actually hold up under scrutiny. AuraOne uses scientific specialists to grade...  ...work reviewing AI-assisted research tooling and its failure modes. Multilingual... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  •  ...Quant Finance Reasoning Model Evaluator is a remote review track for evaluating...  ...for fuzzy reasoning. AuraOne uses experienced finance-and-...  ...assisted close, audit, or risk tooling and its failure modes....  ...eligible. Remote · Independent specialist contractor. Employment type... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  •  ...Pull Request Reasoning Evaluation Specialist is a remote review track for evaluating...  ...the right next step so the modeling team can train on it. Why...  ...actual day at work. AuraOne uses experienced operators to grade...  ...with AI-assisted workflow tooling and its failure modes. Bilingual... 
    Suggested
    Remote job
    Hourly pay
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  • $20 per hour

     ...Responsibilities Conduct fact-checking using trusted public sources and external tools. Generate high-quality human evaluation data by identifying...  ...inaccuracies. Assess reasoning quality, clarity, tone, and...  ...completeness of responses. Ensure model responses align with... 
    Suggested
    Remote job
    Contract work
    Part time
    Summer work

    Mercor

    New York, NY
    2 days ago
  • $350k

     ...the future of AI-powered legal reasoning. This position focuses on the...  ...of large language models, agentic systems, and legal workflows...  ...the development of rigorous evaluation frameworks to measure and enhance...  ...failure modes across legal use cases. Publish internal research... 
    Suggested
    Remote job
    Full time

    SaidGig

    United States
    3 days ago
  • $85 per hour

     ...and renewable energy siting to evaluate AI-generated geospatial...  ...expert-level training examples. Use your practical knowledge of typical...  ...GIS Analyst, Interconnection Specialist, Solar Design Engineer, or...  ...one or more of the following tools: ArcGIS, Oracle MDM, Aurora Solar... 
    Hourly pay
    Contract work
    Part time
    Work at office
    Remote work
    Flexible hours

    SaidGig

    United States
    3 days ago
  • $125 per hour

     ...QGIS and spatial analysis skills to evaluate AI-generated GIS outputs and help improve how models handle spatial data workflows. In this contract role you will use hands-on experience with...  ...workflows using QGIS and related tools. Evaluate large language model... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    3 days ago
  • $60 per hour

     ...help advance AI development. AI models are increasingly capable of...  ...analytical and scientific reasoning — but these systems still need...  ...-art AI models on tasks like evaluating AI-generated quantitative analysis...  ...solve quantitative problems used to train and benchmark AI... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Little Rock, AR
    1 day ago
  • $30 - $90 per hour

     ...next generation of developer tools alongside a high-caliber engineering...  ...Actively test new AI-powered models in Cursor, providing...  ...schema design. ~ Extensive use of AI tools for coding; familiarity...  ...). Experience designing or evaluating experimental tooling and developer... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    3 days ago
  • $40 per hour

     ...our team to help train AI models. In this role, you will evaluate AI-generated security content...  ...to improve how AI systems reason about real-world threats...  ...to building more reliable tools for the cybersecurity...  ...focused technical problems used to train AI systems Write... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Juneau, AK
    1 day ago
  • $5,540 - $5,780 per month

     ...Admissions and Program Evaluations (GAPE) in the College...  ...schedule. The Evaluation Specialist is responsible for...  ...applying for graduation; use appropriate catalog...  ...management, and communication tools Demonstrated ability...  ...on and off campus). Reasonable accommodation is made... 
    Permanent employment
    Full time
    Work experience placement
    Internship
    Work at office
    Visa sponsorship

    Opt For Healthy Living

    California, MO
    4 days ago
  • $80 per hour

     ...refinery engineering expertise to evaluate AI-generated technical content...  ..., project-based contract role uses your hands-on experience with...  ...and data outputs using the tools and workflows from your professional...  ...building price forecast models or developing PPA documentation... 
    Hourly pay
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    3 days ago
  • $85 per hour

     ...planning, and energy data systems to evaluate AI-generated content and...  ...training data. You will use your operational knowledge to...  ...judge accuracy and relevance of model outputs and deliver clear feedback...  ...data systems and common industry tools to validate AI outputs.... 
    Hourly pay
    Contract work
    Part time
    Work at office
    Remote work
    Flexible hours

    SaidGig

    United States
    6 days ago
  • $17 per hour

     ...AI Evaluation Specialists contribute to the advancement of Large Language Models (LLMs) by testing and providing feedback in collaboration with leading AI labs. This role...  ...labs. A belief that your knowledge and reasoning skills can challenge today’s most advanced AI... 
    Temporary work
    Part time
    Remote work

    SaidGig

    United States
    9 days ago
  •  ...Role Overview Psychiatrists evaluate AI-generated psychiatric content, use clinical judgment to assess model responses, and deliver clear, structured feedback that improves...  ..., treatment recommendations, diagnostic reasoning, and patient-facing language. Develop prompts... 
    Hourly pay
    Temporary work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    5 days ago
  •  ...build cutting-edge foundation AI models and end-to-end products that...  ...Paris. Join us!Why this role?Evaluation is critical to making progress...  ...superhuman in many real-world use cases, we must continue to develop...  ....Build scalable and reusable tools for digging into model... 
    Full time
    Part time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    7 hours ago
  • $15 - $20 per hour

     ...Responsibilities Conduct fact-checking using trusted public sources and external tools. Generate high-quality human evaluation data by identifying...  ...inaccuracies. Assess reasoning quality, clarity, tone, and...  ...of responses. Ensure model responses align with... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    8 days ago
  • $15 - $20 per hour

     ...Description ~Conduct fact-checking using trusted public sources and external tools. ~Generate high-quality human evaluation data by identifying...  ...factual inaccuracies. ~Assess reasoning quality, clarity, tone, and...  ...of responses. ~Ensure model responses align with... 
    Part time
    Summer work

    Mercor

    Remote
    19 days ago
  •  ...Policy Preference Reward Model Evaluator is a remote review...  ...citation accuracy, statutory reasoning, and policy adherence;...  ...on precision. AuraOne uses qualified legal and...  ...or compliance tooling. Bilingual experience...  ...Remote · Independent specialist contractor. Employment... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  • $75 - $150 per hour

     ...opportunity to influence the future of legal reasoning. This role involves evaluating AI-generated legal research and...  ...legal research. ~ Regular use of Westlaw as part of your current or...  ...regulations, and legal citations using tools such as KeyCite and other Westlaw research... 
    Hourly pay
    Remote work

    SaidGig

    United States
    6 days ago
  • $20 per hour

     ...You will develop complex prompts to test AI models, write high-quality responses to demonstrate excellence, and evaluate different model outputs based on accuracy...  ...communications and are interested in how AI tools may be used in real-world work environments. No previous... 
    Hourly pay
    Full time
    Contract work
    Part time
    For contractors
    Self employment
    Freelance
    Remote work

    DataAnnotation

    Wyoming, OH
    1 day ago
  • $20 - $30 per hour

     ...As an Image Evaluation Generalist, you will play a pivotal role in supporting an image...  ...expertise will be essential in shaping how models learn, reason, and perform by providing high-quality...  .... Comfort working with review tools or annotation platforms tailored for digital... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Indiana
    a month ago
  • $125 per hour

     ...ParaView specialists leverage their expertise in scientific...  ...In this role, you will evaluate AI-generated content and...  ...responses from large language models (LLMs). Contribute to...  ...of 1 year of experience using ParaView or other related technical tools. Ability to... 
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    3 days ago
  • $60 - $90 per hour

     ...end-to-end expert examples the model learns from, and shape the criteria used to evaluate AI outputs. This role is remote...  ...CRM, communication, and billing tools. Produce polished, senior-level...  ...Opportunity employer and will provide reasonable accommodations for qualified... 
    Remote job
    Hourly pay
    Full time
    Part time
    Freelance
    Work at office

    SaidGig

    Remote
    1 day ago
  • $30 - $90 per hour

     ...As a Go Developer, you will play a crucial role in evaluating and training next-generation AI coding tools during their highly confidential alpha stages. This...  ...performance optimization. Test and evaluate alpha AI models in Cursor over multiple 4-day, 5+ hour daily bursts... 
    Remote job
    Hourly pay
    Contract work
    Part time

    SaidGig

    United States
    7 days ago
  • $70 - $90 per hour

     ...and low-level programming experts to apply their knowledge in systems programming and security concepts to enhance AI models'' ability to detect and reason about potential threats. The opportunity begins with a work trial and may extend into a two-month project based on... 
    Hourly pay
    Remote work

    SaidGig

    United Kingdom
    26 days ago
  • $20 - $70 per hour

     ...crucial role in shaping how AI models learn and perform by providing...  ...to annotate, label, and evaluate designs for accuracy and clarity...  ...input based on real-world CAD use cases. Apply industry best...  ...Ability to quickly adapt to new tools and workflows as needed for diverse... 
    Remote job
    Hourly pay
    Contract work

    SaidGig

    United States
    5 days ago
  • $60 per hour

     ...company is looking for experienced quantitative professionals to evaluate AI-generated analyses and solve complex problems. This fully...  .... Joining this team means shaping the future of AI systems used for reasoning about data and analytics. #J-18808-Ljbffr DataAnnotation
    Remote job
    Hourly pay
    Flexible hours

    DataAnnotation

    New York, NY
    5 days ago
  • $80 - $150 per hour

     .... As a Physics Expert (Postdoc / Junior Professor), you will play a critical role in evaluating and enhancing the training of next-generation AI models, ensuring they learn and reason effectively through your domain knowledge. Key Responsibilities Critically evaluate... 
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Remote
    14 days ago
  •  ...a Clinic Assistant to join our Veterans Evaluation Services (VES) team. This is a hybrid position...  ...role. New hires will not be exempt from using company provided equipment. Must...  ...compensation. Accommodations Maximus provides reasonable accommodations to individuals requiring... 
    Full time
    Contract work
    Currently hiring
    Work at office
    Local area
    Remote work
    Home office
    Shift work
    Weekend work

    MAXIMUS

    Sandy Springs, GA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Tool-Use Reasoning Model Evaluation Specialist [Remote]. Be the first to apply!