Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Tool-Use Reasoning Model Evaluation Specialist [Remote]

AuraOne Human Data

Remote
  • Remote job

Tool-Use Reasoning Model Evaluation Specialist is a remote evaluation track for reviewing tool use reasoning model evaluation evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.

Why this role matters

AI data reviewers help turn tool use reasoning model evaluation evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.

Responsibilities

  • Evaluate tool use reasoning model evaluation evaluation model outputs against a versioned rubric and assign severity tags for Tool-Use Reasoning Model Evaluation Specialist assignments.
  • Compare paired responses and pick the stronger answer with a written rationale.
  • Label hallucinations, instruction-following failures, and unsafe content with structured tags.
  • Capture ambiguous prompts and route them back to the program team for rubric updates.
  • Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
  • Document recurring failure modes so the modeling team can target them in the next training run.

Qualifications

  • Prior evaluation, annotation, or human-rater experience on tool use reasoning model evaluation evaluation or adjacent content for Tool-Use Reasoning Model Evaluation Specialist work.
  • Comfort applying multi-page rubrics consistently across long batches.
  • Clear written reasoning that names the issue and the rubric clause being applied.
  • Strong attention to detail and the ability to flag when a prompt itself is the problem.
  • Reliable async availability for at least 10 hours per week.

Example tasks

  • Compare two tool use reasoning model evaluation evaluation model responses to the same prompt and pick the stronger one with rationale.
  • Tag an unsafe response with the correct policy category and severity.
  • Audit a 50-row batch for rubric consistency and report drift to the program lead.
  • Propose a rubric clarification after spotting a recurring failure mode.

Nice to have

  • Background in linguistics, content moderation, or trust & safety review.
  • Experience with inter-rater agreement metrics and calibration cycles.
  • Domain expertise that lets you spot subject-matter errors automated checks miss.

Skills

  • Model output evaluation
  • Rubric-based annotation
  • Severity tagging
  • Inter-rater calibration
  • Tool Use Reasoning Model Evaluation evaluation
  • Frontier evaluation
  • Rubric calibration
  • Failure analysis
  • Tool
  • Reasoning

Work model

Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.

Compensation

Hourly rate confirmed after the interview process.

Application process

Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.

Vacancy posted 8 days ago
Similar jobs that could be interesting for youBased on the Tool-Use Reasoning Model Evaluation Specialist [Remote] in Remote vacancy
  • $20 - $60 per hour

     ...iteratively with an advanced language model to create, edit, and evaluate Office Open XML files, with a focus...  ...Provide cross-functional expertise for use cases such as reporting, presenting,...  ...Experience with LLMs or AI-powered tools in iterative conversational formats is... 
    Suggested
    Hourly pay
    Contract work
    For contractors
    Work at office
    Remote work

    SaidGig

    United States
    1 day ago
  • $20 per hour

     ...Responsibilities Conduct fact-checking using trusted public sources and external tools. Generate high-quality human evaluation data by identifying...  ...inaccuracies. Assess reasoning quality, clarity, tone, and...  ...completeness of responses. Ensure model responses align with... 
    Suggested
    Remote job
    Contract work
    Part time
    Summer work

    Mercor

    New York, NY
    4 days ago
  •  ...Apply advanced mathematical reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math...  ...implementation, and the use of formal theorem proving to...  ...by learning to leverage AI tools to be a more effective analyst... 
    Suggested
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    a month ago
  •  ...can actually afford to use. Inference is where that...  ...product built on a model is bounded by what it costs...  ...~ This role builds the evaluation and decision systems...  ...level across task success, tool use, quality, cost,...  ...windows, prompt caching, reasoning controls, and inference... 
    Suggested
    Full time

    Bitdeer Technologies Group

    Austin, TX
    18 days ago
  • $60 - $90 per hour

     ...improve how advanced AI models reason through real-world...  ..., and create rigorous evaluations grounded in industry practice...  ...domain-specific tools with the research team...  ...researchers and adjacent-domain specialists to calibrate...  ...Hands-on professional use of large language models... 
    Suggested
    Hourly pay
    Full time
    Remote work

    SaidGig

    United States
    1 day ago
  • $65 - $105 per hour

     ...Help advance frontier AI models by bringing rigorous...  ...into task design, evaluation, and model improvement...  ...to ensure models can reason credibly about real scientific...  ...within the client''s tools using client-issued accounts...  ...and adjacent-field specialists, translating expert scientific... 
    Hourly pay
    Full time
    Freelance
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    28 days ago
  • $90 - $175 per hour

     ...assurance expertise to evaluate technical AI outputs and...  ...AI systems learn, reason, and perform. This remote...  ...precise reproduction steps using structured tracking...  ...frameworks and test management tools, such as Selenium,...  ..., RLHF, AI response or model evaluation, or rubric-... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    10 days ago
  • $65 - $105 per hour

     ...judgment to help frontier AI models reason more accurately about...  ...engineering work, evaluate model performance, and...  ...-specific skills and tools with the research team...  ...with researchers and specialists in adjacent disciplines...  ...Hands-on professional use of large language... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    28 days ago
  •  ...highly motivated Program Evaluation Specialist to lead program...  ...evaluation plans, logic models, performance indicators...  ...evaluation. Analyze data using statistical software...  ...platforms with tools such as Tableau. Excellent...  ...any time, and for any reason, throughout the course... 
    Full time
    Temporary work
    Part time
    Work at office
    Remote work
    Worldwide
    Flexible hours
    Afternoon shift
    3 days per week
    Early shift

    University of Michigan

    Ann Arbor, MI
    5 days ago
  • $68k - $113k

     ...and implementation of program evaluation and performance measurement...  ...Analyze and visualize data using tools such as R and Power BI to support...  ...evaluation concepts, logic models, performance measurement,...  ...required to provide needed reasonable accommodation. All communication... 
    Temporary work
    Remote work
    Flexible hours

    Guidehouse

    United States
    2 days ago
  • $70 - $110 per hour

     ...advanced AI systems reason about real...  ...research team to evaluate medical knowledge...  ...measure meaningful model improvement. Key...  ...specific skills and tools with the research...  ...and adjacent-domain specialists to calibrate standards...  ...-on experience using large language models... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    28 days ago
  • $20 - $60 per hour

     ...Help train next-generation AI systems by creating rigorous, real-world evaluations that test how well advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates, advanced-degree holders, and professionals from any background... 
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    10 hours ago
  • $50 - $70 per hour

     ...improve frontier AI systems by evaluating the quality of professional...  ...correctness. Provide clear, well reasoned written feedback and ratings....  ...rank AI generated outputs using defined evaluation criteria....  ...using common productivity tools, including documents, slides,... 
    Hourly pay
    Remote work

    SaidGig

    United Kingdom
    3 days ago
  • $5,540 - $5,780 per month

     ...Admissions and Program Evaluations (GAPE) in the College...  ...schedule. The Evaluation Specialist is responsible for...  ...applying for graduation; use appropriate catalog...  ...management, and communication tools. Demonstrated ability...  ...by the university. Reasonable accommodation is made... 
    Permanent employment
    Full time
    Work experience placement
    Internship
    Work at office
    Visa sponsorship

    The California State University

    California, MO
    5 days ago
  • $70 - $80 per hour

     ...AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data...  ...Author and review evaluation tasks using DSURs, PSURs/PBRERs, associated safety data...  ...pharmacovigilance documentation and analysis tools. Qualifications At least 5 years of... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    a month ago
  •  ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets... 
    For contractors
    Remote work

    SaidGig

    United States
    22 days ago
  • $70 - $90 per hour

     ...tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task...  ...will provide clear written feedback using defined evaluation rubrics. Key Responsibilities...  ...and benchmarking performance with tools such as Nsight, NCU, roofline... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    24 days ago
  •  ...As the Manager of Model Validation & Verification (VnV) for Behavior Autonomy...  ...data science team responsible for evaluating, benchmarking, and validating the machine...  ...experiences, and skills. We may use artificial intelligence (AI) tools to support parts of the hiring... 
    Full time
    Temporary work
    Relocation package

    Zoox

    California
    4 days ago
  •  ...specializes in monitoring, evaluation, research, and...  ...Senior Evaluation Specialist Department/...  ..., project logic-model framework, evaluation...  ...practical use by professional evaluators...  ...the framework and tools support cross-...  ...objects up to 20 lbs. Reasonable accommodations may... 
    Temporary work
    Part time
    Local area
    Remote work

    Integrated Business & Technical Consultants

    Vienna, VA
    21 days ago
  •  ...Operations Research Model Evaluator is a remote review track for evaluating...  ...model research review reasoning, calculations, and...  ...up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a...  ...reviewing AI-assisted research tooling and its failure modes.... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    a month ago
  •  ...Policy Preference Reward Model Evaluator is a remote review...  ...citation accuracy, statutory reasoning, and policy adherence;...  ...on precision. AuraOne uses qualified legal and...  ...or compliance tooling. Bilingual experience...  ...Remote · Independent specialist contractor. Employment... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    14 days ago
  • $60 - $90 per hour

     ...Learning Engineer — Model Evaluation & Experimentation is...  ...evaluation experimentation specialist operations workflows....  ...day at work. AuraOne uses experienced operators...  .... Clear written reasoning that names the policy...  ...AI-assisted workflow tooling and its failure modes... 
    Remote job
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    7 days ago
  • $50 per hour

     ...role focuses on improving and evaluating large language models through advanced mathematical reasoning, clear written solutions, and computational...  ...AI systems for enterprise use. Key Responsibilities...  ...Pacific Standard Time (PST). Tools and skills: Work will involve Python... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    18 days ago
  •  ...Program Evaluation Specialist Location: Nashville, Tennessee - 37243 Contract: 12+ months from the...  ...and quality improvement strategies, tools, and data collection plans to track outcomes...  ...analysing data and creating reports using available data. Desired Skills, Knowledge... 
    Contract work
    Work experience placement
    Work at office

    Software Technology Inc

    Nashville, TN
    5 days ago
  • $20 - $60 per hour

     ...Help improve how large language models create, understand, and modify...  ...scenarios, produce and evaluate complex .xlsx, .docx, and .pptx...  ...scenarios that require advanced use of Microsoft Excel, PowerPoint...  ...writing long-form prompts for AI tools is a strong plus, but not... 
    Hourly pay
    For contractors
    Work at office
    Remote work

    SaidGig

    United States
    11 days ago
  •  ...image-based work that helps develop and evaluate advanced AI models for clinical image interpretation....  ...outputs reflect real-world clinical reasoning and dermatologic standards of care....  ...and describe dermatologic conditions using precise medical terminology. English... 
    Hourly pay
    Remote work

    SaidGig

    United States
    9 days ago
  •  ...establish the clinical reference standards used to evaluate advanced AI systems on real radiology...  .... This role draws on the diagnostic reasoning that matters most in practice, including...  ...cases that test the limits of strong AI models. Qualifications Board-certified... 
    Hourly pay
    Remote work

    SaidGig

    United States
    4 days ago
  • $228.7k - $343.1k

     ...builds simple, powerful tools that make progress...  ...enormous scale, and one bad model can mean millions in...  ...scale, so you critically evaluate what it produces and own...  ...run in parallel. Reason about ML systems end to...  ...ambiguity. Technologies We Use and Teach Python... 
    Remote job
    Full time
    Local area
    Shift work

    Block

    New York, NY
    10 hours ago
  •  ...Title Associate Specialist, BPS Operations...  ...between broker-dealers using ACATS and other industry tools. Mutual Fund...  ...opportunity employer. We evaluate qualified...  ...for purposes of ADA reasonable accommodation. All...  ...basis. Sourcing Model Recruitment at... 
    Full time

    FIS

    Brown Deer, WI
    1 day ago
  •  ...DevOps and SRE Evaluation Specialist is a remote engineering review track...  ...write the unit test the model should have written, and...  ..., RSpec, or whatever you use. Clear written reasoning — your review note has to...  ...target language's standard tooling, linters, and idiomatic... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    14 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Tool-Use Reasoning Model Evaluation Specialist [Remote]. Be the first to apply!