Tool-Use Reasoning Model Evaluation Specialist [Remote]
AuraOne Human Data
- Remote job
Tool-Use Reasoning Model Evaluation Specialist is a remote evaluation track for reviewing tool use reasoning model evaluation evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
Why this role matters
AI data reviewers help turn tool use reasoning model evaluation evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.
Responsibilities
- Evaluate tool use reasoning model evaluation evaluation model outputs against a versioned rubric and assign severity tags for Tool-Use Reasoning Model Evaluation Specialist assignments.
- Compare paired responses and pick the stronger answer with a written rationale.
- Label hallucinations, instruction-following failures, and unsafe content with structured tags.
- Capture ambiguous prompts and route them back to the program team for rubric updates.
- Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
- Document recurring failure modes so the modeling team can target them in the next training run.
Qualifications
- Prior evaluation, annotation, or human-rater experience on tool use reasoning model evaluation evaluation or adjacent content for Tool-Use Reasoning Model Evaluation Specialist work.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that names the issue and the rubric clause being applied.
- Strong attention to detail and the ability to flag when a prompt itself is the problem.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Compare two tool use reasoning model evaluation evaluation model responses to the same prompt and pick the stronger one with rationale.
- Tag an unsafe response with the correct policy category and severity.
- Audit a 50-row batch for rubric consistency and report drift to the program lead.
- Propose a rubric clarification after spotting a recurring failure mode.
Nice to have
- Background in linguistics, content moderation, or trust & safety review.
- Experience with inter-rater agreement metrics and calibration cycles.
- Domain expertise that lets you spot subject-matter errors automated checks miss.
Skills
- Model output evaluation
- Rubric-based annotation
- Severity tagging
- Inter-rater calibration
- Tool Use Reasoning Model Evaluation evaluation
- Frontier evaluation
- Rubric calibration
- Failure analysis
- Tool
- Reasoning
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$20 - $60 per hour
...iteratively with an advanced language model to create, edit, and evaluate Office Open XML files, with a focus... ...Provide cross-functional expertise for use cases such as reporting, presenting,... ...Experience with LLMs or AI-powered tools in iterative conversational formats is...SuggestedHourly payContract workFor contractorsWork at officeRemote work$20 per hour
...Responsibilities Conduct fact-checking using trusted public sources and external tools. Generate high-quality human evaluation data by identifying... ...inaccuracies. Assess reasoning quality, clarity, tone, and... ...completeness of responses. Ensure model responses align with...SuggestedRemote jobContract workPart timeSummer work- ...Apply advanced mathematical reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math... ...implementation, and the use of formal theorem proving to... ...by learning to leverage AI tools to be a more effective analyst...SuggestedContract workFor contractorsFreelanceRemote work
- ...can actually afford to use. Inference is where that... ...product built on a model is bounded by what it costs... ...~ This role builds the evaluation and decision systems... ...level across task success, tool use, quality, cost,... ...windows, prompt caching, reasoning controls, and inference...SuggestedFull time
$60 - $90 per hour
...improve how advanced AI models reason through real-world... ..., and create rigorous evaluations grounded in industry practice... ...domain-specific tools with the research team... ...researchers and adjacent-domain specialists to calibrate... ...Hands-on professional use of large language models...SuggestedHourly payFull timeRemote work$65 - $105 per hour
...Help advance frontier AI models by bringing rigorous... ...into task design, evaluation, and model improvement... ...to ensure models can reason credibly about real scientific... ...within the client''s tools using client-issued accounts... ...and adjacent-field specialists, translating expert scientific...Hourly payFull timeFreelanceLive inRelocationRelocation package$90 - $175 per hour
...assurance expertise to evaluate technical AI outputs and... ...AI systems learn, reason, and perform. This remote... ...precise reproduction steps using structured tracking... ...frameworks and test management tools, such as Selenium,... ..., RLHF, AI response or model evaluation, or rubric-...Hourly payContract workRemote work$65 - $105 per hour
...judgment to help frontier AI models reason more accurately about... ...engineering work, evaluate model performance, and... ...-specific skills and tools with the research team... ...with researchers and specialists in adjacent disciplines... ...Hands-on professional use of large language...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package- ...highly motivated Program Evaluation Specialist to lead program... ...evaluation plans, logic models, performance indicators... ...evaluation. Analyze data using statistical software... ...platforms with tools such as Tableau. Excellent... ...any time, and for any reason, throughout the course...Full timeTemporary workPart timeWork at officeRemote workWorldwideFlexible hoursAfternoon shift3 days per weekEarly shift
$68k - $113k
...and implementation of program evaluation and performance measurement... ...Analyze and visualize data using tools such as R and Power BI to support... ...evaluation concepts, logic models, performance measurement,... ...required to provide needed reasonable accommodation. All communication...Temporary workRemote workFlexible hours$70 - $110 per hour
...advanced AI systems reason about real... ...research team to evaluate medical knowledge... ...measure meaningful model improvement. Key... ...specific skills and tools with the research... ...and adjacent-domain specialists to calibrate standards... ...-on experience using large language models...Hourly payFull timeLive inRelocationRelocation package$20 - $60 per hour
...Help train next-generation AI systems by creating rigorous, real-world evaluations that test how well advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates, advanced-degree holders, and professionals from any background...Hourly payContract workFor contractorsRemote work$50 - $70 per hour
...improve frontier AI systems by evaluating the quality of professional... ...correctness. Provide clear, well reasoned written feedback and ratings.... ...rank AI generated outputs using defined evaluation criteria.... ...using common productivity tools, including documents, slides,...Hourly payRemote work$5,540 - $5,780 per month
...Admissions and Program Evaluations (GAPE) in the College... ...schedule. The Evaluation Specialist is responsible for... ...applying for graduation; use appropriate catalog... ...management, and communication tools. Demonstrated ability... ...by the university. Reasonable accommodation is made...Permanent employmentFull timeWork experience placementInternshipWork at officeVisa sponsorship$70 - $80 per hour
...AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data... ...Author and review evaluation tasks using DSURs, PSURs/PBRERs, associated safety data... ...pharmacovigilance documentation and analysis tools. Qualifications At least 5 years of...Hourly payContract workRemote work- ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets...For contractorsRemote work
$70 - $90 per hour
...tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task... ...will provide clear written feedback using defined evaluation rubrics. Key Responsibilities... ...and benchmarking performance with tools such as Nsight, NCU, roofline...Hourly payRemote work- ...As the Manager of Model Validation & Verification (VnV) for Behavior Autonomy... ...data science team responsible for evaluating, benchmarking, and validating the machine... ...experiences, and skills. We may use artificial intelligence (AI) tools to support parts of the hiring...Full timeTemporary workRelocation package
- ...specializes in monitoring, evaluation, research, and... ...Senior Evaluation Specialist Department/... ..., project logic-model framework, evaluation... ...practical use by professional evaluators... ...the framework and tools support cross-... ...objects up to 20 lbs. Reasonable accommodations may...Temporary workPart timeLocal areaRemote work
- ...Operations Research Model Evaluator is a remote review track for evaluating... ...model research review reasoning, calculations, and... ...up under scrutiny. AuraOne uses scientific specialists to grade outputs the way a... ...reviewing AI-assisted research tooling and its failure modes....Remote jobHourly payFor contractors10 hours per week
- ...Policy Preference Reward Model Evaluator is a remote review... ...citation accuracy, statutory reasoning, and policy adherence;... ...on precision. AuraOne uses qualified legal and... ...or compliance tooling. Bilingual experience... ...Remote · Independent specialist contractor. Employment...Remote jobHourly payFor contractors10 hours per week
$60 - $90 per hour
...Learning Engineer — Model Evaluation & Experimentation is... ...evaluation experimentation specialist operations workflows.... ...day at work. AuraOne uses experienced operators... .... Clear written reasoning that names the policy... ...AI-assisted workflow tooling and its failure modes...Remote jobFor contractorsWork experience placement10 hours per week$50 per hour
...role focuses on improving and evaluating large language models through advanced mathematical reasoning, clear written solutions, and computational... ...AI systems for enterprise use. Key Responsibilities... ...Pacific Standard Time (PST). Tools and skills: Work will involve Python...Contract workFor contractorsFreelanceRemote work- ...Program Evaluation Specialist Location: Nashville, Tennessee - 37243 Contract: 12+ months from the... ...and quality improvement strategies, tools, and data collection plans to track outcomes... ...analysing data and creating reports using available data. Desired Skills, Knowledge...Contract workWork experience placementWork at office
$20 - $60 per hour
...Help improve how large language models create, understand, and modify... ...scenarios, produce and evaluate complex .xlsx, .docx, and .pptx... ...scenarios that require advanced use of Microsoft Excel, PowerPoint... ...writing long-form prompts for AI tools is a strong plus, but not...Hourly payFor contractorsWork at officeRemote work- ...image-based work that helps develop and evaluate advanced AI models for clinical image interpretation.... ...outputs reflect real-world clinical reasoning and dermatologic standards of care.... ...and describe dermatologic conditions using precise medical terminology. English...Hourly payRemote work
- ...establish the clinical reference standards used to evaluate advanced AI systems on real radiology... .... This role draws on the diagnostic reasoning that matters most in practice, including... ...cases that test the limits of strong AI models. Qualifications Board-certified...Hourly payRemote work
$228.7k - $343.1k
...builds simple, powerful tools that make progress... ...enormous scale, and one bad model can mean millions in... ...scale, so you critically evaluate what it produces and own... ...run in parallel. Reason about ML systems end to... ...ambiguity. Technologies We Use and Teach Python...Remote jobFull timeLocal areaShift work- ...Title Associate Specialist, BPS Operations... ...between broker-dealers using ACATS and other industry tools. Mutual Fund... ...opportunity employer. We evaluate qualified... ...for purposes of ADA reasonable accommodation. All... ...basis. Sourcing Model Recruitment at...Full time
- ...DevOps and SRE Evaluation Specialist is a remote engineering review track... ...write the unit test the model should have written, and... ..., RSpec, or whatever you use. Clear written reasoning — your review note has to... ...target language's standard tooling, linters, and idiomatic...Remote jobHourly payFor contractors10 hours per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Tool-Use Reasoning Model Evaluation Specialist [Remote]. Be the first to apply!
- welding specialist Remote
- transportation specialist Remote
- title specialist Remote
- e learning specialist Remote
- employment specialist Remote
- employment placement specialist Remote
- deployment specialist Remote
- localization specialist Remote
- order entry specialist Remote
- electronic health record specialist Remote




