Model Grading Model Evaluation Specialist [Remote]
AuraOne Human Data
- Remote job
Model Grading Model Evaluation Specialist is a remote evaluation track for reviewing model grading model evaluation evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
Why this role matters
AI data reviewers help turn model grading model evaluation evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.
Responsibilities
- Evaluate model grading model evaluation evaluation model outputs against a versioned rubric and assign severity tags for Model Grading Model Evaluation Specialist assignments.
- Compare paired responses and pick the stronger answer with a written rationale.
- Label hallucinations, instruction-following failures, and unsafe content with structured tags.
- Capture ambiguous prompts and route them back to the program team for rubric updates.
- Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
- Document recurring failure modes so the modeling team can target them in the next training run.
Qualifications
- Prior evaluation, annotation, or human-rater experience on model grading model evaluation evaluation or adjacent content for Model Grading Model Evaluation Specialist work.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that names the issue and the rubric clause being applied.
- Strong attention to detail and the ability to flag when a prompt itself is the problem.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Compare two model grading model evaluation evaluation model responses to the same prompt and pick the stronger one with rationale.
- Tag an unsafe response with the correct policy category and severity.
- Audit a 50-row batch for rubric consistency and report drift to the program lead.
- Propose a rubric clarification after spotting a recurring failure mode.
Nice to have
- Background in linguistics, content moderation, or trust & safety review.
- Experience with inter-rater agreement metrics and calibration cycles.
- Domain expertise that lets you spot subject-matter errors automated checks miss.
Skills
- Model output evaluation
- Rubric-based annotation
- Severity tagging
- Inter-rater calibration
- Model Grading Model Evaluation evaluation
- Frontier evaluation
- Rubric calibration
- Failure analysis
- Model
- Grading
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
- Dorado is seeking a Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems, probe model reasoning at the frontier, and help identify where models fail under rigorous scientific scrutiny....SuggestedRemote job
$100 - $150 per hour
...Role Overview Evaluate how well AI systems perform real-world technical sales work by defining excellence and judging completed work... ...producing sales deliverables yourself, you will create task-specific grading rubrics and assign scores to AI-generated and human submissions...SuggestedHourly payRemote work- A leading AI company is seeking a Biology Specialist to help fine-tune large language models. Ideal candidates will be pursuing or hold a Ph.D. in Biology or a related field and possess strong research skills. The role involves solving complex biological problems and collaborating...SuggestedRemote job
$20 - $60 per hour
...AI systems. You will design realistic Fortune 500 style scenarios and interact iteratively with an advanced language model to create, edit, and evaluate Office Open XML files, with a focus on .pptx deliverables. No prior AI experience is required, your domain knowledge...SuggestedHourly payContract workFor contractorsWork at officeRemote work- ...Causal Reasoning Model Evaluation Specialist is a remote review track for evaluating AI outputs across causal reasoning model evaluation research... ...reasoning, calculations, and research workflows. Reviewers grade derivations and assumptions, reproduce key results, and document...SuggestedRemote jobHourly payFor contractors10 hours per week
$80 - $110 per hour
...AI systems through accurate, practical evaluation and feedback. This part time contract role... ...concise written feedback that improves model performance. Verify documentation and... ...develop and refine evaluation rubrics and grading criteria for tax related outputs....Hourly payContract workPart timeFor contractorsRemote work$75 - $150 per hour
...everyone’s full potential. Treliant is looking for Credit Risk Modelers for remote, project-based opportunities. Responsibilities... ...credit decisioning and related consumer lending models. Rigorously evaluate predictive accuracy of model assumptions against actual...Work experience placementWork at officeRemote workFlexible hours- ...first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world... ...City, Montreal, Seoul, Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling intelligence. As...Full timeWork at officeLocal areaRemote workHome office
- ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote workFlexible hours
$40 per hour
A technology company in Massachusetts is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal candidates will hold advanced qualifications in biology or biochemistry. This position allows...Hourly payFull timePart timeRemote work$40 - $80 per hour
...This remote contractor role focuses on creating, reviewing, and evaluating materials that reflect authentic Department of Defense... ...Develop high-quality example responses for AI training. Create grading rubrics and evaluation guidance based on federal and military...Hourly payFor contractorsRemote work$100 per hour
...-matter expertise to improve the performance of large language models on finance tasks. You will work with AI researchers to identify... ...accelerator focused on advanced AI systems. Key Responsibilities Evaluate LLM performance in finance areas where models struggle,...Hourly payContract workFor contractorsFreelanceRemote work10 hours per weekFlexible hours$70 - $110 per hour
...AI research team improve how advanced AI models reason about real clinical work. In this... ...clinical tasks, model answers, and evaluation standards alongside research and program... ...standards with researchers and adjacent-domain specialists, turning implicit clinical judgment into...Hourly payFull timeFreelanceLive inRelocationRelocation package$70 - $90 per hour
...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness, CUDA-to-NKI migration fidelity, and whether implementations are well suited to...Hourly payRemote work- ...Role Overview Work with a leading AI lab to evaluate outputs from generative music models in German and English. This role focuses on listening, scoring, and annotating AI-generated music and lyrics across genres, using music production and audio engineering vocabulary...Hourly payPart timeImmediate startRemote work10 hours per week
$46 per hour
...of a partner company, who manages all applications and next steps. Our partner is looking for a Legal Domain Expert (SME) – AI Model Evaluation based in the United States. This is a remote, flexible opportunity for an experienced legal professional to help evaluate...Full timeContract workRemote workFlexible hours$15 per hour
...Role Overview Apply your Punjabi music expertise to evaluate AI-generated music and lyrics across a broad range of genres. You will assess outputs against detailed quality standards in both Punjabi and English. Key Responsibilities Compare AI-generated lyrics with...Hourly payImmediate startRemote workFlexible hours$20 - $60 per hour
...Help train next-generation AI systems by creating rigorous, real-world evaluations that test how well advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates, advanced-degree holders, and professionals from any background...Hourly payContract workFor contractorsRemote work$15 per hour
...Evaluate AI-generated music and lyrics in Malayalam and English, helping assess outputs across a broad range of genres against detailed quality standards. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities. Rate lyrics...Hourly payImmediate startRemote workFlexible hours$11 - $19 per hour
...Evaluate AI-generated music and lyrics in Telugu and English, helping assess outputs across a broad range of genres against detailed quality standards. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities. Rate lyrics for...Hourly payImmediate startRemote workFlexible hours- ...Evaluate AI-generated music and lyrics across a broad range of genres, applying your Bengali music expertise to help assess quality, originality, and natural expression. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities...Hourly payImmediate startRemote workFlexible hours
$35 - $62 per hour
...Apply your Japanese music expertise to evaluate AI-generated music and lyrics across a wide range of genres. You will assess outputs against detailed quality standards in both Japanese and English. Key Responsibilities Compare AI-generated lyrics with published songs...Hourly payFor contractorsImmediate startRemote workFlexible hours$50 - $70 per hour
...Role Overview Help improve frontier AI models by evaluating the quality of real-world professional materials and AI-generated work. You will apply careful judgment across documents, presentations, spreadsheets, and other written content, providing feedback that helps...Hourly payRemote work$20 - $60 per hour
...Role Overview Help train and evaluate next-generation AI systems by creating rigorous, real-world assessments that test how advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates and other researchers and writers with...Hourly payContract workFor contractorsRemote work$75 per hour
...Join a project focused on evaluating AI models in the architecture domain, specifically in visual document understanding and instruction-following. This role involves authoring complex, grounded tasks that include a clear ground-truth output and an objective rubric....Hourly payRemote work$70 - $90 per hour
...Review GPU and accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical correctness, completeness, fair performance benchmarking, appropriate scope, and whether kernels...Hourly payRemote work$28 - $60 per hour
...Evaluate AI-generated music and lyrics across a wide range of genres, applying your knowledge of the Dutch music scene and strong editorial judgment to detailed quality standards. Key Responsibilities Assess AI-generated music and rate it against detailed quality...Hourly payFor contractorsImmediate startRemote workFlexible hours- Senior Research Scientist, Model Evaluation Cohere | Posted Mar 2 | Full-time | New York | Negotiable | Unknown Why this role? Evaluation is critical to making progress in scaling intelligence. As models continue to become superhuman in many real-world use cases, we must...Full timeWork at officeRemote workFlexible hours
- Mercor is seeking a Generalist who can operate in English and Punjabi. This contract, remote position focuses on evaluating AI outputs and supporting model evaluation tasks. You will conduct fact-checking, assess reasoning, clarity, tone and completeness, and provide actionable...Remote jobContract work
- Dorado is seeking an experienced Investment Banking SME to support the development, evaluation, and improvement of advanced AI models in finance. You will assess AI-generated analyses for accuracy, reasoned judgments, and data integrity, guiding model refinements and prompts...Remote job
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Model Grading Model Evaluation Specialist [Remote]. Be the first to apply!
- cost specialist Remote
- strategic sourcing specialist Remote
- absence management specialist Remote
- authorization specialist Remote
- treasury specialist Remote
- print production specialist Remote
- workforce management specialist Remote
- wellness specialist Remote
- helpdesk specialist Remote
- program eligibility specialist Remote



