Public Sector Practice Reviewer — AI Model Evaluation (Remote)
$50 - $55 per hourAuraOne Human Data
- Remote job
Public Sector Practice Reviewer — AI Model Evaluation (Remote) is a remote evaluation track for reviewing public sector practice evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
Why this role matters
AI data reviewers help turn public sector practice evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.
Responsibilities
- Evaluate public sector practice evaluation model outputs against a versioned rubric and assign severity tags for Public Sector Practice Reviewer — AI Model Evaluation (Remote) assignments.
- Compare paired responses and pick the stronger answer with a written rationale.
- Label hallucinations, instruction-following failures, and unsafe content with structured tags.
- Capture ambiguous prompts and route them back to the program team for rubric updates.
- Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
- Document recurring failure modes so the modeling team can target them in the next training run.
Qualifications
- Prior evaluation, annotation, or human-rater experience on public sector practice evaluation or adjacent content for Public Sector Practice Reviewer — AI Model Evaluation (Remote) work.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that names the issue and the rubric clause being applied.
- Strong attention to detail and the ability to flag when a prompt itself is the problem.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Compare two public sector practice evaluation model responses to the same prompt and pick the stronger one with rationale.
- Tag an unsafe response with the correct policy category and severity.
- Audit a 50-row batch for rubric consistency and report drift to the program lead.
- Propose a rubric clarification after spotting a recurring failure mode.
Nice to have
- Background in linguistics, content moderation, or trust & safety review.
- Experience with inter-rater agreement metrics and calibration cycles.
- Domain expertise that lets you spot subject-matter errors automated checks miss.
Skills
- Model output evaluation
- Rubric-based annotation
- Severity tagging
- Inter-rater calibration
- Public Sector Practice evaluation
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
$50–$55 / hr
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$100 per hour
...expertise to improve AI-generated content... ...a customer-facing evaluation and optimization... ...providing high-quality, practical feedback. Prior AI... .... Author, review, and improve... ...refine large language model prompts using structured... ...Work Terms Remote, part-time...Remote workHourly payPart timeFor contractors- ...solving to improve and evaluate large language models. You will design... ...results with code, and review model outputs to identify... ...thinking, practical Python implementation... ...Accelerate frontier AI research by contributing... ...work independently in a remote setting. Technical...Remote workContract workFor contractorsFreelance
$60 - $90 per hour
...technical talent with leading AI research labs. Headquartered... ...Machine Learning Engineer — Model Evaluation & Experimentation Type... ...$90/hour Location: Remote Commitment: 35 hours... ...mercor.com PS: Our team reviews applications daily. Please complete...Remote workFull timeContract workSummer work$60 - $100 per hour
...Finance Expert — AI Model Reviewer (Python & Financial Modeling) $60-100/hr Remote Freelance STEM About the Role What if your deep knowledge of financial modeling... ...logic errors, ambiguity, or clarity gaps Evaluate whether AI outputs reflect correct application...Remote workHourly payOngoing contractContract workFreelanceFlexible hours$70 - $110 per hour
...Shape how advanced AI systems reason about... ...research team to evaluate medical knowledge tasks... ...measure meaningful model improvement. Key... ...Review clinical knowledge... ...real-world medical practice. Design challenging... .... This is not a remote role. Candidates relocating...SuggestedHourly payFull timeLive inRelocationRelocation package- ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote workFlexible hours
- ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities...Remote workHourly payFlexible hours
$60 per hour
...Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with... ...and offers a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing...Remote jobHourly payWork from homeFlexible hours$60 per hour
...Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing experimental designs....Remote jobHourly payWork from homeFlexible hours$50 - $70 per hour
...Help improve frontier AI systems by evaluating the everyday professional materials they produce. You... ...the work. Key Responsibilities Review documents, slides, spreadsheets, and related... ...will be provided. Work Terms Remote, hourly engagement. Applicants must...Remote workHourly pay- ...Obsidian is hiring expert Evaluators for a remote, hourly role focused on Compliance and regulatory response with financial-services AI. You'll assess AI-generated work products for accuracy and domain quality, leveraging your expertise. The ideal candidate has over 5...Remote workHourly payWork at office
$70 - $90 per hour
...Role Overview Help evaluate Neuron Kernel Interface development tasks... ...and evaluation of advanced AI models. You will assess kernel... ...hardware appropriateness. Review CUDA-to-NKI migrations for fidelity... ...instances. Work Terms Remote role, open to candidates located...Remote workHourly pay$70 - $90 per hour
...Review GPU and accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical correctness, completeness... .../XLA custom calls. Work Terms Remote role open to candidates located in...Remote workHourly pay$70 - $80 per hour
...expertise to help improve next-generation AI systems through high-quality evaluations, safety-report analysis, and structured feedback. This remote contractor opportunity is designed for... ...with experience authoring and reviewing complex pharmacovigilance reports. Prior...Remote workHourly payFor contractors$15 per hour
...Evaluate AI-generated music and lyrics in Malayalam and English, helping assess outputs across a broad range of genres against detailed... ...contemporary Malayalam genres, sub-genres, and artists. Work Terms Remote, hourly engagement with an immediate start. Flexible...Remote workHourly payImmediate startFlexible hours- ...research-grade expertise to help evaluate and improve AI reasoning across technical and humanities disciplines. This remote contractor role supports AI-model training through rigorous analysis... ...analysis, technical writing, peer review, and research methodology skills....Remote workHourly payFor contractors
$11 - $19 per hour
...Evaluate AI-generated music and lyrics across a broad range of genres, applying your Bengali music expertise to help assess quality, originality... ...Bengali genres, sub-genres, and artists. Work Terms Remote, hourly engagement with an immediate start. Flexible...Remote workHourly payImmediate startFlexible hours$28 - $60 per hour
...Evaluate AI-generated music and lyrics across a wide range of genres, applying your knowledge of the Dutch music scene and strong editorial... ...Dutch genres, sub-genres, and artists. Work Terms Remote, independent-contractor engagement. Immediate start, with work...Remote workHourly payFor contractorsImmediate startFlexible hours$100 per hour
...the performance of large language models on finance tasks. You will work with AI researchers to identify model weaknesses... ...experience is required. This is a remote, US-based contract opportunity... .... Key Responsibilities Evaluate LLM performance in finance areas where...Remote workHourly payContract workFor contractorsFreelance10 hours per weekFlexible hours$20 - $60 per hour
...Role Overview Help train and evaluate next-generation AI systems by creating rigorous, real... ...that test how advanced models learn, reason, and perform. This remote contract opportunity is open to... ...Revise content in response to reviewer feedback while following project...Remote workHourly payContract workFor contractors$100 - $120 per hour
...Immigration Law Reviewer — AI Output Evaluation (Remote) is a remote review track for evaluating AI outputs in... ...document the corrected analysis so the modeling team can train on it. Why this... ...qualification in legal review or an adjacent practice for Immigration Law Reviewer — AI...Remote jobFor contractors10 hours per week$70 - $80 per hour
...expertise to author and review complex safety reports and... ...help train next-generation AI systems for... ...contractor role is fully remote and focuses on high-quality evaluation of DSURs, PSURs/PBRERs, aggregate... ...with pharmacovigilance best practices. Verify template and structural...Remote workHourly payFor contractors- Obsidian is seeking expert Evaluators in Biology and environmental science to review AI-generated outputs for accuracy and domain quality. Work is remote and paid hourly, leveraging deep subject-matter knowledge to assess and grade documents, spreadsheets, and slide decks...Remote jobHourly pay
$110 per hour
...clinical expertise to help shape and evaluate advanced AI systems through remote, hourly contract opportunities... ...Responsibilities Train and evaluate AI models used in medicine. Create tasks... ...your work location. Credential review and an AI interview are required...Remote workHourly payContract work$60 - $80 per hour
...connects mathematicians with AI labs and companies to shape and evaluate cutting-edge AI in... ...contribute domain expertise to model training and evaluation,... ...Work independently on remote projects, managing your time... ...steps and pass a review process. Apply once by...Remote workHourly payContract workImmediate start$100 - $150 per hour
...considered for future projects evaluating how well AI systems perform real-world... ...-produced analyses and models, document decisions in writing... ...on evaluations with senior reviewers when projects arise. There... ...Work Terms Location: Remote Engagement type: hourly...Remote workHourly payImmediate start$110 per hour
...expertise to help train, evaluate, and shape medical AI systems by joining a... ...and evaluate AI models in medical and... ...real-world clinical practice. Provide domain-specific... ...independently in a remote environment. Work... .... Credential review and verification are...Remote workHourly payContract work$60 - $90 per hour
...rigorous, real-world materials science and engineering tasks that evaluate an AI model’s expert reasoning. You will create prompts, supporting data... ...and why an incorrect answer fails. Work Terms ~ Remote, hourly engagement. Compensation ~$60 to $90 per hour...Remote workHourly payWork at office- General Motors, through Embodied AI, seeks a Senior Engineer to measure and visualize AV model performance. You will design and implement evaluation workflows, collaborate across Data, Infra... ...with real-world behavior. This remote role offers opportunities to influence...Remote job
$100 per hour
...technology firm is seeking finance experts to enhance AI models. Responsibilities include evaluating performance in capital markets and creating assessment... ...banking and possess strong financial knowledge. This remote role offers competitive hourly pay around $100, with flexibility...Remote jobHourly pay10 hours per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Public Sector Practice Reviewer — AI Model Evaluation (Remote). Be the first to apply!
- senior network engineer remote Remote
- business intelligence analyst remote Remote
- document specialist remote Remote
- remote legal intern Remote
- remote accounts receivable Remote
- revit remote Remote
- remote construction estimator Remote
- remote customer service chat Remote
- junior designer remote Remote
- remote html Remote




