Image-Text Grounding Model Evaluator [Remote]
AuraOne Human Data
- Remote job
Image-Text Grounding Model Evaluator is a remote evaluation track for reviewing image text grounding model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
Why this role matters
AI data reviewers help turn image text grounding model evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.
Responsibilities
- Evaluate image text grounding model evaluation model outputs against a versioned rubric and assign severity tags for Image-Text Grounding Model Evaluator assignments.
- Compare paired responses and pick the stronger answer with a written rationale.
- Label hallucinations, instruction-following failures, and unsafe content with structured tags.
- Capture ambiguous prompts and route them back to the program team for rubric updates.
- Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
- Document recurring failure modes so the modeling team can target them in the next training run.
Qualifications
- Prior evaluation, annotation, or human-rater experience on image text grounding model evaluation or adjacent content for Image-Text Grounding Model Evaluator work.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that names the issue and the rubric clause being applied.
- Strong attention to detail and the ability to flag when a prompt itself is the problem.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Compare two image text grounding model evaluation model responses to the same prompt and pick the stronger one with rationale.
- Tag an unsafe response with the correct policy category and severity.
- Audit a 50-row batch for rubric consistency and report drift to the program lead.
- Propose a rubric clarification after spotting a recurring failure mode.
Nice to have
- Background in linguistics, content moderation, or trust & safety review.
- Experience with inter-rater agreement metrics and calibration cycles.
- Domain expertise that lets you spot subject-matter errors automated checks miss.
Skills
- Model output evaluation
- Rubric-based annotation
- Severity tagging
- Inter-rater calibration
- Image Text Grounding Model evaluation
- Multimodal evaluation
- Cross-modal reasoning
- Grounding review
- Image
- Text
- Grounding
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$20 per hour
...Datatricks AI is actively hiring remote evaluators for **Project Edgehog**, a... ...role, you will analyze multimodal model generations, perform side-by-side image evaluations, and rate system accuracy... ...* Evaluate descriptive text prompts alongside model-generated...SuggestedHourly payContract workRemote workFlexible hours- ...Job Description Job Description Breast Imager, Northern California, University Town Flexible Work Arrangement Hybrid Model, 1099 vs W2, Part Time or Full Time Join a physician-owned group committed to early cancer detection and precision diagnostics. This...SuggestedFull timePart timeWork at officeRemote workWork from homeRelocation packageFlexible hours2 days per week3 days per week
- ...Job Description Job Description Breast Imager, Northern California, University Town Flexible Work Arrangement Hybrid Model, 1099 vs W2, Part Time or Full Time Join a physician-owned group committed to early cancer detection and precision diagnostics. This...SuggestedPermanent employmentFull timeTemporary workPart timeWork at officeRemote workWork from homeRelocation packageFlexible hours2 days per week3 days per week
$20 per hour
...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths, areas for improvement,... ...quality, clarity, tone, and completeness of responses. Ensure model responses align with expected conversational behavior and system...SuggestedRemote jobContract workPart timeSummer work- ...Scientific Figure Understanding Model Evaluator is a remote review track for evaluating AI outputs across scientific figure understanding... ...review Multimodal evaluation Cross-modal reasoning Grounding review Scientific Figure Understanding Work model...SuggestedRemote jobHourly payFor contractors10 hours per week
$114.6k - $252.1k
...Principal AI/ML Engineer (Large Language Model)Job Category: ScienceTime Type: Full... ...Analyze large multi-domain datasets such as images, text and/or graph data, to identify... ...success, and find opportunities to break new ground — in your career and in our legacy.Pay Range...Contract workWork experience placementLocal areaRemote workFlexible hours$130k - $260k
...continuously improve through rigorous evaluation and model post-training. This role exists to... ...operations, data, and technology—bringing grounded technical insight to strategic... ...multimodal inputs and outputs across text, images, and documents—context limits, structured...Full timeContract workTemporary workPart time$65 per hour
Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience...Hourly paySelf employmentWork from homeFlexible hours- ...Medical Imaging Safety Evaluator is a remote clinical-review track for evaluating AI outputs that touch clinical review. Reviewers grade differential... ...; and document the corrected clinical reasoning so the modeling team can close the gap. Why this role matters Clinical...Remote jobHourly payFor contractors10 hours per week
- ...Quant Finance Reasoning Model Evaluator is a remote review track for evaluating AI outputs across finance and risk workflows. Reviewers grade calculations, narrative reasoning, and policy adherence; flag compliance and reconciliation issues; and document the correct treatment...Remote jobHourly payFor contractorsWork experience placement10 hours per week
- ...Frontier Model Misuse Risk Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety team...Remote jobHourly payFor contractors10 hours per week
- ...Policy Preference Reward Model Evaluator is a remote review track for evaluating AI outputs in policy review workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can...Remote jobHourly payFor contractors10 hours per week
- ...Apply advanced chemistry expertise to improve large language models by creating, solving, and clearly explaining complex chemistry... ...comprehension, and multimodal scientific communication using text, images, chemical equations, diagrams, and visual representations....Contract workFor contractorsRemote work
- ...language processing, reinforcement learning, and large language models. We offer generous relocation benefits for eligible... ....ccp, vllm etc) Analyze large multi-domain datasets such as images, text and/or graph data, to identify statistically relevant features...Full timeTemporary workWork at officeLocal areaRemote workVisa sponsorshipRelocation packageFlexible hours
- ...Seeking a full-time Remote AI Research Evaluator with a PhD in Quantitative Finance to assess and enhance AI models' capabilities in financial reasoning and quantitative... ...relevance of quantitative and finance-related text produced by AI models Crafting and answering questions...Full timeContract workRemote workFlexible hours
$30 per hour
...Designers to act as Domain Experts for a high-level AI evaluation project. AI models are evolving beyond simple image generation and now tackle complex design tasks... ...design dimensions, including readability, text density, information hierarchy, and overall professionalism...Remote jobWork from homeFlexible hours- ...demanding AI workloads. We provide high-performance GPU compute and Model API services, enabling AI companies, research labs, and... ...APIs to drive revenue growth and build our sales function from the ground up. You will own two core revenue lines: GPU compute sales (on-demand...Remote jobFull timeFlexible hours
- ...fundamentals – Fine-tuning (SFT/preference), prompting, model-based eval, and the failure modes of each. At... ...the field. Multimodal dataset experience – Building image/video/audio/text datasets to train or evaluate models. Strong Python + ML tooling – PyTorch (or...
$22 per hour
...in their respective markets. Job Overview As a Content Quality Evaluator, English, you will assess the quality of English-language content... ...knowledge of online communication. Experience with long-form text evaluation Bachelor's degree or higher. Background in English,...Work at office$12.6 per hour
...through natural Thai voice interactions, then documenting and evaluating the results. This remote, hourly assignment is designed... .... Direct the AI to complete tasks such as generating images, formatting text, applying spreadsheet formulas, and summarizing documents....Hourly payImmediate startRemote work$147.5k - $211k
...Corporate Vice President, Data Scientist - Model Validation and AI Governance will play a... ...challenging model methodologies, evaluation approaches, controls, and monitoring strategies... ...and AI-enabled organization, we remain grounded in the values that drive lasting impact....Local area3 days per week$160k - $200k
...for Remote Work: ORA_ON_SITE Description SAIC is seeking an Model Based Systems Engineer in Chantilly, VA to support SAIC's large SETA program, supporting the NRO's Ground Enterprise Directorate (GED), Advanced Ground Office. This role puts...Full timeFor contractorsWork at officeRemote workShift work- ...Description Job Description As the Manager of Model Validation & Verification (VnV) for... ...and data science team responsible for evaluating, benchmarking, and validating the machine... ...About Zoox Zoox is developing the first ground-up, fully autonomous vehicle fleet and...Temporary workRelocation package
$208.73k - $279.57k
...AI infrastructure company built from the ground up, we own and operate each layer of the... ...The Staff Software Engineer for the AI Model Lifecycle team will play a crucial role in... ...experiment management: versioning, lineage, evaluation, and reproducible fine-tuning at scale....Full timeTemporary work$50 per hour
...Help improve advanced Spanish-language AI by evaluating model outputs, curating linguistic data, and providing the expert feedback that strengthens... ...meet linguistic and cultural standards. Annotate Spanish text for grammatical, syntactic, and semantic features. Identify...Hourly payRemote work$50 per hour
...your German linguistic and literary expertise to help evaluate and improve advanced language models. You will assess model-generated German content, curate... ...Evaluate the quality and accuracy of German text generated by language models across translation, summarization...Hourly payRemote work$45 - $95 per hour
...systems through accurate language content, evaluation, and contextual guidance. This remote... ...knowledge to projects that support how AI models learn, reason, and perform. Key Responsibilities... ..., review, and proofread Sanskrit texts for accuracy and clarity. Create high-...Hourly payContract workRemote work$50 per hour
...Help evaluate and improve advanced AI language systems by applying expert English-language judgment to model outputs and training data. This fully remote role focuses on assessing how... ...stylistic, and cultural standards. Annotate text for grammatical, syntactic, semantic,...Hourly payRemote workFlexible hours- ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets,... .... Create expert-level prompts and reference answers grounded in real deal experience. Flag errors, weak reasoning,...For contractorsRemote work
- Work from Home | Internet Analyst | Social Media Evaluator At Appen, we work with 8 out of the top 10 global technology companies in the... ...Morphology, Phonology, Translation, Transcription, Proof-reading, or Text and Voice Data Collection and more! Your contribution will...Part timeWork from homeWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Image-Text Grounding Model Evaluator [Remote]. Be the first to apply!





