Image-Text Grounding Model Evaluator [Remote]
AuraOne Human Data
- Remote job
Image-Text Grounding Model Evaluator is a remote evaluation track for reviewing image text grounding model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback the modeling team can use to retrain.
Why this role matters
AI data reviewers help turn image text grounding model evaluation outputs into auditable labels, rationales, and regression cases for AuraOne Human Data.
Responsibilities
- Evaluate image text grounding model evaluation model outputs against a versioned rubric and assign severity tags for Image-Text Grounding Model Evaluator assignments.
- Compare paired responses and pick the stronger answer with a written rationale.
- Label hallucinations, instruction-following failures, and unsafe content with structured tags.
- Capture ambiguous prompts and route them back to the program team for rubric updates.
- Maintain reviewer-quality scores by calibrating against gold-standard examples each week.
- Document recurring failure modes so the modeling team can target them in the next training run.
Qualifications
- Prior evaluation, annotation, or human-rater experience on image text grounding model evaluation or adjacent content for Image-Text Grounding Model Evaluator work.
- Comfort applying multi-page rubrics consistently across long batches.
- Clear written reasoning that names the issue and the rubric clause being applied.
- Strong attention to detail and the ability to flag when a prompt itself is the problem.
- Reliable async availability for at least 10 hours per week.
Example tasks
- Compare two image text grounding model evaluation model responses to the same prompt and pick the stronger one with rationale.
- Tag an unsafe response with the correct policy category and severity.
- Audit a 50-row batch for rubric consistency and report drift to the program lead.
- Propose a rubric clarification after spotting a recurring failure mode.
Nice to have
- Background in linguistics, content moderation, or trust & safety review.
- Experience with inter-rater agreement metrics and calibration cycles.
- Domain expertise that lets you spot subject-matter errors automated checks miss.
Skills
- Model output evaluation
- Rubric-based annotation
- Severity tagging
- Inter-rater calibration
- Image Text Grounding Model evaluation
- Multimodal evaluation
- Cross-modal reasoning
- Grounding review
- Image
- Text
- Grounding
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$123k - $131k
...delivery for the planet. If you're ready to do the greatest work of your life, come join us. About the Role Wing is looking for a Ground Crew Evaluator to join our Aviation Training and Operational Readiness team based in Dallas, TX . As a Ground Evaluator, you are...SuggestedFull timeTraineeshipLocal areaRemote work$20 - $26 per hour
Prolific seeks fluent Kannada speakers to act as evaluators for AI language data. You will assess text and voice segments, rate naturalness, and help identify... ...contextual reviews to ensure authentic Kannada speech in AI models. Flexible hours and work-from-home options are...SuggestedRemote jobHourly payWork from homeFlexible hours$20 per hour
...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths, areas for improvement,... ...quality, clarity, tone, and completeness of responses. Ensure model responses align with expected conversational behavior and system...SuggestedRemote jobContract workPart timeSummer work- ...Scientific Figure Understanding Model Evaluator is a remote review track for evaluating AI outputs across scientific figure understanding... ...review Multimodal evaluation Cross-modal reasoning Grounding review Scientific Figure Understanding Work model...SuggestedRemote jobHourly payFor contractors10 hours per week
- ...Medical Document OCR Model Evaluator is a remote clinical-review track for evaluating AI outputs that touch clinical review. Reviewers grade... ...review Multimodal evaluation Cross-modal reasoning Grounding review Medical Document Work model Remote — US-...SuggestedRemote jobHourly payFor contractors10 hours per week
- **Job Title: AI Trainer || Image Quality Evaluator || English** **Location**: Remote | Work from Home **Employment Type:** Project-based | Contract We are looking for detail-oriented Image Quality Evaluator for a multilingual AI data Annotation and Transcription Specialists...Contract workRemote workWork from homeMonday to FridayDay shift
$185k - $400k
...We are seeking accomplished Research Scientists in Foundation Models with expertise in pre-training and mid-training large-scale multimodal... ...for large-scale multimodal pre-training/mid-training (text, image, audio, and video), and drive innovative approaches for foundational...Remote work$114.6k - $252.1k
...Principal AI/ML Engineer (Large Language Model)Job Category: ScienceTime Type: Full... ...Analyze large multi-domain datasets such as images, text and/or graph data, to identify... ...success, and find opportunities to break new ground — in your career and in our legacy.Pay Range...Contract workWork experience placementLocal areaRemote workFlexible hours$130k - $260k
...continuously improve through rigorous evaluation and model post-training. This role exists to... ...operations, data, and technology—bringing grounded technical insight to strategic... ...multimodal inputs and outputs across text, images, and documents—context limits, structured...Full timeContract workTemporary workPart time$65 per hour
Prolific is seeking registered nurses in Las Vegas, NV, to assist in training AI models. You will review AI responses, rate their accuracy and safety, and write feedback to enhance AI learning. Candidates must be verified registered nurses, have recent clinical experience...Hourly paySelf employmentWork from homeFlexible hours$10 per hour
...both spoken audio and on-screen text. By validating content within... ...will generate high-quality "ground truth" data. Job Type: Freelance... ...hour What You'll Do Media Evaluation: Watch and analyze short-form... ...training automated machine learning models. What we're looking for:...Hourly payPart timeFreelanceImmediate startWork from homeFlexible hours- ...short-form video media and classify languages present in audio and on-screen text from their native locale. This freelance role focuses on producing high-quality ground truth data for AI models, with flexible 4-hour daily commitments and immediate start. Remote work...Remote jobPart timeFreelanceImmediate startFlexible hours
- ...Job Description Job Description Breast Imager, Northern California, University Town Flexible Work Arrangement Hybrid Model, 1099 vs W2, Part Time or Full Time Join a physician-owned group committed to early cancer detection and precision diagnostics. This...Full timePart timeWork at officeRemote workWork from homeRelocation packageFlexible hours2 days per week3 days per week
$23.2 per hour
...set of guidelines to help us improve Roblox systems. As a Human Evaluator you will have the opportunity to provide us direct feedback on... ...classification includes the review and classification of text, image, video, scripts, and audio Track and document insights and trends...Hourly payFull timeContract workWork experience placementWork at officeLocal areaRemote workMonday to Friday- ...Job Description Job Description Breast Imager, Northern California, University Town Flexible Work Arrangement Hybrid Model, 1099 vs W2, Part Time or Full Time Join a physician-owned group committed to early cancer detection and precision diagnostics. This...Permanent employmentFull timeTemporary workPart timeWork at officeRemote workWork from homeRelocation packageFlexible hours2 days per week3 days per week
- Turing is seeking a Software Engineering Evaluator to create cutting‑edge datasets for training, benchmarking, and advancing large language models, collaborating closely with researchers. This includes curating code examples, providing precise solutions, and making corrections...Remote job
- Workada is seeking a Data Labeling Specialist for a remote contract role based in the United States. You'll evaluate AI outputs across text, images, video, and documents against detailed criteria and provide clear justification for your judgments. The ideal candidate has...Remote jobContract work
- ...Refusal Preference Reward Model Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts. Reviewers craft attack scenarios, document the failure mode, and pair each successful jailbreak with the rubric clause it violated so the safety...Remote jobHourly payFor contractors10 hours per week
- ...Quant Finance Reasoning Model Evaluator is a remote review track for evaluating AI outputs across finance and risk workflows. Reviewers grade calculations, narrative reasoning, and policy adherence; flag compliance and reconciliation issues; and document the correct treatment...Remote jobHourly payFor contractorsWork experience placement10 hours per week
- ...Policy Preference Reward Model Evaluator is a remote review track for evaluating AI outputs in policy review workflows. Reviewers grade citation accuracy, statutory reasoning, and policy adherence; flag risk; and document the corrected analysis so the modeling team can...Remote jobHourly payFor contractors10 hours per week
$86.8k - $198k
Model and Simulation Software EngineerThe Opportunity: You will play... ...‑generation M&S systems by evaluating new frameworks, enhancing simulation... ...Generator (VRSG), and various Image Generators (IG) for data... ...weapons systems including aircraft, ground‑based radars, and surface‑to‑...Full timeContract workPart timeWork at officeLocal areaRemote work- ...Apply advanced chemistry expertise to improve large language models by creating, solving, and clearly explaining complex chemistry... ...comprehension, and multimodal scientific communication using text, images, chemical equations, diagrams, and visual representations....Contract workFor contractorsRemote work
- Position: Voice Actor / Narrator AI Text-to-Speech Voice Capture Type: Short-Term Contract Location: Remote Commitment: Full-time (30 40 hours/week, including 4-hour overlap with PST) Engagement Length: 2 months Start Date: Immediate Role Responsibilities Record high-quality...Remote jobFull timeTemporary workImmediate start
- ...language processing, reinforcement learning, and large language models. We offer generous relocation benefits for eligible... ....ccp, vllm etc) Analyze large multi-domain datasets such as images, text and/or graph data, to identify statistically relevant features...Full timeTemporary workWork at officeLocal areaRemote workVisa sponsorshipRelocation packageFlexible hours
$60k - $100k
The ETF and Model Portfolio Business Development Associate will support the ETF and Model... ...skills, including the ability to evaluate market trends, sales data, product positioning... ...family status, disability, or any other ground protected by applicable law.It is our priority...Full timeTemporary workLocal areaHome officeFlexible hours- ...demanding AI workloads. We provide high-performance GPU compute and Model API services, enabling AI companies, research labs, and... ...APIs to drive revenue growth and build our sales function from the ground up. You will own two core revenue lines: GPU compute sales (on-demand...Remote jobFull timeFlexible hours
$30 per hour
...Designers to act as Domain Experts for a high-level AI evaluation project. AI models are evolving beyond simple image generation and now tackle complex design tasks... ...design dimensions, including readability, text density, information hierarchy, and overall professionalism...Remote jobWork from homeFlexible hours- ...fundamentals – Fine-tuning (SFT/preference), prompting, model-based eval, and the failure modes of each. At... ...the field. Multimodal dataset experience – Building image/video/audio/text datasets to train or evaluate models. Strong Python + ML tooling – PyTorch (or...
- ...looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team... ...Deepgram's models — across speech-to-text , text-to-speech , and increasingly... ...evaluation pipelines, define pass/fail criteria grounded in Research benchmarks, and build the...Full time
- ...Sports Model Position We're more than an e-commerce retailer - we're a destination for athletes, adventurers, and sports enthusiasts... ...and accessories for professional product photography. These images will be featured across our company websites to showcase our products...Work at officeShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Image-Text Grounding Model Evaluator [Remote]. Be the first to apply!


