Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Model Policy Trainer, Image Evaluation - Remote US

$93.6k - $114.4k
Full-time

Handshake

About Handshake

Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions.

In 2025, we started Handshake AI and built the fastest-growing AI data business in history. We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We've grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals every month.

Why join Handshake now:

  • Shape how every career evolves in the AI economy, at global scale, with impact your friends, family and peers can see and feel

  • Partner hand-in-hand with world-class AI labs, Fortune 500 partners and the world's top educational institutions

  • Work together with engineers, scientists, operators, and more from Palantir, Meta, Scale AI, and former YC founders

  • Build a massive, fast-growing business with billions in revenue

About Handshake AI

Human data is the core infrastructure to AI advancement. Frontier AI labs currently improve model capabilities with various data-intensive post-training techniques. We believe that data spend for AI training will increase by 3-5x in the next few years and continue for much longer as models take on new domains. Handshake AI supports all of the frontier AI labs, working on their most complex data at the largest scale.

About the Role

As an AI Image Evaluator, you will help image generation models learn two things at once: what a good image is, and what an acceptable image is.

You will look at prompts and the images a model produced from them, then answer questions like: Did the image do what the prompt asked? Is it well made, or does it have the hands, lighting, text, and anatomy problems that give generated images away? Which of two images is better, and why? Does the image violate the customer's content policy, and if so, which category and how severely? Does it depict a real person, a protected brand, or a minor in a way the policy does not allow?

The interesting cases are the close ones. Two images that look nearly identical until you notice one has a logo in the background. A stylized nude that is fine as figure study and not fine with one change of pose. A prompt that asked for "a realistic photo of a senator" and a model that complied. A beautiful image that ignored half the prompt, next to an ugly one that nailed it.

We are looking for people who already see images critically, whether that came from photography, illustration, design, years inside Midjourney and Stable Diffusion, or moderating visual content at scale. You do not need all of these. You need one deep, and the judgment to learn the rest.

This is not rote annotation. Rubrics cannot anticipate every image, and good evaluators do not apply them mechanically. You will balance the rubric's text and intent with customer expectations, precedent, and team calibration, and you will explain your reasoning clearly enough that it can train a model.

What You Will Do

  • Evaluate generated images against their prompts for adherence, composition, realism, style consistency, and technical defects

  • Compare images side by side and select the stronger one with a clear, evidence-based rationale

  • Classify images against customer content policies covering sexual content, violence, hate symbols, real-person likeness, intellectual property, and depictions of minors

  • Select the most defensible classification when an image is genuinely ambiguous, and write concise rationales that cite rubric language and specific visual details

  • Distinguish "I do not like this" from "this fails the prompt" from "this violates policy," and keep those judgments separate

  • Write and refine prompts that probe where a model's quality or safety behavior breaks down

  • Identify rubric gaps, contradictions, and emerging edge cases, and raise them with project leads and policy teams

  • Participate actively in calibration discussions; challenge interpretations respectfully and update your judgment when stronger reasoning emerges

  • Apply customer policy consistently without substituting personal taste or personal beliefs for the standard

  • Maintain accuracy and consistency across hundreds of visually similar evaluations

You May Be a Fit If

  • You have a trained eye from photography, illustration, concept art, art direction, retouching, photo editing, VFX, or visual design, and you can say precisely why one image is better than another

  • You use generative image tools heavily (Midjourney, Stable Diffusion, ComfyUI, Flux, DALL-E, Ideogram) and know their failure modes, their prompt quirks, and how their safety filters get bypassed

  • You have moderated or reviewed visual content at scale and have applied a policy taxonomy to borderline images under time pressure

  • You notice small details: an extra finger, a mismatched shadow, a brand mark, a face that is a little too familiar

  • You can hold a rubric steady across a long session of near-identical images

  • You can hold a strong opinion without becoming attached to being right

  • You explain judgment calls clearly enough that another person can audit your reasoning

  • You can separate your personal taste from the standard a customer has asked you to apply

  • You communicate clearly and precisely in writing

  • You treat sensitive imagery and difficult subject matter with maturity and sound judgment

Strong candidates may come from photography, illustration, graphic or UX design, art direction, photo editing, animation or VFX, game art, trust and safety, content moderation, brand or IP enforcement, ad review, or art education. We care more about how you see and how you reason than where you learned to do it. A degree and a technical background are not required.

Nice to Have

  • A public portfolio, publication credits, or a body of generative work (Civitai, Discord communities, LoRA or model training, published prompt work)

  • Experience judging images comparatively: portfolio review, photo competition judging, creative A/B testing, art school critique

  • Formal training in anatomy, color, lighting, or composition

  • Content moderation or trust and safety experience on an image-heavy platform

  • Working knowledge of copyright, trademark, and right-of-publicity basics

  • Prior work in AI evaluation, RLHF, image labeling, or data annotation

  • Familiarity with calibration sessions, inter-rater agreement, or adjudication workflows

Prior AI evaluation experience is helpful, but it is not required.

Sensitive-Content Notice

This role involves regular and deliberate engagement with sensitive imagery. Depending on the project, evaluations may include sexual content and nudity, graphic violence and gore, hate symbols, self-harm, and depictions of real people and of minors in contexts that must be assessed against policy. Some of this material is disturbing by design, because the purpose of the work is to teach models not to produce it.

The work is conducted within structured evaluation frameworks and professional guidelines, with exposure limits, content rotation, mandatory reporting protocols for illegal material, and access to mental health support. Candidates must be able to engage with this material carefully, responsibly, and sustainably while maintaining sound judgment and consistent work quality.

Role Details

  • Location: Remote, US

  • Compensation: $45-55/hr

  • Employment classification: W-2

  • Schedule: 8AM - 5PM PT

  • Weekly commitment: M-F

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI Model Policy Trainer, Image Evaluation - Remote US in Remote vacancy
  •  ...Map and Geospatial Image Model Evaluator is a remote evaluation track for reviewing map...  ...Why this role matters AI data reviewers help turn map...  ...response with the correct policy category and severity. Audit...  ...Work model Remote — US-eligible. Remote · Independent... 
    Remote job
    Policy
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  • $93.6k - $114.4k

     ..., we started Handshake AI and built the fastest-growing...  ...researchers to create evaluations, publish benchmarks,...  ...labs currently improve model capabilities with...  ...About the Role As an AI Policy Specialist on the Violence...  ...Details Location: Remote, US Compensation: $45-55... 
    Remote work
    Policy
    Full time
    Shift work

    Handshake

    Remote
    1 day ago
  • $60 - $90 per hour

     ...Machine Learning Engineer — Model Evaluation & Experimentation is a remote review track for evaluating AI outputs across machine learning...  ...Reviewers grade workflow correctness, policy adherence, and stakeholder fit...  ...Work model Remote — US-eligible. Remote · Independent... 
    Remote job
    Policy
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    a month ago
  • $70 - $110 per hour

     ...Financial and Investment Analyst — AI Model Output Evaluation (Remote) is a remote review track for evaluating...  ..., narrative reasoning, and policy adherence; flag compliance and reconciliation...  ...Investment analysis Work model Remote — US-eligible. Remote · Independent... 
    Remote job
    Policy
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  • $86.8k - $198k

    Model and Simulation Software Engineer...  ...by integrating AI‑enabled models that...  ...M&S systems by evaluating new frameworks,...  ...environments. Join us. The world can't...  ...), and various Image Generators (IG)...  ...AI Usage Policy AI is a part...  ...during meetings. Remote : If this position... 
    Remote work
    Policy
    Full time
    Contract work
    Part time
    Work at office
    Local area

    Booz Allen Hamilton

    Suffolk, VA
    1 day ago
  • $60 - $70 per hour

     ...alignment, and overall quality of frontier AI model outputs on complex, policy sensitive, and ambiguous "grey area" topics. Work through structured evaluations to identify unsafe behavior,...  ...content. Work Terms Location: Remote. Employment type: hourly. Compensation... 
    Remote work
    Policy
    Hourly pay

    SaidGig

    United States
    2 days ago
  • $145k - $200k

     ...expertise in enabling ML models in production. We deploy AI models to run in...  ...customers rely on us for frontier AI capabilities...  ...ability to quickly evaluate and integrate new...  ...that allow for “Remote” work on an exceptional...  ...by Palantir, please see our Privacy Policy.
    Remote work
    Policy
    Full time
    Work experience placement
    Work at office
    Work from home
    Relocation package

    Palantir Technologies

    Palo Alto, CA
    1 day ago
  • $46 per hour

     ...partner is looking for a Legal Domain Expert (SME) – AI Model Evaluation based in the United States. This is a remote, flexible opportunity for an experienced legal...  ...environment. ~ Priority expertise includes US Corporate Law , US Securities Law , US Employment... 
    Remote work
    Full time
    Contract work
    Flexible hours

    jobgether

    United States
    7 days ago
  • $300k - $320k

     ...a Technical Program Manager to lead our AI model evaluation initiatives across multiple workstreams....  ...Trust & Safety, Frontier Redteaming, and Policy teams, you will drive high-priority...  ...roles may require more time in our offices. US visa sponsorship: We do sponsor visas! However... 
    Policy
    Work at office
    Home office
    Visa sponsorship
    Relocation package

    Anthropic

    Seattle, WA
    4 days ago
  •  ...Location US-IA-Cedar Rapids...  ...infrastructure behind AI, cloud computing, and...  ...client worksite safety policies and procedures....  ...Performs "Train the Trainer" activities to...  ...Optics BICSI ATFs and remote E2 Optics sites....  ...requests. ~ Ability to evaluate training needs,... 
    Remote work
    Policy
    Full time
    Traineeship
    Work at office
    Local area

    E2 Optics

    Cedar Rapids, IA
    3 days ago
  • £45k - £55k per year

     ...focused on delivering AI-powered solutions that...  ...prostate, and thyroid imaging AI, deployed across enterprise...  ...and post-training evaluations to ensure learning objectives...  ...Follows all DeepHealth policies and procedures. •...  ...Hybrid/Remote Accommodations... 
    Remote work
    Policy
    Local area

    DeepHealth

    Boston, MA
    1 day ago
  • $100 per hour

     ...the performance of large language models on finance tasks. You will work with AI researchers to identify model...  ...experience is required. This is a remote, US-based contract opportunity with a...  ...systems. Key Responsibilities Evaluate LLM performance in finance areas... 
    Remote work
    Hourly pay
    Contract work
    For contractors
    Freelance
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $60 per hour

     ...and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with...  ...offers a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing... 
    Remote work
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    Dallas, TX
    2 days ago
  • $60 - $90 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract...  ...Compensation: $60–$90/hour Location: Remote Commitment: 35 hours/week... 
    Remote work
    Full time
    Contract work
    Summer work

    Mercor

    Remote
    1 day ago
  •  ...US Litigation Attorney — AI Evaluation & Legal Output Review is a remote review track for evaluating AI outputs in litigation workflows...  ..., statutory reasoning, and policy adherence; flag risk; and document...  ...the corrected analysis so the modeling team can train on it. Why... 
    Remote job
    Policy
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    15 hours ago
  • $238k - $302k

     ...across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in...  ...this full-time position across US locations is listed below. Actual...  ...or, if the role can be performed remote, the specific salary range for your... 
    Remote work
    Full time

    Waymo

    San Francisco, CA
    1 day ago
  • $164.78k - $314.96k

     ...members. Be part of what truly makes us special and impactful.We are...  ...military spouses. USAA roles may offer remote or hybrid flexibility for active-...  ...consistent with applicable policy and business needs.The OpportunityAs the AI Model Governance & Monitoring Lead, you... 
    Remote work
    Policy
    Full time
    H1b
    Work at office
    Home office
    Relocation package
    Flexible hours

    USAA - United Services Automobile Association

    San Antonio, TX
    8 hours ago
  • $40 per hour

     ...Massachusetts is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal...  ...biochemistry. This position allows full-time or part-time remote work, with projects that pay hourly starting at $40+.... 
    Remote work
    Hourly pay
    Full time
    Part time

    DataAnnotation

    United States
    4 days ago
  • $174.72k - $295.68k

     ...integrating advanced AI and autonomous...  ...Scientist to drive the modeling and algorithmic...  ...unlabeled fleet data (images, video, LiDAR, CAN...  ...temporal reasoning, policy distillation, imitation...  ...ablation, evaluation, and visualization...  ...position across all US locations. Within the... 
    Policy
    Full time

    XPENG Motors

    Santa Clara, CA
    8 hours ago
  • $60 per hour

    Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    1 day ago
  •  ...Refusal Preference Reward Model Evaluator is a remote red-team track for stress-testing AI systems against adversarial prompts....  ...known weakness classes (jailbreak, policy bypass, prompt injection) for...  ...Preference Work model Remote — US-eligible. Remote · Independent... 
    Remote job
    Policy
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    a month ago
  • Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote job
    Flexible hours

    Prolific

    Charlotte, NC
    5 days ago
  • $11 - $19 per hour

     ...Evaluate AI-generated music and lyrics across a broad range of genres, applying your Bengali music expertise to help assess quality, originality...  ...Bengali genres, sub-genres, and artists. Work Terms Remote, hourly engagement with an immediate start. Flexible... 
    Remote work
    Hourly pay
    Immediate start
    Flexible hours

    SaidGig

    United States
    21 days ago
  • $70 - $90 per hour

     ...accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical...  ...kernels, or JAX/XLA custom calls. Work Terms Remote role open to candidates located in the United States.... 
    Remote work
    Hourly pay

    SaidGig

    Remote
    8 days ago
  • $15 per hour

     ...Evaluate AI-generated music and lyrics in Malayalam and English, helping assess outputs across a broad range of genres against detailed...  ...contemporary Malayalam genres, sub-genres, and artists. Work Terms Remote, hourly engagement with an immediate start. Flexible... 
    Remote work
    Hourly pay
    Immediate start
    Flexible hours

    SaidGig

    United States
    10 days ago
  • $15 - $25 per hour

     ...help improve next-generation AI systems through accurate,...  ...analysis and feedback. This remote, part-time contractor role focuses...  ...quality financial insights, models, and evaluations. Prior AI experience is not...  ...and organizational policies. Use advanced Excel functions... 
    Remote work
    Policy
    Hourly pay
    Part time
    For contractors

    SaidGig

    United States
    a month ago
  • $70 - $90 per hour

     ...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness...  ...or Trn2 instances. Work Terms Remote role, open to candidates located in the... 
    Remote work
    Hourly pay

    SaidGig

    Remote
    8 days ago
  • $800 per unit

     ...annuity expertise to create and evaluate high-quality insurance...  ...work that improves advanced AI systems. This remote opportunity focuses on realistic...  ...advisory, underwriting, and policy-service scenarios involving...  ...feedback used to improve model behavior. Participate in... 
    Remote work
    Policy
    Work at office
    Immediate start

    SaidGig

    United Kingdom
    5 days ago
  • $14 - $42 per hour

     ...Evaluate AI generated music and lyrics across a broad range of genres, applying your knowledge of Urdu music and language to detailed quality standards. This remote, hourly opportunity combines critical listening with bilingual lyric evaluation in Urdu and English. Key... 
    Remote work
    Hourly pay
    Immediate start
    Flexible hours

    SaidGig

    United States
    21 days ago
  •  ...Role Overview Apply research-grade expertise to help evaluate and improve AI reasoning across technical and humanities disciplines. This remote contractor role supports AI-model training through rigorous analysis, high-quality feedback, and clearly articulated academic... 
    Remote work
    Hourly pay
    For contractors

    SaidGig

    United States
    7 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Model Policy Trainer, Image Evaluation - Remote US. Be the first to apply!