AI Evaluators for Gemini (US only)
$20 per hourGramian Consulting Group
About Us Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs. Role overview We are looking for detail-oriented professionals based in the United States to support AI model improvement through high-quality data review, annotation, classification, and validation. This is a short-term remote contract opportunity for candidates who can follow detailed guidelines, make consistent decisions, and maintain strong accuracy across repetitive review tasks. In this role, you will work on structured AI data tasks used to train and evaluate the Gemini model. CONTRACT: Short-term freelance/contractor assignment COMMITMENT: 40 hours per week, at least 4 hours per day, including 4 hours overlap with PST HOURLY RATE: $20/h LOCATIONS: Remote, Alabama, Arizona, Colorado, Florida, Idaho, Iowa, Kentucky, Michigan, Minnesota, Mississippi, Missouri, Montana, New York, North Carolina, North Dakota, Oklahoma, Pennsylvania, South Carolina, South Dakota, Texas, Virginia, Wisconsin, Wyoming Responsibilities Review, label, classify, and validate data according to detailed project guidelines Ensure accuracy, consistency, and quality across assigned tasks Apply instructions objectively when evaluating or categorizing content Identify edge cases, discrepancies, and annotation issues Escalate unclear or inconsistent cases when needed Perform quality checks and provide clear feedback where required Meet productivity and quality expectations in a remote work environment Work independently while following project standards and timelines Candidates must be in the US and eligible for 1099 contract Candidates must be students, graduates, or recent graduates Candidates must have at least 20 recent chat sessions with Gemini Candidates must feel comfortable sharing their Gemini history for the purpose of the project #J-18808-Ljbffr Gramian Consulting Group
$60 per hour
A leading project-based AI consultancy is seeking legal consultants to evaluate AI systems and improve their reasoning. This non-permanent role requires a law degree and at least 2 years of experience in US law. Candidates must have strong written English skills and a...SuggestedPermanent employmentPart timeFreelanceFlexible hours$20.3 per hour
Contract type: Freelance Hourly rate: $20.30 Language: English (US) Estimated volume: 8 - 10 hours Start date: The project runs... ...are looking for qualified raters to support a project focused on evaluating personalized search and location recommendations. This role...SuggestedHourly payContract workFreelanceImmediate start$14.5 per hour
Join to apply for the AI Web Search Evaluator role at Welo Data Welo Data works with technology companies to provide datasets that are high-quality... ...and spoken) Strong understanding of pop culture in the US Reliable computer system and internet connection Familiar...SuggestedHourly payPart timeImmediate startRemote workWork from home10 hours per weekFlexible hours- About the Opportunity A leading AI research organization is seeking advanced LLM power... ...personal life tasks. This project focuses on evaluating how well AI systems handle personalized,... ...Looking For Strong candidates will have: US-based only Strong MCP experience and plug...SuggestedTrial period
- ...BAM Ventures is seeking Swedish-speaking remote annotators to evaluate AI-generated content, ensuring that it's coherent and aligns with real-world expectations. Your role will involve reviewing outputs, identifying deviations, and providing structured feedback to enhance...SuggestedRemote work
- ...English to Turkish Software Localization MT/LLM Evaluator to support large-scale IT/software... ...of Turkish content, you will assess MT and AI outputs and provide feedback within established... ...quality frameworks. Remote work across US, Canada, and Latin America time zones,...Remote jobFreelance
- Welocalize is seeking a Maps Visual Design Relevance Evaluator (US English). You will assess digital map design quality, readability, and relevance of location recommendations, ensuring precise geospatial framing and an intuitive user experience. The freelance, remote role...Remote jobHourly payFreelance
- Centraprise is looking for detail-oriented and motivated AI Data Annotators proficient in Canadian French to enhance AI-powered conversational... ...and collaboration within a supportive team culture. Join us to gain valuable experience in the fast-growing AI industry! #J-1...Remote jobFlexible hours
- ...Obsidian is hiring expert Evaluators in real estate, hospitality, and events to review AI-generated work products for accuracy and quality. This is a remote position requiring strong expertise in the relevant domains. Ideal candidates should have over 5 years of professional...Work at officeRemote work
$400 per month
...the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models... ...such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools. Ability to evaluate...- Mercor is hiring experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards. Start date: immediate; Duration: up to...Immediate startFlexible hours
- Cincinnatus LLC is recruiting a senior character animator to help shape evaluation standards for an AI-driven performance transfer model. The role focuses on facial animation, performance capture, and translating complex visuals into precise data captions, with part-time...Part time
- AIUC is seeking experienced AI Safety Practitioners to evaluate safety, quality, and alignment of frontier AI models on policy-sensitive topics. You will assess AI-generated responses and provide structured feedback to improve model behavior. Join a collaborative team of...
- Obsidian is seeking expert Evaluators for a remote, hourly role focused on reviewing AI-generated work products for quality and accuracy. The successful candidates will leverage their subject matter expertise to assess various documents, spreadsheets, and slide decks....Remote jobHourly payWork at office
- Obsidian is seeking experienced evaluators to test AI systems on complex personal workflows across health, travel, planning, home services, and career search. You will create realistic prompts, execute tasks with screen recording, and use your own plugins to complete actions...
- Mercor seeks experienced brand designers to define grading criteria for real-world brand deliverables and to score AI-generated versus human work with detailed written justifications. You will apply evidence-based judgment, ensure reproducibility, and iteratively incorporate...
- MERIT Beauty is seeking Turkish-speaking annotators for a contract role evaluating AI-generated content. You will review content for coherence, consistency, and alignment with real-world expectations in Turkish. The ideal candidates are native or fluent Turkish speakers...Remote jobContract workTemporary workImmediate start
$20 - $30 per hour
A remote-focused technology firm is looking for an individual proficient in evaluating large language model outputs. The role involves assessing AI systems, reviewing workflows, and providing actionable feedback to enhance product quality. Ideal candidates will have strong...Remote jobHourly pay- Productive Playhouse seeks a Portugal Portuguese AI Product Evaluator for an upcoming project testing and evaluating leading AI chatbots. You will interact with AI models to assess capability, safety, and helpfulness, informing model development. Open to freelancers outside...Remote jobFor contractorsFreelanceFlexible hours
$50 - $60 per hour
A leading company in AI research is seeking candidates with a PhD or Master's degree in Chemistry to evaluate and improve large language models (LLMs). The position is part-time, around 30 hours a week, and operates in a remote and flexible environment. Responsibilities...Remote jobHourly payPart timeFlexible hours- Mercor is seeking expert Evaluators in Clinical / biomedical / pharma to review and assess AI-generated work products (documents, spreadsheets, and slide decks) for accuracy, rigor, and domain quality. This is a remote, hourly engagement. You will apply deep subject-matter...Remote jobHourly pay
- Obsidian is hiring expert Evaluators based in New York, United States for a remote role focusing on reviewing AI-generated work products like documents, spreadsheets, and slide decks. The ideal candidate will have over 5 years of experience in BI dashboards and performance...Hourly payWork at officeRemote work
$30 per hour
...Location: Remote Commitment: 10-40 hours/week Role Responsibilities Evaluate outputs from large language models and autonomous agent systems... ...stakeholders. Requirements Strong experience in LLM evaluation, AI output analysis, QA/testing, UX research, or similar analytical...Remote jobHourly payContract work- A talent marketplace is seeking soccer experts to evaluate live soccer games. The role involves assessing AI-generated commentary, scoring performance, and providing feedback. Qualified candidates will have deep expertise in soccer, strong analytical and communication...Contract work
- Mercor is hiring expert Evaluators in Operations / inventory / capacity planning to review AI-generated documents, spreadsheets, and slide decks for accuracy and quality. You will apply deep subject-matter expertise to grade outputs and provide actionable feedback. This...Remote jobHourly payWork at office
$24 per hour
Prolific is seeking an AI Trainer with advanced Tamil fluency to evaluate AI models' understanding of the Tamil language's emotional and cultural nuances. Responsibilities include assessing audio clips, reviewing tones for cultural relevancy, and ensuring quality control...Remote jobFlexible hours- Mercor is seeking experienced musicians to evaluate generative music AI models, in partnership with a leading AI lab. You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Hindi and English. Responsibilities...
- Dorado is seeking expert Evaluators in program management / implementation planning to review AI-generated documents, spreadsheets, and slide decks for accuracy, rigor, and domain quality. You will apply deep subject-matter expertise to grade outputs. This is a remote,...Remote jobHourly payWork at office
- Mercor is seeking expert Evaluators in Clinical/biomedical/pharma to review AI-generated documents, spreadsheets, and slide decks for accuracy and domain quality. Remote, hourly engagement, requiring 5+ years of relevant experience and native/professional English fluency...Remote jobHourly payWork at office
- YO AI Labs is seeking an experienced Adobe Marketing Technology Expert for a remote contractor role. The candidate will test marketing workflows and evaluate AI-generated outcomes across Adobe Workfront, AEM, CJA, Analytics, and Experience Cloud. You will identify workflow...Remote jobFor contractors
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluators for Gemini (US only). Be the first to apply!
- quality evaluator New York, NY
- clinical evaluator New York, NY
- work from home web search evaluator New York, NY
- evaluator New York, NY
- program evaluator New York, NY
- ai evaluator New York, NY
- education evaluator New York, NY
- social media evaluator New York, NY
- work from home social media evaluator New York, NY
- vocational evaluator

