Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Bilingual Language Model Evaluator

$15 - $20 per hour
Part-time

Mercor

Role Description

  • Conduct fact-checking using trusted public sources and external tools.
  • Generate high-quality human evaluation data by identifying response strengths, areas for improvement, and factual inaccuracies.
  • Assess reasoning quality, clarity, tone, and completeness of responses.
  • Ensure model responses align with expected conversational behavior and system guidelines.
  • Work independently and asynchronously to meet deadlines while improving AI model performance.

Qualifications

  • Must-Have:
  • Bachelor's degree.
  • Native speaker in Urdu.
  • Significant experience using large language models (LLMs).
  • Excellent writing skills in English.
  • Strong attention to detail.
  • Background or experience in domains requiring structured analytical thinking.
  • Preferred:
  • Prior experience with RLHF, model evaluation, or data annotation work.
  • Experience writing or editing high-quality written content.
  • Experience comparing multiple outputs and making fine-grained qualitative judgments.

Company Description

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.

Vacancy posted 10 days ago
Similar jobs that could be interesting for youBased on the Bilingual Language Model Evaluator in Remote vacancy
  • $15 - $20 per hour

     ...tools . Generate high-quality human evaluation data by identifying response strengths,...  ...and completeness of responses. Ensure model responses align with expected conversational...  .... Significant experience using large language models (LLMs). Excellent writing... 
    Bilingual
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    19 days ago
  • $15 - $20 per hour

     ...Generate high-quality human evaluation data by identifying response...  ...completeness of responses. Ensure model responses align with expected...  ...experience using large language models (LLMs) and an understanding...  .... Candidates must complete a Bilingual Competency interview in... 
    Bilingual
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    13 days ago
  •  ...healthcare technology company is seeking a Pyxis Technician to train AI models and ensure their medical accuracy. Ideal candidates will have...  ...bonuses for high-quality work. Candidates must be fluently bilingual in English and located in the United States. This is an... 
    Bilingual
    Hourly pay
    For contractors
    Remote work

    DataAnnotation

    Wyoming, OH
    2 days ago
  • $40 per hour

     ...cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content, solve technical...  ...experience required ~ Fluency in English (native or bilingual level) ~ Strong writing and analytical skills ~ A bachelor... 
    Bilingual
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Juneau, AK
    2 days ago
  • $500 per day

     ...We're looking to hire NYS certified Speech Language Pathologists (SLPs), to conduct CSE (school-age) Speech and Language Evaluations in person. Job Type:  Flexible Schedule...  ...turnaround time. We have monolingual and bilingual evaluators in all areas of the city and we... 
    Bilingual
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Work from home
    Flexible hours

    AngelCare ABA Therapy

    New York, NY
    18 days ago
  • $15 per hour

     ...opportunity to engage in analytical tasks that enhance the capabilities of large language models (LLMs). As a Business Analyst, you will utilize your strong analytical skills and bilingual proficiency in English and French to dissect complex content, conduct thorough online... 
    Bilingual
    Contract work
    Remote work
    Flexible hours

    SaidGig

    United States
    17 days ago
  • $45 - $95 per hour

     ...Role Overview As a Tongan Bilingual Expert, you will play a crucial role in a dynamic project...  ...AI systems, shaping how these models learn, reason, and perform through high-...  ...users. Provide detailed feedback on language usage, tone, and linguistic nuances present... 
    Bilingual
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Indiana
    19 days ago
  • $45 - $95 per hour

     ...As a Lingala Bilingual Expert, you will leverage your language skills to contribute to the training of next-generation...  ...a crucial role in shaping how AI models learn and perform by providing high...  ...intent of speakers. Analyze and evaluate grammar, tone, and overall... 
    Bilingual
    Remote job

    SaidGig

    United States
    a month ago
  • $15 - $95 per hour

     ...As a Dioula Bilingual Expert, you will play a crucial role in training...  ...directly influence how AI models learn, reason, and perform, making...  ...structure, and tone in both languages to improve language model...  ...Strong analytical skills to evaluate grammar and provide comprehensive... 
    Bilingual
    Remote work

    SaidGig

    United States
    17 days ago
  • $20 - $33 per hour

     ...Role Overview As a Hebrew Language Expert, you will play a crucial role in training next-generation AI systems by providing high-quality...  ...-world input. Your expertise will directly influence how AI models learn, reason, and perform, making your contributions essential... 
    Remote job
    Hourly pay
    Contract work

    SaidGig

    United States
    a month ago
  • $80 - $120 per hour

     ...Role Overview As an Evaluator in Public-sector procurement and RFI response, you will leverage your expertise to review and assess AI-generated work products, including documents, spreadsheets, and slide decks. Your role is crucial in ensuring the accuracy, rigor, and... 
    Hourly pay
    Work at office
    Remote work

    SaidGig

    United States
    1 day ago
  • $80 - $120 per hour

     ...In this role, you will leverage your expertise in General Sales and Go-To-Market (GTM) strategies to evaluate and assess AI-generated work products, including documents, spreadsheets, and slide decks. Your primary focus will be on ensuring accuracy, rigor, and domain... 
    Remote job
    Hourly pay
    Work at office

    SaidGig

    United States
    a month ago
  • $80 - $120 per hour

     ...As an Evaluator in Personal Finance and Consumer Planning, you will play a crucial role in reviewing and assessing AI-generated work products, including documents, spreadsheets, and slide decks. Your expertise will ensure the accuracy, rigor, and overall quality of these... 
    Remote job
    Hourly pay
    Work at office

    SaidGig

    United States
    a month ago
  • $80 - $120 per hour

     ...Role Overview As an Evaluator in Market Research and Competitive Intelligence, you will leverage your expertise to review and assess AI-generated work products, including documents, spreadsheets, and slide decks. Your role is crucial in ensuring the accuracy, rigor, and... 
    Remote job
    Hourly pay
    Work at office

    SaidGig

    United States
    a month ago
  • $80 - $120 per hour

     ...As a Process Improvement / SOPs Evaluator, you will leverage your expertise to review and assess AI-generated work products, including documents, spreadsheets, and slide decks. Your role is crucial in ensuring the accuracy, rigor, and quality of these outputs by applying... 
    Remote job
    Hourly pay
    Work at office

    SaidGig

    United States
    a month ago
  • $80 - $120 per hour

     ...Role Overview As a Cybersecurity / IT GRC Evaluator, you will leverage your expertise to review and assess AI-generated work products, including documents, spreadsheets, and slide decks. Your role is crucial in ensuring the accuracy, rigor, and quality of these outputs... 
    Remote job
    Hourly pay
    Work at office

    SaidGig

    United States
    a month ago
  • $10 - $40 per hour

     ...As a Turkish Language Expert, you will play a crucial role in training next-generation AI...  ...expertise will directly influence how AI models learn, reason, and perform by providing...  ...and consistency in written Turkish. Evaluate and refine audio clips for phonetic clarity... 
    Remote job

    SaidGig

    United States
    a month ago
  • $80 - $120 per hour

     ...This role involves evaluating and assessing AI-generated work products, including documents, spreadsheets, and slide decks, specifically within the context of nonprofit, philanthropy, and community programs. As an expert evaluator, you will leverage your deep subject-... 
    Remote job
    Hourly pay
    Work at office

    SaidGig

    United States
    a month ago
  • $80 - $120 per hour

     ...In this role, expert Evaluators in Clinical, Biomedical, or Pharma will review and assess AI-generated work products, including documents, spreadsheets, and slide decks, ensuring accuracy, rigor, and domain quality. Your deep subject-matter expertise will be crucial in... 
    Remote job
    Hourly pay
    Work at office

    SaidGig

    United States
    a month ago
  •  ...their team in a remote capacity. The role involves training AI models by providing coding challenges and assessing the output for quality...  .... Ideal candidates will have proficiency in programming languages like Python or JavaScript and be detail-oriented. Benefits include... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Hartford, CT
    2 days ago
  •  ...Welocalize is seeking a Data Quality Associate in New York to evaluate AI model outputs and provide structured feedback. This full-time position requires native-level language proficiency and a university degree. You will contribute to improving evaluation frameworks while... 
    Full time
    Remote work

    Welocalize

    New York, NY
    2 days ago
  •  ...Prolific is seeking fluent Thai speakers to act as evaluators for AI language models. You will assess how naturally Thai is spoken in AI-generated text and audio, ensuring cultural and contextual accuracy. This is a remote, paid task with flexible hours and competitive... 
    Remote work
    Flexible hours

    Prolific

    Los Angeles, CA
    3 days ago
  • $70 per hour

     ...Position: AI Model Assessment Specialist Type: Contract Compensation: $22 - $70/hour Location: Remote Commitment: 10-40 hrs/week Role Responsibilities Evaluate and critique the performance and accuracy of AI-generated content across various domains. Deliver detailed,... 
    Contract work
    Remote work

    Crossing Hurdles

    New York, NY
    2 days ago
  • $25 per hour

    A leading tech company is seeking a Language Specialist to improve AI chatbots by evaluating their progress and teaching them conversational skills in English and Japanese. This role allows flexibility with remote work and the choice of projects. Candidates should have... 
    Hourly pay
    Remote work

    DataAnnotation

    New York, NY
    3 days ago
  • $80 - $120 per hour

     ...Role Overview As an Evaluator specializing in Product launch and experiment readiness, you will play a crucial role in reviewing and assessing AI-generated work products, including documents, spreadsheets, and slide decks. Your expertise will ensure that these outputs... 
    Remote job
    Hourly pay
    Work at office

    SaidGig

    United States
    a month ago
  • $40 per hour

    An AI training company in the United States is looking for a Statistician to help improve AI models. You will evaluate the mathematics logic behind chatbots and assess their performance. Candidates should hold expertise in various branches of mathematics and strong attention... 
    Hourly pay
    Contract work
    Remote work

    DataAnnotation

    Honolulu, HI
    3 days ago
  • $40 per hour

     ...data annotation company is seeking a Statistician to join their team. This remote role involves training AI models by posing complex mathematical problems, evaluating outputs, and assessing the model's performance. Candidates must be detail-oriented and proficient in... 
    Hourly pay
    Remote work

    DataAnnotation

    New York, NY
    2 days ago
  • $80 - $120 per hour

     ...fundraising, and pitchbooks to review and assess AI-generated documents, spreadsheets, and slide decks. Your primary focus will be on evaluating the accuracy, rigor, and overall quality of these outputs, ensuring they meet high standards of excellence. Key Responsibilities... 
    Hourly pay
    Work at office
    Remote work

    SaidGig

    United States
    2 days ago
  • $30 per hour

    A technology company is seeking a Web Platform Engineer to evaluate AI chatbots and enhance model performance. This role requires proficiency in programming languages like Python and JavaScript. You will assess AI outputs from coding challenges and writing tasks, ensuring... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Jackson, MS
    2 days ago
  • $30 - $40 per hour

    An AI training company is seeking a Web Platform Engineer to evaluate AI chatbots' outputs and improve their logic. The role allows for...  ...be fluent in English and have experience with programming languages including Python and JavaScript. A Bachelor's degree is preferred... 
    Hourly pay
    For contractors
    Remote work

    DataAnnotation

    Little Rock, AR
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Bilingual Language Model Evaluator. Be the first to apply!