Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Bilingual Language Model Evaluator

$15 - $20 per hour

Mercor

Job Description

Job Description

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .

Position: Generalist - English & Urdu
Type: Contract
Compensation: $15–$20/hour
Location: Remote

Role Responsibilities

  • Conduct fact-checking using trusted public sources and external tools .
  • Generate high-quality human evaluation data by identifying response strengths, areas for improvement, and factual inaccuracies.
  • Assess reasoning quality, clarity, tone, and completeness of responses.
  • Ensure model responses align with expected conversational behavior and system guidelines.
  • Work independently and asynchronously to meet deadlines while improving AI model performance .

Qualifications

Must-Have

  • Bachelor's degree .
  • Native speaker in Urdu .
  • Significant experience using large language models (LLMs).
  • Excellent writing skills in English .
  • Strong attention to detail .
  • Background or experience in domains requiring structured analytical thinking .

Preferred

  • Prior experience with RLHF, model evaluation, or data annotation work .
  • Experience writing or editing high-quality written content .
  • Experience comparing multiple outputs and making fine-grained qualitative judgments .

Application Process (Takes 20–30 mins to complete)

  • Upload resume
  • AI interview based on your resume
  • Submit form

Resources & Support

PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.

Vacancy posted 22 days ago
Similar jobs that could be interesting for youBased on the Bilingual Language Model Evaluator in San Francisco, CA vacancy
  • $218.5k - $288k

     ...The OpportunityAs an Applied Scientist specializing in Small Language Models and AI Training, you will lead research and development efforts...  ...small and efficient language models.Design, implement, and evaluate model training experiments to improve performance, robustness... 
    Suggested
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    3 days ago
  •  ...remotely. The role focuses on fact-checking and generating evaluation data, requiring native fluency in Urdu and strong English writing...  ...a bachelor’s degree and significant experience with large language models. Responsibilities also include independently assessing... 
    Bilingual
    Remote job

    Mercor

    San Francisco, CA
    3 days ago
  • Mercor is seeking experienced musicians to evaluate generative musical AI models in collaboration with a leading AI lab. You will assess model outputs across different categories of music in your bilingual language and contribute to structured taxonomy annotations. Ideal... 
    Bilingual
    Part time
    Immediate start
    10 hours per week

    Mercor

    San Francisco, CA
    1 day ago
  • Mercor is hiring experienced musicians to evaluate generative musical AI models in partnership with a leading AI lab. You will assess model outputs...  ..., voice generation, and other standards, using your bilingual language skills. Ideal candidates have 3+ years as a music... 
    Bilingual
    Part time
    Immediate start
    10 hours per week

    Mercor

    San Francisco, CA
    1 day ago
  • Mercor is hiring experienced Musicians to evaluate generative musical AI models in partnership with a leading AI lab. You will assess model outputs across in different categories of music in your bilingual language. Key Responsibilities Evaluate AI model output lyrics,... 
    Bilingual
    Part time
    Immediate start
    10 hours per week

    Mercor

    San Francisco, CA
    5 days ago
  • $50 - $75 per hour

    A leading tech company based in Australia is seeking an AI Model Evaluator on a contract basis. The role involves evaluating AI-generated responses, writing prompts, and providing justifications based on specific criteria. Ideal candidates will hold a Master's degree in... 
    Hourly pay
    Contract work

    Mercor

    San Francisco, CA
    5 days ago
  • Mercor is hiring experienced musicians to evaluate generative musical AI models in partnership with a leading AI lab. You will assess model outputs across different categories of music in your bilingual language. Ideal candidates have 3+ years in music production or audio... 
    Bilingual
    Part time
    Immediate start
    10 hours per week

    Mercor

    San Francisco, CA
    1 day ago
  • Welo Data is seeking Data Labeling Associates in California to evaluate AI outputs and ensure cultural context and safety in Arabic datasets. This role requires professional-level proficiency in Portuguese (Brazil), a bachelor's degree, and at least 2 years of experience... 
    Bilingual

    Welo Data

    San Francisco, CA
    5 days ago
  • A cutting-edge AI firm in San Francisco seeks a VLM Post-Training Owner to lead enterprise engagements and enhance vision-language models. The role combines project ownership with technical execution, ensuring quality data generation and customer satisfaction in AI solutions... 

    Liquid AI

    San Francisco, CA
    5 days ago
  • $15 per hour

     ...months Commitment: 20+ hours/week Role Responsibilities Evaluate AI-generated music across various genres and assess quality...  ...20–30 mins to complete) Upload your resume. Complete the Bilingual Competency Interview in Malayalam. Receive next steps and... 
    Bilingual
    Contract work
    Summer work
    Immediate start
    Remote work
    Flexible hours

    Mercor

    San Francisco, CA
    8 days ago
  • $85 per hour

     ...Location: Remote Role Responsibilities Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving cloud platforms , Kubernetes , CI/CD systems , and infrastructure... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    a month ago
  •  ...seeks experienced music producers and audio engineers to evaluate generative music AI models in collaboration with a leading AI lab. You'll assess AI...  ..., and rate quality against detailed standards. Work bilingual in Korean and English with flexible schedule for up to six... 
    Bilingual
    Immediate start
    Flexible hours

    Mercor

    San Francisco, CA
    2 days ago
  • $85 per hour

     ...$85/hour Location: Remote Role Responsibilities Use frontier AI coding agents to complete and evaluate complex engineering tasks. Review model-generated mobile application code for correctness, quality, maintainability, and performance. Identify bugs... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    22 days ago
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing infrastructure engineering tasks and reviewing model-generated implementations on cloud platforms, Kubernetes,... 

    Mercor

    San Francisco, CA
    2 days ago
  • $400 per month

     ...Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic infrastructure engineering workflows... 

    Mercor

    San Francisco, CA
    2 days ago
  • Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. You will assess AI-generated music across genres and rate it against detailed quality standards, working in Turkish and English... 
    Bilingual
    Immediate start

    Obsidian

    San Francisco, CA
    3 days ago
  • $75 - $85 per hour

     ...services utilizing the Parent Coaching Model (Rush & Sheldon) Follow each child’s Individualized...  ...) Requirements: California Speech Language Pathology license Pediatric experience...  ...Strong clinical & interpersonal skills Bilingual in Spanish a plus! Sunny Days is an... 
    Bilingual
    Part time
    Local area
    Work from home
    Relocation package
    Flexible hours

    Sunny Days

    San Francisco, CA
    4 days ago
  • Engineering Manager, Foundation Model Inference (FMAPI)RDQ427R519At Databricks, we are driven by a passion to empower data teams in...  ...ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political... 
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $89k - $105.7k

     ...Schools. We offer a Speech Language Pathology team lead that can...  ...practices for special education. Bilingual (Spanish or Vietnamese)...  ...Ability to perform diagnostic evaluations according to CA law and ASHA...  ...literacy achievement and to model scaffolding strategies. Support... 
    Bilingual
    Full time
    Temporary work
    Work at office

    KIPP Public Schools Northern California

    San Francisco, CA
    1 day ago
  • $80.1k

    Speech Language Pathologist - Clinical Fellow (CF-SLP) Department:...  ...Intervention Conducts speech-language evaluations and analyzes assessment...  ...professional learning and models communication strategies for...  .... Preferred Qualifications Bilingual in Spanish or Vietnamese. Experience... 
    Bilingual
    Full time
    Temporary work
    Work at office

    KIPP Public Schools Northern California

    San Francisco, CA
    2 days ago
  • $298k - $368k

     ...continuously learning from large scale real-world data, to (2) develop models and model training at scale, to (3) analyze real-world behavior...  ...and technical constraints ~ Experience applying large language models or foundation models in complex, safety-critical domains... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $100k - $115k

     ...Description Job Description Speech and Language Pathologist Felton ECE Programs...  ...Program Specific Responsibilities Screen, evaluate and treat children with communication...  ...assessing potential language delays Bilingual in Spanish/English required Additional... 
    Bilingual
    Full time
    Contract work
    Local area

    Felton Institute

    San Francisco, CA
    a month ago
  • $166k - $225k

     ...use deep data insights to improve their business. Databricks’ Model Serving product provides enterprises with a unified, scalable,...  ...models — from traditional ML to fine-tuned and proprietary large language models. It offers real-time, low-latency inference, governance,... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  • $172.43k - $230.95k

     ...The Senior Software Engineer for the AI Model Lifecycle team will play a crucial role...  ...Machine Learning models, including Large Language Models (LLMs).What You’ll Be Working On:...  ...experiment management: versioning, lineage, evaluation, and reproducible fine-tuning at scale.... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  • $119.77k - $140.9k

     ...Compliance organization. Specifically, this position supports the Model Risk Management (“MRM”) program at the Bank. The overall MRM...  ...necessity (ability to explain complex ideas in simple, non-technical language).· Able to perform complex mathematical analysis utilizing... 
    Full time
    Local area
    3 days per week

    US Bank

    San Francisco, CA
    4 hours ago
  • Obsidian is seeking a Spanish Audio Generalist Evaluator Expert to contribute to a high-impact audio AI research project. You will handle...  ..., and evaluation tasks to help train and benchmark advanced language models. The ideal candidate should have strong writing skills,... 
    Part time
    10 hours per week

    Obsidian

    San Francisco, CA
    4 days ago
  • $17 per hour

     ...Commitment: 10+ hours/week Role Responsibilities Evaluate AI model output lyrics , voice generation, and other standards in various...  .... Past experience writing lyrical music in the domain language listed in the title. Start Date ~ Immediate Application... 
    Remote job
    Contract work
    Summer work
    Immediate start

    Mercor

    San Francisco, CA
    7 days ago
  • $192k - $260k

     ...can use deep data insights to improve their business. Foundation Model Serving is the API Product for hosting and serving frontier AI...  ..., family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  • $29.81 - $35.4 per hour

    Speech Language Pathologist Assistant (SLP-A) - San Francisco Speech Language Pathologist...  ...Pathologist. Supports professional learning and modeling of speech and language strategies for...  ...'s mission. Preferred Qualifications Bilingual in Spanish or Vietnamese. Experience... 
    Bilingual
    Hourly pay
    Full time
    Internship
    Work at office
    Local area

    KIPP Public Schools Northern California

    San Francisco, CA
    5 days ago
  • $190k - $265k

     ...and products that power everything from data apps, AI agents, model training, model serving, and Vector Search. You'll be joining a...  ...Foundation Model APIs) team — the unified serving layer for large language models across real-time and batch inference, powering model... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Bilingual Language Model Evaluator. Be the first to apply!