Bilingual Language Model Evaluator
$15 - $20 per hourMercor
Job Description
Job Description
About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .
Position: Generalist - English & Urdu
Type: Contract
Compensation: $15–$20/hour
Location: Remote
Role Responsibilities
- Conduct fact-checking using trusted public sources and external tools .
- Generate high-quality human evaluation data by identifying response strengths, areas for improvement, and factual inaccuracies.
- Assess reasoning quality, clarity, tone, and completeness of responses.
- Ensure model responses align with expected conversational behavior and system guidelines.
- Work independently and asynchronously to meet deadlines while improving AI model performance .
Qualifications
Must-Have
- Bachelor's degree .
- Native speaker in Urdu .
- Significant experience using large language models (LLMs).
- Excellent writing skills in English .
- Strong attention to detail .
- Background or experience in domains requiring structured analytical thinking .
Preferred
- Prior experience with RLHF, model evaluation, or data annotation work .
- Experience writing or editing high-quality written content .
- Experience comparing multiple outputs and making fine-grained qualitative judgments .
Application Process (Takes 20–30 mins to complete)
- Upload resume
- AI interview based on your resume
- Submit form
Resources & Support
- For details about the interview process and platform information, please check:
- For any help or support, reach out to: View email address on ziprecruiter.com
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
$218.5k - $288k
...The OpportunityAs an Applied Scientist specializing in Small Language Models and AI Training, you will lead research and development efforts... ...small and efficient language models.Design, implement, and evaluate model training experiments to improve performance, robustness...SuggestedWork at officeFlexible hours3 days per week- ...remotely. The role focuses on fact-checking and generating evaluation data, requiring native fluency in Urdu and strong English writing... ...a bachelor’s degree and significant experience with large language models. Responsibilities also include independently assessing...BilingualRemote job
- Mercor is seeking experienced musicians to evaluate generative musical AI models in collaboration with a leading AI lab. You will assess model outputs across different categories of music in your bilingual language and contribute to structured taxonomy annotations. Ideal...BilingualPart timeImmediate start10 hours per week
- Mercor is hiring experienced musicians to evaluate generative musical AI models in partnership with a leading AI lab. You will assess model outputs... ..., voice generation, and other standards, using your bilingual language skills. Ideal candidates have 3+ years as a music...BilingualPart timeImmediate start10 hours per week
- Mercor is hiring experienced Musicians to evaluate generative musical AI models in partnership with a leading AI lab. You will assess model outputs across in different categories of music in your bilingual language. Key Responsibilities Evaluate AI model output lyrics,...BilingualPart timeImmediate start10 hours per week
$50 - $75 per hour
A leading tech company based in Australia is seeking an AI Model Evaluator on a contract basis. The role involves evaluating AI-generated responses, writing prompts, and providing justifications based on specific criteria. Ideal candidates will hold a Master's degree in...Hourly payContract work- Mercor is hiring experienced musicians to evaluate generative musical AI models in partnership with a leading AI lab. You will assess model outputs across different categories of music in your bilingual language. Ideal candidates have 3+ years in music production or audio...BilingualPart timeImmediate start10 hours per week
- Welo Data is seeking Data Labeling Associates in California to evaluate AI outputs and ensure cultural context and safety in Arabic datasets. This role requires professional-level proficiency in Portuguese (Brazil), a bachelor's degree, and at least 2 years of experience...Bilingual
- A cutting-edge AI firm in San Francisco seeks a VLM Post-Training Owner to lead enterprise engagements and enhance vision-language models. The role combines project ownership with technical execution, ensuring quality data generation and customer satisfaction in AI solutions...
$15 per hour
...months Commitment: 20+ hours/week Role Responsibilities Evaluate AI-generated music across various genres and assess quality... ...20–30 mins to complete) Upload your resume. Complete the Bilingual Competency Interview in Malayalam. Receive next steps and...BilingualContract workSummer workImmediate startRemote workFlexible hours$85 per hour
...Location: Remote Role Responsibilities Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving cloud platforms , Kubernetes , CI/CD systems , and infrastructure...Contract workSummer workRemote work- ...seeks experienced music producers and audio engineers to evaluate generative music AI models in collaboration with a leading AI lab. You'll assess AI... ..., and rate quality against detailed standards. Work bilingual in Korean and English with flexible schedule for up to six...BilingualImmediate startFlexible hours
$85 per hour
...$85/hour Location: Remote Role Responsibilities Use frontier AI coding agents to complete and evaluate complex engineering tasks. Review model-generated mobile application code for correctness, quality, maintainability, and performance. Identify bugs...Contract workSummer workRemote work- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing infrastructure engineering tasks and reviewing model-generated implementations on cloud platforms, Kubernetes,...
$400 per month
...Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic infrastructure engineering workflows...- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. You will assess AI-generated music across genres and rate it against detailed quality standards, working in Turkish and English...BilingualImmediate start
$75 - $85 per hour
...services utilizing the Parent Coaching Model (Rush & Sheldon) Follow each child’s Individualized... ...) Requirements: California Speech Language Pathology license Pediatric experience... ...Strong clinical & interpersonal skills Bilingual in Spanish a plus! Sunny Days is an...BilingualPart timeLocal areaWork from homeRelocation packageFlexible hours- Engineering Manager, Foundation Model Inference (FMAPI)RDQ427R519At Databricks, we are driven by a passion to empower data teams in... ...ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political...Worldwide
$89k - $105.7k
...Schools. We offer a Speech Language Pathology team lead that can... ...practices for special education. Bilingual (Spanish or Vietnamese)... ...Ability to perform diagnostic evaluations according to CA law and ASHA... ...literacy achievement and to model scaffolding strategies. Support...BilingualFull timeTemporary workWork at office$80.1k
Speech Language Pathologist - Clinical Fellow (CF-SLP) Department:... ...Intervention Conducts speech-language evaluations and analyzes assessment... ...professional learning and models communication strategies for... .... Preferred Qualifications Bilingual in Spanish or Vietnamese. Experience...BilingualFull timeTemporary workWork at office$298k - $368k
...continuously learning from large scale real-world data, to (2) develop models and model training at scale, to (3) analyze real-world behavior... ...and technical constraints ~ Experience applying large language models or foundation models in complex, safety-critical domains...Full timeRemote work$100k - $115k
...Description Job Description Speech and Language Pathologist Felton ECE Programs... ...Program Specific Responsibilities Screen, evaluate and treat children with communication... ...assessing potential language delays Bilingual in Spanish/English required Additional...BilingualFull timeContract workLocal area$166k - $225k
...use deep data insights to improve their business. Databricks’ Model Serving product provides enterprises with a unified, scalable,... ...models — from traditional ML to fine-tuned and proprietary large language models. It offers real-time, low-latency inference, governance,...Local areaWorldwide$172.43k - $230.95k
...The Senior Software Engineer for the AI Model Lifecycle team will play a crucial role... ...Machine Learning models, including Large Language Models (LLMs).What You’ll Be Working On:... ...experiment management: versioning, lineage, evaluation, and reproducible fine-tuning at scale....Temporary work$119.77k - $140.9k
...Compliance organization. Specifically, this position supports the Model Risk Management (“MRM”) program at the Bank. The overall MRM... ...necessity (ability to explain complex ideas in simple, non-technical language).· Able to perform complex mathematical analysis utilizing...Full timeLocal area3 days per week- Obsidian is seeking a Spanish Audio Generalist Evaluator Expert to contribute to a high-impact audio AI research project. You will handle... ..., and evaluation tasks to help train and benchmark advanced language models. The ideal candidate should have strong writing skills,...Part time10 hours per week
$17 per hour
...Commitment: 10+ hours/week Role Responsibilities Evaluate AI model output lyrics , voice generation, and other standards in various... .... Past experience writing lyrical music in the domain language listed in the title. Start Date ~ Immediate Application...Remote jobContract workSummer workImmediate start$192k - $260k
...can use deep data insights to improve their business. Foundation Model Serving is the API Product for hosting and serving frontier AI... ..., family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation...Local areaWorldwide$29.81 - $35.4 per hour
Speech Language Pathologist Assistant (SLP-A) - San Francisco Speech Language Pathologist... ...Pathologist. Supports professional learning and modeling of speech and language strategies for... ...'s mission. Preferred Qualifications Bilingual in Spanish or Vietnamese. Experience...BilingualHourly payFull timeInternshipWork at officeLocal area$190k - $265k
...and products that power everything from data apps, AI agents, model training, model serving, and Vector Search. You'll be joining a... ...Foundation Model APIs) team — the unified serving layer for large language models across real-time and batch inference, powering model...Local areaWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Bilingual Language Model Evaluator. Be the first to apply!
- work from home web search evaluator San Francisco, CA
- social media evaluator San Francisco, CA
- evaluator San Francisco, CA
- quality evaluator San Francisco, CA
- education evaluator San Francisco, CA
- ai evaluator San Francisco, CA
- bilingual spanish receptionist San Francisco, CA
- bilingual spanish virtual assistant San Francisco, CA
- bilingual caregiver San Francisco, CA
- customer service bilingual San Francisco, CA



