AI Evaluation Specialist
Crossing Hurdles
A leading AI organization in Australia is seeking individuals with strong writing and analytical skills to evaluate and improve AI outputs. The ideal candidate must possess the ability to assess emotional nuances and detail while adhering to structured guidelines. Responsibilities include ensuring clarity and empathy in evaluations, identifying weaknesses in AI reasoning, and contributing to a safe AI environment. Immediate availability post-onboarding is required for this role. #J-18808-Ljbffr
- Mercor is seeking experienced music producers and audio engineers to evaluate generative music AI models. You will assess AI-generated tracks, rate musicality, creativity, and production quality, and label songs by genre and instruments. The role requires native Norwegian...SuggestedImmediate startFlexible hours
- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards. Join a collaborative...Suggested
- Dorado is seeking language professionals for a six‑month, independent contractor engagement to contribute to an AI evaluation project. Review AI‑generated content, annotate language data, and provide feedback to improve accuracy and reliability. The role requires reliability...SuggestedFor contractors
- ...create role-play scenarios across domains such as travel, financial services, telecoms and technical support, contributing to diverse evaluation datasets. Responsibilities include evaluating performance with metrics on task completion, naturalness, audio comprehension, and...SuggestedRemote jobContract work
- AuraOne is seeking a Physical Sciences Research Assistant for a Remote AI Evaluation track. You will review AI outputs in physics, reproduce key derivations, and document correct methods to help train modeling systems. This contractor-style role emphasizes rigorous reasoning...SuggestedRemote jobPart timeFor contractors
- Dorado is seeking Speech AI Evaluation Specialists to support AI content improvement. This freelance, part-time role is based remotely from Malaysia, with 10+ hours per week and a starting date immediately. You will evaluate Vietnamese-language responses and provide structured...Remote jobPart timeFreelanceImmediate start10 hours per week
- A tech company specializing in AI projects is seeking skilled LibreSprite users to assist in evaluating AI-generated visual content. As an independent contractor, you can work flexibly from anywhere, contributing around 5-20 hours per week depending on project needs. Ideal...Remote jobFor contractors
- ...Human Resources Expert for a remote contract in the United States. You will design realistic HR scenarios, craft model prompts, and evaluate AI-generated HR outputs while ensuring alignment with employment laws and best practices. You will collaborate with the team to...Remote jobContract work
- Welo Data is hiring Data Labeling Associates in New York City for Project Perseus. The role involves evaluating Arabic language AI outputs and ensuring AI safety, requiring professional proficiency in Arabic and experience in writing and AI safety. You’ll critique models...
- Mercor is seeking experienced music producers and audio engineers to evaluate generative music AI models in partnership with a leading AI lab. You will assess AI-generated music across a wide range of genres, rate it against detailed quality standards, and perform labeling...Flexible hours
- Cincinnatus LLC, the employer of record, is seeking a seasoned Retail SME to help train AI models by evaluating outputs against rubrics and applying real-world retail judgment. The role requires 8+ years in retail and the ability to commit at least 35 hours per week, with...
- CNTXT AI is seeking a fully remote, hourly contractor to support AI data and language projects on a flexible, project-based schedule. The role involves content generation, data annotation, LLM evaluation, and localization QA across diverse topics. Ideal candidates are native...Remote jobHourly payFor contractorsFlexible hours
- ...Job TitleAI Evaluation EngineerLocationHybrid / RemoteEmployment TypeFull-timeJob SummaryWe are seeking an AI Evaluation Engineer to design, implement, and maintain evaluation frameworks for AI and machine learning systems, with a focus on Large Language Models (LLMs)...
$161.6k - $200k
...AI Evaluation EngineerLocations: Charlotte, North Carolina, United States; Denver, Colorado, United States; New York, New York, United StatesAbout Judi HealthJudi Health is an enterprise health technology company providing a comprehensive suite of solutions for employers...Local areaFlexible hours$50 - $190 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Remote Commitment: 20+ hours/week Role Responsibilities Evaluate AI systems on complex personal workflows, including personal...Hourly payContract workFor contractorsSummer workRemote workTrial period$80 per hour
...human intelligence to ethically shape the future of AI. What We Do The Mindrift platform connects specialists with AI projects from major tech innovators. Our... ...for someone who can design realistic and structured evaluation scenarios for LLM-based agents. You'll create test...Part timeFreelanceRemote workFlexible hours$80 per hour
...CV in English and indicate your level of English proficiency. Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment...Permanent employmentTemporary work$229.9k - $262.4k
...Senior Lead AI Engineer (SDK's: Gen AI Evaluation and MCP) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized...Full timePart timeLocal area$80 - $120 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Jack Dorsey . Position: Data analysis / quantitative readouts Evaluator Type: Contract Compensation: $80–$120/hour...Contract workSummer workWork at officeRemote work$80 - $120 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...and Jack Dorsey . Position: Biology / environmental science Evaluator Type: Contract Compensation: $80–$120/hour Location...Contract workSummer workWork at officeRemote work$50 - $100 per hour
...progressive technology company is looking for expert biologists to join their team part-time. Responsibilities include annotating and evaluating AI tool outputs, generating evaluation tasks, and conducting literature searches. Candidates should have a PhD or Masters in...Hourly payPart time- ...Member of Technical Staff - Evals located in New York, NY, where you will ensure our AI-powered features are high quality and reliable. Your responsibilities include designing evaluation frameworks, building automated tests, and developing tools for seamless evaluations....
- Scale AI, Inc. is looking for a Research Scientist focused on Frontier Risk Evaluations to develop evaluation measures and datasets for assessing AI risks. As a key team member, you will engage with government agencies and publish methodologies that influence AI safety...
$50 - $70 per hour
United States Digital Space LLC is seeking a Data Scientist for AI Evaluation Analytics to work as a remote contractor on long-term projects. You will develop evaluation metrics, validate datasets, analyze signals, and apply statistical methods to measure AI system performance...Remote jobHourly payFor contractors$152k - $241.5k
...us! We believe open-weight models are foundational to American AI leadership and cybersecurity, and that trust in AI grows through... ...scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find, validate, and patch software...Full timeRemote work$30 - $50 per hour
A tech company is seeking an AI Researcher to support end-to-end research for modern AI systems. This remote role involves designing experiments, defining evaluation protocols, and improving evaluation rigor for large language models. Key responsibilities include developing...Remote jobHourly pay- Rex.zone is seeking an AI Research Scientist to lead applied AI research projects for US-based customers, translating open-ended questions into measurable experiments in LLM evaluation and RLHF data design. You will evaluate prompts, design datasets, and work with cross...Remote jobHourly payFlexible hours
- Mercor seeks experienced music producers and audio engineers to evaluate generative music AI models, working in Korean and English to assess AI-generated tracks against detailed quality standards. Responsibilities include comparing songs for musicality, creativity, adherence...Contract workImmediate startFlexible hours
- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Italian...
- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models. You will assess AI-generated music across genres and rate it against detailed quality standards, working in Hindi and English. Responsibilities include head-to-head...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Specialist. Be the first to apply!


