AI Model Evaluation Lead: Metrics, Bias & Fairness
MERIT Beauty
We are seeking an expert to evaluate and improve our AI models through comprehensive testing and analysis. You will be responsible for designing evaluation frameworks, conducting model assessments, and providing actionable insights for model improvement. Key Responsibilities Design and implement evaluation metrics for AI models Conduct thorough testing of model performance across different scenarios Analyze model outputs for bias, fairness, and accuracy Collaborate with ML engineers to implement improvements Document findings and recommendations Ideal Candidate PhD or Masters in Computer Science, ML, or related field 5+ years experience in AI/ML model evaluation Strong background in statistical analysis Experience with evaluation frameworks and metrics Excellent communication skills #J-18808-Ljbffr MERIT Beauty
- ...interpretable, and steerable AI systems. We want AI to... ...Engineers to build the evaluations that tell us — and the... ...into clear, defensible metrics that researchers,... ...leadership use to monitor model health during training,... ...infrastructure A bias toward picking up slack...SuggestedWork at officeVisa sponsorshipFlexible hours
$124k - $177k
...Intelligence & Data (AI&D) organization,... ...to ensure proper modeling processes are... ...practical controls and evaluations (e.g., bias testing, stress... ..., and evaluation metrics.• Parallel... ...for drift, bias/fairness, stability, hallucinations... ...and agents are leading the industry and...SuggestedLocal area3 days per week- ...independent assurance and evaluating the company's risk... ...scientists and AI developers who... ...for auditing models, including criteria... ...like robustness, fairness, interpretability,... ...vulnerabilities including bias, fairness... ...strategies.- Measurement Metrics & Statistical...Suggested
- Anthropic in the United States is hiring Research Engineers to build evaluations that quantify Claude's capabilities and measure reasoning,... ...at scale. You will design end-to-end experiments, define metrics, and develop the infrastructure to run evaluations using live...Suggested
$182k - $242k
CoreWeave, the AI Hyperscaler™, acquired Weights & Biases to create the most powerful end-to-... ...together CoreWeave's industry-leading cloud infrastructure with... ...experiment tracking and model optimization to high-... ...visualize complex training metrics in real-time. About the...SuggestedPermanent employmentTemporary workCasual workWork at officeRemote workFlexible hours$228.7k - $343.1k
...scale, and one bad model can mean millions in... ...unreported, or a fair lending violation.... ...clear every headline metric and still be broken... ...models applies to AI. We build the tooling... ...so you critically evaluate what it produces and... ...contributor, you lead through technical depth...Remote jobFull timeLocal areaShift work$118.98k - $195.47k
...join our team as a Finance Model & AI Solutions Lead.The colleague in this role... ...that automatically create and evaluate multiple planning... ...ethical guidelines, including bias mitigation and explainability... ...retrievalUnderstanding of AI model evaluation metrics and performance...Full timeH1bWork at officeVisa sponsorshipWork visaFlexible hours3 days per week- ...Role: AI Team Lead Location: NYC, NY Project description... ...convert research and models into scalable,... ...model training, evaluation, versioning, and monitoring... ...observability: metrics, logging, model drift... ...validation techniques, and bias/fairness considerations....
- A leading AI development company is seeking experienced quantitative professionals for remote work evaluating AI-generated quantitative analysis. Ideal candidates will have a robust background in fields like data science, economics, or biostatistics, with at least 2 years...Remote work
$40 per hour
A growing technology company is seeking an R&D Biologist to join their team. This role involves training AI models by evaluating chatbot outputs on complex biology questions. Applicants should possess strong expertise in biology and related fields. You will work remotely...Hourly payRemote work$40 per hour
A leading data annotation company is seeking a Statistician to join their team. This remote role involves training AI models by posing complex mathematical problems, evaluating outputs, and assessing the model's performance. Candidates must be detail-oriented and proficient...Hourly payRemote work- Dorado is seeking an experienced Investment Banking SME to support the development, evaluation, and improvement of advanced AI models in finance. You will assess AI-generated analyses for accuracy, reasoned judgments, and data integrity, guiding model refinements and prompts...Remote job
- AuraOne is seeking a remote Spoken Instruction Conversation Evaluator to review prompts and responses against its quality rubric. You will... ...edge cases, and provide structured feedback for retraining the model. As an independent contractor, you will evaluate model outputs,...Remote jobHourly payContract workFor contractors
- Productive Playhouse seeks AI Evaluators to support evaluating AI chatbots by interacting with models, assessing capabilities, safety, and usefulness. This is a project-based, task-based engagement with flexible hours and batch deliveries. Open to freelancers outside the...Remote jobFreelanceFlexible hours
- SME Careers is seeking biologists to contribute to an AI training project that involves reviewing AI-generated responses and providing... ...hold a MS or PhD in a relevant field and have experience in evaluating complex biology content. Strong communication skills and proficient...Immediate start
- Dorado is seeking a Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems, probe model reasoning at the frontier, and help identify where models fail under rigorous scientific scrutiny....Remote job
- Obsidian is seeking Insurance domain SMEs to join a cutting-edge AI training program. You will evaluate AI model outputs against real underwriting practice and rubrics, guiding research and engineering teams to close knowledge gaps in underwriting, claims, and risk assessment...
- Mercor is seeking experienced Musicians to evaluate generative musical AI models in partnership with a leading AI lab. In this role, you will assess model outputs across categories of music in your bilingual language and contribute to improving AI-driven music systems....
- Who are we? Cohere is the leading security-first enterprise AI company. We build cutting‑edge foundation AI models and end‑to‑end products that are designed to solve real-world... ...and Paris. Join us! Why this role? Evaluation is critical to making progress in scaling...Full timeWork at officeLocal areaRemote workHome office
- ...Quant Model Risk Vice President Bring your expertise to JPMorganChase. As part of Risk... ...for specific products and structures. Evaluate model behavior and ensure the suitability... ...and maintain robust model performance metrics to compare and monitor the outcomes of various...
$200k - $275k
IT Operating Model and Strategy Advisory Lead Join Mizuho as the IT Operating Model and Strategy Advisory Lead! The Office of Strategic Advisory is the... ...and complexity while positioning the firm for scalable, AI‑enabled growth. Operating with delegated Co‑CIO authority...Contract workWork at officeLocal areaRemote workWorldwide- ...seeking a remote Kotlin Engineer to review AI-generated responses and create high-... ...optimizing AI performance, and ensuring model accuracy. The ideal candidate has a Bachelor... ...an expert network and requires critical evaluation of technical concepts. #J-18808-Ljbffr SME...Remote job
- We are seeking an experienced AI/ML Model Validation and Governance professional to join a... ...bonus. This position will independently evaluate predictive, machine learning,... ...testing. Assess model explainability, bias, fairness, robustness, and regulatory compliance....Full timeRelocation package3 days per week
$107.5k - $188.4k
...seeking a hands-on Lead Product Manager to... ...define, build, and ship AI-enabled platform... ...growing library of Model Context Protocol (... ...defining product success metrics and operating with... ...working with evaluation frameworks for AI systems... ...maintain a fair and genuine hiring...Full timeWork at office$130.97k - $183.68k
...Know the Opportunity:The Lead AI Product Manager is a senior... ...in large language models (LLMs), agentic systems, AI evaluation methodologies, and regulated... ...intelligence.Establish success metrics and own business outcomes... ...committed to pay that’s fair and equitable, which...Full timePart timeWork experience placementLocal areaWork from homeFlexible hours$185k - $200k
...AI Governance Lead This is an opportunity to join Ascot Group... ..., Predictive Modeling & AI Strategy, the AI... ...scoring frameworks (e.g., fairness, explainability, reversibility... ...explainability, Bias/fairness testing,... ...-service guidance Evaluate governance maturity of...Temporary workWork at officeLocal areaFlexible hours- ...creating a diverse, fair and respectful culture... ...2017, Wayve is the leading developer of Embodied AI technology. Our... ...software and foundation models enable vehicles to perceive... ...pretraining and evaluation Manage and mentor a... ...(e.g., Weights & Biases, ClearML, Labelbox,...Full timeWork at officeRemote workWork from homeVisa sponsorshipRelocation packageFlexible hours
$89.25k - $150.25k
...strategy and coverage for AI/Gen AI. The Manager will... ...planning and execution, evaluating the design and... ...Team: Our Data Science and Model Risk/AI team plays a critical... ...Gen AI risks, including bias, ethics, compliance, privacy... ...solutions to the team Lead audit client meetings...Work experience placementLocal areaWorldwide- ...Developer Experience Lead At Lead, we... ...feedback, and adoption metrics, and close the loop... ...our object model and API design — into... ...developer showcases. AI & Documentation... ...content itself. Evaluate emerging AI tools and... ...the San Francisco Fair Chance Ordinance, we...Flexible hours
$40 per hour
A leading data annotation company is seeking a Statistician for a remote role. You will train AI models by measuring progress and solving problems related to AI chatbots. Candidates should have proficiency in various mathematical fields and fluency in English. This position...Hourly payRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Model Evaluation Lead: Metrics, Bias & Fairness. Be the first to apply!


