AI Evaluation Scientist
Convergenz
We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements. This role is essential for establishing trust in AI solutions and supporting continuous improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment.
- Implement evaluation frameworks for AI models, including accuracy, robustness, relevance, bias, hallucination rate, and safety metrics.
- Build and maintain automated evaluation scripts, tests, and pipelines that assess AI model outputs and detect performance drift over time.
- Develop benchmark datasets, challenge sets, and scenario-based test cases tailored to mission and user needs.
- Perform structured error analysis and behavioral audits of LLMs, retrieval-augmented generation (RAG) systems, and predictive models, documenting findings and improvement recommendations.
- Collaborate with AI Developers, LLMOps Engineers, and Data Scientists to support iterative experimentation, model hardening, and quality improvements.
- Contribute to the design of human-in-the-loop evaluation workflows, integrating qualitative and quantitative insight into evaluation reports.
- Assist in mapping evaluation outcomes to responsible AI principles such as fairness, transparency, reliability, and safety.
- Partner with AI Governance Analysts to ensure evaluation outputs support compliance, documentation, and risk assessments.
- Stay current with emerging evaluation tools, frameworks, metrics, and research related to LLM assessment and generative AI reliability.
- Document evaluation processes, criteria, and results for both technical and non-technical audiences.
- Ability to hold a position of public trust with the U.S. government.
- Bachelor's degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field and 5+ years of experience.
- Master's degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field and 3+ years of experience.
- 2+ years of experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred).
- Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design for AI systems.
- Proficiency in Python and relevant libraries such as PyTorch, Hugging Face, scikit-learn, LangChain.
- Proficiency in AI evaluation frameworks such as Ragas or DeepEval.
- Proficiency in AI traceability/observability tools based on the OpenTelemetry protocol.
- Experience analyzing structured and unstructured data, including text, documents, and embeddings.
- Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures.
- Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness, transparency, reliability).
- Strong analytical skills with the ability to interpret model weaknesses, extract insights, and recommend actionable improvements.
- Excellent written and verbal communication skills, with the ability to present evaluation findings clearly to technical and non-technical stakeholders.
- Experience working in agile or iterative development environments is a plus.
- Familiarity with OWASP LLM Top 10 Risks
- Relevant certifications (helpful but not required): NIST AI RMF (AISIC), INFORMS CAP, AWS/Azure/Google ML Certifications.
- ...AI Evaluation Specialist Role Type: Contractor Location: Remote (US, CA, UK, IE, AU, NZ) Micro1 is engaging AI Evaluation Specialists to assess and elevate the quality of AI assistant outputs for an enterprise AI training initiative. In this role, you'll apply...SuggestedFor contractorsRemote work
- ...A leading AI organization in Australia is seeking individuals with strong writing and analytical skills to evaluate and improve AI outputs. The ideal candidate must possess the ability to assess emotional nuances and detail while adhering to structured guidelines. Responsibilities...SuggestedImmediate start
$120 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...hour Location: Remote Role Responsibilities Evaluate complex technical tasks using deep language expertise in Scala...SuggestedHourly payWeekly payFull timeContract workFor contractorsSummer workRemote work- ...Join a fast-paced AI evaluation initiative supporting one of the world's leading AI research organizations. We are seeking detail-oriented professionals to evaluate AI-generated outputs by applying structured grading rubrics with precision and consistency. This is...SuggestedTemporary workImmediate start
- ...Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red...SuggestedFull timeContract workFor contractorsRemote workFlexible hours
- ...About the Opportunity We are looking for experienced professionals with strong expertise in the video game industry to support AI evaluation initiatives. The role involves creating gaming-related evaluation scenarios, reviewing AI-generated responses, and providing...
- ...About the Opportunity We are looking for experienced entertainment professionals and subject matter experts to support AI evaluation initiatives focused on Movies and TV content. The role involves creating challenging evaluation scenarios, reviewing AI-generated responses...
$25 - $30 per hour
...Bilingual Simplified Chinese AI Evaluation Specialist is a remote Chinese specialist track for evaluating chinese evaluation outputs against native-speaker standards. Reviewers spot fluency, register, and cultural-context errors that automated checks miss, and write structured...For contractorsRemote work10 hours per week$60 per hour
...seeking contributors for a part-time QA project focused on autonomous AI agents. This flexible remote opportunity requires strong... ...familiarity with structured data formats. Candidates will review evaluation tasks, identify inconsistencies, and help define expected AI behaviors...Part timeRemote workFlexible hours$150k - $250k
...About Distyl AI Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect... ...What We Are Looking For At Distyl, we build AI systems using Evaluation-Driven Development —an approach where evaluation is not an afterthought...Full timeWork at officeFlexible hours3 days per week- ...experienced frontend and full-stack software developers across eligible global regions to support leading AI labs in training frontier models on frontend code evaluation. This role focuses on leveraging your web development expertise to evaluate and grade AI-generated...Temporary workFor contractorsRemote work
$20 - $80 per hour
...Role Overview Help improve next-generation AI systems by supplying precise, real-world evaluation, annotation, and feedback. This remote contractor role focuses on how AI models learn, reason, and perform across diverse subject areas. Key Responsibilities Evaluate...Hourly payContract workFor contractorsRemote work$36 - $72 per hour
...power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions. Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise. Role Details Location:...Hourly payFull timeMonday to FridayFlexible hours- ...Remote | Work from Home Employment Type: Project-based | Contract We are looking for detail-oriented Image Quality Evaluator for a multilingual AI data Annotation and Transcription Specialists with strong proficiency in English. In this role, you will support AI/ML...Contract workRemote workWork from homeMonday to FridayDay shift
- ...Summary This is a fully remote, hourly contractor role supporting AI data and language projects on a project-based, flexible hour... ..., and other content to support AI training datasets. LLM evaluation: reviewing AI-generated responses for accuracy, reasoning quality...Hourly payFor contractorsRemote workFlexible hours
- ...AI Data Scientist Fully Remote-United States Job Type Full-time Description Overview Tanaq Technical Services (TTS), a... ...quality assessments to support AI and analytics initiatives. Evaluate model performance and continuously refine algorithms to...Full timeContract workWork at officeLocal areaRemote work
- ...Job Description Our client is seeking an AI Data Scientist for a direct hire opportunity in North Phoenix, AZ or Hillsboro, OR. This... ...performance. Model Development & Continuous Improvement Evaluate model performance using real-world inspection data and...Remote work
- ...Senior AI/ML Data Scientist Key Required Skills: Strong knowledge of AI/ML/LLM, RAG LLM, Python, NLP, Generative AI.Position Description... ...Determine the nature of analytic problems, evaluate options, and offer recommendations for resolution....Contract work
$60k - $75k
...Research on Aging is seeking an exceptional, highly motivated AI Data Scientist / Agentic AI Engineer to join a collaborative research team... ..., hypothesis generation, and biological interpretation Evaluating the performance, limitations, and reliability of AI-enabled...Full timeVisa sponsorshipWork visa$130.7k - $205.2k
...AI Data Scientist About the Position The HP Enterprise AI & Machine Learning organization is a centralized team of data scientists and... ...innovation. Manages relationships with business partners to evaluate and foster data driven innovation, provides domain-specific...Full timeTemporary workLocal areaRelocationFlexible hoursShift work$219k - $246k
...technology company purpose-built to power AI-enabled precision health solutions that... .... Description As an AI Applied Scientist, you will occupy a unique, highly impactful... ...and accuracy. Develop automated evaluation pipelines and \"LLM-as-a-judge\" grading...Full timeRemote work$35 - $120 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Improve task and rubric quality through structured review. Evaluate the accuracy and depth of AI-generated content to strengthen reasoning...Full timeContract workSummer workRemote work$60 - $90 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation:...Hourly payWeekly payFull timeContract workFor contractorsSummer workRemote work- ...About the Role We're looking for a freelance AI agent developer or data scientist to help us build a suite of intelligent automation tools across... ...be a match for this role. Please note that a writing evaluation may be required as part of our application process....Contract workTemporary workFreelance
$111k - $151k
...looking for 6+ years of experience in a Data Scientist, Applied Scientist, Machine Learning... ...We value hands-on experience building and evaluating LLM-based or machine learning solutions,... ...rigorous evaluation frameworks and metrics for AI/ML systems, covering offline and online...Full timeWorldwide$126k - $234k
...AI/ML Scientist - Oncology Data Science TeamLocation: Cambridge or Basel Onsite Relocation is offered for this role. The Oncology Data... ...image-derived biomarkers and biological findings.Contribute to evaluation, validation, and benchmarking of image analysis algorithms...RelocationFlexible hours$229.9k - $262.4k
Senior Lead AI Engineer (SDK's: Gen AI Evaluation and MCP) Overview: At Capital One, we are creating responsible and reliable AI systems, changing... ...Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to...Full timePart timeLocal area$170k - $240k
...Xaira is an innovative biotech startup focused on leveraging AI to transform drug discovery and development. The company is leading... ...generation Build scalable systems for data processing, training, and evaluation Work in a research-driven environment with evolving problem...Flexible hours$126k - $234k
...oncology drug discovery, computational biology, AI/ML, and data engineering. We are seeking an enthusiastic AI/ML scientist with deep expertise in digital pathology and... ...and biological findings. Contribute to evaluation, validation, and benchmarking of image analysis...Local areaRelocationFlexible hours$86.32k - $154.96k
...ecosystem where experimental, clinical, and data scientists collaborate to accelerate discoveries, support scientific... ...is seeking a skilled and collaborative Agentic AI Scientist to lead the design, development, and evaluation of agentic AI systems that support research and...Work experience placementWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Scientist. Be the first to apply!



