AI Evaluation Scientist
Convergenz
We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements. This role is essential for establishing trust in AI solutions and supporting continuous improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment. Implement evaluation frameworks for AI models, including accuracy, robustness, relevance, bias, hallucination rate, and safety metrics. Build and maintain automated evaluation scripts, tests, and pipelines that assess AI model outputs and detect performance drift over time. Develop benchmark datasets, challenge sets, and scenario-based test cases tailored to mission and user needs. Perform structured error analysis and behavioral audits of LLMs, retrieval-augmented generation (RAG) systems, and predictive models, documenting findings and improvement recommendations. Collaborate with AI Developers, LLMOps Engineers, and Data Scientists to support iterative experimentation, model hardening, and quality improvements. Contribute to the design of human-in-the-loop evaluation workflows, integrating qualitative and quantitative insight into evaluation reports. Assist in mapping evaluation outcomes to responsible AI principles such as fairness, transparency, reliability, and safety. Partner with AI Governance Analysts to ensure evaluation outputs support compliance, documentation, and risk assessments. Stay current with emerging evaluation tools, frameworks, metrics, and research related to LLM assessment and generative AI reliability. Document evaluation processes, criteria, and results for both technical and non-technical audiences. Qualifications Ability to hold a position of public trust with the U.S. government. Bachelor's degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field and 5+ years of experience. Master's degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field and 3+ years of experience. 2+ years of experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred). Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design for AI systems. Proficiency in Python and relevant libraries such as PyTorch, Hugging Face, scikit-learn, LangChain. Proficiency in AI evaluation frameworks such as Ragas or DeepEval. Proficiency in AI traceability/observability tools based on the OpenTelemetry protocol. Experience analyzing structured and unstructured data, including text, documents, and embeddings. Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures. Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness, transparency, reliability). Strong analytical skills with the ability to interpret model weaknesses, extract insights, and recommend actionable improvements. Excellent written and verbal communication skills, with the ability to present evaluation findings clearly to technical and non-technical stakeholders. Experience working in agile or iterative development environments is a plus. Familiarity with OWASP LLM Top 10 Risks Relevant certifications (helpful but not required): NIST AI RMF (AISIC), INFORMS CAP, AWS/Azure/Google ML Certifications. Convergenz
$30 - $50 per hour
A tech company is seeking an AI Researcher to support end-to-end research for modern AI systems. This remote role involves designing experiments, defining evaluation protocols, and improving evaluation rigor for large language models. Key responsibilities include developing...SuggestedRemote jobHourly pay- A leading AI organization in Australia is seeking individuals with strong writing and analytical skills to evaluate and improve AI outputs. The ideal candidate must possess the ability to assess emotional nuances and detail while adhering to structured guidelines. Responsibilities...SuggestedImmediate start
- ...the United States. You will label and review multi-modal data for AI training, including text, images, audio, and video, and perform... ...vision annotation with bounding boxes and segmentation, content safety labeling, and RLHF-style evaluation tasks. #J-18808-Ljbffr REXSuggestedRemote job
- Welo Data is hiring Data Labeling Associates in New York City for Project Perseus. The role involves evaluating Arabic language AI outputs and ensuring AI safety, requiring professional proficiency in Arabic and experience in writing and AI safety. You’ll critique models...Suggested
- Obsidian is collaborating with AI labs to find experienced health insurance professionals to enhance AI systems related to coverage... ...assess AI performance on health insurance tasks. The role includes evaluating AI outputs, creating health insurance scenarios, and providing...Suggested
$11 - $30.65 per hour
Meridial is seeking contractors to evaluate advanced agentic audio models by simulating realistic customer service interactions across multiple domains. You will contribute to developing diverse datasets and assess model performance using various metrics. The role requires...Remote jobHourly payFor contractors$80 per hour
...collective human intelligence to ethically shape the future of AI. What We Do The Mindrift platform connects specialists with AI... ...'re looking for someone who can design realistic and structured evaluation scenarios for LLM-based agents. You'll create test cases that simulate...Part timeFreelanceRemote workFlexible hours- ...the RoleLexisNexis Legal & Professional is hiring an Applied AI Data Scientist to help shape the next generation of AI-powered legal... ...with internal teams and enterprise stakeholders to design, evaluate, and continuously improve the AI capabilities that power legal...Full timeRemote workWorldwide
- ...Job Description Job Description Job Title: AI Data Science Domain Expert Job Type: Contractor (Part-Time) Location: Remote... ...advancing next-generation AI systems. In this role, you will review, evaluate, and refine AI-generated technical and analytical content to...Part timeFor contractorsRemote work
$111k - $151k
...looking for 6+ years of experience in a Data Scientist, Applied Scientist, Machine Learning... ...We value hands-on experience building and evaluating LLM-based or machine learning solutions,... ...rigorous evaluation frameworks and metrics for AI/ML systems, covering offline and online...Full timeWorldwide$180k - $240k
Applied AI Data Scientist LA office / US remote AE Studio builds products, platforms, and internal ventures that increase agency for all conscious... ..., etc.), vector databases, multimodality, fine-tuning, and evaluation. Proven data science experience: You've been a key...Work at officeRemote workFlexible hours- YO AI Labs is seeking an experienced Management Consultant to support AI training and evaluation projects. This contractor role emphasizes structured problem-solving, business analysis, and consulting expertise to assess and improve AI model performance. Remote work enables...Remote jobFor contractors
- ...business intelligence, data engineering, and visualization.As AI Data Scientist Senior Associate within Payments Data and Analytics, you... ...and reliability of our intelligent solutions through testing, evaluation, and monitoring. Your contributions will help streamline operations...Work experience placement
$61k - $101k
...experience working with Large Language Models (LLMs) and agentic AI approaches. We need the ability to collaborate effectively... ...requirements with stakeholders. Preferred: familiarity with evaluation and monitoring concepts for GenAI solutions, including test cases...Full time- ...world's leading enterprises orchestrate AI-powered work. Our vision is to expand human... ...in the world. As an AI research scientist, you'll be at the center of that work. You... ...early hypothesis through model training, evaluation, and production deployment Design and...Full timeWork at officeLocal areaFlexible hours
- ...Job Title AI Research Scientist Location Hybrid / Remote Employment Type Full-time Job Summary We are seeking an AI... ...reinforcement learning, or generative AI. Design, develop, and evaluate client AI models, algorithms, and architectures....Full timeRemote work
$153k - $376k
...translating designs into code, or iterating with AI. From idea to product, Figma empowers... ..., join us!We’re looking for applied scientists with a Machine Learning and Artificial... ...feedback into requirements for AI systemsBuild evaluation systems to measure and improve quality...Minimum wageFull timeTemporary workLocal areaRemote workFlexible hours$175k - $250k
...AI ScientistAbout MillenniumMillennium is a global, diversified alternative investment firm, founded in 1989. Defined by evolution... ...optimization, and agentic workflows.Conduct applied research to evaluate new AI and machine learning techniques for complex business problems...Flexible hours- ...Labs in New York is seeking a Research Scientist focused on Safety Post-Training to develop... ...techniques that make frontier AI systems safer and better understood by researchers... ..., develop interpretability-informed evaluations, and collaborate with policymakers, engineers...
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks and participate in a new benchmark for scientific computing. You will author original, executable research problems that frontier models cannot solve, focusing on materials science and related subdomains...Part timeImmediate start
- SupportFinity™ is seeking a Mathematical Modeling Specialist to train AI models and measure their progress. This role includes working on mathematics problems to evaluate AI chatbot outputs for quality and performance. The position is remote and offers flexibility in choosing...Remote jobHourly pay
- Mercor is seeking experimental scientists and engineers across inorganic synthesis, characterization, superconductors... ..., and semiconductors to support a frontier AI research lab. You will generate, structure, and evaluate scientific data to train models that reason about...Remote job
$320k - $400k
As a Research Scientist on our team, you will partner with Research Engineers, working on fundamental... ...products.Building on our track record of AI-powered solutions (e.g., Bits AI, Bits... ...environments, RL training loops, and evaluation infrastructure needed to train agents...$100k - $143k
...Role Overview Join Enterprise Intelligence as an AI Product Scientist supporting Decision Intelligence products that help field professionals... ...design, causal analysis, and performance reporting to evaluate the effectiveness of intelligence products. Develop and optimize...Local area3 days per week- ...Role You'll work alongside senior research scientists on problems at the frontier of LLM... ...post-training methodology, and agentic AI — in one of the few environments where your... ...time compute scaling, and systematic model evaluation — grounded in financial and crypto-...InternshipWork from homeShift work
$300k
...discovery experience. We are looking for a seasoned Machine Learning Scientist to design and develop innovative Machine Learning (ML)... ...problems such as summarization quality and building LLM-based evaluators ("judges") that assess whether generated content meets quality...Hourly payFull timeImmediate startFlexible hours$204.44k - $324.99k
...APIs, and develop and test Prompt Builder templates and grounded AI experiences using Salesforce data, Data Cloud, knowledge, and... ...augmented generation (RAG) patterns Implement agent testing, evaluation, observability, and guardrails, including accuracy, hallucination...Full timeH1bLocal area- Senior AI Researcher LOCATION: Full Remote. About: a global leader on a mission to create a world where cancer can't hide by providing... ...for breast cancer detection, density assessment and risk evaluation. The ProFound Breast Health Suite is cleared by the U.S. Food &...Remote work
- Working remotely in a full-time capacity, the AI Research Scientist will focus on developing agentic systems for cybersecurity by training models... ...agentic planning and establish benchmarking criteria for evaluating systems Required qualifications Strong foundations in...Full timeRemote work
$117.6k - $176.4k
Data Scientist At Schneider Electric, we are committed to solving real-world problems to create... ...and sustainability. Within our Global AI Hub we combine our long-standing manufacturing... ...experiments (baselines, ablations, evaluation protocols) and communicate findings clearly...Full timeTemporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Scientist. Be the first to apply!



