AI Evaluation Scientist
Convergenz
We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements. This role is essential for establishing trust in AI solutions and supporting continuous improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment.
- Implement evaluation frameworks for AI models, including accuracy, robustness, relevance, bias, hallucination rate, and safety metrics.
- Build and maintain automated evaluation scripts, tests, and pipelines that assess AI model outputs and detect performance drift over time.
- Develop benchmark datasets, challenge sets, and scenario-based test cases tailored to mission and user needs.
- Perform structured error analysis and behavioral audits of LLMs, retrieval-augmented generation (RAG) systems, and predictive models, documenting findings and improvement recommendations.
- Collaborate with AI Developers, LLMOps Engineers, and Data Scientists to support iterative experimentation, model hardening, and quality improvements.
- Contribute to the design of human-in-the-loop evaluation workflows, integrating qualitative and quantitative insight into evaluation reports.
- Assist in mapping evaluation outcomes to responsible AI principles such as fairness, transparency, reliability, and safety.
- Partner with AI Governance Analysts to ensure evaluation outputs support compliance, documentation, and risk assessments.
- Stay current with emerging evaluation tools, frameworks, metrics, and research related to LLM assessment and generative AI reliability.
- Document evaluation processes, criteria, and results for both technical and non-technical audiences.
- Ability to hold a position of public trust with the U.S. government.
- Bachelor's degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field and 5+ years of experience.
- Master's degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field and 3+ years of experience.
- 2+ years of experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred).
- Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design for AI systems.
- Proficiency in Python and relevant libraries such as PyTorch, Hugging Face, scikit-learn, LangChain.
- Proficiency in AI evaluation frameworks such as Ragas or DeepEval.
- Proficiency in AI traceability/observability tools based on the OpenTelemetry protocol.
- Experience analyzing structured and unstructured data, including text, documents, and embeddings.
- Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures.
- Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness, transparency, reliability).
- Strong analytical skills with the ability to interpret model weaknesses, extract insights, and recommend actionable improvements.
- Excellent written and verbal communication skills, with the ability to present evaluation findings clearly to technical and non-technical stakeholders.
- Experience working in agile or iterative development environments is a plus.
- Familiarity with OWASP LLM Top 10 Risks
- Relevant certifications (helpful but not required): NIST AI RMF (AISIC), INFORMS CAP, AWS/Azure/Google ML Certifications.
$105k - $145k
OverviewWe are looking for an AI Evaluation Scientistto design and execute evaluation processes that ensure our predictive and generative... ...improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance...SuggestedLocal area- Scale AI, Inc. is looking for a Research Scientist focused on Frontier Risk Evaluations to develop evaluation measures and datasets for assessing AI risks. As a key team member, you will engage with government agencies and publish methodologies that influence AI safety...Suggested
- Studyfetch Beverly Hills, CA, is seeking an AI Research Scientist focused on learning and evaluation. You will build the evaluation layer under model training and learning science across Learn Engine and the Honen platform, driving metrics that matter for student outcomes...Suggested
- ...seeks a highly skilled researcher to ensure AI features perform across languages and cultures. You will drive multilingual evaluation development from design to implementation,... ...scalable methods with measurement scientists and ML researchers. You will publish novel...Suggested
- Apple is seeking a Research Scientist in AI Evaluation Science to advance evaluation methodology for LLMs, agentic systems, and human‑AI interaction. You will conduct original research, publish findings, and partner with platform engineers to productionize methods into...Suggested
- Johns Hopkins University is seeking a PREP Research Associate to advance agentic AI research within the NIST PREP program. You will design, develop, and evaluate AI agents, build experimental platforms, and contribute to benchmarks and methodologies for safe, trustworthy...
- Steampunk, Inc. seeks an AI Evaluation Scientist to design and implement evaluation frameworks ensuring accuracy, safety, and alignment with mission requirements for predictive and generative AI systems. You will collaborate with engineers, data scientists, governance analysts...
- ...Arena Intelligence Arena Intelligence is the open platform for evaluating how AI models perform in the real world. Created by researchers... ...Arena Intelligence is seeking a variety of Machine Learning Scientist to help advance how we evaluate and understand AI models. You...Permanent employmentWork at office
- Apple Inc. in Seattle, WA, seeks a Senior Applied Scientist to advance AI evaluation quality systems. You will build scalable ground-truth pipelines, calibration frameworks, and autonomous QA agents to ensure data powering AI/ML systems meets high accuracy, consistency...
$50 - $70 per hour
United States Digital Space LLC is seeking a Data Scientist for AI Evaluation Analytics to work as a remote contractor on long-term projects. You will develop evaluation metrics, validate datasets, analyze signals, and apply statistical methods to measure AI system performance...Remote jobHourly payFor contractors$105k - $145k
Steampunk is seeking an AI Evaluation Scientist in McLean, Virginia, to design evaluation processes for AI systems ensuring accuracy and reliability. Responsibilities include implementing evaluation frameworks and performing error analysis. Ideal candidates will have a...- Mercor is seeking advanced STEM Researchers to join a leading AI lab's research team. The role requires a PhD and involves guiding teams to tackle complex problems in STEM domains. You will evaluate solutions produced by AI agents, develop rigorous domain tasks, and collaborate...
- AI Research Scientist, Learning & Evaluation Studyfetch Beverly Hills, California, United States About this position About Studyfetch StudyFetch is the #1 AI-native learning platform globally, transforming how millions of students learn through personalized AI-powered...Work at officeWorldwide
$30 - $50 per hour
A tech company is seeking an AI Researcher to support end-to-end research for modern AI systems. This remote role involves designing experiments, defining evaluation protocols, and improving evaluation rigor for large language models. Key responsibilities include developing...Remote jobHourly pay- Rex.zone is seeking an AI Research Scientist to lead applied AI research projects for US-based customers, translating open-ended questions into measurable experiments in LLM evaluation and RLHF data design. You will evaluate prompts, design datasets, and work with cross...Remote jobHourly payFlexible hours
- ...Mercor is hiring for a STEM Computational Scientific Software & Evaluation Design position in Astrophysics & Cosmology. This role is remote... ...requires candidates to design computational problems and evaluate AI systems for research purposes. A graduate degree in a...Part timeRemote work
- ...To support AI development, the part-time AI Evaluation Specialist will review and assess AI-generated outputs for quality and usability while collaborating with teams to refine evaluation standards in a remote contract role. Key responsibilities Review and critically...Contract workPart timeRemote work
- ...candidate to conduct critical comparative analysis to advance our understanding of model capabilities. You will build and refine evaluation systems that create tight feedback loops between data, evals, and model behavior, and develop generalizable frameworks for reasoning...
- ...candidate for the position of STEM Computational Scientific Software & Evaluation Design in Astrophysics & Cosmology. This position is remote with... ...include designing computational problems, evaluating AI systems, and developing strategies for data recovery. Interested...Remote jobContract work
- ...Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red...Full timeContract workFor contractorsRemote workFlexible hours
$70 per hour
...clients Compensation: $70 per hour Join a cutting-edge AI research initiative focused on improving the quality, accuracy,... ...with exceptional critical thinking and communication skills to evaluate AI-generated responses across a variety of topics. In this role...Hourly payWeekly payContract workFor contractorsRemote workFlexible hours$20 per hour
A leading AI training company is seeking analytical, detail-oriented individuals to remotely teach AI chatbots. Responsibilities include developing prompts, writing high-quality responses, and evaluating different AI models. The role is ideal for those with experience in...Hourly payFull timePart timeRemote workFlexible hours$25 - $30 per hour
...independent contractors in the Town of Vermont, Wisconsin to help train AI chatbots. Ideal candidates are proactive individuals who can... ...include creating complex prompts, writing responses, and evaluating AI models. Candidates are required to have strong writing and research...Hourly payFor contractorsRemote workFlexible hours- A leading research accelerator is seeking a Geospatial Expert to enhance AI systems through advanced geospatial analysis. This entry-level, remote role involves evaluating geospatial datasets and supporting tasks aligned with crisis management and agriculture. Candidates...Contract workTemporary workRemote work
$20 per hour
A tech firm specializing in AI seeks analytical and detail-oriented individuals for remote work in training AI chatbots. The role involves... ...diverse conversation prompts, writing high-quality answers, and evaluating different AI models. Ideal candidates should have a bachelor's...For contractorsFreelanceRemote workFlexible hours$20 per hour
A leading AI training company is seeking remote workers to assist in training AI chatbots. Responsibilities include developing conversations, writing responses, and evaluating AI performance. The ideal candidates are analytical and detail-oriented with strong communication...Hourly payRemote workFlexible hours- ...A leading AI organization in Australia is seeking individuals with strong writing and analytical skills to evaluate and improve AI outputs. The ideal candidate must possess the ability to assess emotional nuances and detail while adhering to structured guidelines. Responsibilities...Immediate start
$25 - $30 per hour
...Bilingual Korean AI Evaluation Specialist is a remote Korean specialist track for evaluating korean evaluation outputs against native-speaker standards. Reviewers spot fluency, register, and cultural-context errors that automated checks miss, and write structured rationale...For contractorsRemote work10 hours per week- ...Join a fast-paced AI evaluation initiative supporting one of the world's leading AI research organizations. We are seeking detail-oriented professionals to evaluate AI-generated outputs by applying structured grading rubrics with precision and consistency. This is...Temporary workImmediate start
$20 per hour
A growing AI development company is seeking individuals for a remote position to help train AI chatbots. You will create diverse conversations, provide high-quality responses, and evaluate AI model performance. This role is perfect for those with experience in managing...Hourly payRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Scientist. Be the first to apply!


