Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Evaluation Scientist

Convergenz

We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements. This role is essential for establishing trust in AI solutions and supporting continuous improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment.

  • Implement evaluation frameworks for AI models, including accuracy, robustness, relevance, bias, hallucination rate, and safety metrics.
  • Build and maintain automated evaluation scripts, tests, and pipelines that assess AI model outputs and detect performance drift over time.
  • Develop benchmark datasets, challenge sets, and scenario-based test cases tailored to mission and user needs.
  • Perform structured error analysis and behavioral audits of LLMs, retrieval-augmented generation (RAG) systems, and predictive models, documenting findings and improvement recommendations.
  • Collaborate with AI Developers, LLMOps Engineers, and Data Scientists to support iterative experimentation, model hardening, and quality improvements.
  • Contribute to the design of human-in-the-loop evaluation workflows, integrating qualitative and quantitative insight into evaluation reports.
  • Assist in mapping evaluation outcomes to responsible AI principles such as fairness, transparency, reliability, and safety.
  • Partner with AI Governance Analysts to ensure evaluation outputs support compliance, documentation, and risk assessments.
  • Stay current with emerging evaluation tools, frameworks, metrics, and research related to LLM assessment and generative AI reliability.
  • Document evaluation processes, criteria, and results for both technical and non-technical audiences.
Qualifications
  • Ability to hold a position of public trust with the U.S. government.
  • Bachelor's degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field and 5+ years of experience.
    • Master's degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field and 3+ years of experience.
  • 2+ years of experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred).
  • Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design for AI systems.
  • Proficiency in Python and relevant libraries such as PyTorch, Hugging Face, scikit-learn, LangChain.
  • Proficiency in AI evaluation frameworks such as Ragas or DeepEval.
  • Proficiency in AI traceability/observability tools based on the OpenTelemetry protocol.
  • Experience analyzing structured and unstructured data, including text, documents, and embeddings.
  • Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures.
  • Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness, transparency, reliability).
  • Strong analytical skills with the ability to interpret model weaknesses, extract insights, and recommend actionable improvements.
  • Excellent written and verbal communication skills, with the ability to present evaluation findings clearly to technical and non-technical stakeholders.
  • Experience working in agile or iterative development environments is a plus.
  • Familiarity with OWASP LLM Top 10 Risks
  • Relevant certifications (helpful but not required): NIST AI RMF (AISIC), INFORMS CAP, AWS/Azure/Google ML Certifications.
Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the AI Evaluation Scientist in United States vacancy
  • $105k - $145k

    OverviewWe are looking for an AI Evaluation Scientistto design and execute evaluation processes that ensure our predictive and generative...  ...improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance... 
    Suggested
    Local area

    Steampunk

    McLean, VA
    1 day ago
  • Scale AI, Inc. is looking for a Research Scientist focused on Frontier Risk Evaluations to develop evaluation measures and datasets for assessing AI risks. As a key team member, you will engage with government agencies and publish methodologies that influence AI safety... 
    Suggested

    Scale AI, Inc.

    New York, NY
    4 days ago
  • Studyfetch Beverly Hills, CA, is seeking an AI Research Scientist focused on learning and evaluation. You will build the evaluation layer under model training and learning science across Learn Engine and the Honen platform, driving metrics that matter for student outcomes... 
    Suggested

    Studyfetch

    Beverly Hills, CA
    3 days ago
  •  ...seeks a highly skilled researcher to ensure AI features perform across languages and cultures. You will drive multilingual evaluation development from design to implementation,...  ...scalable methods with measurement scientists and ML researchers. You will publish novel... 
    Suggested

    Socket.dev

    Seattle, WA
    10 hours ago
  • Apple is seeking a Research Scientist in AI Evaluation Science to advance evaluation methodology for LLMs, agentic systems, and human‑AI interaction. You will conduct original research, publish findings, and partner with platform engineers to productionize methods into... 
    Suggested

    Apple Inc.

    Seattle, WA
    1 day ago
  • Johns Hopkins University is seeking a PREP Research Associate to advance agentic AI research within the NIST PREP program. You will design, develop, and evaluate AI agents, build experimental platforms, and contribute to benchmarks and methodologies for safe, trustworthy... 

    Johns Hopkins University

    Gaithersburg, MD
    1 day ago
  • Steampunk, Inc. seeks an AI Evaluation Scientist to design and implement evaluation frameworks ensuring accuracy, safety, and alignment with mission requirements for predictive and generative AI systems. You will collaborate with engineers, data scientists, governance analysts... 

    Steampunk.com

    Mc Lean, VA
    4 days ago
  •  ...Arena Intelligence Arena Intelligence is the open platform for evaluating how AI models perform in the real world. Created by researchers...  ...Arena Intelligence is seeking a variety of Machine Learning Scientist to help advance how we evaluate and understand AI models. You... 
    Permanent employment
    Work at office

    Arena

    San Francisco, CA
    4 days ago
  • Apple Inc. in Seattle, WA, seeks a Senior Applied Scientist to advance AI evaluation quality systems. You will build scalable ground-truth pipelines, calibration frameworks, and autonomous QA agents to ensure data powering AI/ML systems meets high accuracy, consistency... 

    Apple Inc.

    Seattle, WA
    1 day ago
  • $50 - $70 per hour

    United States Digital Space LLC is seeking a Data Scientist for AI Evaluation Analytics to work as a remote contractor on long-term projects. You will develop evaluation metrics, validate datasets, analyze signals, and apply statistical methods to measure AI system performance... 
    Remote job
    Hourly pay
    For contractors

    United States Digital Space LLC

    New York, NY
    2 days ago
  • $105k - $145k

    Steampunk is seeking an AI Evaluation Scientist in McLean, Virginia, to design evaluation processes for AI systems ensuring accuracy and reliability. Responsibilities include implementing evaluation frameworks and performing error analysis. Ideal candidates will have a... 

    Steampunk

    Mc Lean, VA
    1 day ago
  • Mercor is seeking advanced STEM Researchers to join a leading AI lab's research team. The role requires a PhD and involves guiding teams to tackle complex problems in STEM domains. You will evaluate solutions produced by AI agents, develop rigorous domain tasks, and collaborate... 

    AIToolboard

    Kannapolis, NC
    1 hour ago
  • AI Research Scientist, Learning & Evaluation Studyfetch Beverly Hills, California, United States About this position About Studyfetch StudyFetch is the #1 AI-native learning platform globally, transforming how millions of students learn through personalized AI-powered... 
    Work at office
    Worldwide

    Studyfetch

    Beverly Hills, CA
    3 days ago
  • $30 - $50 per hour

    A tech company is seeking an AI Researcher to support end-to-end research for modern AI systems. This remote role involves designing experiments, defining evaluation protocols, and improving evaluation rigor for large language models. Key responsibilities include developing... 
    Remote job
    Hourly pay

    Rex.zone

    New York, NY
    4 days ago
  • Rex.zone is seeking an AI Research Scientist to lead applied AI research projects for US-based customers, translating open-ended questions into measurable experiments in LLM evaluation and RLHF data design. You will evaluate prompts, design datasets, and work with cross... 
    Remote job
    Hourly pay
    Flexible hours

    AIToolboard

    New York, NY
    10 hours ago
  •  ...Mercor is hiring for a STEM Computational Scientific Software & Evaluation Design position in Astrophysics & Cosmology. This role is remote...  ...requires candidates to design computational problems and evaluate AI systems for research purposes. A graduate degree in a... 
    Part time
    Remote work

    Mercor

    New York, NY
    1 day ago
  •  ...To support AI development, the part-time AI Evaluation Specialist will review and assess AI-generated outputs for quality and usability while collaborating with teams to refine evaluation standards in a remote contract role. Key responsibilities Review and critically... 
    Contract work
    Part time
    Remote work

    Virtual Vocations Inc

    United States
    5 days ago
  •  ...candidate to conduct critical comparative analysis to advance our understanding of model capabilities. You will build and refine evaluation systems that create tight feedback loops between data, evals, and model behavior, and develop generalizable frameworks for reasoning... 

    Visa Hunt

    Brooklyn, NY
    2 days ago
  •  ...candidate for the position of STEM Computational Scientific Software & Evaluation Design in Astrophysics & Cosmology. This position is remote with...  ...include designing computational problems, evaluating AI systems, and developing strategies for data recovery. Interested... 
    Remote job
    Contract work

    Mercor

    New York, NY
    4 days ago
  •  ...Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red... 
    Full time
    Contract work
    For contractors
    Remote work
    Flexible hours

    Weekday

    Remote
    13 days ago
  • $70 per hour

     ...clients Compensation: $70 per hour Join a cutting-edge AI research initiative focused on improving the quality, accuracy,...  ...with exceptional critical thinking and communication skills to evaluate AI-generated responses across a variety of topics. In this role... 
    Hourly pay
    Weekly pay
    Contract work
    For contractors
    Remote work
    Flexible hours

    Weekday AI

    United States
    5 days ago
  • $20 per hour

    A leading AI training company is seeking analytical, detail-oriented individuals to remotely teach AI chatbots. Responsibilities include developing prompts, writing high-quality responses, and evaluating different AI models. The role is ideal for those with experience in... 
    Hourly pay
    Full time
    Part time
    Remote work
    Flexible hours

    DataAnnotation

    Providence, RI
    1 day ago
  • $25 - $30 per hour

     ...independent contractors in the Town of Vermont, Wisconsin to help train AI chatbots. Ideal candidates are proactive individuals who can...  ...include creating complex prompts, writing responses, and evaluating AI models. Candidates are required to have strong writing and research... 
    Hourly pay
    For contractors
    Remote work
    Flexible hours

    DataAnnotation

    Vermont
    1 day ago
  • A leading research accelerator is seeking a Geospatial Expert to enhance AI systems through advanced geospatial analysis. This entry-level, remote role involves evaluating geospatial datasets and supporting tasks aligned with crisis management and agriculture. Candidates... 
    Contract work
    Temporary work
    Remote work

    Turing

    Washington DC
    1 day ago
  • $20 per hour

    A tech firm specializing in AI seeks analytical and detail-oriented individuals for remote work in training AI chatbots. The role involves...  ...diverse conversation prompts, writing high-quality answers, and evaluating different AI models. Ideal candidates should have a bachelor's... 
    For contractors
    Freelance
    Remote work
    Flexible hours

    DataAnnotation

    Lansing, MI
    1 day ago
  • $20 per hour

    A leading AI training company is seeking remote workers to assist in training AI chatbots. Responsibilities include developing conversations, writing responses, and evaluating AI performance. The ideal candidates are analytical and detail-oriented with strong communication... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Madison, WI
    3 days ago
  •  ...A leading AI organization in Australia is seeking individuals with strong writing and analytical skills to evaluate and improve AI outputs. The ideal candidate must possess the ability to assess emotional nuances and detail while adhering to structured guidelines. Responsibilities... 
    Immediate start

    Crossing Hurdles

    New York, NY
    4 days ago
  • $25 - $30 per hour

     ...Bilingual Korean AI Evaluation Specialist is a remote Korean specialist track for evaluating korean evaluation outputs against native-speaker standards. Reviewers spot fluency, register, and cultural-context errors that automated checks miss, and write structured rationale... 
    For contractors
    Remote work
    10 hours per week

    AuraOne Human Data

    Remote
    21 days ago
  •  ...Join a fast-paced AI evaluation initiative supporting one of the world's leading AI research organizations. We are seeking detail-oriented professionals to evaluate AI-generated outputs by applying structured grading rubrics with precision and consistency. This is... 
    Temporary work
    Immediate start

    Weekday

    Remote
    19 days ago
  • $20 per hour

    A growing AI development company is seeking individuals for a remote position to help train AI chatbots. You will create diverse conversations, provide high-quality responses, and evaluate AI model performance. This role is perfect for those with experience in managing... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Columbia, SC
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Evaluation Scientist. Be the first to apply!