Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Evaluation Scientist

Convergenz

We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements. This role is essential for establishing trust in AI solutions and supporting continuous improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment. Implement evaluation frameworks for AI models, including accuracy, robustness, relevance, bias, hallucination rate, and safety metrics. Build and maintain automated evaluation scripts, tests, and pipelines that assess AI model outputs and detect performance drift over time. Develop benchmark datasets, challenge sets, and scenario-based test cases tailored to mission and user needs. Perform structured error analysis and behavioral audits of LLMs, retrieval-augmented generation (RAG) systems, and predictive models, documenting findings and improvement recommendations. Collaborate with AI Developers, LLMOps Engineers, and Data Scientists to support iterative experimentation, model hardening, and quality improvements. Contribute to the design of human-in-the-loop evaluation workflows, integrating qualitative and quantitative insight into evaluation reports. Assist in mapping evaluation outcomes to responsible AI principles such as fairness, transparency, reliability, and safety. Partner with AI Governance Analysts to ensure evaluation outputs support compliance, documentation, and risk assessments. Stay current with emerging evaluation tools, frameworks, metrics, and research related to LLM assessment and generative AI reliability. Document evaluation processes, criteria, and results for both technical and non-technical audiences. Qualifications Ability to hold a position of public trust with the U.S. government. Bachelor's degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field and 5+ years of experience. Master's degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field and 3+ years of experience. 2+ years of experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred). Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design for AI systems. Proficiency in Python and relevant libraries such as PyTorch, Hugging Face, scikit-learn, LangChain. Proficiency in AI evaluation frameworks such as Ragas or DeepEval. Proficiency in AI traceability/observability tools based on the OpenTelemetry protocol. Experience analyzing structured and unstructured data, including text, documents, and embeddings. Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures. Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness, transparency, reliability). Strong analytical skills with the ability to interpret model weaknesses, extract insights, and recommend actionable improvements. Excellent written and verbal communication skills, with the ability to present evaluation findings clearly to technical and non-technical stakeholders. Experience working in agile or iterative development environments is a plus. Familiarity with OWASP LLM Top 10 Risks Relevant certifications (helpful but not required): NIST AI RMF (AISIC), INFORMS CAP, AWS/Azure/Google ML Certifications. Convergenz

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI Evaluation Scientist in New York, NY vacancy
  • $30 - $50 per hour

    A tech company is seeking an AI Researcher to support end-to-end research for modern AI systems. This remote role involves designing experiments, defining evaluation protocols, and improving evaluation rigor for large language models. Key responsibilities include developing... 
    Suggested
    Remote job
    Hourly pay

    REX

    New York, NY
    3 days ago
  • A leading AI organization in Australia is seeking individuals with strong writing and analytical skills to evaluate and improve AI outputs. The ideal candidate must possess the ability to assess emotional nuances and detail while adhering to structured guidelines. Responsibilities... 
    Suggested
    Immediate start

    Crossing Hurdles

    New York, NY
    3 days ago
  •  ...the United States. You will label and review multi-modal data for AI training, including text, images, audio, and video, and perform...  ...vision annotation with bounding boxes and segmentation, content safety labeling, and RLHF-style evaluation tasks. #J-18808-Ljbffr REX
    Suggested
    Remote job

    REX

    New York, NY
    5 days ago
  • Welo Data is hiring Data Labeling Associates in New York City for Project Perseus. The role involves evaluating Arabic language AI outputs and ensuring AI safety, requiring professional proficiency in Arabic and experience in writing and AI safety. You’ll critique models... 
    Suggested

    Welo Data

    New York, NY
    1 day ago
  • Obsidian is collaborating with AI labs to find experienced health insurance professionals to enhance AI systems related to coverage...  ...assess AI performance on health insurance tasks. The role includes evaluating AI outputs, creating health insurance scenarios, and providing... 
    Suggested

    Obsidian

    New York, NY
    4 days ago
  • $11 - $30.65 per hour

    Meridial is seeking contractors to evaluate advanced agentic audio models by simulating realistic customer service interactions across multiple domains. You will contribute to developing diverse datasets and assess model performance using various metrics. The role requires... 
    Remote job
    Hourly pay
    For contractors

    Meridial

    New York, NY
    5 days ago
  • $80 per hour

     ...collective human intelligence to ethically shape the future of AI. What We Do The Mindrift platform connects specialists with AI...  ...'re looking for someone who can design realistic and structured evaluation scenarios for LLM-based agents. You'll create test cases that simulate... 
    Part time
    Freelance
    Remote work
    Flexible hours

    Mindrift

    New York, NY
    3 days ago
  •  ...the RoleLexisNexis Legal & Professional is hiring an Applied AI Data Scientist to help shape the next generation of AI-powered legal...  ...with internal teams and enterprise stakeholders to design, evaluate, and continuously improve the AI capabilities that power legal... 
    Full time
    Remote work
    Worldwide

    RELX Group

    New York, NY
    7 days ago
  •  ...Job Description Job Description Job Title: AI Data Science Domain Expert Job Type: Contractor (Part-Time) Location: Remote...  ...advancing next-generation AI systems. In this role, you will review, evaluate, and refine AI-generated technical and analytical content to... 
    Part time
    For contractors
    Remote work

    YO AI Labs

    New York, NY
    28 days ago
  • $111k - $151k

     ...looking for 6+ years of experience in a Data Scientist, Applied Scientist, Machine Learning...  ...We value hands-on experience building and evaluating LLM-based or machine learning solutions,...  ...rigorous evaluation frameworks and metrics for AI/ML systems, covering offline and online... 
    Full time
    Worldwide

    LexisNexis

    New York, NY
    12 days ago
  • $180k - $240k

    Applied AI Data Scientist LA office / US remote AE Studio builds products, platforms, and internal ventures that increase agency for all conscious...  ..., etc.), vector databases, multimodality, fine-tuning, and evaluation. Proven data science experience: You've been a key... 
    Work at office
    Remote work
    Flexible hours

    AE Studio

    New York, NY
    1 day ago
  • YO AI Labs is seeking an experienced Management Consultant to support AI training and evaluation projects. This contractor role emphasizes structured problem-solving, business analysis, and consulting expertise to assess and improve AI model performance. Remote work enables... 
    Remote job
    For contractors

    YO AI Labs

    New York, NY
    4 days ago
  •  ...business intelligence, data engineering, and visualization.As AI Data Scientist Senior Associate within Payments Data and Analytics, you...  ...and reliability of our intelligent solutions through testing, evaluation, and monitoring. Your contributions will help streamline operations... 
    Work experience placement

    JP Morgan Chase

    New York, NY
    7 days ago
  • $61k - $101k

     ...experience working with Large Language Models (LLMs) and agentic AI approaches. We need the ability to collaborate effectively...  ...requirements with stakeholders. Preferred: familiarity with evaluation and monitoring concepts for GenAI solutions, including test cases... 
    Full time

    J.P. Morgan

    New York, NY
    16 days ago
  •  ...world's leading enterprises orchestrate AI-powered work. Our vision is to expand human...  ...in the world. As an AI research scientist, you'll be at the center of that work. You...  ...early hypothesis through model training, evaluation, and production deployment Design and... 
    Full time
    Work at office
    Local area
    Flexible hours

    Writer Corporation

    New York, NY
    2 days ago
  •  ...Job Title AI Research Scientist Location Hybrid / Remote Employment Type Full-time Job Summary We are seeking an AI...  ...reinforcement learning, or generative AI. Design, develop, and evaluate client AI models, algorithms, and architectures.... 
    Full time
    Remote work

    Ova Technologies

    New York, NY
    3 days ago
  • $153k - $376k

     ...translating designs into code, or iterating with AI. From idea to product, Figma empowers...  ..., join us!We’re looking for applied scientists with a Machine Learning and Artificial...  ...feedback into requirements for AI systemsBuild evaluation systems to measure and improve quality... 
    Minimum wage
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    Figma

    New York, NY
    a month ago
  • $175k - $250k

     ...AI ScientistAbout MillenniumMillennium is a global, diversified alternative investment firm, founded in 1989. Defined by evolution...  ...optimization, and agentic workflows.Conduct applied research to evaluate new AI and machine learning techniques for complex business problems... 
    Flexible hours

    Millennium Management

    New York, NY
    a month ago
  •  ...Labs in New York is seeking a Research Scientist focused on Safety Post-Training to develop...  ...techniques that make frontier AI systems safer and better understood by researchers...  ..., develop interpretability-informed evaluations, and collaborate with policymakers, engineers... 

    United States Digital Space LLC

    New York, NY
    2 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks and participate in a new benchmark for scientific computing. You will author original, executable research problems that frontier models cannot solve, focusing on materials science and related subdomains... 
    Part time
    Immediate start

    Obsidian

    New York, NY
    5 days ago
  • SupportFinity™ is seeking a Mathematical Modeling Specialist to train AI models and measure their progress. This role includes working on mathematics problems to evaluate AI chatbot outputs for quality and performance. The position is remote and offers flexibility in choosing... 
    Remote job
    Hourly pay

    SupportFinity™

    New York, NY
    5 days ago
  • Mercor is seeking experimental scientists and engineers across inorganic synthesis, characterization, superconductors...  ..., and semiconductors to support a frontier AI research lab. You will generate, structure, and evaluate scientific data to train models that reason about... 
    Remote job

    Obsidian

    New York, NY
    4 days ago
  • $320k - $400k

    As a Research Scientist on our team, you will partner with Research Engineers, working on fundamental...  ...products.Building on our track record of AI-powered solutions (e.g., Bits AI, Bits...  ...environments, RL training loops, and evaluation infrastructure needed to train agents... 

    Datadog

    New York, NY
    a month ago
  • $100k - $143k

     ...Role Overview Join Enterprise Intelligence as an AI Product Scientist supporting Decision Intelligence products that help field professionals...  ...design, causal analysis, and performance reporting to evaluate the effectiveness of intelligence products. Develop and optimize... 
    Local area
    3 days per week

    New York Life Insurance Co

    New York, NY
    more than 2 months ago
  •  ...Role You'll work alongside senior research scientists on problems at the frontier of LLM...  ...post-training methodology, and agentic AI — in one of the few environments where your...  ...time compute scaling, and systematic model evaluation — grounded in financial and crypto-... 
    Internship
    Work from home
    Shift work

    binance

    New York, NY
    1 day ago
  • $300k

     ...discovery experience. We are looking for a seasoned Machine Learning Scientist to design and develop innovative Machine Learning (ML)...  ...problems such as summarization quality and building LLM-based evaluators ("judges") that assess whether generated content meets quality... 
    Hourly pay
    Full time
    Immediate start
    Flexible hours

    Netflix

    New York, NY
    1 day ago
  • $204.44k - $324.99k

     ...APIs, and develop and test Prompt Builder templates and grounded AI experiences using Salesforce data, Data Cloud, knowledge, and...  ...augmented generation (RAG) patterns  Implement agent testing, evaluation, observability, and guardrails, including accuracy, hallucination... 
    Full time
    H1b
    Local area

    KPMG

    New York, NY
    11 days ago
  • Senior AI Researcher LOCATION: Full Remote. About: a global leader on a mission to create a world where cancer can't hide by providing...  ...for breast cancer detection, density assessment and risk evaluation. The ProFound Breast Health Suite is cleared by the U.S. Food &... 
    Remote work

    Syntricate Technologies

    New York, NY
    1 day ago
  • Working remotely in a full-time capacity, the AI Research Scientist will focus on developing agentic systems for cybersecurity by training models...  ...agentic planning and establish benchmarking criteria for evaluating systems Required qualifications Strong foundations in... 
    Full time
    Remote work

    Virtual Vocations Inc

    New York, NY
    5 days ago
  • $117.6k - $176.4k

    Data Scientist At Schneider Electric, we are committed to solving real-world problems to create...  ...and sustainability. Within our Global AI Hub we combine our long-standing manufacturing...  ...experiments (baselines, ablations, evaluation protocols) and communicate findings clearly... 
    Full time
    Temporary work

    Schneider Electric

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Evaluation Scientist. Be the first to apply!