Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Model Evaluation Specialist

BAM VENTURES LLC

We are seeking an expert to evaluate and improve our AI models through comprehensive testing and analysis. You will be responsible for designing evaluation frameworks, conducting model assessments, and providing actionable insights for model improvement. Key Responsibilities Design and implement evaluation metrics for AI models Conduct thorough testing of model performance across different scenarios Analyze model outputs for bias, fairness, and accuracy Collaborate with ML engineers to implement improvements Document findings and recommendations Ideal Candidate PhD or Masters in Computer Science, ML, or related field 5+ years experience in AI/ML model evaluation Strong background in statistical analysis Experience with evaluation frameworks and metrics Excellent communication skills #J-18808-Ljbffr BAM VENTURES LLC

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the AI Model Evaluation Specialist in New York, NY vacancy
  • Cincinnatus LLC is recruiting for an Insurance Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting judgments, with placement at the client... 
    Suggested
    Weekday work

    Mercor

    New York, NY
    1 day ago
  • Xperteez Technology seeks a data science and ML-focused specialist to evaluate AI model outputs across statistics, ML, and quantitative problems. You will assess quality, identify errors, and provide actionable feedback to improve model capability. Responsibilities include... 
    Suggested

    Xperteez Technology

    New York, NY
    18 hours ago
  • SME Careers is seeking biologists to contribute to an AI training project that involves reviewing AI-generated responses and providing...  ...hold a MS or PhD in a relevant field and have experience in evaluating complex biology content. Strong communication skills and proficient... 
    Suggested
    Immediate start

    SME Careers

    New York, NY
    2 days ago
  • Obsidian is seeking Insurance domain SMEs to join a cutting-edge AI training program. You will evaluate AI model outputs against real underwriting practice and rubrics, guiding research and engineering teams to close knowledge gaps in underwriting, claims, and risk assessment... 
    Suggested

    Obsidian

    New York, NY
    3 days ago
  • Mercor is hiring Legal Experts to evaluate AI-generated responses for employment and labor law scenarios. This fully remote, hourly contract...  ...week. You will assess accuracy, provide feedback to improve model behavior and participate in calibration sessions. Requirements... 
    Suggested
    Remote job
    Hourly pay
    Contract work
    Flexible hours

    Intellerzone

    New York, NY
    18 hours ago
  •  ...PhD in Math or Physics to craft and review challenging math or physics problems. The role involves supporting training for large AI models while ensuring problem quality and relevance. Qualified candidates must possess a Master's or PhD from a top university, showcase... 
    Remote job
    Hourly pay
    Contract work

    Crossing Hurdles

    New York, NY
    3 days ago
  •  ...seeking a remote Kotlin Engineer to review AI-generated responses and create high-...  ...optimizing AI performance, and ensuring model accuracy. The ideal candidate has a Bachelor...  ...an expert network and requires critical evaluation of technical concepts. #J-18808-Ljbffr SME... 
    Remote job

    SME Careers

    New York, NY
    3 days ago
  • $400 per month

    About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic data engineering workflows... 

    Mercor Inc

    New York, NY
    18 hours ago
  • Turing is seeking a Software Engineering Evaluator to create cutting‑edge datasets for training...  ..., and advancing large language models, collaborating closely with researchers....  ...Java, Rust, and Go; evaluating and refining AI-generated code for efficiency, scalability... 
    Remote job

    Turing

    New York, NY
    4 days ago
  • Rise Data Labs is seeking advanced Mathematics and Statistics experts to support the training and evaluation of state-of-the-art AI systems. We need subject-matter experts who can apply deep quantitative knowledge to AI evaluation problems, assess AI-generated reasoning... 
    Remote job
    Contract work
    Immediate start
    Flexible hours

    BAM Ventures

    New York, NY
    4 days ago
  • Meridial is seeking a dedicated freelance Japanese Voice Actor to produce and evaluate training data for AI models. The role involves recording scripted material and assessing outputs for accuracy and naturalness while providing expert feedback. The ideal candidate will... 
    Remote job
    Hourly pay
    Freelance

    Meridial

    New York, NY
    18 hours ago
  • SME Careers is seeking chemists to join an AI training project. Your role will involve reviewing AI-generated responses, providing expert feedback, and ensuring accuracy in chemical content. This position will open access to future projects within our expert network. Ideal... 

    SME Careers

    New York, NY
    2 days ago
  • Mercor partners with a leading AI research lab to support a Frontier Code Agents project, focusing on realistic data engineering workflows and model evaluation. Contributors help evaluate and improve frontier AI coding models through structured technical assessments and... 

    Mercor Inc

    New York, NY
    18 hours ago
  • SME Careers is looking for a remote R Engineer to review AI-generated responses and create high-quality R and data-analysis content. The position requires a strong background in R programming, applied statistics, and excellent writing skills to document analyses. Candidates... 
    Remote job

    SME Careers

    New York, NY
    3 days ago
  • $75 - $150 per hour

     ...everyone’s full potential. Treliant is looking for Credit Risk Modelers for remote, project-based opportunities. Responsibilities...  ...credit decisioning and related consumer lending models. Rigorously evaluate predictive accuracy of model assumptions against actual... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Treliant (Acquired by Huron - 2025)

    New York, NY
    3 days ago
  • $130k - $160k

     ...PropositionAs Lead, Operating Model Design, you will serve as an internal...  ...for Global Technology’s AI-First operating model. You...  ...from delivery leaders. This is a specialist individual-contributor role: your...  ...incorporate field feedback, evaluate model effectiveness, and align... 
    Full time
    Temporary work
    Work experience placement
    Work at office
    Local area
    Relocation package
    3 days per week

    Metropolitan Life Insurance Company

    New York, NY
    7 days ago
  •  ...we?Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve...  ...Montreal, Seoul, Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    3 days ago
  • $228.7k - $343.1k

     ...financial crime at enormous scale, and one bad model can mean millions in credit losses,...  ...same scrutiny you apply to models applies to AI. We build the tooling that lets a lean team validate at scale, so you critically evaluate what it produces and own the evaluation... 
    Remote job
    Full time
    Local area
    Shift work

    Block

    New York, NY
    18 hours ago
  • $175k - $215k

     ...across 15+ U.S. states. The mission of the Waymo AI Foundations team is to develop machine learning...  ..., learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. In this hybrid role, you will report to a Senior... 
    Full time
    Remote work

    Waymo

    New York, NY
    18 hours ago
  • $118.98k - $195.47k

     ...seeking a motivated individual to join our team as a Finance Model & AI Solutions Lead.The colleague in this role will design, build,...  ...scenario generation capabilities that automatically create and evaluate multiple planning scenariosImplement AI-assisted data validation... 
    Full time
    H1b
    Work at office
    Visa sponsorship
    Work visa
    Flexible hours
    3 days per week

    Guardian Life Insurance

    New York, NY
    3 days ago
  •  ...for providing independent assurance and evaluating the company's risk management, governance...  ...completion.We are looking for data scientists and AI developers who will power our mission by...  ...in frameworks for auditing models, including criteria like robustness, fairness... 

    TikTok

    New York, NY
    5 days ago
  • $405k

     ...create reliable, interpretable, and steerable AI systems. We want AI to be safe and...  ...You'll architect the systems, tooling, and evaluation infrastructure that determine how quickly...  ...Architect eval frameworks that measure model capabilities across diverse coding tasks... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    18 hours ago
  • $60 - $90 per hour

     ...character animation expertise to shape how a performance transfer model is evaluated, ensuring actor timing, emotional intent, and subtle physical...  ...pool that scores model outputs for a leading generative AI research effort. Key Responsibilities Define evaluation... 
    Hourly pay
    Part time
    Freelance

    SaidGig

    New York, NY
    5 days ago
  •  ...Corporate Vice President, Model Validation & AI Governance About the Company A regulated insurer focused on model validation and responsible...  ...validation of AI solutions, particularly focusing on the evaluation of agentic and multi-step AI systems, and establishing... 

    Confidential

    New York, NY
    4 days ago
  • We’re seeking a future team member for the role of Specialist II, Program & Project Management (Model Risk Validation) to join our Model Risk Validation team....  ...investible assets. Every day, our teams harness cutting-edge AI and breakthrough technologies to collaborate with... 
    Worldwide
    Flexible hours

    The Bank of New York Mellon

    New York, NY
    3 days ago
  • The Browser Company is building a new era of AI-assisted web experiences. This hands-on role sits at the intersection of product and...  ...owning evals that measure quality and cost. You’ll work with the Model Behavior Lead, collaborate with engineering and product, and ship... 
    Remote work

    The Browser Company

    New York, NY
    1 day ago
  •  ...ll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and perform through high-quality, real-...  ...and problem-solving strategies in backend engineering. Evaluate real-world Node.js scenarios to help models learn... 
    Remote job
    Contract work

    YO IT Consulting

    New York, NY
    18 hours ago
  • Senior Research Scientist, Model Evaluation Cohere | Posted Mar 2 | Full-time | New York | Negotiable | Unknown Why this role? Evaluation is...  ...high data quality. You are obsessive about rigorously measuring AI capabilities, and also about making sure your measurements... 
    Full time
    Work at office
    Remote work
    Flexible hours

    SupportFinity™

    New York, NY
    3 days ago
  • Thermodynamic Model Developer Location: Parsippany, NJ Join Our Innovative...  ...and executed. We’re seeking a Specialist in Chemical Engineering...  ...Pro-II, OLI) Experience with AI/ML applications (i.e.:Hybrid,...  ...for all. Applicants will be evaluated through a structured, rubric-based... 

    OLI Systems, Inc.

    New York, NY
    2 days ago
  •  ...tech company in Canada is seeking individuals for a role focused on evaluating AI-generated responses. The responsibilities include assessing reasoning quality, providing structured feedback for model improvements, and ensuring outputs are accurate and human-aligned. Candidates... 
    Full time
    Contract work
    Part time
    Flexible hours

    Crossing Hurdles

    New York, NY
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Model Evaluation Specialist. Be the first to apply!