Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Evaluation Engineer

Ova Technologies

Job TitleAI Evaluation EngineerLocationHybrid / RemoteEmployment TypeFull-timeJob SummaryWe are seeking an AI Evaluation Engineer to design, implement, and maintain evaluation frameworks for AI and machine learning systems, with a focus on Large Language Models (LLMs) and generative AI applications. The ideal candidate will develop robust evaluation methodologies, assess model performance across quality and safety dimensions, and collaborate with AI engineers, data scientists, and product teams to improve AI system reliability and user experience.Key ResponsibilitiesDesign and implement evaluation frameworks for AI, machine learning, and generative AI systems.Develop automated and manual evaluation pipelines to assess model quality and performance.Define evaluation metrics for accuracy, relevance, factuality, consistency, completeness, latency, and user satisfaction.Create benchmark datasets, test suites, and evaluation scenarios for AI models.Evaluate LLMs and AI applications for hallucinations, bias, toxicity, fairness, robustness, and safety.Measure Retrieval-Augmented Generation (RAG) performance, including retrieval quality and response grounding.Conduct A/B testing and comparative evaluations of models, prompts, and AI workflows.Analyze evaluation results and provide actionable recommendations for model improvement.Collaborate with AI engineers, data scientists, product managers, and QA teams throughout the AI development lifecycle.Monitor production AI systems and identify performance degradation, data drift, and model drift.Document evaluation methodologies, findings, and best practices.Ensure compliance with organizational AI governance, privacy, security, and responsible AI policies.Required QualificationsBachelor's degree in Computer Science, Artificial Intelligence, Data Science, Statistics, Mathematics, Software Engineering, or a related field.3–5+ years of experience in AI, machine learning, software engineering, data science, or model evaluation.Strong understanding of machine learning concepts and evaluation methodologies.Experience evaluating AI or generative AI applications.Proficiency in Python and SQL.Experience with REST APIs and cloud-based applications.Preferred QualificationsMaster's degree in AI, Machine Learning, Data Science, Computer Science, or a related field.Experience evaluating Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems.Knowledge of Responsible AI, AI safety, and AI governance principles.Experience with MLOps and model lifecycle management.AI or cloud certifications (AWS, Azure, Google Cloud).Technical SkillsPythonSQLMachine learning fundamentalsLLM evaluation methodologiesPrompt engineeringRAG evaluationAI benchmarking techniquesStatistical analysisData visualization (Power BI, Tableau, Matplotlib)Git and CI/CDREST APIsJSONAI evaluation frameworks (DeepEval, Ragas, LangSmith, Promptfoo)MLflowDocker and KubernetesCloud platforms (AWS, Azure, Google Cloud)Soft SkillsAnalytical thinkingCritical thinkingProblem-solvingStrong communication skillsTechnical documentationCollaborationAttention to detailTime managementContinuous learningPreferred ExperienceGenerative AI applicationsConversational AI and chatbotsEnterprise AI platformsRetrieval-Augmented Generation (RAG)Machine learning model validationAI-powered SaaS applicationsRegulated industries such as healthcare, finance, or insuranceSuccess MetricsAI evaluation coverageBenchmark quality and completenessImprovement in model accuracy and reliabilityReduction in hallucinations and unsafe responsesEvaluation pipeline automationDetection of model regressions before releaseProduction AI performance and stabilityStakeholder satisfactionCompliance with Responsible AI and governance standardsNice-to-Have SkillsExplainable AI (XAI)Model monitoring and observability toolsSynthetic data generationVector databasesLangChain or similar AI orchestration frameworksData annotation toolsExperiment tracking platformsA/B testing methodologiesAI red teamingFamiliarity with standards such as NIST AI Risk Management Framework (AI RMF) or ISO/IEC 42001

Vacancy posted 2 hours ago
Similar jobs that could be interesting for youBased on the AI Evaluation Engineer in New York, NY vacancy
  • $161.6k - $200k

     ...AI Evaluation EngineerLocations: Charlotte, North Carolina, United States; Denver, Colorado, United States; New York, New York, United...  ...we deserve. To learn more, visit SummaryAs an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics... 
    Suggested
    Local area
    Flexible hours

    Judi Health

    New York, NY
    20 hours ago
  • $229.9k - $262.4k

     ...Senior Lead AI Engineer (SDK's: Gen AI Evaluation and MCP) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    3 days ago
  •  ...Member of Technical Staff - Evals located in New York, NY, where you will ensure our AI-powered features are high quality and reliable. Your responsibilities include designing evaluation frameworks, building automated tests, and developing tools for seamless evaluations.... 
    Suggested

    Entendre

    New York, NY
    1 day ago
  • $152k - $241.5k

     ...believe open-weight models are foundational to American AI leadership and cybersecurity, and that trust in AI grows...  ...and broad scientific scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find, validate, and patch software... 
    Suggested
    Full time
    Remote work

    Nvidia

    New York, NY
    1 day ago
  • Mercor seeks experienced music producers and audio engineers to evaluate generative music AI models, working in Korean and English to assess AI-generated tracks against detailed quality standards. Responsibilities include comparing songs for musicality, creativity, adherence... 
    Suggested
    Contract work
    Immediate start
    Flexible hours

    Obsidian

    New York, NY
    4 days ago
  • Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Italian... 

    Mercor

    New York, NY
    2 days ago
  • Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models. You will assess AI-generated music across genres and rate it against detailed quality standards, working in Hindi and English. Responsibilities include head-to-head... 
    Remote work

    Obsidian

    New York, NY
    4 days ago
  • Mercor is seeking experienced music producers and audio engineers to evaluate generative music AI models in collaboration with a leading AI lab. You will assess AI-generated music in German and English across a wide range of genres and rate it against detailed quality standards... 

    Mercor

    New York, NY
    2 days ago
  • $400 per month

    Obsidian is seeking contributors for a Frontier Code Agents project, focused on evaluating AI coding models in fraud and risk engineering. Candidates will use AI coding tools to handle complex tasks and provide technical assessments. The role requires 2+ years of experience... 

    Obsidian

    New York, NY
    3 days ago
  •  ...Job description Rengo AI is building the intelligence layer for fund management —...  ...strategies . The Role As a Founding AI Engineer , you will build the core system that...  ...in actual portfolio data ~ Build evaluation frameworks for correctness of financial narratives... 
    Full time
    Shift work

    Decircle

    New York, NY
    20 hours ago
  • $135k - $200k

     ...missing children, and more. The Role Forward Deployed AI Engineers work directly with customers owning Gen AI strategy and...  ...landscape. Strong foundation in Machine Learning basics (Evaluation, Training, Problem Decomposition). Strong engineering background... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package
    Flexible hours

    Palantir Technologies

    New York, NY
    20 hours ago
  • $160k - $230k

     ...Standard Template Labs is a stealth-mode, AI-native startup reimagining the future of...  ...Role Our team is currently seeking AI Engineers that will design, build, and deploy machine...  ...that differentiate our platform. Evaluate performance, optimize models, and ensure... 
    Full time
    Work at office
    Local area

    Standard Template Labs

    New York, NY
    20 hours ago
  • $149.6k - $184k

     ...clinical operations, and internal tooling. We're hiring an AI Engineer to join a new team we're building from the ground up...  ...your work will involve building orchestration layers, prompts, evaluations, retrieval systems, and tooling that make LLMs reliable enough... 
    Full time
    Internship
    Local area

    Weight Loss, Better Sex, Fuller Hair, Improved Skin And More...

    New York, NY
    20 hours ago
  • $155k - $200k

     ...the disciplined collaboration and transcendent thinking as an AI Engineer at Capstone Investment Advisors here. Responsibilities and...  ...applications using Python Engage with domain experts to identify, evaluate, and execute high-value AI use cases Our future colleague... 
    Minimum wage
    Full time

    Capstone Investment Advisors

    New York, NY
    20 hours ago
  • $200k - $300k

    About Farsight Farsight is the agentic AI platform for financial services, currently...  ...Ventures, supercharged by scalable engineering and AI skills from companies including Amazon...  ...to polished output. Build the evaluation and quality systems behind generated deliverables... 
    Full time
    Local area
    Remote work

    Farsight Ai

    New York, NY
    20 hours ago
  •  ...about sciemo Sciemo builds AI for consumer goods: technology that helps businesses...  ...our products. We are looking for an AI Engineer to build and deploy the generative and agent...  ...prompt and architecture design through evaluation, deployment, monitoring, and iteration in... 
    Full time
    Work at office
    Remote work

    Sciemo

    New York, NY
    20 hours ago
  • $180k - $240k

    Forward Deployed Software Engineer — US / NYC About Indicium AI Indicium AI was Anthropic’s first European launch partner, a Preferred Anthropic Partner...  ..., real-time context retrieval systems, and LLM evaluation architectures. This is a high-agency "builder-consultant... 
    Full time
    Flexible hours

    Indicium Ai

    New York, NY
    20 hours ago
  • $175k - $275k

     ...changing that. We're building the agentic AI platform designed exclusively for...  ...Role Overview We're looking for an AI Engineer to help build the agentic platform at the...  ...that keep outputs reliable Fine-tune and evaluate models — fine-tune in-house and open-source... 
    Full time
    Work at office

    Translucent Inc

    New York, NY
    20 hours ago
  • $107.25k - $214.5k

     ...building the next layer of that platform: AI that can reason across everything we know...  .... We're looking for an Agentic AI Engineer who has already shipped production AI systems...  ...outputs. Develop systems that evaluate confidence, uncertainty and consequence before... 
    Full time

    Catapult Sports

    New York, NY
    20 hours ago
  •  ...users engage — and sometimes abuse. Our AI agents, integrated workflow platform, and...  ...-starters. Cinder is seeking an AI/ML engineer to architect and deploy production AI systems...  ...and fine-tuning through rigorous evaluation and production deployment. Develop automated... 
    Full time
    Work at office
    Immediate start

    Cinder

    New York, NY
    20 hours ago
  •  ...one of the hardest problems in enterprise AI: AI models are generic but company...  ...London. We're a deeply technical team of engineers, AI researchers, and strategists with a high...  ...agent systems in production, or designed evaluation frameworks for enterprise deployments, we... 
    Full time

    Edra

    New York, NY
    20 hours ago
  •  ...Founding AI Engineer (LLM + Production) Location: New York, NY (In‑Person) Compensation: Top of Market + Equity + Benefits About...  ...production‑grade LLMs and AI agents. You’ll bring rigor to evaluations, fine‑tuning, benchmarking, and synthetic data pipelines — and... 
    Full time
    Weekend work

    Everstar

    New York, NY
    20 hours ago
  • $90k - $95k

     .... The Role Are you someone who builds AI agents and thinks about how to automate the...  ...you. We're looking for a full-time AI Engineer to help shape and execute Fanatics' AI strategy...  ...and best practices, including tool evaluation, security and InfoSec coordination, and tracking... 
    Full time
    Temporary work
    Work at office

    Fanatics Inc

    New York, NY
    20 hours ago
  • $150k - $275k

     ...cracks. Joyful Health is building the AI-powered financial operating system for healthcare...  ...talented and ambitious machine learning engineer with 5+ years of experience building...  ..., and payer policy information Create evaluation frameworks and feedback loops to... 
    Full time
    Private practice
    Work at office
    3 days per week

    Joyful Health

    New York, NY
    20 hours ago
  •  ...We're looking for an experienced AI Engineer to join Optiver's Applied AI and Platform Engineering team. In this role, you' ll design...  ...agent platforms, AI assistants, code review harnesses, evaluation frameworks, and agentic research pipelines—using Python and modern... 
    Full time
    Work at office

    Optiver

    New York, NY
    20 hours ago
  •  ...thinking organization, apply now.We are currently seeking a Gen AI Engineer to join our team in a hybrid basis either out of our NYC or...  ..., and least-privilege access.· Productionize LLMs: Build evaluation framework for open-source and foundational LLMs; implement retrieval... 
    Work at office
    Remote work
    Flexible hours
    3 days per week

    NTT DATA

    New York, NY
    2 days ago
  •  ...FinTechSelling Points Contribute to innovative AI systems for enterprise finance...  ...development, focusing on Python and production engineering tasks.Support customer implementations...  ....Develop and maintain model pipelines, evaluation sets, and monitoring tools.Analyze... 

    Green Key Resources

    New York, NY
    1 day ago
  • $175k - $250k

    AI EngineerAbout MillenniumMillennium is a global, diversified alternative investment...  ...incentive workflowsDesign and implement evaluation frameworks to measure solution quality, robustness...  ...closely with product managers, data engineers, and business stakeholders to translate... 
    Flexible hours

    Millennium Management

    New York, NY
    4 days ago
  • $133k - $166k

     ...Description & ResponsibilitiesPersistent Systems is looking for an AI Engineer to design, build, and deploy secure AI applications that help...  ...safety, security, governance, and monitoring practicesCreate evaluation frameworks to measure AI quality, accuracy, latency, cost,... 
    Flexible hours

    Persistent Systems

    New York, NY
    3 days ago
  • $161.8k - $184.6k

    Senior AI Engineer Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking...  ...model inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage... 
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    New York, NY
    20 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Evaluation Engineer. Be the first to apply!