AI Evaluation Engineer
Ova Technologies
Job TitleAI Evaluation EngineerLocationHybrid / RemoteEmployment TypeFull-timeJob SummaryWe are seeking an AI Evaluation Engineer to design, implement, and maintain evaluation frameworks for AI and machine learning systems, with a focus on Large Language Models (LLMs) and generative AI applications. The ideal candidate will develop robust evaluation methodologies, assess model performance across quality and safety dimensions, and collaborate with AI engineers, data scientists, and product teams to improve AI system reliability and user experience.Key ResponsibilitiesDesign and implement evaluation frameworks for AI, machine learning, and generative AI systems.Develop automated and manual evaluation pipelines to assess model quality and performance.Define evaluation metrics for accuracy, relevance, factuality, consistency, completeness, latency, and user satisfaction.Create benchmark datasets, test suites, and evaluation scenarios for AI models.Evaluate LLMs and AI applications for hallucinations, bias, toxicity, fairness, robustness, and safety.Measure Retrieval-Augmented Generation (RAG) performance, including retrieval quality and response grounding.Conduct A/B testing and comparative evaluations of models, prompts, and AI workflows.Analyze evaluation results and provide actionable recommendations for model improvement.Collaborate with AI engineers, data scientists, product managers, and QA teams throughout the AI development lifecycle.Monitor production AI systems and identify performance degradation, data drift, and model drift.Document evaluation methodologies, findings, and best practices.Ensure compliance with organizational AI governance, privacy, security, and responsible AI policies.Required QualificationsBachelor's degree in Computer Science, Artificial Intelligence, Data Science, Statistics, Mathematics, Software Engineering, or a related field.3–5+ years of experience in AI, machine learning, software engineering, data science, or model evaluation.Strong understanding of machine learning concepts and evaluation methodologies.Experience evaluating AI or generative AI applications.Proficiency in Python and SQL.Experience with REST APIs and cloud-based applications.Preferred QualificationsMaster's degree in AI, Machine Learning, Data Science, Computer Science, or a related field.Experience evaluating Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems.Knowledge of Responsible AI, AI safety, and AI governance principles.Experience with MLOps and model lifecycle management.AI or cloud certifications (AWS, Azure, Google Cloud).Technical SkillsPythonSQLMachine learning fundamentalsLLM evaluation methodologiesPrompt engineeringRAG evaluationAI benchmarking techniquesStatistical analysisData visualization (Power BI, Tableau, Matplotlib)Git and CI/CDREST APIsJSONAI evaluation frameworks (DeepEval, Ragas, LangSmith, Promptfoo)MLflowDocker and KubernetesCloud platforms (AWS, Azure, Google Cloud)Soft SkillsAnalytical thinkingCritical thinkingProblem-solvingStrong communication skillsTechnical documentationCollaborationAttention to detailTime managementContinuous learningPreferred ExperienceGenerative AI applicationsConversational AI and chatbotsEnterprise AI platformsRetrieval-Augmented Generation (RAG)Machine learning model validationAI-powered SaaS applicationsRegulated industries such as healthcare, finance, or insuranceSuccess MetricsAI evaluation coverageBenchmark quality and completenessImprovement in model accuracy and reliabilityReduction in hallucinations and unsafe responsesEvaluation pipeline automationDetection of model regressions before releaseProduction AI performance and stabilityStakeholder satisfactionCompliance with Responsible AI and governance standardsNice-to-Have SkillsExplainable AI (XAI)Model monitoring and observability toolsSynthetic data generationVector databasesLangChain or similar AI orchestration frameworksData annotation toolsExperiment tracking platformsA/B testing methodologiesAI red teamingFamiliarity with standards such as NIST AI Risk Management Framework (AI RMF) or ISO/IEC 42001
$161.6k - $200k
...AI Evaluation EngineerLocations: Charlotte, North Carolina, United States; Denver, Colorado, United States; New York, New York, United... ...we deserve. To learn more, visit SummaryAs an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics...SuggestedLocal areaFlexible hours$229.9k - $262.4k
...Senior Lead AI Engineer (SDK's: Gen AI Evaluation and MCP) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized...SuggestedFull timePart timeLocal area- ...Member of Technical Staff - Evals located in New York, NY, where you will ensure our AI-powered features are high quality and reliable. Your responsibilities include designing evaluation frameworks, building automated tests, and developing tools for seamless evaluations....Suggested
$152k - $241.5k
...believe open-weight models are foundational to American AI leadership and cybersecurity, and that trust in AI grows... ...and broad scientific scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find, validate, and patch software...SuggestedFull timeRemote work- Mercor seeks experienced music producers and audio engineers to evaluate generative music AI models, working in Korean and English to assess AI-generated tracks against detailed quality standards. Responsibilities include comparing songs for musicality, creativity, adherence...SuggestedContract workImmediate startFlexible hours
- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. You will assess AI-generated music across a wide range of genres and rate it against detailed quality standards, working in Italian...
- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models. You will assess AI-generated music across genres and rate it against detailed quality standards, working in Hindi and English. Responsibilities include head-to-head...Remote work
- Mercor is seeking experienced music producers and audio engineers to evaluate generative music AI models in collaboration with a leading AI lab. You will assess AI-generated music in German and English across a wide range of genres and rate it against detailed quality standards...
$400 per month
Obsidian is seeking contributors for a Frontier Code Agents project, focused on evaluating AI coding models in fraud and risk engineering. Candidates will use AI coding tools to handle complex tasks and provide technical assessments. The role requires 2+ years of experience...- ...Job description Rengo AI is building the intelligence layer for fund management —... ...strategies . The Role As a Founding AI Engineer , you will build the core system that... ...in actual portfolio data ~ Build evaluation frameworks for correctness of financial narratives...Full timeShift work
$135k - $200k
...missing children, and more. The Role Forward Deployed AI Engineers work directly with customers owning Gen AI strategy and... ...landscape. Strong foundation in Machine Learning basics (Evaluation, Training, Problem Decomposition). Strong engineering background...Full timeWork experience placementWork at officeRemote workWork from homeRelocation packageFlexible hours$160k - $230k
...Standard Template Labs is a stealth-mode, AI-native startup reimagining the future of... ...Role Our team is currently seeking AI Engineers that will design, build, and deploy machine... ...that differentiate our platform. Evaluate performance, optimize models, and ensure...Full timeWork at officeLocal area$149.6k - $184k
...clinical operations, and internal tooling. We're hiring an AI Engineer to join a new team we're building from the ground up... ...your work will involve building orchestration layers, prompts, evaluations, retrieval systems, and tooling that make LLMs reliable enough...Full timeInternshipLocal area$155k - $200k
...the disciplined collaboration and transcendent thinking as an AI Engineer at Capstone Investment Advisors here. Responsibilities and... ...applications using Python Engage with domain experts to identify, evaluate, and execute high-value AI use cases Our future colleague...Minimum wageFull time$200k - $300k
About Farsight Farsight is the agentic AI platform for financial services, currently... ...Ventures, supercharged by scalable engineering and AI skills from companies including Amazon... ...to polished output. Build the evaluation and quality systems behind generated deliverables...Full timeLocal areaRemote work- ...about sciemo Sciemo builds AI for consumer goods: technology that helps businesses... ...our products. We are looking for an AI Engineer to build and deploy the generative and agent... ...prompt and architecture design through evaluation, deployment, monitoring, and iteration in...Full timeWork at officeRemote work
$180k - $240k
Forward Deployed Software Engineer — US / NYC About Indicium AI Indicium AI was Anthropic’s first European launch partner, a Preferred Anthropic Partner... ..., real-time context retrieval systems, and LLM evaluation architectures. This is a high-agency "builder-consultant...Full timeFlexible hours$175k - $275k
...changing that. We're building the agentic AI platform designed exclusively for... ...Role Overview We're looking for an AI Engineer to help build the agentic platform at the... ...that keep outputs reliable Fine-tune and evaluate models — fine-tune in-house and open-source...Full timeWork at office$107.25k - $214.5k
...building the next layer of that platform: AI that can reason across everything we know... .... We're looking for an Agentic AI Engineer who has already shipped production AI systems... ...outputs. Develop systems that evaluate confidence, uncertainty and consequence before...Full time- ...users engage — and sometimes abuse. Our AI agents, integrated workflow platform, and... ...-starters. Cinder is seeking an AI/ML engineer to architect and deploy production AI systems... ...and fine-tuning through rigorous evaluation and production deployment. Develop automated...Full timeWork at officeImmediate start
- ...one of the hardest problems in enterprise AI: AI models are generic but company... ...London. We're a deeply technical team of engineers, AI researchers, and strategists with a high... ...agent systems in production, or designed evaluation frameworks for enterprise deployments, we...Full time
- ...Founding AI Engineer (LLM + Production) Location: New York, NY (In‑Person) Compensation: Top of Market + Equity + Benefits About... ...production‑grade LLMs and AI agents. You’ll bring rigor to evaluations, fine‑tuning, benchmarking, and synthetic data pipelines — and...Full timeWeekend work
$90k - $95k
.... The Role Are you someone who builds AI agents and thinks about how to automate the... ...you. We're looking for a full-time AI Engineer to help shape and execute Fanatics' AI strategy... ...and best practices, including tool evaluation, security and InfoSec coordination, and tracking...Full timeTemporary workWork at office$150k - $275k
...cracks. Joyful Health is building the AI-powered financial operating system for healthcare... ...talented and ambitious machine learning engineer with 5+ years of experience building... ..., and payer policy information Create evaluation frameworks and feedback loops to...Full timePrivate practiceWork at office3 days per week- ...We're looking for an experienced AI Engineer to join Optiver's Applied AI and Platform Engineering team. In this role, you' ll design... ...agent platforms, AI assistants, code review harnesses, evaluation frameworks, and agentic research pipelines—using Python and modern...Full timeWork at office
- ...thinking organization, apply now.We are currently seeking a Gen AI Engineer to join our team in a hybrid basis either out of our NYC or... ..., and least-privilege access.· Productionize LLMs: Build evaluation framework for open-source and foundational LLMs; implement retrieval...Work at officeRemote workFlexible hours3 days per week
- ...FinTechSelling Points Contribute to innovative AI systems for enterprise finance... ...development, focusing on Python and production engineering tasks.Support customer implementations... ....Develop and maintain model pipelines, evaluation sets, and monitoring tools.Analyze...
$175k - $250k
AI EngineerAbout MillenniumMillennium is a global, diversified alternative investment... ...incentive workflowsDesign and implement evaluation frameworks to measure solution quality, robustness... ...closely with product managers, data engineers, and business stakeholders to translate...Flexible hours$133k - $166k
...Description & ResponsibilitiesPersistent Systems is looking for an AI Engineer to design, build, and deploy secure AI applications that help... ...safety, security, governance, and monitoring practicesCreate evaluation frameworks to measure AI quality, accuracy, latency, cost,...Flexible hours$161.8k - $184.6k
Senior AI Engineer Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking... ...model inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage...Full timePart timeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Engineer. Be the first to apply!


