AI Evaluation Engineer
$161.6k - $200kJudi Health
AI Evaluation Engineer
Locations: Charlotte, North Carolina, United States; Denver, Colorado, United States; New York, New York, United States
About Judi Health
Judi Health is an enterprise health technology company providing a comprehensive suite of solutions for employers and health plans, including:
- Judi Rx, a public benefit corporation delivering full-service pharmacy benefit management (PBM) solutions to self-insured employers,
- Judi Health™, which offers full-service health benefit management solutions to employers, TPAs, and health plans, and
- Judi®, the industry's leading proprietary Enterprise Health Platform (EHP), which consolidates all claim administration-related workflows in one scalable, secure platform.
Together with our clients, we're rebuilding trust in healthcare in the U.S. and deploying the infrastructure we need for the care we deserve. To learn more, visit
Position Summary
As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and autonomous agents in production. This role bridges the gap between model development and real-world usage by translating ambiguous product goals into measurable quality targets.
We're looking for someone to lead evaluation end-to-end — from unit and integration testing to offline, online, and statistical evaluations of probabilistic systems. What we need is someone who can design and operate robust evaluation frameworks, partner with scientists and engineers, and ensure we can confidently answer questions like: "Did this change improve or degrade quality, safety, or user outcomes?"
What You'll Build
Evaluation & Quality Pipelines
- Build data evaluation pipelines that collect production conversations and agent interactions
- Reconstruct full sessions from traces, logs, recordings, and transcripts
- Apply labeling and scoring using human feedback signals (surveys, sentiment, outcomes) and automated evaluators (e.g., LLM-as-judge)
Continuous Quality & Safety Benchmarking
- Own weekly and on-demand automated evaluation runs against staging and production
- Define benchmarks that track accuracy, reliability, and safety-related signals
- Produce trend dashboards that clearly answer: "Did this deploy change quality or risk?"
Unified Evaluation Framework
- Design and extend a standardized evaluation framework that supports multiple agent types and workflows
- Translate high-level product expectations into concrete success criteria and metrics
- Ensure new agents and features can be evaluated consistently with minimal friction
Self Service Evaluation Tooling
- Build APIs and internal tools so data scientists and engineers can go from "interesting scenario" to "included in the eval suite" quickly
- Enable scenario curation, dataset management, and eval execution without deep infrastructure knowledge
Experiment Tracking & Visibility
- Provide shared visibility into prompt, model, and agent experiments
- Enable reproducibility and comparison across runs so teams can build on each other's work instead of operating in silos
Position Responsibilities:
Data Engineering
- Build and maintain ETL pipelines for heterogeneous data sources (traces, logs, transcripts, user feedback)
- Implement complex data stitching and session reconstruction logic
- Manage dataset versioning, provenance, and lifecycle
Platform & Observability
- Develop dashboards and monitoring tools for AI quality metrics
- Integrate evaluations into CI/CD pipelines for scheduled and gated runs
- Implement alerting on quality and safety signals, not just infrastructure health
AI / ML Evaluation Tooling
- Apply and extend LLM-as-judge evaluation patterns
- Design metrics and scoring approaches suitable for stochastic, non-deterministic systems
- Use tools like LangSmith to track runs, traces, experiments, and evaluation results
Collaboration
- Partner closely with data science, engineering, and product teams
- Translate between research goals, product intent, and engineering constraints
- Help define what "good" looks like for AI behavior in production
- Advocate for strong developer experience and usability in the tools you build
- Responsible for adherence to the Judi Health Code of Conduct including the reporting of non-compliance.
Required Qualifications
- 4+ years of experience in data engineering, ML engineering, or software engineering
- Bachelor's or Master's degree in Computer Science, Machine Learning, or a related quantitative field • Strong proficiency in Python • Experience building and maintaining production data pipelines • Strong SQL skills • Experience working with at least one cloud platform (AWS preferred)
Nice to Haves
- Prior work on LLM or agent evaluation infrastructure
- Familiarity with designing metrics for safety, reliability, or quality in AI systems
- Experience with voice or call-center data (audio, transcripts, sentiment)
- Experience with browser automation tools (e.g., Playwright) for end-to-end evals
- Deep SQL expertise
New York, NY Salary Range: $161,600 - $200,000 USD
Denver, CO Salary Range: $148,400 - $185,000 USD
Charlotte, NC Salary Range: $134,800 - $168,500 USD
All employees are responsible for adherence to the Judi Health Code of Conduct including the reporting of non-compliance. This position description is designed to be flexible, allowing management the opportunity to assign or reassign duties and responsibilities as needed to best meet organizational goals.
We provide equal employment opportunities to all employees and applicants for employment and prohibit discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, medical condition, genetic information, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.
By submitting an application, you agree to the retention of your personal data for consideration for a future position at Judi Health. More details about Judi Health's privacy practices can be found at
$150k - $250k
Slingshot Aerospace is seeking a Senior AI Engineer to join our AI and Data Science team. This role involves developing evaluation frameworks for intelligent systems in mission-critical space operations. Responsibilities include maintaining our validation SDK, designing...SuggestedRemote job$40 per hour
A technology consulting company is seeking experienced cybersecurity professionals for a remote position. In this role, you'll evaluate AI-generated cybersecurity content, solve technical problems, and provide valuable feedback to enhance AI models. Ideal candidates should...SuggestedHourly payRemote work$40 per hour
A cybersecurity firm in the United States seeks experienced professionals to evaluate AI-generated security content and solve technical cybersecurity problems. In this remote role, you'll work on your own schedule, contributing to the next generation of AI security systems...SuggestedHourly payRemote work$40 per hour
A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated content and solve technical problems in cybersecurity. Candidates should have at least 2 years of relevant experience in areas such as penetration testing or incident response. This...SuggestedHourly payRemote workFlexible hours- A cybersecurity training company is seeking experienced cybersecurity professionals to evaluate AI-generated security content and tackle technical cybersecurity challenges. Candidates should have at least 2 years of hands-on experience in cybersecurity, along with some...SuggestedFull timePart timeRemote workFlexible hours
$40 per hour
A leading AI-focused cybersecurity firm is looking for experienced cybersecurity professionals to evaluate AI-generated content and solve technical security problems. In this flexible role, you can work remotely and choose your projects. Ideal candidates will have 2+ years...Remote jobHourly payFlexible hours- A cybersecurity solutions company is looking for experienced cybersecurity professionals to help train AI models. You will work remotely to evaluate AI-generated security content, solve technical problems, and provide feedback to improve AI systems. Ideal candidates have...Remote jobFlexible hours
- A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated cybersecurity content and solve technical security problems. You will play a significant role in training AI models, providing critical feedback, and improving system accuracy. This...Remote jobFlexible hours
- ...hours Role Responsibilities Review and refine AI-generated prompts, responses, and code... ...or language Support benchmarking efforts to evaluate and compare model capabilities Requirements Strong experience in software engineering, technical research, or educational content...Part timeRemote work
$40 per hour
We are looking for a Web Application Developer to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the quality of each model. To apply to this role, you will need to be proficient...Remote jobHourly payFull timeContract workPart time$300k - $350k
...At FMX Services, LLC, we are looking for an experienced AI Engineer to join our dynamic team in the Fenics UST Development department... ...expertise will be instrumental in developing agentic AI workflows, evaluating LLM-powered systems, and contributing to our technical...$160k - $230k
...AI Engineer Standard Template Labs is a stealth-mode, AI-native startup reimagining the future of IT Service and Configuration Management... ...define AI-first features that differentiate our platform. Evaluate performance, optimize models, and ensure reliability,...Work at officeLocal area$40k
...Agentic AI Development Design and implement autonomous AI agents capable of reasoning... ...related frameworks. Implement prompt engineering, memory systems, and reasoning chains.... ...AI Systems Implement observability, evaluation, and guardrails for agent behavior....Full timeFor contractors- ...AI Engineer Starbridge helps go-to-market teams win in the public sector. Government buyers signal their priorities constantly, in... ...work on projects that sit at the core of our product—building, evaluating, and deploying LLM-driven features that make it easier for...Work at office
- ...AI Engineer Our client is a leading global independent investment banking advisory firm that provides strategic advice on mergers... ...Experience working with structured and unstructured data AI model evaluation, tuning, and performance optimization Ability to translate...Worldwide
$175k - $275k
...AI Engineer Vatic is looking for an AI engineer with proven experience with large foundational models. Our environment is highly collaborative... ...large-scale experiments that contribute towards rigorous evaluation and in-depth model analysis. We employ a team of...Work at officeNight shift$160k - $210k
...Join to apply for the Founding AI Engineer role at Interfere (YC S25) in New York, NY. Base pay range $160,000.00/yr - $210,000.00/yr... ...understand when to act autonomously and when to defer to humans. Evaluation frameworks that measure what matters—real‑world reliability,...Full timeVisa sponsorship$141.3k - $185.4k
...AI Engineer | Data Science & AI Engineering Full-Time Hybrid (3 days/per week in office) The Opportunity MassMutual's AI & Data... ...challenges, working independently to scope problems, build and evaluate solutions, and bring them into production. At this level, you...Full timeWork at office3 days per week- ...AI Engineer Location: San Francisco or New York City About Pathwork Pathwork is redesigning life & health insurance jobs... ...owning model orchestration, retrieval, real-time inference, evaluation, and production infrastructure, and collaborating closely with...Work at office
- ...is simple: Be extraordinary, together. ABOUT THE ROLE The AI Enablement Engineer plays a critical role in accelerating the adoption and practical... ...developments and recommendations Provide input on vendor evaluation, tool consolidation, and platform investment decisions...
$154k - $175k
...Email Join our Americas Data and AI team and help shape the future of AI at... ...technical expert, bridging strategy, design, engineering, and change adoption to accelerate... ...Engineering and Assurance teams through evaluation frameworks, observability, and reviews of...Temporary workWork from homeFlexible hours- ...Job Title: Applied AI Engineer (GenAI / LLM Platforms) Location: New York New York 10019 ( Hybrid 3 days onsite 2 days remote... ...platforms, RAG architectures, AI orchestration frameworks, evaluation methodologies, and enterprise AI governance. Key Responsibilities...Contract workRemote work
$110k - $160k
...role, you will design, build, and deploy AI-powered systems that accelerate AI... ...role will partner directly with Product, Engineering, Data, Security, Legal, Compliance, Operations... ...cases end-to-end, from discovery and evaluation through launch, monitoring, and iteration...Flexible hours$160k
...About the job AI Engineer Job Title: AI Engineer Agentic & RAG Systems Location: Remote Department: AI & Data Platforms... ...-agent orchestration, Retrieval-Augmented Generation (RAG), evaluation frameworks, and AI guardrails to build safe, reliable, and...Remote work- ...Founding AI Engineer at Hera (YC S25) We're building the world's first AI motion designer that creates professional animations in seconds... ...-linear movement, visual hierarchy). Develop comprehensive evaluation frameworks to assess animation quality, combining automated...Work at officeRelocation
$135k - $200k
...locate missing children, and more. The Role Forward Deployed AI Engineers work directly with customers owning Gen AI strategy and implementation... ...AI landscape. Strong foundation in Machine Learning basics (Evaluation, Training, Problem Decomposition). Strong engineering...Work experience placementWork at officeRemote workWork from homeRelocation packageFlexible hours$269.1k - $307.2k
...Distinguished AI Engineer Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for... ...language model inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc....Full timePart timeLocal area$200k - $300k
...AI Engineer Title of Role: AI Engineer Location: New York, onsite Company Stage of Funding: Seed — Software Development Office... ...solutions. Engage in prompt engineering and develop evaluation frameworks to enhance AI tooling. Proactively identify and...Work at office$140k - $180k
...radically raise the bar on what agentic AI, CTV, eCommerce, social, and mobile can do... ...building it. Mission As the AI Engineer at Kargo, you will architect, build, and... ...) to accelerate velocity; familiar with evaluation and observability frameworks for LLM...Work experience placementLocal area- ...About the job AI ENGINEER Role Description We're looking for a talented and ambitious machine learning engineer with 5+... ...documentation, billing records, and payer policy information Create evaluation frameworks and feedback loops to continuously improve agent...Full timeRelocation
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Engineer. Be the first to apply!

