Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Evaluation Engineer

$161.6k - $200k

Capital Rx

About Judi Health

Judi Health is a health technology company providing benefit administration solutions to employers, unions, health plans, and government entities. Judi Health replaces fragmented, outdated systems with the industry's first Unified Claims Processing architecture, seamlessly consolidating pharmacy and medical benefit administration on a single, secure platform. By delivering true price transparency, eliminating unnecessary middleman fees, and leveraging advanced AI-powered care delivery, Judi Health helps clients achieve unprecedented operational efficiency and service levels.

At Judi Health, we're deploying the infrastructure our country needs to deliver the healthcare we all deserve. We are the intelligence platform powering benefits plans for millions of Americans and proudly leading the next generation of care. To learn more, visit

Hybrid 3 days (offices in NYC, Denver, CO and Charlotte, NC area)

Position Summary

As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and autonomous agents in production. This role bridges the gap between model development and realworld usage by translating ambiguous product goals into measurable quality targets.

We're looking for someone to lead evaluation end-to-end - from unit and integration testing to offline, online, and statistical evaluations of probabilistic systems. What we need is someone who can design and operate robust evaluation frameworks, partner with scientists and engineers, and ensure we can confidently answer questions like: "Did this change improve or degrade quality, safety, or user outcomes?"

What You'll Build

Evaluation & Quality Pipelines

  • Build data evaluation pipelines that collect production conversations and agent interactions
  • Reconstruct full sessions from traces, logs, recordings, and transcripts
  • Apply labeling and scoring using human feedback signals (surveys, sentiment, outcomes) and automated evaluators (e.g., LLMasjudge)

Continuous Quality & Safety Benchmarking

  • Own weekly and ondemand automated evaluation runs against staging and production
  • Define benchmarks that track accuracy, reliability, and safetyrelated signals
  • Produce trend dashboards that clearly answer: "Did this deploy change quality or risk?"

Unified Evaluation Framework

  • Design and extend a standardized evaluation framework that supports multiple agent types and workflows
  • Translate highlevel product expectations into concrete success criteria and metrics
  • Ensure new agents and features can be evaluated consistently with minimal friction

Self Service Evaluation Tooling

  • Build APIs and internal tools so data scientists and engineers can go from "interesting scenario" to "included in the eval suite" quickly
  • Enable scenario curation, dataset management, and eval execution without deep infrastructure knowledge

Experiment Tracking & Visibility

  • Provide shared visibility into prompt, model, and agent experiments
  • Enable reproducibility and comparison across runs so teams can build on each other's work instead of operating in silos

Position Responsibilities:

Data Engineering

  • Build and maintain ETL pipelines for heterogeneous data sources (traces, logs, transcripts, user feedback)
  • Implement complex data stitching and session reconstruction logic
  • Manage dataset versioning, provenance, and lifecycle

Platform & Observability

  • Develop dashboards and monitoring tools for AI quality metrics
  • Integrate evaluations into CI/CD pipelines for scheduled and gated runs
  • Implement alerting on quality and safety signals, not just infrastructure health

AI / ML Evaluation Tooling

  • Apply and extend LLMasjudge evaluation patterns
  • Design metrics and scoring approaches suitable for stochastic, nondeterministic systems
  • Use tools like LangSmith to track runs, traces, experiments, and evaluation results

Collaboration

  • Partner closely with data science, engineering, and product teams
  • Translate between research goals, product intent, and engineering constraints
  • Help define what "good" looks like for AI behavior in production
  • Advocate for strong developer experience and usability in the tools you build

Required Qualifications

  • 4+ years of experience in data engineering, ML engineering, or software engineering
  • Bachelor's or Master's degree in Computer Science, Machine Learning, or a related quantitative field
    * Strong proficiency in Python
    * Experience building and maintaining production data pipelines
    * Strong SQL skills
    * Experience working with at least one cloud platform (AWS preferred)

Nice to Haves

  • Prior work on LLM or agent evaluation infrastructure
  • Familiarity with designing metrics for safety, reliability, or quality in AI systems
  • Experience with voice or callcenter data (audio, transcripts, sentiment)
  • Experience with browser automation tools (e.g., Playwright) for endtoend evals
  • Deep SQL expertise
New York, NY Salary Range $161,600—$200,000 USD Denver, CO Salary Range $148,400—$185,000 USD Charlotte, NC Salary Range $134,800—$168,500 USD

All employees are responsible for adherence to the Judi Health Code of Conduct including the reporting of non-compliance. This position description is designed to be flexible, allowing management the opportunity to assign or reassign duties and responsibilities as needed to best meet organizational goals.

We provide equal employment opportunities to all employees and applicants for employment and prohibit discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, medical condition, genetic information, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

By submitting an application, you agree to the retention of your personal data for consideration for a future position at Judi Health. More details about Judi Health's privacy practices can be found at

Vacancy posted 19 hours ago
Similar jobs that could be interesting for youBased on the AI Evaluation Engineer in New York, NY vacancy
  •  ...Job Title AI Evaluation Engineer Location Hybrid / Remote Employment Type Full-time Job Summary We are seeking an AI Evaluation Engineer to design, implement, and maintain evaluation frameworks for AI and machine learning systems, with a focus on... 
    Suggested
    Full time
    Remote work

    Ova Technologies

    New York, NY
    2 days ago
  • $40 per hour

     ...cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content, solve technical...  ...penetration testing, red teaming, incident response, detection engineering, DFIR, malware analysis, threat intelligence, or... 
    Suggested
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    New York, NY
    19 hours ago
  • $40 per hour

    A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated content and solve technical problems in cybersecurity. Candidates should have at least 2 years of relevant experience in areas such as penetration testing or incident response. This... 
    Suggested
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    New York, NY
    3 days ago
  • $152k - $241.5k

     ...believe open-weight models are foundational to American AI leadership and cybersecurity, and that trust in AI grows...  ...and broad scientific scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find, validate, and patch software... 
    Suggested
    Full time
    Remote work

    Nvidia

    New York, NY
    15 hours ago
  • A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated cybersecurity content and solve technical security problems. You will play a significant role in training AI models, providing critical feedback, and improving system accuracy. This... 
    Suggested
    Remote job
    Flexible hours

    DataAnnotation

    New York, NY
    1 day ago
  •  ...FinTechSelling Points Drive innovation in AI systems in a hybrid work environment....  ...into impactful solutions.Mentor and guide engineering teams, fostering growth and enhancing...  ...and cloud infrastructure.Implement robust evaluation frameworks to ensure continuous improvement... 
    Work at office
    Remote work

    Green Key Resources

    New York, NY
    3 days ago
  •  ...At FMX Services, LLC, we are looking for an experienced AI Engineer to join our dynamic team in the Fenics UST Development department...  ...expertise will be instrumental in developing agentic AI workflows, evaluating LLM-powered systems, and contributing to our technical... 

    Cantor Fitzgerald

    New York, NY
    4 days ago
  • $200k

    We're looking for an experienced AI Engineer to join Optiver's Applied AI and Platform Engineering team. In this role, you'll design,...  ...including agent platforms, AI assistants, code review harnesses, evaluation frameworks, and agentic research pipelines—using Python and... 
    Work at office

    Optiver

    New York, NY
    3 days ago
  • $133k - $166k

     ...Description & ResponsibilitiesPersistent Systems is looking for an AI Engineer to design, build, and deploy secure AI applications that help...  ...safety, security, governance, and monitoring practicesCreate evaluation frameworks to measure AI quality, accuracy, latency, cost,... 
    Flexible hours

    Persistent Systems

    New York, NY
    1 day ago
  • $175k - $250k

    AI EngineerAbout MillenniumMillennium is a global, diversified alternative investment...  ...incentive workflowsDesign and implement evaluation frameworks to measure solution quality, robustness...  ...closely with product managers, data engineers, and business stakeholders to translate... 
    Flexible hours

    Millennium Management

    New York, NY
    3 days ago
  •  ...Job description Rengo AI is building the intelligence layer for fund management —...  ...strategies . The Role As a Founding AI Engineer , you will build the core system that...  ...in actual portfolio data ~ Build evaluation frameworks for correctness of financial narratives... 
    Full time
    Shift work

    Decircle

    New York, NY
    19 hours ago
  • $135k - $200k

     ...missing children, and more. The Role Forward Deployed AI Engineers work directly with customers owning Gen AI strategy and...  ...landscape. Strong foundation in Machine Learning basics (Evaluation, Training, Problem Decomposition). Strong engineering background... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package
    Flexible hours

    Palantir Technologies

    New York, NY
    19 hours ago
  • $175k - $275k

     ...changing that. We're building the agentic AI platform designed exclusively for...  ...Role Overview We're looking for an AI Engineer to help build the agentic platform at the...  ...that keep outputs reliable Fine-tune and evaluate models — fine-tune in-house and open-source... 
    Full time
    Work at office

    Translucent Inc

    New York, NY
    19 hours ago
  • $180k - $240k

    Forward Deployed Software Engineer — US / NYC About Indicium AI Indicium AI was Anthropic’s first European launch partner, a Preferred Anthropic Partner...  ..., real-time context retrieval systems, and LLM evaluation architectures. This is a high-agency "builder-consultant... 
    Full time
    Flexible hours

    Indicium Ai

    New York, NY
    19 hours ago
  •  ...users engage — and sometimes abuse. Our AI agents, integrated workflow platform, and...  ...-starters. Cinder is seeking an AI/ML engineer to architect and deploy production AI systems...  ...and fine-tuning through rigorous evaluation and production deployment. Develop automated... 
    Full time
    Work at office
    Immediate start

    Cinder

    New York, NY
    19 hours ago
  • $150k - $275k

     ...cracks. Joyful Health is building the AI-powered financial operating system for healthcare...  ...talented and ambitious machine learning engineer with 5+ years of experience building...  ..., and payer policy information Create evaluation frameworks and feedback loops to... 
    Full time
    Private practice
    Work at office
    3 days per week

    Joyful Health

    New York, NY
    19 hours ago
  •  ...one of the hardest problems in enterprise AI: AI models are generic but company...  ...London. We're a deeply technical team of engineers, AI researchers, and strategists with a high...  ...agent systems in production, or designed evaluation frameworks for enterprise deployments, we... 
    Full time

    Edra

    New York, NY
    19 hours ago
  •  ...Founding AI Engineer (LLM + Production) Location: New York, NY (In‑Person) Compensation: Top of Market + Equity + Benefits About...  ...production‑grade LLMs and AI agents. You’ll bring rigor to evaluations, fine‑tuning, benchmarking, and synthetic data pipelines — and... 
    Full time
    Weekend work

    Everstar

    New York, NY
    19 hours ago
  • $90k - $95k

     .... The Role Are you someone who builds AI agents and thinks about how to automate the...  ...you. We're looking for a full-time AI Engineer to help shape and execute Fanatics' AI strategy...  ...and best practices, including tool evaluation, security and InfoSec coordination, and tracking... 
    Full time
    Temporary work
    Work at office

    Fanatics Inc

    New York, NY
    19 hours ago
  • $160k - $230k

     ...Standard Template Labs is a stealth-mode, AI-native startup reimagining the future of...  ...Role Our team is currently seeking AI Engineers that will design, build, and deploy machine...  ...that differentiate our platform. Evaluate performance, optimize models, and ensure... 
    Full time
    Work at office
    Local area

    Standard Template Labs

    New York, NY
    19 hours ago
  • $149.6k - $184k

     ...clinical operations, and internal tooling. We're hiring an AI Engineer to join a new team we're building from the ground up...  ...your work will involve building orchestration layers, prompts, evaluations, retrieval systems, and tooling that make LLMs reliable enough... 
    Full time
    Internship
    Local area

    Weight Loss, Better Sex, Fuller Hair, Improved Skin And More...

    New York, NY
    19 hours ago
  • $200k - $300k

    About Farsight Farsight is the agentic AI platform for financial services, currently...  ...Ventures, supercharged by scalable engineering and AI skills from companies including Amazon...  ...to polished output. Build the evaluation and quality systems behind generated deliverables... 
    Full time
    Local area
    Remote work

    Farsight Ai

    New York, NY
    19 hours ago
  • $155k - $200k

     ...the disciplined collaboration and transcendent thinking as an AI Engineer at Capstone Investment Advisors here. Responsibilities and...  ...applications using Python Engage with domain experts to identify, evaluate, and execute high-value AI use cases Our future colleague... 
    Minimum wage
    Full time

    Capstone Investment Advisors

    New York, NY
    19 hours ago
  • $142.32k - $213.48k

     ...is seeking a versatile Senior GenAI Platform Engineer to spearhead the design and development of cutting-edge Generative AI platforms. This role is a key driver in creating...  ...problems, leading projects through in-depth evaluation of complex business processes, system... 
    Full time

    Citigroup

    New York, NY
    4 days ago
  •  ...AI Engineer | IntePros IntePros is partnering with a rapidly growing applied AI organization that is helping enterprise customers...  ...translate business problems into scalable technical solutions Evaluate and refine AI concepts to determine feasibility, scalability,... 
    Contract work

    IntePros

    New York, NY
    2 days ago
  • $120k - $165k

     ...future of our communities. This is a Software Engineering III position at the Director level, which...  ...Analytics, Machine learning and Gen AI Platform team(s), across multiple project...  ...accelerate technology and business roadmaps Evaluate state-of-art Gen AI centric technologies... 
    Temporary work

    Morgan Stanley

    New York, NY
    1 day ago
  • $100k

     ...Job Description We are seeking a highly skilled AI Prompt Engineer / AI Application Engineer to design, develop, and deploy AI-driven digital...  ...and policy documentation. Implement prompt testing, evaluation, and optimization frameworks to improve accuracy,... 
    Work from home

    Afficiency

    New York, NY
    4 days ago
  • $100k

     ...and growing fast. We're building the AI-powered revenue platform for modern finance...  ...we're building. We're looking for engineers who are ambitious and take ownership end-...  ...opinions about the tooling for building, evaluating, and securely running multi-provider AI systems... 
    Live in
    Work at office
    Visa sponsorship
    Flexible hours

    Sequence Inc

    New York, NY
    2 days ago
  • $175k

    Position SummaryAssured Guaranty is seeking a Senior AI Software Engineer to build production-grade AI applications using our emerging agentic...  ...production LLM-based applications (prompting, integration, evaluation)Exposure to agent-based systems or multi-step AI... 

    Assured Guaranty

    New York, NY
    3 days ago
  • $204k - $255k

    Joining Collibra’s Unstructured AI TeamWork at the forefront of context engineering - shaping how AI systems retrieve, structure, and leverage context to...  ...cataloging, or document AI workflows.Knowledge of model evaluation best practices.Experience with search relevance.A... 
    Work experience placement
    Work at office
    Flexible hours
    2 days per week

    Collibra

    New York, NY
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Evaluation Engineer. Be the first to apply!