AI Evaluation Engineer
$161.6k - $200kCapital Rx
Judi Health is an enterprise health technology company providing a comprehensive suite of solutions for employers and health plans, including:
- Judi Rx , a public benefit corporation delivering full-service pharmacy benefit management (PBM) solutions to self-insured employers,
- Judi Health , which offers full-service health benefit management solutions to employers, TPAs, and health plans, and
- Judi , the industry's leading proprietary Enterprise Health Platform (EHP), which consolidates all claim administration-related workflows in one scalable, secure platform.
Together with our clients, we're rebuilding trust in healthcare in the U.S. and deploying the infrastructure we need for the care we deserve. To learn more, visit
Hybrid 3 days (offices in NYC, Denver, CO and Charlotte, NC area)
Position Summary
As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and autonomous agents in production. This role bridges the gap between model development and realworld usage by translating ambiguous product goals into measurable quality targets.
We're looking for someone to lead evaluation end-to-end - from unit and integration testing to offline, online, and statistical evaluations of probabilistic systems. What we need is someone who can design and operate robust evaluation frameworks, partner with scientists and engineers, and ensure we can confidently answer questions like: "Did this change improve or degrade quality, safety, or user outcomes?"
What You'll Build
Evaluation & Quality Pipelines
- Build data evaluation pipelines that collect production conversations and agent interactions
- Reconstruct full sessions from traces, logs, recordings, and transcripts
- Apply labeling and scoring using human feedback signals (surveys, sentiment, outcomes) and automated evaluators (e.g., LLMasjudge)
Continuous Quality & Safety Benchmarking
- Own weekly and ondemand automated evaluation runs against staging and production
- Define benchmarks that track accuracy, reliability, and safetyrelated signals
- Produce trend dashboards that clearly answer: "Did this deploy change quality or risk?"
Unified Evaluation Framework
- Design and extend a standardized evaluation framework that supports multiple agent types and workflows
- Translate highlevel product expectations into concrete success criteria and metrics
- Ensure new agents and features can be evaluated consistently with minimal friction
Self Service Evaluation Tooling
- Build APIs and internal tools so data scientists and engineers can go from "interesting scenario" to "included in the eval suite" quickly
- Enable scenario curation, dataset management, and eval execution without deep infrastructure knowledge
Experiment Tracking & Visibility
- Provide shared visibility into prompt, model, and agent experiments
- Enable reproducibility and comparison across runs so teams can build on each other's work instead of operating in silos
Position Responsibilities:
Data Engineering
- Build and maintain ETL pipelines for heterogeneous data sources (traces, logs, transcripts, user feedback)
- Implement complex data stitching and session reconstruction logic
- Manage dataset versioning, provenance, and lifecycle
Platform & Observability
- Develop dashboards and monitoring tools for AI quality metrics
- Integrate evaluations into CI/CD pipelines for scheduled and gated runs
- Implement alerting on quality and safety signals, not just infrastructure health
AI / ML Evaluation Tooling
- Apply and extend LLMasjudge evaluation patterns
- Design metrics and scoring approaches suitable for stochastic, nondeterministic systems
- Use tools like LangSmith to track runs, traces, experiments, and evaluation results
Collaboration
- Partner closely with data science, engineering, and product teams
- Translate between research goals, product intent, and engineering constraints
- Help define what "good" looks like for AI behavior in production
- Advocate for strong developer experience and usability in the tools you build
Required Qualifications
- 4+ years of experience in data engineering, ML engineering, or software engineering
- Bachelor's or Master's degree in Computer Science, Machine Learning, or a related quantitative field
* Strong proficiency in Python
* Experience building and maintaining production data pipelines
* Strong SQL skills
* Experience working with at least one cloud platform (AWS preferred)
Nice to Haves
- Prior work on LLM or agent evaluation infrastructure
- Familiarity with designing metrics for safety, reliability, or quality in AI systems
- Experience with voice or callcenter data (audio, transcripts, sentiment)
- Experience with browser automation tools (e.g., Playwright) for endtoend evals
- Deep SQL expertise
All employees are responsible for adherence to the Judi Health Code of Conduct including the reporting of non-compliance. This position description is designed to be flexible, allowing management the opportunity to assign or reassign duties and responsibilities as needed to best meet organizational goals.
We provide equal employment opportunities to all employees and applicants for employment and prohibit discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, medical condition, genetic information, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.
By submitting an application, you agree to the retention of your personal data for consideration for a future position at Judi Health. More details about Judi Health's privacy practices can be found at
$150k - $250k
Slingshot Aerospace is seeking a Senior AI Engineer to join our AI and Data Science team. This role involves developing evaluation frameworks for intelligent systems in mission-critical space operations. Responsibilities include maintaining our validation SDK, designing...SuggestedRemote job- A cybersecurity training company is seeking experienced cybersecurity professionals to evaluate AI-generated security content and tackle technical cybersecurity challenges. Candidates should have at least 2 years of hands-on experience in cybersecurity, along with some...SuggestedFull timePart timeRemote workFlexible hours
$40 per hour
A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated content and solve technical problems in cybersecurity. Candidates should have at least 2 years of relevant experience in areas such as penetration testing or incident response. This...SuggestedHourly payRemote workFlexible hours$40 per hour
A cybersecurity firm in the United States seeks experienced professionals to evaluate AI-generated security content and solve technical cybersecurity problems. In this remote role, you'll work on your own schedule, contributing to the next generation of AI security systems...SuggestedHourly payRemote work$40 per hour
A technology consulting company is seeking experienced cybersecurity professionals for a remote position. In this role, you'll evaluate AI-generated cybersecurity content, solve technical problems, and provide valuable feedback to enhance AI models. Ideal candidates should...SuggestedHourly payRemote work$40 per hour
A leading AI-focused cybersecurity firm is looking for experienced cybersecurity professionals to evaluate AI-generated content and solve technical security problems. In this flexible role, you can work remotely and choose your projects. Ideal candidates will have 2+ years...Remote jobHourly payFlexible hours- A cybersecurity solutions company is looking for experienced cybersecurity professionals to help train AI models. You will work remotely to evaluate AI-generated security content, solve technical problems, and provide feedback to improve AI systems. Ideal candidates have...Remote jobFlexible hours
- A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated cybersecurity content and solve technical security problems. You will play a significant role in training AI models, providing critical feedback, and improving system accuracy. This...Remote jobFlexible hours
$150k - $275k
...cracks. Joyful Health is building the AI-powered financial operating system for healthcare... ...talented and ambitious machine learning engineer with 5+ years of experience building... ..., and payer policy information Create evaluation frameworks and feedback loops to...Full timePrivate practiceWork at office3 days per week$135k - $200k
...missing children, and more. The Role Forward Deployed AI Engineers work directly with customers owning Gen AI strategy and... ...landscape. Strong foundation in Machine Learning basics (Evaluation, Training, Problem Decomposition). Strong engineering background...Full timeWork experience placementWork at officeRemote workWork from homeRelocation packageFlexible hours- ...Job description Rengo AI is building the intelligence layer for fund management —... ...strategies . The Role As a Founding AI Engineer , you will build the core system that... ...in actual portfolio data ~ Build evaluation frameworks for correctness of financial narratives...Full timeShift work
$160k - $230k
...Standard Template Labs is a stealth-mode, AI-native startup reimagining the future of... ...Role Our team is currently seeking AI Engineers that will design, build, and deploy machine... ...that differentiate our platform. Evaluate performance, optimize models, and ensure...Full timeWork at officeLocal area$149.6k - $184k
...clinical operations, and internal tooling. We're hiring an AI Engineer to join a new team we're building from the ground up... ...your work will involve building orchestration layers, prompts, evaluations, retrieval systems, and tooling that make LLMs reliable enough...Full timeInternshipLocal area$155k - $200k
...the disciplined collaboration and transcendent thinking as an AI Engineer at Capstone Investment Advisors here. Responsibilities and... ...applications using Python Engage with domain experts to identify, evaluate, and execute high-value AI use cases Our future colleague...Minimum wageFull time- ...users engage — and sometimes abuse. Our AI agents, integrated workflow platform, and... ...-starters. Cinder is seeking an AI/ML engineer to architect and deploy production AI systems... ...and fine-tuning through rigorous evaluation and production deployment. Develop automated...Full timeWork at officeImmediate start
- ...one of the hardest problems in enterprise AI: AI models are generic but company... ...London. We're a deeply technical team of engineers, AI researchers, and strategists with a high... ...agent systems in production, or designed evaluation frameworks for enterprise deployments, we...Full time
$90k - $95k
.... The Role Are you someone who builds AI agents and thinks about how to automate the... ...you. We're looking for a full-time AI Engineer to help shape and execute Fanatics' AI strategy... ...and best practices, including tool evaluation, security and InfoSec coordination, and tracking...Full timeTemporary workWork at office- ...Founding AI Engineer (LLM + Production) Location: New York, NY (In‑Person) Compensation: Top of Market + Equity + Benefits About... ...production‑grade LLMs and AI agents. You’ll bring rigor to evaluations, fine‑tuning, benchmarking, and synthetic data pipelines — and...Full timeWeekend work
$110k - $160k
...role, you will design, build, and deploy AI-powered systems that accelerate AI... ...role will partner directly with Product, Engineering, Data, Security, Legal, Compliance, Operations... ...cases end-to-end, from discovery and evaluation through launch, monitoring, and iteration...Flexible hours$160k - $210k
...Join to apply for the Founding AI Engineer role at Interfere (YC S25) in New York, NY. Base pay range $160,000.00/yr - $210,000.00/yr... ...understand when to act autonomously and when to defer to humans. Evaluation frameworks that measure what matters—real‑world reliability,...Full timeVisa sponsorship- ...Founding AI Engineer at Hera (YC S25) We're building the world's first AI motion designer that creates professional animations in seconds... ...-linear movement, visual hierarchy). Develop comprehensive evaluation frameworks to assess animation quality, combining automated...Work at officeRelocation
$135k - $200k
...locate missing children, and more. The Role Forward Deployed AI Engineers work directly with customers owning Gen AI strategy and implementation... ...AI landscape. Strong foundation in Machine Learning basics (Evaluation, Training, Problem Decomposition). Strong engineering...Work experience placementWork at officeRemote workWork from homeRelocation packageFlexible hours- ...AI Engineer Founded in 1917, the National Hockey League (NHL®) is the premier professional ice hockey league in the world and is one... ...building parallel systems. Define success criteria, build evaluation frameworks, and run structured tests before any system goes...Full timeWork at office
$154k - $175k
...Join our Americas Data and AI team and help shape the future of AI at Macquarie.... ...technical expert, bridging strategy, design, engineering, and change adoption to accelerate... ...Engineering and Assurance teams through evaluation frameworks, observability, and reviews of...Temporary workWork from homeFlexible hours$170k - $200k
...AI Engineer New York, New York, United States Genius Sports is enabling a new era of sports for fans worldwide, delivering experiences... ...inference pipelines that orchestrate them, and rigorously evaluating output quality against messy, real-world data. You'll work...Work at officeWorldwide- ...We have a contract opportunity for AI Engineer with one of our clients. Detailed job descriptions are provided below. Title:... ...including governance of which systems are exposed via MCP. Evaluate and integrate external AI platforms (e.g., Azure OpenAI,...Contract work
$100k
...and growing fast. We're building the AI-powered revenue platform for modern finance... ...we're building. We're looking for engineers who are ambitious and take ownership end-... ...opinions about the tooling for building, evaluating, and securely running multi-provider AI systems...Live inWork at officeVisa sponsorshipFlexible hours- ...We're looking for a junior-level AI developer with a solid software engineering foundation and a drive to build and deploy AI-powered applications that... ...JS), back-end APIs, and cloud-hosted AI services • Evaluate, integrate, and maintain AI platforms, APIs, and toolchains...
- ...Role : AI Engineer Location : New York Its day onsite role Skills - AI, AWS, TypeScript, JavaScript, Python, Go... ...pipelines • LLM infrastructure, inference, and model gateways • Evaluation, observability, and safety tooling for autonomous systems •...
$120k - $140k
...the business. Fractal is a strategic AI partner to Fortune 500 companies with a vision... ...Fractal. Job Title: Senior Agentic AI Engineer / AI-ML Engineer Location... ...exception handling, and rule recommendation. Evaluate the stability, reliability, and performance...Hourly payFull timeLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Engineer. Be the first to apply!


