AI Evaluation Engineer
$161.6k - $200kCapital Rx
About Judi Health
Judi Health is a health technology company providing benefit administration solutions to employers, unions, health plans, and government entities. Judi Health replaces fragmented, outdated systems with the industry's first Unified Claims Processing architecture, seamlessly consolidating pharmacy and medical benefit administration on a single, secure platform. By delivering true price transparency, eliminating unnecessary middleman fees, and leveraging advanced AI-powered care delivery, Judi Health helps clients achieve unprecedented operational efficiency and service levels.
At Judi Health, we're deploying the infrastructure our country needs to deliver the healthcare we all deserve. We are the intelligence platform powering benefits plans for millions of Americans and proudly leading the next generation of care. To learn more, visitHybrid 3 days (offices in NYC, Denver, CO and Charlotte, NC area)
Position Summary
As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and autonomous agents in production. This role bridges the gap between model development and realworld usage by translating ambiguous product goals into measurable quality targets.
We're looking for someone to lead evaluation end-to-end - from unit and integration testing to offline, online, and statistical evaluations of probabilistic systems. What we need is someone who can design and operate robust evaluation frameworks, partner with scientists and engineers, and ensure we can confidently answer questions like: "Did this change improve or degrade quality, safety, or user outcomes?"
What You'll Build
Evaluation & Quality Pipelines
- Build data evaluation pipelines that collect production conversations and agent interactions
- Reconstruct full sessions from traces, logs, recordings, and transcripts
- Apply labeling and scoring using human feedback signals (surveys, sentiment, outcomes) and automated evaluators (e.g., LLMasjudge)
Continuous Quality & Safety Benchmarking
- Own weekly and ondemand automated evaluation runs against staging and production
- Define benchmarks that track accuracy, reliability, and safetyrelated signals
- Produce trend dashboards that clearly answer: "Did this deploy change quality or risk?"
Unified Evaluation Framework
- Design and extend a standardized evaluation framework that supports multiple agent types and workflows
- Translate highlevel product expectations into concrete success criteria and metrics
- Ensure new agents and features can be evaluated consistently with minimal friction
Self Service Evaluation Tooling
- Build APIs and internal tools so data scientists and engineers can go from "interesting scenario" to "included in the eval suite" quickly
- Enable scenario curation, dataset management, and eval execution without deep infrastructure knowledge
Experiment Tracking & Visibility
- Provide shared visibility into prompt, model, and agent experiments
- Enable reproducibility and comparison across runs so teams can build on each other's work instead of operating in silos
Position Responsibilities:
Data Engineering
- Build and maintain ETL pipelines for heterogeneous data sources (traces, logs, transcripts, user feedback)
- Implement complex data stitching and session reconstruction logic
- Manage dataset versioning, provenance, and lifecycle
Platform & Observability
- Develop dashboards and monitoring tools for AI quality metrics
- Integrate evaluations into CI/CD pipelines for scheduled and gated runs
- Implement alerting on quality and safety signals, not just infrastructure health
AI / ML Evaluation Tooling
- Apply and extend LLMasjudge evaluation patterns
- Design metrics and scoring approaches suitable for stochastic, nondeterministic systems
- Use tools like LangSmith to track runs, traces, experiments, and evaluation results
Collaboration
- Partner closely with data science, engineering, and product teams
- Translate between research goals, product intent, and engineering constraints
- Help define what "good" looks like for AI behavior in production
- Advocate for strong developer experience and usability in the tools you build
Required Qualifications
- 4+ years of experience in data engineering, ML engineering, or software engineering
- Bachelor's or Master's degree in Computer Science, Machine Learning, or a related quantitative field
* Strong proficiency in Python
* Experience building and maintaining production data pipelines
* Strong SQL skills
* Experience working with at least one cloud platform (AWS preferred)
Nice to Haves
- Prior work on LLM or agent evaluation infrastructure
- Familiarity with designing metrics for safety, reliability, or quality in AI systems
- Experience with voice or callcenter data (audio, transcripts, sentiment)
- Experience with browser automation tools (e.g., Playwright) for endtoend evals
- Deep SQL expertise
All employees are responsible for adherence to the Judi Health Code of Conduct including the reporting of non-compliance. This position description is designed to be flexible, allowing management the opportunity to assign or reassign duties and responsibilities as needed to best meet organizational goals.
We provide equal employment opportunities to all employees and applicants for employment and prohibit discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, medical condition, genetic information, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.
By submitting an application, you agree to the retention of your personal data for consideration for a future position at Judi Health. More details about Judi Health's privacy practices can be found at
- ...Job Title AI Evaluation Engineer Location Hybrid / Remote Employment Type Full-time Job Summary We are seeking an AI Evaluation Engineer to design, implement, and maintain evaluation frameworks for AI and machine learning systems, with a focus on...SuggestedFull timeRemote work
$40 per hour
...cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content, solve technical... ...penetration testing, red teaming, incident response, detection engineering, DFIR, malware analysis, threat intelligence, or...SuggestedHourly payFull timePart timeRemote work$40 per hour
A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated content and solve technical problems in cybersecurity. Candidates should have at least 2 years of relevant experience in areas such as penetration testing or incident response. This...SuggestedHourly payRemote workFlexible hours$152k - $241.5k
...believe open-weight models are foundational to American AI leadership and cybersecurity, and that trust in AI grows... ...and broad scientific scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find, validate, and patch software...SuggestedFull timeRemote work- A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated cybersecurity content and solve technical security problems. You will play a significant role in training AI models, providing critical feedback, and improving system accuracy. This...SuggestedRemote jobFlexible hours
- ...FinTechSelling Points Drive innovation in AI systems in a hybrid work environment.... ...into impactful solutions.Mentor and guide engineering teams, fostering growth and enhancing... ...and cloud infrastructure.Implement robust evaluation frameworks to ensure continuous improvement...Work at officeRemote work
- ...At FMX Services, LLC, we are looking for an experienced AI Engineer to join our dynamic team in the Fenics UST Development department... ...expertise will be instrumental in developing agentic AI workflows, evaluating LLM-powered systems, and contributing to our technical...
$200k
We're looking for an experienced AI Engineer to join Optiver's Applied AI and Platform Engineering team. In this role, you'll design,... ...including agent platforms, AI assistants, code review harnesses, evaluation frameworks, and agentic research pipelines—using Python and...Work at office$133k - $166k
...Description & ResponsibilitiesPersistent Systems is looking for an AI Engineer to design, build, and deploy secure AI applications that help... ...safety, security, governance, and monitoring practicesCreate evaluation frameworks to measure AI quality, accuracy, latency, cost,...Flexible hours$175k - $250k
AI EngineerAbout MillenniumMillennium is a global, diversified alternative investment... ...incentive workflowsDesign and implement evaluation frameworks to measure solution quality, robustness... ...closely with product managers, data engineers, and business stakeholders to translate...Flexible hours- ...Job description Rengo AI is building the intelligence layer for fund management —... ...strategies . The Role As a Founding AI Engineer , you will build the core system that... ...in actual portfolio data ~ Build evaluation frameworks for correctness of financial narratives...Full timeShift work
$135k - $200k
...missing children, and more. The Role Forward Deployed AI Engineers work directly with customers owning Gen AI strategy and... ...landscape. Strong foundation in Machine Learning basics (Evaluation, Training, Problem Decomposition). Strong engineering background...Full timeWork experience placementWork at officeRemote workWork from homeRelocation packageFlexible hours$175k - $275k
...changing that. We're building the agentic AI platform designed exclusively for... ...Role Overview We're looking for an AI Engineer to help build the agentic platform at the... ...that keep outputs reliable Fine-tune and evaluate models — fine-tune in-house and open-source...Full timeWork at office$180k - $240k
Forward Deployed Software Engineer — US / NYC About Indicium AI Indicium AI was Anthropic’s first European launch partner, a Preferred Anthropic Partner... ..., real-time context retrieval systems, and LLM evaluation architectures. This is a high-agency "builder-consultant...Full timeFlexible hours- ...users engage — and sometimes abuse. Our AI agents, integrated workflow platform, and... ...-starters. Cinder is seeking an AI/ML engineer to architect and deploy production AI systems... ...and fine-tuning through rigorous evaluation and production deployment. Develop automated...Full timeWork at officeImmediate start
$150k - $275k
...cracks. Joyful Health is building the AI-powered financial operating system for healthcare... ...talented and ambitious machine learning engineer with 5+ years of experience building... ..., and payer policy information Create evaluation frameworks and feedback loops to...Full timePrivate practiceWork at office3 days per week- ...one of the hardest problems in enterprise AI: AI models are generic but company... ...London. We're a deeply technical team of engineers, AI researchers, and strategists with a high... ...agent systems in production, or designed evaluation frameworks for enterprise deployments, we...Full time
- ...Founding AI Engineer (LLM + Production) Location: New York, NY (In‑Person) Compensation: Top of Market + Equity + Benefits About... ...production‑grade LLMs and AI agents. You’ll bring rigor to evaluations, fine‑tuning, benchmarking, and synthetic data pipelines — and...Full timeWeekend work
$90k - $95k
.... The Role Are you someone who builds AI agents and thinks about how to automate the... ...you. We're looking for a full-time AI Engineer to help shape and execute Fanatics' AI strategy... ...and best practices, including tool evaluation, security and InfoSec coordination, and tracking...Full timeTemporary workWork at office$160k - $230k
...Standard Template Labs is a stealth-mode, AI-native startup reimagining the future of... ...Role Our team is currently seeking AI Engineers that will design, build, and deploy machine... ...that differentiate our platform. Evaluate performance, optimize models, and ensure...Full timeWork at officeLocal area$149.6k - $184k
...clinical operations, and internal tooling. We're hiring an AI Engineer to join a new team we're building from the ground up... ...your work will involve building orchestration layers, prompts, evaluations, retrieval systems, and tooling that make LLMs reliable enough...Full timeInternshipLocal area$200k - $300k
About Farsight Farsight is the agentic AI platform for financial services, currently... ...Ventures, supercharged by scalable engineering and AI skills from companies including Amazon... ...to polished output. Build the evaluation and quality systems behind generated deliverables...Full timeLocal areaRemote work$155k - $200k
...the disciplined collaboration and transcendent thinking as an AI Engineer at Capstone Investment Advisors here. Responsibilities and... ...applications using Python Engage with domain experts to identify, evaluate, and execute high-value AI use cases Our future colleague...Minimum wageFull time$142.32k - $213.48k
...is seeking a versatile Senior GenAI Platform Engineer to spearhead the design and development of cutting-edge Generative AI platforms. This role is a key driver in creating... ...problems, leading projects through in-depth evaluation of complex business processes, system...Full time- ...AI Engineer | IntePros IntePros is partnering with a rapidly growing applied AI organization that is helping enterprise customers... ...translate business problems into scalable technical solutions Evaluate and refine AI concepts to determine feasibility, scalability,...Contract work
$120k - $165k
...future of our communities. This is a Software Engineering III position at the Director level, which... ...Analytics, Machine learning and Gen AI Platform team(s), across multiple project... ...accelerate technology and business roadmaps Evaluate state-of-art Gen AI centric technologies...Temporary work$100k
...Job Description We are seeking a highly skilled AI Prompt Engineer / AI Application Engineer to design, develop, and deploy AI-driven digital... ...and policy documentation. Implement prompt testing, evaluation, and optimization frameworks to improve accuracy,...Work from home$100k
...and growing fast. We're building the AI-powered revenue platform for modern finance... ...we're building. We're looking for engineers who are ambitious and take ownership end-... ...opinions about the tooling for building, evaluating, and securely running multi-provider AI systems...Live inWork at officeVisa sponsorshipFlexible hours$175k
Position SummaryAssured Guaranty is seeking a Senior AI Software Engineer to build production-grade AI applications using our emerging agentic... ...production LLM-based applications (prompting, integration, evaluation)Exposure to agent-based systems or multi-step AI...$204k - $255k
Joining Collibra’s Unstructured AI TeamWork at the forefront of context engineering - shaping how AI systems retrieve, structure, and leverage context to... ...cataloging, or document AI workflows.Knowledge of model evaluation best practices.Experience with search relevance.A...Work experience placementWork at officeFlexible hours2 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Engineer. Be the first to apply!


