AI Evaluation Engineer
$161.6k - $200kCapital Rx
Judi Health is an enterprise health technology company providing a comprehensive suite of solutions for employers and health plans, including:
- Judi Rx , a public benefit corporation delivering full-service pharmacy benefit management (PBM) solutions to self-insured employers,
- Judi Health , which offers full-service health benefit management solutions to employers, TPAs, and health plans, and
- Judi , the industry's leading proprietary Enterprise Health Platform (EHP), which consolidates all claim administration-related workflows in one scalable, secure platform.
Together with our clients, we're rebuilding trust in healthcare in the U.S. and deploying the infrastructure we need for the care we deserve. To learn more, visit
Position Summary
As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and autonomous agents in production. This role bridges the gap between model development and realworld usage by translating ambiguous product goals into measurable quality targets.
We're looking for someone to lead evaluation end-to-end - from unit and integration testing to offline, online, and statistical evaluations of probabilistic systems. What we need is someone who can design and operate robust evaluation frameworks, partner with scientists and engineers, and ensure we can confidently answer questions like: "Did this change improve or degrade quality, safety, or user outcomes?"
What You'll Build
Evaluation & Quality Pipelines
- Build data evaluation pipelines that collect production conversations and agent interactions
- Reconstruct full sessions from traces, logs, recordings, and transcripts
- Apply labeling and scoring using human feedback signals (surveys, sentiment, outcomes) and automated evaluators (e.g., LLMasjudge)
Continuous Quality & Safety Benchmarking
- Own weekly and ondemand automated evaluation runs against staging and production
- Define benchmarks that track accuracy, reliability, and safetyrelated signals
- Produce trend dashboards that clearly answer: "Did this deploy change quality or risk?"
Unified Evaluation Framework
- Design and extend a standardized evaluation framework that supports multiple agent types and workflows
- Translate highlevel product expectations into concrete success criteria and metrics
- Ensure new agents and features can be evaluated consistently with minimal friction
Self Service Evaluation Tooling
- Build APIs and internal tools so data scientists and engineers can go from "interesting scenario" to "included in the eval suite" quickly
- Enable scenario curation, dataset management, and eval execution without deep infrastructure knowledge
Experiment Tracking & Visibility
- Provide shared visibility into prompt, model, and agent experiments
- Enable reproducibility and comparison across runs so teams can build on each other's work instead of operating in silos
Position Responsibilities:
Data Engineering
- Build and maintain ETL pipelines for heterogeneous data sources (traces, logs, transcripts, user feedback)
- Implement complex data stitching and session reconstruction logic
- Manage dataset versioning, provenance, and lifecycle
Platform & Observability
- Develop dashboards and monitoring tools for AI quality metrics
- Integrate evaluations into CI/CD pipelines for scheduled and gated runs
- Implement alerting on quality and safety signals, not just infrastructure health
AI / ML Evaluation Tooling
- Apply and extend LLMasjudge evaluation patterns
- Design metrics and scoring approaches suitable for stochastic, nondeterministic systems
- Use tools like LangSmith to track runs, traces, experiments, and evaluation results
Collaboration
- Partner closely with data science, engineering, and product teams
- Translate between research goals, product intent, and engineering constraints
- Help define what "good" looks like for AI behavior in production
- Advocate for strong developer experience and usability in the tools you build
- Responsible for adherence to the Capital Rx Code of Conduct including the reporting of non-compliance.
Required Qualifications
- 4+ years of experience in data engineering, ML engineering, or software engineering
- Bachelor's or Master's degree in Computer Science, Machine Learning, or a related quantitative field
* Strong proficiency in Python
* Experience building and maintaining production data pipelines
* Strong SQL skills
* Experience working with at least one cloud platform (AWS preferred)
Nice to Haves
- Prior work on LLM or agent evaluation infrastructure
- Familiarity with designing metrics for safety, reliability, or quality in AI systems
- Experience with voice or callcenter data (audio, transcripts, sentiment)
- Experience with browser automation tools (e.g., Playwright) for endtoend evals
- Deep SQL expertise
All employees are responsible for adherence to the Judi Health Code of Conduct including the reporting of non-compliance. This position description is designed to be flexible, allowing management the opportunity to assign or reassign duties and responsibilities as needed to best meet organizational goals.
We provide equal employment opportunities to all employees and applicants for employment and prohibit discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, medical condition, genetic information, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.
By submitting an application, you agree to the retention of your personal data for consideration for a future position at Judi Health. More details about Judi Health's privacy practices can be found at
$116k - $175k
...continuous professional development. The Prompt + Skills Engineer is the hands‑on builder in Cherry Bekaert’s AI Center of Excellence — the person who writes the... ...gets done. Participates in use case intake, evaluating submitted ideas from across the Firm for technical...SuggestedWork experience placementLocal area$126.8k - $158.5k
...shape wherever we go next. You create your future and ours. AI Engineer JOB SUMMARY We are seeking a motivated AI Engineer... ...MLOps practices appropriate to the solution (versioning, evaluation, monitoring signals, and safe rollout patterns). Reviews and...SuggestedInternshipWork at office$60 per hour
...A leading AI development firm is seeking proficient programmers to join their remote team. You'll tackle diverse coding challenges, create applications, and provide critical evaluations of AI-generated code. Successful candidates will be fluent in English and proficient...SuggestedRemote workFlexible hours- ...A leading AI research accelerator is seeking a skilled software engineer to evaluate AI-generated code and improve its efficiency and reliability. The role involves collaboration with cross-functional teams to enhance coding solutions, requiring a minimum of 5 years of...SuggestedContract workFor contractorsRemote work10 hours per weekFlexible hours
- A leading AI research accelerator is looking for a contractor to evaluate AI-generated code and enhance AI-driven coding solutions. The ideal candidate will have over 5 years of software engineering experience, including time at a top-tier company, and possess strong skills...SuggestedContract workFor contractorsRemote work10 hours per weekFlexible hours
$90k - $105k
...Senior Life Sciences Knowledge Engineer Company: Norstella Location: Remote, United... ...and critical global life sciences data and AI solutions provider dedicated to improving... ...unites market-leading brands - Citeline, Evaluate, MMIT, Panalgo, Skipta and The Dedham Group...Full timeTemporary workWork at officeLocal areaRemote workFlexible hours$150k - $184k
...per week. The Opportunity The Generative AI Innovation Team is transforming how Litera... ...competitive advantage. As a Senior AI Engineer, you will play a critical role in shaping... ...technical direction across engineering. Evaluate, prototype, and document emerging tools,...Work experience placement3 days per week$92.5k - $209.5k
...Job Description As a Senior AI Software Engineer in an AI Innovation organization within OCI, you will help build AI capabilities into... ...inference systems, model serving, AI workflow orchestration, evaluation, and observability. Build production-grade services for...Temporary workFlexible hours$55 per hour
Freelance AI Trainer - Civil Engineering & Python 1 day ago Be among the first 25 applicants This opportunity is only for candidates currently residing... ...Civil Engineers with Python skills to train and evaluate AI models on realistic civil engineering problems. This role...Part timeFreelanceRemote workFlexible hours$155k - $235k
...Catalog in Databricks, to the semantic and AI layers that sit on top. This high‑impact... ...standards, and ensuring data works for engineers, analysts, and business users alike. About... ...how modern LLMs are trained, aligned and evaluated (RLHF, fine‑tuning, prompt engineering, retrieval...Home officeFlexible hours$73.5k - $212.28k
...At PwC, our people in data and analytics engineering focus on leveraging advanced technologies... ...Opportunity As part of the People Tech & AI team you will lead the design, build, and... ...closely with team members. We evaluate these factors thoughtfully to establish a...Full timeWork experience placementH1bRemote work$73.5k - $212.28k
...At PwC, our people in data and analytics engineering focus on leveraging advanced technologies... ...will lead the development of innovative AI solutions that drive remarkable client... ...collaborating closely with team members. We evaluate these factors thoughtfully to establish a...Full timeH1b$40 per hour
A tech-driven cybersecurity company is seeking experienced professionals to evaluate AI-generated security content and solve technical cybersecurity problems. This remote role allows flexible scheduling and project selection, paying $40+ per hour. Candidates should have...Hourly payRemote workFlexible hours- A leading AI training firm is seeking an Audit Defense Specialist to improve AI models related to healthcare. This remote position involves evaluating the logic and accuracy of AI chatbot responses while ensuring medical correctness. Applicants should possess a healthcare...Hourly payRemote workFlexible hours
- A leading AI research accelerator is seeking a Software Engineer with over 5 years of experience. The role involves evaluating AI-generated code and working with various teams to enhance coding solutions. Candidates must have strong full-stack development skills and excellent...Remote jobContract workFor contractors10 hours per weekFlexible hours
- A leading AI research firm is seeking a Software Engineer with over 5 years of experience to evaluate AI-generated code and collaborate on enhancing coding solutions. The role requires strong full-stack development skills and excellent communication abilities. This is a...Remote jobFor contractors10 hours per weekFlexible hours
- ...daily. One of our co-founders has led our AI and systems work to this point. This role... ...-Driven Improvement (approximately 70%) Evaluate how work gets done across retail, finance... ...FOR 3 to 8 years in a technology, data, engineering, operations, or systems role, or a less traditional...Immediate start
- ...redefining security operations with Agentic AI automation that empowers organizations to... ...are looking for a Principal AI Systems Engineer to act as a pathfinder and shape the next... ..., and workflow automation. Own Evaluation (Evals): Create test sets, define success...
$90k - $150k
...of our advanced software development and AI capabilities. Under direction, works collaboratively... ...-edge information technologies, data engineering, machine learning, cloud platforms, and... ...techniques to support project delivery Evaluate whether a use case is best solved through...Temporary workFlexible hours$52 - $56 per hour
...Immediate need for a talented AI/LLM Engineers. This is a 12+ Months contract opportunity with long-term potential and is in Denver, CO (Onsite). Please review the job description below and contact me ASAP if you are interested. Job Diva ID: 26-16106...Contract workLocal areaImmediate start$130k - $170k
...data and analytics ecosystem by embedding AI and Generative AI capabilities across... ...and Administrative systems. As the AI Lead Engineer – AWS Platform, you will play a key role... ...data governance, and security frameworks. Evaluate new AWS services (Amazon Q, Bedrock Agents...Full timeTemporary workWork at officeRemote workHome officeFlexible hours$110k - $150k
...Description Your role at GEI. The AI Engineer is responsible for the development of AI solutions, typically leveraging pretrained models and copilots, to support GEI's priority digital and AI initiatives. This role focuses on building, deploying...Work at officeFlexible hours$147k - $202k
...of logistics but building what comes next. Job Title: Lead AI Engineer Company: Prologis Title: Lead AI Engineer Location: Denver... ...-augmented generation, tool calling, workflow orchestration, evaluation, and human review. Build an AI-native building knowledge...Full timeWork at office- ...Head of Product and AI Engineering (CTO) About the Company Expanding market research technology company Industry Market... ...will also be instrumental in establishing processes for the evaluation, prioritization, and launch of new capabilities, and in ensuring...
$20 per hour
A leading AI training company is looking for independent contractors to help teach AI chatbots. This role involves developing prompts, writing responses, and evaluating AI outputs. The company offers a flexible remote work schedule, with pay starting at $20 per hour, increasing...Remote jobHourly payFor contractorsFlexible hours$104.8k - $152k
Generative AI Developer At Stantec, we have some of the world’s leading professionals... ...will work closely with data scientists, engineers, enterprise architects, and business stakeholders... ..., including model lifecycle management, evaluation, deployment, and monitoring. Partner...Full timeTemporary workPart timeLocal areaFlexible hours$60 per hour
A leading AI development firm is seeking experienced quantitative professionals for a fully remote position. You will evaluate AI-generated quantitative work and help shape future AI systems through technical accuracy and real-world validity. Ideal candidates have a quantitative...Hourly payRemote workFlexible hours$130k - $160k
...Job Description Job Description AI Engineer Location: Centennial, CO (In-Office) Company: Yield Solutions Group Reports... ...Establish internal standards for AI development, including model evaluation frameworks, validation protocols, and responsible deployment...Work at officeShift work$50 - $100 per hour
...DataAnnotation is seeking a Software Engineer to engage in innovative AI development. As part of a diverse coding team, you'll tackle programming challenges, interact with AI models, and contribute to evolving intelligent systems. This role offers flexibility to choose...Hourly payContract workRemote workWork from home$130k - $170k
...you opportunities to learn and grow? A position at Xcel Energy could be just what you’re looking for. Position Summary The Sr. AI Engineer plays a pivotal role in advancing Xcel’s AI vision by leading the technical delivery of innovative IT and AI-powered solutions that...Temporary workFor contractors
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Engineer. Be the first to apply!

