AI Evaluation Engineer
$161.6k - $200kCapital Rx
Judi Health is a health technology company providing benefit administration solutions to employers, unions, health plans, and government entities. Judi Health replaces fragmented, outdated systems with the industry's first Unified Claims Processing architecture, seamlessly consolidating pharmacy and medical benefit administration on a single, secure platform. By delivering true price transparency, eliminating unnecessary middleman fees, and leveraging advanced AI-powered care delivery, Judi Health helps clients achieve unprecedented operational efficiency and service levels.
At Judi Health, we're deploying the infrastructure our country needs to deliver the healthcare we all deserve. We are the intelligence platform powering benefits plans for millions of Americans and proudly leading the next generation of care. To learn more, visitHybrid 3 days (offices in NYC, Denver, CO and Charlotte, NC area)
Position Summary
As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and autonomous agents in production. This role bridges the gap between model development and realworld usage by translating ambiguous product goals into measurable quality targets.
We're looking for someone to lead evaluation end-to-end - from unit and integration testing to offline, online, and statistical evaluations of probabilistic systems. What we need is someone who can design and operate robust evaluation frameworks, partner with scientists and engineers, and ensure we can confidently answer questions like: "Did this change improve or degrade quality, safety, or user outcomes?"
What You'll Build
Evaluation & Quality Pipelines
- Build data evaluation pipelines that collect production conversations and agent interactions
- Reconstruct full sessions from traces, logs, recordings, and transcripts
- Apply labeling and scoring using human feedback signals (surveys, sentiment, outcomes) and automated evaluators (e.g., LLMasjudge)
Continuous Quality & Safety Benchmarking
- Own weekly and ondemand automated evaluation runs against staging and production
- Define benchmarks that track accuracy, reliability, and safetyrelated signals
- Produce trend dashboards that clearly answer: "Did this deploy change quality or risk?"
Unified Evaluation Framework
- Design and extend a standardized evaluation framework that supports multiple agent types and workflows
- Translate highlevel product expectations into concrete success criteria and metrics
- Ensure new agents and features can be evaluated consistently with minimal friction
Self Service Evaluation Tooling
- Build APIs and internal tools so data scientists and engineers can go from "interesting scenario" to "included in the eval suite" quickly
- Enable scenario curation, dataset management, and eval execution without deep infrastructure knowledge
Experiment Tracking & Visibility
- Provide shared visibility into prompt, model, and agent experiments
- Enable reproducibility and comparison across runs so teams can build on each other's work instead of operating in silos
Position Responsibilities:
Data Engineering
- Build and maintain ETL pipelines for heterogeneous data sources (traces, logs, transcripts, user feedback)
- Implement complex data stitching and session reconstruction logic
- Manage dataset versioning, provenance, and lifecycle
Platform & Observability
- Develop dashboards and monitoring tools for AI quality metrics
- Integrate evaluations into CI/CD pipelines for scheduled and gated runs
- Implement alerting on quality and safety signals, not just infrastructure health
AI / ML Evaluation Tooling
- Apply and extend LLMasjudge evaluation patterns
- Design metrics and scoring approaches suitable for stochastic, nondeterministic systems
- Use tools like LangSmith to track runs, traces, experiments, and evaluation results
Collaboration
- Partner closely with data science, engineering, and product teams
- Translate between research goals, product intent, and engineering constraints
- Help define what "good" looks like for AI behavior in production
- Advocate for strong developer experience and usability in the tools you build
Required Qualifications
- 4+ years of experience in data engineering, ML engineering, or software engineering
- Bachelor's or Master's degree in Computer Science, Machine Learning, or a related quantitative field
* Strong proficiency in Python
* Experience building and maintaining production data pipelines
* Strong SQL skills
* Experience working with at least one cloud platform (AWS preferred)
Nice to Haves
- Prior work on LLM or agent evaluation infrastructure
- Familiarity with designing metrics for safety, reliability, or quality in AI systems
- Experience with voice or callcenter data (audio, transcripts, sentiment)
- Experience with browser automation tools (e.g., Playwright) for endtoend evals
- Deep SQL expertise
All employees are responsible for adherence to the Judi Health Code of Conduct including the reporting of non-compliance. This position description is designed to be flexible, allowing management the opportunity to assign or reassign duties and responsibilities as needed to best meet organizational goals.
We provide equal employment opportunities to all employees and applicants for employment and prohibit discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, medical condition, genetic information, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.
By submitting an application, you agree to the retention of your personal data for consideration for a future position at Judi Health. More details about Judi Health's privacy practices can be found at
- ...practical knowledge of agile development methodologies and engineering best practices. As an AI Engineer , you’ll play an integral role using your... ...with global internal teams in an agile environment. Evaluate tech options, build PoCs, and review code and designs....SuggestedFull timeWork at office3 days per week
$130k - $180k
...$180k# Job Description We're currently seeking an **Agentic AI Engineer** to design, build, deploy, and maintain AI-powered agents and... ...DevOps practices - Prompt engineering and model tuning - Agent evaluation frameworks - Knowledge graph implementation - Familiarity...SuggestedFull time- ...Southeastern United States is looking to bring on its first AI engineer. This is a ground-floor opportunity to architect and own the... ...Science, Engineering, or related field Experience building and evaluating LLM systems in production Vector databases (Pinecone, Weaviate...SuggestedFull time
- ...efficiency. We specialize in leveraging advanced AI and analytics to create innovative... ...to address business challenges. • Build, evaluate, and deploy models, standardize code, and... ....Expertise in ML model development, data engineering, and software engineering principles....SuggestedFull timeTemporary workRelocation
$116k - $175k
...continuous professional development. The Prompt + Skills Engineer is the hands-on builder in Cherry Bekaert’s AI Center of Excellence — the person who writes the... ...gets done.Participates in use case intake, evaluating submitted ideas from across the Firm for technical...SuggestedFull timeWork experience placementLocal area- ...efficiency. We specialize in leveraging advanced AI and analytics to create innovative... ...to address business challenges. • Build, evaluate, and deploy models, standardize code, and... ...Expertise in ML model development, data engineering, and software engineering principles.Knowledge...Full timeTemporary workRelocation
$51.59 - $61.59 per hour
job summary: We are seeking a highly skilled AI Engineer with a Master's degree in Computer Science, Artificial Intelligence, or a... ...for agent continuity and long-running workflows. Implement evaluation pipelines, prompt engineering strategies, and guardrails to ensure...Hourly payContract workTemporary workWork experience placement- ...Job Title: Generative AI Engineer (Java/Python) Location: Charlotte, NC Job Summary We are seeking a skilled Generative AI... ...latest trends and advancements in Generative AI. Monitor, evaluate, and improve AI systems using appropriate metrics. Required...
$50.5k - $112.5k
...The Opportunity As an AI Engineer, you will be at the forefront of transforming raw data into actionable insights, enabling informed... ...generative AI techniques, including prompt engineering, LLM evaluation, and fine-tuning, to develop production-ready applications powered...Work experience placementH1bLocal area- ...AI Prompt Engineer Location: Chandler, AZ, Jersey City, NJ & Charlotte, NC FTE Job Description Objective of Role: We are seeking... ...Platform (Power Apps, Power Automate). Exposure to model evaluation, monitoring, and AI observability tools. Knowledge of AI...
$140k - $175k
...Title: Agentic AI Engineer Location/Schedule: Charlotte, NC with a hybrid weekly schedule (3 Days Onsite/2 Days Remote) Status:... ...audit logging, least-privilege access, prompt injection defense, evaluations, and human approval mechanisms for high-risk actions....Full timeRemote work3 days per week- ...sponsorship of an employment visa at this time. Job Specs – AI Engineering Lead About this role Wells fFargo is seeking a AI... ...· Guide prompt engineering, skill engineering, and model evaluation approaches including guardrails, observability, red teaming,...Work experience placementWork visa
- ...AI Engineer Technical Leadership & System Architecture Lead the architecture and development of LLM-driven applications, AI agents... ...prompt engineering, prompt routing, safety guardrails, and evaluation metrics. Data & Vector Search Engineering Build data pipelines...
$110.7k - $372.9k
...afford the care they need. Deloitte has a new AI-first effort, backed by $1B in committed... ...months, not into a lab. As an Agentic AI Engineer, you will design, build, and... ...information at the right time. Reliability, evaluation & safety • Implement observability and tracing...Local areaVisa sponsorship- ...thinking services company at the forefront of AI-native innovation. We partner with... ...next-generation, agent-powered workflows engineered to scale in real-world settings. Our... ...policy-based routing, tool invocation, evaluation harnesses, and lifecycle observability.Implement...Full timeWork experience placementLive inWork at officeLocal area
- ...America)Please review the following job description:The Senior AI Agentic Engineer is a senior production builder responsible for designing,... .../context management, retrieval-augmented generation (RAG), evaluation, observability, and governed deployment. The engineer works...Full timePart timeShift workDay shift
$150k - $220k
...application and origination processes around data and AI systems, and we’re hiring a Senior AI Engineer to join our Innovation team. In this role, you’ll... ...Experience with end-to-end model development, deployment, evaluation, and monitoring workflows, including LLM-based...Full timeLocal areaRemote workMonday to FridayFlexible hours$128k - $252.5k
...span from account executives and data scientists to AI strategists, machine learning specialists, and data engineers. SFL Scientific, a Deloitte Business, is looking... ...to executives, helping them understand and evaluate new and essential areas for AI investment and identify...Local areaVisa sponsorship$144.25k - $256.25k
...(if applicable) + benefitsJob Function: Engineering & ArchitectureSchedule: Full timeWorkplace... ...technology and data insights.At American Express, AI is reshaping the future of commerce and... ...initiatives from concept to production.Evaluate emerging models, techniques, and agentic...Work at officeVisa sponsorship3 days per week$96k - $181k
...TechnologyExperience: 4+ yearsAbout the RoleKeyBank is seeking a Lead AI Engineer to design, build, and modernize AI-powered conversational... ...to enhance conversational experiences.Develop, test, evaluate, and optimize prompts, workflows, and AI orchestration strategies...Full timeWork at officeRemote workFlexible hours- ...consider a career in Advisory.KPMG is currently seeking a Manager, AI Engineer to join our Advisory Services practice.Responsibilities:End-to... ...with emerging AI/ML technologies, frameworks, and tools, and evaluate their applicability for client needsAct with integrity,...H1bLocal area
- We Are:Accenture's Oracle Business Group AI Center of Excellence is one of the most... ...client conversation.You might come from engineering, consulting, product, or pre-sales — what... ..., including retrieval pipelines and evaluation frameworksAccelerate AI-assisted sales pursuits...Full timeWork experience placementLive inWork at officeLocal area
- ...America)Please review the following job description:• The Lead AI Security Engineer is a senior hands-on security engineer responsible for... ...red teaming support, misuse-case validation, model or prompt evaluation, or safety monitoring for AI-enabled systems.Experience in...Full timePart timeShift workDay shift
- ...the forefront of a new era in enterprise AI — one defined not by model capability... ...of frontier AI research and production engineering — investigating the foundational challenges... ...Design and execute rigorous benchmarking and evaluation methodologies scoped to production-...Full timeWork experience placementLive inWork at officeLocal areaRelocation
$126.82k - $149.2k
...discover what you excel at—all from Day One.Job DescriptionThe AI Red Team Lead Engineer leads the execution and evolution of offensive security... ...(APIs, plugins, agents, RAG pipelines)Training, evaluation, and inference pipelinesData ingestion, labeling, and governance...Full timeLocal area3 days per week- ...solutions while pioneering the next generation of AI-empowered software development. You’ll lead teams in adopting AI-first engineering practices—bringing clarity, structure, and... ...and codify repeatable practices such as evaluation harnesses, code-review agents, context...Temporary workLocal area
- ...Accenture is helping companies use generative AI and semantic layer to reinvent their... ...and Power Ecosystem.You Are:As a Knowledge Engineer, you formulate real-world problems into... ...needed by the specific problem, you design, evaluate, and maintain ontologies.As a significant...Full timeWork experience placementLive inWork at officeLocal area
- ...DoWe are seeking a hands-on Financial Services Consultant with AI experience to design, build, and scale GenAI-powered... ...semantic search, and knowledge-driven applicationsImplement prompt engineering, evaluation frameworks, and guardrailsPerform testing, validation, of...Full timeTemporary workWork experience placement
- ...On-site 4/1 in Charlotte, NC Our client seeks a Senior AI Engineer to lead IVR and conversational AI initiatives for large-scale... ...deploy LLM workflows including prompt engineering, RAG, and evaluation for voice and chat use cases. Lead development of agentic...Hourly payLocal area
- ...We are seeking a highly motivated and technically proficient AI Engineer to join our growing Data & Analytics team. In this role, you will... ...frameworks like LangChain, LlamaIndex, or Semantic Kernel. Evaluate and select appropriate foundation models (OpenAI, Anthropic,...Work experience placementWork at officeLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Engineer. Be the first to apply!



