Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Evaluation Lead

$120k - $140k
Full-time

fintentional.ai

Role Description

Hence is hiring an AI Evaluation Lead to own how we measure the quality of AI-generated financial advice. Getting the advice right matters. A bad output here has real consequences for real people, and this role owns making sure we catch it.

You will work with an AI-generated test case library and automated scoring infrastructure that is already in place. Your job is to:

  • Make sure we are measuring the right things.
  • Interpret what the results are telling us.
  • Determine what needs to change to keep the system performing well as it scales.

This is not a monitoring and reporting role. It requires genuine judgment about AI system behavior, advice quality, and what the data is and is not capturing. You will report to our Head of Revenue & Compliance, work closely with the AI/ML team and founders, and partner with subject matter experts who provide domain judgment on complex or ambiguous cases. You need enough personal finance literacy to make first-pass quality assessments independently and know when to escalate.

What You’ll Do (The Day-to-Day)

  • Define and validate the evaluation set: what cases we should be testing, whether coverage is sufficient across domains, and where the current framework has gaps.
  • Analyze scoring results to identify highest-frequency case types, patterns in what is performing well versus poorly, and anomalies that warrant closer review.
  • Assess whether current measures are detecting the right failure modes or whether new measures are needed.
  • Review flagged cases and make judgment calls on what the results mean and what should be done about them, drawing on both data and domain knowledge.
  • Own the criteria and calibration for when human review is triggered: defining what rises to that level, what does not, and ensuring the threshold stays well-calibrated as the platform scales.
  • Partner with subject matter experts on cases that require deeper domain judgment, and incorporate their input into evaluation design.
  • Ensure evaluation coverage keeps pace with new domain additions and model changes before they ship.
  • Translate findings into specific, actionable recommendations for the AI/ML team on what needs to change in the system.
  • Evolve the evaluation framework as the system grows, new domains are added, and user patterns shift.

Qualifications

  • You have worked on AI or ML system quality in a context where outputs had real stakes.
  • You think analytically about what data is and is not telling you.
  • You are comfortable making judgment calls in ambiguous situations rather than waiting for the answer to be obvious.
  • You have enough AI/ML fluency to reason about why a system is producing what it is producing, not just whether the output looks right.
  • You bring enough personal finance literacy to read an advice response and have a genuine opinion about whether it is directionally sound.
  • You do not need formal credentials or deep expertise across every domain the system covers.
  • Fluency with how LLM-based systems behave in production, including output variance, failure modes, and the limits of automated scoring.
  • Ability to assess whether an eval framework is measuring the right things, not just whether it is running correctly.
  • Comfortable working with behavioral and interaction data to surface patterns and quality signals.
  • Familiarity with evaluation and observability tooling.

Requirements

  • Model evaluation or QA on a consumer-facing AI product, particularly in a regulated or high-stakes context.
  • Model risk or validation with LLM or generative AI exposure.
  • Data science or analytics with ownership of production AI system quality.
  • Operations quality control built around AI- or ML-generated outputs.
  • Financial services or fintech product roles where you developed both analytical depth and personal finance domain familiarity.

This is probably not the right role for you if:

  • Your background is primarily in building models rather than evaluating what they produce.
  • Personal finance is entirely unfamiliar territory.
  • You are looking for a well-defined role with stable processes.
  • You default to manual review rather than thinking systematically about what should be automated and what requires human judgment.

How we work

We are a fully remote, distributed team. Periodic in-person get-togethers will be integral to our operating cadence. We’re adults who prioritize outcomes and output over set schedules. We value clear writing, high ownership, fast iteration, direct communication, and thoughtful async collaboration.

As an early team member, you should expect broad ownership, frequent context shifts, and a high degree of autonomy. You will help shape not just the product, but also the technical standards and operating cadence of the company.

Compensation

Salary: 120-140k, plus early-stage option equity. Final compensation will depend on level, experience, location, and scope of responsibility. This role is open to candidates based in the United States.

Equal Opportunity & Accommodations

Hence is proud to be an equal opportunity employer. We do not discriminate in hiring or any employment decision based on race, color, religion, national origin, age, sex (including pregnancy, childbirth, or related medical conditions), marital status, ancestry, physical or mental disability, genetic information, veteran status, gender identity or expression, sexual orientation, or other applicable legally protected characteristic. Hence is also committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or an accommodation due to a disability, please let your recruiter know.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI Evaluation Lead in Remote vacancy
  • $120k - $140k

    Role Description Our client is hiring an AI Evaluation Lead to own how we measure the quality of AI-generated financial advice. Getting the advice right matters. A bad output here has real consequences for real people, and this role owns making sure we catch it. ~You... 
    Suggested
    Full time
    Remote work
    Shift work

    elly

    Remote
    2 days ago
  • $97k - $114k

    Role Description At Fetch, we’re building AI and automation systems that make our work smarter, faster, and more scalable...  ...reliability, and measurable impact. As an Automation Lead, you’ll own automation and evaluation programs across AI Operations. You’ll translate business... 
    Suggested
    Full time
    Work at office
    Remote work
    Flexible hours

    Fetch

    Remote
    16 hours ago
  • $146.2k - $261.4k

     ...Position Description RAND's Center on AI, Security, and Technology (CAST), part of...  ...and policy analysis projects, and leading multidisciplinary teams of policy researchers...  ...scientists. Your team will build systems to evaluate how AI models perform across the full... 
    Suggested
    Fixed term contract
    Work experience placement
    Remote work
    Work from home

    RAND

    Sherman Oaks, CA
    3 days ago
  • $84 per hour

     ...Role Overview Drive and lead clinical documentation integrity efforts that shape next-generation AI tools for healthcare documentation and coding. In this role you will evaluate AI-generated CDI suggestions, guide physician query processes, and ensure documentation captures... 
    Suggested
    Hourly pay
    Remote work

    SaidGig

    United States
    16 hours ago
  • $100 per hour

     ...Role Overview Lead revenue cycle analytics and decision support to evaluate and shape AI tools that improve revenue cycle intelligence, financial reporting, and operational decision-making across the full patient revenue lifecycle. This role combines hands-on RCM analytics... 
    Suggested
    Hourly pay
    Remote work

    SaidGig

    United States
    16 hours ago
  • $100 - $150 per hour

     ...Role Overview Lead the clinical evaluation of AI systems built to support utilisation review, medical necessity determination, and care coordination. In this role you will apply hands-on UM and case management leadership experience to review AI recommendations, validate... 
    Hourly pay
    Remote work
    Night shift

    SaidGig

    United States
    16 hours ago
  • $136k - $190.08k

    Job Description: Merkle is seeking a Lead Enterprise Architect to provide senior technical...  ..., content, and journey orchestration. Evaluate integration approaches, architectural trade...  ...similar platforms. Familiarity with generative AI concepts, AI-enabled customer experience... 
    Permanent employment
    Full time
    Contract work
    Temporary work
    Work at office
    Local area
    Remote work
    Flexible hours
    2 days per week
    3 days per week

    dentsu

    New York, NY
    4 days ago
  • $117.23k - $195.39k

     ...apply now. We are currently seeking a Lead Enterprise Architect to join our team in Arlington...  ...cases and roadmaps, and help the client evaluate and integrate new technologies. This role...  .... We are one of the world's leading AI and digital infrastructure providers, with... 
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA Services

    Virginia
    3 days ago
  • $140k - $180k

    Role Description The AI Enablement & Governance – Security & Controls Lead enables secure, responsible, and scalable AI adoption by defining, implementing, and evaluating AI‑specific security and risk controls across the AI lifecycle. This role serves as a bridge between... 
    Full time
    Flexible hours

    Alight

    Remote
    3 days ago
  • $180k - $260k

    Role Description AI buyers have changed. From mid-market SaaS companies fine-tuning open...  ...to frontier AI labs running large-scale evaluations, the question is no longer “is AI useful”...  .... That is the opportunity as ML Lead. Rolling up directly to the VP of Commercial... 
    Full time
    For contractors
    Immediate start
    Flexible hours
    Shift work

    NewtonX

    Remote
    7 days ago
  • Role Description The AI Implementation Lead is the Data Team's hands-on builder, deploying AI solutions where Data Team work meets the business...  ...OpenAI, or comparable), with hands-on prompt engineering, evaluation, and iteration on production prompts ~Working knowledge... 
    Full time
    Immediate start
    Remote work

    Industrial Electric Manufacturing

    Remote
    1 day ago
  • $131k - $165k

     ...Commercial Scaled Intelligence (CSI) team is an AI-first team dedicated to delivering...  ...across the company. As an Ads AI Analytics Lead II, you will own the intelligence behind our...  ...search/vector services. ~Design and evaluate retrieval workflows (RAG) with existing services... 
    Full time
    Remote work
    Flexible hours

    Instacart

    Remote
    6 days ago
  •  ...an exciting opportunity for a highly experienced builder to run AI productivity and executive Power BI for Astrion. You will deploy...  ..., Dataverse, and custom pipelines where required. ~Identify, evaluate, pilot, and roll out the right AI tool for the right business function... 
    Full time
    Contract work

    Astrion

    Remote
    1 day ago
  • Role Description We're looking for an AI Analytics Lead — a strong product analyst / analytics lead who'll take ownership of analytics at...  ...tools for product, growth, or business teams. ~Experience evaluating the quality of AI-agent responses: hallucinations, guardrails... 
    Contract work
    Remote work
    Shift work

    App EWA (Lithium Lab)

    Remote
    3 days ago
  • $120k - $160k

     ...programmable, buyers increasingly need guidance on how to evaluate, architect, and operate payment infrastructure. Search and AI systems are becoming a critical part of that...  .... We’re hiring an SEO & AI Search Discovery Lead to help shape how the market discovers and... 
    Full time
    Shift work

    Modern Treasury

    Remote
    26 days ago
  • $161k - $176k

    Role Description As a Modeling Lead, you will work closely with the Head of Modeling to design...  ...analytics. You will also help leverage AI-enabled tooling to accelerate the...  ...human actuarial judgment in control. ~Evaluate new approaches pragmatically, prioritizing... 
    Full time
    Temporary work
    Flexible hours

    Security Benefit Business Services / Everly Life

    Remote
    4 days ago
  •  ...between GCO, the WE DO Corp IT Platforms team, and the Automation and AI COEs, fostering a collaborative and transparent working...  ...Technology Implementation: ~Partner closely with IT teams to design, evaluate, and implement technology solutions – including RPA, AI, machine... 
    Full time
    Remote work

    WTW

    Remote
    2 days ago
  • Role Description The AI Enablement Team is the catalyst for internal transformation and...  .... We are looking for an AI Enablement Lead who will architect, scale, and own our...  ...Infrastructure: Establish strict guardrails, evaluation frameworks, and monitoring tools to track... 
    Full time
    Remote work
    Flexible hours

    Actian Corporation

    Remote
    16 hours ago
  • Role Description We are looking for a hands-on AI Lead or Lead GenAI Engineer to design and implement real-world Generative AI solutions...  ...up and manage AI/ML infrastructure on AWS/Azure ~Implement evaluation frameworks to measure LLM performance Qualifications ~... 
    Full time

    SolGenie

    Remote
    16 hours ago
  •  ...be the go-to person for C-level and group leads on infrastructure matters. If you thrive...  ...resource allocation. ~Adapt the best of AI technology to the daily engineering and...  ...hiring efforts: sourcing, interviewing, and evaluating candidates to grow the team with top... 
    Full time
    Remote work
    Flexible hours

    Nethermind

    Remote
    6 days ago
  •  ...will help us transform and evolve in the AI era. This is not a traditional role with...  ...automations, agents, or workflows ~Capable of evaluating business value, calculating ROI, and...  ...rapid prototyping tools ~Experience leading change initiatives or managing technical... 
    Contract work
    Summer work
    Remote work

    INFUSE

    Remote
    5 days ago
  • Role Description Liquid Agency is looking for an AI Enablement Lead to help transform how work gets done across the organization. This role...  ...identify opportunities for AI-enabled workflow improvements. ~Evaluate, test, and refine new ways of working that increase... 
    Full time
    Work at office
    Remote work

    Liquid Agency

    Remote
    2 days ago
  • $160k - $170k

     ...Description Jorie Healthcare Partners is seeking an Enterprise AI Enablement Lead to drive the responsible adoption of Artificial Intelligence...  ...initiatives across business and technical departments. ~Evaluate business processes to identify opportunities for AI-assisted... 
    Full time
    Work at office

    Jorie AI

    Remote
    2 days ago
  •  ...handling more than 120 currencies, we are a leading processor of USD payments with daily...  ...trillions. As a Vice President, Applied AI/ML Lead (Sr Level IC role) within JPMorgan...  ...constraints. Define rigorous evaluation and measurement: offline metrics, calibration... 

    Next Frontier Capital

    Palo Alto, CA
    2 days ago
  •  ...Cloud Artificial Intelligence Security Lead This role at Arbitration Forums is as unique...  ...a crucial role in ensuring that our AI products and solutions uphold the highest...  ...execution of test plans and strategies for evaluating the compliance of AI systems, including... 
    Work experience placement
    Remote work
    Flexible hours

    Arbitration Forums

    United States
    2 days ago
  •  ...unlock liquidity for the world. Backed by leading investors like PayPal Ventures, Paradigm,...  ..., we’re building toward a future where AI is embedded into how we operate, not layered...  ...patterns Standardize how we build, evaluate, and scale AI solutions internally Ensure... 
    Work at office
    Remote work
    2 days per week

    B Capital

    San Francisco, CA
    2 days ago
  •  ...AI Enablement & Field Intelligence Lead (AEC) US Location: Remote (US only) + 60–80% travel to jobsites nationwide A Note Before You Read Further...  ...to AI-assisted estimating to reality capture — and can evaluate where each piece might actually fit in the real... 
    Remote work

    Human Agency

    United States
    2 days ago
  • $95k - $115k

     ...AI Enablement Lead Department: Corporate Employment Type: Full Time Location: Chicago, IL Compensation: $95,000 - $115,000 / year Description...  ..., and maintainable. Stay current on AI developments and evaluate new tools, models, and approaches that could benefit the organization... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Greenwood Project

    Chicago, IL
    3 days ago
  • $20 per hour

    A healthcare technology company is seeking a Medical Billing Manager to enhance AI models by providing complex healthcare-related problems, evaluating responses, and ensuring medical accuracy. Candidates should be fluent in English and hold a relevant degree in healthcare... 
    Hourly pay
    Remote work

    DataAnnotation

    Columbia, SC
    2 days ago
  • $20 per hour

     ...States is seeking a Medical Billing Manager to join their team. This role involves training AI models by providing them with complex healthcare-related challenges, evaluating AI outputs for accuracy and logic, and ensuring the quality of responses. Candidates should be... 
    Hourly pay
    Remote work

    DataAnnotation

    Salt Lake City, UT
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Evaluation Lead. Be the first to apply!