AI Evaluation Lead
$120k - $140kelly
Role Description
Our client is hiring an AI Evaluation Lead to own how we measure the quality of AI-generated financial advice. Getting the advice right matters. A bad output here has real consequences for real people, and this role owns making sure we catch it.
- You will work with an AI-generated test case library and automated scoring infrastructure that is already in place.
- Your job is to make sure we are measuring the right things, interpreting what the results are telling us, and determining what needs to change to keep the system performing well as it scales.
- This is not a monitoring and reporting role. It requires genuine judgment about AI system behavior, advice quality, and what the data is and is not capturing.
- You will report to our Head of Revenue & Compliance, work closely with the AI/ML team and founders, and partner with subject matter experts who provide domain judgment on complex or ambiguous cases.
- You need enough personal finance literacy to make first-pass quality assessments independently and know when to escalate.
Qualifications
- You have worked on AI or ML system quality in a context where outputs had real stakes.
- You think analytically about what data is and is not telling you.
- You are comfortable making judgment calls in ambiguous situations rather than waiting for the answer to be obvious.
- You have enough AI/ML fluency to reason about why a system is producing what it is producing, not just whether the output looks right.
- You bring enough personal finance literacy to read an advice response and have a genuine opinion about whether it is directionally sound.
- You do not need formal credentials or deep expertise across every domain the system covers—you will partner with subject matter experts for the complex judgment calls.
- Your review is substantive rather than mechanical, and you can have an informed conversation with those experts about what you are seeing in the data.
- Fluency with how LLM-based systems behave in production, including output variance, failure modes, and the limits of automated scoring.
- Ability to assess whether an eval framework is measuring the right things, not just whether it is running correctly.
- Comfortable working with behavioral and interaction data to surface patterns and quality signals.
- Familiarity with evaluation and observability tooling.
Requirements
- Model evaluation or QA on a consumer-facing AI product, particularly in a regulated or high-stakes context.
- Model risk or validation with LLM or generative AI exposure.
- Data science or analytics with ownership of production AI system quality.
- Operations quality control built around AI- or ML-generated outputs.
- Financial services or fintech product roles where you developed both analytical depth and personal finance domain familiarity.
Benefits
- Salary: 120-140k, plus early-stage option equity.
- Final compensation will depend on level, experience, location, and scope of responsibility.
- This role is open to candidates based in the United States.
How we work
- We are a fully remote, distributed team.
- Periodic in-person get-togethers will be integral to our operating cadence.
- We prioritize outcomes and output over set schedules.
- We value clear writing, high ownership, fast iteration, direct communication, and thoughtful async collaboration.
- As an early team member, you should expect broad ownership, frequent context shifts, and a high degree of autonomy.
- You will help shape not just the product, but also the technical standards and operating cadence of the company.
AI Interview
We expect a high volume of applications for this role. To help candidates showcase more than what's on their resume, you'll have the opportunity to complete an AI interview as part of the application process.
As an AI-first company, we embrace AI throughout the hiring process and are excited to meet candidates who are equally curious about and enthusiastic about the technology. This interview is your chance to demonstrate your experience, communication skills, and potential beyond your resume.
$50 - $60 per hour
A data-driven technology company is seeking an M&A Integration Manager to train AI models. This role involves evaluating AI chatbots and improving their logic and performance. Candidates should have strong financial reasoning skills, with proficiency in financial analysis...SuggestedRemote jobHourly payFull timePart timeFlexible hours- RWS Group Deutschland sucht eine/n Speech AI Evaluation Specialist (m/w/d) für Home Office. Zu den Aufgaben gehört das Führen kurzer Sprachdialoge mit KI-Modellen, das Durchlaufen von Szenarien und das Abgeben objektiver Bewertungen gemäß Vorgaben. Erforderlich sind Muttersprache...SuggestedRemote jobHome officeFlexible hours
$300k - $320k
About the role: We are seeking a Technical Program Manager to lead our AI model evaluation initiatives across multiple workstreams. This role will be crucial in assessing the performance, capabilities, limitations, and potential risks of our AI models. Working closely...SuggestedWork at officeHome officeVisa sponsorshipRelocation package$84 per hour
...documentation accuracy and coding integrity workflows through the evaluation of AI tools. As a Clinical Documentation Integrity (CDI) Leader, you... ..., compliance, and revenue integrity. Key Responsibilities Lead clinical documentation integrity programs for inpatient and/or...SuggestedRemote work$110 per hour
...will directly influence the development of AI tools designed to improve documentation... ...clinical accuracy. Key Responsibilities Lead risk adjustment and HCC coding operations... ..., and/or ACA risk adjustment programs. Evaluate AI-generated HCC coding assignments and risk...SuggestedHourly payRemote work$80 per hour
...management professionals can leverage their expertise in operations and inventory management to contribute to AI research projects. This role involves evaluating AI-generated content and providing insights that enhance AI''s understanding of retail store management...Contract workPart timeRemote workFlexible hours$85 per hour
...can leverage their expertise in asset management, maintenance planning, and grid operations to contribute to AI research projects. This role involves evaluating AI-generated content and providing critical feedback to enhance AI's understanding of maintenance workflows and...Contract workPart timeRemote workFlexible hours$120k - $140k
Role Description Hence is hiring an AI Evaluation Lead to own how we measure the quality of AI-generated financial advice. Getting the advice right matters. A bad output here has real consequences for real people, and this role owns making sure we catch it. You will...Full timeRemote workShift work$150k - $200k
Role Description The AI Quality & Evaluation lead enables responsible and scalable AI adoption by defining technical quality standards, evaluation framework and control requirements across the AI lifecycle. As a function owner of AI quality and robustness standards, the...Full timeFlexible hours$25 - $30 per hour
DataAnnotation is seeking a Lead Product Designer to train AI models by evaluating their outputs and providing critique on designs. Your expertise will enhance the next generation of AI tools and ensure high-quality models. This position offers the flexibility to choose...Remote work- DataAnnotation is seeking a Lead Product Designer to evaluate and improve AI models, particularly in UI/UX design. In this role, you'll review AI-generated visuals, and provide feedback to enhance the models' understanding of design principles. The position offers flexibility...Work from home
$25 - $40 per hour
DataAnnotation is seeking a Lead Product Designer to help train AI models in Idaho, United States. You will evaluate and critique AI-generated UI/UX designs, ensuring they meet visual and usability standards. Your insights will help shape AI tools to better support designers...For contractorsWork from home- ...processing (NLP) technologies. As a Vice President and Applied AI/ML Lead, you'll play a pivotal role in building innovative solutions... ..., and signals from unstructured text and documents. Define evaluation strategies and success metrics, including offline validation,...
$50 - $90 per hour
...project aimed at enhancing next-generation AI systems. Your expertise will be... ...threat landscapes. Design and validate evaluation frameworks for offensive security, focusing... ...engineer, exploit developer, cloud red-team lead, malware reverse-engineer, or security researcher...Remote jobHourly payFor contractors- Feitong Buke is hiring a Lead AI Trainer to oversee and enhance the quality of AI model dialogues with users. The role involves reviewing datasets for accuracy, providing feedback to annotators, and validating AI model outputs to ensure high production quality. Candidates...Remote jobFull time
- ...goal is to build the next generation of AI: autonomous agents that can reason, plan,... ...solve critical problems for an industry leading financial institution. We are looking for... ...delivery, shaping how applied AI is designed, evaluated, and deployed at scale. You will partner...
- ...handling more than 120 currencies, we are a leading processor of USD payments with daily... ...trillions. As a Vice President, Applied AI/ML Lead (Sr Level IC role) within JPMorgan... ...constraints. Define rigorous evaluation and measurement: offline metrics, calibration...
$95k - $115k
...AI Enablement Lead Department: Corporate Employment Type: Full Time Location: Chicago,IL Compensation: $95,000 - $115,000 /... ...operate • Stay current on AI developments and continuously evaluate new tools, models, and approaches that could benefit the organization...Permanent employmentFull timeWork at officeRemote workFlexible hours$190k - $230k
...PC company with a full‑stack portfolio of AI-enabled, AI-ready, and AI-optimized devices... ...hiring an AI User Experience Reliability Lead to define and drive the technical strategy... ...direction for how Qira’s intelligence is evaluated, monitored, and improved — ensuring users...Local areaRemote work- ...AI Enablement Lead The AI Enablement Lead plays a pivotal role in accelerating enterprise-wide AI adoption across non-engineering teams... ...Demonstrated ability to define measures of success and use data to evaluate outcomes and drive accountability. Preferred Skills:...Remote workFlexible hours
- ...will have a lasting impact on society. Job Summary: The AI Enablement Lead is responsible for driving the adoption and scaled delivery... ....#LI-SM2 Major Responsibilities: Lead intake, evaluation, and prioritization of AI initiatives across the GBU, aligned...Full timeWork experience placementWork at officeLocal areaRemote workRelocation
$20 per hour
A healthcare technology company is seeking a Medical Billing Manager to help train AI models. This role involves providing complex healthcare-related problems to AI chatbots, evaluating their outputs for accuracy, and ensuring high-quality responses. Candidates should...Hourly payFor contractorsRemote workFlexible hours$20 per hour
A technology company specializing in AI is seeking a Credentialing Manager to train AI models. The role requires diverse healthcare expertise and focuses on evaluating AI outputs for accuracy and performance. Responsibilities include solving complex healthcare-related...Hourly payRemote workFlexible hours- ...focused data company is seeking a Medical Billing Manager to train AI models. The role is remote, offering flexibility and hourly pay... ...degree. You will engage AI chatbots with healthcare problems, evaluate their performance, and ensure the accuracy of responses. This is...Hourly payFor contractorsRemote work
$94k - $110k
...and high-impact. Job Overview: We are hiring an AI Adoption Lead to transform how Airbel works by embedding AI into the... ...solutions (e.g., quality checks, human review expectations, evaluation approaches, and documentation requirements) and ensure new solutions...Work at officeLocal areaImmediate startRemote work- ...AI Enablement & Field Intelligence Lead (AEC) We're scaling rapidly and have a growing pipeline of opportunities that demand exceptional talent across... ...to AI-assisted estimating to reality capture — and can evaluate where each piece might actually fit in the real...Remote work
$141k - $307k
...develop algorithms and automated processes to evaluate large data sets from disparate sources.... .... What You’ll Do Design end-to-end AI solutions on Lam's shared AI platform — AI... ...platform capabilities are applied consistently. Lead solution and design reviews, and define and...Work at officeLocal areaRemote workFlexible hours2 days per week3 days per week1 day per week$20 per hour
A healthcare technology company is seeking a Medical Billing Manager to enhance AI models by providing complex healthcare-related problems, evaluating responses, and ensuring medical accuracy. Candidates should be fluent in English and hold a relevant degree in healthcare...Hourly payRemote work- ..., performing virtual testing, or training AI and autonomy for complex systems, we know... .... Job Overview The Strategic Partnerships Lead is responsible for building, managing, and... ...Strategic Partnership Development: Identify, evaluate, and establish strategic relationships...Permanent employmentFor contractorsRemote workHome officeFlexible hours
$20 per hour
...strong background in healthcare and be responsible for training AI models used in healthcare applications. Responsibilities include providing complex healthcare-related problems to AI chatbots, evaluating the outputs for accuracy, and ensuring the responses are medically...Hourly payRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Lead. Be the first to apply!





