Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff AI Evaluation Lead

$129k - $152k

Fetch

About the Role: At Fetch, we’re building AI and automation systems that make our work smarter, faster, and more scalable. The AI Operations team ensures our models, automations, and LLM systems perform with quality, reliability, and measurable impact. As a Staff AI Evaluation Lead, you’ll own automation and evaluation programs across AI Operations. You’ll translate business and functional goals into scalable systems, define how we measure quality, and ensure automation and evaluation become durable, high-impact capabilities across the organization. This role is ideal for someone who combines deep technical problem-solving with systems-level thinking and strong cross-functional leadership. This is a full-time role that can be held from one of our US offices or remotely in the United States. Role Responsibilities: Own programs: Lead complex, high-impact automation and evaluation initiatives across workflows or teams Design scalable solutions: Architect end-to-end workflows integrating datasets, evaluations, automations, and HITL processes; build the harness and reusable components other analysts build against; build evaluation pipelines that run against production without hands-on operation Establish standards: Define dataset standards, evaluation methodology, failure taxonomy, and quality measurement across AI Operations; verify that projects you don't run are meeting them; own the process by which those standards get made and used Own metric definitions: Create and maintain the metric definitions the org evaluates against, validate them against real production behavior, and revise them as models, tooling, and system architectures change Set the bar for production: Define the quality bar a system clears before it reaches production and where human review stays in the loop, and pull a system back when it stops meeting the bar Allocate evaluation depth by risk: Decide which systems get what depth of evaluation given finite capacity, name the risk accepted on the rest, and make that tradeoff visible to project stakeholders Keep the evaluation stack current: Define when a change to a model or platform requires re-baselining across the org and own the process for doing it; evaluate new models and AI capabilities as they ship and decide what the org adopts Improve systems at scale: Lead redesign of workflows, tooling, and processes to improve performance and durability across AI Operations, not only within your own programs Drive cross-functional alignment: Influence priorities and partner with Engineering, Product, and AI teams to deliver solutions Measure and communicate impact: Define success metrics and communicate performance and recommendations to leadership Elevate the team: Raise the bar through your work and actively mentor others by sharing approaches, guiding problem-solving, and enabling the team to build stronger automation and evaluation capabilities Drive innovation: Stay current on emerging tools and approaches; pilot and translate them into actionable improvements for the team Minimum Requirements: 8+ years of professional experience AI, machine learning, operational automations, or a related field. Proven ability to lead complex automation or evaluation initiatives across systems or teams Experience designing, building, and scaling evaluation frameworks, datasets, and quality systems at scale for LLMs or AI products Fluency in SQL, JSON, APIs, and scripting with AI assistance, and ownership of the technical direction of the evaluation stack Deep working knowledge of LLM and agentic system behavior, automation platforms, and system design, demonstrated in systems you have built Experience with data pipelines, APIs, and production systems Experience in influencing cross-functional stakeholders and aligning priorities Demonstrated ability to define metrics and drive measurable business impact Preferred Requirements: Experience mentoring or leading technical contributors Experience evaluating agentic systems in production at scale Experience setting technical standards adopted across an organization Compensation: At Fetch, we offer competitive compensation packages including base, equity, and benefits to the exceptional folks we hire. The base salary range for this position is $129,000 - $152,000. Discover our benefits and how our employees live rewarded at #J-18808-Ljbffr Fetch

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Staff AI Evaluation Lead in Brooklyn, NY vacancy
  • $142.32k - $213.48k

     ...Infrastructure, ProfessionalCompany: CitiAbout Citi:Citi, the leading global bank, has approximately 200 million...  ...enable growth and progress together.The Role:The AI Partnerships Lead (Technology Vendor & Start-up Evaluation) sits within Citi's Head of AI Organization and... 
    Suggested
    Full time
    Work at office

    Citigroup

    Jersey City, NJ
    5 days ago
  • Uncover is building a central evaluation framework to measure model quality across teams. You will design reusable pipelines, benchmarks, and baselines, and build visualization tools to turn results into actionable insights for stakeholders. The role demands strong software... 
    Suggested

    Uncover

    Brooklyn, NY
    5 days ago
  • Literally Media in New York is launching a new YouTube channel focused on AI and its impact on women's lives and careers. We’re seeking a creator who can shape the editorial direction, identify the stories that matter, and turn complex developments into compelling, potentially... 
    Suggested

    44 Ventures

    Brooklyn, NY
    5 days ago
  • Virginia Tech National Security Institute (VTNSI) seeks a senior research faculty member to lead AI Test & Evaluation efforts, coordinating rapid development, prototyping, and transition of AI/ML capabilities for national security. The role blends technical leadership with... 
    Suggested

    Southwest Virginia Higher Education Center

    Brooklyn, NY
    5 days ago
  • $147k - $200k

    Lead Maritime Test and Evaluation Director Arlington, VA About the Team: The Sensors Division within STR focuses on the development and analysis of...  ...of software-intensive system testing, including autonomy, AI/ML model validation, sensor fusion, or similar disciplines... 
    Suggested
    Full time
    Local area
    Night shift

    STR

    Brooklyn, NY
    5 days ago
  •  ...Artificial Intelligence and Machine Learning Lead at JPMorganChase within Markets...  ...delivery of machine learning and generative AI solutions that measurably improve Markets...  ...enforce best practices for model monitoring, evaluation, and performance optimization in... 

    JP Morgan Chase

    Jersey City, NJ
    1 day ago
  •  ...help shape how analytics and generative AI are delivered at enterprise scale—turning...  ...learning.As an Applied AI and Machine Learning Lead at JPMorganChase within Corporate...  ...pipelines and frameworks for model training, evaluation, optimization, monitoring, and production... 
    Work at office

    JP Morgan Chase

    Jersey City, NJ
    4 days ago
  •  ...peak times. Leadership & Team Supervision: Lead by example, providing guidance and support to junior bartenders and bar staff,maintaininga positive work environment. Train...  ...expectations, conduct regular performance evaluations, and offer feedback to team members to... 
    Night shift

    Proper Hospitality

    Brooklyn, NY
    2 days ago
  • $15.5 - $25.5 per hour

     ...associates in the Starbucks Department. Is responsible for assisting the Department Manager with the overall direction, coordination, and evaluation of this department. Carry out supervisory responsibilities in accordance with Harris Teeter's policies and standards.... 
    Full time
    Local area
    Flexible hours

    Harris Teeter

    Brooklyn, NY
    2 days ago
  •  ...The Role AI is changing what it means to deploy enterprise software. Configuring a...  ...s APAC entity. As Omnea's AI Solutions Lead, you own deployment of the Omnea platform...  ...to write clear, specific instructions and evaluate what comes back ~ Client-facing experience... 
    Contract work
    Shift work
    Day shift

    Omnea

    Astoria, NY
    4 days ago
  •  ...Clearance Required: Active Public Trust Job Description Summary The AI Lead will set the technical direction and lead the delivery of...  ...and oversee MLOps and LLMOps practices for automated testing, evaluation, release management, observability, traceability,... 
    Temporary work
    Flexible hours

    Dovel Technologies

    Brooklyn, NY
    4 days ago
  • $123.21k - $205.35k

     ...are currently seeking a MUMPS Developer Lead to join our team in Jersey City, New Jersey...  ...related to Profile upgrades, including evaluation of approximately 25 years of customized application...  .... We are one of the world's leading AI and digital infrastructure providers,... 
    Full time
    Temporary work
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Jersey City, NJ
    3 days ago
  •  ...objectives.As the Applied ML and Generative Lead within J.P.Morgan, you will operate as a...  ...production-grade ML and Generative AI services, while setting technical direction...  ...summarization and text generation.Conduct thorough evaluations of generative models (e.g., GPT-4.1),... 

    JP Morgan Chase

    Jersey City, NJ
    3 days ago
  • Appalachian Regional Healthcare (ARH) is seeking a Staff Physical Therapist to conduct and supervise physical therapy programs that...  ...disability, and empower families to participate in home care. You will evaluate patients, provide therapy, supervise staff in the absence of... 

    Appalachian Regional Healthcare (ARH)

    Brooklyn, NY
    6 days ago
  •  ...execute projects that require Generative AI and machine learning development to support...  ...the bank operates.  As an Applied AI/ML Lead, you will apply sophisticated machine learning...  ...training frameworks, and to outline and evaluate intrinsic and extrinsic metrics for model... 
    Work at office

    JPMorgan Chase & Co.

    Jersey City, NJ
    2 days ago
  • $61k - $101k

     ...deploying machine learning and generative AI solutions with Python Proven ability...  ...Experience implementing monitoring and evaluation approaches for machine learning and generative...  ..., and overall workflow performance Lead semantic modeling strategy, including ontology... 
    Full time
    Work at office

    J.P. Morgan

    Jersey City, NJ
    1 day ago
  •  ...and staying active and healthy. How best to support her: Stay calm and respectful Ask permission Validate feelings A successful lead staff will have strong boundaries, experience with mental health and conflict resolution, the ability to not take things personally, and... 
    Hourly pay
    Full time
    Work at office
    Immediate start
    Relocation
    Monday to Friday
    Flexible hours
    Shift work
    Night shift
    Weekend work

    Harbor SLS

    Brooklyn, NY
    5 days ago
  • Harmony Biosciences is recruiting for an Associate Director, Business Development (Search, Evaluation & Analytics) in Chicago. You will build the analytical foundation for assessing new opportunities, translating data into strategic insights to inform BD prioritization... 

    Paragon Biosciences LLC

    Brooklyn, NY
    3 days ago
  • Anthropic is seeking a Strategic Deals Lead in San Francisco to shape infrastructure strategy, evaluate compute options, and structure multi-layered silicon transactions...  ..., and strategic leadership in a fast-moving AI ecosystem. Hybrid policy applies at select offices... 

    Uncover

    Brooklyn, NY
    2 days ago
  •  ...Industries (founded by Jimmy Donaldson, MrBeast) seeks its first AI Enablement Lead to drive how AI is adopted across the company, co-building...  ...across the stack: prompt engineering, data architecture, evaluation, deployment Strong product and systems thinking Excellent communication... 

    Workman Labs

    Brooklyn, NY
    3 days ago
  •  ...work. Why this role is on the menu Most companies are layering AI onto existing processes. This role is about something different:...  ...teams Brief quality and standardization — an AI review layer that evaluates briefs before they leave the building, flags what's missing or... 
    Work at office
    Work from home
    Flexible hours

    EngineersOfAI

    Brooklyn, NY
    3 days ago
  • $157k - $184k

     ...any of our customers in the dark. So, join us as a Lead Solution Architect and find your superpower. We...  ...sponsorship. National Grid utilizes an assessment that evaluates the job qualifications/characteristics using AI or statistically based scoring. For more... 
    Local area
    Flexible hours

    National Grid USA

    Brooklyn, NY
    5 days ago
  • $61k - $101k

     ...demonstrated experience using enterprise-approved AI capabilities in architecture and...  ...data sensitivity. You should be able to evaluate and integrate AI-enabled capabilities into...  ...work aligned to business outcomes. We lead and contribute to architecture governance... 
    Full time

    J.P. Morgan

    Jersey City, NJ
    6 days ago
  • OpenAI’s Enablement Lead, Builder role is a post‑sales technical enablement specialist...  ...trainings on OpenAI APIs, Codex, agents, evaluations, and related capabilities. You will work...  ...audiences from hands-on builders to executives, driving #J-18808-Ljbffr AI Chopping Block

    AI Chopping Block

    Brooklyn, NY
    4 days ago
  • $105.79k - $141.05k

    Lumen is the trusted network for the AI‑powered world, connecting people, data, and applications...  ...quickly, securely, and effortlessly. As a Lead IT Systems Analyst (Software Asset...  ...maximize return on software investments. Evaluate vendor proposals, licensing models, and contract... 
    Contract work
    Temporary work
    Work at office
    Remote work

    Lumen Technologies

    Brooklyn, NY
    6 days ago
  • Ripple in San Francisco, CA seeks a Senior Staff Technical Program Manager for Information Security to lead the Post-Quantum Cryptography (PQC) migration program. You will work with InfoSec, XRPL Core, Payments, Custody, Stablecoin, Platform, Compliance, and Legal to ensure... 

    Ripple

    Brooklyn, NY
    6 days ago
  • Optima Care Fountains in New Jersey is seeking an experienced Infection Preventionist / Staff Development Coordinator to support regulatory compliance, staff competency, and resident safety. This dual-role position is responsible for assessing staff educational needs; planning... 

    Optima Care Castle Hill

    Brooklyn, NY
    2 days ago
  • $15.68 - $23.51 per hour

     ...Leads environmental services associates in performing a variety of environmental services to maintain assigned areas in a clean and orderly condition. MINIMUM QUALIFICATIONS: One (1) year of housekeeping experience. PREFERRED QUALIFICATIONS: High school diploma or equivalent... 
    Local area
    Shift work

    WVU Medicine

    Brooklyn, NY
    1 day ago
  • $73.15k - $101.01k

     ...Commerce Exchange. Overview The Lead Business Systems Analyst will...  ...as a liaison between client staff and internal technical delivery...  ...Able to assemble, analyze and evaluate data and be able to make appropriate...  ...of the SDLC Familiarity with AI and generative AI tools (e.g.,... 
    Full time
    Temporary work
    Freelance
    Work at office
    Local area
    Remote work
    Worldwide
    Work visa
    Flexible hours
    Shift work

    Publicis Groupe Holdings B.V

    Brooklyn, NY
    1 day ago
  •  ...The AI Governance Lead will serve as the day‑to‑day driver of an enterprise AI governance implementation programme, co‑leading delivery with...  ..., and controls Establish and operationalise AI intake, risk evaluation, approval, and monitoring workflows Translate regulatory... 

    Siri InfoSolutions

    Jersey City, NJ
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff AI Evaluation Lead. Be the first to apply!