Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Finance Expert - AI Evaluation Specialist

$60 - $100 per hour

Mercor

Job Description

Job Description

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .

Position: Finance Domain Expert — AI Training & Evaluation
Type: Contract
Compensation: $60–$100/hour
Location: Hybrid, Bay Area, California
Commitment: 40 hours/week

Role Responsibilities

  • Evaluate the quality of finance knowledge work tasks and AI model outputs . Identify missing behaviors, flawed assumptions, and ensure professional scrutiny.
  • Design high-quality instruction specs and produce golden solutions to financial problems. Define new finance tasks reflecting real-world practices.
  • Develop challenging finance tasks and evaluation sets. Collaborate with research teams to build finance-specific skills and tools .
  • Collaborate with client researchers and specialists to maintain consistent standards. Translate tacit financial judgment into explicit, teachable criteria.

Qualifications

Must-Have

  • 5+ years of professional finance experience at a recognized institution.
  • Specialization in a core finance discipline such as corporate finance, investment banking, or quantitative finance.
  • Seniority at a leadership level with real ownership of analysis and decisions.
  • An advanced degree or recognized professional credential such as CFA , CPA , or FRM .
  • Hands-on experience with large language models in professional work.
  • Ability to commit to 40 hours/week for an initial engagement of 6 months.
  • Reside in the Bay Area, California , and work on-site as required.

Compensation & Legal

  • Hourly contractor
  • W-2 employment

Application Process (Takes 20–30 mins to complete)

  • Upload resume
  • AI interview based on your resume
  • Submit form

Resources & Support

PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.

Vacancy posted 10 days ago
Similar jobs that could be interesting for youBased on the Finance Expert - AI Evaluation Specialist in San Francisco, CA vacancy
  • Cincinnatus LLC is recruiting for a finance SME to join a leading AI lab's GenAI team, evaluating model outputs against rubrics and guiding financial judgment in AI training data. This is a W-2 employment placement with the option to be placed at a premier AI Lab as part... 
    Suggested
    Weekday work

    Obsidian

    San Francisco, CA
    19 hours ago
  • $1,750 - $2,150 per month

    Obsidian is looking for experienced cybersecurity professionals to review AI systems' threat detection and vulnerability assessments. Responsibilities include evaluating AI outputs and creating realistic cybersecurity scenarios. Ideal candidates should have over 3 years... 
    Suggested

    Obsidian

    San Francisco, CA
    1 day ago
  • $20 - $26 per hour

    Prolific is seeking fluent Kannada speakers to act as evaluators who compare text and voice samples to assess naturalness and authenticity. You will listen to audio clips, rate quality, and flag any mismatches in tone or pronunciation, with emphasis on cultural context... 
    Suggested
    Remote job
    Flexible hours

    Prolific

    San Francisco, CA
    1 day ago
  • $50 - $90 per hour

     ...and technical talent with leading AI research labs. Headquartered in San...  ...and Jack Dorsey . Position: Finance Domain Expert — AI Training & Evaluation Type: Contract...  ...Collaborate with client researchers and specialists to maintain consistent standards.... 
    Suggested
    Contract work
    Summer work

    Mercor

    San Francisco, CA
    2 days ago
  • Welo Data in San Francisco seeks a full-time AI Evaluator with professional proficiency in Portuguese (Portugal) and experience in Generative AI safety. The role involves critiquing AI outputs, identifying biases, and refining evaluation frameworks. Candidates should possess... 
    Suggested
    Full time

    Welo Data

    San Francisco, CA
    2 days ago
  • Welo Data is looking for a Data Labeling Associate in San Francisco to evaluate AI systems' handling of Arabic nuances. The role requires professional-level proficiency in Arabic and 2 years of AI safety experience. Responsibilities include critiquing Arabic AI outputs... 

    Welo Data

    San Francisco, CA
    19 hours ago
  • A forward-thinking tech company is seeking an AI Trainer specializing in visual and graphic design to evaluate AI outputs. Responsibilities include assessing design quality and providing feedback to ensure high professional standards. Applicants should have formal education... 
    Remote job
    Flexible hours

    Prolific

    San Francisco, CA
    4 days ago
  • $150k - $250k

    About Distyl AI Distyl is an applied AI technology company partnering with the world’...  ...ForAt Distyl, we build AI systems using Evaluation-Driven Development—an approach where evaluation...  ..., aligning automated judgments with expert human assessments. They investigate where... 
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    4 days ago
  • $240k - $280k

    A leading software monitoring company is seeking a Senior Software Engineer on its AI/ML team to build evaluation infrastructure for measuring the performance of AI systems. This role involves designing datasets, creating benchmarks, and ensuring AI features behave reliably... 

    Sentry

    San Francisco, CA
    3 days ago
  • Cincinnatus LLC is seeking a senior Insurance SME to join a leading AI lab's GenAI team in San Francisco. You will guide underwriting-focused evaluation of AI model outputs, develop scoring rubrics, and work with cross-functional teams to ensure high-quality training data... 

    Mercor

    San Francisco, CA
    3 days ago
  • Cincinnatus LLC is seeking an Insurance SME to join a leading AI lab's GenAI team in San Francisco. You will evaluate AI outputs against rubrics and contribute real-world underwriting judgment to training data for foundational AI models. The role requires 8+ years in insurance... 
    Weekday work

    Obsidian

    San Francisco, CA
    2 days ago
  • Synthires is seeking experienced Legal Experts in the United States to contribute to advanced AI research and evaluation projects. You will evaluate AI-generated legal content, apply real-world legal reasoning, and help improve how next-generation AI systems analyze employment... 
    Remote job

    Synthires

    San Francisco, CA
    1 day ago
  • Obsidian is partnering with an AI research initiative to support a Frontier Code Agents project in San Francisco. You will evaluate frontier AI coding models by performing realistic data engineering tasks, reviewing ETL pipelines, data warehouses, and distributed systems... 

    Obsidian

    San Francisco, CA
    1 day ago
  • $80 - $120 per hour

     ...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our...  ..., Larry Summers , and Jack Dorsey . Position: Finance operations / audit support Evaluator Type: Contract Compensation: $80–$120/hour... 
    Contract work
    Summer work
    Work at office
    Remote work

    Mercor

    San Francisco, CA
    18 days ago
  •  ...people think about and interact with personal finance.We’re a next-generation financial...  ...financial world.The role:SoFi's Associate AI Engineer, Finance Transformation is a hands...  ...stakeholders, with support.Nice to have:Exposure to evaluation, tracing, or observability practices for... 
    Internship
    Remote work

    SoFi

    San Francisco, CA
    1 day ago
  • $130k - $220k

     ...Artificial Analysis is the leading independent AI benchmarking and insights company. They...  ...** This role is best described as an AI Evaluation Engineer / Technical Generalist. It is...  ...success bar for this role is becoming a world expert in modern AI technologies. This is a... 
    Full time
    Worldwide

    Aurora Jobs ApS

    San Francisco, CA
    a month ago
  • Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. The work requires evaluating AI-generated music in Hebrew and English to assess quality across genres and styles. You will rate... 
    Immediate start

    Obsidian

    San Francisco, CA
    2 days ago
  • A leading AI research organization in San Francisco is seeking a Research Engineer focused on pushing the boundaries of frontier models...  ...offers the opportunity to shape innovative AI measurements and build evaluation environments that drive progress. #J-18808-Ljbffr OpenAI

    OpenAI

    San Francisco, CA
    4 days ago
  • B Capital seeks a talented individual for an AI Evaluation role in San Francisco. This position involves conducting critical comparative analysis, refining evaluation systems, and collaborating with various teams to enhance model capabilities. The ideal candidate will have... 

    B Capital

    San Francisco, CA
    3 days ago
  •  ...San Francisco is seeking an innovative Quality Engineer for their AI products. This role blends ops, strategy, and analytics to...  ...in leading labs, and ensure user satisfaction through effective evaluation baselines. Competitive salary and benefits offered, with a focus... 

    Notion

    San Francisco, CA
    3 days ago
  • Scale Labs seeks a Research Scientist focused on Frontier Risk Evaluations to design evaluation measures, harnesses and datasets for measuring risks posed by frontier AI systems. You will build harnesses to test models, collaborate with government agencies to scope evaluations... 

    Scale

    San Francisco, CA
    2 days ago
  • OpenAI is seeking a researcher to advance frontier evaluations and environments for safe AGI/ASI. You will help design north star model environments and steer major training runs so that research outputs translate into real-world products. Collaborate with researchers,... 

    Neura Market

    San Francisco, CA
    1 day ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering with leading AI labs. You will author original, executable research problems that frontier models cannot solve. Engagement lasts... 
    Part time
    Immediate start

    Mercor

    San Francisco, CA
    3 days ago
  • $216k - $270k

    Scale AI, Inc. is looking for a Research Scientist specializing in Frontier Risk Evaluations to develop measures for assessing risks of advanced AI systems. In this role, you will design testing harnesses, collaborate with agencies, and publish reports to inform policymakers... 

    Scale AI, Inc.

    San Francisco, CA
    4 days ago
  •  ...tech company in New York is seeking a Research Scientist focused on Frontier Risk Evaluations. The ideal candidate will contribute to designing and creating evaluation measures for assessing AI risks and will have strong experience in machine learning and technical... 

    Scale AI

    San Francisco, CA
    4 days ago
  • HeyMilo AI is hiring a Research Engineer to join our Applied AI team in San Francisco. You’ll design and build reinforcement learning...  ...-world use cases, creating simulators, reward functions, and evaluation harnesses to measure model performance. The role blends ML research... 

    HeyMilo AI

    San Francisco, CA
    4 days ago
  • Anthropic in San Francisco seeks a bio safety researcher to design and run capability evaluations for biology-focused models, build and curate datasets for safety classifiers, and iterate on those classifiers with ML engineers. You will work at the intersection of applied... 

    Anthropic

    San Francisco, CA
    3 days ago
  •  ...hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. You will assess...  .... Start date: immediate. Duration: up to 6 months. Most experts work around 20 hours per week; there is no cap. We... 
    Immediate start

    Obsidian

    San Francisco, CA
    4 days ago
  • $400 per month

    About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic infrastructure engineering... 

    Mercor Inc

    San Francisco, CA
    3 days ago
  • $150k

     ...engineer to enhance their machine intelligence systems in San Francisco. As part of the team, you'll be responsible for building evaluation infrastructure, designing data pipelines, and implementing fine-tuning processes. Ideal candidates have expertise in Python and machine... 

    Tzafon

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Finance Expert - AI Evaluation Specialist. Be the first to apply!