Finance Expert - AI Evaluation Specialist
$60 - $100 per hourMercor
Job Description
Job Description
About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .
Position: Finance Domain Expert — AI Training & Evaluation
Type: Contract
Compensation: $60–$100/hour
Location: Hybrid, Bay Area, California
Commitment: 40 hours/week
Role Responsibilities
- Evaluate the quality of finance knowledge work tasks and AI model outputs . Identify missing behaviors, flawed assumptions, and ensure professional scrutiny.
- Design high-quality instruction specs and produce golden solutions to financial problems. Define new finance tasks reflecting real-world practices.
- Develop challenging finance tasks and evaluation sets. Collaborate with research teams to build finance-specific skills and tools .
- Collaborate with client researchers and specialists to maintain consistent standards. Translate tacit financial judgment into explicit, teachable criteria.
Qualifications
Must-Have
- 5+ years of professional finance experience at a recognized institution.
- Specialization in a core finance discipline such as corporate finance, investment banking, or quantitative finance.
- Seniority at a leadership level with real ownership of analysis and decisions.
- An advanced degree or recognized professional credential such as CFA , CPA , or FRM .
- Hands-on experience with large language models in professional work.
- Ability to commit to 40 hours/week for an initial engagement of 6 months.
- Reside in the Bay Area, California , and work on-site as required.
Compensation & Legal
- Hourly contractor
- W-2 employment
Application Process (Takes 20–30 mins to complete)
- Upload resume
- AI interview based on your resume
- Submit form
Resources & Support
- For details about the interview process and platform information, please check:
- For any help or support, reach out to: View email address on ziprecruiter.com
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
- Cincinnatus LLC is recruiting for a finance SME to join a leading AI lab's GenAI team, evaluating model outputs against rubrics and guiding financial judgment in AI training data. This is a W-2 employment placement with the option to be placed at a premier AI Lab as part...SuggestedWeekday work
$1,750 - $2,150 per month
Obsidian is looking for experienced cybersecurity professionals to review AI systems' threat detection and vulnerability assessments. Responsibilities include evaluating AI outputs and creating realistic cybersecurity scenarios. Ideal candidates should have over 3 years...Suggested$20 - $26 per hour
Prolific is seeking fluent Kannada speakers to act as evaluators who compare text and voice samples to assess naturalness and authenticity. You will listen to audio clips, rate quality, and flag any mismatches in tone or pronunciation, with emphasis on cultural context...SuggestedRemote jobFlexible hours$50 - $90 per hour
...and technical talent with leading AI research labs. Headquartered in San... ...and Jack Dorsey . Position: Finance Domain Expert — AI Training & Evaluation Type: Contract... ...Collaborate with client researchers and specialists to maintain consistent standards....SuggestedContract workSummer work- Welo Data in San Francisco seeks a full-time AI Evaluator with professional proficiency in Portuguese (Portugal) and experience in Generative AI safety. The role involves critiquing AI outputs, identifying biases, and refining evaluation frameworks. Candidates should possess...SuggestedFull time
- Welo Data is looking for a Data Labeling Associate in San Francisco to evaluate AI systems' handling of Arabic nuances. The role requires professional-level proficiency in Arabic and 2 years of AI safety experience. Responsibilities include critiquing Arabic AI outputs...
- A forward-thinking tech company is seeking an AI Trainer specializing in visual and graphic design to evaluate AI outputs. Responsibilities include assessing design quality and providing feedback to ensure high professional standards. Applicants should have formal education...Remote jobFlexible hours
$150k - $250k
About Distyl AI Distyl is an applied AI technology company partnering with the world’... ...ForAt Distyl, we build AI systems using Evaluation-Driven Development—an approach where evaluation... ..., aligning automated judgments with expert human assessments. They investigate where...Work at office3 days per week$240k - $280k
A leading software monitoring company is seeking a Senior Software Engineer on its AI/ML team to build evaluation infrastructure for measuring the performance of AI systems. This role involves designing datasets, creating benchmarks, and ensuring AI features behave reliably...- Cincinnatus LLC is seeking a senior Insurance SME to join a leading AI lab's GenAI team in San Francisco. You will guide underwriting-focused evaluation of AI model outputs, develop scoring rubrics, and work with cross-functional teams to ensure high-quality training data...
- Cincinnatus LLC is seeking an Insurance SME to join a leading AI lab's GenAI team in San Francisco. You will evaluate AI outputs against rubrics and contribute real-world underwriting judgment to training data for foundational AI models. The role requires 8+ years in insurance...Weekday work
- Synthires is seeking experienced Legal Experts in the United States to contribute to advanced AI research and evaluation projects. You will evaluate AI-generated legal content, apply real-world legal reasoning, and help improve how next-generation AI systems analyze employment...Remote job
- Obsidian is partnering with an AI research initiative to support a Frontier Code Agents project in San Francisco. You will evaluate frontier AI coding models by performing realistic data engineering tasks, reviewing ETL pipelines, data warehouses, and distributed systems...
$80 - $120 per hour
...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our... ..., Larry Summers , and Jack Dorsey . Position: Finance operations / audit support Evaluator Type: Contract Compensation: $80–$120/hour...Contract workSummer workWork at officeRemote work- ...people think about and interact with personal finance.We’re a next-generation financial... ...financial world.The role:SoFi's Associate AI Engineer, Finance Transformation is a hands... ...stakeholders, with support.Nice to have:Exposure to evaluation, tracing, or observability practices for...InternshipRemote work
$130k - $220k
...Artificial Analysis is the leading independent AI benchmarking and insights company. They... ...** This role is best described as an AI Evaluation Engineer / Technical Generalist. It is... ...success bar for this role is becoming a world expert in modern AI technologies. This is a...Full timeWorldwide- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. The work requires evaluating AI-generated music in Hebrew and English to assess quality across genres and styles. You will rate...Immediate start
- A leading AI research organization in San Francisco is seeking a Research Engineer focused on pushing the boundaries of frontier models... ...offers the opportunity to shape innovative AI measurements and build evaluation environments that drive progress. #J-18808-Ljbffr OpenAI
- B Capital seeks a talented individual for an AI Evaluation role in San Francisco. This position involves conducting critical comparative analysis, refining evaluation systems, and collaborating with various teams to enhance model capabilities. The ideal candidate will have...
- ...San Francisco is seeking an innovative Quality Engineer for their AI products. This role blends ops, strategy, and analytics to... ...in leading labs, and ensure user satisfaction through effective evaluation baselines. Competitive salary and benefits offered, with a focus...
- Scale Labs seeks a Research Scientist focused on Frontier Risk Evaluations to design evaluation measures, harnesses and datasets for measuring risks posed by frontier AI systems. You will build harnesses to test models, collaborate with government agencies to scope evaluations...
- OpenAI is seeking a researcher to advance frontier evaluations and environments for safe AGI/ASI. You will help design north star model environments and steer major training runs so that research outputs translate into real-world products. Collaborate with researchers,...
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering with leading AI labs. You will author original, executable research problems that frontier models cannot solve. Engagement lasts...Part timeImmediate start
$216k - $270k
Scale AI, Inc. is looking for a Research Scientist specializing in Frontier Risk Evaluations to develop measures for assessing risks of advanced AI systems. In this role, you will design testing harnesses, collaborate with agencies, and publish reports to inform policymakers...- ...tech company in New York is seeking a Research Scientist focused on Frontier Risk Evaluations. The ideal candidate will contribute to designing and creating evaluation measures for assessing AI risks and will have strong experience in machine learning and technical...
- HeyMilo AI is hiring a Research Engineer to join our Applied AI team in San Francisco. You’ll design and build reinforcement learning... ...-world use cases, creating simulators, reward functions, and evaluation harnesses to measure model performance. The role blends ML research...
- Anthropic in San Francisco seeks a bio safety researcher to design and run capability evaluations for biology-focused models, build and curate datasets for safety classifiers, and iterate on those classifiers with ML engineers. You will work at the intersection of applied...
- ...hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. You will assess... .... Start date: immediate. Duration: up to 6 months. Most experts work around 20 hours per week; there is no cap. We...Immediate start
$400 per month
About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic infrastructure engineering...$150k
...engineer to enhance their machine intelligence systems in San Francisco. As part of the team, you'll be responsible for building evaluation infrastructure, designing data pipelines, and implementing fine-tuning processes. Ideal candidates have expertise in Python and machine...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Finance Expert - AI Evaluation Specialist. Be the first to apply!
- financial planner cfp San Francisco, CA
- senior financial advisor San Francisco, CA
- strategic finance associate San Francisco, CA
- financial advisor part time San Francisco, CA
- oracle financial functional consultant San Francisco, CA
- associate leveraged finance San Francisco, CA
- entry level financial advisor San Francisco, CA
- wealth advisor associate San Francisco, CA
- junior oracle financial functional consultant San Francisco, CA
- financial advisor development program San Francisco, CA


