Technical Lead, AI Evaluation Platform
$300k - $385kHarvey
Harvey is a secure AI platform for professionals in law, tax, and finance that augments productivity and automates complex workflows. Harvey uses algorithms with reasoning-adept LLMs that have been customized by our expert team of lawyers, engineers and research scientists. We’ve found product market fit and are scaling our team very quickly. Some reasons to join Harvey are: Exceptional product market fit: We have partnered with the largest law firms and professional service providers in the world like A&O, PwC, and many others. Strategic investors: Raised over $100 million from strategic investors including Sequoia, Kleiner Perkins, and the OpenAI Startup Fund. World-class team: Harvey is hiring the best technical and non-technical talent from DeepMind, Google Brain, Stripe, FAIR, Tesla Autopilot, Superhuman, Glean, etc. Partnerships: Our engineers and researchers work directly with OpenAI to build the future of generative AI and redefine professional services. Value: Top of market cash and equity compensation. Role We are looking for a technical lead who can own the development of our evaluation platform. In this role, you will: Build a team of 10-20 researchers and engineers with experience evaluating LLMs and large-scale AI systems. Lead research and development of novel model-based evaluation methods and language model programs for evaluating complex tasks in legal and professional services. Design and implement a red-teaming pipeline for our custom models and collaborate with other research teams to fine-tune models from human feedback. Train reward models that accurately reflect the preferences of top-tier domain experts. Experiment with synthetic data generation and LLM-based data augmentation to complement human-generated eval benchmarks. Impact: Lead research and development of Harvey’s evaluation platform. Contribute to a product that transforms the nature of professional services. Help define what it means for LLMs to effectively perform complex knowledge work tasks. Work directly with our founders, research, and product teams, as well as foundation model providers like OpenAI. Tackle unsolved research and engineering problems, including the hardest in the world relevant to LLMs in production. Qualifications 5+ years experience leading highly-technical teams composed of both researchers and engineers. Experience evaluating large-scale AI systems in high-stakes settings. Technical: can serve as a tech lead and contribute substantially to our codebase as necessary. Ability to communicate complex technical outcomes to diverse stakeholders. Strong conviction in setting technical direction. Compensation The expected range of compensation for this role is between $300,000 and $385,000. Additionally, this role is eligible to participate in our equity plan. The successful candidate’s starting salary will be determined based on non-discriminatory factors such as skills, experience, and geographic location. #J-18808-Ljbffr Harvey
$235.03k - $352.29k
...profound opportunity for AI to drive positive... ...a universal autonomy platform: self-driving for all... ...Rowe Price, and other leading investors. About the... ...Autonomy Leader to drive the technical roadmap for the... ...lead the development of evaluation tooling that ensures our...PlatformImmediate startFlexible hours- ...About the Team The Applied AI team works across research, engineering... ...transcending geographic, economic, or platform barriers. Our commitment is to... ...the Role In this role, you’ll lead development of the systems we use to evaluate the quality of our AI models and products...PlatformFull timeWork at officeRelocation package
$175k - $215k
...applied to a range of vehicle platforms and product use cases. The Waymo... ...state-of-the-art Generative AI to create a training ground for... ...Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge... ...will work closely with Tech Leads to refine our evaluation...PlatformFull timeRemote work$170k - $216k
...applied to a range of vehicle platforms and product use cases. The... ...role, you will report to a Technical Lead Manager. You Will Dive... ...mission-critical automation and evaluation frameworks that establish... ...of experience in industrial AI applications involving the creation...PlatformFull timeRemote work$229.9k - $262.4k
...Senior Lead AI Engineer (SDK’s: Gen AI Evaluation and MCP) Overview: At Capital One, we are creating responsible... ...of customers. Our AI models and platforms empower teams across Capital One to... ...of engineers, research scientists, technical program managers, and product...PlatformFull timePart time$230k - $270k
...democratizing access to cutting-edge AI innovation to enable any... ...to explore our AI Hub platform features extensively, enabling... ...data into insights instantly.Technical Lead Manager (TLM) — AI Systems &... ...custom API interfaces.Build Evaluation & Trajectory Testing Harnesses...PlatformWork at officeFlexible hours$116.2k - $229.1k
...Position SummaryJoin our AI & Engineering team in transforming technology platforms, driving innovation,... ...Senior Consultant, you will lead Commercial Banking... ...business conversations and technical discussions without... ...architecture diagrams, and evaluate design trade-...PlatformLocal areaImmediate start$245k - $295k
Ironclad is the leading AI contracting platform that transforms agreements into assets. Contracts move faster... ...-scale performance.We're hiring a Technical Lead Manager to lead the engineering... ..., prompt and workflow design, model evaluation, and production quality controls....PlatformFull timeContract workWork at office- ...One in San Francisco, CA is seeking a Senior Lead AI Engineer to architect, build, and scale AI-powered solutions across platforms. You will collaborate with engineers,... ...production-grade ML components, including model evaluation, governance, and observability. The role...Platform
$165k - $225k
Backed by leading Silicon Valley investors, Peregrine helps public safety organizations... ...speed and accuracy. Our AI-enabled platform turns siloed and disconnected data... ...safety operators to federal evaluators to commercial technical buyers. The work starts from technical...PlatformWork experience placementLocal area$151k - $178k
...CheckrCheckr is building the data platform to power safe and fair... ...of people rely on Checkr for AI verification in the moments that... ...606 and U.S. GAAP. You'll lead technical revenue accounting analyses,... ...with demonstrated experience evaluating complex customer contracts, non...PlatformContract workWork at officeLocal areaRemote workRelocation3 days per week- ...Artificial Analysis is the leading independent AI benchmarking company. We support... ...and hiring a Member of Technical Staff to lead it. We are... ...substrate for training and evaluating robot policies at scale, and... ...our leading AI benchmarking platform and working with our...Platform
$400 per month
...Mercor is partnering with a leading AI research lab to support a Frontier... ...project. Contributors help evaluate and improve frontier AI... ...coding models through structured technical assessments. The work... ...implementations involving cloud platforms, Kubernetes, CI/CD systems,...Platform- ...innovates at the frontier of AI infrastructure, search,... .... The the company API Platform brings our technology... .... The initiatives you lead will bolster the... .... Partner with fellow technical staff, GTM, support, and... ...developers learn about, evaluate, and select API products...Platform
$130k - $220k
...Artificial Analysis is the leading independent AI benchmarking and insights company... .... **Position: Member of Technical Staff** **What This Role... ...is best described as an AI Evaluation Engineer / Technical Generalist... ...product direction of the platform. The success bar for this...PlatformFull timeWorldwide$238k - $302k
...applied to a range of vehicle platforms and product use cases.... .... The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With... ...Experience navigating complex technical and product landscapes,... ...Tensorflow) Experience leading a team of Engineers The...PlatformFull timeRemote work- Member of Technical Staff, Lead Researcher San Francisco, CA; Sunnyvale, CA About... ...DoorDash is building an AI Research org from the ground... ...Partner across DoorDash with ML platform, product, and operations... ...training runs, and large-scale evaluation sweeps Full research...PlatformLocal area
$238k - $302k
...to a range of vehicle platforms and product use cases.... ...Driver Understanding and Evaluation team at Waymo develops... ...You will: Lead a top-tier applied ML team... ...deep learning and Gen AI. Lead the development... ...millions of miles. Drive technical direction, and provide...PlatformFull timeRemote work$200k
Dormont Manufacturing Co is seeking a Member of Technical Staff to manage our internal evaluations platform, essential for improving AI model performance. You will be responsible for designing and validating evaluation tasks, ensuring reproducibility and reliability in...Platform- Kindredventures is seeking research engineers to build a central evaluation framework and scalable evaluation pipelines for models and datasets. You will design benchmarks, implement baselines, and create dashboards that translate results into actionable insights for the...Platform
- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing infrastructure engineering... ...model-generated implementations on cloud platforms, Kubernetes, and CI/CD tools. This sprint-...Platform
$85 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San... ...AI coding agents to complete and evaluate complex engineering tasks. Review... ...details about the interview process and platform information, please check: For...PlatformContract workSummer workRemote work- ...communities around the globe. Our AI-powered labor marketplace... ...physical AI. We’re working with leading frontier labs to create the... ...teams, and create and drive the technical requirements that tell IRL what... ..., and more. Our AI-powered platform serves thousands of businesses...PlatformHourly payInternshipLocal areaShift work
- Mercor is partnering with a leading AI research lab to support a Frontier... ...Agents project. You will help evaluate frontier AI coding models through structured technical assessments and focus on... ...pipelines, data warehouses, analytics platforms, and distributed data systems,...Platform
$151.5k - $244.2k
...people. Discovery Technology and Platforms (DTP) accelerates molecule... ...team within DTP that enables AI-native drug discovery through... ...for scientific or technical applications.Preferred QualificationsPharmaceutical... ...) and Women’s Initiative for Leading at Lilly (WILL).Actual...PlatformFlexible hours$117.4k - $176k
.... By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re... ...Harvey is looking for a Technical Revenue Accounting Lead to help us scale revenue recognition... ...and usage-based deal structures Evaluate the revenue accounting implications...PlatformContract work- ...help shape the future of AI. Our mission: make... ...development, deployment, and evaluation. Sharding, replication,... .... As Shared Services Lead, you'll provide... ...this layer, guide the technical and architectural decisions... ...their needs into reliable platform primitives. You'll stay...PlatformWork at officeVisa sponsorship
- ...are proficient in Python and willing to learn new technologies. The role focuses on building and improving simulation and evaluation platforms for AI agents in San Francisco. Successful candidates will have 5-10 years of experience in software engineering, particularly...Platform
- ...powering the next generation of AI products. We build the... ..., but practical: a unified platform where high-performance inference... ...Overview We're hiring a Technical Accounting Lead to own fal's technical... ...stand up to auditor review Evaluate accounting implications of...PlatformContract work
$290.4k - $363k
Scale's LLM post-training platform team builds our internal distributed... ...and automatic training and evaluation of LLMs. It also serves as... ...at the heart of the field of AI as an indispensable provider... ...technologies that power the world's leading models, and help enterprises...PlatformFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Technical Lead, AI Evaluation Platform. Be the first to apply!
- technical leader San Francisco, CA
- technical lead San Francisco, CA
- technical lead manager San Francisco, CA
- digital platform specialist San Francisco, CA
- power platform San Francisco, CA
- director of digital platform San Francisco, CA
- platform product manager San Francisco, CA
- platform manager San Francisco, CA
- java tech lead
- technical leader




