Research Engineer, Benchmarks
$150k - $250kClera
About the Role
Join a small, highly technical team of researchers and engineers — including International Olympiad medalists and published AI researchers — at an early-stage startup building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. As a Research Engineer, Benchmarks , you'll own the design and implementation of evaluations that frontier labs and enterprise customers rely on to measure real-world agent performance. This is a critical, high-ownership role at the intersection of research rigor and engineering execution.
The company operates in the AI/ML evaluation and reinforcement learning infrastructure space, providing a platform for building, running, and scaling RL environments and post-training datasets. The team is based in San Francisco, CA and works on-site. Visa sponsorship is available.
What You'll Do
Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
Partner with subject-matter experts to define realistic workflows and tasks for domain-specific evaluations.
Build reliable infrastructure to run models and agents against benchmark tasks at scale.
Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.
Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations.
Write clear documentation and benchmark reports that make results legible and credible to technical audiences.
What We're Looking For
Required
2–4 years of experience in research engineering, ML engineering, or related roles — with a focus on building and delivering AI benchmarks, evaluation infrastructure, or agent environments.
Demonstrated experience designing, implementing, and running benchmarks or evaluation environments for AI agents or large language models.
Strong proficiency in Python, Docker, and Linux environments for building research or production infrastructure.
Experience building and operating infrastructure to reliably run AI models or agents against benchmark or evaluation tasks at scale.
Experience developing metrics, statistical analyses, or validation studies to assess benchmark difficulty, reliability, and real-world correlation.
Experience collaborating with subject-matter experts to translate domain workflows into benchmark tasks and evaluation criteria.
Experience analyzing workflows across diverse technical or business domains to inform task design.
Strong technical writing skills — able to produce benchmark reports and documentation for research and engineering audiences.
Nice to Have
Published papers or technical blog posts on AI benchmarking, model evaluation, or model failure modes.
Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation.
Background at frontier AI labs, research institutions, or involvement in widely used public benchmark projects.
Traits We Value
Deep curiosity about how workflows operate across varied domains.
Sharp attention to detail — a habit of spotting subtle inconsistencies and edge cases in task design.
Ability to reason from first principles about task design, scoring, and failure modes.
Comfort thriving in unstructured problem spaces and working independently in a fast-paced, early-stage environment.
Excellent communication skills for collaborating across time zones and with technical teams.
Compensation & Benefits
Salary: $150,000 – $250,000 USD annually, depending on experience.
Early-stage equity participation.
Visa sponsorship available.
Location
This is an on-site role based in San Francisco, CA, United States . Candidates must be willing and able to work from the office. Fully remote arrangements are not available for this position.
- ..., seeks a leader to drive our in-house research and post-training efforts. You will train... ...models, evaluate quality, and build benchmarks while shaping the public research narrative... ...to platform tooling, and help grow the engineering team. Expect startup speed, strong...Suggested
- A leading technology company located in San Francisco is seeking a Machine Learning Research Engineer to design and develop safe AI benchmarking methodologies. This role involves collaboration with various teams to implement responsible evaluation techniques. Candidates...Suggested
$110.7k - $379.2k
Position Summary Research Engineer — Post-Training & Small Language Models (SLMs), Healthcare AI Three hundred fifty million Americans... .... • Build healthcare-grade evaluation — held-out clinical benchmarks, deployment regression gates, calibration and uncertainty,...SuggestedLocal areaVisa sponsorship$200k - $225k
About AlembicAlembic is where top engineers are solving marketing's hardest problem: proving... ...to leverage advanced analyticsDocument research and implementation decisions for... ...ensures new hires earn at or above real-time benchmarks. Ownership: Our generous equity program...Suggested- ...2025. The Impact You'll Make Our research team is expanding to keep pace with a wave... ...edge of the field. As a Research Engineer, you'll take a research direction and run with it – finding the right papers, benchmarks, and prior work, reimplementing what's relevant...SuggestedFull time
$150k - $250k
...International Olympiad medalists, serial AI startup founders, and researchers with publications at top venues including ICLR and NeurIPS. As a Research Engineer , you will work across agent training environments, benchmarks, and synthetic data pipelines — shaping how AI agents...Work at officeVisa sponsorship- ...teammates (we've accomplished a lot as just one engineer and one designer!) to a clan of around... ...with leading performance on relevant benchmarks, generating novel insights along the way... ...for your life". Specifically, in an AI research engineer role, we are looking for the...
$150k - $250k
...About the Role Join a small, high-caliber engineering team at a fast-moving AI infrastructure... ...-training data. As a Forward Deployed Research Engineer , you will own end-to-end... ...environments. ~ Experience working on benchmarks and evals for RL training data and AI...Full timeWork at officeRemote workRelocationVisa sponsorship- ...building Agentic AI that empowers software engineers by automating production engineering and... ...powered workflows end‑to‑end, balancing research and engineering to create production‑... ...training and evaluation Design and execute benchmarks to evaluate AI models, improve...Full timeWork at officeVisa sponsorshipFlexible hours
$140k - $200k
...Center for AI Safety (CAIS) is a leading research and advocacy organization focused on... ...introducing the first state-of-the-art benchmarks for measuring it. More recently, we've been... ...policymakers. About the role As a Research Engineer (RE) or Research Scientist (RS) at CAIS,...Work at officeLocal area$200k - $350k
...training), second-time technical founders, engineers that made 100+ games for Voodoo,... ...engaging games & 3D environments. Our current research spans: Distributed multi-agent... ...and engagement modeling. Define new benchmarks for fun, retention, and interactive intelligence...Visa sponsorshipRelocation package- ...Research Engineer On Physical Ai Team Hedra is a pioneering generative modeling company — first models to market — now building a Physical... ...action sequences Evaluate model performance using both benchmark datasets and real-world deployment metrics Contributions...Work at office
- ...Applied Scientist / Research Engineer â Speech AI Â HIGHLIGHTS Location: Â San Francisco, CA OR REMOTE Position Type... ...training improvements. Compare approaches using clear benchmarks and production-relevant quality measures. Move quickly...Remote work
$264.8k - $331k
...complex agents in enterprises around the world. The Enterprise ML Research Lab works on the front lines of this AI revolution. We are... ...algorithms to real life enterprise datasets across our clients + benchmarks. This will involve creating best-in-class Agents that achieve...Full time$165k - $310k
Senior Research Engineer, LLM Training & Post-Training New York, New York, United States; Remote; San Francisco, California, United States... ...performance bottlenecks. Design evaluation methodologies, benchmark models, analyze failure modes, and guide model improvements...For contractorsFor subcontractorWork at officeRemote workWork from homeFlexible hours2 days per week- ...ownership. Every applied AI company we benchmark against like Decagon, Harvey, Sierra, Cursor... ...scale, every day. We see exactly where research meets production and where the data is... ...alongside elite and competitive engineering minds. Translate findings into infrastructure...Relocation
$350k
...a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build... ...production ML systems Have experience designing evals or benchmarks for LLMs Have domain expertise in a vertical where we...Work at officeVisa sponsorshipFlexible hours$75 - $90 per hour
...Description About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey ....Contract workFor contractorsSummer workRemote work$150k - $250k
...About the Role We're a ~15-person engineering team — made up of Olympiad medalists and published researchers — building infrastructure that aligns AI to real-world workflows... ..., or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or...Full timeRelocationVisa sponsorship$200k - $350k
...stage AI company in San Francisco is seeking a Machine Learning Research Engineer to own end-to-end research cycles. The role involves training models across creative domains, developing evaluation benchmarks, and collaborating with creative experts and AI labs. The...$164.6k - $313.3k
...s Sound Design AI group (SODA) is looking for a driven Data/ML engineer to push the boundaries of audio GenAI. Join the team behind Firefly... ...in Adobe products.We’re a small, collaborative and efficient research team looking for highly motivated candidates of all levels with...Full timeTemporary workLocal areaWorldwide$200k - $350k
...deeply curious—building at the intersection of research, product, and creativity . The Role As a Machine Learning Research Engineer , you’ll own end-to-end research cycles—... ...(writing, design, visual style) Develop benchmarks and evaluation methods for subjective tasks...- Factory is seeking innovative Research Engineers to design and integrate advanced AI and ML capabilities that revolutionize productivity and accelerate innovation within software organizations.What you will do and achieve:Design, develop, and deploy AI-driven agentic systems...Work at office
$197.3k - $313.7k
...TeamSalesforce AI is looking for talented software and platform engineers to embed in our AI team to bridge the gap between frontier AI... ...where your engineering skills directly enable world-class research and products used by millions?At Salesforce, we are driving the...Full time$150k - $250k
...ability to scale. You'll join a ~15-person engineering team made up of Olympiad medalists, AI startup founders, and published researchers. This role is ideal for someone who... ...Strongly Required: Experience working on benchmarks and evaluations for RL training data,...For contractorsRemote workVisa sponsorship- ...Luma AI is seeking a Research Scientist/Engineer to advance multimodal agent models across research and product integrations. You will explore modeling, data, systems and evaluation to push state-of-the-art capabilities. Join a team driving large-scale training with PyTorch...
$140k - $250k
...This is the company's top hiring priority and a genuinely hard research problem. Because data flows through a decentralized marketplace... ...scale is the single biggest bottleneck to growth. As a Research Engineer, you will build the automated systems that verify and assure...Full timeWork at officeRemote work- ...Manning, Michael Ovitz, Michael Abbott, Cory Levy, Kevin Hartz, and others. About the Role: They are hiring an Agent Product Engineer to build high-taste products for self-learning agents. This role is crucial for scaling the product to meet customer demand and...Full timeWork at officeRelocation
- ...Preference Model is seeking Research Engineers or Research Scientists to advance self-directed learning in AI. The role involves training and evaluating models within proprietary RL environments and optimizing ML infrastructure. Candidates will benefit from competitive...
- ...Job Description Job Description We are Genmo, a research lab dedicated to building open, state-of-the-art models for video generation... .... About the Role We're seeking an exceptional Software Engineer to join our research team in advancing the frontiers of visual...Work at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Engineer, Benchmarks. Be the first to apply!
- ai research engineer San Francisco, CA
- deep learning research engineer San Francisco, CA
- research engineer San Francisco, CA
- research assistant engineering San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- research programmer San Francisco, CA
- research software engineer San Francisco, CA
- senior research engineer San Francisco, CA
- historical research San Francisco, CA
- biology research San Francisco, CA


