Research Engineer, Benchmarks
$150k - $250kClera
About the Role
This role sits at the core of a small, technical team building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. You'll own the design and implementation of evaluations that frontier labs and enterprise customers rely on to understand true agent capability. Your work directly shapes the credibility and rigor of a product at the frontier of AI evaluation.
What You'll Do
Design, implement, and own the quality of internal benchmarks for evaluating frontier AI agents on domain-specific tasks.
Partner with subject-matter experts to define realistic workflows and translate them into well-scoped evaluation tasks.
Build reliable infrastructure to run models and agents against benchmark tasks at scale.
Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.
Validate that benchmark performance correlates meaningfully with real-world agent behavior and customer needs.
Write clear documentation and benchmark reports that make results legible and credible to technical audiences.
What We're Looking For
2–4 years of experience in research engineering or ML engineering, with a focus on building and delivering AI benchmarks, evaluation infrastructure, or agent environments.
Strong proficiency in Python, Docker, and Linux environments for research or production infrastructure.
Experience designing and running evaluations for AI agents or large language models.
Experience building and operating infrastructure to reliably run AI models or agents against benchmark tasks at scale.
Experience developing metrics and validation studies to assess benchmark difficulty, reliability, and real-world correlation.
Ability to collaborate with subject-matter experts and translate domain workflows into evaluation criteria.
Strong attention to detail and a habit of spotting subtle inconsistencies and edge cases.
Comfort working independently in fast-paced, early-stage startup environments with unstructured problem spaces.
Excellent written communication skills for cross-functional collaboration across time zones.
Bonus: Published papers or technical writing on AI benchmarking, model evaluation, or failure modes; experience with RL training pipelines or widely used public benchmark projects.
Compensation & Benefits
Salary range: $150,000 – $250,000 USD annually. Visa sponsorship is available.
Location
On-site in San Francisco, CA, United States .
- ...Responsibilities Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning.... ...generation, augmentation, and curation. Collaborate with AI researchers, applied AI teams, and data producers to align evaluations...SuggestedFull timeWork at officeRelocation package
- A leading technology company located in San Francisco is seeking a Machine Learning Research Engineer to design and develop safe AI benchmarking methodologies. This role involves collaboration with various teams to implement responsible evaluation techniques. Candidates...Suggested
$200k - $275k
...applied AI and data company, seeks a lead researcher to head in-house post-training and research... ...train models, publish findings, drive benchmarks, contribute to platform tooling, and help build a high-performance engineering culture from the ground up, with a focus on...SuggestedRelocation- Obsidian in San Francisco seeks a Computational Statistics and Applied Mathematics Expert to design benchmark problems that test AI systems' ability to solve real scientific workflows. You will craft graduate-level problems requiring advanced software, run simulations,...Suggested
$110.7k - $379.2k
Position Summary Research Engineer — Post-Training & Small Language Models (SLMs), Healthcare AI Three hundred fifty million Americans... .... • Build healthcare-grade evaluation — held-out clinical benchmarks, deployment regression gates, calibration and uncertainty,...SuggestedLocal areaVisa sponsorship$200k - $225k
About AlembicAlembic is where top engineers are solving marketing's hardest problem: proving... ...to leverage advanced analyticsDocument research and implementation decisions for... ...ensures new hires earn at or above real-time benchmarks. Ownership: Our generous equity program...$180k - $280k
...Research EngineerSuperAnnotate helps the world's leading AI teams build responsible, next... ...cutting edge of the field. As a Research Engineer, you'll take a research direction and run with it – finding the right papers, benchmarks, and prior work, reimplementing what's...Full time$200k - $350k
...training), second-time technical founders, engineers that made 100+ games for Voodoo,... ...engaging games & 3D environments. Our current research spans:Distributed multi-agent... ...orchestration and engagement modeling.Define new benchmarks for fun, retention, and interactive intelligence...Visa sponsorshipRelocation package- ...Research EngineerHUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace... ...to-end without a fully prescribed roadmapExperience working on benchmarks and evals - you can reason about what makes a task realistic,...Full timeRemote workRelocationVisa sponsorship
$140k - $200k
...Center for AI Safety (CAIS) is a leading research and advocacy organization focused on... ...introducing the first state-of-the-art benchmarks for measuring it. More recently, we've been... ...policymakers. About the role As a Research Engineer (RE) or Research Scientist (RS) at CAIS,...Work at officeLocal area$180k - $250k
...AI Research EngineerLocation: San Francisco, CA Company Stage of Funding: Series A AI... ...synthetic data generation pipelines, and benchmarking frameworks for document AI applications... ...and benchmarks.Collaborate closely with engineering and product teams to deploy research...Work at officeRelocation packageMonday to Friday- ...Research EngineerWe believe that software is the foundation of modern civilization - yet... ...Infrastructure.We're seeking an experienced Research Engineer to join our effort in building and... ..., experience in model evaluation, and benchmarks. Reinforcement Learning experience is a...Full timeWork at office
$180k - $340k
...Research EngineerYou'll own the quality of AI across everything Gamma creates. As our Research Engineer, you'll design evaluation frameworks that measure AI output quality, systematically... ..., validate changes against quality benchmarks, and ensure our AI gets smarter with...Full timeWork at officeWork from home$210k - $275k
...Research EngineerYou'll build the evaluation systems that tell us whether Firecrawl actually... ...where you inherit a framework and run benchmarks. You'll design the metrics, build the... ..."good" actually means and have the engineering depth to measure it, this is the role.About...Full timeTemporary workFor contractorsRemote workFlexible hours- ...Research EngineerUnlock New Career Opportunities with SA Technologies Inc. We're Hiring: Research Engineer At SA Technologies Inc., we believe in turning career aspirations into reality... ...with it - finding the right papers, benchmarks, and prior work, reimplementing what's...Full time
- ...makers of Devin, the first AI software engineer.Our team is extremely talent-dense. Among... ...future systems can do. This role blends deep research and hands-on engineering. We don't... ...production performance, not just isolated benchmarks.Evaluation Design and Integrity: Build...Shift work
$220k - $300k
...Research Engineer, Post-TrainingVizcom is where design teams at companies like Nike, GM, New Balance, and Hasbro bring ideas from sketch... ...stipendCompensationBase salary: $220,000–$300,000 USD + equityWe regularly benchmark compensation against relevant peer companies using current...- ...remembers, forgets, and learns over time. This isn't research bolted onto a product team: you'll take an... ...into research hypotheses; implement and benchmark ideas straight out of the literature; and ship them with Engineering at SOTA latency, reliability, and cost. You'll...Work at officeRemote work
$150k - $250k
...About the Role This is a Research Engineer role at an early-stage AI infrastructure startup building the technical foundation for training... ...frontier AI agents. You'll work across QC automation, benchmarks, and synthetic data — sitting at the intersection of research...Full timeVisa sponsorship$150k - $250k
...About the Role This is a hands-on engineering role at the intersection of applied research and customer deployment, sitting within a small, high-caliber team... ...and Linux environments. ~ Experience working on benchmarks and evals, with sound judgment about what makes a...Full timeWork at officeRemote workVisa sponsorship$264.8k - $331k
...complex agents in enterprises around the world. The Enterprise ML Research Lab works on the front lines of this AI revolution. We are... ...algorithms to real life enterprise datasets across our clients + benchmarks. This will involve creating best-in-class Agents that achieve...Full time- ...ownership. Every applied AI company we benchmark against like Decagon, Harvey, Sierra, Cursor... ...scale, every day. We see exactly where research meets production and where the data is... ...alongside elite and competitive engineering minds. Translate findings into infrastructure...Relocation
$165k - $310k
Senior Research Engineer, LLM Training & Post-Training New York, New York, United States; Remote; San Francisco, California, United States... ...performance bottlenecks. Design evaluation methodologies, benchmark models, analyze failure modes, and guide model improvements...For contractorsFor subcontractorWork at officeRemote workWork from homeFlexible hours2 days per week- ...building Agentic AI that empowers software engineers by automating production engineering and... ...powered workflows end‑to‑end, balancing research and engineering to create production‑... ...training and evaluation Design and execute benchmarks to evaluate AI models, improve...Full timeWork at officeVisa sponsorshipFlexible hours
$350k
...a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build... ...production ML systems Have experience designing evals or benchmarks for LLMs Have domain expertise in a vertical where we...Work at officeVisa sponsorshipFlexible hours- ...harnessing, and deployment—and connect that research to the patients, clinicians, and real-... ...improvements rather than just higher benchmark scores.Create meaningful, trustworthy, and... ....Work closely with researchers, engineers, clinicians, and product teams to bring...Work at officeRelocation package
$250k - $300k
...infrastructure that powers breakthrough AI models at leading research labs and enterprises. Since 2018, we've been pioneering data-centric... ...Your Impact Create frameworks and tools to construct, train, benchmark and evaluate autonomous agent capabilities. Design agent-...Work at officeFlexible hours2 days per week- ...safe, and under control. The Role We are seeking a Staff Research Engineer, AI/ML & Cybersecurity to serve as a core technical pillar... ...observability, and reproducibility Develop model evaluation, benchmarking, and stress-testing workflows Cybersecurity & Governance Conduct...
$164.6k - $313.3k
...s Sound Design AI group (SODA) is looking for a driven Data/ML engineer to push the boundaries of audio GenAI. Join the team behind Firefly... ...in Adobe products.We’re a small, collaborative and efficient research team looking for highly motivated candidates of all levels with...Full timeTemporary workLocal areaWorldwide$200k - $350k
...stage AI company in San Francisco is seeking a Machine Learning Research Engineer to own end-to-end research cycles. The role involves training models across creative domains, developing evaluation benchmarks, and collaborating with creative experts and AI labs. The...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Engineer, Benchmarks. Be the first to apply!
- research programmer San Francisco, CA
- research engineer San Francisco, CA
- senior research engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- deep learning research engineer San Francisco, CA
- research software engineer San Francisco, CA
- research assistant engineering San Francisco, CA
- ai research engineer San Francisco, CA
- ultrasound research San Francisco, CA
- research editor San Francisco, CA


