Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Engineer, Benchmarks

$150k - $250k
Full-time

Clera

About the Role

This role sits at the core of a small, technical team building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. You'll own the design and implementation of evaluations that frontier labs and enterprise customers rely on to understand true agent capability. Your work directly shapes the credibility and rigor of a product at the frontier of AI evaluation.

What You'll Do

  • Design, implement, and own the quality of internal benchmarks for evaluating frontier AI agents on domain-specific tasks.

  • Partner with subject-matter experts to define realistic workflows and translate them into well-scoped evaluation tasks.

  • Build reliable infrastructure to run models and agents against benchmark tasks at scale.

  • Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.

  • Validate that benchmark performance correlates meaningfully with real-world agent behavior and customer needs.

  • Write clear documentation and benchmark reports that make results legible and credible to technical audiences.

What We're Looking For

  • 2–4 years of experience in research engineering or ML engineering, with a focus on building and delivering AI benchmarks, evaluation infrastructure, or agent environments.

  • Strong proficiency in Python, Docker, and Linux environments for research or production infrastructure.

  • Experience designing and running evaluations for AI agents or large language models.

  • Experience building and operating infrastructure to reliably run AI models or agents against benchmark tasks at scale.

  • Experience developing metrics and validation studies to assess benchmark difficulty, reliability, and real-world correlation.

  • Ability to collaborate with subject-matter experts and translate domain workflows into evaluation criteria.

  • Strong attention to detail and a habit of spotting subtle inconsistencies and edge cases.

  • Comfort working independently in fast-paced, early-stage startup environments with unstructured problem spaces.

  • Excellent written communication skills for cross-functional collaboration across time zones.

  • Bonus: Published papers or technical writing on AI benchmarking, model evaluation, or failure modes; experience with RL training pipelines or widely used public benchmark projects.

Compensation & Benefits

Salary range: $150,000 – $250,000 USD annually. Visa sponsorship is available.

Location

On-site in San Francisco, CA, United States .

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Research Engineer, Benchmarks in San Francisco, CA vacancy
  •  ...Responsibilities Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning....  ...generation, augmentation, and curation. Collaborate with AI researchers, applied AI teams, and data producers to align evaluations... 
    Suggested
    Full time
    Work at office
    Relocation package

    Mercor

    San Francisco, CA
    4 days ago
  • A leading technology company located in San Francisco is seeking a Machine Learning Research Engineer to design and develop safe AI benchmarking methodologies. This role involves collaboration with various teams to implement responsible evaluation techniques. Candidates... 
    Suggested

    Jobleads-US

    San Francisco, CA
    18 hours ago
  • $200k - $275k

     ...applied AI and data company, seeks a lead researcher to head in-house post-training and research...  ...train models, publish findings, drive benchmarks, contribute to platform tooling, and help build a high-performance engineering culture from the ground up, with a focus on... 
    Suggested
    Relocation

    Halluminate

    San Francisco, CA
    1 day ago
  • Obsidian in San Francisco seeks a Computational Statistics and Applied Mathematics Expert to design benchmark problems that test AI systems' ability to solve real scientific workflows. You will craft graduate-level problems requiring advanced software, run simulations,... 
    Suggested

    Obsidian

    San Francisco, CA
    1 day ago
  • $110.7k - $379.2k

    Position Summary Research Engineer — Post-Training & Small Language Models (SLMs), Healthcare AI Three hundred fifty million Americans...  .... • Build healthcare-grade evaluation — held-out clinical benchmarks, deployment regression gates, calibration and uncertainty,... 
    Suggested
    Local area
    Visa sponsorship

    Deloitte

    San Francisco, CA
    3 days ago
  • $200k - $225k

    About AlembicAlembic is where top engineers are solving marketing's hardest problem: proving...  ...to leverage advanced analyticsDocument research and implementation decisions for...  ...ensures new hires earn at or above real-time benchmarks. Ownership: Our generous equity program... 

    Alembic

    San Francisco, CA
    4 days ago
  • $180k - $280k

     ...Research EngineerSuperAnnotate helps the world's leading AI teams build responsible, next...  ...cutting edge of the field. As a Research Engineer, you'll take a research direction and run with it – finding the right papers, benchmarks, and prior work, reimplementing what's... 
    Full time

    SuperAnnotate AI

    San Francisco, CA
    18 hours ago
  • $200k - $350k

     ...training), second-time technical founders, engineers that made 100+ games for Voodoo,...  ...engaging games & 3D environments. Our current research spans:Distributed multi-agent...  ...orchestration and engagement modeling.Define new benchmarks for fun, retention, and interactive intelligence... 
    Visa sponsorship
    Relocation package

    ROAM

    San Francisco, CA
    18 hours ago
  •  ...Research EngineerHUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace...  ...to-end without a fully prescribed roadmapExperience working on benchmarks and evals - you can reason about what makes a task realistic,... 
    Full time
    Remote work
    Relocation
    Visa sponsorship

    Hud (yc W25)

    San Francisco, CA
    18 hours ago
  • $140k - $200k

     ...Center for AI Safety (CAIS) is a leading research and advocacy organization focused on...  ...introducing the first state-of-the-art benchmarks for measuring it. More recently, we've been...  ...policymakers. About the role As a Research Engineer (RE) or Research Scientist (RS) at CAIS,... 
    Work at office
    Local area

    AI Safety, Inc

    San Francisco, CA
    4 days ago
  • $180k - $250k

     ...AI Research EngineerLocation: San Francisco, CA Company Stage of Funding: Series A AI...  ...synthetic data generation pipelines, and benchmarking frameworks for document AI applications...  ...and benchmarks.Collaborate closely with engineering and product teams to deploy research... 
    Work at office
    Relocation package
    Monday to Friday

    Recruiting from Scratch

    San Francisco, CA
    18 hours ago
  •  ...Research EngineerWe believe that software is the foundation of modern civilization - yet...  ...Infrastructure.We're seeking an experienced Research Engineer to join our effort in building and...  ..., experience in model evaluation, and benchmarks. Reinforcement Learning experience is a... 
    Full time
    Work at office

    DepthFirst

    San Francisco, CA
    18 hours ago
  • $180k - $340k

     ...Research EngineerYou'll own the quality of AI across everything Gamma creates. As our Research Engineer, you'll design evaluation frameworks that measure AI output quality, systematically...  ..., validate changes against quality benchmarks, and ensure our AI gets smarter with... 
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    18 hours ago
  • $210k - $275k

     ...Research EngineerYou'll build the evaluation systems that tell us whether Firecrawl actually...  ...where you inherit a framework and run benchmarks. You'll design the metrics, build the...  ..."good" actually means and have the engineering depth to measure it, this is the role.About... 
    Full time
    Temporary work
    For contractors
    Remote work
    Flexible hours

    Firecrawl

    San Francisco, CA
    18 hours ago
  •  ...Research EngineerUnlock New Career Opportunities with SA Technologies Inc. We're Hiring: Research Engineer At SA Technologies Inc., we believe in turning career aspirations into reality...  ...with it - finding the right papers, benchmarks, and prior work, reimplementing what's... 
    Full time

    SA Technologies Inc

    San Francisco, CA
    18 hours ago
  •  ...makers of Devin, the first AI software engineer.Our team is extremely talent-dense. Among...  ...future systems can do. This role blends deep research and hands-on engineering. We don't...  ...production performance, not just isolated benchmarks.Evaluation Design and Integrity: Build... 
    Shift work

    Cognition AI

    San Francisco, CA
    18 hours ago
  • $220k - $300k

     ...Research Engineer, Post-TrainingVizcom is where design teams at companies like Nike, GM, New Balance, and Hasbro bring ideas from sketch...  ...stipendCompensationBase salary: $220,000–$300,000 USD + equityWe regularly benchmark compensation against relevant peer companies using current... 

    Vizcom

    San Francisco, CA
    3 days ago
  •  ...remembers, forgets, and learns over time. This isn't research bolted onto a product team: you'll take an...  ...into research hypotheses; implement and benchmark ideas straight out of the literature; and ship them with Engineering at SOTA latency, reliability, and cost. You'll... 
    Work at office
    Remote work

    Mem0

    San Francisco, CA
    5 days ago
  • $150k - $250k

     ...About the Role This is a Research Engineer role at an early-stage AI infrastructure startup building the technical foundation for training...  ...frontier AI agents. You'll work across QC automation, benchmarks, and synthetic data — sitting at the intersection of research... 
    Full time
    Visa sponsorship

    Clera

    San Francisco, CA
    2 days ago
  • $150k - $250k

     ...About the Role This is a hands-on engineering role at the intersection of applied research and customer deployment, sitting within a small, high-caliber team...  ...and Linux environments. ~ Experience working on benchmarks and evals, with sound judgment about what makes a... 
    Full time
    Work at office
    Remote work
    Visa sponsorship

    Clera

    San Francisco, CA
    2 days ago
  • $264.8k - $331k

     ...complex agents in enterprises around the world. The Enterprise ML Research Lab works on the front lines of this AI revolution. We are...  ...algorithms to real life enterprise datasets across our clients + benchmarks. This will involve creating best-in-class Agents that achieve... 
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  •  ...ownership. Every applied AI company we benchmark against like Decagon, Harvey, Sierra, Cursor...  ...scale, every day. We see exactly where research meets production and where the data is...  ...alongside elite and competitive engineering minds. Translate findings into infrastructure... 
    Relocation

    Rox Data Corp

    San Francisco, CA
    3 days ago
  • $165k - $310k

    Senior Research Engineer, LLM Training & Post-Training New York, New York, United States; Remote; San Francisco, California, United States...  ...performance bottlenecks. Design evaluation methodologies, benchmark models, analyze failure modes, and guide model improvements... 
    For contractors
    For subcontractor
    Work at office
    Remote work
    Work from home
    Flexible hours
    2 days per week

    Lightning AI

    San Francisco, CA
    2 days ago
  •  ...building Agentic AI that empowers software engineers by automating production engineering and...  ...powered workflows end‑to‑end, balancing research and engineering to create production‑...  ...training and evaluation Design and execute benchmarks to evaluate AI models, improve... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Resolve AI

    San Francisco, CA
    18 hours ago
  • $350k

     ...a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build...  ...production ML systems Have experience designing evals or benchmarks for LLMs Have domain expertise in a vertical where we... 
    Work at office
    Visa sponsorship
    Flexible hours

    Neura Market

    San Francisco, CA
    4 days ago
  •  ...harnessing, and deployment—and connect that research to the patients, clinicians, and real-...  ...improvements rather than just higher benchmark scores.Create meaningful, trustworthy, and...  ....Work closely with researchers, engineers, clinicians, and product teams to bring... 
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    1 day ago
  • $250k - $300k

     ...infrastructure that powers breakthrough AI models at leading research labs and enterprises. Since 2018, we've been pioneering data-centric...  ...Your Impact Create frameworks and tools to construct, train, benchmark and evaluate autonomous agent capabilities. Design agent-... 
    Work at office
    Flexible hours
    2 days per week

    Labelbox

    San Francisco, CA
    3 days ago
  •  ...safe, and under control. The Role We are seeking a Staff Research Engineer, AI/ML & Cybersecurity to serve as a core technical pillar...  ...observability, and reproducibility Develop model evaluation, benchmarking, and stress-testing workflows Cybersecurity & Governance Conduct... 

    Ephapsys

    San Francisco, CA
    4 days ago
  • $164.6k - $313.3k

     ...s Sound Design AI group (SODA) is looking for a driven Data/ML engineer to push the boundaries of audio GenAI. Join the team behind Firefly...  ...in Adobe products.We’re a small, collaborative and efficient research team looking for highly motivated candidates of all levels with... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    2 days ago
  • $200k - $350k

     ...stage AI company in San Francisco is seeking a Machine Learning Research Engineer to own end-to-end research cycles. The role involves training models across creative domains, developing evaluation benchmarks, and collaborating with creative experts and AI labs. The... 

    Coders Connect

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Engineer, Benchmarks. Be the first to apply!