Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Engineer, Benchmarks

$150k - $250k
Full-time

Clera

About the Role

Join a small, highly technical team of researchers and engineers — including International Olympiad medalists and published AI researchers — at an early-stage startup building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. As a Research Engineer, Benchmarks , you'll own the design and implementation of evaluations that frontier labs and enterprise customers rely on to measure real-world agent performance. This is a critical, high-ownership role at the intersection of research rigor and engineering execution.

The company operates in the AI/ML evaluation and reinforcement learning infrastructure space, providing a platform for building, running, and scaling RL environments and post-training datasets. The team is based in San Francisco, CA and works on-site. Visa sponsorship is available.

What You'll Do

  • Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.

  • Partner with subject-matter experts to define realistic workflows and tasks for domain-specific evaluations.

  • Build reliable infrastructure to run models and agents against benchmark tasks at scale.

  • Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.

  • Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations.

  • Write clear documentation and benchmark reports that make results legible and credible to technical audiences.

What We're Looking For

Required

  • 2–4 years of experience in research engineering, ML engineering, or related roles — with a focus on building and delivering AI benchmarks, evaluation infrastructure, or agent environments.

  • Demonstrated experience designing, implementing, and running benchmarks or evaluation environments for AI agents or large language models.

  • Strong proficiency in Python, Docker, and Linux environments for building research or production infrastructure.

  • Experience building and operating infrastructure to reliably run AI models or agents against benchmark or evaluation tasks at scale.

  • Experience developing metrics, statistical analyses, or validation studies to assess benchmark difficulty, reliability, and real-world correlation.

  • Experience collaborating with subject-matter experts to translate domain workflows into benchmark tasks and evaluation criteria.

  • Experience analyzing workflows across diverse technical or business domains to inform task design.

  • Strong technical writing skills — able to produce benchmark reports and documentation for research and engineering audiences.

Nice to Have

  • Published papers or technical blog posts on AI benchmarking, model evaluation, or model failure modes.

  • Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation.

  • Background at frontier AI labs, research institutions, or involvement in widely used public benchmark projects.

Traits We Value

  • Deep curiosity about how workflows operate across varied domains.

  • Sharp attention to detail — a habit of spotting subtle inconsistencies and edge cases in task design.

  • Ability to reason from first principles about task design, scoring, and failure modes.

  • Comfort thriving in unstructured problem spaces and working independently in a fast-paced, early-stage environment.

  • Excellent communication skills for collaborating across time zones and with technical teams.

Compensation & Benefits

  • Salary: $150,000 – $250,000 USD annually, depending on experience.

  • Early-stage equity participation.

  • Visa sponsorship available.

Location

This is an on-site role based in San Francisco, CA, United States . Candidates must be willing and able to work from the office. Fully remote arrangements are not available for this position.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Research Engineer, Benchmarks in San Francisco, CA vacancy
  •  ..., seeks a leader to drive our in-house research and post-training efforts. You will train...  ...models, evaluate quality, and build benchmarks while shaping the public research narrative...  ...to platform tooling, and help grow the engineering team. Expect startup speed, strong... 
    Suggested

    Halluminate

    San Francisco, CA
    3 days ago
  • A leading technology company located in San Francisco is seeking a Machine Learning Research Engineer to design and develop safe AI benchmarking methodologies. This role involves collaboration with various teams to implement responsible evaluation techniques. Candidates... 
    Suggested

    Jobleads-US

    San Francisco, CA
    2 days ago
  • $110.7k - $379.2k

    Position Summary Research Engineer — Post-Training & Small Language Models (SLMs), Healthcare AI Three hundred fifty million Americans...  .... • Build healthcare-grade evaluation — held-out clinical benchmarks, deployment regression gates, calibration and uncertainty,... 
    Suggested
    Local area
    Visa sponsorship

    Deloitte

    San Francisco, CA
    21 hours ago
  • $200k - $225k

    About AlembicAlembic is where top engineers are solving marketing's hardest problem: proving...  ...to leverage advanced analyticsDocument research and implementation decisions for...  ...ensures new hires earn at or above real-time benchmarks. Ownership: Our generous equity program... 
    Suggested

    Alembic

    San Francisco, CA
    1 day ago
  • $180k - $280k

     ...Research EngineerSuperAnnotate helps the world's leading AI teams build responsible, next...  ...cutting edge of the field. As a Research Engineer, you'll take a research direction and run with it – finding the right papers, benchmarks, and prior work, reimplementing what's... 
    Suggested
    Full time

    SuperAnnotate AI

    San Francisco, CA
    2 days ago
  •  ...troubleshooting have become a massive tax of engineering velocity. Resolve AI is solving this by...  ...workflows end-to-end, balancing research and engineering to create production-ready...  ...and evaluationDesign and execute benchmarks to evaluate AI models, improve performance... 
    Work at office

    Resolve AI

    San Francisco, CA
    2 days ago
  • $180k - $340k

     ...Research EngineerYou'll own the quality of AI across everything Gamma creates. As our Research Engineer, you'll design evaluation frameworks that measure AI output quality, systematically...  ..., validate changes against quality benchmarks, and ensure our AI gets smarter with... 
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    2 days ago
  •  ...Research EngineerUnlock New Career Opportunities with SA Technologies Inc. We're Hiring: Research Engineer At SA Technologies Inc., we believe in turning career aspirations into reality...  ...with it - finding the right papers, benchmarks, and prior work, reimplementing what's... 
    Full time

    SA Technologies Inc

    San Francisco, CA
    2 days ago
  •  ...makers of Devin, the first AI software engineer.Our team is extremely talent-dense. Among...  ...future systems can do. This role blends deep research and hands-on engineering. We don't...  ...production performance, not just isolated benchmarks.Evaluation Design and Integrity: Build... 
    Shift work

    Cognition AI

    San Francisco, CA
    2 days ago
  • $210k - $275k

     ...Research EngineerYou'll build the evaluation systems that tell us whether Firecrawl actually...  ...where you inherit a framework and run benchmarks. You'll design the metrics, build the...  ..."good" actually means and have the engineering depth to measure it, this is the role.About... 
    Full time
    Temporary work
    For contractors
    Remote work
    Flexible hours

    Firecrawl

    San Francisco, CA
    2 days ago
  •  ...Research EngineerHUD is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace...  ...to-end without a fully prescribed roadmapExperience working on benchmarks and evals - you can reason about what makes a task realistic,... 
    Full time
    Remote work
    Relocation
    Visa sponsorship

    Hud (yc W25)

    San Francisco, CA
    2 days ago
  •  ...Research EngineerWe believe that software is the foundation of modern civilization - yet...  ...Infrastructure.We're seeking an experienced Research Engineer to join our effort in building and...  ..., experience in model evaluation, and benchmarks. Reinforcement Learning experience is a... 
    Full time
    Work at office

    DepthFirst

    San Francisco, CA
    2 days ago
  • $180k - $250k

     ...AI Research EngineerLocation: San Francisco, CA Company Stage of Funding: Series A AI...  ...synthetic data generation pipelines, and benchmarking frameworks for document AI applications...  ...and benchmarks.Collaborate closely with engineering and product teams to deploy research... 
    Work at office
    Relocation package
    Monday to Friday

    Recruiting from Scratch

    San Francisco, CA
    2 days ago
  • $140k - $200k

     ...Center for AI Safety (CAIS) is a leading research and advocacy organization focused on...  ...introducing the first state-of-the-art benchmarks for measuring it. More recently, we've been...  ...policymakers. About the role As a Research Engineer (RE) or Research Scientist (RS) at CAIS,... 
    Work at office
    Local area

    AI Safety, Inc

    San Francisco, CA
    1 day ago
  •  ...Research EngineerDatacurve provides the frontier coding data that powers the world's most...  ...create the world's first autonomous data engine, allowing us to teach the next...  ...of the same domain. You will produce the benchmarks, artifacts, and technical narratives that... 
    Shift work

    Datacurve

    San Francisco, CA
    2 days ago
  • $350k

     ...Research Engineer, Visual Knowledge WorkNew York City, NY; San Francisco, CA; Seattle, WAAbout AnthropicAnthropic's mission is to create...  ...Candidates May Also Have Experience With:Designing evals or benchmarks for LLMs or vision language modelsLarge-scale pretraining, SL... 
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    2 days ago
  •  ...Memory Features Research ScientistOwn the end-to-end lifecycle of memory features—from research to production....  ...pain points into research hypotheses; implement and benchmark ideas from papers; and ship with Engineering to state-of-the-art latency, reliability, and cost.... 

    Mem0

    San Francisco, CA
    2 days ago
  • $264.8k - $331k

     ...complex agents in enterprises around the world. The Enterprise ML Research Lab works on the front lines of this AI revolution. We are...  ...algorithms to real life enterprise datasets across our clients + benchmarks. This will involve creating best-in-class Agents that achieve... 
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  •  ...building Agentic AI that empowers software engineers by automating production engineering and...  ...powered workflows end‑to‑end, balancing research and engineering to create production‑...  ...training and evaluation Design and execute benchmarks to evaluate AI models, improve... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Resolve AI

    San Francisco, CA
    2 days ago
  •  ...ownership. Every applied AI company we benchmark against like Decagon, Harvey, Sierra, Cursor...  ...scale, every day. We see exactly where research meets production and where the data is...  ...alongside elite and competitive engineering minds. Translate findings into infrastructure... 
    Relocation

    Rox Data Corp

    San Francisco, CA
    21 hours ago
  • The role As an applied research engineer, you’ll own customer engagements end-to-end: understanding their data, their workflows, and their...  ...are fortunate to be backed by partners like Kleiner Perkins, Benchmark, Sequoia, Lux, and Greenoaks. Who Thrives Here : We're... 
    Work at office
    Visa sponsorship
    Relocation package

    Applied Compute

    San Francisco, CA
    11 hours ago
  • $350k

     ...a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build...  ...production ML systems Have experience designing evals or benchmarks for LLMs Have domain expertise in a vertical where we... 
    Work at office
    Visa sponsorship
    Flexible hours

    Neura Market

    San Francisco, CA
    1 day ago
  • $200k - $350k

     ...stage AI company in San Francisco is seeking a Machine Learning Research Engineer to own end-to-end research cycles. The role involves training models across creative domains, developing evaluation benchmarks, and collaborating with creative experts and AI labs. The... 

    Coders Connect

    San Francisco, CA
    1 day ago
  • $75 - $90 per hour

     ...Description About the job Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .... 
    Contract work
    For contractors
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    a month ago
  •  ...safe, and under control. The Role We are seeking a Staff Research Engineer, AI/ML & Cybersecurity to serve as a core technical pillar...  ...observability, and reproducibility Develop model evaluation, benchmarking, and stress-testing workflows Cybersecurity & Governance Conduct... 

    Ephapsys

    San Francisco, CA
    1 day ago
  • $250k - $300k

     ...infrastructure that powers breakthrough AI models at leading research labs and enterprises. Since 2018, we've been pioneering data-centric...  ...Your Impact Create frameworks and tools to construct, train, benchmark and evaluate autonomous agent capabilities. Design agent-... 
    Work at office
    Flexible hours
    2 days per week

    Labelbox

    San Francisco, CA
    11 hours ago
  • $200k - $350k

     ...deeply curious—building at the intersection of research, product, and creativity . The Role As a Machine Learning Research Engineer , you’ll own end-to-end research cycles—...  ...(writing, design, visual style) Develop benchmarks and evaluation methods for subjective tasks... 

    Coders Connect

    San Francisco, CA
    1 day ago
  • Factory is seeking innovative Research Engineers to design and integrate advanced AI and ML capabilities that revolutionize productivity and accelerate innovation within software organizations.What you will do and achieve:Design, develop, and deploy AI-driven agentic systems... 
    Work at office

    The San Francisco AI Factory

    San Francisco, CA
    1 day ago
  • $197.3k - $313.7k

     ...TeamSalesforce AI is looking for talented software and platform engineers to embed in our AI team to bridge the gap between frontier AI...  ...where your engineering skills directly enable world-class research and products used by millions?At Salesforce, we are driving the... 
    Full time

    Salesforce

    San Francisco, CA
    21 hours ago
  • $164.6k - $313.3k

     ...s Sound Design AI group (SODA) is looking for a driven Data/ML engineer to push the boundaries of audio GenAI. Join the team behind Firefly...  ...in Adobe products.We’re a small, collaborative and efficient research team looking for highly motivated candidates of all levels with... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Engineer, Benchmarks. Be the first to apply!