Research Engineer, Benchmarks
$150k - $250kClera
Job Description
Job Description
About the Role
This is a high-ownership research engineering role on a small, technical team focused on designing and building benchmarks that evaluate frontier AI agents on realistic, domain-specific workflows. You'll work alongside researchers and engineers to ensure evaluations are rigorous, credible, and trusted by AI labs and enterprise customers. The quality of these benchmarks directly shapes how the world measures and improves AI agent performance.
What You'll DoDesign, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
Partner with subject-matter experts to define realistic workflows and translate them into evaluation tasks and criteria.
Build and operate reliable infrastructure to run models and agents against benchmark tasks at scale.
Develop metrics and statistical analyses to measure benchmark difficulty, reliability, and failure modes.
Validate that benchmark performance correlates with real-world evaluations and frontier lab expectations.
Write clear technical documentation and benchmark reports for research and engineering audiences.
2–4 years of experience in research engineering, ML engineering, or related roles — with a focus on AI benchmarks, evaluation infrastructure, or agent environments.
Strong proficiency in Python, Docker, and Linux environments for building research or production infrastructure.
Experience designing and running benchmarks or evaluation environments for AI agents or large language models.
Experience developing metrics, statistical analyses, or validation studies to assess benchmark quality and real-world correlation.
Ability to collaborate with domain experts to translate complex workflows into well-scoped evaluation tasks.
Deep intuition for what makes a benchmark realistic, reliable, and practically useful.
Strong attention to detail with a habit of catching subtle inconsistencies and edge cases.
Comfort working independently in fast-moving, unstructured, early-stage environments.
Excellent written communication skills; able to make technical results legible across time zones and teams.
Nice to have: Published papers or technical writing on AI benchmarking, model evaluation, or failure modes; experience with RL training pipelines or data generation; background at frontier AI labs, research institutions, or on widely used public benchmark projects.
Salary range: $150,000 – $250,000 USD annually . Equity offered. Visa sponsorship is available.
LocationOn-site in San Francisco, CA, USA . In-person presence is expected.
- ...Responsibilities Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning.... ...generation, augmentation, and curation. Collaborate with AI researchers, applied AI teams, and data producers to align evaluations...SuggestedFull timeWork at officeRelocation package
- A leading technology company located in San Francisco is seeking a Machine Learning Research Engineer to design and develop safe AI benchmarking methodologies. This role involves collaboration with various teams to implement responsible evaluation techniques. Candidates...Suggested
$200k - $275k
...applied AI and data company, seeks a lead researcher to head in-house post-training and research... ...train models, publish findings, drive benchmarks, contribute to platform tooling, and help build a high-performance engineering culture from the ground up, with a focus on...SuggestedRelocation- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. In this six-week, part-time role (20+ hours/week), you will source material, write prompts...SuggestedPart timeImmediate start
$200k - $225k
About AlembicAlembic is where top engineers are solving marketing's hardest problem: proving... ...to leverage advanced analyticsDocument research and implementation decisions for... ...ensures new hires earn at or above real-time benchmarks. Ownership: Our generous equity program...Suggested$110.7k - $379.2k
Position Summary Research Engineer — Post-Training & Small Language Models (SLMs), Healthcare AI Three hundred fifty million Americans... .... • Build healthcare-grade evaluation — held-out clinical benchmarks, deployment regression gates, calibration and uncertainty,...Local areaVisa sponsorship$150k - $250k
...goods, and global social organizations.We research and deploy technologies that power AI-... ...What We Are Looking ForAt Distyl, Research Engineers build the bridge between frontier AI... ...why behavior changes, not just whether a benchmark improvesAI Systems Mindset: You understand...Work at office3 days per week- ...troubleshooting have become a massive tax of engineering velocity. Resolve AI is solving this by... ...workflows end-to-end, balancing research and engineering to create production-ready... ...and evaluation Design and execute benchmarks to evaluate AI models, improve performance...Work at officeVisa sponsorshipFlexible hours
$9.7k - $19k
...Center for AI Safety (CAIS) is a leading research and advocacy organization focused on... ...AI CEOs. The Role As a research engineer intern here, you will work very closely... ...security, machine ethics, AI alignment, and benchmarking AI risks. We will assign you a...Full timeInternshipLocal area$180k - $280k
...Research Engineer SuperAnnotate helps the world's leading AI teams build responsible, next-generation models powered by high-quality human... ...direction and run with it – finding the right papers, benchmarks, and prior work, reimplementing what's relevant, and building...Full time$180k - $340k
...Research Engineer You'll own the quality of AI across everything Gamma creates. As our Research Engineer, you'll design evaluation frameworks... ...that enable rapid testing, validate changes against quality benchmarks, and ensure our AI gets smarter with every iteration. You'll...Full timeWork at officeWork from home$210k - $275k
...Research Engineer You'll build the evaluation systems that tell us whether Firecrawl actually works. That sounds simple. It isn't. Our... ...isn't an eval role where you inherit a framework and run benchmarks. You'll design the metrics, build the pipelines, generate the...Full timeTemporary workFor contractorsRemote workVisa sponsorshipFlexible hours- ...Research Engineer Datacurve provides the frontier coding data that powers the world's most advanced models. We absorb and standardize... ...different surfaces of the same domain. You will produce the benchmarks, artifacts, and technical narratives that define our work. You...Shift work
- ...Research Engineer We believe that software is the foundation of modern civilization - yet vulnerabilities threaten its integrity, security... ...with strong intuition, experience in model evaluation, and benchmarks. Reinforcement Learning experience is a plus. Your work will...Full timeWork at office
$200k - $350k
...training), second-time technical founders, engineers that made 100+ games for Voodoo,... ...engaging games & 3D environments. Our current research spans:Distributed multi-agent... ...orchestration and engagement modeling.Define new benchmarks for fun, retention, and interactive intelligence...Visa sponsorshipRelocation package- ...Job Description Job Description Research Engineer — AI Alignment & Evaluation AI Safety / Research Engineering | San Francisco, CA |... ...Nice to Have Experience building evaluation frameworks, benchmarks, simulation environments, or agent-based systems. Exposure...Full timeWork at officeRelocationVisa sponsorship
$150k - $250k
...ownership role on a small, high-caliber engineering team building infrastructure for AI... ...partner problems quickly. Coordinate with research and go-to-market teams to keep... ...environments. ~ Experience working on benchmarks and evals — with solid judgment about task...RelocationVisa sponsorship- Fundamental AI Research Institute Come join one of the only research institutions globally with resources to compete with top AI companies... ...developing agentic and evaluation harnesses Develop benchmarks and evaluations for reasoning and agentic capabilities Contribute...
- ...building Agentic AI that empowers software engineers by automating production engineering and... ...powered workflows end‑to‑end, balancing research and engineering to create production‑... ...training and evaluation Design and execute benchmarks to evaluate AI models, improve...Full timeWork at officeVisa sponsorshipFlexible hours
$165k - $310k
Senior Research Engineer, LLM Training & Post-Training New York, New York, United States; Remote; San Francisco, California, United States... ...performance bottlenecks. Design evaluation methodologies, benchmark models, analyze failure modes, and guide model improvements...For contractorsFor subcontractorWork at officeRemote workWork from homeFlexible hours2 days per week$264.8k - $331k
...complex agents in enterprises around the world. The Enterprise ML Research Lab works on the front lines of this AI revolution. We are... ...algorithms to real life enterprise datasets across our clients + benchmarks. This will involve creating best-in-class Agents that achieve...Full time$200k - $350k
...deeply curious—building at the intersection of research, product, and creativity . The Role As a Machine Learning Research Engineer , you’ll own end-to-end research cycles—... ...(writing, design, visual style) Develop benchmarks and evaluation methods for subjective tasks...$140k - $200k
...Center for AI Safety (CAIS) is a leading research and advocacy organization focused on... ...introducing the first state-of-the-art benchmarks for measuring it. More recently, we've... ...policymakers. About the role As a Research Engineer (RE) or Research Scientist (RS) at CAIS,...Work at officeLocal area- ...harnessing, and deployment—and connect that research to the patients, clinicians, and real‑... ...improvements rather than just higher benchmark scores. Create meaningful, trustworthy,... ...decisions. Work closely with researchers, engineers, clinicians, and product teams to bring...Work at officeRelocation package
- ...safe, and under control. The Role We are seeking a Staff Research Engineer, AI/ML & Cybersecurity to serve as a core technical pillar... ...observability, and reproducibility Develop model evaluation, benchmarking, and stress-testing workflows Cybersecurity & Governance Conduct...
$200k - $350k
...stage AI company in San Francisco is seeking a Machine Learning Research Engineer to own end-to-end research cycles. The role involves training models across creative domains, developing evaluation benchmarks, and collaborating with creative experts and AI labs. The...$197.3k - $313.7k
...TeamSalesforce AI is looking for talented software and platform engineers to embed in our AI team to bridge the gap between frontier AI... ...where your engineering skills directly enable world-class research and products used by millions?At Salesforce, we are driving the...Full time- Factory is seeking innovative Research Engineers to design and integrate advanced AI and ML capabilities that revolutionize productivity and accelerate innovation within software organizations.What you will do and achieve:Design, develop, and deploy AI-driven agentic systems...Work at office
$164.6k - $313.3k
...s Sound Design AI group (SODA) is looking for a driven Data/ML engineer to push the boundaries of audio GenAI. Join the team behind Firefly... ...in Adobe products.We’re a small, collaborative and efficient research team looking for highly motivated candidates of all levels with...Full timeTemporary workLocal areaWorldwide$150k - $250k
...Job Description Job Description About the Role This is a Research Engineer role focused on building synthetic data pipelines for AI... ...designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language...Remote workVisa sponsorship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Engineer, Benchmarks. Be the first to apply!
- ai research engineer San Francisco, CA
- research assistant engineering San Francisco, CA
- senior research engineer San Francisco, CA
- research programmer San Francisco, CA
- research software engineer San Francisco, CA
- deep learning research engineer San Francisco, CA
- research engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- research economist San Francisco, CA
- education policy research San Francisco, CA


