Research Engineer, Benchmarks
$150k - $250kClera
Job Description
Job Description
About the Role
Join a small, technically elite team — including International Olympiad medalists and published AI researchers — building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. As a Research Engineer, Benchmarks , you'll own the design and implementation of evaluations that frontier labs and enterprise customers trust. This role is central to ensuring our benchmarks are rigorous, credible, and tightly aligned with real-world agent performance.
This is an on-site role based in San Francisco, CA . Visa sponsorship is available.
What You'll DoDesign, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
Partner with subject-matter experts to define realistic workflows and tasks for domain-specific evaluations.
Build reliable infrastructure to run models and agents against benchmark tasks at scale.
Develop metrics and analyses that measure benchmark difficulty, reliability, and failure modes.
Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations.
Write clear documentation and benchmark reports that make results legible and credible to technical audiences.
Required
2–4 years of experience in software engineering, ML engineering, or research roles.
Strong proficiency in Python, Docker, and Linux environments.
Experience building environments, evaluations, or benchmarks for AI systems.
Published research or technical writing on topics such as public benchmarks, model failure modes, or evaluation methodology.
Deep understanding of what makes a benchmark realistic, reliable, and practically useful.
Curiosity and genuine ability to understand how real-world workflows operate across diverse domains.
Strong attention to detail — a habit of spotting subtle inconsistencies and edge cases in task design.
Ability to reason from first principles about task design, scoring, and failure modes.
Comfort thriving in unstructured problem spaces and working independently in fast-paced, early-stage environments.
Excellent communication skills for collaborating across time zones and with technical teams.
Salary: $150,000 – $250,000 USD annually, depending on experience.
Visa sponsorship available.
Opportunity for significant early-stage equity and career growth within a high-impact, research-driven team.
This is a full-time, on-site position in San Francisco, CA . Candidates must be willing and able to work in-office.
$150k - $250k
...Olympiad medalists, serial AI startup founders, and researchers with publications at top venues (ICLR, NeurIPS, and similar). We're looking for Research Engineers to work across agent quality control automation, benchmarks, and synthetic data — shaping how AI agents learn...SuggestedFull timeVisa sponsorship- ...teammates (we've accomplished a lot as just one engineer and one designer!) to a clan of around... ...with leading performance on relevant benchmarks, generating novel insights along the way... ...for your life". Specifically, in an AI research engineer role, we are looking for the...Suggested
$200k - $350k
...training), second-time technical founders, engineers that made 100+ games for Voodoo,... ...engaging games & 3D environments. Our current research spans: Distributed multi-agent... ...and engagement modeling. Define new benchmarks for fun, retention, and interactive intelligence...SuggestedVisa sponsorshipRelocation package$140k - $200k
...Center for AI Safety (CAIS) is a leading research and advocacy organization focused on... ...introducing the first state-of-the-art benchmarks for measuring it. More recently, we've been... ...policymakers. About the role As a Research Engineer (RE) or Research Scientist (RS) at CAIS,...SuggestedWork at officeLocal area- ...Research Engineer On Physical Ai Team Hedra is a pioneering generative modeling company — first models to market — now building a Physical... ...action sequences Evaluate model performance using both benchmark datasets and real-world deployment metrics Contributions...SuggestedWork at office
- ...things we have to think about as well. If you want to have your research come into contact with reality, Ando is the place for it.... ...track, or intervene. Evaluating that judgment means building benchmarks where the ground truth includes silence. Existing agent benchmarks...Work from home
- ...Cerebro in San Francisco is seeking a Research Engineer for Post-Training & Reasoning to push the frontiers of AI research. You will develop... ..., experience with RL methods, LLM post-training, evaluation benchmarks, and PyTorch, and are excited by equity-backed early-stage...Work experience placement
- ...building Agentic AI that empowers software engineers by automating production engineering and... ...powered workflows end‑to‑end, balancing research and engineering to create production‑... ...training and evaluation Design and execute benchmarks to evaluate AI models, improve...Full timeWork at officeVisa sponsorshipFlexible hours
$200k - $350k
...curious—building at the intersection of research, product, and creativity . The Role As a Machine Learning Research Engineer , you’ll own end-to-end research cycles—... ..., design, visual style) Develop benchmarks and evaluation methods for subjective...$350k
...want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: You want to...Full timeWork at officeVisa sponsorshipFlexible hours$140k - $250k
...AI Training Data Infrastructure Research Engineer Location: San Francisco, CA Work Model: On-site Industry: AI training data infrastructure Compensation: $140K-$250K base, plus equity Our partner is a YC-backed company building a new kind of marketplace...Full timeWork at office$160k - $250k
...Join to apply for the Founding Research Engineer role at Adam Join to apply for the Founding Research Engineer role at Adam This range is provided by Adam. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range...Full time$225k
...believe the most promising path to safe AGI lies in automating research and code generation to improve models and solve alignment more... ...compute to achieve this goal. About the role As a Research Engineer, you'll work on training, evaluating, and serving large AI...RelocationVisa sponsorship- ...Research Engineer Factory is seeking innovative Research Engineers to design and integrate advanced AI and ML capabilities that revolutionize productivity and accelerate innovation within software organizations. What You Will Do And Achieve: Design, develop...Work at office
$120k - $200k
...We are actively seeking a Research Engineer specializing in Machine Learning and AI to play a pivotal role in pioneering advanced solutions. In this role, you will lead end-to-end research projects and contribute technical expertise to build scalable systems, all within...Casual workWork at office$200k - $400k
...Senior Research Engineer Decagon is the leading conversational AI platform empowering every brand to deliver concierge customer experiences. Our technology enables industry-defining enterprises like Avis Budget Group, Block's Cash App and Square, Chime, Oura Health,...Full timeWork at officeLocal area- ...Adam Founding Research Team Opportunity We're building the founding research team at Adam. At Adam, we're tackling a frontier problem: training AI models to intelligently interpret and edit parametric CAD in 3D space. This demands creativity, deep technical ability...
- ...generative models, concentrating on GPU-level optimizations and distributed systems. The role involves direct collaboration with researchers to implement efficient training solutions. Candidates should possess strong skills in PyTorch, experience with large-scale training...
- ...Mindset You pick work that teaches you something. You’re the kind of researcher who reads the paper, questions the setup, builds the missing system, and shows the work. Original thinking matters to you, but it is not enough on its own. You turn ideas into code, experiments...
- ...A technology company in San Francisco is looking for a candidate to drive research initiatives that influence engineering solutions. You'll build evaluations using real tool data, tackle search challenges for tools, and train systems for improved accuracy. Ideal candidates...
- ...Luma AI is seeking a Research Scientist/Engineer to advance multimodal agent models across research and product integrations. You will explore modeling, data, systems and evaluation to push state-of-the-art capabilities. Join a team driving large-scale training with PyTorch...
- ...reinventing life sciences the same way it reinvented software engineering, and Chai is at the forefront of this shift. Leading pharmaceutical... ...the road ahead. About the role We are seeking an AI Research Engineer to help design, train, evaluate, and optimize Chai's...Shift work
$120k - $250k
...Research Engineer London, England, United Kingdom; New York, New York, United States; San Francisco, California, United States; Seattle, Washington, United States Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end...Work at officeWork from homeFlexible hours2 days per week$225k - $400k
...Research Engineer Title of Role: Research Engineer Location: San Francisco, onsite Company Stage of Funding: Venture-Backed — Software Development, AI, Devtools, Data, Enterprise, B2B Office Type: Onsite Salary: $225K–$400K Company Description We're...Work at office$100k - $300k
...individuals who are eager to explore uncharted waters and contribute to our innovative projects. Position Overview We are hiring Research Engineers to develop scalable robotic systems aimed at achieving general-purpose robotic intelligence. In this role, you'll work with...Full time$264.8k - $331k
...training algorithms to reach the performance necessary for complex agents in enterprises around the world. The Enterprise ML Research Lab works on the front lines of this AI revolution. We are working on an arsenal of proprietary research, tools, and resources that...Full time$150k - $250k
...align AI models to real-world workflows. As a Forward Deployed Research Engineer , you'll own end-to-end resolution of urgent, ambiguous... ..., Docker, and Linux environments. ~ Experience working on benchmarks and evals, with sound judgment about what makes a task realistic...Remote workRelocationVisa sponsorship- ...2025. The Impact You'll Make Our research team is expanding to keep pace with a wave... ...edge of the field. As a Research Engineer, you'll take a research direction and run with it – finding the right papers, benchmarks, and prior work, reimplementing what's relevant...Full time
- ...Perception AI Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI agents that answer any question, anywhere in their environments. Our systems can automatically detect and reason about any physical...
$250k - $425k
...About the job ML Research Engineer Junior/Mid/Senior/Staff ML Research Engineer Posted by Transparent Search Group on behalf of Decagon . About Decagon Decagon builds the leading conversational AI platform that empowers brands to deliver concierge...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Engineer, Benchmarks. Be the first to apply!
- research engineer San Francisco, CA
- research programmer San Francisco, CA
- research software engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- deep learning research engineer San Francisco, CA
- ai research engineer San Francisco, CA
- research editor San Francisco, CA
- brain research San Francisco, CA
- research economist San Francisco, CA
- undergraduate summer research internship San Francisco, CA

