Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff AI Research Engineer — Frontier Model Evaluations

Intelligence

Intelligence, the product lab behind DesignArena, invites you to join a talent-dense team in San Francisco with 5.5M+ users and a rapidly growing platform. You will define how frontier AI models are measured, design new benchmarks, run experiments, and publish analyses that become industry gold standards. We sponsor visas and relocation; compensation is competitive with meaningful equity. #J-18808-Ljbffr Intelligence

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Staff AI Research Engineer — Frontier Model Evaluations in San Francisco, CA vacancy
  • $136.44k - $265.11k

     ...rebuilding biotech for the AI era.When a...  ...run AI agents and models directly in their...  ...focused on making frontier AI models better at...  ...the datasets, evaluations, and systems that...  ...intersection of software engineering, biology, and...  ...scientists, and external research partners.Desire to... 
    Suggested
    Work at office
    Local area
    Monday to Friday
    Shift work

    Benchling

    San Francisco, CA
    4 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. The role emphasizes material sourcing, prompt design, and robust grading criteria against... 
    Suggested
    Part time
    Immediate start

    Obsidian

    San Francisco, CA
    5 days ago
  • A leading AI research organization in San Francisco is seeking a Research Engineer focused on pushing the boundaries of frontier models and AGI/ASI measurement. Candidates should have strong engineering...  ...AI measurements and build evaluation environments that drive progress.... 
    Suggested

    OpenAI

    San Francisco, CA
    4 days ago
  • Obsidian is partnering with an AI research initiative to support a Frontier Code Agents project in San Francisco. You will evaluate frontier AI coding models by performing realistic data engineering tasks, reviewing ETL pipelines, data warehouses, and distributed systems... 
    Suggested

    Obsidian

    San Francisco, CA
    1 day ago
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing infrastructure engineering tasks and reviewing model-generated implementations on cloud platforms, Kubernetes,... 
    Suggested

    Mercor Inc

    San Francisco, CA
    2 days ago
  •  ...seeking an exceptional health AI researcher to build frontier capabilities that translate into...  ...within health-focused models and products. The candidate will...  ...pretraining, RL/post-training, and evaluation, working with researchers, engineers, clinicians, and product teams... 

    AI Chopping Block

    San Francisco, CA
    2 days ago
  •  ...Analysis is seeking a Member of Technical Staff to design frontier evaluations for language models and publish results used by AI labs and enterprises. You will build datasets...  ...while collaborating with leading researchers and #J-18808-Ljbffr Artificial Analysis,... 
    Worldwide
    Flexible hours

    Artificial Analysis, Inc.

    San Francisco, CA
    2 days ago
  • HeyMilo AI is hiring a Research Engineer to join our Applied AI team in San Francisco. You’ll design and build reinforcement learning...  ..., creating simulators, reward functions, and evaluation harnesses to measure model performance. The role blends ML research with systems... 

    HeyMilo AI

    San Francisco, CA
    4 days ago
  • Anthropic in San Francisco seeks a bio safety researcher to design and run capability evaluations for biology-focused models, build and curate datasets for safety...  ...classifiers, and iterate on those classifiers with ML engineers. You will work at the intersection of applied... 

    Anthropic

    San Francisco, CA
    2 days ago
  • $200k - $400k

    Simile is seeking a Member of Technical Staff, Model Evaluations, to develop measurement systems that gauge the accuracy of simulations of human behavior. You will design evaluation metrics and collaborate with modeling teams to ensure models are effective. Ideal candidates... 

    Simile

    San Francisco, CA
    2 days ago
  • OpenAI is seeking an exceptional researcher to build frontier health capabilities and translate them into measurable impact...  ...experimentation and deployment into frontier models and products. The candidate should have strong ML/AI research depth, be a hands‑on builder, and... 

    OpenAI

    San Francisco, CA
    2 days ago
  • A leading AI platform company in San Francisco is seeking a software engineer to drive model integrations and deliver high-quality evaluations. The ideal candidate will have 4+ years of software...  ...mindset. You will work closely with researchers and product teams to solve high... 

    Arena Intelligence, Inc.

    San Francisco, CA
    5 days ago
  • OpenAI Health team is seeking an exceptional researcher to build frontier health capabilities and scale impact. You will turn underdefined problems...  ...experiments, and drive work from concept to measurable model improvements used across products. We value research excellence... 

    Triwill Group

    San Francisco, CA
    4 days ago
  • $150k - $250k

    About Distyl AI Distyl is an applied AI...  ...operations for the frontier of AI. Our...  ...organizations.We research and deploy technologies...  ...AI systems using Evaluation-Driven Development...  ...production.AI Evaluation Engineers focus on designing...  ..., agent logic, model selection, and release... 
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    4 days ago
  • $214k - $285k

     ...About the RoleHex is an AI-powered platform for...  ...Analytics workflows. AI Research Engineers at Hex partner with...  ...experiments, fine-tune models, deploy AI infrastructure...  ...it's pushing the frontiers of SQL generation quality or using LLMs to evaluate agent threads, your job... 
    Full time
    Work at office
    Flexible hours

    HEX

    San Francisco, CA
    5 days ago
  • $227.2k - $284k

     ...develop reliable AI systems for the...  ...most advanced models — fueling...  ...combining rigorous evaluation with full-stack...  ...on pushing the frontier of what agentic...  ...with applied ML research, design, and evaluation...  ...the RoleAs a Staff Machine Learning Research Engineer, you will... 
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  •  ...AI Research Engineer (Robot Learning) San Francisco AI & Software In office Full-...  ...Engineer (Robot Learning) you will drive frontier AI model development and data flywheel...  ...collection, model training and model evaluation in the real world. As an early employee... 
    Full time
    Work at office
    Immediate start

    Software Engineering, Data Science

    San Francisco, CA
    more than 2 months ago
  • $130k - $220k

     ...leading independent AI benchmarking...  .... They help engineers, enterprises, investors...  ...just measure frontier AI — they...  ...Member of Technical Staff** **What This...  ...as an AI Evaluation Engineer / Technical...  ...not a pure ML research position. It combines...  ..., foundation models, and hardware... 
    Full time
    Worldwide

    Aurora Jobs ApS

    San Francisco, CA
    more than 2 months ago
  •  ...Job Description Job Description Research Engineer — AI Alignment & Evaluation AI Safety / Research Engineering | San Francisco, CA | Hybrid / In...  ...research organization working at the intersection of frontier model evaluation, AI safety, and security. The team develops... 
    Full time
    Work at office
    Relocation
    Visa sponsorship

    W3 Sourcing

    San Francisco, CA
    11 days ago
  •  ...re building next‑generation AI systems that help military...  ...planners explore, compare, and evaluate operational courses of action. Our work combines frontier language models, simulation, planning, and...  .... As an Applied AI Research Engineer, you’ll focus on human‑machine... 
    Remote work
    Relocation package
    Flexible hours

    Code Metal

    San Francisco, CA
    2 days ago
  • $400 per month

     ...Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical...  ...focuses on realistic infrastructure engineering workflows and model evaluation.... 

    Mercor Inc

    San Francisco, CA
    3 days ago
  • Obsidian is seeking a candid evaluator to assess the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train frontier AI labs. You will review repository-level...  ...-based assessments to guide model training and #J-18808-Ljbffr Obsidian

    Obsidian

    San Francisco, CA
    3 days ago
  • Ephapsys seeks a Staff Research Engineer to advance the SDK and platform, focusing on production-grade AI systems, secure model deployment, and governance. You will work with the founder to translate AI research into hardened, scalable infrastructure. You will lead AI/... 

    Ephapsys

    San Francisco, CA
    4 days ago
  •  ...Analysis, Inc. in San Francisco is seeking a Member of Technical Staff (Applied AI Research). You will design novel frontier evaluations, build datasets and infrastructure, and evaluate every major model as it is released. This applied research role directly influences... 

    Artificial Analysis, Inc.

    San Francisco, CA
    4 days ago
  • An innovative tech company in New York is seeking a Research Scientist focused on Frontier Risk Evaluations. The ideal candidate will contribute to designing and creating evaluation measures for assessing AI risks and will have strong experience in machine learning and... 

    Scale AI

    San Francisco, CA
    4 days ago
  •  ...invites applications for a Research Scientist to design novel benchmarks and evaluate frontier language models and agents. You will lead research...  ...collaborate with research engineers, foundation-model developers...  ...that influence how leading AI #J-18808-Ljbffr Cerebro
    Relocation

    Cerebro

    San Francisco, CA
    5 days ago
  • Scale Labs seeks a Research Scientist focused on Frontier Risk Evaluations to design evaluation measures, harnesses and datasets for measuring risks posed by frontier AI systems. You will build harnesses to test models, collaborate with government agencies to scope evaluations... 

    Scale

    San Francisco, CA
    2 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific...  ...leading AI labs. You will author original, executable research problems that frontier models cannot solve. Engagement lasts 6 weeks, part-time with... 
    Part time
    Immediate start

    Mercor

    San Francisco, CA
    2 days ago
  •  ...hiring a Member of Technical Staff for the Frontier Data team. You will build RL environments, evaluations, and datasets that close...  ...gaps in frontier models, while developing scalable...  ...You’ll work directly with researchers and engineers to shape dataset development... 

    Roboflow

    San Francisco, CA
    2 days ago
  •  ...Francisco, California. The Role: As a Research Engineer - Language Model Pre-Training , you'll shape our...  ...Dataset collection, processing, and evaluation Architecture and methodology...  ...enjoy what we do and love discussing AI Benefits and Perks: Comprehensive... 
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    12 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff AI Research Engineer — Frontier Model Evaluations. Be the first to apply!