Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff AI Research Engineer — Frontier Model Evaluations

Intelligence

Intelligence, the product lab behind DesignArena, invites you to join a talent-dense team in San Francisco with 5.5M+ users and a rapidly growing platform. You will define how frontier AI models are measured, design new benchmarks, run experiments, and publish analyses that become industry gold standards. We sponsor visas and relocation; compensation is competitive with meaningful equity. #J-18808-Ljbffr Intelligence

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff AI Research Engineer — Frontier Model Evaluations in San Francisco, CA vacancy
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models through structured technical assessments and focus on realistic data engineering workflows. The role involves reviewing model... 
    Suggested

    Mercor

    San Francisco, CA
    2 days ago
  • A leading AI research organization in San Francisco is seeking a Research Engineer focused on pushing the boundaries of frontier models and AGI/ASI measurement. Candidates should have strong engineering...  ...AI measurements and build evaluation environments that drive progress.... 
    Suggested

    OpenAI

    San Francisco, CA
    2 days ago
  •  ...located in San Francisco is seeking an innovative Quality Engineer for their AI products. This role blends ops, strategy, and analytics to...  ...leading labs, and ensure user satisfaction through effective evaluation baselines. Competitive salary and benefits offered, with a... 
    Suggested

    Notion

    San Francisco, CA
    1 day ago
  • Anthropic in San Francisco seeks a bio safety researcher to design and run capability evaluations for biology-focused models, build and curate datasets for safety...  ...classifiers, and iterate on those classifiers with ML engineers. You will work at the intersection of applied... 
    Suggested

    Anthropic

    San Francisco, CA
    2 days ago
  • HeyMilo AI is hiring a Research Engineer to join our Applied AI team in San Francisco. You’ll design and build reinforcement learning...  ..., creating simulators, reward functions, and evaluation harnesses to measure model performance. The role blends ML research with systems... 
    Suggested

    HeyMilo AI

    San Francisco, CA
    3 days ago
  • $350k

    Mirendil is seeking research engineers in San Francisco to build the post-training stack for frontier reasoning models. You will engage in designing experiments and iterating on reinforcement...  ...scalable infrastructure for large-scale AI research. The position offers a base... 

    Mirendil

    San Francisco, CA
    5 days ago
  • $350k

    Mirendil is a tech-first company located in San Francisco, California, seeking a Research Engineer to develop evaluation infrastructure for AI models. You'll design frameworks that measure model capabilities, implement automated pipelines, and create workflows for inspecting... 

    Mirendil

    San Francisco, CA
    5 days ago
  • $200k - $400k

    Simile is seeking a Member of Technical Staff, Model Evaluations, to develop measurement systems that gauge the accuracy of simulations of human behavior. You will design evaluation metrics and collaborate with modeling teams to ensure models are effective. Ideal candidates... 

    Simile

    San Francisco, CA
    1 day ago
  • $150k - $250k

    About Distyl AI Distyl is an applied AI...  ...operations for the frontier of AI. Our...  ...organizations.We research and deploy technologies...  ...AI systems using Evaluation-Driven Development...  ...production.AI Evaluation Engineers focus on designing...  ..., agent logic, model selection, and release... 
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    3 days ago
  • Luma AI in the San Francisco Bay Area and beyond is seeking a Research Scientist / Engineer to shape the foundation of multimodal AI. You will bridge frontier research with shipped products like Dream Machine...  ...GPU clusters, rigorous evaluation, and a fast, reproducible... 
    Remote job

    Luma AI

    San Francisco, CA
    5 days ago
  •  ...AI Research Engineer (Robot Learning) San Francisco AI & Software In office Full-...  ...Engineer (Robot Learning) you will drive frontier AI model development and data flywheel...  ...collection, model training and model evaluation in the real world. As an early employee... 
    Full time
    Work at office
    Immediate start

    Software Engineering, Data Science

    San Francisco, CA
    more than 2 months ago
  • $130k - $220k

     ...leading independent AI benchmarking...  .... They help engineers, enterprises, investors...  ...just measure frontier AI — they...  ...Member of Technical Staff** **What This...  ...as an AI Evaluation Engineer / Technical...  ...not a pure ML research position. It combines...  ..., foundation models, and hardware... 
    Full time
    Worldwide

    Aurora Jobs ApS

    San Francisco, CA
    a month ago
  • $214k - $285k

     ...the Role Hex is an AI-powered platform for modern...  ...workflows. AI Research Engineers at Hex partner with product...  ..., fine-tune models, deploy AI infrastructure...  ...Whether it's pushing the frontiers of SQL generation quality or using LLMs to evaluate agent threads, your job... 
    Full time
    Work at office
    Flexible hours

    HEX

    San Francisco, CA
    1 day ago
  • $238k - $302k

     ...simulation across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in Large Language...  ...critical. We are looking for quantitatively-minded engineers to research and propose new ways to assess the ML models... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $85 per hour

     ...technical talent with leading AI research labs. Headquartered in...  .... Position: iOS Engineer (Coding Agent Experience)...  ...Responsibilities Use frontier AI coding agents to complete and evaluate complex engineering tasks. Review model-generated mobile application... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    13 days ago
  • Cacheflow is seeking a Senior Applied Research Engineer to enhance the effectiveness of our AI systems through focused research and experimentation. This role involves designing information retrieval strategies and collaborating with engineers to turn validated approaches... 
    Flexible hours

    Cacheflow

    San Francisco, CA
    1 day ago
  •  ...based in San Francisco, is advancing frontier AI through design arena benchmarks and large-scale evaluation platforms. You will invent scalable...  ...improve supervision for multimodal models, with opportunities to collaborate with researchers on forward-deployed strategies. The... 

    Intelligence

    San Francisco, CA
    1 day ago
  • Halluminate, an applied AI and data company in San Francisco,...  ...leader to drive our in-house research and post-training efforts. You will train frontier models, evaluate quality, and build benchmarks...  ...platform tooling, and help grow the engineering team. Expect startup speed,... 

    Halluminate

    San Francisco, CA
    5 days ago
  • $320k

     ...interpretable, and steerable AI systems. We want...  ...of committed researchers, engineers, policy experts,...  ...the Team The Frontier Red Team (FRT) is...  ...the year where models reach expert-level...  ...to elicit and evaluate autonomous AI cyber...  ...Currently, we expect all staff to be in one of... 
    Work at office
    Relocation
    Visa sponsorship
    Flexible hours

    Neura Market

    San Francisco, CA
    3 days ago
  •  ...Analysis, Inc. in San Francisco is seeking a Member of Technical Staff (Applied AI Research). You will design novel frontier evaluations, build datasets and infrastructure, and evaluate every major model as it is released. This applied research role directly influences... 

    Artificial Analysis, Inc.

    San Francisco, CA
    3 days ago
  • Ephapsys seeks a Staff Research Engineer to advance the SDK and platform, focusing on production-grade AI systems, secure model deployment, and governance. You will work with the founder to translate AI research into hardened, scalable infrastructure. You will lead AI/... 

    Ephapsys

    San Francisco, CA
    3 days ago
  • $400 per month

     ...Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical...  ...The work focuses on realistic data engineering workflows and model evaluation.... 

    Mercor

    San Francisco, CA
    2 days ago
  •  ...passion for discovering how AI can solve real-world...  ...works to advance the frontiers of human-machine interaction...  ...computing. Our applied research explores ways that...  ...intellectually curious AI research engineer with strong technical...  ...with the latest AI models, Agentic frameworks,... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    San Francisco, CA
    3 days ago
  • United States Digital Space LLC is seeking a bio safety researcher to design and run capability evaluations for frontier models in biology, curate training data for safety classifiers, and collaborate with ML engineers to ensure robustness and practical deployment. The... 

    United States Digital Space LLC

    San Francisco, CA
    4 days ago
  • OpenAI is seeking a researcher to advance frontier evaluations and environments for safe AGI/ASI. You will help design north star model environments and steer major training runs so that research...  .... Collaborate with researchers, engineers, product and safety teams to decide... 

    Neura Market

    San Francisco, CA
    5 days ago
  • $216k - $270k

    Scale AI, Inc. is looking for a Research Scientist specializing in Frontier Risk Evaluations to develop measures for assessing risks of advanced AI systems. In this role, you will design testing harnesses, collaborate with agencies, and publish reports to inform policymakers... 

    Scale AI, Inc.

    San Francisco, CA
    3 days ago
  • Prime Intellect is building an open frontier AI platform unifying compute, environments, evaluations, and deployment for researchers and engineers worldwide. This generalist software engineering...  ...used to train and deploy frontier models. You will own features end-to-end from... 
    Worldwide

    AI Chopping Block

    San Francisco, CA
    1 day ago
  •  ...leading independent AI benchmarking...  ...We support labs, engineers and enterprises to...  ...actively shaping the frontier. Our benchmarks and...  ...Opportunity Language model evaluation is the sharpest...  ...Members of Technical Staff to build the next...  ...with their research teams; our commercial... 

    Artificial Analysis

    San Francisco, CA
    3 days ago
  • $350k

     ...goal is to democratize frontier AI R&D across scientific disciplines...  ...building a frontier AI research company and training our own models end-to-end. Our work...  ...researchers and engineers from Anthropic, Google DeepMind...  ...engineer to build the evaluation infrastructure that tells... 

    Mirendil

    San Francisco, CA
    5 days ago
  •  ...Origin is building Physical AI for the built world - starting...  ...building Construction Action Models which allows our modular robots...  ...deployment on Jetson AGX Orin. Every research project will have a deployment...  ...detection. Design offline evaluation metrics that predict real-... 

    Origin

    San Francisco, CA
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff AI Research Engineer — Frontier Model Evaluations. Be the first to apply!