Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Scientist, AI Evaluation & Benchmarking

Arena Intelligence, Inc.

A leading AI evaluation firm based in San Francisco seeks a Machine Learning Scientist to foster understanding of AI model performance. You'll engage in designing and analyzing comprehensive experiments while collaborating across teams. Applicants should possess a PhD in a relevant field and hands-on experience with large-scale models. The role offers competitive compensation and comprehensive benefits, fostering a culture centered around transparency and community impact. #J-18808-Ljbffr Arena Intelligence, Inc.

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the ML Scientist, AI Evaluation & Benchmarking in San Francisco, CA vacancy
  • $196k - $230k

     ...AreNotion is the collaborative AI workspace where teams and...  ...Researcher to define and scale how we evaluate Notion’s AI-powered...  ...help teams spot regressions, benchmark improvements, and understand when...  ...and working with Data Science/ML partners on measurement strategy... 
    Suggested
    Local area
    Shift work

    Notion Labs

    San Francisco, CA
    3 days ago
  • $216k - $270k

    Scale Labs, Research Scientist — AI Controls and MonitoringAs the leading data and evaluation partner for frontier AI companies, Scale...  ...to establish standards and benchmarks for AI monitoring and escalation...  ...experience addressing sophisticated ML problems, whether in a... 
    Suggested
    Full time

    Scale AI

    San Francisco, CA
    7 hours ago
  • $150k - $250k

     ...signal training data and evaluation infrastructure for frontier AI labs, with a founding team...  ...that go beyond static benchmarks. Small team where individual...  ...people · Industry: AI / ML — training data & evaluation...  ...with the other Research Scientists to build shared... 
    Suggested
    Full time
    Visa sponsorship
    Shift work

    David Joseph & Company

    San Francisco, CA
    5 days ago
  • $216k - $270k

    Scale Labs, Research Scientist - AI Controls and Monitoring As the leading data and evaluation partner for frontier AI companies, Scale...  ...to establish standards and benchmarks for AI monitoring and escalation...  ...experience addressing sophisticated ML problems, whether in a... 
    Suggested
    Full time

    Scale AI, Inc.

    San Francisco, CA
    3 days ago
  • $150k - $250k

    About Distyl AI Distyl is an applied AI technology company partnering with the world...  ...to drive incremental improvements on benchmarks or optimize an existing process but instead...  ...is measured. Researchers design evaluation frameworks that capture reasoning depth,... 
    Suggested
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    3 days ago
  • $180k - $260k

     ...About the RoleWe’re looking for an Applied Scientist, AI to turn messy, high-stakes healthcare...  ...ll build strong baselines, design honest evaluations, run careful error analysis, and iterate...  ...also be able to partner closely with ML engineering to productionize models, work... 
    Temporary work
    Work at office
    Monday to Friday
    Monday to Thursday

    Sprinter Health

    San Francisco, CA
    2 days ago
  • $141.1k - $262.1k

     ...makes us Roche.Advances in AI, data, and computational...  ...Intelligence (AI) to assist our scientists in both pRED and gRED to...  ...-edge machine learning (ML) techniques. We are...  ...objectives, training signals, and evaluation criteria.Evaluation & Benchmarks: Design and implement... 
    Full time
    Work experience placement
    Local area
    Worldwide
    Relocation package

    Genentech

    South San Francisco, CA
    3 days ago
  • $167.4k - $310.8k

     ...makes us Roche.Advances in AI, data, and...  ...Intelligence (AI) to assist our scientists in both pRED and gRED to...  ...edge machine learning (ML) techniques. We are...  ...training strategies, and evaluation methodologies.Model Capability...  ...of rigorous reasoning benchmarks.You act as a technical... 
    Full time
    Local area
    Worldwide
    Relocation package

    Genentech

    South San Francisco, CA
    3 days ago
  • $160k - $220k

     ...the RoleWe’re looking for an AI Research Scientist to advance the...  ...architectures, new training or evaluation techniques, long-horizon research...  ...validation standards are higher than benchmark culture alone, and you are...  ...and engineers on rigorous ML research practices.External... 
    Temporary work
    Work at office
    Monday to Friday
    Monday to Thursday

    Sprinter Health

    San Francisco, CA
    7 hours ago
  •  ...applications, processes, and AI into a single, governed...  ...AI Research Scientist to join our growing team...  ...optimised RAG, tool‑use evaluation, and multi‑agent collaboration.Prototype and benchmark models; present findings...  ...PhD in Computer Science, ML, or related field—or equivalent... 
    Remote work
    Flexible hours

    Workato

    San Francisco, CA
    4 days ago
  • DataAnnotation is seeking a Clinical Data Scientist for a remote contract role to evaluate AI-generated quantitative analyses and create benchmark problems for training AI systems. You will assess AI outputs, develop training problems across forecasting, experiment design... 
    Remote job
    Contract work

    DataAnnotation

    San Francisco, CA
    2 days ago
  • $234.3k - $349k

     ...enterprises orchestrate AI-powered work. Our...  ...world. As an AI research scientist, you'll be at the center...  ...through model training, evaluation, and production deploymentDesign...  ...novel evaluation benchmarks and methodologies that...  ...7+ years of hands-on ML research experience, with... 
    Full time
    Work at office
    Local area

    Writer

    San Francisco, CA
    7 hours ago
  •  ...looking for a talented researcher to work on improving AI tutoring systems. You will shape experiments, evaluate tutoring quality, and turn pedagogical principles...  ...candidates have a PhD or equivalent experience in CS, ML, or related fields. Responsibilities include... 
    Contract work
    Summer work

    Heyaristotle

    San Francisco, CA
    4 days ago
  • $84.13 - $91.34 per hour

    AI Researcher - Efficient AI (Contractor) Step into the innovative...  ...workflows. • Propose and evaluate novel compression methods (PTQ...  ...vision, reasoning, and agentic benchmarks. • Contribute to publications,...  ...or engineering experience in ML, efficient AI, model optimization... 
    Full time
    Contract work
    Temporary work
    For contractors
    Local area
    Immediate start

    LG Electronics

    San Francisco, CA
    5 days ago
  • $200k - $325k

     ...anything. We're building the AI that finally changes that. Ivo...  ...accurate on legal- specific tasks. Evaluate emerging work in agentic...  ...Design and maintain datasets, benchmarks, and evals for training and measuring...  ...at top venues — e.g., ML/AI conferences (NeurIPS, AAAI,... 
    Contract work
    Work at office
    Immediate start
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Ivo

    San Francisco, CA
    4 days ago
  • We are seeking an Edge AI Research Scientist to develop next-generation speech and audio AI...  ...frameworks. Performance Evaluation Develop rigorous benchmarking methodologies for edge AI systems....  ...deployment frameworks such as: Core ML ExecuTorch ONNX Runtime LiteRT / TensorFlow... 

    Huxley

    San Francisco, CA
    3 days ago
  • $200k - $280k

     ...high‑performance computing for ML. Are comfortable working from...  ...‑scale rollout collection and evaluation cheaper. Use these pipelines...  ...as needed. Establish metrics, benchmarks, and experimentation...  ...engineering. About Together AI Together AI is a research-driven... 
    Full time

    Togetherai

    San Francisco, CA
    5 days ago
  • Carnaby Fox is seeking a Member of Technical Staff (AI Research) in San Francisco to help shape the...  ...with world-class researchers to design experiments, evaluate LLMs, and improve data quality for high-stakes AI benchmarks. The role emphasizes independent ownership,... 

    Carnaby Fox

    San Francisco, CA
    1 day ago
  • the company is seeking a Research Scientist to advance measurable recursive-self-improvement in large models. You will design evaluations, build models of capability growth, and interpret...  ..., and a track record in evaluating AI systems. #J-18808-Ljbffr United States... 

    United States Digital Space LLC

    San Francisco, CA
    4 days ago
  •  ...Department Technical About the Role Generative AI is transforming what's computationally...  ...path through these bottlenecks. As an ML Research Scientist, you'll work at the frontier of...  ...and likelihood estimation Develop and benchmark novel solver methods for diffusion ODEs... 
    Full time
    Casual work
    Visa sponsorship

    Wheel the World

    San Francisco, CA
    1 day ago
  •  ...looking for an exceptional Research Scientist to develop next-generation AI technologies, focusing on user...  ...applications. Design, prototype, evaluate, and deploy transformer-based generative...  ...and establish reproducible benchmarking pipelines. Work closely with product... 

    Predactiv

    San Francisco, CA
    1 day ago
  •  ...About the Role We’re looking for an Applied Scientist, AI to turn messy, high‑stakes healthcare...  ...build strong baselines, design honest evaluations, run careful error analysis, and iterate...  ...should also be able to partner closely with ML engineering to productionize models,... 
    Temporary work
    Work at office
    Relocation package
    Monday to Friday
    Monday to Thursday
    Flexible hours

    Sprinter Health

    San Francisco, CA
    5 days ago
  •  ..., we are building a team of world-class scientists, ML researchers, and engineers to work together...  ...frontier of model architectures for AI x Chemistry: developing world models for...  ...symmetries and constraints. Prototype, benchmark, and iterate rapidly to transform research... 
    Work at office

    Achira

    San Francisco, CA
    1 day ago
  •  ...is hiring a Senior Machine Learning Scientist to lead autonomous life science systems...  ...design architectures, workflows, and evaluation methods enabling AI to propose experiments, incorporate...  ...-on role sits at the intersection of ML, biology, and automation, requiring translating... 

    Lilasciences

    San Francisco, CA
    1 day ago
  • As a Research Scientist , you'll lead cutting-edge research that advances the state of generative AI for long-form storytelling. You'll work at the intersection...  ...horizon generation. Build Novel Evaluation Frameworks Design robust benchmarks and evaluation methodologies for... 
    Worldwide

    Pocket FM

    San Francisco, CA
    3 days ago
  •  ...infrastructure layer for enterprise AI agents, and we are hiring a Research Scientist to advance the neuro-symbolic...  ...Context Graph. Publish at top AI and ML conferences (NeurIPS, ICML, ICLR,...  ...components with full observability and benchmarking. Engage with the Bay Area academic... 
    Work at office
    Relocation

    Rippletide SAS

    San Francisco, CA
    4 days ago
  •  ...cutting-edge research into production-grade AI systems. You will design and build the...  ..., including retrieval, orchestration, and evaluation, and you’ll take techniques from research...  ...client-ready solutions. Ideal candidates ship ML-powered products end-to-end, balancing... 

    Crucibl

    San Francisco, CA
    3 days ago
  • Anthropic in San Francisco seeks a Research Scientist to measure recursive-self-improvement in large models. You will design evaluations, build models of capability growth, and...  ...RL, and policy teams to advance safe and reliable AI systems. #J-18808-Ljbffr Anthropic

    Anthropic

    San Francisco, CA
    5 days ago
  • $147k - $180k

    Profluent is an AI‑first protein design company. Founded...  ...Role We are expanding the ML Design Evaluation (MDE) program; the cross‑functional...  .... We are looking for a scientist with deep expertise in...  ...Contribute to campaign charters, benchmarking assay design, and post‑... 

    Profluent Bio

    Emeryville, CA
    3 days ago
  •  ...building quantum-accelerated AI servers to exponentially speed...  ...We are looking for a Research Scientist who can help define Quantum AI...  ...the intersection of frontier AI/ML, quantum algorithms,...  ...test, and refine hypotheses. Benchmarking frameworks that reveal when a... 
    Casual work
    Visa sponsorship

    Sygaldry

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Scientist, AI Evaluation & Benchmarking. Be the first to apply!