Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Scientist, AI Evaluation & Benchmarking

Arena Intelligence, Inc.

A leading AI evaluation firm based in San Francisco seeks a Machine Learning Scientist to foster understanding of AI model performance. You'll engage in designing and analyzing comprehensive experiments while collaborating across teams. Applicants should possess a PhD in a relevant field and hands-on experience with large-scale models. The role offers competitive compensation and comprehensive benefits, fostering a culture centered around transparency and community impact. #J-18808-Ljbffr Arena Intelligence, Inc.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the ML Scientist, AI Evaluation & Benchmarking in San Francisco, CA vacancy
  •  ...Fleet AI, Inc. is seeking a Research Scientist to join their core research team in San Francisco. This role focuses on investigating how environments...  ...labs. Key responsibilities include generating benchmarks to evaluate frontier models, automating environment... 
    Suggested

    Fleet AI, Inc.

    San Francisco, CA
    19 hours ago
  • $167.4k - $310.8k

     ...us Roche. Advances in AI, data, and computational...  ...(AI) to assist our scientists in both pRED and gRED to...  ...edge machine learning (ML) techniques. We are seeking...  ...strategies, and evaluation methodologies. Model Capability...  ...of rigorous reasoning benchmarks. You act as a... 
    Suggested
    Local area
    Worldwide
    Relocation package

    Genentech

    San Francisco, CA
    2 days ago
  • $196k - $230k

     ...Notion is the collaborative AI workspace where teams and agents...  ...Researcher to define and scale how we evaluate Notion’s AI-powered...  ...help teams spot regressions, benchmark improvements, and understand when...  ...and working with Data Science/ML partners on measurement strategy... 
    Suggested
    Local area
    Shift work

    Fixed Frames

    San Francisco, CA
    19 hours ago
  • $216k - $270k

    Scale Labs, Research Scientist - AI Controls and Monitoring As the leading data and evaluation partner for frontier AI companies, Scale...  ...to establish standards and benchmarks for AI monitoring and escalation...  ...experience addressing sophisticated ML problems, whether in a... 
    Suggested
    Full time

    Scale AI, Inc.

    San Francisco, CA
    4 days ago
  •  ...looking for a talented researcher to work on improving AI tutoring systems. You will shape experiments, evaluate tutoring quality, and turn pedagogical principles...  ...candidates have a PhD or equivalent experience in CS, ML, or related fields. Responsibilities include... 
    Suggested
    Contract work
    Summer work

    Heyaristotle

    San Francisco, CA
    19 hours ago
  • About Thorin Thorin is an applied AI company born out of 8VC’s...  ...architectures, training paradigms, and evaluation techniques tailored to...  ...Design, implement, and test ML / AI methods that improve model...  ...appropriate (writing, talks, shared benchmarks). What You’ll Bring Advanced... 

    8VC

    San Francisco, CA
    19 hours ago
  •  ..., we are building a team of world-class scientists, ML researchers, and engineers to work together...  ...frontier of model architectures for AI x Chemistry: developing world models for...  ...symmetries and constraints. Prototype, benchmark, and iterate rapidly to transform research... 
    Work at office

    Achira

    San Francisco, CA
    3 days ago
  • $160k - $280k

     ...time. We are a team of musicians and AI experts, including alumni from...  ...build and deploy our state of the art ML models trained with an H100/scientist ratio of >100x. Check out our Suno...  ...engineering, designing, training and evaluating machine learning models Track... 
    Work at office
    Flexible hours

    Menlo Ventures

    San Francisco, CA
    4 days ago
  •  ...Department Technical About the Role Generative AI is transforming what's computationally...  ...path through these bottlenecks. As an ML Research Scientist, you'll work at the frontier of...  ...and likelihood estimation Develop and benchmark novel solver methods for diffusion ODEs... 
    Full time
    Casual work
    Visa sponsorship

    Wheel the World

    San Francisco, CA
    2 days ago
  • Knowtex is seeking an ML Scientist (Research) to enhance its voice AI and clinical NLP capabilities within healthcare. You will develop and evaluate innovative machine learning solutions aimed at advancing medical speech recognition and clinical language understanding.... 

    Knowtex

    San Francisco, CA
    4 days ago
  •  ...infrastructure layer for enterprise AI agents, and we are hiring a Research Scientist to advance the neuro-symbolic...  ...Context Graph. Publish at top AI and ML conferences (NeurIPS, ICML, ICLR,...  ...components with full observability and benchmarking. Engage with the Bay Area academic... 
    Work at office
    Relocation

    Rippletide SAS

    San Francisco, CA
    19 hours ago
  •  ...mining to translate scientific insights into actionable AI solutions. Conduct experiments, evaluate model performance, and iterate rapidly to improve...  .... Hands‑on experience building and deploying scalable ML models on cloud platforms (e.g., AWS). Demonstrated ability... 

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    4 days ago
  • $141.1k - $262.1k

     ...makes us Roche. Advances in AI, data, and computational...  ...(AI) to assist our scientists in both pRED and gRED to...  ...cutting‑edge machine learning (ML) techniques. We are...  ...objectives, training signals, and evaluation criteria. Evaluation & Benchmarks: Design and implement... 
    Work experience placement
    Local area
    Worldwide
    Relocation package

    Genentech

    South San Francisco, CA
    4 days ago
  • $147k - $180k

    Scientist II, ML - Guided Protein Design Evaluation Emeryville, California, United States; Hybrid (2-3 days on‑site) Profluent is an AI‑first protein design company. Founded in 2022, we develop deep...  ...to campaign charters, benchmarking assay design, and post‑campaign... 

    Profluent

    Emeryville, CA
    4 days ago
  •  ...building quantum-accelerated AI servers to exponentially speed...  ...We are looking for a Research Scientist who can help define Quantum AI...  ...the intersection of frontier AI/ML, quantum algorithms,...  ...test, and refine hypotheses. Benchmarking frameworks that reveal when a... 
    Casual work
    Visa sponsorship

    Sygaldry

    San Francisco, CA
    2 days ago
  •  ...About the role AI research at WRITER isn't just...  ...world. As an AI research scientist, you'll be at the center...  ...model training, evaluation, and production deployment...  ...Build novel evaluation benchmarks and methodologies that...  ...need 7+ years of hands‑on ML research experience, with... 
    Full time
    Local area
    Flexible hours

    Writer Corporation

    San Francisco, CA
    4 days ago
  • At Goaly, our mission is to make custom AI affordable for every business. Our...  ...at scale. Design domain‑tailored eval benchmarks: Build evaluation frameworks that capture real‑world performance...  ...on Hugging Face or competitive ML achievements (Kaggle medals, competition... 
    Full time
    Work at office

    Goaly

    San Francisco, CA
    19 hours ago
  • $70 - $100 per hour

    Mercor is seeking a STEM Computational Scientific Software & Evaluation Design specialist to tackle complex computational problems remotely...  ...scientific software libraries. You will design challenges to assess AI models and collaborate with research teams to refine these... 
    Remote job
    Flexible hours

    Mercor

    San Francisco, CA
    4 days ago
  •  ...Role We’re looking for an AI Research Scientist to advance the methodological...  ...architectures, new training or evaluation techniques, long-horizon...  ...validation standards are higher than benchmark culture alone, and you are...  ...and engineers on rigorous ML research practices.... 
    Full time
    Temporary work
    Work at office
    Monday to Friday
    Monday to Thursday
    Flexible hours

    Sprinter Health

    San Francisco, CA
    1 day ago
  •  ...LMArena is the open platform for evaluating how AI models perform in the real...  ...variety of Machine Learning Scientist to help advance how we...  ...learning. Strong foundation in ML and statistics, with a track...  ...that go beyond traditional benchmarks Analyze large-scale human voting... 
    Permanent employment
    Work at office

    Arena Intelligence, Inc.

    San Francisco, CA
    19 hours ago
  •  ...Blank Bio is an applied AI research lab focused on...  ...a technical team of AI scientists and engineers from...  ...Institute. The Role As an ML research scientist, you...  ...ML methods and develop benchmarks that reflect clinically...  ...Develop benchmarks and evaluation frameworks for tasks spanning... 

    blank

    San Francisco, CA
    1 day ago
  •  ...are partnering with a AI‑native therapeutics company...  ...a Machine Learning Scientist to conduct original, high...  ...implement experiments, evaluate results, and clearly...  ...modalities Define meaningful benchmark tasks and evaluation...  ...scientists, and ML researchers Evaluate modern... 

    Harnham

    San Francisco, CA
    3 days ago
  • $50 per hour

    A leading AI research firm is seeking PhDs in Mathematics or related fields for a fully remote contract role. The successful candidate...  ...will design advanced math problems to test AI performance and evaluate outputs for accuracy. Strong mathematical reasoning, problem-... 
    Remote job
    Contract work
    Flexible hours

    Turing

    San Francisco, CA
    19 hours ago
  •  ...further notice. Meet the Team At Foundation AI, we are leading frontier AI research...  ...training methods, inference optimization, evaluation techniques, and data pipelines. Together,...  ...software engineering practices, and common AI/ML libraries Experience with large language models... 

    Cisco Systems, Inc.

    San Francisco, CA
    1 day ago
  •  ...systems / high‑performance computing for ML. Are comfortable working from...  ...make large‑scale rollout collection and evaluation cheaper. Use these pipelines to train, evaluate...  ...and APIs as needed. Establish metrics, benchmarks, and experimentation frameworks to validate... 
    Full time

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    19 hours ago
  • AI Researcher (Computer Vision/Multimodal/Generative AI) About the Role We are hiring ML Researchers to develop novel approaches that advance the frontier of multimodal vision...  ...aligned with product differentiation. Evaluate new model paradigms for scalability and efficiency... 

    SpreeAI

    San Francisco, CA
    1 day ago
  •  ...revolutionizing software development with AI-powered formal verification. We've developed...  ...’ve won a well-known formal verification benchmark called PutnamBench, which consists of 672...  ...algorithm and LLMs Build effective and efficient ML pipelines Collaborate with other teams to... 
    Contract work

    Logical Intelligence

    San Francisco, CA
    1 day ago
  • Patronus AI is seeking an Applied Researcher in San Francisco to lead foundational research...  ...how agentic AI systems are trained, evaluated and improved. You will work at the intersection...  ...research. A BS, MS, or PhD in CS/ML is required; strong Python and ML framework... 

    Doist

    San Francisco, CA
    19 hours ago
  •  ...creating trustworthy and reliable AI systems, changing banking for...  ..., our applications of AI & ML bring humanity and simplicity...  ...cross‑functional team of data scientists, software engineers, machine learning...  ...from design through training, evaluation, validation, and... 
    Flexible hours

    Capital One

    San Francisco, CA
    3 days ago
  •  ...consumer-grade agents that redefine human-AI collaboration for millions. Software shouldn...  ...our research rigor, safety metrics, and evaluation pipelines for everything that calls itself...  ...Qualifications PhD or equivalent track record in ML / NLP / RL / systems 3+ strong papers or... 
    Relocation package

    Agi,-Inc.

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Scientist, AI Evaluation & Benchmarking. Be the first to apply!