Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Evaluation Scientist — Research, Metrics & Impact Equity

Arena

About Arena Intelligence Arena Intelligence is the open platform for evaluating how AI models perform in the real world. Created by researchers from UC Berkeley’s SkyLab, our mission is to measure and advance the frontier of AI for real-world use. Millions of people use Arena Intelligence each month to explore how frontier systems perform — and we use our community’s feedback to build transparent, rigorous, and human-centered model evaluations. Leading enterprises and AI labs rely on our evaluations to understand real-world reliability, alignment, and impact. Our leaderboards are the gold standard for AI performance — trusted by leaders across the AI community and shaping the global conversation on model reliability and progress. We’re a team of researchers, engineers, academics, and builders from places like UC Berkeley, Google, Stanford, DeepMind, and Discord. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We’re building a company where thoughtful, curious people from all backgrounds can do their best work. Everyone on our team is a deep expert in their field — our office radiates excellence, energy, and focus. About the Role Arena Intelligence is seeking a variety of Machine Learning Scientist to help advance how we evaluate and understand AI models. You’ll help design and analyse experiments that uncover what makes models useful, trustworthy and capable through human preference signals. Your work will contribute to the scientific foundations of understanding AI at scale. This role is deeply interdisciplinary. You’ll work closely with engineers, product teams, marketing and the broader research community to develop new methods for comparing models, analyzing preference data, and disentangling performance factors like style, reasoning, and robustness. Your work will inform both the public leaderboard and the tools we provide to model developers. If you’re excited by open-ended questions, rigorous evaluation, and research that’s grounded in real-world impact, you’ll find a meaningful home here. We’re looking for: Hands-on experience training large-scale models, including reward models, preference models, and fine-tuning LLMs with methods like RLHF, DPO, and contrastive learning. Strong foundation in ML and statistics, with a track record of designing novel training objectives, evaluation schemes, or statistical frameworks to improve model reliability and alignment. Fluent in the full experimental stack, from dataset design and large-batch training to rigorous evaluation and ablation, with an eye for what scales to production. Deeply collaborative mindset, working closely with engineers to productionize research insights and iterating with product teams to align modeling goals with user needs. You’ll Design and conduct experiments to evaluate AI model behavior across reasoning, style, robustness, and user preference dimensions Develop new metrics, methodologies, and evaluation protocols that go beyond traditional benchmarks Analyze large-scale human voting and interaction data to uncover insights into model performance and user preferences Collaborate with engineers to implement and scale research findings into production systems Prototype and test research ideas rapidly, balancing rigor with iteration speed Author internal reports and external publications that contribute to the broader ML research community Partner with model providers to shape evaluation questions and support responsible model testing Contribute to the scientific integrity and transparency of the Arena Intelligence leaderboard and tools You’ll have PhD or equivalent research experience in Machine Learning, Natural Language Processing, Statistics, or a related field Strong understanding of LLMs and modern deep learning architectures (e.g., Transformers, diffusion models, reinforcement learning with human feedback) Proficiency in Python and ML research libraries such as PyTorch, JAX, or TensorFlow Demonstrated ability to design and analyze experiments with statistical rigor Experience publishing research or working on open-source projects in ML, NLP, or AI evaluation Comfortable working with real-world usage data and designing metrics beyond standard benchmarks Ability to translate research questions into practical systems and collaborate across engineering and product teams Passion for open science, reproducibility, and community-driven research. What we offer We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location. Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs. The opportunity to work on cutting-edge AI with a small, mission-driven team A culture that values transparency, trust, and community impact Come help build the space where anyone can explore and help shape the future of AI. Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities. #J-18808-Ljbffr Arena

Vacancy posted 21 hours ago
Similar jobs that could be interesting for youBased on the AI Evaluation Scientist — Research, Metrics & Impact Equity in San Francisco, CA vacancy
  • $270k - $340k

    Principal AI Research Scientist, Research Director - AI ScalingP...  ...into production.The Impact You Will HaveLead and...  ...‑the‑art methods and evaluating trade‑offs in quality...  ...large models.Establish metrics, evaluation protocols...  ...performance bonus, equity, and the benefits listed... 
    Employment Equity
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $357k

    Workato is looking for a Lead AI Research Scientist to redefine AI with real-world impact. This role will involve leading research initiatives, mentoring a team...  ...based in San Francisco with a starting salary of $357,000 plus benefits and equity. #J-18808-Ljbffr Workato
    Employment Equity

    Workato

    San Francisco, CA
    21 hours ago
  • $160k - $220k

     ...the RoleWe’re looking for an AI Research Scientist to advance the...  ...architectures, new training or evaluation techniques, long-horizon research...  ...meaningful and practically impactful.Hybrid & Office ExperienceWe...  ...HealthEngineeringCompensationSF Bay AreaEstimated Base Salary $160K - $220K • Offers Equity
    Employment Equity
    Temporary work
    Work at office
    Monday to Friday
    Monday to Thursday

    Sprinter Health

    San Francisco, CA
    3 days ago
  • $234.3k - $349k

     ...enterprises orchestrate AI-powered work. Our...  ...About the roleAI research at WRITER isn't...  ...As an AI research scientist, you'll be at the...  ...You'll drive a high-impact research agenda...  ...through model training, evaluation, and production...  ...can also include equity, an immersive... 
    Employment Equity
    Full time
    Work at office
    Local area

    Writer

    San Francisco, CA
    3 days ago
  • $188k - $215k

     ...Collate is an AI document generation...  ...of Lever. Our AI researchers, engineers, and designers...  ...an AI Research Scientist to push the...  ...safe, reliable, and impactful products.   This...  ...production. Develop evaluation frameworks to...  ...to have meaningful equity and career-defining... 
    Employment Equity

    Collate

    San Francisco, CA
    11 days ago
  • $150k - $250k

    About Distyl AI Distyl is an applied AI...  ...organizations.We research and deploy technologies...  ...AI systems using Evaluation-Driven Development...  ...of overfitting to metrics that don’t reflect...  ...for meaningful equity, along with a comprehensive...  ...of high-impact projects across top... 
    Employment Equity
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    1 day ago
  • $238k - $302k

     ...states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With...  ...quantitatively-minded engineers to research and propose new ways to...  ...: Develop novel metrics and sampling techniques...  ...annual bonus program, equity incentive plan, and generous... 
    Employment Equity
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $211k - $290.5k

     ...Faire’s user facing AI bet within the...  ...Senior Applied AI/ML Scientist on the Compass...  ...quality through data, evaluation, and modeling,...  ...carry the highest impact probability-of-success...  ..., LLM-as-judge metrics, and quality criteria...  ...be eligible for equity and benefits. Actual... 
    Employment Equity
    Work experience placement
    Work at office
    Local area
    Immediate start
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire

    San Francisco, CA
    3 days ago
  • $196k - $230k

     ...the collaborative AI workspace where teams...  ...an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered...  ...guidelines, and observable metrics.AI fluency and...  ...communication and impact orientation: You...  ...cash compensation, equity, and benefits. The... 
    Employment Equity
    Local area
    Shift work

    Notion Labs

    San Francisco, CA
    1 day ago
  •  ...record of exceptional research or engineering achievement...  ...systems… About P-1 AI At P-1 AI, we are building...  ...AI Research Scientist to join our small team...  ...from data generation to evaluation to product integration....  ...Competitive salary, meaningful equity ownership, healthcare,... 
    Employment Equity
    Relocation package

    Namely

    San Francisco, CA
    1 day ago
  • $164.7k - $339.08k

     ....At Pinterest, AI isn't just a feature...  ...amplifies our impact, and we’re...  ...for a Staff Data Scientist for our Ads...  ...Engineering, Design, Research, Product...  ...scalable ML and evaluation frameworks—spanning...  ..., and metric design; bridge...  ...also eligible for equity. Final salary is... 
    Employment Equity
    Temporary work
    Work at office
    Local area
    Relocation package

    Pinterest

    San Francisco, CA
    8 hours ago
  • $130k - $220k

     ...Salario:** $130K - $220K **Equity:** $60K - $120K options...  ...leading independent AI benchmarking and...  ...where they can have real impact. We are currently helping...  ...best described as an AI Evaluation Engineer / Technical Generalist...  ...role and not a pure ML research position. It combines... 
    Employment Equity
    Full time
    Worldwide

    Aurora Jobs ApS

    San Francisco, CA
    a month ago
  • $172.5k - $260.1k

     ...SalesforceSalesforce is the #1 AI CRM, where...  ...Staff Data Scientist to support...  ...'s highest-impact business questions...  ...behavior, evaluate strategic initiatives...  ...Leaders, and Researchers to identify opportunities...  ...and success metrics to evaluate...  ...compensation, equity, and benefits.... 
    Employment Equity
    Full time

    Salesforce

    San Francisco, CA
    1 day ago
  • $216k - $270k

    Scale Labs, Research Scientist — Frontier Risk EvaluationsAs the leading data and evaluation partner for frontier AI companies, Scale plays an integral role...  ...include base salary, equity, and benefits. The range...  ...applications that deliver real impact. We work closely with... 
    Employment Equity
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $238k - $302k

     ...states. Rigorous behavioral evaluation of the Waymo Driver is a...  ...are having positive intended impacts. Conduct deep dive analysis...  ...and prototype improvements in metrics, sampling strategy, statistical...  ...annual bonus program, equity incentive plan, and generous... 
    Employment Equity
    Full time
    Remote work

    Latent Logic

    San Francisco, CA
    4 days ago
  • $164.7k - $339.08k

     ...defined, high-impact business problems...  ...ML and evaluation frameworks spanning...  ...instrumentation, and metric design,...  ..., Engineering, Research, Business, and...  ...junior and senior scientists and help foster...  ...Spark Web AI Support More...  ..., along with equity, and we are committed... 
    Employment Equity
    Full time
    Temporary work
    Work at office
    Remote work

    Pinterest, Inc.

    San Francisco, CA
    1 day ago
  • $274k - $376.2k

     ...Every Identity, from AI to HumanIdentity...  ...Applied AI Scientist to take product ideas...  ...it out without a research org behind you. This...  ...modelsCreate and maintain evaluation and training...  ...addition, Okta offers equity (where applicable)...  ...Social Impact Developing Talent... 
    Employment Equity
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $153k - $376k

     ...or iterating with AI. From idea to product...  ...for applied scientists with a Machine Learning...  ...and applied research in this area. You...  ...for AI systemsBuild evaluation systems to measure...  ...experience, potential impact, and scope of role...  .... Figma offers equity to employees, as well... 
    Employment Equity
    Minimum wage
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    Figma

    San Francisco, CA
    1 day ago
  • $139.23k - $163.8k

     ...Bank is seeking a Principal AI Research Scientist to join the Artificial Intelligence...  ...Excellence (AI CoE), a high-impact organization responsible for...  ...candidate will continuously evaluate emerging advances in...  ...incentive and recognition programs, equity stock purchase 401(k)... 
    Employment Equity
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours
    3 days per week

    US Bank

    San Francisco, CA
    1 day ago
  • A leading AI research company in San Francisco, backed by top investors, is looking for a highly technical candidate to join their engineering...  ...quickly. The position offers competitive compensation and equity, along with a thriving team dynamic focused on precision and... 
    Employment Equity

    Listen Labs

    San Francisco, CA
    2 days ago
  • $164.7k - $339.08k

     ....At Pinterest, AI isn't just a feature...  ...amplifies our impact, and we’re...  ...experienced Staff Data Scientist to join our...  ...interpret AI research (papers, model...  ...preparation, and metric diagnostics (e....  ...-accuracy evaluation, and seasonality...  ...also eligible for equity. Final salary... 
    Employment Equity
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    4 days ago
  • $218.5k - $288k

     ...OpportunityAs an Applied Scientist specializing in...  ...Models and AI Training, you will lead research and development efforts...  ...interpretable, and impactful across diverse usage...  ...Design, implement, and evaluate model training...  ...plus a competitive equity package. Actual compensation... 
    Employment Equity
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    1 day ago
  • $180k - $250k

     ...AI Research EngineerLocation: San Francisco, CA Company Stage of Funding...  ...for document AI applications.Evaluate foundation models and...  ...BenefitsBase salary: $180,000–$250,000.Equity: 0.1–0.4%.Hybrid work...  ...innovation, experimentation, and shipping impactful AI products.... 
    Employment Equity
    Work at office
    Relocation package
    Monday to Friday

    Recruiting from Scratch

    San Francisco, CA
    2 days ago
  • $176k - $253k

     ...Senior Member of Technical Staff, AI Quality, in San Francisco. Your...  ...agent quality into quantifiable metrics, ensuring high standards through robust evaluation processes. You'll build capability...  ...salary of $176,000-$253,000, with equity options and benefits like meals and... 
    Employment Equity

    Harper

    San Francisco, CA
    3 days ago
  •  ...fully. If you want to make a meaningful impact in an environment that's as affirming as...  ...the position? The Manager of Impact and Evaluation has a vital role on our Strategy Team...  ...methodology skills, a passion for data equity, and the ability to work across teams to... 
    Employment Equity
    Hourly pay
    Full time
    Work at office
    Afternoon shift

    San Francisco AIDS Foundation

    San Francisco, CA
    a month ago
  •  ...building alongside leading AI companies. We're...  ...monitoring Build and improve evaluation pipelines to measure,...  ...pipelines — designing metrics, running systematic...  ...Competitive salary + equity package, commensurate...  ...milestones and customer impact. The Hiring Journey... 
    Employment Equity
    Full time
    Shift work
    Night shift
    Weekend work

    Hilbert's Ai

    San Francisco, CA
    1 day ago
  • $192.6k - $344.85k

    ## AI Research Manager/Scientist, Reinforcement LearningApplylocations: San Francisco...  ...training workflows### ### Evaluation, Alignment & *Model*...  ...opportunities that unlock product impact## ## Qualifications### ###...  ...usefulness* Evaluation metrics are trusted and adopted across... 
    Remote work

    Autodesk, Inc.

    San Francisco, CA
    21 hours ago
  • $210k - $385k

     ...- $385K • Offers Equity U.S. Benefits Full...  ...automated evaluation pipelines to assess...  ...specifically to measure the impact of tool calls (...  ...your evaluation metrics directly shape product...  ...and using AI-assisted development...  ...at scale A strong research background, with experience... 
    Employment Equity
    Full time
    Local area

    Perplexity

    San Francisco, CA
    5 days ago
  • $147k - $211k

     ...promising areas of future research at the...  ...machine learning and AI models in production...  ....Ability to evaluate, analyze, and improve...  ...experiments and delivers impactful findings for both...  ...a Research Data Scientist to help us solve...  ...5% bonus target + equity + benefitsLearn more... 
    Employment Equity
    Work experience placement

    Google

    San Francisco, CA
    4 days ago
  •  ...entertainment — provide a distinctive research environment: rich structured...  ...tasks, and real-world evaluation grounded in professional workflows...  ...research advances to product impact at scale. This is not a role...  ...— whether in an academic lab, AI research organization, or industry... 
    For contractors
    Remote work

    Autodesk

    San Francisco, CA
    8 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Evaluation Scientist — Research, Metrics & Impact Equity. Be the first to apply!