Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Scientist, Agentic Data & Benchmarking

Institute of Foundation Models

Job Description

Job Description

About the Institute of Foundation Models 

The Institute of Foundation Models (IFM) is a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy. 

As part of our team, you'll work at the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You'll help build groundbreaking AI systems with the potential to reshape entire industries, and contribute to establishing MBZUAI as a global hub for high-performance computing and deep learning. 

About the role 

The Agents team trains advanced agentic language models that use reasoning and tool use to complete real tasks on a computer. This is a specialist role at the center of the loop that drives those models: the data we train on and the benchmarks we measure against. 

You'll own the agentic data pipeline end-to-end — sourcing and generating high-quality trajectories, tool-use data, and RL environments — and the evaluation suite that tells us, rigorously and reproducibly, what our agents can actually do. These two halves are inseparable: benchmarks expose where models fail, and targeted data closes the gap. The agents are only as good as the data they learn from and the evals that keep us honest, and this role owns both. 

This is a research scientist position for someone who wants depth in data and measurement rather than breadth across the whole stack. You should be the kind of person who reads through datasets line by line, distrusts a metric until it's been validated, and gets satisfaction from making an eval suite that nobody questions. 

Key responsibilities

Benchmarking & evaluation 

  • Design and run evaluations of agentic capabilities — multi-step reasoning, tool use, long-horizon planning, computer use, and safety properties — turning ambiguous notions of "intelligence" into defensible, reproducible metrics. 

  • Build and harden evaluation harnesses so benchmarks run reliably at scale against training checkpoints, with clear signal on regressions and model health. 

  • Run experiments characterizing how prompting, sampling, scaffolding, and environment design affect agentic performance on internal and public benchmarks. 

  • Diagnose anomalous eval results mid-training-run — determine whether the cause is the model, the data, the harness, or the infrastructure — and communicate the answer clearly. 

Agentic data 

  • Source, generate, and curate high-quality agentic training data: trajectories, tool-use traces, and task datasets for new capabilities. 

  • Design and scale RL environments and reward signals, and measure their impact on model performance. 

  • Manage technical relationships with external data vendors and domain experts, evaluating data quality and iterating quickly on feedback. 

  • Develop QA frameworks that catch reward hacking, label noise, and contamination, keeping data and benchmark quality high. 

Across both 

  • Contribute to technical reports, research publications, and open-source benchmarks and tooling. 

  • Partner with research and product teams to translate capability goals into measurable data and evaluation artifacts. 

Qualifications

Academic qualifications 

  • BS, MS, or PhD (or equivalent experience) in Computer Science, Machine Learning, or a related field. 

Minimum qualifications 

  • 2+ years of experience with a clear emphasis on evaluations and/or training-data curation for ML systems (related areas: LLM training/fine-tuning, RL, or distributed ML systems). 

  • Strong Python and PyTorch development experience. 

  • Demonstrated experience designing and deep-diving into evaluations, or curating and generating training datasets — ideally both. 

  • Hands-on experience using LLM agents in your personal or professional work. 

  • A habit of reading through raw data and trajectories to understand them and spot issues, and an instinct to distrust a metric until it's validated. 

Preferred qualifications 

  • Experience with reinforcement learning, reward design, or RL environment construction for LLMs. 

  • Background in statistics and experimental design — a feel for signal-to-noise, statistical power, and contamination in evaluations. 

  • Experience with large-scale dataset sourcing, curation, and processing, including working with external vendors or domain experts. 

  • Strong knowledge of the literature on agent evaluation, RL, LLM reasoning, and tool use. 

  • Experience building or operating data pipelines and evaluation infrastructure reliable at scale (e.g., PyTorch, Ray). 

  • Experience evaluating or generating data for software-engineering or computer-use agents. 

  • Contributions to published research, public benchmarks, and/or open-source ML software. 

Representative projects

  • Stand up a new agentic benchmark from scratch — define the task, build the dataset and scoring, validate against known signals, and ship a view that makes the result legible to researchers and leadership. 

  • Build an RL environment for a new high-value capability: design the reward, generate and QA the trajectory data, and measure the lift on model performance. 

  • Diagnose a mid-training regression: an eval suite returns anomalous numbers and you determine whether it's the model, the harness, the data, or the infrastructure. 

  • Partner with an external data vendor or domain expert to source high-quality trajectories, then build the QA framework that keeps reward hacking and contamination out. 

  • Take a flaky distributed eval pipeline and make it reliable — better retries, better observability, faster feedback to researchers. 

Salary Range

 

The posted salary range represents the company’s good faith estimate of the compensation for this position upon hire. The actual compensation offered may vary within this range depending on individual qualifications, including but not limited to relevant skills, experience, education, certifications, geographic location, and specific business needs.  

We encourage you to apply even if you don't meet every qualification listed. Strong candidates rarely match every line, and we'd rather hear from you than have you rule yourself out. 

Vacancy posted 29 days ago
Similar jobs that could be interesting for youBased on the Research Scientist, Agentic Data & Benchmarking in Sunnyvale, CA vacancy
  • $192.2k - $260k

     ...and inventive Applied Scientist with a strong machine learning...  ..., image and structured data sources, and large-...  ...of our customers by researching and building innovative solutions using Agentic AI.Agentic AI drives innovation...  ...methodologies, create benchmarks, and build evaluation... 
    Data
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...worldWe are looking for an outstanding Senior Agentic AI Applied Researcher to build groundbreaking mutli-modal agentic AI solutions for data science and machine learning. As a...  ...workflowsDevelop and maintain agentic AI benchmarks to evaluate performance across diverse... 
    Data
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $192k - $304.75k

     ...generative AI. We are looking for a research scientist / engineer who is passionate...  ...intersection of the areas: 1) Synthetic data and algorithmic research for agentic RL 2) Data and training...  ...models by developing training data, benchmarks, LLMs and software (including NeMo... 
    Data
    Full time

    Nvidia

    Santa Clara, CA
    14 hours ago
  • $168k - $264.5k

     ...computational engine for groundbreaking research, enabling scientists to model complex biological...  ..., and digital twins.Designing benchmarks and evaluation methods for agentic systems in this domain, and...  ...Experience with clinical trial data, electronic health records, or... 
    Data
    Full time
    Remote work
    Shift work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k

     ...technology platform powered by data and machine learning...  ...in AI-driven agentic systems. We're dedicated...  ...safety across diverse benchmarks and production scenariosLead...  ..., engineering, and research teams to translate...  ...senior ML engineers and scientists, providing technical... 
    Data
    Full time
    Worldwide

    Expedia

    San Jose, CA
    1 day ago
  • $171.6k - $222.2k

     ...talented, and inventive Applied Scientist with a strong machine learning background...  ..., text, image and structured data sources, and large-scale...  ...impact millions of our customers by researching and building innovative solutions using Agentic AI.Agentic AI drives innovation... 
    Data
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    1 day ago
  •  ...computing experiences—from AI and data centers, to PCs, gaming and...  ...hiring Forward Deployed AI Research Scientist to help bring advanced AI...  ...tasks, including agentic workflows, tool use, code generation...  ...teams to define metrics, evals, benchmarks, and acceptance criteria for... 
    Data

    AMD

    Santa Clara, CA
    4 days ago
  • $176k - $253.5k

     ...time /HybridAt Toyota Research Institute (TRI), we’re...  ...looking for an AI Research Scientist or Senior AI Scientist...  ..., evaluation, and benchmarking. This role operates at...  ...(e.g., SFT, RLHF) and agentic systems.Proficiency in...  ...about how your data is processed, please contact... 
    Data
    Full time
    Temporary work
    Local area
    Shift work

    Toyota Research Institute

    Los Altos, CA
    2 days ago
  • $201.3k - $352.3k

     ...platform brings together any AI, any data, and any workflow— helping 85% of...  ...for customers. We are a group of researchers, applied scientists, engineers, and product managers...  ...driven Senior Engineering Manager, Agentic & GenAI Benchmarking and Evaluations to establish and lead... 
    Data
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours
    Shift work

    ServiceNow

    Santa Clara, CA
    3 days ago
  •  ...Google DeepMind, we’re a team of scientists, engineers, machine learning...  ...models. The role of the Research Scientist / Research Engineer...  ...will be to apply and develop data and algorithmic cutting edge...  ...audio-to-text modalities and agentic capabilitiesExploring data, reasoning... 
    Data

    DeepMind

    Mountain View, CA
    1 day ago
  • $143k - $275k

     ...information, visit Quantum Topological Research Scientist - Technology Architecture is a senior...  ...systemsContribute to experimental design, data analysis, and failure learning cycles...  ...heterogeneous integration)Establish technical benchmarks, requirements, and success metrics... 
    Data
    Full time
    Local area

    Globalfoundries

    Santa Clara, CA
    1 day ago
  • $165k - $238k

     ...aspiration and riskiness of research with the speed and ambition of...  ...the role:As a Senior Research Scientist you will be a key architect...  ...on building next-generation agentic systems designed to perform complex...  ...corpora of unstructured data.We look for engineers who bridge... 
    Data
    Full time
    Work at office
    3 days per week

    X Company

    Mountain View, CA
    2 days ago
  • $207k - $301k

     ...and ML efficiency. Drive new research ideas from conception, experimentation...  ...project work by defining the data structure, framework, design,...  ...types of work. As a Research Scientist, you'll setup large-scale...  ...and RecSys. We focus on: 1. benchmarking and improving Gemini models... 
    Data
    Shift work

    Google

    Mountain View, CA
    4 days ago
  • $192k - $304.75k

     ...'re looking for a passionate scientist at the intersection of quantum...  .... As a Sr. Quantum Applied Research Scientist, you will help design...  ...develop physics-informed data synthesis pipelines, post-trainable...  ...architectures, and practical benchmarks that the quantum community... 
    Data
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $137.7k - $275.4k

     ...AI Incubation team seeks a Research Scientist to advance AI-driven innovation...  ...understanding, and Agentic AI. Together, you'll build systems...  ...over complex, heterogeneous data to surface actionable insights...  ...top performers in OpenASR benchmarks, while its federated AI architecture... 
    Data
    Full time
    Work at office
    Remote work
    Worldwide

    Zoom

    San Jose, CA
    14 hours ago
  • $147k - $211k

    SnapshotWe are seeking strong Research Scientists with expertise in AI research and experience in...  ...LLMs using RLExperience with developing agentic AI solutions to complex problemsExcited...  ...integrating information from different data types (e.g., vision, audio, text).The US... 
    Data
    Full time

    DeepMind

    Mountain View, CA
    2 days ago
  • $180k - $258.75k

     .../Full-time /HybridAt Toyota Research Institute (TRI), we’re on a mission...  ...modeling, new experimental data, artificial intelligence, and...  ...are hiring a Senior Research Scientist to help deepen the chemistry...  ...other appropriate scientific benchmarks.Communicate scientific... 
    Data
    Full time
    Local area
    Shift work

    Toyota Research Institute

    Los Altos, CA
    3 days ago
  • $184k - $287.5k

     ...is hiring Senior Deep Learning Scientists to advance our efforts in streaming and agentic multimodal AI. You will demonstrate...  ...:Apply fundamental and applied research to develop, train, fine-tune,...  ...collection, development, and benchmarking of multimodal datasets, ensuring... 
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...generation computing experiences—from AI and data centers, to PCs, gaming and embedded...  ...career. THE ROLE:We are hiring a AI Research Scientist, Hardware AI Systems, to develop AI systems...  ...one-off wins into reusable models, benchmarks, and transfer protocols that compound across... 
    Data
    Shift work

    AMD

    Santa Clara, CA
    14 hours ago
  • $192k - $304.75k

     ...computing. We're looking for a passionate AI research scientist with deep quantum computing expertise...  ...models, curated datasets, and rigorous benchmarks that advance the state of the art and...  ...and hardware-derived syndrome data, enabling the community to train and evaluate... 
    Data
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $165k - $195k

    Company DescriptionThe Bosch Research and Technology Center North America with offices in Sunnyvale...  ...& Mixed Reality, Cloud Robotics, Big Data Visual Analytics, Explainable AI (XAI),...  ...-centric AI, synthetic data generation, agentic AIProficiency with version control... 
    Data
    Full time
    Work experience placement
    Local area
    Worldwide

    Robert Bosch

    Sunnyvale, CA
    14 hours ago
  •  ...Role We’re looking for Applied Scientists to join Wayve Labs and help...  ...Wayve, we are a high‑conviction research team with the strategic...  ...long contexts, and scale with data and compute. Cross‑Embodiment...  ...evolve Evaluation Frameworks and Benchmarks for long‑horizon prediction,... 
    Data
    Full time
    Work at office
    Work from home
    Visa sponsorship
    Relocation package
    Flexible hours

    Icehouseventures

    Sunnyvale, CA
    4 days ago
  • $207k - $304k

     ...to serve as a founding pillar of Tapestry’s new NYC Engineering Hub. As the Engineering Manager for the Data Platform team, you will lead the evolution of our Agentic Data Platform—the engine that transforms raw grid data, drone imagery, and street-level perception into... 
    Data
    Full time
    Flexible hours

    X Company

    Mountain View, CA
    3 days ago
  •  ...the order of listing. What you’ll do As a Research Scientist at Simular, you will: Shape the future of agentic AI by pioneering new research directions in...  ...and execute experiments end-to-end: from data collection and benchmarking, to model training and evaluation. Develop... 
    Data

    Simular

    Palo Alto, CA
    2 days ago
  •  ...Models We are a dedicated research lab for building,...  ...alongside world-class researchers, data scientists, and engineers, tackling the...  ...development of the PAN (Physical, Agentic, and Networked) world models...  ...metrics and evaluation benchmarks to better assess model performance... 
    Data
    Visa sponsorship

    Institute of Foundation Models

    Sunnyvale, CA
    24 days ago
  • $207k - $301k

     ...develop, and deploy scalable and agentic AI solutions for enterprise...  ...system changes or training data enhancements.Implement, optimize...  ...team of developers, researchers, engaging in design and code...  ...types of work. As a Research Scientist, you'll setup large-scale tests... 
    Data

    Google

    Sunnyvale, CA
    14 hours ago
  • $150k - $200k

     ...safety by developing novel evaluation benchmarks and alignment techniques for enterprise...  ...language models. Responsibilities As a Research Scientist/ ML Engineer, you will play a crucial role...  ...judges for safety, reliability, and data curation. Your technical skills will accelerate... 
    Data
    Full time

    Collinear AI

    Mountain View, CA
    4 days ago
  •  ...Models We are a dedicated research lab for building, understanding...  ...world-class researchers, data scientists, and engineers, tackling the...  ...understanding, reasoning, and agentic capabilities. You will work...  ...post-training, and evaluation benchmarks. The role combines cutting-... 
    Data

    Institute of Foundation Models

    Sunnyvale, CA
    9 days ago
  • $132k - $189k

    Define, plan, and conduct quantitative research to evaluate user behavior on CE platforms.Apply...  ...design) to analyze large-scale log data, survey data, and other behavioral data.Partner...  ..., or CRM platforms, and evaluating agentic AI workflows and architectures.Experience... 
    Data

    Google

    Mountain View, CA
    4 days ago
  • $174k - $253k

     ...Innovate on algorithmic interventions and data curation strategies to improve model...  ...experience.2 years of experience leading a research agenda.Preferred qualifications:2 years of...  ...influencing other researchers.As a Research Scientist, you will be responsible for improving... 
    Data

    Google

    Mountain View, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Scientist, Agentic Data & Benchmarking. Be the first to apply!