Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Head of Evaluations Research

Doist

About Aaru

Aaru builds simulations of human behavior. Each simulation contains a population of agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing, from product launches and policy changes to critical communications. Because the agents are simulated rather than recruited, they can reason through complex hypotheticals without fatigue or the response effects common in human studies.

The role

Evaluation Research determines whether Aaru's populations, predictions, and simulations correspond closely enough to the real world to support consequential decisions. The function defines what should be measured, develops the methods for measuring it, and produces the evidence Aaru uses to improve its systems and describe their capabilities. As Head of Evaluation Research, you will set Aaru's evaluation charter and lead the research needed to test its core systems. The strongest evaluations will be grounded in observed behavior and outcomes, including transactions, product usage, operational records, resolved events, and longitudinal decisions. You will work across Population Research, Prediction Research, Simulation Engineering, and customer-facing teams while preserving the independence needed to identify inconvenient results. This is a hands‑on research leadership role. You will design studies, construct evaluation datasets, write analysis code, inspect individual failures, and develop new measurements when existing benchmarks are inadequate. You will also recruit and lead a small team that can combine scientific rigor with a practical understanding of how research and engineering systems improve.

What you will do

Define a coherent evaluation agenda across population construction, predictive systems, agent behavior, group dynamics, and end-to-end simulations. Turn broad questions about realism, accuracy, and usefulness into measurable constructs, decisive experiments, and clear decision criteria. Build tests of population quality that assess whether each generated person forms a coherent whole, whether the population reproduces important relationships in the data, and whether rare but plausible profiles are represented. Evaluate forecasts and other predictive outputs using temporal holdouts, prospective outcomes, calibration, ranking quality, subgroup performance, and the real cost of different errors. Compare simulations with transactions, behavioral traces, product adoption, operational outcomes, market movements, and other records of what people actually did. Design longitudinal and interaction-based evaluations that test how agents change over time, respond to new information, and influence one another. Develop end‑to‑end studies that show whether better components lead to better answers for the decisions customers use Aaru to make. Find failures hidden by aggregate metrics, especially those concentrated in important subgroups, rare cases, or changing environments. Establish strong baselines, clean holdouts, contamination controls, and statistical standards appropriate to each research question. Create diagnostic evaluations that help researchers identify why a system failed and whether a proposed fix generalizes. Work with Simulation Engineering to make evaluations repeatable and versioned, and convert production outcomes and customer failures into durable test cases. Produce evidence for customers and the public that is reproducible, appropriately scoped, and clear about uncertainty and limitations. Communicate negative and inconclusive results with the same care as positive findings. Hire, mentor, and lead exceptional evaluation researchers and research engineers while remaining a direct contributor.

Representative research directions

Construct a population using data available at one point in time, then test whether its future transactions, choices, or behavioral outcomes match what later occurred. Develop methods for detecting contradictions, impossible combinations, unstable attributes, and unsupported specificity within individual profiles. Recreate a historical decision environment using only information available before the outcome, then compare the simulation with the observed result. Run prospective evaluations in which Aaru makes predictions before outcomes are known and tracks performance as those outcomes resolve. Determine whether component measures such as row-wise coherence or overall accuracy predict the quality of an end‑to‑end simulation. Compare agent‑based simulation with direct forecasting, statistical models, and simpler segment‑level approaches on the same outcome. Measure whether simulated groups reproduce observed patterns in information diffusion, coordination, influence, or collective decision‑making. Study how population size, heterogeneity, interaction structure, model capability, and computation affect simulation fidelity. Build tests for memorization, leakage, prompt sensitivity, unsupported certainty, and benchmark‑specific overfitting. Develop evaluation methods for settings where outcomes are delayed, noisy, only partially observed, or open to more than one reasonable interpretation.

How we work

A useful evaluation measures something consequential and helps the company learn. We compare systems with strong alternatives, protect held‑out data, quantify uncertainty, and separate exploratory findings from evidence used to support a claim. Ground truth is often imperfect, so the quality and limits of the outcome data are part of the research problem. Evaluation should make research faster by giving teams clear signals about what improved and what did not. It should also make Aaru more trustworthy by exposing failures early and keeping external claims aligned with the available evidence.

You might thrive in this role if

You have developed an original evaluation or measurement agenda in machine learning, behavioral science, computational social science, statistics, economics, psychometrics, or an environment with a comparable bar for rigor. You have built evaluations that changed a research direction, model capability, product decision, or scientific conclusion. You can define a difficult construct precisely enough to measure it without losing the underlying question. You are comfortable with experimental design, observational data, sampling, uncertainty, statistical power, leakage, and condition shift. You can write code, analyze large datasets, design studies, and inspect individual model failures. You can work closely with the teams building a system while reaching independent conclusions about its quality. You care more about an accurate result than a favorable one and are willing to revise your own evaluation when it proves inadequate. You can explain technical evidence clearly to researchers, engineers, customers, company leadership, and the public. You have led researchers or a major technical direction while remaining directly involved in the work. You want to build in person, in New York, at high speed.

Strong candidates may also have

Work in ML evaluation, model behavior, forecasting, econometrics, psychometrics, causal inference, experimental economics, or measurement theory. Experience evaluating LLM agents, multi‑agent systems, synthetic populations, recommender systems, probabilistic models, or decision‑support tools. Experience with longitudinal records, transaction data, product analytics, field experiments, prospective studies, or validation against operational outcomes. Experience building evaluation platforms, regression suites, experiment‑tracking systems, or shared research datasets. A record of finding an important failure that standard metrics missed and developing a better way to measure it. Experience communicating scientific or technical results in customer‑facing, public, policy, or regulatory settings.

Success in this role looks like

Aaru has a clear and widely trusted definition of quality across populations, predictions, and end‑to‑end simulations. Core evaluations are grounded in observed outcomes and reveal whether performance generalizes across time, domains, and groups. Researchers receive measurements that are diagnostic enough to guide improvement while protected holdouts preserve the integrity of final results. Important failures are found early, explained clearly, and converted into durable tests. Product and company decisions rely on evidence that is reproducible, appropriately uncertain, and connected to real‑world value. Customers and the public can understand what Aaru has demonstrated, where the evidence is limited, and how confidence should be interpreted. A small, exceptional team develops a reputation for evaluation work that is scientifically rigorous, practically useful, and unusually honest.

#J-18808-Ljbffr

Vacancy posted 9 hours ago
Similar jobs that could be interesting for youBased on the Head of Evaluations Research in New York, NY vacancy
  • Attendance Works is seeking a Director of Evaluation and Research to lead measurement of impact across programs and initiatives, guiding data-informed decision making for continuous improvement. The role collaborates with the CEO, VPs, and directors to shape the organization... 
    Suggested

    Attendance Works

    New York, NY
    2 days ago
  • $115k - $130k

    Senior Manager, Applied Research & Evaluation New York, NY About Understood Understood is a nonprofit focused on shaping the world for difference. We raise awareness of the challenges and strengths of people who learn and think differently. Our resources help people navigate... 
    Suggested
    Permanent employment
    Work at office
    Local area
    Visa sponsorship
    Work visa
    3 days per week

    Understood For All, Inc.

    New York, NY
    4 days ago
  • $250k - $300k

    Our client, an innovative AI and Data company, is hiring a Head of Research to join the team in New York. The successful candidate will transform...  ...data streams into a gold standard for AI training and evaluation by shaping the company’s research strategy, collaborating... 
    Suggested
    Full time

    Alldus International Consulting Ltd

    New York, NY
    a month ago
  • Mercor is seeking a qualified candidate for the position of STEM Computational Scientific Software & Evaluation Design in Astrophysics & Cosmology. This position is remote with a commitment of 15-20 hours per week. Candidates must have substantial STEM training and Python... 
    Suggested
    Remote job
    Contract work

    Mercor

    New York, NY
    3 days ago
  • Head of Special Situations Research On site in Manhattan, NY An established investment research organization is seeking a highly accomplished investment...  ...judgment. Build detailed financial models that evaluate business fundamentals, transaction outcomes, and valuation... 
    Suggested
    Work at office

    Blue-Signal-Search

    New York, NY
    4 days ago
  • $40 per hour

    AIToolboard in the United States is seeking a part-time Remote Chemistry Research Scientist for AI Evaluation to join our expanding AI research team. You will train AI models related to chemical data, test outputs, and provide rigorous evaluation to improve model accuracy... 
    Remote job
    Hourly pay
    Part time
    Flexible hours

    AIToolboard

    New York, NY
    4 days ago
  • $160k - $230k

     ...financial oil markets worldwide. The Oil Market Research function sits at the intersection of...  ...and oil-producing region.The Global Head of Oil Market Research is a senior, externally...  ...of Midland WTI into the Brent complex, evaluation of new index proposals, and the ICE HOU... 
    Full time
    Worldwide
    Night shift

    Intercontinental Exchange

    New York, NY
    2 days ago
  • $230k - $320k

    Head of YouTube Primary Research and Insights Copy link YouTube place San Bruno, CA, USA ; Mountain View, CA, USA ; +3 more ; +2 more Advanced Experience...  ...models and conduct scenario and sensitivity analyses to evaluate business outcomes and mitigate risks for new initiatives.... 

    Google Inc.

    New York, NY
    1 day ago
  • The Role You'll define the research agenda that makes AI trustworthy for enterprise decisions. At Kepler, we've solved hallucination...  ...agentic systems, memory architectures, retrieval mechanisms, and evaluation frameworks. You'll have access to completely differentiated... 
    Work at office

    Kepler (formerly Keru.ai)

    New York, NY
    5 days ago
  • $198.2k - $368k

    Director, Applied Research Are you excited about working at the forefront of applied research in an industry setting? Thomson Reuters...  ...graphs, or neuro-symbolic approachesExperience with AI agent evaluation, language model training with verifiable rewards, and synthetic... 
    Full time
    Local area
    Flexible hours

    Thomson Reuters

    New York, NY
    9 hours ago
  • $206.93k - $289k

     ...Specifics Duties Description : The New York State Institute for Basic Research in Developmental Disabilities (IBR) is recruiting for an...  ...of a new Genomics Core Facility to facilitate the proper evaluation and diagnosis for patients with a suspected IDD. The institute... 
    Permanent employment
    Full time
    Temporary work
    Part time
    Work at office
    Local area
    Remote work

    State of New York, USA

    New York, NY
    9 hours ago
  • Anthropic in the United States is hiring Research Engineers to build evaluations that quantify Claude's capabilities and measure reasoning, knowledge, and safety properties at scale. You will design end-to-end experiments, define metrics, and develop the infrastructure... 

    SignalAI

    New York, NY
    3 days ago
  •  ...as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together...  ...The Role We\'re looking for Research Engineers to build the evaluations that tell us — and the world — what Claude can actually do.... 
    Work at office
    Visa sponsorship
    Flexible hours

    SignalAI

    New York, NY
    3 days ago
  •  ...the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about...  ...Montreal, Seoul, Germany and Paris. Join us! Why this role? Evaluation is critical to making progress in scaling intelligence. As models... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    cohere

    New York, NY
    2 days ago
  • $198k - $247k

     ...experienced, strategic and mission-driven leader to drive the policy research priorities of Scale Labs. This is a unique opportunity to lead...  ...of the AI safety space.Extensive knowledge of frontier risk evaluations, AI control, and preparedness research.Established connections... 
    Full time
    Work experience placement

    Scale AI

    New York, NY
    9 hours ago
  • $65k

     ...Description Research Coordinator US-NY-Bronx Job ID: 2026-18188 Employee Classification: Exempt Department: Medicine - Nephrology Position...  ...attachments associated with grant and contract proposals. Evaluate study resource requirements and provide recommendations to investigators... 
    Full time
    Contract work
    Local area

    Stryker

    New York, NY
    3 days ago
  •  ...inspired. From strong benefits and wellness programs to meaningful collaboration and team connection. About the Role As Head of Threat Research at Netcraft, you will lead the team responsible for publishing research grounded in Netcraft's unique data sets, including... 
    Work at office
    Remote work
    Flexible hours

    Netcraft

    New York, NY
    a month ago
  • $66.3k - $70k

     ...improving the human condition through medical education, scientific research, and direct patient care. At NYU Langone Health, equity and...  ...referral, advertisement and directly scheduling a visit to evaluate the patient/subject. Reviews all the elements of the screening... 
    Traineeship
    Work at office

    NYU Grossman School of Medicine

    New York, NY
    1 day ago
  •  ...accomplishments, internal equity, budget, and subject to Fair Market Value evaluation. The hiring range listed is a good faith determination of...  ...Responsibilities: Coordinates clinical research activities for the ARJR service under the direction of the Industry... 
    Full time
    Local area
    Shift work

    HSS

    New York, NY
    5 days ago
  •  ...is looking to add a Senior Director to its Structured Finance research and origination team focused on the non-traditional commercial...  ...monitor transactionsPrior experience in credit research, risk evaluation, collateral performance assumptions with scenario and sensitivity... 
    Local area

    New York Life Insurance Company

    New York, NY
    9 hours ago
  • $258k - $348k

     ...shape the future of design and collaboration, join us!The Figma Research team is hiring an Director, Research - AI Evals to own how we...  ...to ship.The ideal candidate brings deep, hands-on experience evaluating AI/LLM-powered products — blending human evaluation with automated... 
    Minimum wage
    Full time
    For contractors
    Local area
    Remote work
    Flexible hours

    Figma

    New York, NY
    4 days ago
  • $250k - $350k

     ...build an exceptional career.Position Summary:The Director of Research & Intelligence provides strategic leadership and operational oversight...  ..., and effective use of research tools and resources.Lead the evaluation, implementation, and optimization of AI-powered legal research... 
    Work at office

    Fried, Frank, Harris, Shriver & Jacobson LLP

    New York, NY
    3 days ago
  • $140k - $230k

     ...improvement assistance when needed, and evaluates and develops logical work flows where applicable...  ....5. Collaborates with the division head or chairman to develop short term and long...  ...funding and in the 99th percentile in research dollars per investigator according to the... 
    Temporary work
    Traineeship
    Local area

    Mount Sinai Health System

    New York, NY
    9 hours ago
  •  ...Google, and Ramp — go to market with unique data, signals, and AI research. In 2025, we raised a $100M Series C backed by world-class...  ...more. Hear from our employees directly on our Glassdoor page! Head of ARC (Activation, Research & Community) Recruiting is an... 

    Clay Labs

    New York, NY
    9 hours ago
  •  ...The National Kidney Foundation is seeking a Vice President of Research Strategy, a leadership role in New York, NY. This position drives...  ..., including implementation science, market research, program evaluation, and external research funding, reporting directly to the... 

    Jobleads-US

    New York, NY
    2 days ago
  • $315k

    We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand what AI Safety Level (ASL) to assign to models. Research leads on this team collaborate with engineers in one of our focus areas: CBRN, Cyber, Autonomy... 
    Currently hiring
    Work at office
    Immediate start
    Home office
    Visa sponsorship
    Relocation package

    Anthropic

    New York, NY
    2 days ago
  • DescriptionThe Clinical Research Coordinator is an entry human subjects researcher, responsible for conducting and assisting in clinical...  ...Americans. Participants will undergo a standard dementia evaluation that includes a medical exam, cognitive testing, clinical assessment... 
    Hourly pay
    Traineeship
    Work at office
    Local area

    Mount Sinai Health System

    New York, NY
    2 days ago
  •  ...requirements. *Participates in site selection including feasibility evaluation. *Familiar with IP needs for project (IP packager/labeler,...  ...*5+ years of experience in the pharmaceutical and/or clinical research industry. *3+ years of experience in project management... 
    For contractors
    Local area
    Remote work

    Planet Pharma

    New York, NY
    2 days ago
  • Reporting to our Director of Ecosystem Research. Rate - 125k-140k You are a researcher who works at the intersection of market intelligence...  ...performing rigorous field scans and secondary research to evaluate markets for product creation. Expert experience building... 
    Work at office
    3 days per week

    ZRG Careers

    New York, NY
    4 days ago
  • $100k - $130k

     ...development and engagement initiatives. The Consumer Insights & Research function within the team focuses on uncovering strategic...  ...recommendations for League and Club stakeholders Managing creative evaluation research to assess campaign, content, and messaging... 
    Hourly pay
    Work experience placement
    Work at office
    Local area
    Remote work
    Relocation
    Flexible hours

    Triwill Group

    New York, NY
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Head of Evaluations Research. Be the first to apply!