Head of Evaluations Research
Doist
About Aaru
Aaru builds simulations of human behavior. Each simulation contains a population of agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing, from product launches and policy changes to critical communications. Because the agents are simulated rather than recruited, they can reason through complex hypotheticals without fatigue or the response effects common in human studies.
The role
Evaluation Research determines whether Aaru's populations, predictions, and simulations correspond closely enough to the real world to support consequential decisions. The function defines what should be measured, develops the methods for measuring it, and produces the evidence Aaru uses to improve its systems and describe their capabilities. As Head of Evaluation Research, you will set Aaru's evaluation charter and lead the research needed to test its core systems. The strongest evaluations will be grounded in observed behavior and outcomes, including transactions, product usage, operational records, resolved events, and longitudinal decisions. You will work across Population Research, Prediction Research, Simulation Engineering, and customer-facing teams while preserving the independence needed to identify inconvenient results. This is a hands‑on research leadership role. You will design studies, construct evaluation datasets, write analysis code, inspect individual failures, and develop new measurements when existing benchmarks are inadequate. You will also recruit and lead a small team that can combine scientific rigor with a practical understanding of how research and engineering systems improve.
What you will do
Define a coherent evaluation agenda across population construction, predictive systems, agent behavior, group dynamics, and end-to-end simulations. Turn broad questions about realism, accuracy, and usefulness into measurable constructs, decisive experiments, and clear decision criteria. Build tests of population quality that assess whether each generated person forms a coherent whole, whether the population reproduces important relationships in the data, and whether rare but plausible profiles are represented. Evaluate forecasts and other predictive outputs using temporal holdouts, prospective outcomes, calibration, ranking quality, subgroup performance, and the real cost of different errors. Compare simulations with transactions, behavioral traces, product adoption, operational outcomes, market movements, and other records of what people actually did. Design longitudinal and interaction-based evaluations that test how agents change over time, respond to new information, and influence one another. Develop end‑to‑end studies that show whether better components lead to better answers for the decisions customers use Aaru to make. Find failures hidden by aggregate metrics, especially those concentrated in important subgroups, rare cases, or changing environments. Establish strong baselines, clean holdouts, contamination controls, and statistical standards appropriate to each research question. Create diagnostic evaluations that help researchers identify why a system failed and whether a proposed fix generalizes. Work with Simulation Engineering to make evaluations repeatable and versioned, and convert production outcomes and customer failures into durable test cases. Produce evidence for customers and the public that is reproducible, appropriately scoped, and clear about uncertainty and limitations. Communicate negative and inconclusive results with the same care as positive findings. Hire, mentor, and lead exceptional evaluation researchers and research engineers while remaining a direct contributor.
Representative research directions
Construct a population using data available at one point in time, then test whether its future transactions, choices, or behavioral outcomes match what later occurred. Develop methods for detecting contradictions, impossible combinations, unstable attributes, and unsupported specificity within individual profiles. Recreate a historical decision environment using only information available before the outcome, then compare the simulation with the observed result. Run prospective evaluations in which Aaru makes predictions before outcomes are known and tracks performance as those outcomes resolve. Determine whether component measures such as row-wise coherence or overall accuracy predict the quality of an end‑to‑end simulation. Compare agent‑based simulation with direct forecasting, statistical models, and simpler segment‑level approaches on the same outcome. Measure whether simulated groups reproduce observed patterns in information diffusion, coordination, influence, or collective decision‑making. Study how population size, heterogeneity, interaction structure, model capability, and computation affect simulation fidelity. Build tests for memorization, leakage, prompt sensitivity, unsupported certainty, and benchmark‑specific overfitting. Develop evaluation methods for settings where outcomes are delayed, noisy, only partially observed, or open to more than one reasonable interpretation.
How we work
A useful evaluation measures something consequential and helps the company learn. We compare systems with strong alternatives, protect held‑out data, quantify uncertainty, and separate exploratory findings from evidence used to support a claim. Ground truth is often imperfect, so the quality and limits of the outcome data are part of the research problem. Evaluation should make research faster by giving teams clear signals about what improved and what did not. It should also make Aaru more trustworthy by exposing failures early and keeping external claims aligned with the available evidence.
You might thrive in this role if
You have developed an original evaluation or measurement agenda in machine learning, behavioral science, computational social science, statistics, economics, psychometrics, or an environment with a comparable bar for rigor. You have built evaluations that changed a research direction, model capability, product decision, or scientific conclusion. You can define a difficult construct precisely enough to measure it without losing the underlying question. You are comfortable with experimental design, observational data, sampling, uncertainty, statistical power, leakage, and condition shift. You can write code, analyze large datasets, design studies, and inspect individual model failures. You can work closely with the teams building a system while reaching independent conclusions about its quality. You care more about an accurate result than a favorable one and are willing to revise your own evaluation when it proves inadequate. You can explain technical evidence clearly to researchers, engineers, customers, company leadership, and the public. You have led researchers or a major technical direction while remaining directly involved in the work. You want to build in person, in New York, at high speed.
Strong candidates may also have
Work in ML evaluation, model behavior, forecasting, econometrics, psychometrics, causal inference, experimental economics, or measurement theory. Experience evaluating LLM agents, multi‑agent systems, synthetic populations, recommender systems, probabilistic models, or decision‑support tools. Experience with longitudinal records, transaction data, product analytics, field experiments, prospective studies, or validation against operational outcomes. Experience building evaluation platforms, regression suites, experiment‑tracking systems, or shared research datasets. A record of finding an important failure that standard metrics missed and developing a better way to measure it. Experience communicating scientific or technical results in customer‑facing, public, policy, or regulatory settings.
Success in this role looks like
Aaru has a clear and widely trusted definition of quality across populations, predictions, and end‑to‑end simulations. Core evaluations are grounded in observed outcomes and reveal whether performance generalizes across time, domains, and groups. Researchers receive measurements that are diagnostic enough to guide improvement while protected holdouts preserve the integrity of final results. Important failures are found early, explained clearly, and converted into durable tests. Product and company decisions rely on evidence that is reproducible, appropriately uncertain, and connected to real‑world value. Customers and the public can understand what Aaru has demonstrated, where the evidence is limited, and how confidence should be interpreted. A small, exceptional team develops a reputation for evaluation work that is scientifically rigorous, practically useful, and unusually honest.
#J-18808-Ljbffr- Attendance Works is seeking a Director of Evaluation and Research to lead measurement of impact across programs and initiatives, guiding data-informed decision making for continuous improvement. The role collaborates with the CEO, VPs, and directors to shape the organization...Suggested
$115k - $130k
Senior Manager, Applied Research & Evaluation New York, NY About Understood Understood is a nonprofit focused on shaping the world for difference. We raise awareness of the challenges and strengths of people who learn and think differently. Our resources help people navigate...SuggestedPermanent employmentWork at officeLocal areaVisa sponsorshipWork visa3 days per week$250k - $300k
Our client, an innovative AI and Data company, is hiring a Head of Research to join the team in New York. The successful candidate will transform... ...data streams into a gold standard for AI training and evaluation by shaping the company’s research strategy, collaborating...SuggestedFull time- Mercor is seeking a qualified candidate for the position of STEM Computational Scientific Software & Evaluation Design in Astrophysics & Cosmology. This position is remote with a commitment of 15-20 hours per week. Candidates must have substantial STEM training and Python...SuggestedRemote jobContract work
- Head of Special Situations Research On site in Manhattan, NY An established investment research organization is seeking a highly accomplished investment... ...judgment. Build detailed financial models that evaluate business fundamentals, transaction outcomes, and valuation...SuggestedWork at office
$40 per hour
AIToolboard in the United States is seeking a part-time Remote Chemistry Research Scientist for AI Evaluation to join our expanding AI research team. You will train AI models related to chemical data, test outputs, and provide rigorous evaluation to improve model accuracy...Remote jobHourly payPart timeFlexible hours$160k - $230k
...financial oil markets worldwide. The Oil Market Research function sits at the intersection of... ...and oil-producing region.The Global Head of Oil Market Research is a senior, externally... ...of Midland WTI into the Brent complex, evaluation of new index proposals, and the ICE HOU...Full timeWorldwideNight shift$230k - $320k
Head of YouTube Primary Research and Insights Copy link YouTube place San Bruno, CA, USA ; Mountain View, CA, USA ; +3 more ; +2 more Advanced Experience... ...models and conduct scenario and sensitivity analyses to evaluate business outcomes and mitigate risks for new initiatives....- The Role You'll define the research agenda that makes AI trustworthy for enterprise decisions. At Kepler, we've solved hallucination... ...agentic systems, memory architectures, retrieval mechanisms, and evaluation frameworks. You'll have access to completely differentiated...Work at office
$198.2k - $368k
Director, Applied Research Are you excited about working at the forefront of applied research in an industry setting? Thomson Reuters... ...graphs, or neuro-symbolic approachesExperience with AI agent evaluation, language model training with verifiable rewards, and synthetic...Full timeLocal areaFlexible hours$206.93k - $289k
...Specifics Duties Description : The New York State Institute for Basic Research in Developmental Disabilities (IBR) is recruiting for an... ...of a new Genomics Core Facility to facilitate the proper evaluation and diagnosis for patients with a suspected IDD. The institute...Permanent employmentFull timeTemporary workPart timeWork at officeLocal areaRemote work- Anthropic in the United States is hiring Research Engineers to build evaluations that quantify Claude's capabilities and measure reasoning, knowledge, and safety properties at scale. You will design end-to-end experiments, define metrics, and develop the infrastructure...
- ...as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together... ...The Role We\'re looking for Research Engineers to build the evaluations that tell us — and the world — what Claude can actually do....Work at officeVisa sponsorshipFlexible hours
- ...the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about... ...Montreal, Seoul, Germany and Paris. Join us! Why this role? Evaluation is critical to making progress in scaling intelligence. As models...Full timeWork at officeLocal areaRemote workHome office
$198k - $247k
...experienced, strategic and mission-driven leader to drive the policy research priorities of Scale Labs. This is a unique opportunity to lead... ...of the AI safety space.Extensive knowledge of frontier risk evaluations, AI control, and preparedness research.Established connections...Full timeWork experience placement$65k
...Description Research Coordinator US-NY-Bronx Job ID: 2026-18188 Employee Classification: Exempt Department: Medicine - Nephrology Position... ...attachments associated with grant and contract proposals. Evaluate study resource requirements and provide recommendations to investigators...Full timeContract workLocal area- ...inspired. From strong benefits and wellness programs to meaningful collaboration and team connection. About the Role As Head of Threat Research at Netcraft, you will lead the team responsible for publishing research grounded in Netcraft's unique data sets, including...Work at officeRemote workFlexible hours
$66.3k - $70k
...improving the human condition through medical education, scientific research, and direct patient care. At NYU Langone Health, equity and... ...referral, advertisement and directly scheduling a visit to evaluate the patient/subject. Reviews all the elements of the screening...TraineeshipWork at office- ...accomplishments, internal equity, budget, and subject to Fair Market Value evaluation. The hiring range listed is a good faith determination of... ...Responsibilities: Coordinates clinical research activities for the ARJR service under the direction of the Industry...Full timeLocal areaShift work
- ...is looking to add a Senior Director to its Structured Finance research and origination team focused on the non-traditional commercial... ...monitor transactionsPrior experience in credit research, risk evaluation, collateral performance assumptions with scenario and sensitivity...Local area
$258k - $348k
...shape the future of design and collaboration, join us!The Figma Research team is hiring an Director, Research - AI Evals to own how we... ...to ship.The ideal candidate brings deep, hands-on experience evaluating AI/LLM-powered products — blending human evaluation with automated...Minimum wageFull timeFor contractorsLocal areaRemote workFlexible hours$250k - $350k
...build an exceptional career.Position Summary:The Director of Research & Intelligence provides strategic leadership and operational oversight... ..., and effective use of research tools and resources.Lead the evaluation, implementation, and optimization of AI-powered legal research...Work at office$140k - $230k
...improvement assistance when needed, and evaluates and develops logical work flows where applicable... ....5. Collaborates with the division head or chairman to develop short term and long... ...funding and in the 99th percentile in research dollars per investigator according to the...Temporary workTraineeshipLocal area- ...Google, and Ramp — go to market with unique data, signals, and AI research. In 2025, we raised a $100M Series C backed by world-class... ...more. Hear from our employees directly on our Glassdoor page! Head of ARC (Activation, Research & Community) Recruiting is an...
- ...The National Kidney Foundation is seeking a Vice President of Research Strategy, a leadership role in New York, NY. This position drives... ..., including implementation science, market research, program evaluation, and external research funding, reporting directly to the...
$315k
We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand what AI Safety Level (ASL) to assign to models. Research leads on this team collaborate with engineers in one of our focus areas: CBRN, Cyber, Autonomy...Currently hiringWork at officeImmediate startHome officeVisa sponsorshipRelocation package- DescriptionThe Clinical Research Coordinator is an entry human subjects researcher, responsible for conducting and assisting in clinical... ...Americans. Participants will undergo a standard dementia evaluation that includes a medical exam, cognitive testing, clinical assessment...Hourly payTraineeshipWork at officeLocal area
- ...requirements. *Participates in site selection including feasibility evaluation. *Familiar with IP needs for project (IP packager/labeler,... ...*5+ years of experience in the pharmaceutical and/or clinical research industry. *3+ years of experience in project management...For contractorsLocal areaRemote work
- Reporting to our Director of Ecosystem Research. Rate - 125k-140k You are a researcher who works at the intersection of market intelligence... ...performing rigorous field scans and secondary research to evaluate markets for product creation. Expert experience building...Work at office3 days per week
$100k - $130k
...development and engagement initiatives. The Consumer Insights & Research function within the team focuses on uncovering strategic... ...recommendations for League and Club stakeholders Managing creative evaluation research to assess campaign, content, and messaging...Hourly payWork experience placementWork at officeLocal areaRemote workRelocationFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Head of Evaluations Research. Be the first to apply!
- director of research New York, NY
- associate director market research New York, NY
- research supervisor New York, NY
- director institutional research New York, NY
- qualitative research director New York, NY
- research project manager New York, NY
- research operations manager New York, NY
- research manager New York, NY
- account manager market research New York, NY
- clinical research manager remote New York, NY



