Evaluation Researcher
Aaru
About Aaru Aaru builds simulations of human behavior. Each simulation contains a population of AI agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing—from product launches and pricing decisions to strategic communications and policy changes. Building a useful simulation requires more than generating plausible text. Populations must represent real people and groups; predictions must be calibrated; simulations must remain coherent as conditions change; and the product must make the resulting evidence legible enough to support real decisions. We are a small, in-person team in New York. We work with urgency, high ownership, and intellectual honesty. We expect people to surface inconvenient evidence, change their minds quickly, and carry important work all the way to a result. About Evaluation Research Evaluation Research determines whether Aaru's populations, predictions, and end-to-end simulations correspond closely enough to the real world to support consequential decisions. The team defines what should be measured, develops the methods for measuring it, and produces the evidence Aaru uses to improve its systems and describe their capabilities. The function builds both rails and carts. Rails are reusable evaluation infrastructure: datasets, harnesses, libraries, experiment standards, leaderboards, reporting systems, and ways to translate technical evidence into decisions. Carts are the specific evaluations that run on those rails: a historical backtest, a prospective forecast study, a population-coherence test, a reproduction of an observed behavioral pattern, or an end-to-end comparison with a resolved outcome. Evaluation Research is not conventional QA and it is not benchmark administration. It is an independent research function. The work requires understanding the systems deeply, collaborating closely with their builders, and remaining willing to conclude that an attractive method did not improve what matters. The role As an Evaluation Researcher, you will own difficult measurement problems at the boundary of machine learning, statistics, behavioral science, and product decision-making. You will define constructs, assemble or create evaluation data, design studies, write analysis and evaluation code, inspect individual failures, quantify uncertainty, and communicate what the evidence does and does not support. Some projects will build reusable rails used across the research organization. Others will be focused studies intended to resolve one important uncertainty. In both cases, the goal is the same: create an evaluation that is valid enough to trust, diagnostic enough to guide improvement, and clear enough to inform a real decision. You will work closely with Population Research, Prediction Research, Simulation Engineering, Product Engineering, Research Product, and Deployment while protecting the independence and integrity of final measurements. What you will do Own an important evaluation or measurement area across population construction, predictive systems, individual agent behavior, group dynamics, or end-to-end simulations. Turn broad questions about realism, accuracy, calibration, usefulness, and decision quality into measurable constructs and explicit decision criteria. Design studies using historical backtests, temporal holdouts, prospective outcomes, observational records, controlled experiments, expert judgment, or mixed methods as the problem requires. Build tests of individual-profile quality, including internal coherence, contradictions, impossible combinations, unsupported specificity, stability, and whether a profile induces behavior consistent with the represented person. Build tests of population quality, including marginal and joint distributions, conditional relationships, coverage of rare but plausible profiles, subgroup fidelity, and sensitivity to sampling choices. Evaluate predictions using calibration, proper scoring rules, ranking quality, selective prediction, temporal validity, subgroup performance, and the decision cost of different errors. Compare simulations with transactions, product usage, behavioral traces, operational outcomes, market movements, resolved events, surveys, and longitudinal decisions. Design longitudinal and interaction-based evaluations that test how agents change over time, respond to new information, and influence one another. Build end-to-end studies that determine whether a component improvement actually changes the quality of the conclusion a customer receives. Establish strong baselines and compare agent-based simulation with direct forecasting, conventional statistical models, simpler segment-level methods, and human or market benchmarks where appropriate. Find failures hidden by aggregate metrics, especially failures concentrated in important subgroups, rare cases, changing environments, or ambiguous labels. Build diagnostic evaluations that localize why a system failed and whether a proposed fix generalizes beyond the development set. Develop reusable evaluation datasets, harnesses, libraries, graders, experiment schemas, leaderboards, and reporting tools where shared infrastructure will accelerate future research. Work with Simulation Engineering to version and automate evaluations while keeping protected holdouts and final claims insulated from development leakage. Convert customer surprises, production incidents, and resolved real-world outcomes into durable test cases. Write clear technical reports that distinguish exploratory evidence from claim-supporting evidence and communicate uncertainty, limitations, and alternative interpretations. Report negative, null, and inconclusive findings with the same care as positive results. Representative research directions You might investigate questions such as: Construct a population using data available at one point in time, then test whether its future transactions, choices, or behavioral outcomes match what later occurred. Recreate a historical decision environment using only information available before the outcome and compare the simulation with the observed result. Run prospective evaluations in which Aaru records predictions before outcomes are known and tracks performance as those outcomes resolve. Determine whether profile-coherence scores, population-distribution metrics, or forecast calibration predict end-to-end simulation quality. Measure whether simulated groups reproduce observed patterns in information diffusion, coordination, influence, polarization, or collective choice. Study how population size, heterogeneity, interaction structure, model capability, context length, and computation affect simulation fidelity. Build tests for memorization, leakage, prompt sensitivity, unsupported certainty, judge bias, and benchmark-specific overfitting. Develop evaluation methods for settings where outcomes are delayed, noisy, partially observed, selected, or open to more than one reasonable interpretation. Compare automated graders, expert review, behavioral outcomes, and human-subject measurements to determine when each is a valid proxy. Estimate whether the measured gain is large enough to matter for the decision, not merely statistically distinguishable from zero. How we work A useful evaluation measures something consequential and helps the company learn. We compare systems with strong alternatives, protect held-out data, quantify uncertainty, and separate exploratory findings from evidence used to support a claim. When ground truth is imperfect, the quality and limits of the outcome data are part of the research problem rather than a footnote. Evaluation should make research faster by giving teams clear signals about what improved, what did not, and why. It should also make Aaru more trustworthy by exposing failures early and keeping product and external claims aligned with the evidence. Researchers own the full arc of the work: construct definition, data provenance, implementation, analysis, failure inspection, interpretation, communication, and the decision that follows. A polished metric without a valid construct is not success. Strong candidates may also have Work in forecasting evaluation, econometrics, psychometrics, causal inference, survey methodology, experimental economics, measurement theory, or model behavior. Experience evaluating LLM agents, multi-agent systems, synthetic populations, recommender systems, probabilistic models, simulations, or decision-support tools. Experience with longitudinal records, transaction data, product analytics, field experiments, prospective studies, or operational validation. Experience building evaluation platforms, regression suites, experiment-tracking systems, shared datasets, model scorecards, or scientific reporting tools. Experience with automated graders, human evaluation, rubric design, inter-rater reliability, benchmark contamination, or adversarial evaluation. A record of finding an important failure that standard metrics missed and developing a better way to measure it. Experience communicating scientific results in customer-facing, public, policy, legal, or regulatory settings. Candidates need not have prior experience in population simulation or a job title containing the word "evaluation." A PhD, provided you can demonstrate equivalent research depth and empirical rigor. Expertise in every statistical or machine-learning method listed above. We care most about measurement judgment, technical execution, and intellectual honesty. What success looks like You create evaluations that resolve important uncertainty rather than merely producing additional metrics. Your measurements are valid enough to support decisions, diagnostic enough to guide improvement, and reproducible enough for others to challenge. Important failures are found early, explained clearly, and converted into durable datasets, tests, or research questions. Protected evidence remains trustworthy while development teams still receive fast feedback. Reusable rails reduce duplicated work and make results comparable across projects, systems, and versions. Product and research claims become more precise because their supporting evidence and limitations are explicit. Your work helps Aaru distinguish a genuine general improvement from overfitting, leakage, a proxy failure, or a favorable anecdote. Colleagues trust your conclusions because you combine scientific rigor with a practical understanding of how systems improve. Location and benefits This role is based in New York City. Aaru is an in-person company, working five days a week in the office. Candidates should be located in the New York metropolitan area or open to relocation. Final compensation depends on level and experience and is set within Aaru's internal bands. competitive base salary equity participation comprehensive medical coverage vision coverage dental coverage visa sponsorship relocation support other benefits and perks #J-18808-Ljbffr Aaru
$196k - $230k
...deeply about giving our customers more time for their life’s work.About the Role:We’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered experiences—focusing on what “good” looks like not only for model output quality, but for...SuggestedLocal areaShift work- Aaru is seeking an Evaluation Researcher to tackle challenging measurement problems at the intersection of machine learning and behavioral science. In this role, you will design studies, build tests, and communicate evidence to impact decision-making. Located in New York...Suggested
- Aaru in New York is hiring an Evaluation Research Manager to lead a focused team of researchers and engineers. You will translate the evaluation charter into a coherent portfolio of studies, datasets, and reusable infrastructure, ensuring quality and usefulness of the team...Suggested
- Aaru in New York City seeks an Evaluation Researcher to tackle measurement problems at the boundary of ML, statistics, and behavioral science. You will define constructs, assemble data, design studies, write analysis code, and clearly communicate uncertainty. You will build...Suggested
- Notion is looking for an experienced UX Researcher in New York City to define and create evaluation frameworks for AI products. The role involves translating user insights into actionable guidelines and collaborating with cross-functional teams to improve product experiences...Suggested
$196k - $230k
Notion is seeking an experienced UX Researcher based in either San Francisco or New York City. This role focuses on evaluating AI-powered experiences, ensuring quality through defined evaluation criteria and user insights. The successful candidate will work closely with...$196k - $230k
...deeply about giving our customers more time for their life’s work. About The Role We’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI‑powered experiences—focusing on what “good” looks like not only for model output quality, but for...Local areaShift work$196k - $230k
...deeply about giving our customers more time for their life’s work. About the Role: We’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered experiences—focusing on what “good” looks like not only for model output quality, but for...Local areaShift work- ...A health tech company in New York City is hiring a Research Scientist to evaluate the impact of ambient AI on healthcare outcomes. The role emphasizes designing studies, engaging with health systems, and fostering collaboration across product teams. A PhD in a relevant...Work at office
$180.6k - $225.75k
...to provide high quality data and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers... ...expertise in LLM post-training (SFT, RLHF, reward modeling) and evaluation. This role is on the evaluation pod within the GenAI Research...Full time$216k - $270k
Scale Labs, Research Scientist — Frontier Risk EvaluationsAs the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs...Full time$200k - $300k
...provide liquidity to the global markets.As a Machine Learning Researcher at Virtu, you'll pursue high-impact research opportunities within... ...of what industry you come from. The RoleInvestigate, evaluate, and prototype innovative algorithmic solutions using novel machine...$150k - $200k
Hudson River Trading (HRT) is seeking a Quantitative Researcher focused on Treasury Optimization and Research to join our PostTrade team... ...instructed or agreed upon. We employ various methods to evaluate the authenticity of candidate responses. If we determine that...Work at officeImmediate start- ...Description SummaryBalyasny Asset Management is seeking a Quantitative Researcher to join our Commodities team. This front-office role will... ...building models used directly in risk platforms, trade evaluation, or idea generation.• Familiarity with SQL/NoSQL databases, distributed...
$200k - $300k
Hudson River Trading (HRT) is seeking an LLM-focused AI Researcher to join the HAIL team. HAIL (HRT AI Labs) is the team at HRT responsible... ..., dataset curation and mixing, pretraining, post-training, evaluation design, inference, and live trading. HAIL researchers have...Work at officeLocal areaImmediate start$200k - $300k
Hudson River Trading (HRT) is hiring an AI Researcher to join the HAIL team. HAIL (HRT AI Labs) is the team at HRT responsible for developing... ...instructed or agreed upon. We employ various methods to evaluate the authenticity of candidate responses. If we determine that...Work experience placementWork at officeLocal areaImmediate start- OpenRouter in New York seeks a Research Scientist to advance how the world understands, evaluates, and routes large language models. You will design experiments, build evaluation frameworks, and publish findings that influence rankings and routing decisions. You will collaborate...
- ...the SoHo neighborhood of New York, and East Liberty in Pittsburgh. The Role Abridge is hiring Research Scientists to join our Strategic Research team to rigorously evaluate and advance the real-world impact of ambient AI on patient outcomes, care quality, and provider...Hourly payFull timeWork at officeRelocation packageFlexible hours
- UJA-Federation of New York is seeking a Research and Evaluation Analyst to partner with CPER in assessing the effectiveness of programs. You will design studies, analyze data, and communicate findings to guide decisions and improve impact. The ideal candidate has strong...
- ...the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about... ...City, Montreal, Seoul, Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling intelligence. As...Full timeWork at officeLocal areaRemote workHome office
$175k - $250k
...Job Title: Quantitative Researcher Department: Global Markets Location: New York Corporate Title: Associate/Vice President Pay range:... ...risk from client facilitation and market making Create tools for evaluating trade-offs between risk reduction, capital efficiency, and...Visa sponsorshipRelocation package$278.4k - $317.7k
Distinguished Applied Researcher Overview: At Capital One, we are creating trustworthy and reliable AI systems, changing banking for... ...through all phases of development, from design through training, evaluation, validation, and implementation. ~ Engage in high impact...Full timePart timeLocal areaFlexible hours- ...Pantera and Mark Cuban. The Role We are looking for a Quantitative Researcher to fit into our existing highly-skilled NY-based quantitative... ...across strategies and venues. Run post-trade analytics to evaluate execution quality, slippage, and market impact. Develop risk metrics...Contract workImmediate startHome officeFlexible hours
- ...Build AI foundation models through all phases of development, from design through training, evaluation, validation, and implementation.* Engage in high impact applied research to take the latest AI developments and push them into the next generation of customer experiences...Full timePart timeFlexible hours
- ...a Post-Doctoral Fellow in the Center for Data Analytics, Innovation, and Rigor (DAIR Center). The role focuses on designing and evaluating multimodal data analyses, developing ML models, and publishing findings in peer-reviewed journals. The position is exempt, full-time...Full timeWork at office
- ...Clinical Researcher Level IIGrade: 23Salary: $21,625.50 to $26,100The Research Foundation for Mental Hygiene is seeking a qualified candidate... ...per week) are currently being invited. Our work is focused on evaluating research intervention for adult patients with anorexia nervosa...Part time
$225k - $325k
...ML Researcher | Healthcare AI | New York | $225,000–$325,000 + Equity One of the most exciting AI opportunities I've worked on this year... ..., they're creating the reinforcement learning environments, evaluation frameworks and verification systems that help frontier AI labs...Work at officeVisa sponsorship$262.5k - $299.6k
...at Capital One to life. Our work touches every aspect of the research life cycle, from partnering with academia to building production... ...all phases of development, from design through training, evaluation, validation, and implementation. Engage in high impact applied...Full timePart timeLocal area$218.7k - $249.6k
...at Capital One to life. Our work touches every aspect of the research life cycle, from partnering with academia to building production... ...all phases of development, from design through training, evaluation, validation, and implementation. Engage in high‑impact applied...Full timePart timeLocal areaFlexible hours$180k - $280k
...edge AI and impactful learning outcomes. As a Machine Learning Researcher , you will play a pivotal role in pushing the boundaries of... ...engineering, Retrieval-Augmented Generation (RAG), fine-tuning, and evaluation of large language model applications. Proficiency in Python...Permanent employmentFull timeLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Evaluation Researcher. Be the first to apply!
- senior design researcher New York, NY
- trend researcher New York, NY
- vulnerability researcher New York, NY
- researcher New York, NY
- music researcher New York, NY
- legal researcher New York, NY
- remote researcher New York, NY
- lead researcher New York, NY
- title researcher New York, NY
- product researcher New York, NY

