Evaluation Research Manager
Aaru
About Aaru Aaru builds simulations of human behavior. Each simulation contains a population of AI agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing—from product launches and pricing decisions to strategic communications and policy changes. Building a useful simulation requires more than generating plausible text. Populations must represent real people and groups; predictions must be calibrated; simulations must remain coherent as conditions change; and the product must make the resulting evidence legible enough to support real decisions. We are a small, in-person team in New York. We work with urgency, high ownership, and intellectual honesty. We expect people to surface inconvenient evidence, change their minds quickly, and carry important work all the way to a result. About Evaluation Research Evaluation Research determines whether Aaru's populations, predictions, and end-to-end simulations correspond closely enough to the real world to support consequential decisions. The team defines what should be measured, develops the methods for measuring it, and produces the evidence Aaru uses to improve its systems and describe their capabilities. The function builds both rails and carts. Rails are reusable evaluation infrastructure: datasets, harnesses, libraries, experiment standards, leaderboards, reporting systems, and ways to translate technical evidence into decisions. Carts are the specific evaluations that run on those rails: a historical backtest, a prospective forecast study, a population-coherence test, a reproduction of an observed behavioral pattern, or an end-to-end comparison with a resolved outcome. Evaluation Research is not conventional QA and it is not an internal approval service. It is an independent research function that works closely with the teams building Aaru's systems while preserving the ability to reach and communicate inconvenient conclusions. The role As Evaluation Research Manager, you will lead a focused team of Evaluation Researchers and research engineers. You will translate Aaru's evaluation charter into a coherent portfolio of studies, datasets, and shared infrastructure, and you will be accountable for the quality, pace, and usefulness of the team's work. Managers at Aaru remain researchers. You will design studies, write analysis code, inspect individual failures, review statistical and measurement choices, and directly contribute to the hardest evaluations. You will also set priorities, hire exceptional people, develop the team, provide candid feedback, and create the operating mechanisms that keep protected evidence independent while making diagnostic evidence available quickly. You will work across Population Research, Prediction Research, Simulation Engineering, Product Engineering, Research Product, Deployment, and company leadership. The job requires both scientific independence and practical judgment: an evaluation must be rigorous enough to support a real claim and diagnostic enough to help a team improve the system. What you will do Build, lead, and develop a high-performing team of Evaluation Researchers and research engineers. Turn broad questions about realism, accuracy, calibration, usefulness, and decision quality into measurable constructs, decisive experiments, and explicit decision criteria. Set a focused evaluation agenda across population construction, predictive systems, individual agent behavior, group dynamics, and end-to-end simulations. Decide which evaluation infrastructure should become a reusable organizational rail and which questions require a purpose-built study. Establish standards for baselines, temporal holdouts, prospective testing, contamination control, statistical power, uncertainty, subgroup analysis, and reproducibility. Build tests of population quality that assess individual coherence, joint and conditional distributions, representation of rare but plausible profiles, and whether a profile induces behavior consistent with the person it represents. Evaluate forecasts and other predictive outputs using calibration, proper scoring rules, ranking quality, selective prediction, temporal validity, subgroup performance, and the real cost of different errors. Compare simulations with transactions, product usage, behavioral traces, operational outcomes, market movements, resolved events, and longitudinal decisions. Design end-to-end studies that determine whether improvements to a component actually improve the decision-relevant output customers receive. Find failures hidden by aggregate metrics, especially those concentrated in important subgroups, rare cases, changing environments, or ambiguous ground truth. Create diagnostic evaluations that help researchers localize why a system failed and distinguish a real general improvement from benchmark-specific optimization. Partner with Simulation Engineering to make evaluations repeatable, versioned, scalable, and integrated into development and release workflows without compromising protected holdouts. Convert production incidents, customer surprises, and deployment failures into durable test cases and better measurement methods. Review evidence used in product, customer, or public claims and ensure that conclusions are reproducible, appropriately scoped, and honest about uncertainty and limits. Communicate negative, null, and inconclusive results with the same precision and urgency as positive findings. Recruit exceptional researchers, set clear expectations, provide direct feedback, develop independent research judgment, and address performance problems early. Representative research and leadership problems You might be responsible for situations such as: Two teams disagree about whether a system improved because they use different metrics and test sets. Identify the underlying construct, choose the right evidence, and create a shared evaluation that resolves the disagreement. A population looks plausible one profile at a time but fails to reproduce important real-world relationships. Build measurements that expose the gap and help Population Research identify its cause. A prediction method is well calibrated overall but systematically overconfident for a high-value subgroup. Determine whether the issue is data coverage, model structure, selection, condition shift, or the evaluation itself. An offline benchmark has become a development target and is beginning to leak into decisions. Redesign the evaluation system so teams can iterate quickly without exhausting the integrity of the final holdout. A component metric improves, but customer decisions do not. Determine whether the metric is invalid, the effect is too small, downstream components erase the gain, or the product is presenting the result incorrectly. Ground truth is delayed, noisy, incomplete, or open to multiple interpretations. Design a study that remains useful without pretending that the label is cleaner than it is. A customer outcome contradicts the simulation. Reconstruct the information available at decision time, identify the relevant comparison, and convert the failure into a fair and reproducible test. Evaluation work is becoming a collection of bespoke notebooks. Choose the common abstractions, data contracts, and reporting systems that should become durable rails without freezing the research too early. A favorable result is strategically important but methodologically weak. Communicate the limitation clearly, resist pressure to overclaim, and propose the fastest credible path to stronger evidence. How we work Useful evaluation measures something consequential and helps the company learn. We compare systems with strong alternatives, protect held-out data, quantify uncertainty, and separate exploratory findings from evidence used to support a claim. When ground truth is imperfect, the quality and limits of the outcome data are part of the research problem rather than a footnote. Evaluation should accelerate research by giving teams clear, diagnostic signals about what improved, what did not, and why. It should also make Aaru more trustworthy by exposing failures early and keeping product and external claims aligned with the available evidence. Management in this function requires independence without isolation. The team must understand the systems deeply enough to measure them well, collaborate closely enough to make the results useful, and remain willing to conclude that an attractive idea did not work. You might thrive in this role if You have led evaluation, measurement, or empirical research in machine learning, behavioral science, computational social science, statistics, economics, psychometrics, or a comparably rigorous environment. You have built evaluations that changed a research direction, model capability, product decision, or scientific conclusion. You can define a difficult construct precisely enough to measure it without reducing away the underlying question. You are comfortable with experimental design, observational data, sampling, statistical power, uncertainty, causal threats, leakage, and condition shift. You can write code, analyze large datasets, design studies, inspect individual failures, and review the technical work of researchers and engineers. You can work closely with builders while reaching independent conclusions about the quality of their systems. You care more about an accurate result than a favorable one and are willing to revise your own evaluation when evidence shows it is inadequate. You can prioritize a research portfolio and choose which uncertainty is most important to resolve next. You have managed or technically led strong researchers, give clear feedback, and can develop independent judgment rather than creating dependence on your review. You can explain technical evidence clearly to researchers, engineers, product teams, customers, company leadership, and non-specialists. You want to work in person in New York with a team that moves quickly and takes truth-seeking seriously. Strong candidates may also have Work in ML evaluation, model behavior, forecasting, econometrics, psychometrics, causal inference, experimental economics, survey methodology, or measurement theory. Experience evaluating LLM agents, multi-agent systems, synthetic populations, recommender systems, probabilistic models, simulations, or decision-support tools. Experience with longitudinal records, transaction data, product analytics, field experiments, prospective studies, backtesting, or validation against operational outcomes. Experience building evaluation platforms, regression suites, experiment-tracking systems, shared research datasets, model scorecards, or scientific reporting tools. A record of finding an important failure that standard metrics missed and developing a better measurement method. Experience communicating scientific results in customer-facing, public, policy, legal, or regulatory settings. Experience hiring and leading a small, high-talent research team through ambiguous work with short feedback cycles. There is no requirement for a career spent exclusively in AI evaluation or prior experience with Aaru's exact domain. A PhD, provided you have equivalent evidence of rigorous empirical work and research leadership. Managed managers or a large organization; this role is about leading a focused team and remaining directly involved in the research. Location and benefits This role is based in New York City. Aaru is an in-person company, working five days a week in the office. Candidates should be located in the New York metropolitan area or open to relocation. Aaru offers a competitive base salary, equity participation, comprehensive medical, vision, and dental coverage, visa sponsorship and relocation support, and other benefits and perks. Final compensation depends on level and experience and is set within Aaru's internal bands. #J-18808-Ljbffr Aaru
- Attendance Works is seeking a Director of Evaluation and Research to lead measurement of impact across programs and initiatives, guiding data-informed decision making for continuous improvement. The role collaborates with the CEO, VPs, and directors to shape the organization...Suggested
$275k - $425k
...incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the... ...newest Snorkeler!ABOUT THE ROLE We're looking for a manager to lead a team of researchers to focus on data evaluation, error analysis and data valuation methods to...Suggested$115k - $130k
Senior Manager, Applied Research & Evaluation New York, NY About Understood Understood is a nonprofit focused on shaping the world for difference. We raise awareness of the challenges and strengths of people who learn and think differently. Our resources help people navigate...SuggestedPermanent employmentWork at officeLocal areaVisa sponsorshipWork visa3 days per week$58.66k - $146.02k
DescriptionThe Research Manager oversees research projects within a department or division. This individual assists the Principal Investigator... ...for human resources, includes of recruitment, hiring, evaluating, retention, discipline and termination. 8. Develops and maintains...SuggestedTraineeshipLocal area- Reporting to our Director of Ecosystem Research. Rate - 125k-140k You are a researcher... ...rigorous field scans and secondary research to evaluate markets for product creation. Expert... ...has meaningful results Create and manage processes for both our team and cross‑functional...SuggestedWork at office3 days per week
- .... Position Details: As the Site Manager for the Healthy Brain Network (HBN), you... ...efficiency and administrative duties of the research site. You will oversee all daily... ...clinical direction into daily site operations, evaluate business procedures, and implement...Full timeWork experience placementWork at officeLocal areaFlexible hours
$124k - $335k
...strategic leadership for major tax innovation platforms. As a Senior Manager, you will lead large projects, innovate processes, and maintain... ...assets, or collaborating closely with team members. We evaluate these factors thoughtfully to establish a secure and trusted workplace...Full timeH1b$40 per hour
AIToolboard in the United States is seeking a part-time Remote Chemistry Research Scientist for AI Evaluation to join our expanding AI research team. You will train AI models related to chemical data, test outputs, and provide rigorous evaluation to improve model accuracy...Remote jobHourly payPart timeFlexible hours$140k - $170k
...reliability. Guided by our vision to help researchers rapidly move from the lab to life-... ...everything possible.The Innovation and AI Manager plays a key role in bringing AI strategy... ...both technical teams and business partners.Evaluate and pilot emerging AI tools and...Full timeRemote workWork from homeFlexible hours- ...Clinical Research Manager The candidate will be an integral member of the senior leadership team within the Clinical Protocol & Data... ...providing supervision and delegation of work assignments and evaluation of the CPDM Clinical Research Coordinators (CRCs), Data Coordinators...Local areaFlexible hours
$105k - $120k
...Job Overview We are looking to add an experienced Clinical Research Manager to a growing national organization. This position is fully remote... ...personnel management and employee relations, including evaluating the performance of CRC personnel; approving and submitting all...Full timeContract workWork experience placementLive inLocal areaRemote work- ...the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about... ...Montreal, Seoul, Germany and Paris. Join us! Why this role? Evaluation is critical to making progress in scaling intelligence. As models...Full timeWork at officeLocal areaRemote workHome office
$206.93k - $289k
...The New York State Institute for Basic Research in Developmental Disabilities (IBR) is... ...Core Facility to facilitate the proper evaluation and diagnosis for patients with a suspected... ...Director will be responsible for the management, administration and oversight of all...Permanent employmentFull timeTemporary workPart timeWork at officeLocal areaRemote work$198.2k - $368k
Director, Applied Research Are you excited about working at the forefront of applied research... ...and end-to-end solutions for data management, filing, audit, and compliance. Come... ...symbolic approachesExperience with AI agent evaluation, language model training with...Full timeLocal areaFlexible hours$99k - $266k
...responsible for coaching, leveraging team member’s unique strengths, and managing performance to deliver on client expectations. With your... ...assets, or collaborating closely with team members. We evaluate these factors thoughtfully to establish a secure and trusted workplace...Full timeH1b$140k - $185k
...are seeking an intellectually curious, highly analytical Research Product Manager to lead research and product definition for emerging AP Intelligence... ...with data and product outputs, testing hypotheses, evaluating results and learning from customers. This is not a...Remote work$125k - $140k
...build confidence, and find community. About the Role Senior Research Manager (reports to Director of Ecosystem Research). This is a hybrid... ...experience performing rigorous field scans and secondary research to evaluate markets for product creation. Expert experience building...Permanent employmentWork at officeLocal areaVisa sponsorshipWork visa3 days per week$124k - $335k
...practice, enhancing performance and capabilities. As a Senior Manager, you will leverage your skills and professional networks to deliver... ...assets, or collaborating closely with team members. We evaluate these factors thoughtfully to establish a secure and trusted workplace...Full timeH1b$124k - $335k
...SummaryThe OpportunityAs a Tax Innovation - Data Engineer - Senior Manager, you will focus on designing and building data infrastructure... ...assets, or collaborating closely with team members. We evaluate these factors thoughtfully to establish a secure and trusted workplace...Full timeH1b$90k - $115k
...Saatchi World Services US is seeking a dedicated Insights Manager. The Insights Manager will sit within the Insights team... ...Saatchi World Services’ communications campaigns through research, monitoring, evaluation and learning. The role will work closely with programme management...Work at officeLocal area$100k - $130k
...engagement initiatives. The Consumer Insights & Research function within the team focuses on... ...key areas across the business. The Manager of Consumer Insights position is responsible... ...and Club stakeholders Managing creative evaluation research to assess campaign, content,...Hourly payWork experience placementWork at officeLocal areaRemote workRelocationFlexible hours$90 per hour
...Insights Division of our organization is expanding its nationwide research team. We are seeking dependable individuals to support ongoing... ...feedback through standardized online questionnaires Evaluate consumer products, mobile applications, or digital platforms as...Extra incomeRemote workHome officeFlexible hours$66.3k - $70.5k
...reported outcome instruments, and physical evaluations and histories and enter data into... ...fellows, visiting fellows, residents, other research coordinators and the Division Leadership... ...assist with the writing, submission, and managing of grants and other research funding in...Temporary workAfternoon shift- ...budget, and subject to Fair Market Value evaluation. The hiring range listed is a good... ...Responsibilities: Coordinates clinical research activities for the ARJR service under... ...the field of clinical research project management. o May handle and ship biospecimens...Full timeLocal areaShift work
- ...Director of Research The Agency You'll Join: The New York City Mayor's Office is responsible... ...public agencies and departments, and managing public property. The administration is... ...for leading the Office's research, evaluation, and evidence-building agenda, ensuring...Work at office
$66.3k - $70k
...condition through medical education, scientific research, and direct patient care. At NYU Langone... ...Nurses, Research Pharmacists, Program Managers, Medical Technicians, Clinical... ...advertisement and directly scheduling a visit to evaluate the patient/subject. Reviews all the...TraineeshipWork at office$65.89k - $98.83k
...Description The Research Project Manager I oversees the operational aspects and scope of a specific project or ongoing department initiatives... ...phases of project met. May contribute to the performance evaluation of the employees associated with the project scope and may...TraineeshipWork experience placementLocal area$130.6k - $150.1k
...Title: Research Biostatistician Manager Location: Upper East Side Org Unit: MCC Laboratories Work Days: Weekly Hours: 35.00... ...position assists faculty in the development of study design, evaluation tools and methods associated with various projects. Statistically...Full timeLocal area- ...company powering the acceleration of clinical research to transform patient outcomes. We built... ...worldwide. The Research Finance Manager is a senior leader at the intersection of... ...Participate in site selection discussions by evaluating financial feasibility, budget...Contract workTemporary workWork at officeWorldwide2 days per week
- Scale AI, Inc. is seeking a Research Scientist Manager to lead a world-class team focused on GenAI research, including evaluation, post-training, and RL environments. You will define strategy, mentor researchers, and drive projects from prototyping to deployment. You will...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Evaluation Research Manager. Be the first to apply!
- research coordinator remote New York, NY
- clinical research director New York, NY
- research program manager New York, NY
- associate director market research New York, NY
- qualitative research director New York, NY
- research manager New York, NY
- research supervisor New York, NY
- director institutional research New York, NY
- research project manager New York, NY
- associate director clinical research New York, NY



