Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Manager, Evaluation

$280k - $425k

Aaru Inc.

  • Research Manager, EvaluationTechnical StaffResearchNYC$280K – $425K • Offers Equity## About AaruAaru builds simulations of human behavior. Each simulation contains a population of AI agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing—from product launches and pricing decisions to strategic communications and policy changes.Building a useful simulation requires more than generating plausible text. Populations must represent real people and groups; predictions must be calibrated; simulations must remain coherent as conditions change; and the product must make the resulting evidence legible enough to support real decisions.We are a small, in-person team in New York. We work with urgency, high ownership, and intellectual honesty. We expect people to surface inconvenient evidence, change their minds quickly, and carry important work all the way to a result.## About Evaluation ResearchEvaluation Research determines whether Aaru's populations, predictions, and end-to-end simulations correspond closely enough to the real world to support consequential decisions. The team defines what should be measured, develops the methods for measuring it, and produces the evidence Aaru uses to improve its systems and describe their capabilities.The function builds both **rails** and **carts**. Rails are reusable evaluation infrastructure: datasets, harnesses, libraries, experiment standards, leaderboards, reporting systems, and ways to translate technical evidence into decisions. Carts are the specific evaluations that run on those rails: a historical backtest, a prospective forecast study, a population-coherence test, a reproduction of an observed behavioral pattern, or an end-to-end comparison with a resolved outcome.Evaluation Research is not conventional QA and it is not an internal approval service. It is an independent research function that works closely with the teams building Aaru's systems while preserving the ability to reach and communicate inconvenient conclusions.## The roleAs Evaluation Research Manager, you will lead a focused team of Evaluation Researchers and research engineers. You will translate Aaru's evaluation charter into a coherent portfolio of studies, datasets, and shared infrastructure, and you will be accountable for the quality, pace, and usefulness of the team's work.Managers at Aaru remain researchers. You will design studies, write analysis code, inspect individual failures, review statistical and measurement choices, and directly contribute to the hardest evaluations. You will also set priorities, hire exceptional people, develop the team, provide candid feedback, and create the operating mechanisms that keep protected evidence independent while making diagnostic evidence available quickly.You will work across Population Research, Prediction Research, Simulation Engineering, Product Engineering, Research Product, Deployment, and company leadership. The job requires both scientific independence and practical judgment: an evaluation must be rigorous enough to support a real claim and diagnostic enough to help a team improve the system.## What you will do* Build, lead, and develop a high-performing team of Evaluation Researchers and research engineers.* Turn broad questions about realism, accuracy, calibration, usefulness, and decision quality into measurable constructs, decisive experiments, and explicit decision criteria.* Set a focused evaluation agenda across population construction, predictive systems, individual agent behavior, group dynamics, and end-to-end simulations.* Decide which evaluation infrastructure should become a reusable organizational rail and which questions require a purpose-built study.* Establish standards for baselines, temporal holdouts, prospective testing, contamination control, statistical power, uncertainty, subgroup analysis, and reproducibility.* Build tests of population quality that assess individual coherence, joint and conditional distributions, representation of rare but plausible profiles, and whether a profile induces behavior consistent with the person it represents.* Evaluate forecasts and other predictive outputs using calibration, proper scoring rules, ranking quality, selective prediction, temporal validity, subgroup performance, and the real cost of different errors.* Compare simulations with transactions, product usage, behavioral traces, operational outcomes, market movements, resolved events, and longitudinal decisions.* Design end-to-end studies that determine whether improvements to a component actually improve the decision-relevant output customers receive.* Find failures hidden by aggregate metrics, especially those concentrated in important subgroups, rare cases, changing environments, or ambiguous ground truth.* Create diagnostic evaluations that help researchers localize why a system failed and distinguish a real general improvement from benchmark-specific optimization.* Partner with Simulation Engineering to make evaluations repeatable, versioned, scalable, and integrated into development and release workflows without compromising protected holdouts.* Convert production incidents, customer surprises, and deployment failures into durable test cases and better measurement methods.* Review evidence used in product, customer, or public claims and ensure that conclusions are reproducible, appropriately scoped, and honest about uncertainty and limits.* Communicate negative, null, and inconclusive results with the same precision and urgency as positive findings.* Recruit exceptional researchers, set clear expectations, provide direct feedback, develop independent research judgment, and address performance problems early.## Representative research and leadership problemsYou might be responsible for situations such as:* Two teams disagree about whether a system improved because they use different metrics and test sets. Identify the underlying construct, choose the right evidence, and create a shared evaluation that resolves the disagreement.* A population looks plausible one profile at a time but fails to reproduce important real-world relationships. Build measurements that expose the gap and help Population Research identify its cause.* A prediction method is well calibrated overall but systematically overconfident for a high-value subgroup. Determine whether the issue is data coverage, model structure, selection, condition shift, or the evaluation itself.* An offline benchmark has become a development target and is beginning to leak into decisions. Redesign the evaluation system so teams can iterate quickly without exhausting the integrity of the final holdout.* A component metric improves, but customer decisions do not. Determine whether the metric is invalid, the effect is too small, downstream components erase the gain, or the product is presenting the result incorrectly.* Ground truth is delayed, noisy, incomplete, or open to multiple interpretations. Design a study that remains useful without pretending that the label is cleaner than it is.* A customer outcome contradicts the simulation. Reconstruct the information available at decision time, identify the relevant comparison, and convert the failure into a fair and reproducible test.* Evaluation work is becoming a collection of bespoke notebooks. Choose the common abstractions, data contracts, and reporting systems that should become durable rails without freezing the research too early.* A favorable result is strategically important but methodologically weak. Communicate the limitation clearly, resist pressure to overclaim, and propose the fastest credible path to stronger evidence.## How we workA useful evaluation measures something consequential and helps the company learn. We compare systems with strong alternatives, protect held-out data, quantify uncertainty, and separate exploratory findings from evidence used to support a claim. When ground truth is imperfect, the quality and limits of the outcome data are part of the research problem rather than a footnote.Evaluation should accelerate research by giving teams clear, diagnostic signals about what improved, what did not, and why. It should also make Aaru more trustworthy by exposing failures early and keeping product and external claims aligned with the available evidence.Management in this function requires independence without isolation. The team must understand the systems deeply enough to measure them well, collaborate closely enough to make the results useful, and remain willing to conclude that an attractive idea did not work.## You might thrive in this role if* You have led evaluation, measurement, or empirical research in machine learning, behavioral science, computational social science, statistics, economics, psychometrics, or a comparably rigorous environment.* You have built evaluations that changed a research direction, model capability, product decision, or scientific conclusion.* You can define a difficult construct precisely enough to measure it without reducing away the underlying question.* You are comfortable with experimental design, observational data, sampling, statistical power, uncertainty, causal threats, leakage, and condition shift.* You can write code, analyze large datasets, design studies, inspect individual failures, and review the technical work of researchers and engineers.* You can work closely with builders while reaching independent conclusions about the quality of their systems.* You care more about an accurate result than a favorable one and are willing to revise your own evaluation when evidence shows it is inadequate.* You can prioritize a research portfolio and choose which uncertainty is most important to resolve next.* You have managed or technically led strong researchers, give clear feedback, and can develop independent judgment rather than creating dependence on your review.* You can explain technical evidence clearly to researchers, engineers, product teams, customers, company leadership, and non-specialists.* You want to work in person in New York with a team that moves quickly and takes truth-seeking seriously.## Strong candidates may also have* Work in ML evaluation, model behavior, forecasting, econometrics, psychometrics, causal inference, experimental economics, survey methodology, or measurement theory.* Experience evaluating LLM agents, multi-agent systems, synthetic populations, recommender systems, probabilistic models, simulations, or decision-support tools.* Experience with longitudinal records, transaction data, product analytics, field experiments, prospective studies, backtesting, or validation against operational outcomes.* Experience building evaluation platforms, regression suites, experiment-tracking systems, shared research datasets, model scorecards, or scientific reporting tools.* A record of finding an important failure that standard metrics missed and developing a better measurement method.* Experience communicating scientific results in customer-facing, public, policy, legal, or regulatory settings.* Experience hiring and leading a small, high-talent research team through ambiguous work with short feedback cycles.## Candidates need not have* A career spent exclusively in AI evaluation or prior experience with Aaru's exact domain.* A PhD, provided you have equivalent evidence of rigorous empirical work and research leadership.* Managed managers or a large organization; this role is about leading a focused team and remaining directly involved in the research.## What success looks like* Aaru has clear, widely trusted definitions of quality across populations, predictions, and end-to-end simulations.* Core evaluations are grounded in observed outcomes and reveal whether performance generalizes across time, domains, populations, and subgroups.* Researchers receive fast, diagnostic measurements while protected holdouts preserve the integrity of final results.* Important failures are found early, explained clearly, and converted into durable datasets, tests, or research questions.* Reusable evaluation rails reduce duplicated work and make results comparable across projects and versions.* Product and company decisions rely on evidence that is reproducible, appropriately uncertain, and connected to real-world value.* Customers and external audiences can understand what Aaru has demonstrated, where the evidence is limited, and how confidence should be interpreted.* The team develops a reputation inside and outside Aaru for evaluation work that is scientifically rigorous, practically useful, and unusually honest.* Evaluation Researchers grow into independent owners of important measurement areas, and the team maintains a high bar as it expands.## Location and benefitsThis role is based in New York City. Aaru is an in-person company, working five days a week in the office. Candidates should be located in the New York metropolitan area or open to relocation.Aaru offers a competitive base salary, equity participation, comprehensive medical, vision, and dental coverage, visa sponsorship and relocation support, and other benefits and perks. Final compensation depends on level and experience and is set within Aaru's internal bands.
  • J-18808-Ljbffr Aaru Inc.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Research Manager, Evaluation in New York, NY vacancy
  • $280k - $425k

    # Research Manager, PredictionTechnical StaffResearchNYC$280K - $425K • Offers Equity## About AaruAaru builds simulations of human behavior...  ...evidence.You will work closely with Population Research, Evaluation Research, Simulation Engineering, Product Engineering, Research... 
    Suggested
    Work at office
    Relocation
    Visa sponsorship
    Relocation package
    Shift work

    Aaru Inc.

    New York, NY
    2 days ago
  • $86k - $131.38k

     ..., internal equity, budget, and subject to Fair Market Value evaluation. The hiring range listed is a good faith determination of potential...  ...may be modified in the future. What you will be doing Research Manager , The Center for Regenerative Medicine Position Summary... 
    Suggested
    Full time
    Contract work
    Work at office
    Local area
    Shift work

    Hospital for Special Surgery

    New York, NY
    2 days ago
  • $315k

    We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand what AI Safety Level (ASL) to assign...  ...skills (e.g. experience in Python) Have experience managing research programs of dozens of technical and non-technical... 
    Suggested
    Currently hiring
    Work at office
    Immediate start
    Home office
    Visa sponsorship
    Relocation package

    Anthropic

    New York, NY
    2 days ago
  • $40 per hour

    AIToolboard in the United States is seeking a part-time Remote Chemistry Research Scientist for AI Evaluation to join our expanding AI research team. You will train AI models related to chemical data, test outputs, and provide rigorous evaluation to improve model accuracy... 
    Suggested
    Remote job
    Hourly pay
    Part time
    Flexible hours

    AIToolboard

    New York, NY
    4 days ago
  • $115k - $130k

     ...scales rapidly, we’re looking for a strategic and hands-on Manager of Science & Research to play a critical role in shaping the scientific...  ...this role, you will: Claims Research & Substantiation : Evaluate existing and emerging ingredients to develop new, evidence... 
    Suggested
    Fixed term contract
    Remote work

    Grüns

    New York, NY
    1 day ago
  • $100k - $110k

     ...senior leadership team within the Clinical Protocol & Data Management (CPDM) Office of the Herbert Irving Comprehensive...  ...providing supervision and delegation of work assignments and evaluation of the CPDM Clinical Research Coordinators (CRCs), Data Coordinators(DCs), Research... 
    Temporary work
    Local area
    Flexible hours

    Columbia University in the City of New York

    New York, NY
    2 days ago
  • Mercor connects elite creative and technical talent with leading AI research labs. A remote, 8-10 week contract invites a PhD physicist to help align AI outputs with physical principles. You will evaluate AI-generated physics content, develop benchmarks, and provide... 
    Remote job
    Contract work

    Remote Jobs

    New York, NY
    1 day ago
  • $198.2k - $368k

    Job Description: Director of Applied Research, Information RetrievalThomson Reuters Labs...  ...and contextual retrieval techniquesBuild evaluation frameworks and benchmarks for search quality...  ...a High-Performing TeamHire, train, and manage a highly skilled team with expertise... 
    Full time
    Local area
    Flexible hours

    Thomson Reuters

    New York, NY
    2 days ago
  • $150k

     ...fostering an environment where people and technology thrive together - Managing stakeholder relationships and confirming the delivery of...  ...assets, or collaborating closely with team members. We evaluate these factors thoughtfully to establish a secure and trusted workplace... 
    Full time
    H1b

    PwC

    New York, NY
    4 days ago
  • $105k - $120k

     ...Job Overview We are looking to add an experienced Clinical Research Manager to a growing national organization. This position is fully remote...  ...personnel management and employee relations, including evaluating the performance of CRC personnel; approving and submitting all... 
    Full time
    Contract work
    Work experience placement
    Live in
    Local area
    Remote work

    Medix™

    New York, NY
    3 days ago
  •  ...our members. We’re transforming wealth management through the perfect blend of cutting-edge...  ...role As Range's Head of Quantitative Research, you will own the research and modeling...  ...and rebalancing behavior after launch Evaluate new data sources, market data vendors, and... 
    Work at office
    Relocation
    Monday to Friday

    Range

    New York, NY
    4 days ago
  • $110k - $125k

     ...Join to apply for the Innovation Brand Manager role at izzio artisan bakery Get AI-powered...  ...Validation & Market Analysis Evaluate commercial opportunities for innovation,...  ...management. Identify and support consumer research initiatives as needed, translating... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Izzio Artisan Bakery

    New York, NY
    3 days ago
  • $110k - $130k

     ...position reports to the Vice President of Sports & Entertainment Research and will play a key role in gathering, analyzing, and...  ...across different platforms, generating key metrics and reports to evaluate content effectiveness. Present findings and recommendations... 
    Full time
    Work at office
    Local area
    3 days per week

    Versant

    New York, NY
    1 day ago
  • $176.6k - $294.3k

     ...The role combines organizational, project management, strategic and tactical planning...  ...develop and implement team priorities, evaluate strategic trade-off decisions, and identify...  ...reimbursement, and patient advocacy insights and research to help gauge the impact of evolving... 
    Permanent employment
    Temporary work
    Work experience placement
    H1b
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Work visa
    Relocation package

    Pfizer

    New York, NY
    21 hours ago
  • Latham & Watkins is seeking a Research Services Manager - Operations, AI Platforms to lead evaluation, testing, and deployment of AI-powered research tools across the firm’s New York or Washington, DC offices, with a hybrid in-office and flexible schedule. You will coordinate... 
    Work at office
    Flexible hours

    careers-lw

    New York, NY
    1 day ago
  • Latham & Watkins is seeking a Research Services Manager - Operations, AI Platforms to lead evaluation, testing, and implementation of AI-powered research services. You will develop testing protocols, coordinate pilots, collect feedback, and advise on adoption strategies... 
    Work at office

    Careers

    New York, NY
    1 day ago
  • $66.3k - $70k

     ...condition through medical education, scientific research, and direct patient care. At NYU Langone...  ...Nurses, Research Pharmacists, Program Managers, Medical Technicians, Clinical...  ...Making and Problems Solving Combines and evaluates information and data to make decisions about... 
    Work at office

    NYU Langone Health

    New York, NY
    1 day ago
  • $167k - $235k

     ...Practice Solutions Manager Build your big career with the firm that does Big Law, Better. McDermott Will & Schulte is a leading...  ...attorney team to help identify practice-level needs and support the evaluation and implementation of solutions to address them Support... 
    Full time
    Work at office

    McDermott Will & Schulte

    New York, NY
    5 days ago
  • $75.8k - $89.79k

    Clinical Research Project Manager B University Overview The University of Pennsylvania, the largest private employer in Philadelphia, is a world-renowned leader in education, research, and innovation. This historic, Ivy League school consistently ranks among the top 10... 
    Local area
    Work from home
    Flexible hours

    University of Pennsylvania

    New York, NY
    1 day ago
  • $55.24 - $61.38 per hour

     ...advance the company's policy objectives. Develop and project manage innovative policy programs and campaigns, including those for new...  ...awesome benefits! Aquent is an equal-opportunity employer. We evaluate qualified applicants without regard to race, color, religion,... 
    Hourly pay
    Temporary work

    Skill

    New York, NY
    3 days ago
  • $130k - $160k

    Research & Insights Director Orange Barrel Media, LLC New York, New York, United States...  ...objectives. What you'll do Own day-to-day management of research and insights projects for...  ...considering budget, scope and projected impact Evaluate new data sources and tools, advising... 
    Work at office

    Orange Barrel Media, LLC

    New York, NY
    5 days ago
  •  ...to join a tight-knit, high-calibre team at the intersection of research, data science, and creative strategy. The Role This is a senior...  ...lifecycle - from designing randomised controlled trials and managing data quality through to delivering strategic recommendations that... 
    3 days per week

    Metrical Talent

    New York, NY
    1 day ago
  • $66.3k - $70.5k

     ...reported outcome instruments, and physical evaluations and histories and enter data into...  ...fellows, visiting fellows, residents, other research coordinators and the Division Leadership...  ...assist with the writing, submission, and managing of grants and other research funding in... 
    Temporary work
    Afternoon shift

    Columbia University Irving Medical Center

    New York, NY
    1 day ago
  • $165k - $185k

     ...International Rescue Committee’s (IRC’s) research and innovation team -- designs, tests,...  ...We are a leading contributor of impact evaluations and other research in humanitarian settings...  ...Oversee research agendas and manage select research teams to deliver high quality... 
    Work experience placement

    International Rescue Committee

    New York, NY
    5 days ago
  • BE Creative. BE Ambitious. BE a Team Player. Be all that and more at Colliers International. Join our team as a Research Manager in Manhattan, NY. At Colliers International, we help leaders succeed by building amazing workplaces, businesses, and communities worldwide.... 
    Work at office
    Worldwide
    Flexible hours

    Colliers International

    New York, NY
    2 days ago
  • The Director of Research will work collaboratively with colleagues across StartCare departments...  ..., data cleaning and processing, and management plans to ensure the integrity of study/...  ...to the Vice President of Research and Evaluation, the Director of Research leads a team... 
    Local area

    StartCare

    New York, NY
    3 days ago
  • The Director of Research and Engagement provides leadership and management for SFA’s strategic goals to advance research. This position will engage the sarcoma...  ...Scientific Affairs to oversee the design, delivery, evaluation, and quality of SFA’s grant programs. Ensure that... 
    Full time
    Work at office
    Remote work

    Sarcoma Foundation of America

    New York, NY
    3 days ago
  •  ...OVERVIEW JOB TITLE Director of Research Implementation DEPARTMENT...  ...onsite work arrangement as determined by management. FT/PT FULL TIME ☒ PART TIME...  ...professional development opportunities, evaluating performance, delivering feedback and coaching... 
    Full time
    Part time
    Work at office
    Remote work

    Child Mind Institute

    New York, NY
    21 hours ago
  • Manager, Enterprise Talent Services at Motion Recruitment Responsibilities Collaborate with the Vice President and team to design, execute, analyze, and communicate quantitative and qualitative research projects (surveys, online research, focus groups, and more) focused... 

    Motion Recruitment

    New York, NY
    4 days ago
  •  ...budget, and subject to Fair Market Value evaluation. The hiring range listed is a good...  ...JOB DESCRIPTION Job Code/Title: Clinical Research Coordinator Salary: TBD Reports to: Arthroplasty...  ...participate in all aspects of research management and quality assurance of data for the... 
    Full time
    Work at office
    Local area
    Shift work

    Hospital for Special Surgery

    New York, NY
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Manager, Evaluation. Be the first to apply!