Evaluation Researcher: Behavior, ML & Simulation
association of arab universities
Aaru is seeking an Evaluation Researcher to tackle challenging measurement problems at the intersection of machine learning and behavioral science. In this role, you will design studies, build tests, and communicate evidence to impact decision-making. Located in New York City, the position offers a unique opportunity to collaborate closely with cross-functional teams while ensuring the independence of measurement integrity. #J-18808-Ljbffr association of arab universities
- Aaru in New York City seeks an Evaluation Researcher to tackle measurement problems at the boundary of ML, statistics, and behavioral science. You will define constructs, assemble data, design studies, write analysis code, and clearly communicate uncertainty. You will...Suggested
$196k - $230k
...Role:We’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered experiences—... ...expectations, and failure/recovery behaviors, then translate those insights into... ...automation) and working with Data Science/ML partners on measurement strategy...SuggestedLocal areaShift work- About Aaru Aaru builds simulations of human behavior. Each simulation contains a population of AI agents, each representing a person who could... ...carry important work all the way to a result. About Evaluation Research Evaluation Research determines whether Aaru's populations...SuggestedWork at officeRelocationVisa sponsorshipRelocation package
$220k
...intelligence, using AI to simulate and predict human behavior at scale. By generating and... .... About The Role As a Researcher at Aaru, you will design,... ...unstructured profile data Build and evaluate simulation pipelines that... ...You have a PhD in ML, computational social science...SuggestedWork at officeRelocationVisa sponsorshipRelocation package$196k - $230k
...Role We’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI‑powered experiences—... ...expectations, and failure/recovery behaviors, then translate those insights into... ...automation) and working with Data Science/ML partners on measurement strategy...SuggestedLocal areaShift work$196k - $230k
...Role: We’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered experiences—... ...expectations, and failure/recovery behaviors, then translate those insights into... ...automation) and working with Data Science/ML partners on measurement strategy...Local areaShift work$184.7k - $324.8k
...drawn to hard problems where the research and the product are... ...multi-turn RL for coding agents evaluated on SWE-Bench and beyond, scaling... ...learning with publications at top ML or NLP conferences, or a... ...horizon task execution, user simulation Distillation and alignment: on...Relocation$160k - $300k
...remain deeply embedded in AI research, and we’re channeling that scientific... ..., design an intervention, evaluate it against real customer... ...modeling techniques to align AI behavior with real‑world SRE workflows.... ...Some experience shipping AI or ML systems to production Ability...Full timeWork at officeFlexible hours$180k - $350k
Hume AI is seeking talented AI researchers interested in working with... ...learns human preferences from behavior in millions of audio and video... ...‑of‑the‑art audio models and evaluation tools. In this role, you will... ...Python ecosystem and popular ML libraries and tools (e.g....- .... Our applications of AI and ML bring humanity and simplicity... ...touching every aspect of the research lifecycle, from partnering with... ...design through training, evaluation, validation, and implementation... ...EMNLP, NeurIPS, ICML or ICLR Behavioral Models PhD focus on...Full timePart time
- JPMorganChase AI Research is a global team of research scientists,... ...to translate breakthrough AI/ML techniques into deployed solutions... ...innovation and rigorous evaluation to production-scale delivery... ...synthetic personas and agent simulations. Multimodal agent security and...Work at officeShift work
$180.6k - $225.75k
...quality data and accelerate progress in GenAI research. We are looking for Research Scientists... ...(SFT, RLHF, reward modeling) and evaluation. This role is on the evaluation pod within... ...generative AI models.You will:Analyze model behavior to identify, characterize, and diagnose...Full time$262.5k - $326.8k
Applied Researcher II (AI Foundations, LLM Core and Agentic AI) Overview: At Capital One,... ...in real time, our applications of AI & ML are bringing humanity and simplicity to... ...development, from design through training, evaluation, validation, and implementation. Engage...Full timePart timeLocal areaFlexible hours$200k - $300k
...the global markets.As a Machine Learning Researcher at Virtu, you'll pursue high-impact research... ...you come from. The RoleInvestigate, evaluate, and prototype innovative algorithmic solutions... ...impact P&LImplement sophisticated ML approaches for forecasting, feature engineering...$159k - $230k
Drive research strategy by independently identifying, prioritizing,... ...Conduct tactical and strategic evaluations to streamline user journeys,... ...and translate complex behavioral data into frameworks.Establish... ...operates, from accelerating AI/ML deployments to reducing incident...$216k - $270k
Scale Labs, Research Scientist — AI Controls and MonitoringAs... ...the leading data and evaluation partner for frontier... ...methods that track AI behavior in real time to... ...detected;Design red-team simulations to probe weaknesses in... ...sophisticated ML problems, whether in a...Full time$262.5k - $299.6k
...questions in real time, our applications of AI & ML are bringing humanity and simplicity to... .... Our work touches every aspect of the research life cycle, from partnering with academia... ..., from design through training, evaluation, validation, and implementation. Engage in...Full timePart timeLocal area$225k - $325k
...ML Researcher | Healthcare AI | New York | $225,000–$325,000 + Equity One of the most exciting AI opportunities I've worked on this year... ...themselves, they're creating the reinforcement learning environments, evaluation frameworks and verification systems that help frontier AI...Work at officeVisa sponsorship$218.7k - $249.6k
...Applied Researcher I Overview: At Capital One, we are creating trustworthy... ..., our applications of AI & ML are bringing humanity and... ...from design through training, evaluation, validation, and implementation... ...EMNLP, Neurips, ICML or ICLR Behavioral Models PhD focus on topics in...Full timePart timeLocal areaFlexible hours$174k - $240k
...construct and own an AI-forward research roadmap for key product... ...Research: Lead foundational and evaluative research on autonomous AI agents... ...rigor (large-scale surveys, behavioral telemetry, statistical... ...functional partners (Product, Design, ML Engineering, Data Science)....Local areaWorldwideFlexible hours$15k
...apply state-of-the-art AI/ML techniques to construct our... ...creative reinforcement learning researcher to our growing ML research... ...optimization. The behavior of financial markets is noisy... ...conduct experiments to improve simulations and evaluate the success of new models in...Local areaImmediate startRelocationWork visa- ...join a team of scientists, ML researchers, and engineers working together... ...models for molecular simulation that can make the chemical... ...what we train on, what model behaviors matter, and which applications... ...OpenFE or related evaluation efforts. Experience across...
- ...organization prioritizes research in areas poised for... ...and AI for scientific simulation. Science of AI - understanding... ..., training, and evaluating state‑of‑the‑art... ...investments. Strong fluency in ML engineering tools and... ...systems and emergent behavior, AI‑accelerated...Local area
$275k - $300k
...the frontier of applying AI/ML to investment management. We... ...intelligence and machine learning research as well as highly... ...tackling hard problems. The behavior of financial markets is noisy... ...conduct experiments to improve simulations and evaluate the success of new models in...Local areaImmediate startRelocationWork visa- ...datasets used to train, evaluate, and deploy robotic... ...sits directly between research, data, and real-world... ...Research Scientist, RL & Simulation to own the RL +... ...imitation learning + RL: Behavior Cloning, DAgger-style... ...experience) in robotics, ML, or a related field....
- HUG in New York City is seeking an ML Researcher to join the founding team and own contributions in healthcare AI. You will build reinforcement learning environments, verifiers, and evaluation frameworks for real-world clinical workflows, while shaping the product and technical...
- ...Applied AI/ML Researcher As an Applied AI/ML Researcher within the AI4Tech Team at JPMorgan Chase, you will lead technology research... ...observability, and operational resilience. - Design, prototype, and evaluate AI-driven tools and frameworks that streamline engineering...
$197.3k - $313.7k
...Engineering Join a collaborative, diverse team of researchers at Agentforce Operations. The Foundational... ..., or a highly quantitative field with an AI/ML research focus. You possess experience developing, deploying, and evaluating machine learning models in production or top...Immediate start- ...that autonomously perform vulnerability research against real targets: firmware, network stacks... ...tool interfaces for agents. + Build the evaluation and benchmarking infrastructure that... ...across systems programming (C/C++, Rust) and ML infrastructure. + Comfort with binary...
$197.3k - $313.7k
...intelligence and bridges the gap between cutting‑edge research and customer value. The team is part of... ...Mathematics, or a highly quantitative field with an AI/ML research focus. Experience developing, deploying, and evaluating machine learning models in production or top‑...Immediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Evaluation Researcher: Behavior, ML & Simulation. Be the first to apply!
- senior design researcher New York, NY
- trend researcher New York, NY
- vulnerability researcher New York, NY
- researcher New York, NY
- music researcher New York, NY
- legal researcher New York, NY
- remote researcher New York, NY
- lead researcher New York, NY
- title researcher New York, NY
- product researcher New York, NY

