Research Scientist - Frontier Evaluations
$210kAfterQuery
About AfterQuery
AfterQuery is an applied research lab curating data solutions for foundation model development. We serve every frontier AI lab with the mission of delivering the best data to power the best models. In doing so, we can make expertise that once took a lifetime to build available to anyone who needs it.
Our customers are the ones building the foundation models themselves and our work sits directly in the loop of how those systems improve. This is a rare opportunity to join a company at a defining moment in AI. We are YC's fastest unicorn, valued at $3.2 billion. We're based in San Francisco and backed by leading investors including Altos Ventures, BoxGroup, and Y Combinator and angels from Google DeepMind, OpenAI, Anthropic, Meta Superintelligence Labs, and Microsoft AI.
Why Apply
Massive Opportunity: We are YC's fastest unicorn valued at $3.2 billion and we're not slowing down.
Founding Impact: You will own and architect core infrastructure systems that power our platform from the ground up.
Equity & Growth: Competitive salary and meaningful equity. As we scale, you’ll have the opportunity to shape the engineering organization and lead major technical initiatives.
Strong Team: Our founding team has experience from Citadel Securities, Meta, Google, Silver Lake, and Morgan Stanley — work alongside world-class engineers and researchers.
Overview
AfterQuery is hiring Research Scientists to design and publish rigorous evaluations for frontier AI systems. The role spans agentic, coding, and safety evaluations, as well as expert-domain evaluations involving applied AI in healthcare, STEM, finance, and related fields. You will own evaluation development end to end and collaborate across disciplines to turn important capability gaps into rigorous public research.
Responsibilities
Lead the end-to-end design, validation, launch, and continuous improvement of frontier AI benchmarks.
Partner with researchers and domain experts to develop evaluations around meaningful model failures, gaps in existing coverage, and high-priority domains.
Analyze model capabilities and failure modes using rigorous experimental design and statistical methods.
Build reproducible evaluation systems, including harnesses, graders, and benchmark infrastructure.
Collaborate with researchers to post-train models and measure the resulting performance gains.
Communicate results through benchmark reports, technical articles, and research papers.
Required Qualifications
Strong record of publishing benchmarks or research papers.
Clear technical communication and strong scientific writing skills.
Commitment to experimental rigor, including baselines, ablations, statistical validity, and contamination controls.
Ability to take an ambiguous evaluation question from initial scoping through a reproducible public release.
Depth in agentic, coding, and safety evaluations or applied machine learning in an expert domain.
Preferred Qualifications
PhD in a related technical field.
Research publications at leading conferences or peer-reviewed journals.
Interest in multidisciplinary research and the creativity to combine methods and insights from AI, engineering, science, and other expert domains.
Company Benefits (For Eligible Employees):
Health Insurance: Medical, Vision, Dental
401(k) with Employer Match
Daily Meals: Daily UberEats Stipend
Monthly Wellness Stipend
Commute Covered
We are an equal opportunity employer committed to providing a workplace free from discrimination and harassment. Employment decisions are made without regard to legally protected characteristics under applicable federal, state, or local law.
We comply with applicable pay transparency requirements and provide compensation ranges based on the position, qualifications, experience, and other relevant factors. Reasonable accommodations are available to qualified individuals with disabilities and for sincerely held religious beliefs, as required by law. This job description is intended to describe the general nature and level of work performed and is not an exhaustive list of all duties, responsibilities, qualifications, or working conditions associated with the position. We reserve the right to modify this job description as business needs change.
$216k - $270k
Scale Labs, Research Scientist — Frontier Risk EvaluationsAs the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs...SuggestedFull time$176k - $304k
...Research Scientist, Frontier Capabilities Cambridge, MA USA; San Francisco, CA USA Your impact at LILA We're building a talent-dense... ...ablations Experience working with large-scale training or evaluation pipelines Ability to define and pursue research...SuggestedFull timeWork at officeLocal areaFlexible hoursShift work$165.6k - $207k
...and accelerate progress in GenAI research. We are looking for Research Scientists and Research Engineers with expertise... ...(SFT, RLHF, reward modeling) and evaluation. This role is on the evaluation... ...methods that reveal where frontier models fail and why. You will collaborate...SuggestedFull time$380k
...adapt in the real world. Mitigating the frontier risks resulting from these capabilities... ...the RoleWe are seeking exceptional researchers who can push the frontier of safety mitigations... ...(and then continuously refine) the evaluations that enable us to assess the extent of...SuggestedWork at officeLocal areaFlexible hours- ...About idler idler is a frontier data research lab. We build the evals and environments that the... ...here: About the role As a Research Scientist at idler, you'll own measuring and improving... ...scaleable systems for ingesting & evaluating data we are considering buying...SuggestedWork at officeRelocation package
- ...Mercor is seeking computational scientists specializing in atomistic and surface modeling to support a frontier AI research lab building models for materials science and the physical... ...knowledge to generate, structure, and evaluate the scientific data these models learn...Remote work
$150k - $200k
...preference datasets that power benchmarking, evaluation, and post-training for the world's... ...independent creatives, Contra Labs connects frontier AI labs with a global network of top... ...role exists Contra Labs is expanding its research work with frontier AI labs across evaluation...- ...Cerebro invites applications for a Research Scientist to design novel benchmarks and evaluate frontier language models and agents. You will lead research, design experiments and collaborate with research engineers, foundation-model developers and domain experts to turn...
- ...AI is the first audio data research company. We bring an R&D approach... ...on our mission to push the frontier of audio AI. About our... ...this role As a Research Scientist at David AI you'll build cutting... ...gather useful training and evaluation datasets to improve the...Work at office
- ...Discovery Chai Discovery builds frontier AI models to design... ...team brought together leading researchers in this space and top silicon... ...role As an AI Research Scientist, you will conduct groundbreaking... ...Experience training and evaluating large models on protein, antibody...
$160k - $250k
...Overview Research Scientist - Mountain View, CA at Granica. This range is provided by Granica.... ...efficient data systems. By advancing the frontier of how data is represented, stored,... ...fast: prototype new model architectures, evaluate on live datasets, and publish results...Flexible hours- ...Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing. You will craft original, executable research problems that frontier models currently cannot solve. This 6-week, part-time role involves sourcing...Part time
- ...Cincinnatus LLC is recruiting researchers to design and author multi-step evaluation tasks for frontier AI benchmarks. The role emphasizes translating scientific method into practical tasks, with a focus on Python-based analysis, rigorous evaluation, and clear written...Full timePart timeRemote work
- ...Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering... ...AI labs. You will author original, executable research problems that frontier models cannot solve. Engagement lasts 6 weeks,...Part timeImmediate start
- ...Researcher Position at Hedra Hedra is building a world-class Physical AI research team... ...architectures, training objectives, and evaluation frameworks for VLMs, VLAs, and world... ...research into production Stay at the frontier of the field — synthesizing relevant literature...Work at office
$117.2k - $313.7k
...The ExperienceSalesforce AI Research is a global leader in Enterprise... ...we continue to advance the frontier of enterprise AI through... ...for entrepreneurial Research Scientists who want to build, ship, and... ...optimization, model distillation, evaluation, and time series modeling....Full timeWorldwide- ...the right moment. We are looking for a researcher who can turn that ambition into something... ...first phase, your focus will be agent evaluation and harness design , not model training... ...people You do not need to have trained frontier models yourself. You do need to...
- ...NumericalEarth . We are looking for a Research Scientist who can work within the Data Assimilation... ...and benchmarks that allow us to evaluate them against established baselines. Build... ...we want someone excited to be at that frontier. How we work We are a small team that...Remote work
$250k - $400k
...stealth AI start-up building frontier reasoning models for scientific... ...models need to generate, evaluate and refine hypotheses across... ...genuinely novel AI for Science research, combining frontier reasoning... ...Researcher or experienced Research Scientist. What matters most is hands-...$245k - $285k
...growing group of committed researchers, engineers, policy experts,... ...are looking for biological scientists to help build safety and oversight... ...opportunity to shape how frontier AI models handle dual‑use biological... ...and execute capability evaluations ("evals") to assess the...Full timeWork at officeVisa sponsorshipFlexible hoursShift work$160k - $220k
...year runway.About the RoleWe’re looking for an AI Research Scientist to advance the methodological frontier of AI in healthcare. This role is ideal for someone... ...may include novel architectures, new training or evaluation techniques, long-horizon research bets, peer-reviewed...Temporary workWork at officeMonday to FridayMonday to Thursday- ...AfterQuery AfterQuery is an applied research lab curating data solutions for foundation model development. We serve every frontier AI lab with the mission of delivering the... ...Strong familiarity with LLM training and evaluation methodologies. Ability to design lightweight...Local areaShift work
$234.3k - $349k
...work with AI. About the roleAI research at WRITER isn't just about... ...world. As an AI research scientist, you'll be at the center of... ...hypothesis through model training, evaluation, and production... ...representing WRITER at the frontier of the field and contributing...Full timeWork at officeLocal area$150k - $380k
...team in San Francisco building frontier AI and the biological... ...context we expected. We want a scientist who can distinguish those explanations... ...next. You'll own in-vivo research for our designed biological... ...outcomes, rather than evaluating efficacy in isolation....Full time- ...development. Our vast talent network trains frontier AI models in the same way teachers... ...committed team. You’ll work alongside researchers, operators, and AI companies at the forefront... ...a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at the core of...Full timeWork at officeRelocation package
$290.4k - $363k
...intersection of cutting-edge research, large-scale engineering,... ...deployment, partnering with leading frontier labs, enterprises, and... ...the foundational research, evaluation methodologies, and agent/RL... ...and deployed.As a Research Scientist Manager, you will lead a world...Full time$54 - $60 per hour
...lie in enterprise domains, behind closed doors. Our research team's goal is to push the frontier of "domain adaptation" - how can we develop LLMs and... ...domains. This may include: Adapting, improving, and evaluating a method from the literature. Designing an entirely...Hourly payInternshipWorldwide- We are Genmo, a research lab dedicated to building open, state-of-the-art models for video... ...:We are seeking an exceptional Research Scientist to join our team, focusing on developing... ...experiments to validate new ideas and evaluate model performanceCollaborate with cross-...Relocation
$350k
...Research Engineer / Scientist, AlignmentSan Francisco, CAAbout AnthropicAnthropic's mission is to create... ..., Fine-Tuning, and the Frontier Red Team. Our current topics of focus... ..., and coordination with third-party evaluators.Safeguards Research: Developing robust...Work at officeVisa sponsorshipFlexible hours- ...AI Research Scientist (Robot Learning) San Francisco AI & Software In office Full-time... ...(Robot Learning) you will drive frontier AI model development and data flywheel... ...collection, model training and model evaluation in the real world. As an early employee...Full timeWork at officeImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Scientist - Frontier Evaluations. Be the first to apply!
- regulatory scientist San Francisco, CA
- nlp research scientist San Francisco, CA
- scientist biology San Francisco, CA
- applied scientist San Francisco, CA
- health scientist San Francisco, CA
- safety scientist San Francisco, CA
- pharmaceutical scientist San Francisco, CA
- deep learning scientist San Francisco, CA
- cell culture scientist San Francisco, CA
- senior analytical scientist San Francisco, CA




