AI Evaluation Scientist Math PhD Frontier Model Benchmark
Obsidian
Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. The role emphasizes material sourcing, prompt design, and robust grading criteria against leading AI models. Engagement spans 6 weeks, part-time at 20+ hours per week, with immediate start. You will work with Python or R for scientific computing and use GitHub and Docker in a PR-driven workflow. #J-18808-Ljbffr Obsidian
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering with leading AI labs. You will... ..., executable research problems that frontier models cannot solve. Engagement lasts 6 weeks, part...For phdPart timeImmediate start
- ...the open platform for evaluating how AI models perform in the real... ...measure and advance the frontier of AI for real-world... ...of Machine Learning Scientist to help advance how we... ...beyond traditional benchmarks Analyze large-scale... ...and tools You’ll have PhD or equivalent research...For phdPermanent employmentWork at office
- Cerebro invites applications for a Research Scientist to design novel benchmarks and evaluate frontier language models and agents. You will lead research, design experiments... ...to publish and present findings that influence how leading AI #J-18808-Ljbffr CerebroSuggestedRelocation
- ...a talent-dense team in San Francisco with 5.5M+ users and a rapidly growing platform. You will define how frontier AI models are measured, design new benchmarks, run experiments, and publish analyses that become industry gold standards. We sponsor visas and relocation...SuggestedRelocationVisa sponsorship
- OpenAI is seeking a researcher to advance frontier evaluations and environments for safe AGI/ASI. You will help design north star model environments and steer major training runs so that research outputs translate into real-world products. Collaborate with researchers,...Suggested
- Scale Labs seeks a Research Scientist focused on Frontier Risk Evaluations to design evaluation measures, harnesses and datasets for measuring risks posed by frontier AI systems. You will build harnesses to test models, collaborate with government agencies to scope evaluations...
- Snorkel AI in San Francisco is searching for a Research Scientist to lead the development of datasets and benchmarks for AI models. This customer-facing role involves working with academic partners... ...field and a strong focus on AI/ML evaluation and dataset design. With robust...
$216k - $270k
Scale AI, Inc. is looking for a Research Scientist specializing in Frontier Risk Evaluations to develop measures for assessing risks of advanced AI systems. In this role, you will design testing harnesses, collaborate with agencies, and publish reports to inform policymakers...- An innovative tech company in New York is seeking a Research Scientist focused on Frontier Risk Evaluations. The ideal candidate will contribute to designing and creating evaluation measures for assessing AI risks and will have strong experience in machine learning and...
- Artificial Analysis is seeking a Member of Technical Staff to design frontier evaluations for language models and publish results used by AI labs and enterprises. You will build datasets, scoring systems, and evaluation infrastructure applicable across major models released...WorldwideFlexible hours
- ...'re partnering with a frontier AI research company on a... ...open-weight foundation models with a mission to make... ...advanced AI systems are evaluated, stress-tested, and... ...Develop dynamic safety benchmarks that evolve alongside... ...highly desirable MS, PhD, or equivalent practical...For phd
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original... ...research problems that today's frontier models cannot solve. Domains - depth...For phdPart timeImmediate start
- ...improve the security and privacy of frontier intelligence systems.... ...Responsibilities include developing threat models, identifying security threats,... ...qualifications include a PhD in a relevant field and... ...meaningful security improvements in AI systems. #J-18808-Ljbffr...For phd
- ...role! About P-1 AI: At P-1 AI, we... ...custom post-trained models (SFT and RLVR)... ...exceptional AI Research Scientist to join our small... ...data generation to evaluation to product integration... ...About you: A PhD (or equivalent... ...Robotics, Engineering, Math, or a related field...For phdRelocation package
$165k - $195k
Research Scientist - Frontier AI Evaluations Compensation: $165,000-$195,000 base salary... ...company developing rigorous benchmarks and evaluation infrastructure for frontier language models and agents. As AI systems... ...For A Master’s degree, PhD or equivalent research experience...For phdImmediate startRelocationRelocation package$50 per hour
A leading AI research organization is seeking PhDs in Chemistry or related fields for a remote contract. The role... ...researchers. Responsibilities include developing solutions, evaluating AI outputs, and refining benchmarks. The pay rate is $50+/hour, depending on expertise....For phdRemote jobContract work- ...researcher to work on improving AI tutoring systems. You will shape experiments, evaluate tutoring quality, and turn pedagogical... .... Ideal candidates have a PhD or equivalent experience in CS, ML... ...projects and building simulated student models. #J-18808-Ljbffr HeyaristotleFor phdContract workSummer work
- ...Member Of Technical Staff (Language Model Evaluations) Location: San Francisco (preferred),... ...Analysis is the leading independent AI benchmarking company. We support labs, engineers and... ...of AI, they are actively shaping the frontier. Our benchmarks and analysis are trusted...
$70 per hour
...technical talent with leading AI research labs.... ...our investors include Benchmark , General Catalyst ,... ...Position: Material Science PhD Coding Experts Type:... ...input to challenge AI models . Build grading... ...Calibrate tasks against frontier models, ensuring tasks...For phdContract workSummer workImmediate startRemote work- Obsidian is partnering with an AI research initiative to support a Frontier Code Agents project in San Francisco. You will evaluate frontier AI coding models by performing realistic data engineering tasks, reviewing ETL pipelines, data warehouses, and distributed systems...
$400 per month
About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic infrastructure engineering...- Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models by performing infrastructure engineering tasks and reviewing model-generated implementations on cloud platforms, Kubernetes, and...
$50 per hour
A leading AI research accelerator is seeking remote PhD candidates in Mathematics or related fields to design math problems and evaluate AI performance. The role involves collaboration with researchers and offers flexible hours at a pay rate of $50+/hour. Ideal candidates...For phdHourly payRemote workFlexible hours$85 per hour
...technical talent with leading AI research labs.... ...our investors include Benchmark , General Catalyst ,... ...Responsibilities Use frontier AI coding agents to complete and evaluate complex infrastructure engineering... ...tasks. Review model-generated implementations...Contract workSummer workRemote work- ...applications, processes, and AI into a single,... ...exceptional AI Research Scientist to join our growing... ...optimised RAG, tool‑use evaluation, and multi‑agent collaboration... ....Prototype and benchmark models; present findings... .../ Technical SkillsMS/PhD in Computer Science, ML...For phdRemote workFlexible hours
- A leading AI evaluation firm based in San Francisco seeks a Machine Learning Scientist to foster understanding of AI model performance. You'll engage in designing and analyzing comprehensive... ...teams. Applicants should possess a PhD in a relevant field and hands-on experience...For phd
- Mercor is seeking PhD and Master's level scientists to author AI evaluation tasks for Sci Code, collaborating with leading AI labs on a new benchmark for scientific computing. You will craft original... ...problems that current frontier models cannot solve. Responsibilities...For phdPart timeImmediate start
$50 per hour
A leading AI research firm is seeking PhDs in Mathematics or related fields for a fully remote contract role. The successful candidate will design advanced math problems to test AI performance and evaluate outputs for accuracy. Strong mathematical reasoning, problem-solving...For phdRemote jobContract workFlexible hours- ...lead our work on model post-training:... ...from human and AI feedback,... ...modeling, and the evaluation suites that tell... ...Goodharting your own benchmarks Run rigorous... ...modeling) at a frontier‑model lab or... ...NICE TO HAVE PhD in ML, statistics... ...in RL math (policy gradients...For phd
- About: Frontier AI x Biology | Foundation Models | Therapeutic Discovery Stage: Well-funded... ...biologists and experimental scientists to develop the next... ...performance through post-training, evaluation, alignment and fine-... ...biology Qualifications PhD in Machine Learning, Computer...For phd
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Scientist Math PhD Frontier Model Benchmark. Be the first to apply!


