Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Evaluation Scientist: Math PhD for Frontier Benchmarks

Mercor

Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering with leading AI labs. You will author original, executable research problems that frontier models cannot solve. Engagement lasts 6 weeks, part-time with 20+ hours per week, start date immediate. Required: PhD in mathematics or related field, depth in two subdomains, Python for scientific computing, and experience with GitHub and Docker. #J-18808-Ljbffr Mercor

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Evaluation Scientist: Math PhD for Frontier Benchmarks in San Francisco, CA vacancy
  • An innovative tech company in New York is seeking a Research Scientist focused on Frontier Risk Evaluations. The ideal candidate will contribute to designing and creating evaluation measures for assessing AI risks and will have strong experience in machine learning and... 
    Suggested

    Scale AI

    San Francisco, CA
    3 days ago
  • Snorkel AI in San Francisco is searching for a Research Scientist to lead the development of datasets and benchmarks for AI models. This customer-facing role involves working with academic...  ...field and a strong focus on AI/ML evaluation and dataset design. With robust... 
    Suggested

    Snorkel AI

    San Francisco, CA
    3 days ago
  • $216k - $270k

    Scale AI, Inc. is looking for a Research Scientist specializing in Frontier Risk Evaluations to develop measures for assessing risks of advanced AI systems. In this role, you will design testing harnesses, collaborate with agencies, and publish reports to inform policymakers... 
    Suggested

    Scale AI, Inc.

    San Francisco, CA
    3 days ago
  • OpenAI is seeking a researcher to advance frontier evaluations and environments for safe AGI/ASI. You will help design north star model environments and steer major training runs so that research outputs translate into real-world products. Collaborate with researchers,... 
    Suggested

    Neura Market

    San Francisco, CA
    5 days ago
  • Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original...  ...research problems that today's frontier models cannot solve. Domains -... 
    For phd
    Part time
    Immediate start

    Mercor

    San Francisco, CA
    2 days ago
  •  ...working systems… About P-1 AI At P-1 AI, we are...  ...exceptional AI Research Scientist to join our small team...  ...from data generation to evaluation to product integration....  ...consolidation. About you A PhD (or equivalent...  ...Robotics, Engineering, Math, or a related field. Have... 
    For phd
    Relocation package

    Namely

    San Francisco, CA
    1 day ago
  •  ...research to improve the security and privacy of frontier intelligence systems. Responsibilities...  .... Required qualifications include a PhD in a relevant field and experience publishing...  ...drive meaningful security improvements in AI systems. #J-18808-Ljbffr United States Digital... 
    For phd

    United States Digital Space LLC

    San Francisco, CA
    4 days ago
  • $50 per hour

    A leading AI research organization is seeking PhDs in Chemistry or related fields for a remote contract. The role...  ...researchers. Responsibilities include developing solutions, evaluating AI outputs, and refining benchmarks. The pay rate is $50+/hour, depending on expertise.... 
    For phd
    Remote job
    Contract work

    Turing

    San Francisco, CA
    4 days ago
  •  ...looking for a talented researcher to work on improving AI tutoring systems. You will shape experiments, evaluate tutoring quality, and turn pedagogical principles...  ...potential for continuation. Ideal candidates have a PhD or equivalent experience in CS, ML, or related fields... 
    For phd
    Contract work
    Summer work

    Heyaristotle

    San Francisco, CA
    4 days ago
  • A leading AI evaluation firm based in San Francisco seeks a Machine Learning Scientist to foster understanding of AI model performance. You'll engage in designing and analyzing...  ...across teams. Applicants should possess a PhD in a relevant field and hands-on experience with... 
    For phd

    Arena Intelligence, Inc.

    San Francisco, CA
    5 days ago
  •  ...applications, processes, and AI into a single, governed...  ...AI Research Scientist to join our growing team...  ...optimised RAG, tool‑use evaluation, and multi‑agent collaboration.Prototype and benchmark models; present findings...  ...Experience / Technical SkillsMS/PhD in Computer Science, ML... 
    For phd
    Remote work
    Flexible hours

    Workato

    San Francisco, CA
    4 days ago
  • Mercor is seeking PhD and Master's level scientists to author AI evaluation tasks for Sci Code, collaborating with leading AI labs on a new benchmark for scientific computing. You will craft original...  ...research problems that current frontier models cannot solve.... 
    For phd
    Part time
    Immediate start

    Mercor

    San Francisco, CA
    3 days ago
  • $50 per hour

    A leading AI research firm is seeking PhDs in Mathematics or related fields for a fully remote contract role. The successful candidate will design advanced math problems to test AI performance and evaluate outputs for accuracy. Strong mathematical reasoning, problem-solving... 
    For phd
    Remote job
    Contract work
    Flexible hours

    Turing

    San Francisco, CA
    4 days ago
  • $50 per hour

    A leading AI research accelerator is seeking remote PhD candidates in Mathematics or related fields to design math problems and evaluate AI performance. The role involves collaboration with researchers and offers flexible hours at a pay rate of $50+/hour. Ideal candidates... 
    For phd
    Remote job
    Hourly pay
    Flexible hours

    Turing

    San Francisco, CA
    3 days ago
  •  ...remote contract to help fine-tune large language models. You will design problems, test AI solutions, and collaborate on benchmarks with leading AI labs. Requirements include a PhD (pursuing or completed) in mathematics or related fields, strong mathematical reasoning,... 
    For phd
    Remote job
    Contract work
    Flexible hours

    turing

    San Francisco, CA
    4 days ago
  • $250k

     ...AfterQuery builds the training data and evaluation infrastructure that frontier AI labs use to make their models...  ...evaluations that go beyond static benchmarks. We are a small, early team (post Series...  ...’s research (but haven’t done a PhD) Major plus if they’ve worked for/interned... 
    For phd

    afterquery

    San Francisco, CA
    4 days ago
  •  ...learning from human and AI feedback, reward modeling, and the evaluation suites that tell us...  ...your own benchmarks Run rigorous experiments...  ...modeling) at a frontier‑model lab or equivalent...  ...NICE TO HAVE PhD in ML, statistics,...  ...Background in RL math (policy gradients,... 
    For phd

    MakerMaker.AI

    San Francisco, CA
    4 days ago
  • About: Frontier AI x Biology | Foundation Models | Therapeutic Discovery...  ...biologists and experimental scientists to develop the next...  ...performance through post-training, evaluation, alignment and fine-tuning techniques...  ...biology Qualifications PhD in Machine Learning, Computer... 
    For phd

    KenkoTech Futures

    San Francisco, CA
    2 days ago
  • $50 per hour

     ...projects with top AI labs while earning...  ...ChatGPT) using your math and analytical skills...  ...to build better benchmarks. Responsibilities...  ...rigorous logic. Evaluate AI outputs for accuracy...  ...undergraduate to PhD-level math topics....  ...research accelerator for frontier AI labs and a... 
    For phd
    Contract work
    Remote work
    Flexible hours

    Turing

    San Francisco, CA
    5 days ago
  •  ...workflow‑aware planning, cost‑optimised RAG, tool‑use evaluation, and multi‑agent collaboration. Prototype and benchmark models; present findings internally and...  ...Qualifications / Experience / Technical Skills MS/PhD in Computer Science, ML, or related field—or equivalent... 
    For phd

    Workato

    San Francisco, CA
    3 days ago
  •  ...Forward Deployed Research Team in San Francisco. This position focuses on producing quality training data for AI models. The candidate should possess an MS or PhD in a quantitative field, with expertise in fine-tuning large language models and a strong grasp of LLM... 
    For phd

    Labelbox

    San Francisco, CA
    3 days ago
  • Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. You will help evaluate frontier AI coding models through structured technical assessments and focus on realistic data engineering workflows. The role involves reviewing model-... 

    Mercor

    San Francisco, CA
    2 days ago
  •  ...Lila, we don't just use AI to analyze biology; we...  .... We are seeking an ML Scientist Co‑Op to contribute to...  ...to identify patterns, evaluate model outputs, and guide...  ...Currently enrolled as a PhD student in Computer Science...  ...is the most inspiring frontier for AI. Rather than... 
    For phd

    Lila Sciences

    San Francisco, CA
    1 day ago
  • A leading AI research organization in San Francisco is seeking a Research Engineer focused on pushing the boundaries of frontier models and AGI/ASI measurement. Candidates should have strong...  ...AI measurements and build evaluation environments that drive progress. #J... 

    OpenAI

    San Francisco, CA
    3 days ago
  • $150k - $200k

     ...Contra Labs is a human-centered AI lab focused on creative and...  ...datasets that power benchmarking, evaluation, and post-training for the world...  ...creatives, Contra Labs connects frontier AI labs with a global...  ...experience, with a Master’s, PhD, or equivalent industry research... 
    For phd

    ConTra

    San Francisco, CA
    1 day ago
  • $70 per hour

     ...technical talent with leading AI research labs. Headquartered in...  ..., our investors include Benchmark , General Catalyst , Peter...  .... Position: Mathematics PhD Coding Experts Type: Contract...  ...answers. Calibrate tasks against frontier models, ensuring tasks ship... 
    For phd
    Contract work
    Summer work
    Immediate start
    Remote work

    Mercor

    San Francisco, CA
    3 days ago
  • $210k - $270k

     ...help them hire. Senior AI Engineer Location...  ...for model accuracy, evaluation, inference, monitoring,...  ...evaluation frameworks, benchmarks, and monitoring systems...  ...systems Stay current on frontier AI research and...  ...Graduate degree (MS or PhD) in AI, NLP, ML, or related... 
    For phd
    Remote work
    Visa sponsorship

    Recruiting from Scratch

    San Francisco, CA
    4 days ago
  • $55 - $80 per hour

     ...technical talent with leading AI research labs. Headquartered in...  ..., our investors include Benchmark , General Catalyst , Peter...  ...difficult problems in your domain. Evaluate and improve AI model...  ...Qualifications Must-Have PhD or advanced degree (or equivalent... 
    For phd
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    17 days ago
  • $70 per hour

     ...technical talent with leading AI research labs. Headquartered in...  ..., our investors include Benchmark , General Catalyst , Peter...  ...Position: Material Science PhD Coding Experts Type: Contract...  ...tasks. Calibrate tasks against frontier models, ensuring tasks ship... 
    For phd
    Contract work
    Summer work
    Immediate start
    Remote work

    Mercor

    San Francisco, CA
    5 days ago
  •  ...seeking a Member of Technical Staff (Speech) to advance our speech benchmarks. You will build and extend evaluations and arenas for text-to-speech, speech-to-text, and voice cloning, collaborating with leading AI models as they launch and mature. The role offers deep... 

    Artificial Analysis, Inc.

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Evaluation Scientist: Math PhD for Frontier Benchmarks. Be the first to apply!