Mathematics PhD - AI Evaluation Expert
Obsidian
Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) Mercor is partnering with leading AI labs on a new benchmark for scientific computing. You will author original, executable research problems that today's frontier models cannot solve. Domains - depth required in at least two subdomains (with a coding focus) Mathematics — numerical linear algebra, computational mechanics, computational finance What you'll do Source your own material: a published paper, a Kaggle dataset, an open-source repository, or a scenario you design Write scientific prompts based on the input Build the grading criteria that define a correct answer Calibrate against frontier models — a task ships only when strong models fail it more often than they succeed Required PhD in mathematics, applied mathematics, computational mathematics, or a closely related field Demonstrated depth in at least two of the following subdomains: numerical linear algebra, computational mechanics, computational finance Working proficiency in Python or R for scientific computing Comfortable with Git/GitHub and running code in Docker — authoring runs through a pull-request workflow with automated quality checks Preferred Publications in peer-reviewed journals Prior scientific software or research engineering experience Engagement Duration: 6 weeks Commitment: part-time, 20+ hours per week Start date: immediate #J-18808-Ljbffr Obsidian
$70 per hour
...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Angelo , Larry Summers , and Jack Dorsey . Position: Mathematics PhD Coding Experts Type: Contract Compensation: $70/hour...For phdContract workSummer workImmediate startRemote work- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering with leading... ...hours per week, start date immediate. Required: PhD in mathematics or related field, depth in two subdomains, Python...For phdPart timeImmediate start
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code). You will author original, executable research problems that today's frontier models cannot solve. The role emphasizes material sourcing, prompt design, and robust grading criteria against...For phdPart timeImmediate start
$50 per hour
...Remote contract for PhDs in Mathematics, Statistics, or related... ...edge projects with top AI labs while earning $50+... ...rigorous logic. Evaluate AI outputs for accuracy... ...undergraduate to PhD-level math topics. Requirements... ...: Shortlisted experts complete an evaluation...For phdContract workRemote workFlexible hours$90 per hour
...technical talent with leading AI research labs. Headquartered in... ...with researchers and fellow experts to continually improve the benchmark... ...Must-Have ~ MSc or PhD in a STEM field or equivalent... ...-engineering, security, or AI-evaluation roles. ~ Ability to identify...For phdContract workSummer workRemote work$80 - $150 per hour
...are partnering with a leading AI research organisation to develop... ...-informed benchmark for evaluating how AI companion chatbots respond... ...generated conversations and provide expert judgment to calibrate... ...Required Qualifications MD, DO, PhD, PsyD, or equivalent qualification...For phdHourly payTraineeship10 hours per week- ...tasks. You will contribute to creating, evaluating, and refining AI-generated presentations across core... ...figures, and more. Candidates should hold a PhD with 3+ years of active research,... ...strong publication record, and demonstrate expert PowerPoint skills, with superb written...For phdRemote job
- Computational Statistics and Applied Mathematics Expert About the Project We're... ...to test how well advanced AI systems can solve hard scientific... ...-level expertise (MS or PhD required; PhD preferred, or MS... ...Familiarity with benchmark or evaluation design Background in scientific...For phdRemote work
$75 - $100 per hour
Join to apply for the Math PhD - Expert Trainer role at Handshake 1 week... ...in Math to join our AI research community. This program... ...with real-world expertise and evaluate where they excel or fail Get... ...by 2x Get notified about new Mathematics Specialist jobs in San Francisco...For phdContract workPart timeSummer workFreelanceH1bRemote workVisa sponsorship10 hours per weekFlexible hours$164.5k - $219k
...Information Job Title Expert Senior Manager, Data... ...to work with major AI ecosystem partners through... ...to prompt design, evaluation and output quality assuranceDevelop... ...to explain and discuss mathematical and machine learning... ..., or PhD in a technical fieldBackground...For phdPermanent employmentFull timeApprenticeshipLocal areaHome office3 days per week- YO AI Labs in the United States (Remote) seeks seasoned Adobe Marketing Technology Experts to support an AI training and evaluation project focused on enterprise marketing operations. You will test workflows with Adobe Workfront, AEM, CJA, Analytics and Experience Cloud...Remote job
$60 per hour
Prolific seeks Chemistry Experts and Chemical Engineers to join our Expert Network and evaluate AI models using your chemical expertise. Successful candidates will be invited to assess and review AI-generated chemistry tasks and ensure their accuracy. Compensation can reach...Remote jobHourly payFlexible hours- Mercor is seeking experienced musicians to evaluate generative music AI models. You will compare AI-generated lyrics with published songs across genres and rate them against detailed quality standards in Greek and English. The role requires native or near-native Greek,...Remote workFlexible hours
$80 - $150 per hour
...leading behavioral health consulting firm is seeking Senior Behavioral Health Experts to work part-time and remotely on frontier AI research projects. You will be responsible for designing evaluations and testing AI systems in critical mental health contexts. The ideal...Hourly payPart timeRemote work$50 per hour
A leading AI research accelerator is seeking remote PhD candidates in Mathematics or related fields to design math problems and evaluate AI performance. The role involves collaboration with researchers and offers flexible hours at a pay rate of $50+/hour. Ideal candidates...For phdHourly payRemote workFlexible hours$80 - $120 per hour
A leading AI consulting firm is searching for a Senior Mathematics Expert to join their remote team. The ideal candidate will possess a PhD in Mathematics and have a rich experience in solving complex... ...AI models can't solve and evaluating these models. This position offers...For phdRemote job- A leading AI evaluation firm based in San Francisco seeks a Machine Learning Scientist to foster understanding of AI model performance. You'... ...while collaborating across teams. Applicants should possess a PhD in a relevant field and hands-on experience with large-scale models...For phd
$50 per hour
A leading AI research firm is seeking PhDs in Mathematics or related fields for a fully remote contract role. The successful candidate will design advanced math problems to test AI performance and evaluate outputs for accuracy. Strong mathematical reasoning, problem-solving...For phdRemote jobContract workFlexible hours- ...looking for a talented researcher to work on improving AI tutoring systems. You will shape experiments, evaluate tutoring quality, and turn pedagogical principles... ...potential for continuation. Ideal candidates have a PhD or equivalent experience in CS, ML, or related fields...For phdContract workSummer work
- Obsidian is seeking expert Evaluators in FP&A / corporate finance to assess AI-generated work products for accuracy and quality. This role entails deep expertise to grade outputs and provide structured feedback. Candidates should have at least 5 years of relevant experience...Remote jobHourly payWork at office
- Obsidian is seeking experienced AI Safety Practitioners to evaluate safety, quality, and alignment of frontier AI models across complex policy-sensitive topics. You will assess AI-generated responses, apply safety policies, and help improve model behavior through structured...
- Obsidian is seeking expert Evaluators in Biology/environmental science to review and assess AI-generated work products for accuracy and quality. In this remote, hourly role, you will leverage your expertise to provide feedback on documents and presentations, ensuring they...Remote jobHourly pay
- Cincinnatus LLC is seeking a Marketing SME to join a leading GenAI team focused on evaluating AI outputs against rubrics and strengthening brand strategy in AI training data. We require 8+ years of marketing experience with top-tier brands, plus hands-on evaluation of LLM...Weekday work
$160k - $250k
...What Granica does Granica is an AI research and systems company building... ...new model architectures, evaluate on live datasets, and publish results... ...efficient learning. What you’ll bring PhD in Machine Learning, Statistics, Applied Mathematics, or a related field with...For phdFlexible hours- Mercor is seeking a remote Physics PhD to tackle complex physics problems and ensure the accuracy of AI-generated solutions. The role involves independently solving physics challenges and collaborating with AI research teams. Ideal candidates will have a PhD in Physics...For phdRemote jobImmediate start
$128.5k - $171.5k
A data-focused AI engineering role inside the Coro unit, delivering GenAI powered tools... ...relevant Create reproducible training and evaluation pipelines with versioning, experiment tracking... ...program Preferred Qualifications MBA or PhD in a technical field Background in...For phd$200k - $260k
...- two of our co-founders were PhD data scientists at Robinhood. Since... ...graduate studies in physics, mathematics, bioengineering, economics,... ...FluencyRapidly learn and adopt emerging AI tools and workflows to accelerate analysisCritically evaluate AI-generated outputs for...For phdFlexible hours$60 - $70 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...$60–$70/hour Location: Remote Role Responsibilities Evaluate AI-generated responses for safety, factual accuracy, policy compliance...Contract workSummer workRemote work$196k - $230k
...AreNotion is the collaborative AI workspace where teams and agents... ...Researcher to define and scale how we evaluate Notion’s AI-powered experiences—... ...(e.g., Python), or statistical/mathematical software (e.g., R, SAS, Matlab, etc.)Master’s or PhD in HCI, Psychology, Behavioral...For phdLocal areaShift work- ...Description Job Description Mathematics Expert Job Type: Contractor... ...Mathematics Experts to support an AI-focused project involving... ...You will create, review, and evaluate mathematical content with a strong... ...BSc, MSc, or PhD in Mathematics, Applied Mathematics...For phdRemote jobFor contractors
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Mathematics PhD - AI Evaluation Expert. Be the first to apply!
- subject matter expert San Francisco, CA
- fulfillment expert San Francisco, CA
- guest service support expert San Francisco, CA
- technology expert San Francisco, CA
- phd student San Francisco, CA
- phd chemistry San Francisco, CA
- math phd San Francisco, CA
- research phd San Francisco, CA
- phd intern San Francisco, CA
- phd internship San Francisco, CA



