Mathematics LLM Evaluation Expert
OpenTrain AI
About OpenTrain OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. It helps contributors discover specialized projects, build a lasting professional profile, and apply to opportunities in minutes. Creating an OpenTrain account is free. Build a portfolio of AI training and evaluation experience Find opportunities aligned with your subject-matter expertise Work remotely on projects shaping how modern AI systems perform About AI Training and LLM Evaluation Large language models learn and improve through human-created examples, detailed evaluations, and expert feedback. In this fast-growing field, specialists test model outputs, identify weaknesses, and provide the reasoning that helps AI become more accurate, useful, and reliable. Work at the intersection of mathematics and cutting-edge AI Contribute to evaluation benchmarks used to measure model capabilities Use human judgment to assess reasoning, explanations, and solutions The Role OpenTrain is recruiting a Mathematics LLM Evaluation Expert to create and solve challenging mathematical problems that test the capabilities and limitations of large language models. The work covers abstraction, multi-step reasoning, symbolic manipulation, and mathematics topics ranging from early undergraduate study through PhD-level curricula. This is a remote freelance contractor engagement for specialists based in the United States. The opportunity is listed as entry level, but it requires advanced mathematics knowledge appropriate to graduate or PhD-level study. Remote freelance contractor role Candidates must be based in the USA Available commitment options are 30 or 40 hours per week The listing indicates a commitment of 20+ hours per week Continuation may be possible based on performance and project needs What You'll Do You will create rigorous evaluation materials and provide detailed feedback to help assess and improve large language models. Your work will combine advanced mathematical problem solving with clear written explanations and careful analysis of reasoning quality. Design challenging mathematics problems that expose gaps in model reasoning Develop accurate, detailed, step-by-step solutions Explain complex mathematical concepts with accessible language, visuals, and examples Identify weaknesses in abstraction, multi-step reasoning, and symbolic manipulation Work with LLM researchers to align problem types with evaluation goals Contribute to benchmarks based on a broad mathematics curriculum Provide constructive feedback and detailed annotations that support model improvement Requirements and Helpful Background You should have a strong foundation in advanced mathematics and be able to analyze complex problems using a structured, logical approach. Clear, precise communication is essential because solutions, explanations, feedback, and annotations must be understandable and rigorous. Candidates pursuing or holding a Master's, Ph.D., or postdoctoral degree in Mathematics, Applied Mathematics, Statistics, or a related field are encouraged to apply. Experience developing mathematical explanations, evaluating reasoning quality, or providing detailed academic feedback can support success in this work. Advanced mathematics knowledge appropriate to graduate or PhD-level study Strong ability to solve complex problems through structured, logical reasoning Ability to explain mathematical concepts using simple language, visuals, and examples Strong English comprehension and structured written communication Research ability, analytical thinking, and creative problem-solving skills Ability to provide constructive feedback and detailed annotations Ability to work independently and collaborate remotely Reliable computer and internet connection Working Arrangement This role is remote and designed for US-based specialists working as freelance contractors. You must be available for at least four hours per day and four hours of overlap with Pacific Time. Location: United States Engagement: Freelance contractor Schedule options: 30 or 40 hours per week Daily availability: At least four hours Time-zone overlap: At least four hours with Pacific Time Language: English Build Your AI Training Career Mathematics experts are helping shape how AI systems reason, explain solutions, and handle difficult problems. Through OpenTrain, you can turn this work into credible experience, strengthen your professional profile, and grow a portfolio in AI training and data evaluation. Apply through OpenTrain in minutes Showcase specialized mathematics and evaluation experience Work remotely in a rapidly growing technology field #J-18808-Ljbffr OpenTrain AI
$80 - $160 per hour
...Training Work AI systems learn from examples prepared, reviewed, and evaluated by people. In this role, your legal and commercial knowledge... ...is seeking a legal and commercial contract evaluation expert to review domain-specific documents and assess their quality, accuracy...SuggestedHourly payContract workPart timeFor contractorsRemote workFlexible hours$70 - $90 per hour
OpenTrain AI is seeking a Trainium NKI Kernel Expert to evaluate the quality and correctness of NKI development tasks used to train advanced AI models. You will deliver rubric-based feedback on each task, focusing on CUDA-to-NKI migration fidelity and cross-platform numerical...SuggestedPart timeFor contractors$35 - $42 per hour
...join our growing network ofon-call AI Evaluation Specialistssupporting the development,... ...Partners' on-call network of subject matter experts. As client projects become available,... ..., Accounting, Economics, Business, Mathematics, Statistics, Law, Insurance, Risk Management...SuggestedHourly payExtra incomePermanent employmentContract workTemporary workRemote workFlexible hours- Mount Nittany Health in Altoona, PA seeks an audiologist to provide and interpret comprehensive evaluations, including audiometry, tympanometry, OAEs, acoustic reflex testing, and vestibular testing when indicated. The role supports hearing aid evaluation, fitting, and...SuggestedFull time
- Cambium Learning Group seeks an Early Childhood Mathematics Assessment Specialist to lead development, review, and QA of math assessment content for early learners. You will collaborate with content developers, psychometricians, and program teams to ensure alignment with...SuggestedRemote job
$100k - $140k
Job Reference #335608BRCityNew York, WeehawkenJob TypeFull Time Your roleAre you a sharp evaluator of risk? Can you investigate complex situations and propose solutions? We're looking for someone like that who can:• lead or conduct reviews and audits of specific business...Full timeFlexible hours- ...Geneva, Paladyne, Investran, HedgeTek or Oracle Flexcube.At least 4 years of experience in Development/ Configuration/solutions evaluation/ Validation and deployment.QualificationsBachelor’s degree or foreign equivalent required from an accredited institution. Will also...Permanent employmentFull timeH1b
$32.5 - $36 per hour
...Description Overview\n Intuit is seeking highly motivated individuals to join our dynamic team as dedicated year-round TurboTax Retail Experts in one of our TurboTax Retail or Flagship locations across the United States. This unique opportunity combines tax expertise,...Local area$40 - $80 per hour
...portfolio. About AI Training Work AI training is the human side of building modern artificial intelligence. Experts create examples, write reference responses, and evaluate model outputs so AI systems can produce clearer, more accurate, and more useful results. In this role,...Hourly payPart timeFor contractorsRemote workFlexible hours$55 - $70 per hour
...virtual, autonomous care for women in midlife across the United States. You will manage perimenopause and menopause, prescribe HRT, and evaluate pharmacologic weight management in a high-volume, remote setting. The role offers full-time, remote work with a schedule from 7:00...Remote jobHourly payFull timeLocal area$70 - $110 per hour
...the human side of building artificial intelligence. Clinical experts review model responses, identify unsafe or unsupported reasoning... ...systems while applying their clinical judgment to structured evaluation, writing, and benchmark design. The Role OpenTrain is seeking...Hourly payPart timeFor contractors- ...California seeks a senior Biochemistry and Life Sciences Domain Expert for a six-month engagement supporting frontier AI model... ...position requires a PhD and strong research background, with experience in evaluating scientific reasoning and #J-18808-Ljbffr OpenTrain AIContract workPart timeFor contractors
$100 - $150 per hour
...recruiting and contracting for this role, giving specialized experts a way to discover cutting-edge projects, build a professional... ...judgment will help translate tacit legal expertise into explicit evaluation criteria and training data for advanced generative AI systems....Hourly payFull timeContract workFor contractors- ...Practitioner or Physician Assistant (Procedural) for the Reconstructive Oncology Department. The role supports the OR suite, assists in evaluation, management, and treatment of Surgical Oncology disorders, and performs minor procedures in clinics. The ideal candidate is an...Relief
$40 - $65 per hour
...side of building modern artificial intelligence. Subject-matter experts review source materials and model responses, write or improve... ...Data type: Documents and text Work types: Text generation and evaluation or rating Employment type: Part-time contractor Pay: $40-$65...Hourly payPart timeFor contractorsFreelanceRemote work10 hours per weekFlexible hours- Job Title: GIS & Geospatial AI Expert Role Type: Contractor Location: Remote Job Overview We are seeking experienced GIS and Geospatial... ...domain knowledge, document advanced technical workflows, and evaluate AI-generated solutions in real-world geospatial and STEM environments...For contractorsRemote work
- ...(PA/NP) to manage patients in a 24-hour observation pathway. The APP collaborates with emergency medicine and inpatient teams to evaluate, treat, monitor, and disposition patients requiring extended observation following ED evaluation. The role emphasizes protocol-driven...
$35 - $70 per hour
...is the human side of building modern artificial intelligence. Experts review model outputs, create high-quality examples, and provide... ...an expert who can combine advertising judgment with AI model evaluation in a self-directed remote contractor workflow. Role type: Hourly...Remote jobHourly payPart timeFor contractors- Abbott Laboratories in Santa Clara, CA seeks a Senior Biocompatibility Scientist II to support biocompatibility evaluations across product lifecycle. You will develop evaluation plans, review data, and collaborate with cross-functional teams to ensure compliance with ISO...
$151.7k - $222.4k
...Description The Threat Detection & Response Engineer -- Senior Expert defines and builds the technical foundation of that rebuild. This... ...and feature/enrichment pipelines through model selection, evaluation, and safe production deployment. Set the standard for how AI/ML...Contract workWork at officeWork from homeWork visa$39.89 - $56.13 per hour
Pomona Valley Hospital Medical Center is seeking a Medical Coder to review and evaluate medical records to assign accurate diagnosis and procedure codes, ensuring optimal reimbursement while maintaining regulatory compliance. The role requires CCS, RHIT or RHIA credentials...Hourly pay- ...Professional Practice Model and state regulations. The RN will collaborate with the healthcare team to assess, plan, implement, and evaluate care plans for diverse patient populations. The role requires administering medications, documenting assessments, and developing...
- ...to deliver direct and indirect patient care using the nursing process. The role emphasizes assessment, planning, intervention and evaluation, developing care plans with patients, families, and a multidisciplinary team, and contributing to discharge planning for safe...
- ...Registered Nurse for the Pre Anesthesia / Pre Surgery Department in Seattle. The role involves performing complete pre-anesthesia evaluations, patient teaching, and coordinating discharge/transfer planning. Typical schedule includes days, evenings, and weekends with per...Daily paidAfternoon shift
- ...Anchorage to ensure accurate loan decisions and strong customer service. The role requires reviewing credit and property documents, evaluating risk, and ensuring compliance with investor and regulatory standards. The position offers a Monday-Friday schedule and...Monday to Friday
$122 per hour
...community. Creates safety standards, procedures and guidance that meet regulatory compliance and represent best practices. Recognizes, evaluates, recommends and implements PG&E Safety policies and procedures to assure awareness of and compliance with Safety requirements....Contract workFor contractorsWork at office- ...Platte, NE is seeking an RN to provide wound care across patient populations. The role involves assessing, planning, implementing, evaluating, and documenting nursing care per standards and policies. Minimum qualifications include graduation from an accredited nursing...
- Hocking Valley Community Hospital seeks a nursing leader to oversee total nursing care on the unit, directing functions and evaluating outcomes under general supervision. The role emphasizes patient safety, education, and adherence to standards of practice. Requirements...
$45 - $100 per hour
...reading charts — became the training signal that teaches frontier AI how to think in your field? TrainBrain builds expert-authored training data and evaluations for AI labs. We’re hiring practising professionals to write, solve, and grade the tasks advanced models are...For contractorsLocal area$73 per hour
...? Would you like to learn how modern AI systems are tested and evaluated? What the role looks like Analysts, researchers, experienced... ...intended to train the agent to succeed with less hand‑holding. Real expert complexity only. You’re improving the AI tools you’ll...Permanent employmentPart timeFreelanceRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Mathematics LLM Evaluation Expert. Be the first to apply!
- subject matter expert Brooklyn, NY
- fulfillment expert Brooklyn, NY
- guest service support expert Brooklyn, NY
- technology expert Brooklyn, NY
- math phd Brooklyn, NY
- math content creator Brooklyn, NY
- assistant professor mathematics Brooklyn, NY
- math adjunct Brooklyn, NY
- math interventionist Brooklyn, NY
- math degree Brooklyn, NY

