Evaluation Specialist for AI Model Evaluation
$20 - $60 per hourSaidGig
Role Overview
Help train next-generation AI systems by creating rigorous, real-world evaluations that test how advanced models learn, reason, and perform. This remote contractor opportunity is open to recent graduates, advanced-degree holders, and professionals with strong research and writing capabilities. Prior AI experience is not required.
Key Responsibilities
- Develop original, challenging question-and-answer pairs across diverse subjects.
- Conduct in-depth research and triangulate multiple sources to produce accurate, comprehensive, well-documented answers.
- Create multi-step questions that require synthesis and analytical reasoning rather than simple single-source lookups.
- Test questions with AI models, analyze the results, and adjust difficulty as needed.
- Document research methods with clear citations and logical explanations.
- Revise content in response to reviewer feedback while following project guidelines and quality standards.
Qualifications
- Demonstrated ability to independently research and critically assess information from multiple sources.
- Excellent attention to detail, analytical thinking, and precise English writing. English fluency is required, though it need not be your first language.
- Ability to write nuanced, well-structured questions that assess deep understanding.
- Strong self-direction, reliability, and accountability in remote independent work.
- Interest in advanced AI and evaluating current model capabilities.
- Experience with AI training, evaluation, or content creation is helpful but not required.
Work Terms
- Remote contract engagement.
- Work is paid by qualifying task output, and completion time may vary based on experience and workflow.
- Minimum submission requirements apply, including a minimum number of tasks submitted each week.
- Roles are commonly filled within 48 hours. Selected candidates should be prepared to begin their first tasks within 24 to 48 hours after completing onboarding.
Compensation
Compensation is listed at $20 to $60 per hour, with payment determined on a per-task basis for work that meets project specifications.
Application Process
Submit an application using an email address or Google account. Candidates who advance complete onboarding before beginning project tasks. Questions can be reviewed through the available FAQs.
$20 - $60 per hour
...document expertise to a project that trains next-generation AI systems. You will design realistic Fortune 500 style scenarios and interact iteratively with an advanced language model to create, edit, and evaluate Office Open XML files, with a focus on .pptx deliverables....SuggestedHourly payContract workFor contractorsWork at officeRemote work$75 - $115 per hour
...Role Overview Help advance frontier AI models by bringing rigorous pharmaceutical research and development judgment to the evaluation, design, and improvement of domain-specific... ...Partner with researchers and adjacent-domain specialists to build pharmaceutical skills and tools...SuggestedHourly payFull timeContract workLive inRelocationRelocation package$70 - $80 per hour
...Role Overview Apply advanced drug safety expertise to help improve AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data. This remote contract role focuses on the quality, accuracy, and regulatory alignment of complex safety reports...SuggestedHourly payContract workRemote work$70 - $90 per hour
...Review GPU and accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical correctness, completeness, fair performance benchmarking, appropriate scope, and whether kernels...SuggestedHourly payRemote work$100 - $150 per hour
...senior legal judgment to help a leading AI research organization improve how advanced AI models reason about real-world legal work.... ...quality legal tasks, standards, and evaluations. This role is designed for a practicing legal specialist with deep subject-matter expertise,...SuggestedHourly payFull timeFreelanceInternshipLive inRelocationRelocation package- ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets...For contractorsRemote work
$70 - $90 per hour
...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness, CUDA-to-NKI migration fidelity, and whether implementations are well suited to...Hourly payRemote work$60 - $80 per hour
...expertise to help develop advanced large language models. In this role, you will bring practical brand, growth, and campaign judgment to AI training data, partnering with research... ...reasoning quality. Develop and improve evaluation guidelines and scoring rubrics for...Hourly payWeekday work- ...Title: AI Evaluation & Model Risk Lead Location: Bellevue WA Engineer, AI - AI Evaluation & Model Risk Lead Are you ready to join the Un-carrier movement? This role leads how T-Mobile decides which AI models it can trust - the behavioral and model-risk...Work experience placement
- ...Role Title: Research Engineer - Code Generation & Model Evaluation Role Type: Contractor Location: Remote micro1 is engaging Research... ...role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and...Temporary workFor contractorsRemote work
$50 - $100 per hour
...Apply your software engineering expertise to help train and evaluate next-generation AI systems through real-world coding tasks, technical feedback... ...contract role focuses on code generation workflows and model evaluation; prior AI experience is not required. Key Responsibilities...Hourly payContract workFor contractorsRemote work$136.44k - $265.11k
We are rebuilding biotech for the AI era.When a breakthrough is delayed, the world waits... ...structured data, and run AI agents and models directly in their workflows. Over 200,000... ...our work here.You’ll build the datasets, evaluations, and systems that help close that gap....Work at officeLocal areaMonday to FridayShift work- ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically... ...How this work supports customers Accelerate frontier AI research by contributing high quality data and evaluation...Contract workFor contractorsFreelanceRemote work
$60 - $90 per hour
...engineering judgment to improve how advanced AI models reason through real-world engineering... ...correct solutions, and create rigorous evaluations grounded in industry practice. Key... ...with researchers and adjacent-domain specialists to calibrate consistent standards and convert...Hourly payFull timeRemote work$60 - $80 per hour
...mathematics experts considered for future contract opportunities with AI labs and companies. This is an open application, not a posting... ...by applying mathematics knowledge to real-world tasks, model evaluation, and domain-specific feedback. Key Responsibilities...Hourly payContract workRemote work$36 - $72 per hour
...seekers, 1 million+ employers, and 1,600 educational institutions. Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise. Role Details Location: Onsite in Seattle, WA,...Hourly payFull timeMonday to FridayFlexible hours$60 - $80 per hour
...Overview Apply deep insurance expertise to help develop advanced large language models by bringing real-world underwriting, claims, and risk-assessment judgment to AI training and evaluation work. Key Responsibilities Partner with research and engineering teams to...Hourly payWeekday work$100 - $150 per hour
...data scientists who will be considered for future projects evaluating how well AI systems perform real-world data science tasks. Members of this... ...criteria, assess AI or human-produced analyses and models, document decisions in writing, and iterate on evaluations with...Hourly payImmediate startRemote work$110 per hour
...Apply to join a physician talent network supporting AI labs and companies with medical expertise. This is an open application... ...projects by bringing real-world clinical expertise to model development and evaluation. Key Responsibilities Train and evaluate AI models...Hourly payContract workRemote work- Zebra Technologies is seeking an AI Quality Analyst to ensure performance, safety, and reliability of cutting-edge AI/ML models. Design evaluation strategies, identify edge cases, bias sources, and provide actionable insights to drive model improvements across the development...
$60 - $90 per hour
...Role Overview Design rigorous, real-world materials science and engineering tasks that evaluate an AI model’s expert reasoning. You will create prompts, supporting data rooms, and grading criteria, then test and refine each task until it clearly distinguishes strong...Hourly payWork at officeRemote work$65 - $105 per hour
...engineering judgment to help frontier AI models reason more accurately about real-world... ...define high-quality engineering work, evaluate model performance, and turn expert practice... .... Collaborate with researchers and specialists in adjacent disciplines to calibrate...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package- ...Hallucination Failure Model Evaluation Specialist is a remote evaluation track for reviewing hallucination failure model evaluation evaluation prompts... ...team can use to retrain. Why this role matters AI data reviewers help turn hallucination failure model evaluation...Remote jobHourly payFor contractors10 hours per week
$70 - $110 per hour
...Overview Shape how advanced AI systems reason about real clinical... ...with an AI research team to evaluate medical knowledge tasks,... ...benchmarks that measure meaningful model improvement. Key Responsibilities... ...and adjacent-domain specialists to calibrate standards and translate...Hourly payFull timeLive inRelocationRelocation package$100 - $150 per hour
...finance expertise to improve how frontier AI systems reason through real-world... ...team to define high-quality finance tasks, evaluate model performance, and translate professional... ...Calibrate standards with researchers and specialists in adjacent fields, converting tacit financial...Hourly payFull timeLive inRelocationRelocation package- ...Apply your dermatology expertise to image-based work that helps develop and evaluate advanced AI models for clinical image interpretation. This is a non-clinical, remote opportunity with no direct patient care, focused on ensuring AI-generated medical outputs reflect real...Hourly payRemote work
$150 per hour
...Role Overview Apply sell-side equity research judgment to create forecasting and research data that helps evaluate and improve AI models on real financial analysis. You will assess, using only information available at a defined point in time, when a named analyst is...Hourly payFor contractors$60 - $90 per hour
...Role Overview Help shape how an advanced performance-transfer model evaluates character animation, preserving an actor''s timing, emotion,... ...evaluation methods, and help build a reliable human-review process for AI-generated performance results. Key Responsibilities...Hourly payPart time$60 - $80 per hour
...Role Overview Apply deep retail domain expertise to help build and evaluate foundational generative AI models. You will design realistic retail tasks, produce authoritative solutions grounded in real-world practice, and judge model outputs to improve model correctness...Hourly payWeekday work$150 per hour
...authored political forecasting and research data that helps evaluate and improve AI models performing real-world political and financial analysis.... ...campaign strategist, polling or statistical-methodology specialist, or political-risk analyst. A demonstrated record of making...Hourly payOdd job
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Evaluation Specialist for AI Model Evaluation. Be the first to apply!
- welding specialist United States
- revenue cycle specialist United States
- flow cytometry specialist United States
- transportation specialist United States
- imaging specialist United States
- medicaid eligibility specialist United States
- title specialist United States
- e learning specialist United States
- employment specialist United States
- deduction specialist United States


