Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Data Scientist for AI Model Evaluation

$100 per hour

SaidGig

Role Overview

Apply deep domain expertise to train and evaluate next-generation AI systems by producing, refining, and validating high-quality, domain-informed data. This part-time contractor role focuses on improving model outputs through careful content review, prompt refinement, rubric-based evaluation, and clear written reporting. No prior AI employment is required, domain knowledge and rigorous analytical skills are the priority.

Key Responsibilities
  • Review, edit, and refine AI-generated content and data for accuracy, clarity, and domain relevance against project-specific rubrics.
  • Develop and optimize prompts to guide model behavior, using professional writing and technical documentation skills.
  • Perform rubric-based evaluations of AI model outputs, provide structured feedback, and recommend improvements.
  • Annotate data, fact-check outputs, and contribute to quality assurance to ensure adherence to analytic standards.
  • Conduct independent research to validate facts and improve data quality for model training and evaluation.
  • Interpret complex datasets or findings and summarize them into clear, actionable reports and technical summaries.
  • Collaborate asynchronously with project leads and fellow domain experts to share insights and best practices.
Qualifications
  • Required skills: critical thinking, analytical reasoning, attention to detail, quality assurance, written communication, technical documentation, prompt authoring and refinement, AI output evaluation, professional writing, report writing, business communication, problem solving, content review, data interpretation, logical reasoning, professional editing, independent research, data annotation, rubric-based evaluation, content evaluation, AI model evaluation, fact checking, and technical editing.
  • Preferred experience: 3+ years in Data Science, Machine Learning, Applied AI, Statistics, Quantitative Analytics, or Data Analytics.
  • Proven record of producing or reviewing research papers, analytical reports, experiment summaries, notebooks, or technical documentation.
  • Experience with data annotation, content review, or rubric-based evaluation is highly desirable.
  • Background in prompt engineering, AI output evaluation, fact checking, or RLHF is advantageous but not required.
  • Advanced degrees such as a Master’s, JD, MBA, or PhD are preferred; contributors from highly selective universities or leading technology companies are valued.
Work Terms
  • Engagement type: Independent contractor, part-time.
  • Location: Remote.
  • Work is project-based and focused on a customer initiative to advance AI technology; tasks are completed asynchronously and in collaboration with project leads and other experts.
Compensation
  • Pay range: $100 to $200 per hour.
Eligibility and Application
  • No prior AI employment required, domain expertise and strong analytical and writing skills are the primary criteria.
  • Candidates will be screened and vetted through the platform’s talent selection process, which uses an AI-driven recruiter to assess fit for specific projects.
  • Selection for assignments depends on project needs and demonstrated fit to the rubric and domain requirements.
Vacancy posted 13 days ago
Similar jobs that could be interesting for youBased on the Data Scientist for AI Model Evaluation in Remote vacancy
  • $100 per hour

     ...deep domain knowledge to help train and evaluate next-generation AI systems by reviewing, refining,...  ...ensuring accuracy, clarity, and relevance of model outputs through rubric-based...  ..., and refine AI-generated content and data outputs for accuracy, clarity, and domain... 
    Suggested
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    Canada
    13 days ago
  • $80 per hour

     ...Role Overview Evaluate and improve frontier AI coding agents by applying professional data engineering judgment to realistic data infrastructure and pipeline scenarios. You will use and assess model-generated implementations for ETL, data warehouses, analytics platforms... 
    Suggested
    Hourly pay
    Remote work

    SaidGig

    United States
    20 days ago
  • $20 per hour

     ...DataAnnotation is committed to creating quality AI. Join our team to help train AI chatbots while...  ...chatbots. You will develop complex prompts to test AI models, write high-quality responses to demonstrate excellence, and evaluate different model outputs based on accuracy and... 
    Suggested
    Hourly pay
    Full time
    Contract work
    Part time
    For contractors
    Self employment
    Freelance
    Remote work

    DataAnnotation

    Wyoming, OH
    1 day ago
  • Alignerr is seeking a Quantitative Analyst to evaluate and improve AI-generated mathematical outputs for finance applications...  ...about risk and forecasting. You will analyze models for validity, assess performance, and validate data pipelines, communicating findings clearly to... 
    Suggested
    Remote job
    Hourly pay
    Contract work
    Flexible hours

    Alignerr

    Seattle, WA
    19 hours ago
  • $60 - $90 per hour

     ...Help shape rigorous evaluation benchmarks for frontier AI models by turning real-world analytical work into challenging, reproducible tasks. You will partner...  ...to identify where models succeed or fall short in data cleaning, statistical analysis, interpretation, and reporting... 
    Suggested
    Hourly pay
    Full time
    Remote work

    SaidGig

    United States
    24 days ago
  • $238k - $302k

     ...billions in simulation across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in Large Language...  ...inform model development and deployment.  Build data pipelines for signal discovery, data labeling, feature... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    19 hours ago
  • $80 per hour

     ...Role Overview Lead hands-on evaluations of frontier AI coding agents by applying them to realistic data engineering workflows, assessing model-produced ETL, data warehouse, analytics, and distributed system implementations, and surfacing bugs, scalability limits, and... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    1 day ago
  • $80 per hour

     ...and technical talent with leading AI research labs. Headquartered in...  ...and Jack Dorsey . Position: Data Engineer (Coding Agent Experience...  ...AI coding agents to complete and evaluate complex data engineering tasks. Review model-generated implementations... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    6 days ago
  •  ...processes, maximizing our use of technology, integrating data analytics into everything we do, and investing in our...  ...to learn more about you!Support Internal Audit’s evaluation of model and artificial intelligence (AI) risk and governance frameworks and their ability to... 
    Internship
    Monday to Friday

    Navy Federal Credit Union

    Vienna, VA
    4 days ago
  • $60 - $90 per hour

     ...Role Overview Design and execute realistic, research-grade data analysis tasks that serve as ground-truth references for evaluation of frontier generative AI models. You will create one-to-two day, end-to-end analysis challenges that include data cleaning, statistical... 
    Hourly pay
    Full time
    Part time
    Work experience placement
    Freelance
    Remote work

    SaidGig

    United States
    1 day ago
  • $60 - $90 per hour

     ...Author and run rigorous, multi-step machine learning evaluation tasks for a leading generative AI research team. You will take high-level research ideas...  ...experiments, and analyze results to determine where frontier models succeed or fail. Typical tasks require one to two days... 
    Hourly pay
    Full time
    Freelance
    Remote work

    SaidGig

    United States
    16 days ago
  • $120 - $170 per hour

     ...Role Overview Define what excellent, real-world data science work looks like for an AI research project that evaluates how well AI systems perform data science tasks. Instead of producing analyses or models, you will create task-specific grading rubrics and evaluate... 
    Hourly pay
    Remote work

    SaidGig

    United States
    a month ago
  • $150k - $175k

     ...Applied Data Scientist, Health AI Evaluation & Datasets Remote - United States Innodata is a global data engineering company. We believe that...  ...shipping anything real. Innodata partners with foundation model labs, medical AI startups, payers, providers, pharma, and... 
    Remote work
    Shift work

    Innodata Inc.

    United States
    3 days ago
  •  ...Overview Design and write evaluation tasks and reference...  ...Fortune 500 enterprise data science and analytics decision...  ...scenarios, produce model outputs and...  ...including SageMaker, Vertex AI, and MLflow. Apply enterprise...  ...experience as a data scientist, analytics leader, or... 
    Hourly pay
    Remote work

    SaidGig

    United States
    11 days ago
  • $100 - $150 per hour

     ...technical talent with leading AI research labs. Headquartered...  ...Jack Dorsey . Position: Data Scientist Talent Network Type: Contract...  ...analyses , statistical modeling work , machine learning...  ...and A/B test write-ups . Evaluate AI-generated or human-created... 
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    4 days ago
  • $100 - $120 per hour

    A leading AI research lab is seeking a part-time Data Scientist specializing in AI task evaluation and statistical analysis. You will conduct analysis on AI agent performance in finance, transforming data into actionable insights. The role offers $100-$120 per hour and... 
    Hourly pay
    Part time
    Remote work
    Flexible hours

    Call For Referral

    San Francisco, CA
    2 days ago
  •  ...of the highest-stakes domains for generative AI. Numerical accuracy, regulatory compliance, model risk management, auditability, and customer...  ...for financial workflows. As an Applied Data Scientist, Financial AI Evaluation & Datasets , you own the design, measurement... 
    Full time
    Shift work

    Innodata

    Remote
    25 days ago
  •  ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities... 
    Hourly pay
    Remote work
    Flexible hours

    Prolific

    Jacksonville, FL
    7 hours ago
  • $40 per hour

    A biotechnology company is seeking a Biotechnology R&D Scientist to train AI models and evaluate their outputs. This role involves measuring the progress of AI chatbots with complex biology questions and ensuring their performance and correctness. Candidates should have... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Wisconsin
    1 day ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    1 day ago
  • $148k - $184k

    Role Description As a Data Scientist on our team, you will help define how we measure and trust...  ...ll work at the intersection of applied AI evaluation, analytics engineering, and clinical...  ...requirements, then research and build the models and reporting that meet them. This is a... 
    Full time

    Paradigm Health

    Remote
    3 days ago
  • $150k - $175k

    Role Description As an Applied Data Scientist, Financial AI Evaluation & Datasets, you own the design, measurement quality, and domain validity of the...  ...evaluate, and monitor financial-domain LLMs, vision-language models, multimodal document models, and AI agents. ~... 
    Full time

    Innodata Inc.

    Remote
    6 days ago
  • $40 per hour

    A forward-thinking AI development firm seeks experienced quantitative professionals to evaluate AI-generated work, applying their skills in statistical analysis, predictive modeling, and technical writing. This fully remote opportunity offers a flexible schedule and projects... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Sioux Falls, SD
    1 day ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting... 
    Hourly pay
    Remote work

    DataAnnotation

    Juneau, AK
    1 day ago
  • $40 per hour

     ...A leading AI development firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully remote position allows for a flexible schedule, offering competitive pay starting at $40+ per hour. Ideal candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Bismarck, ND
    1 day ago
  • $40 per hour

    A leading AI training company is seeking a Biotechnology R&D Scientist to evaluate and improve AI models by posing complex biological questions. This remote position allows candidates to choose projects and work on their own schedule, with rates starting at $40+ per hour... 
    Hourly pay
    Remote work

    DataAnnotation

    Sioux Falls, SD
    4 days ago
  • $150 per hour

     ...Overview Aerospace engineering professionals apply their domain expertise to evaluate AI-generated outputs, assess technical content, and provide clear, structured feedback that improves models'' understanding of aerospace tasks, terminology, and practices. You will... 
    Hourly pay
    Temporary work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $85 per hour

     ...knowledge to design tasks, create domain-specific prompts, and evaluate large language models to improve their understanding and explanation of psychological concepts and research. This role supports year-round AI research projects that vary by domain and placement. Key... 
    Hourly pay
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $60 per hour

     ...contribute to developing cutting-edge AI systems, while enjoying the...  ...advance AI development. AI models are increasingly capable of...  ...the-art AI models on tasks like evaluating AI-generated quantitative...  ...how these systems reason about data, models, and scientific problems... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Charleston, WV
    1 day ago
  • $65 - $90 per hour

     ...Role Overview Apply your real-world architecture expertise to evaluate and improve how AI systems understand and reason about architecture. In this flexible, part-time, remote role you will review content for technical accuracy, answer domain-specific questions, and provide... 
    Hourly pay
    Part time
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Remote
    16 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Data Scientist for AI Model Evaluation. Be the first to apply!