Data Scientist for AI Model Evaluation
$100 per hourSaidGig
Apply deep domain expertise to train and evaluate next-generation AI systems by producing, refining, and validating high-quality, domain-informed data. This part-time contractor role focuses on improving model outputs through careful content review, prompt refinement, rubric-based evaluation, and clear written reporting. No prior AI employment is required, domain knowledge and rigorous analytical skills are the priority.
Key Responsibilities- Review, edit, and refine AI-generated content and data for accuracy, clarity, and domain relevance against project-specific rubrics.
- Develop and optimize prompts to guide model behavior, using professional writing and technical documentation skills.
- Perform rubric-based evaluations of AI model outputs, provide structured feedback, and recommend improvements.
- Annotate data, fact-check outputs, and contribute to quality assurance to ensure adherence to analytic standards.
- Conduct independent research to validate facts and improve data quality for model training and evaluation.
- Interpret complex datasets or findings and summarize them into clear, actionable reports and technical summaries.
- Collaborate asynchronously with project leads and fellow domain experts to share insights and best practices.
- Required skills: critical thinking, analytical reasoning, attention to detail, quality assurance, written communication, technical documentation, prompt authoring and refinement, AI output evaluation, professional writing, report writing, business communication, problem solving, content review, data interpretation, logical reasoning, professional editing, independent research, data annotation, rubric-based evaluation, content evaluation, AI model evaluation, fact checking, and technical editing.
- Preferred experience: 3+ years in Data Science, Machine Learning, Applied AI, Statistics, Quantitative Analytics, or Data Analytics.
- Proven record of producing or reviewing research papers, analytical reports, experiment summaries, notebooks, or technical documentation.
- Experience with data annotation, content review, or rubric-based evaluation is highly desirable.
- Background in prompt engineering, AI output evaluation, fact checking, or RLHF is advantageous but not required.
- Advanced degrees such as a Master’s, JD, MBA, or PhD are preferred; contributors from highly selective universities or leading technology companies are valued.
- Engagement type: Independent contractor, part-time.
- Location: Remote.
- Work is project-based and focused on a customer initiative to advance AI technology; tasks are completed asynchronously and in collaboration with project leads and other experts.
- Pay range: $100 to $200 per hour.
- No prior AI employment required, domain expertise and strong analytical and writing skills are the primary criteria.
- Candidates will be screened and vetted through the platform’s talent selection process, which uses an AI-driven recruiter to assess fit for specific projects.
- Selection for assignments depends on project needs and demonstrated fit to the rubric and domain requirements.
$100 per hour
...deep domain knowledge to help train and evaluate next-generation AI systems by reviewing, refining,... ...ensuring accuracy, clarity, and relevance of model outputs through rubric-based... ..., and refine AI-generated content and data outputs for accuracy, clarity, and domain...SuggestedHourly payContract workPart timeFor contractorsRemote work$80 per hour
...Role Overview Evaluate and improve frontier AI coding agents by applying professional data engineering judgment to realistic data infrastructure and pipeline scenarios. You will use and assess model-generated implementations for ETL, data warehouses, analytics platforms...SuggestedHourly payRemote work$20 per hour
...DataAnnotation is committed to creating quality AI. Join our team to help train AI chatbots while... ...chatbots. You will develop complex prompts to test AI models, write high-quality responses to demonstrate excellence, and evaluate different model outputs based on accuracy and...SuggestedHourly payFull timeContract workPart timeFor contractorsSelf employmentFreelanceRemote work- Alignerr is seeking a Quantitative Analyst to evaluate and improve AI-generated mathematical outputs for finance applications... ...about risk and forecasting. You will analyze models for validity, assess performance, and validate data pipelines, communicating findings clearly to...SuggestedRemote jobHourly payContract workFlexible hours
$60 - $90 per hour
...Help shape rigorous evaluation benchmarks for frontier AI models by turning real-world analytical work into challenging, reproducible tasks. You will partner... ...to identify where models succeed or fall short in data cleaning, statistical analysis, interpretation, and reporting...SuggestedHourly payFull timeRemote work$238k - $302k
...billions in simulation across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in Large Language... ...inform model development and deployment. Build data pipelines for signal discovery, data labeling, feature...Full timeRemote work$80 per hour
...Role Overview Lead hands-on evaluations of frontier AI coding agents by applying them to realistic data engineering workflows, assessing model-produced ETL, data warehouse, analytics, and distributed system implementations, and surfacing bugs, scalability limits, and...Hourly payContract workRemote work$80 per hour
...and technical talent with leading AI research labs. Headquartered in... ...and Jack Dorsey . Position: Data Engineer (Coding Agent Experience... ...AI coding agents to complete and evaluate complex data engineering tasks. Review model-generated implementations...Contract workSummer workRemote work- ...processes, maximizing our use of technology, integrating data analytics into everything we do, and investing in our... ...to learn more about you!Support Internal Audit’s evaluation of model and artificial intelligence (AI) risk and governance frameworks and their ability to...InternshipMonday to Friday
$60 - $90 per hour
...Role Overview Design and execute realistic, research-grade data analysis tasks that serve as ground-truth references for evaluation of frontier generative AI models. You will create one-to-two day, end-to-end analysis challenges that include data cleaning, statistical...Hourly payFull timePart timeWork experience placementFreelanceRemote work$60 - $90 per hour
...Author and run rigorous, multi-step machine learning evaluation tasks for a leading generative AI research team. You will take high-level research ideas... ...experiments, and analyze results to determine where frontier models succeed or fail. Typical tasks require one to two days...Hourly payFull timeFreelanceRemote work$120 - $170 per hour
...Role Overview Define what excellent, real-world data science work looks like for an AI research project that evaluates how well AI systems perform data science tasks. Instead of producing analyses or models, you will create task-specific grading rubrics and evaluate...Hourly payRemote work$150k - $175k
...Applied Data Scientist, Health AI Evaluation & Datasets Remote - United States Innodata is a global data engineering company. We believe that... ...shipping anything real. Innodata partners with foundation model labs, medical AI startups, payers, providers, pharma, and...Remote workShift work- ...Overview Design and write evaluation tasks and reference... ...Fortune 500 enterprise data science and analytics decision... ...scenarios, produce model outputs and... ...including SageMaker, Vertex AI, and MLflow. Apply enterprise... ...experience as a data scientist, analytics leader, or...Hourly payRemote work
$100 - $150 per hour
...technical talent with leading AI research labs. Headquartered... ...Jack Dorsey . Position: Data Scientist Talent Network Type: Contract... ...analyses , statistical modeling work , machine learning... ...and A/B test write-ups . Evaluate AI-generated or human-created...Contract workSummer workRemote work$100 - $120 per hour
A leading AI research lab is seeking a part-time Data Scientist specializing in AI task evaluation and statistical analysis. You will conduct analysis on AI agent performance in finance, transforming data into actionable insights. The role offers $100-$120 per hour and...Hourly payPart timeRemote workFlexible hours- ...of the highest-stakes domains for generative AI. Numerical accuracy, regulatory compliance, model risk management, auditability, and customer... ...for financial workflows. As an Applied Data Scientist, Financial AI Evaluation & Datasets , you own the design, measurement...Full timeShift work
- ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities...Hourly payRemote workFlexible hours
$40 per hour
A biotechnology company is seeking a Biotechnology R&D Scientist to train AI models and evaluate their outputs. This role involves measuring the progress of AI chatbots with complex biology questions and ensuring their performance and correctness. Candidates should have...Hourly payRemote workFlexible hours- ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote workFlexible hours
$148k - $184k
Role Description As a Data Scientist on our team, you will help define how we measure and trust... ...ll work at the intersection of applied AI evaluation, analytics engineering, and clinical... ...requirements, then research and build the models and reporting that meet them. This is a...Full time$150k - $175k
Role Description As an Applied Data Scientist, Financial AI Evaluation & Datasets, you own the design, measurement quality, and domain validity of the... ...evaluate, and monitor financial-domain LLMs, vision-language models, multimodal document models, and AI agents. ~...Full time$40 per hour
A forward-thinking AI development firm seeks experienced quantitative professionals to evaluate AI-generated work, applying their skills in statistical analysis, predictive modeling, and technical writing. This fully remote opportunity offers a flexible schedule and projects...Hourly payRemote workFlexible hours$40 per hour
A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting...Hourly payRemote work$40 per hour
...A leading AI development firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully remote position allows for a flexible schedule, offering competitive pay starting at $40+ per hour. Ideal candidates...Hourly payRemote workFlexible hours$40 per hour
A leading AI training company is seeking a Biotechnology R&D Scientist to evaluate and improve AI models by posing complex biological questions. This remote position allows candidates to choose projects and work on their own schedule, with rates starting at $40+ per hour...Hourly payRemote work$150 per hour
...Overview Aerospace engineering professionals apply their domain expertise to evaluate AI-generated outputs, assess technical content, and provide clear, structured feedback that improves models'' understanding of aerospace tasks, terminology, and practices. You will...Hourly payTemporary workPart timeRemote workFlexible hours$85 per hour
...knowledge to design tasks, create domain-specific prompts, and evaluate large language models to improve their understanding and explanation of psychological concepts and research. This role supports year-round AI research projects that vary by domain and placement. Key...Hourly payPart timeRemote workFlexible hours$60 per hour
...contribute to developing cutting-edge AI systems, while enjoying the... ...advance AI development. AI models are increasingly capable of... ...the-art AI models on tasks like evaluating AI-generated quantitative... ...how these systems reason about data, models, and scientific problems...Hourly payFull timeRemote workFlexible hours$65 - $90 per hour
...Role Overview Apply your real-world architecture expertise to evaluate and improve how AI systems understand and reason about architecture. In this flexible, part-time, remote role you will review content for technical accuracy, answer domain-specific questions, and provide...Hourly payPart timeRemote work10 hours per weekFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Data Scientist for AI Model Evaluation. Be the first to apply!
- data scientist (hedge fund) Remote
- healthcare data scientist Remote
- data scientist Remote
- genomic data scientist Remote
- entry level data scientist Remote
- part time data scientist Remote
- ai data scientist Remote
- senior data scientist Remote
- principal data scientist Remote
- entry level data scientist remote Remote






