Data Scientist for AI Model Evaluation
$100 - $150 per hourSaidGig
Role Overview
Join a talent network of experienced data scientists who may contribute to future projects evaluating how effectively AI systems perform real-world data science work for leading AI research organizations.
Key Responsibilities
- Create precise, task-specific grading criteria for data science deliverables, including exploratory analyses, statistical modeling, machine learning pipelines, experimentation and A/B test write-ups, feature engineering, and technical reports or notebooks.
- Evaluate AI-generated or human-created work using established criteria.
- Provide detailed written explanations supporting evaluations and scores.
- Apply consistent, evidence-based judgment to produce reproducible, defensible assessments.
- Incorporate structured feedback from senior reviewers and revise submitted work accordingly.
Responsibilities will vary by project.
Qualifications
- At least 1 year of professional data science experience.
- Experience at a leading technology, research, or quantitative firm, such as a top FAANG company, AI lab, top-tier quantitative fund, or equivalent organization.
- Strong Python and SQL skills, plus expertise in statistical modeling, machine learning, experimentation, causal inference, and turning messy real-world data into rigorous analyses.
- Exceptional written communication skills and the ability to explain technical findings clearly.
- A detail-oriented, consistent approach to evaluating complex work.
- Comfort receiving feedback and calibrating judgment to established standards.
Work Terms
- Remote, hourly engagement.
- There is no immediate project opening. Qualified applicants may be contacted as relevant opportunities become available.
Compensation
$100 to $150 per hour.
Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Data Scientist for AI Model Evaluation in United States vacancy
$400 per month
...Mercor is partnering with a leading AI research lab to support a Frontier... ...Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments... .... The work focuses on realistic data engineering workflows and model...Suggested- Obsidian is partnering with an AI research initiative to support a Frontier Code Agents project in San Francisco. You will evaluate frontier AI coding models by performing realistic data engineering tasks, reviewing ETL pipelines, data warehouses, and distributed systems...Suggested
- Mercor partners with a leading AI research lab to support a Frontier Code Agents project, focusing on realistic data engineering workflows and model evaluation. Contributors help evaluate and improve frontier AI coding models through structured technical assessments and...Suggested
- SME Careers is looking for a remote R Engineer to review AI-generated responses and create high-quality R and data-analysis content. The position requires a strong background in R programming, applied statistics, and excellent writing skills to document analyses. Candidates...SuggestedRemote job
- ...for providing independent assurance and evaluating the company's risk management,... ...actions until completion.We are looking for data scientists and AI developers who will power our mission... ...Proficiency in frameworks for auditing models, including criteria like robustness, fairness...Suggested
$60 - $90 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation...Full timeContract workSummer workRemote work$184k - $287.5k
...is redefining what is possible with AI, and the Relational Foundation Model team is helping lead that transformation... ...models: you will design, build, and evaluate novel Transformer and graph neural... ...that generalize across diverse data schemas. Your work will power meaningful...Full time$102.5k - $170.9k
...performance.Work you'll doAs a Data Science Analytics and... ...validating quantitative workforce models that translate... ...Collaborating with Senior Data Scientists and AI engineers to integrate scenario... ...feature engineering, model evaluation, and performance tuning techniquesExperience...Civilian ContractorFull timeLocal area$100 per hour
...Apply data science expertise to help improve next-generation AI systems through rigorous content review, prompt development, model evaluation, and data-quality work. This remote, part-time contract role is designed for specialists who produce or review research papers,...Hourly payContract workPart timeFor contractorsRemote work$123.3k - $140.7k
...Overview Senior Associate, Data Scientist - Model Risk Audit Data is at the center of everything... ...latest emerging NLP and Generative AI technologies. If you love a fast-paced... ..., from design through training, evaluation, validation, and implementation Flex...Full timePart timeLocal areaFlexible hours$147.5k - $211k
...Business Unit Technology, Data, AI and Ventures (TDAV) Within the Tech,... ...The Corporate Vice President, Data Scientist - Model Validation and AI Governance will play... ...independently challenging model methodologies, evaluation approaches, controls, and monitoring strategies...Local area3 days per week$110k - $220k
...summary: Walmart’s Supply Chain AI Lab & Innovation Factory builds... ...This is not an analytics-focused data science role; it is a deeply... ...and continuously improve through evaluation and model post-training. As a Principal Data Scientist, you will turn ambiguous problems...Full timeTemporary workPart time$207k - $300k
...goals.Drive the adoption of next-generation AI/ML techniques to solve ambiguous, high-... ...with researchers to operationalize models and work with product stakeholders to align... ...(e.g., model deployment, model evaluation, data processing, debugging, fine tuning).Experience...Immediate start$130k - $260k
...Opportunity: Walmart’s Supply Chain AI Lab & Innovation Factory is... ...improve through rigorous evaluation and model post-training. This role... ...is not an analytics-focused data science role. It is a deeply... ...with supply chain experts, scientists, data owners, operations, and...Full timeContract workTemporary workPart time$184.7k - $324.8k
...United States Machine Learning and AI Do you get excited by driving... ...impact via measurement and evaluation, for products and services... ...organization, the mission of Data Science and Insights team is to... ...topics) to everyone from data scientists to engineers to business partners...Work experience placementRelocation- Data Scientist and Model Developer, Vice President (Contract) A temporary role (till December 2027) for... ...experience in AML model development, AI/ML and GenAI. Location Kuala Lumpur,... ...including development, ongoing performance evaluation and annual model reviews. Translates...Contract workTemporary work
$150k - $175k
...Applied Data Scientist, Health AI Evaluation & Datasets Innodata is a global data engineering company. We believe that data and Artificial Intelligence... ...anything real. Innodata partners with foundation model labs, medical AI startups, payers, providers, pharma, and...Remote workShift work- ...Aerospace VTI Aerospace builds AI-powered perception and pilot... ...a Senior Software Engineer – Evaluation, you will design and implement... ...(ASR), and small language model (SLM) systems. You will develop... ...closely with machine learning and data engineering teams to evaluate...
- ...Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting... ...and collaborate with other SMEs to ensure accurate training data. #J-18808-Ljbffr MercorWeekday work
- Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should...Weekday work
- ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote workFlexible hours
- ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities...Hourly payRemote workFlexible hours
- Role Description As a Data Scientist on our team, you will help define how we measure and trust... ...ll work at the intersection of applied AI evaluation, analytics engineering, and clinical... ...requirements, then research and build the models and reporting that meet them. This is a...Full time
- ...of the highest-stakes domains for generative AI. Numerical accuracy, regulatory compliance, model risk management, auditability, and customer... ...for financial workflows. As an Applied Data Scientist, Financial AI Evaluation & Datasets , you own the design, measurement...Full timeShift work
$60 - $70 per hour
...Role Overview Shape evaluation tasks for AI systems used in enterprise data and analytics environments. This remote opportunity... ...reflecting the data scale, model complexity, and business-critical... ...5 years of experience as a data scientist, analytics leader, or machine learning...Hourly payRemote work- ...developing the metrics and evaluation frameworks that ensure... ...and simulated driving data, enabling data‑driven... ...seeking a skilled Data Scientist to join our team and play... ...learning (e.g., model evaluation, feature engineering... ...email ****@*****.***.ai. #J-18808-Ljbffr AvrideRemote workRelocation
- Arena Intelligence in San Francisco is seeking a Data Scientist to explore and reason about the data powering AI evaluations on real-world use cases. You will generate,... ...insights that help us understand how frontier models behave in practice. You’ll collaborate with ML...
$184.7k - $324.8k
...United States Machine Learning and AI Imagine what you could do... ...could accomplish. The AIML Evaluation org works with teams across Apple... ...the highest standards for data quality and user privacy compliance. Description As a Data Scientist on the AIML Evaluation team, you...Relocation$136.44k - $265.11k
...rebuilding biotech for the AI era.When a breakthrough... ...structured scientific data and AI are built into... ...for biotech R&D. Scientists use Benchling to design... ...and run AI agents and models directly in their workflows... ...ll build the datasets, evaluations, and systems that help...Work at officeLocal areaMonday to FridayShift work- ...is building the foundation for physical AI - a unified platform that combines high-... ...to build the infrastructure that powers data, models, and AI runtime across Dexmate's Physical... ...model infrastructure for training, evaluation, versioning, deployment, and serving. Build...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Data Scientist for AI Model Evaluation. Be the first to apply!
Related searches
- work from home data scientist United States
- energy data scientist United States
- data scientist United States
- python data scientist (contract) United States
- associate data scientist United States
- data scientist no experience United States
- senior data scientist United States
- entry level data scientist remote United States
- healthcare data scientist United States
- python data scientist United States



