Retail Professional for AI Model Evaluation
$60 - $80 per hourSaidGig
Apply deep retail expertise to help develop and evaluate advanced large language models. In this role, you will bring practical judgment in merchandising, category management, buying, planning, and retail operations to create rigorous training data and improve AI reasoning. Key Responsibilities
- Advise research and engineering teams on retail merchandising, category management, and operations knowledge gaps.
- Create challenging retail-focused tasks and accurate, well-reasoned solutions based on real-world retail practice.
- Assess AI model responses using structured rubrics and provide written feedback on correctness, judgment, and reasoning quality.
- Create and refine retail-specific evaluation guidelines and scoring rubrics.
- Partner with other subject-matter experts to maintain consistency and accuracy across training data.
- At least 8 years of dedicated professional retail experience in areas such as merchandising, category management, retail operations, buying, or planning.
- Experience at a recognized top-tier organization, such as Amazon, Walmart, Target, Nike, Costco, Home Depot, or an equivalent company.
- Hands-on experience evaluating LLM or AI model outputs against rubrics or structured scoring criteria is required. Describe this experience in your application.
- Demonstrated career progression, such as advancing from Category Manager to Senior Manager to Director of Merchandising.
- Strong verbal and written communication, problem-solving, and interpersonal skills.
- United States-based, hourly W-2 employment.
- Reliable weekday availability of at least 35 hours per week is required.
$60 to $80 per hour.
Equal Employment OpportunityEmployment decisions are made without discrimination based on race, religion, color, national origin, sex, including pregnancy, childbirth, reproductive health decisions, related medical conditions, sexual orientation, gender identity, gender expression, age, protected veteran status, disability, genetic information, political views or activity, or any other legally protected characteristic.
$100 per hour
...the performance of large language models on finance tasks. You will work with AI researchers to identify model... ...systems. Key Responsibilities Evaluate LLM performance in finance areas... ...Qualifications Minimum 2 years of professional experience in one or more of the...SuggestedHourly payContract workFor contractorsFreelanceRemote work10 hours per weekFlexible hours$60 - $80 per hour
...help develop advanced large language models by bringing real-world underwriting,... ...claims, and risk-assessment judgment to AI training and evaluation work. Key Responsibilities... ...Qualifications At least 8 years of dedicated professional insurance experience in underwriting,...SuggestedHourly payWeekday work$100 - $150 per hour
...Apply senior finance expertise to improve how frontier AI systems reason through real-world financial work. You... ...research team to define high-quality finance tasks, evaluate model performance, and translate professional judgment into rigorous standards. Role Overview...SuggestedHourly payFull timeLive inRelocationRelocation package$50 per hour
...the performance of large language models on real-world finance tasks by evaluating model outputs, designing assessment... ..., and working directly with AI researchers to shape training and... ...Qualifications Minimum 2 years of professional experience in one or more of the following...SuggestedHourly payContract workRemote work10 hours per weekFlexible hours- ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets...SuggestedFor contractorsRemote work
$60 - $80 per hour
...underwriting and claims judgment, and evaluate large language model outputs against structured rubrics to... ...and claims practice. Evaluate AI model outputs against structured rubrics... ...Qualifications Minimum 8 years of dedicated professional experience in insurance, such as...Hourly payWeekday work$100 - $150 per hour
...Overview Help shape how next-generation AI models perform real financial work by... ...appear plausible but would not survive professional scrutiny. Instruction specs and golden... ...: Design challenging finance tasks and evaluation sets, and collaborate with researchers...Hourly payFull timeLive inRelocationRelocation package$60 - $80 per hour
About OpenTrain OpenTrain AI is the hiring and contracting organization for this... ...language work 20+ hours per week About AI Model Evaluation AI training is the human side of... ...cutting-edge AI development Apply practical professional expertise to model improvement Support...Hourly payPart timeFor contractorsFlexible hours$60 - $80 per hour
...develop advanced large language models. In this role, you will bring... ..., and campaign judgment to AI training data, partnering... ...quality. Develop and improve evaluation guidelines and scoring rubrics... ...least 8 years of dedicated professional marketing experience, such as...Hourly payWeekday work- Prolific is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities...Remote jobHourly payFlexible hours
$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work... ...competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and...Remote jobHourly payWork from homeFlexible hours$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Remote jobHourly payWork from homeFlexible hours- Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote jobFlexible hours
$100 - $150 per hour
...legal judgment to help a leading AI research organization improve how advanced AI models reason about real-world legal... ...quality legal tasks, standards, and evaluations. This role is designed for a... ...that would not withstand professional legal scrutiny. Write detailed...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package$110 - $150 per hour
...Role Overview Help advance frontier AI models by bringing professional finance judgment to the evaluation, design, and improvement of financial knowledge-work tasks. You will work closely with AI research and program teams to define what high-quality financial reasoning...Hourly payFull timeLive inRelocationRelocation package$100 - $150 per hour
...will be considered for future projects evaluating how well AI systems perform real-world data... ...assess AI or human-produced analyses and models, document decisions in writing, and iterate... ...At least 1+ years of professional data science experience Experience...Hourly payImmediate startRemote work$36 - $72 per hour
...educational institutions. Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise. Role... ...evaluation frameworks and professional guidelines, with exposure limits, content...Hourly payFull timeMonday to FridayFlexible hours- A leading AI company is seeking a legal professional for a contractor role focused on evaluating AI model outputs in legal contexts. Candidates must hold a Juris Doctor (J.D.) and have more than 3 years of experience in law. The role involves reviewing complex legal hypotheticals...For contractors10 hours per week
$60 - $100 per hour
...Help advance frontier AI systems by bringing practicing insurance and actuarial judgment into the evaluation, training, and improvement of models used for real insurance work. You will work... ...and responses that would not meet professional standards. Write detailed instruction...Hourly payFull timeLive inRelocationRelocation package$60 - $90 per hour
...engineering judgment to improve how advanced AI models reason through real-world engineering... ...correct solutions, and create rigorous evaluations grounded in industry practice. Key... ...or lead engineering responsibility. A Professional Engineer license is preferred but not...Hourly payFull timeRemote work$110 per hour
...physician talent network supporting AI labs and companies with medical expertise... ...real-world clinical expertise to model development and evaluation. Key Responsibilities Train... ...frontier AI research. Qualifications Professional experience in clinical diagnosis and...Hourly payContract workRemote work$60 - $80 per hour
...for future contract opportunities with AI labs and companies. This is an open application... ...knowledge to real-world tasks, model evaluation, and domain-specific feedback. Key Responsibilities... ...AI research. Qualifications Professional experience in statistical analysis and...Hourly payContract workRemote work$45 per hour
...French Professional Voice Actor — AI Speech Evaluation is a remote evaluation track for reviewing french generalist evaluation prompts and responses against... ...cases, and write the kind of structured feedback the modeling team can use to retrain. Why this role matters AI...Remote jobFor contractors10 hours per week$65 - $105 per hour
...Help advance frontier AI models by bringing rigorous life sciences research judgment into task design, evaluation, and model improvement. You will work closely with an AI lab... ...Hands-on use of large language models in professional work and sound judgment in distinguishing...Hourly payFull timeFreelanceLive inRelocationRelocation package$90 - $175 per hour
...software quality assurance expertise to evaluate technical AI outputs and help improve how next-... ...standards and methods. Qualifications Professional experience as a QA Engineer, SDET,... ...annotation, labeling, RLHF, AI response or model evaluation, or rubric-based grading....Hourly payContract workRemote work$65 - $105 per hour
...engineering judgment to help frontier AI models reason more accurately about real-world... ...define high-quality engineering work, evaluate model performance, and turn expert practice... .... At least 4 years of substantive professional engineering experience building and shipping...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package$60 - $90 per hour
...rigorous, real-world materials science and engineering tasks that evaluate an AI model’s expert reasoning. You will create prompts, supporting... ...Create domain-authentic tasks based on your own professional experience. Develop the data room and objective scoring...Hourly payWork at officeRemote work$70 - $110 per hour
...Overview Shape how advanced AI systems reason about real clinical... ...with an AI research team to evaluate medical knowledge tasks,... ...benchmarks that measure meaningful model improvement. Key... ...or Chief Medical Officer. Professional, hands-on experience using large...Hourly payFull timeLive inRelocationRelocation package$150 per hour
...judgment to create forecasting and research data that helps evaluate and improve AI models on real financial analysis. You will assess, using only... ...Review and grade model-generated analyses against your professional standard. Qualifications Former lead sell-side...Hourly payFor contractors$60 - $100 per hour
...accounting and audit judgment to improve how advanced AI models handle real-world accounting and audit work. You will... ...defining high-quality tasks, correct solutions, and evaluation standards grounded in professional practice. Key Responsibilities Review accounting...Hourly payFull timeLive inRelocationRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Retail Professional for AI Model Evaluation. Be the first to apply!
- retail sales advisor United States
- retail accountant United States
- retail finance United States
- retail customer service United States
- retail boutique United States
- retail clothing United States
- retail store United States
- retail commission sales United States
- retail stocker United States
- premium retail services United States


