Retail Professional for AI Model Evaluation
$60 - $80 per hourSaidGig
Apply deep retail expertise to help improve advanced AI models by creating, assessing, and refining retail-focused training materials. This role brings practical merchandising, category management, buying, planning, and retail operations judgment into AI evaluation work.
Key Responsibilities
- Advise research and engineering teams on retail merchandising, category management, and operations knowledge gaps.
- Create challenging, relevant retail tasks and accurate, well-reasoned solutions based on real-world retail practice.
- Assess AI model responses using structured rubrics and provide written feedback on accuracy, judgment, and reasoning quality.
- Develop and improve retail-specific evaluation guidelines and scoring rubrics.
- Work with other subject-matter experts to maintain consistent, accurate training data.
Qualifications
- At least 8 years of dedicated professional retail experience in areas such as merchandising, category management, retail operations, buying, or planning, gained at a recognized top-tier organization such as Amazon, Walmart, Target, Nike, Costco, Home Depot, or an equivalent company.
- Hands-on experience evaluating LLM or AI model outputs using rubrics or structured scoring criteria is required. Describe this experience in your application.
- Demonstrated career growth, such as progressing from Category Manager to Senior Manager to Director of Merchandising.
- Strong verbal and written communication, problem-solving, and interpersonal skills.
Work Terms
- United States, hourly W-2 employment.
- Reliable weekday availability of at least 35 hours per week is required.
Compensation
- $60 to $80 per hour.
Equal Opportunity
Employment decisions are made without discrimination based on race, religion, color, national origin, sex, pregnancy, childbirth, reproductive health decisions, related medical conditions, sexual orientation, gender identity or expression, age, protected veteran status, disability, genetic information, political views or activity, or any other legally protected characteristic.
$65 - $90 per hour
...deep financial judgment to help improve foundational AI models. In this role, you will create and evaluate finance-focused work that strengthens AI systems''... ...Qualifications At least 8 years of dedicated professional finance experience, such as investment banking,...SuggestedHourly payWeekday work$60 - $80 per hour
...building foundational large language models by applying deep insurance domain expertise to create, evaluate, and refine training data. You... ...claims practice. Evaluate AI model outputs against... ...Qualifications At least 8 years of professional experience in insurance, such...SuggestedHourly payWeekday work$60 - $100 per hour
...Role Overview Help advance frontier AI models by bringing senior insurance and actuarial judgment to the evaluation of real-world insurance work. You will work directly... ...and program management team, translating professional standards into tasks, solutions, benchmarks,...SuggestedHourly payFull timeLive inRelocationRelocation package$110 - $150 per hour
...Role Overview Drive how frontier AI models reason about real-world private equity and... ...that appear plausible but would not pass professional review. Author high-quality instruction... ...Design challenging finance tasks and evaluation sets, and help create finance-specific...SuggestedHourly payFull timeLive inRelocationRelocation package$10 - $14 per hour
...telemetry from PC titles to help train and evaluate next-generation AI systems. This contractor role focuses... ...QA findings from everyday play, so models learn from real player behavior and... ...a track record in competitive or professional play. Experience with gameplay recording...SuggestedHourly payFor contractorsRemote work$60 per hour
...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Hourly payRemote workWork from homeFlexible hours$60 - $80 per hour
...Role Overview Apply deep retail expertise to help develop foundational generative AI models. You will bring practical merchandising... ...judgment to the design and evaluation of AI training data, working... ...least 8 years of dedicated professional retail experience in...Hourly payWeekday work$20 - $60 per hour
...Help train next-generation AI systems by creating rigorous, real-world evaluations that test how well advanced models learn, reason, and perform. This remote contract opportunity... ...graduates, advanced-degree holders, and professionals from any background with strong research...Hourly payContract workFor contractorsRemote work$100 per hour
...finance expertise to help improve AI-driven financial applications... .... Develop, refine, and evaluate prompts related to financial... ...financial data, reports, and model outputs using detailed... ...robustness, and adherence to professional financial standards. Conduct...Hourly payPart timeFor contractorsRemote work$70 - $110 per hour
...Help advance frontier AI systems by bringing rigorous materials... ...engineering judgment to the evaluation, design, and improvement of... ...like in practice and ensure model outputs can withstand technical... ...strongly preferred. Hands-on professional use of large language models...Hourly payFull timeLive inRelocationRelocation package$60 - $80 per hour
...Help shape the training and evaluation of foundational large language models by applying real-world expertise in brand... ...rigorous marketing judgment to AI tasks, model assessments, and training... ...-reasoned solutions grounded in professional marketing practice. Assess AI model...Hourly payWeekday work$70 - $110 per hour
...Role Overview Help a leading AI research team improve how advanced AI models reason about real clinical work. In... ...clinical tasks, model answers, and evaluation standards alongside research and... ...clinical decisions. Hands-on professional use of large language models and...Hourly payFull timeFreelanceLive inRelocationRelocation package$60 per hour
Prolific, located in Arizona, is seeking Biology Experts and Life Science Professionals to join their Expert Network. This role involves evaluating and training AI models with real scientific expertise. Successful candidates will review AI-generated content for accuracy...Hourly pay- Prolific is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities...Remote jobHourly payFlexible hours
$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work... ...competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and...Remote jobHourly payWork from homeFlexible hours- Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote jobFlexible hours
$100 per hour
...Apply deep domain expertise to train and evaluate next-generation AI systems by producing, refining, and... ...contractor role focuses on improving model outputs through careful content... ...prompts to guide model behavior, using professional writing and technical documentation skills...Hourly payPart timeFor contractorsRemote work$100 - $150 per hour
...Role Overview Evaluate how well AI systems perform real-world technical sales work by defining excellence and judging completed work samples... ...Process Submit your resume or a summary of relevant professional experience to begin the review. Qualified candidates may be...Hourly payRemote work- A leading AI company is seeking a legal professional for a contractor role focused on evaluating AI model outputs in legal contexts. Candidates must hold a Juris Doctor (J.D.) and have more than 3 years of experience in law. The role involves reviewing complex legal hypotheticals...For contractors10 hours per week
$65 - $105 per hour
...Overview Help advance frontier AI systems by applying deep engineering... ...benchmarks used to assess how models reason about real software and... ...engineering tasks that reflect professional practice. Design rigorous engineering evaluation sets and contribute to engineering...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package$65 - $105 per hour
...Help improve how frontier AI models reason about real-world life sciences research. In... ...will apply deep scientific judgment to evaluate research tasks and model outputs, define... ...experience using large language models in professional work and the ability to distinguish...Hourly payFull timeLive inRelocationRelocation package$100 per hour
...the performance of large language models on finance tasks. You will work with AI researchers to identify model... ...systems. Key Responsibilities Evaluate LLM performance in finance areas... ...Qualifications Minimum 2 years of professional experience in one or more of the...Hourly payContract workFor contractorsFreelanceRemote work10 hours per weekFlexible hours$400 per month
...Mercor is partnering with a leading AI research lab to support a Frontier... ...project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments... ...strengths and weaknesses. Apply professional engineering judgment to realistic...- ...Opportunity Deepgram is looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team responsible for validating the... ...related field, or equivalent experience. ~5+ years of professional software or QA engineering experience, with a track record...Full time
- ...Role Overview Evaluate and annotate AI-generated music in Arabic and English, providing detailed... ...feedback to improve generative musical models. This role focuses on listening, comparing... ...At least 3 years of professional experience as a music producer, audio...Hourly payPart timeImmediate startRemote work10 hours per week
$224k - $356.5k
...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts... ...computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in crafting the...Full time$80 - $110 per hour
...to help improve next generation AI systems through accurate, practical evaluation and feedback. This part time contract... ...written feedback that improves model performance. Verify... ...the United States. ~5+ years of professional tax preparation or advisory experience...Hourly payContract workPart timeFor contractorsRemote work$400 per month
...Mercor is partnering with a leading AI research lab to support a Frontier... ...Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical... ...their strengths and weaknesses. Apply professional engineering judgment to realistic...- Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should...Weekday work
- Cincinnatus LLC is recruiting for an Insurance Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting judgments, with placement at the client...Weekday work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Retail Professional for AI Model Evaluation. Be the first to apply!



