PowerPoint Specialist for AI Model Evaluation
$20 - $60 per hourSaidGig
Role Overview
Apply advanced business-document expertise to help train next-generation AI systems on Project Navara. You will create realistic, high-complexity office-document scenarios, collaborate with language models, and evaluate their output to improve how they handle real-world business work. Prior AI experience is not required.
Key Responsibilities
- Design realistic, complex tasks involving professional PowerPoint, Excel, and Word use, modeled on Fortune 500 business environments.
- Work with an advanced language model to discuss, edit, and create Office Open XML files, with a primary focus on .pptx files.
- Review and compare two AI-generated deliverables at each task stage, selecting the stronger result based on accuracy and effectiveness.
- Create prompts and assignments that ordinarily require substantial business judgment and multiple hours of work without AI assistance.
- Participate in structured four-turn conversations with the language model, providing feedback and preferences after each round.
- Document observations and recommendations that improve AI handling of real-world business documents.
- Apply expertise across cross-functional business use cases, including reporting, presentations, analysis, and proposals across industries and functions.
Qualifications
- At least 3 years of recent experience delivering complex projects at Fortune 500 companies or top-tier consulting firms.
- Advanced PowerPoint skills, significant Excel and Word experience, and a strong record of creating and reviewing professional business documents.
- Experience executing projects in finance, healthcare, consulting, technology, or retail, and in functions such as strategy, operations, sales, marketing, finance, or HR.
- Ability to develop realistic business scenarios and workflows for modeling, reporting, and executive presentations.
- Strong written and verbal communication skills, including the ability to provide clear, actionable feedback.
- Experience with language models or AI-powered tools, particularly in iterative conversational workflows, is highly desirable.
- Relevant backgrounds include private equity associate, investment banking analyst, consultant, FP&A analyst, HR professional, and marketing strategist.
Work Terms
- Independent contractor engagement.
- Remote work.
Compensation
- $20 to $60 per hour.
- Dorado is seeking a Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems, probe model reasoning at the frontier, and help identify where models fail under rigorous scientific scrutiny....SuggestedRemote job
$20 - $60 per hour
...Help improve how large language models create, understand, and modify... ...expert-level Excel, Word, and PowerPoint skills. You will design... ...business scenarios, produce and evaluate complex .xlsx, .docx, and .pptx... ...model improvements. No prior AI experience is required. Key...SuggestedHourly payFor contractorsWork at officeRemote work$208k - $300k
...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure...SuggestedFull time$60 - $90 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract Compensation...SuggestedFull timeContract workSummer workRemote work$224k - $356.5k
...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts... ...computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in crafting the...SuggestedFull time- Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should...Weekday work
- Cincinnatus LLC is recruiting for an Insurance Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting judgments, with placement at the client...Weekday work
- ...Overview Apply advanced physics knowledge to help improve and evaluate large language models. You will design rigorous problems, produce clear reasoning... ...into accessible explanations while contributing to AI research projects. Key Responsibilities Design and solve...For contractorsFreelanceRemote work
- ...Evaluate generative music AI across a broad range of genres, applying your knowledge of Indonesian music and lyrics to detailed quality standards. You will work in both Indonesian and English to help assess lyric quality and authenticity. Key Responsibilities Compare...Hourly payImmediate startRemote workFlexible hours
$20 - $60 per hour
...Help train next-generation AI systems by creating rigorous, real-world evaluations that test how well advanced models learn, reason, and perform. This remote contract opportunity is open to recent graduates, advanced-degree holders, and professionals from any background...Hourly payContract workFor contractorsRemote work$42 - $78 per hour
...Evaluate generative music and lyrics in Norwegian and English, helping assess output across a broad range of genres against detailed quality standards. Key Responsibilities Compare AI-generated lyrics with published songs to identify similarities. Rate lyrics for...Hourly payFor contractorsImmediate startRemote workFlexible hours$70 - $110 per hour
...Help advance frontier AI systems by bringing rigorous materials... ...engineering judgment to the evaluation, design, and improvement of... ...like in practice and ensure model outputs can withstand technical... ...researchers and adjacent domain specialists to develop materials-focused...Hourly payFull timeLive inRelocationRelocation package$60 - $80 per hour
...Help shape the training and evaluation of foundational large language models by applying real-world expertise in brand strategy, growth marketing, and campaign... .... This role brings rigorous marketing judgment to AI tasks, model assessments, and training data for a leading...Hourly payWeekday work$35 - $62 per hour
...Apply your Japanese music expertise to evaluate AI-generated music and lyrics across a wide range of genres. You will assess outputs against detailed quality standards in both Japanese and English. Key Responsibilities Compare AI-generated lyrics with published songs...Hourly payFor contractorsImmediate startRemote workFlexible hours$65 - $90 per hour
...Apply deep financial judgment to help improve foundational AI models. In this role, you will create and evaluate finance-focused work that strengthens AI systems'' reasoning, analysis, and decision-making capabilities. Role Overview This is a W-2 employment opportunity...Hourly payWeekday work- ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear, logically... ...How this work supports customers Accelerate frontier AI research by contributing high quality data and evaluation...Contract workFor contractorsFreelanceRemote work
$13 - $54 per hour
...Evaluate AI-generated music and lyrics across a wide range of genres, using your knowledge of the Spanish (Mexico) music scene to assess quality against detailed standards. This remote role involves working in both Spanish (Mexico) and English. Key Responsibilities...Hourly payImmediate startRemote workFlexible hours$18 - $42 per hour
...Role Overview Evaluate generative music AI outputs across a wide range of genres, applying your knowledge of Portuguese-language music and lyrics to detailed quality standards. This role combines critical listening with lyric analysis in Portuguese and English. Key...Hourly payImmediate startRemote workFlexible hours$18 per hour
...Role Overview Evaluate AI-generated music and lyrics across a range of genres, applying detailed quality standards in both Thai and English... ...of Thai music, language, and lyrical expression to help assess model outputs. Key Responsibilities Compare AI-generated lyrics...Hourly payFor contractorsImmediate startRemote workFlexible hours$70 - $110 per hour
...Role Overview Help a leading AI research team improve how advanced AI models reason about real clinical work. In... ...clinical tasks, model answers, and evaluation standards alongside research and... ...researchers and adjacent-domain specialists, turning implicit clinical judgment...Hourly payFull timeFreelanceLive inRelocationRelocation package- ...is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities...Remote jobHourly payFlexible hours
$60 per hour
...and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with... ...offers a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing...Remote jobHourly payWork from homeFlexible hours$60 per hour
...in Arizona, is seeking Biology Experts and Life Science Professionals to join their Expert Network. This role involves evaluating and training AI models with real scientific expertise. Successful candidates will review AI-generated content for accuracy and assist in fact...Hourly pay$70 - $110 per hour
...Role Overview Apply your legal practice experience to evaluate and improve AI-generated legal content and workflows. You will assess work grounded in litigation, legal drafting, and client advisory practice. Key Responsibilities Evaluate AI-generated legal memoranda...Hourly payImmediate startRemote work- Cincinnatus LLC is seeking a Marketing SME to join a leading GenAI team focused on evaluating AI outputs against rubrics and strengthening brand strategy in AI training data. We require 8+ years of marketing experience with top-tier brands, plus hands-on evaluation of LLM...Weekday work
- Dorado is seeking an experienced Investment Banking SME to support the development, evaluation, and improvement of advanced AI models in finance. You will assess AI-generated analyses for accuracy, reasoned judgments, and data integrity, guiding model refinements and prompts...Remote job
$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Remote jobHourly payWork from homeFlexible hours- Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote jobFlexible hours
$60 - $90 per hour
...Role Overview Help advance frontier AI research by creating rigorous, real-world data analysis evaluations for generative AI models. You will design and complete complex analytical tasks that mirror practical research work, then use your reference analyses to assess where...Hourly payFull timeFreelanceRemote work$60 - $80 per hour
...GenAI team building foundational large language models by applying deep insurance domain expertise to create, evaluate, and refine training data. You will translate real... ...real underwriting and claims practice. Evaluate AI model outputs against structured rubrics,...Hourly payWeekday work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to PowerPoint Specialist for AI Model Evaluation. Be the first to apply!
- cost specialist United States
- strategic sourcing specialist United States
- absence management specialist United States
- lead sourcing specialist United States
- peer recovery specialist United States
- authorization specialist United States
- treasury specialist United States
- print production specialist United States
- workforce management specialist United States
- wellness specialist United States



