Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Assistant for AI Model Evaluation

SaidGig

Role Overview

Support a research program with a leading AI lab by benchmarking and improving state of the art language models. This part time role focuses on creating evaluation content, formulating research questions and answers, and producing clear, well summarized research outputs to help refine advanced AI systems.

Key Responsibilities
  • Design and produce high quality research questions, model answers, and evaluation materials for language model benchmarking.
  • Perform targeted online research, gather information from diverse sources, and synthesize findings into concise summaries.
  • Evaluate model outputs against established criteria and document results for research teams.
  • Communicate findings clearly in written form and follow structured guidelines for task completion.
  • Work under a defined, low pressure structure that is not equivalent to a high intensity internship or full time consulting engagement.
Qualifications
  • Currently enrolled in or recently completed Masters studies, or 2 to 3 years of relevant experience plus a Bachelors degree from a reputable institution.
  • Strong online research skills and excellent written communication.
  • Ability to gather, distill, and clearly summarize information from multiple sources.
  • Detail oriented, reliable, and able to follow structured task instructions.
  • Interdisciplinary degree or background is a plus.
  • Required device: a Macbook with an M series chip as described below.
Work Terms
  • Engagement type: hourly, part time.
  • Estimated commitment: approximately 10 to 20 hours per week.
  • Location: remote.
  • Preferred start: immediate.
  • Tasks focus on research and evaluation for AI model improvement, with clearly defined deliverables and timelines.
Compensation
  • Pay rate: 50 to 60 hourly.
Eligibility
  • Must have a Macbook device with an Apple Silicon based M series chip to perform tasks.
  • Accepted Mac hardware includes Apple Silicon ARM Macs, M series MacBook Pro models, or MacBook Air models running macOS 15 or higher.
  • Applicants should meet the educational or experience requirements listed under Qualifications.
Application Process
  • Submit your resume to apply.
  • Complete a brief, 15 minute conversation with an AI interviewer that assesses research and reasoning skills.
  • Take a short paid assessment to demonstrate fit for the role.
  • Expect follow up communication within a few days about your application status and next steps.
Vacancy posted 15 days ago
Similar jobs that could be interesting for youBased on the Research Assistant for AI Model Evaluation in United States vacancy
  • $238k - $302k

     ...in simulation across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in Large Language Models...  ...are looking for quantitatively-minded engineers to research and propose new ways to assess the ML models deployed... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $40 per hour

    A healthcare technology firm is seeking medical experts to evaluate AI chatbots' performance and ensure their medical accuracy. This position allows for flexible scheduling and project selection, making it suitable for both full-time and part-time professionals. Candidates... 
    Suggested
    Hourly pay
    Full time
    Part time
    Remote work
    Flexible hours

    DataAnnotation

    Raleigh, NC
    5 days ago
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public...  ...team at Scale deploys advanced AI systems—including LLMs,...  ...systems. Ability to convert research insights into measurable...  ...mental disabilities. If you need assistance and/or a reasonable... 
    Suggested
    Full time

    Scale Ai

    Washington DC
    1 day ago
  • $40 per hour

    A biotechnology company is seeking a Biotechnology R&D Scientist to train AI models and evaluate their outputs. This role involves measuring the progress of AI chatbots with complex biology questions and ensuring their performance and correctness. Candidates should have... 
    Suggested
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Vermont
    5 days ago
  • $224k - $356.5k

     ...tapping into the unlimited potential of AI to define the next era of computing. An...  ...Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful...  ...and communicate effectively across research, engineering, and product teams.Ways to... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    5 days ago
  • $40 per hour

    A forward-thinking AI development firm seeks experienced quantitative professionals to evaluate AI-generated work, applying their skills in statistical analysis, predictive modeling, and technical writing. This fully remote opportunity offers a flexible schedule and projects... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Sioux Falls, SD
    5 days ago
  • $40 per hour

     ...A leading AI development firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully remote position allows for a flexible schedule, offering competitive pay starting at $40+ per hour. Ideal candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Bismarck, ND
    5 days ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting... 
    Hourly pay
    Remote work

    DataAnnotation

    Juneau, AK
    5 days ago
  • $40 per hour

    A leading AI training company is seeking a Biotechnology R&D Scientist to evaluate and improve AI models by posing complex biological questions. This remote position allows candidates to choose projects and work on their own schedule, with rates starting at $40+ per hour... 
    Hourly pay
    Remote work

    DataAnnotation

    Sioux Falls, SD
    3 days ago
  • $150 per hour

     ...Overview Aerospace engineering professionals apply their domain expertise to evaluate AI-generated outputs, assess technical content, and provide clear, structured feedback that improves models'' understanding of aerospace tasks, terminology, and practices. You will... 
    Hourly pay
    Temporary work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...Play a central role on a GenAI research team by applying hands-on...  ...experience to improve how frontier AI models perform real legal work. In this position you will evaluate model outputs, create high-...  ..., senior in-house counsel, or Assistant General Counsel level, with real... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    4 days ago
  • $40 per hour

    A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative work and provide critical feedback. This role offers the flexibility of remote work, allowing you to set your own schedule while focusing on impactful projects... 
    Hourly pay
    Remote work

    DataAnnotation

    Wisconsin
    5 days ago
  • $85 per hour

     ...environmental assessment, GIS, and renewable energy siting expertise to evaluate AI-generated outputs and to create expert-level training data...  ...use in your daily work. Collaborate asynchronously with AI research teams, documenting edge cases, common errors, and context... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...Overview Drive the creation and evaluation of challenging STEM problems...  ...and benchmark large language models. You will design multi-step...  ..., and collaborate with researchers to build evaluation benchmarks...  ...company accelerates frontier AI research and helps enterprises... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to work remotely. In this role, you'll evaluate AI-generated quantitative work and solve technical problems while providing feedback to shape AI systems. Qualifications include 2+ years... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Salt Lake City, UT
    5 days ago
  • $30 - $90 per hour

     ...backend services in Go while evaluating and training alpha-stage AI coding tools. This...  ...structured testing of AI-assisted coding workflows to improve...  ...Test and evaluate alpha AI models using Cursor, conducted...  ...Collaborate with the research team via Slack, providing... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • A leading AI development company is seeking experienced quantitative professionals for remote work evaluating AI-generated quantitative analysis. Ideal candidates will have a robust background in fields like data science, economics, or biostatistics, with at least 2 years... 
    Remote work

    DataAnnotation

    New York, NY
    5 days ago
  • $100 - $150 per hour

     ...subject-matter expertise to a GenAI research team, creating authoritative...  ...specifications, golden solutions, and evaluation benchmarks so frontier AI models reason correctly about real legal...  ...the engagement starts. Relocation assistance is not provided. Provisioning:... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Local area
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    4 days ago
  • $40 per hour

     ...analytics company seeks experienced quantitative professionals to evaluate AI-generated analysis and help advance AI development. This...  ...particularly those with experience in statistical methods and predictive modeling. Join to impact the next generation of AI systems dedicated to... 
    Hourly pay
    Remote work

    DataAnnotation

    El Paso, TX
    5 days ago
  • $85 per hour

     ...Overview Psychology experts apply clinical and research knowledge to design tasks, create domain-specific prompts, and evaluate large language models to improve their understanding and...  ...research. This role supports year-round AI research projects that vary by domain and... 
    Hourly pay
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $60 per hour

     ...contribute to developing cutting-edge AI systems, while enjoying the...  ...advance AI development. AI models are increasingly capable of...  ...the-art AI models on tasks like evaluating AI-generated quantitative...  ...economics, biostatistics, operations research, or any other quantitative... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Charleston, WV
    5 days ago
  •  ...that develops large language models, shaping training data by designing...  ...practice with rigorous evaluation to improve model behavior for...  ...Responsibilities Work with research and engineering teams to close...  ...marketing practice. Evaluate AI model outputs using structured... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    17 days ago
  • $85 per hour

     ...renewable energy generation, REC trading, and portfolio management to evaluate AI-generated content and create expert training material. This...  ...and communicate effectively in writing with a distributed AI research team. Application Process Create a contributor profile... 
    Hourly pay
    Contract work
    Part time
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $65 - $90 per hour

     ...Role Overview Apply your real-world architecture expertise to evaluate and improve how AI systems understand and reason about architecture. In this flexible, part-time, remote role you will review content for technical accuracy, answer domain-specific questions, and provide... 
    Hourly pay
    Part time
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Remote
    15 days ago
  • $60 - $80 per hour

     ...building foundational large language models by applying deep insurance domain expertise to create, evaluate, and refine training data. You...  ...Responsibilities Work with research and engineering teams to close...  ...claims practice. Evaluate AI model outputs against... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    a month ago
  • $40 per hour

     ...A leading AI development firm is seeking experienced quantitative professionals to join their team remotely. The role involves evaluating AI-generated quantitative work, providing insights, and shaping the future of AI systems. Candidates should have over two years of... 
    Hourly pay
    Remote work

    DataAnnotation

    Indiana, PA
    5 days ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated work and help shape future AI systems. This fully remote role offers a flexible schedule, competitive pay starting at $40+ USD per hour, and opportunities to work... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Nashville, TN
    5 days ago
  • $40 per hour

     ...data annotation company is seeking professionals in quantitative fields to enhance AI development. This fully remote role allows individuals to set flexible schedules while evaluating AI-generated analyses and solving complex quantitative problems. Candidates should have... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Oregon, WI
    4 days ago
  • $40 per hour

     ...A forward-thinking AI team is seeking quantitative professionals to evaluate and improve cutting-edge AI systems. The role involves working on AI-generated analyses and providing critical feedback for model enhancement. Candidates should have a strong quantitative background... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Helena, MT
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Assistant for AI Model Evaluation. Be the first to apply!