Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

QA Engineer for AI Model Evaluation

$90 - $175 per hour

SaidGig

Role Overview

Apply software quality assurance expertise to evaluate technical AI outputs and help improve how next-generation AI systems learn, reason, and perform. This remote contract opportunity supports a high-volume project and does not require prior AI experience, but it does require prior paid human-in-the-loop AI evaluation work.

Key Responsibilities
  • Evaluate and score AI-generated technical outputs against established quality criteria and rubrics, identifying responses that appear correct but are inaccurate or incomplete.
  • Design and assess thorough functional, regression, edge-case, negative, and boundary test cases.
  • Review bug reports and test documentation for reproducibility, completeness, and appropriate severity assessment.
  • Identify, isolate, and document defects with precise reproduction steps using structured tracking methods.
  • Provide detailed written feedback and annotations that developers can act on without further clarification.
  • Work with project teams to refine evaluation guidelines and continuously improve testing standards and methods.
Qualifications
  • Professional experience as a QA Engineer, SDET, Test Engineer, QA Analyst, or in a comparable software quality assurance role.
  • Strong knowledge of quality assurance, software testing, test case design, bug tracking, regression testing, and manual and automated testing strategies.
  • Hands-on experience with automation frameworks and test management tools, such as Selenium, Playwright, Cypress, Appium, Postman, Jira, TestRail, Zephyr, BrowserStack, or similar tools.
  • Prior paid experience supporting AI training through human data annotation, labeling, RLHF, AI response or model evaluation, or rubric-based grading. Software QA experience alone does not meet this project requirement.
  • Excellent analytical and problem-solving skills, close attention to detail, and the ability to communicate complex findings clearly in written English at B2 level or above.
  • No formal degree is required. Demonstrable practical testing experience is prioritized.
  • Reliable internet access and readiness to begin promptly.
Work Terms
  • Remote contract engagement.
Compensation
  • $90 to $175 per hour.
Application Process

Apply through the available application flow using email or a Google account. Continuing the application requires agreement to the applicable terms and privacy policies.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the QA Engineer for AI Model Evaluation in United States vacancy
  • $60 - $90 per hour

     ...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Summers , and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract... 
    Suggested
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    19 hours ago
  • $60 per hour

     ...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Suggested
    Hourly pay
    Remote work
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    2 days ago
  • $60 per hour

     ...and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with...  ...offers a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing... 
    Suggested
    Hourly pay
    Remote work
    Work from home
    Flexible hours

    Prolific

    Dallas, TX
    2 days ago
  • $60 - $80 per hour

     ...develop advanced large language models. In this role, you will bring practical...  ..., and campaign judgment to AI training data, partnering with research and engineering teams to improve model...  ...quality. Develop and improve evaluation guidelines and scoring rubrics for... 
    Suggested
    Hourly pay
    Weekday work

    SaidGig

    United States
    a month ago
  • $100 - $150 per hour

     ...Overview Apply senior legal judgment to help a leading AI research organization improve how advanced AI models reason about real-world legal work. You will work...  ...define high-quality legal tasks, standards, and evaluations. This role is designed for a practicing legal... 
    Suggested
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    a month ago
  • $70 - $90 per hour

     ...kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical...  ...including NKI, Pallas, or TPU. Background in compiler engineering, MLIR, or intermediate-representation lowering. Knowledge... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    18 days ago
  •  ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets... 
    For contractors
    Remote work

    SaidGig

    United States
    16 days ago
  • $70 - $80 per hour

     ...Role Overview Apply advanced drug safety expertise to help improve AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data. This remote contract role focuses on the quality, accuracy, and regulatory alignment of complex safety reports... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    a month ago
  • $70 - $90 per hour

     ...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness...  ...NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    18 days ago
  • $100 per hour

     ...expertise to improve the performance of large language models on finance tasks. You will work with AI researchers to identify model weaknesses in areas...  ...focused on advanced AI systems. Key Responsibilities Evaluate LLM performance in finance areas where models... 
    Hourly pay
    Contract work
    For contractors
    Freelance
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...emerging trillion-dollar Voice AI economy, providing real-time...  ...’s voice-native foundation models are accessed through cloud...  ...looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the...  ...and strong engineering and QA practices. What We're Looking... 
    Full time

    Deepgram

    United States
    19 hours ago
  •  ...computational problem solving to improve and evaluate large language models. You will design rigorous math...  ...customers Accelerate frontier AI research by contributing high quality...  ...mathematics at the level expected for engineering entrance exams and for graduate or PhD... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    a month ago
  •  ...reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial...  ...growing group of committed researchers, engineers, policy experts, and business leaders...  ...for Research Engineers to build the evaluations that tell us — and the world — what Claude... 
    Full time

    Anthropic

    New York, NY
    3 days ago
  • $100 - $150 per hour

     ...will be considered for future projects evaluating how well AI systems perform real-world data...  ...assess AI or human-produced analyses and models, document decisions in writing, and iterate...  ...and A/B test write-ups, feature engineering, and technical reports or notebooks... 
    Hourly pay
    Immediate start
    Remote work

    SaidGig

    United States
    a month ago
  • $60 - $80 per hour

     ...expertise to help develop advanced large language models by bringing real-world underwriting, claims, and risk-assessment judgment to AI training and evaluation work. Key Responsibilities Partner with research and engineering teams to address knowledge gaps in... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    more than 2 months ago
  • $36 - $72 per hour

     ...seekers, 1 million+ employers, and 1,600 educational institutions. Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise. Role Details Location: Onsite in Seattle, WA,... 
    Hourly pay
    Full time
    Monday to Friday
    Flexible hours

    Handshake

    Seattle, WA
    6 days ago
  • $60 - $80 per hour

     ...mathematics experts considered for future contract opportunities with AI labs and companies. This is an open application, not a posting...  ...by applying mathematics knowledge to real-world tasks, model evaluation, and domain-specific feedback. Key Responsibilities... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $100 per hour

     ...Apply consulting expertise to evaluate and improve AI-generated business content for a customer-facing project. Your judgment will help AI systems...  ...executive summaries. Develop and refine large language model prompts using structured problem-solving and analytical rigor... 
    Hourly pay
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    22 days ago
  • $110 per hour

     ...Apply to join a physician talent network supporting AI labs and companies with medical expertise. This is an open application...  ...projects by bringing real-world clinical expertise to model development and evaluation. Key Responsibilities Train and evaluate AI models... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    a month ago
  • $65 - $105 per hour

     ...Apply deep engineering judgment to help frontier AI models reason more accurately about real-world software development and engineered systems. You will...  ...management team to define high-quality engineering work, evaluate model performance, and turn expert practice into... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    22 days ago
  • $60 - $90 per hour

     ...Role Overview Design rigorous, real-world materials science and engineering tasks that evaluate an AI model’s expert reasoning. You will create prompts, supporting data rooms, and grading criteria, then test and refine each task until it clearly distinguishes strong from... 
    Hourly pay
    Work at office
    Remote work

    SaidGig

    United States
    12 days ago
  • $65 - $105 per hour

     ...Help advance frontier AI models by bringing rigorous life sciences research judgment into task design, evaluation, and model improvement. You will work closely with an AI lab''s research and program management teams to ensure models can reason credibly about real scientific... 
    Hourly pay
    Full time
    Freelance
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    22 days ago
  •  ...Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced QA and Test Engineers to ensure every benchmark is reliable, reproducible, and accurately measures real AI capabilities... 
    Full time
    Contract work
    For contractors
    Remote work
    Flexible hours

    Weekday

    Remote
    a month ago
  • $100 - $150 per hour

     ...Apply senior finance expertise to improve how frontier AI systems reason through real-world financial work. You will partner closely...  ...with an AI research team to define high-quality finance tasks, evaluate model performance, and translate professional judgment into rigorous... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    a month ago
  • $70 - $110 per hour

     ...Role Overview Shape how advanced AI systems reason about real clinical work. In this...  ...will partner with an AI research team to evaluate medical knowledge tasks, define high-quality...  ...develop benchmarks that measure meaningful model improvement. Key Responsibilities Review... 
    Hourly pay
    Full time
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    22 days ago
  •  ...Apply your dermatology expertise to image-based work that helps develop and evaluate advanced AI models for clinical image interpretation. This is a non-clinical, remote opportunity with no direct patient care, focused on ensuring AI-generated medical outputs reflect real... 
    Hourly pay
    Remote work

    SaidGig

    United States
    3 days ago
  • $60 - $90 per hour

     ...Role Overview Help shape how an advanced performance-transfer model evaluates character animation, preserving an actor''s timing, emotion,...  ...evaluation methods, and help build a reliable human-review process for AI-generated performance results. Key Responsibilities... 
    Hourly pay
    Part time

    SaidGig

    Oregon State
    a month ago
  • $20 - $36 per hour

     ...Role Overview Evaluate AI-generated music in Slovak and English by listening across genres...  ...labels will help improve generative music models. Key Responsibilities Compare...  ...professional experience as a music producer, audio engineer, or mixing engineer Headphones or... 
    Hourly pay
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  • $60 - $80 per hour

     ...Apply deep retail domain expertise to help build and evaluate foundational generative AI models. You will design realistic retail tasks, produce authoritative...  ...practical judgment while working with research and engineering teams at a leading GenAI lab. Key Responsibilities... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    1 day ago
  • $100 - $150 per hour

     ...matter expertise to a GenAI research team, creating authoritative instruction specifications, golden solutions, and evaluation benchmarks so frontier AI models reason correctly about real legal work. This role centers on hands-on legal judgment, translating professional... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Local area
    Relocation
    Relocation package

    SaidGig

    California
    7 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to QA Engineer for AI Model Evaluation. Be the first to apply!