Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

QA/Test Engineer for AI Model Evaluation

$60 - $90 per hour

SaidGig

Role Overview

Validate the integrity of complex, multi-step evaluation tasks used to benchmark frontier generative AI models. You will design and execute rigorous test cases, probe edge cases, and debug task environments so each task is unambiguous, correctly graded, and robust to shortcuts. Individual tasks typically represent one to two days of expert effort and span multiple technical skills. You will operate in a close feedback loop with the lab''s researchers and task authors to ensure benchmark results remain trustworthy.

Key Responsibilities
  • Design checks and test cases that confirm each task behaves as intended, including difficult edge cases.
  • Review tasks and reference solutions in detail before finalization, identifying ambiguity, grading gaps, and missing assumptions.
  • Debug task logic and verification code, actively using Python to investigate and fix failures.
  • Help create simple, repeatable quality checklists and provide actionable feedback that task authors can apply quickly.
  • Monitor AI agent runs for shortcuts or grading weaknesses to protect benchmark validity.
  • Collaborate closely with researchers and task authors in an iterative feedback loop to improve task quality and evaluation processes.
Qualifications
  • MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy or engineering-heavy domain.
  • At least 1 year of experience in test engineering, quality assurance, or a research or software engineering role that included strong quality ownership.
  • Proven ability to design test cases and quality-review processes, and to debug complex systems end-to-end.
  • Working proficiency in Python and Git, with comfort navigating unfamiliar codebases and runtime environments.
  • Exceptional attention to detail and clear written documentation habits.
  • Prior experience with AI training, model evaluation, or quality review of AI-generated outputs is preferred.
  • A perfectionist mindset, creativity in finding what others missed, and the ability to work independently on ambiguous, open-ended problems.
  • Availability to engage reliably for approximately 35 hours per week.
Work Terms
  • Status: Full-time W-2 employment with Cincinnatus LLC, acting as the employer of record for this placement.
  • Placement: Opportunity to be placed at a leading AI lab as part of their extended workforce, working within the client team and enterprise workflows.
  • Location: Fully remote within the United States, candidates must be able to work from the U.S.
  • Schedule: Approximately 35 hours per week.
  • Engagement type: Role-based employment, not a freelance or project-by-project arrangement; integration with client teams and standard enterprise processes is expected.
Compensation
  • Pay rate: $60.00 to $90.00 per hour, paid on an hourly basis.
Eligibility and Hiring Process
  • Employment, onboarding, payroll, and benefits for this role are administered by Cincinnatus LLC, the employer of record.
  • Opportunities may be discovered through third-party talent platforms, but final hiring and employment administration are handled by Cincinnatus LLC.
  • Cincinnatus LLC is an Equal Employment Opportunity employer and provides reasonable accommodations for qualified individuals with disabilities throughout the application process.
Vacancy posted 11 days ago
Similar jobs that could be interesting for youBased on the QA/Test Engineer for AI Model Evaluation in United States vacancy
  •  ...Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced QA and Test Engineers to ensure every benchmark is reliable, reproducible, and accurately measures real AI capabilities... 
    Suggested
    Full time
    Contract work
    For contractors
    Remote work
    Flexible hours

    Weekday

    Remote
    15 days ago
  •  ...Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced QA and Test Engineers to ensure every benchmark is reliable, reproducible, and accurately measures real AI capabilities... 
    Suggested
    Full time
    Contract work
    For contractors
    Remote work
    Flexible hours

    Weekday

    Remote
    7 days ago
  • $90 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ..., and Jack Dorsey . Position: QA/Test Engineer Type: Contract Compensation:...  ...Preferred Experience in AI training , model evaluation, or quality review of AI-generated... 
    Suggested
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    4 days ago
  • $146k - $194k

     ...expertise, technology, and business model of the 21st century’s most...  ...is powered by Lattice OS, an AI-powered operating system that...  ...skilled and experienced Test & Evaluation Manager who is passionate about...  ...for a highly motivated test engineer with emphasis in developmental... 
    Suggested
    Full time
    Work experience placement
    Immediate start
    Remote work

    Anduril Industries

    Costa Mesa, CA
    2 days ago
  • $40 per hour

    A biotechnology company is seeking a Biotechnology R&D Scientist to train AI models and evaluate their outputs. This role involves measuring the progress of AI chatbots with complex biology questions and ensuring their performance and correctness. Candidates should have... 
    Suggested
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Vermont
    1 day ago
  •  ...Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content...  ...Participants enjoy flexible hours and competitive pay, with a brief assessment test prior to joining. #J-18808-Ljbffr... 
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    1 day ago
  • $224k - $356.5k

     ...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our...  ...performance computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in crafting... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $45 - $52 per hour

    DescriptionKforce has a client seeking a QA Test Engineer in Boca Raton, FL to support web application testing initiatives. This role is responsible...  ...to CI/CD processes and continuous testing practices* Leverage AI-assisted tools to improve testing efficiencyRequirements*... 

    KForce

    Boca Raton, FL
    3 days ago
  • $40 per hour

     ...A leading AI development firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully remote position allows for a flexible schedule, offering competitive pay starting at $40+ per hour. Ideal candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Bismarck, ND
    1 day ago
  • $60 per hour

     ...developing cutting-edge AI systems, while...  ...advance AI development. AI models are increasingly capable...  ...models on tasks like evaluating AI-generated quantitative...  ...design (e.g., A/B testing, hypothesis testing, regression...  ...Science, Mathematics, Engineering, or similar); a master... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Kansas City, MO
    2 days ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting... 
    Hourly pay
    Remote work

    DataAnnotation

    Juneau, AK
    1 day ago
  • $40 per hour

     ...A technology company in Mississippi seeks a Full Stack Engineer to improve AI models by providing coding challenges and evaluating performance. Candidates should be proficient in a programming language and fluent in English. This remote position offers flexibility in project... 
    Hourly pay
    Contract work
    Remote work

    DataAnnotation

    Mississippi
    2 days ago
  • $40 per hour

     ...A forward-thinking AI solutions company is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the development of cutting-edge...  ...skills. Join us to directly impact the future of AI analytics and model reasoning. #J-18808-Ljbffr... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Lincoln, NE
    1 day ago
  • $40 per hour

    A leading AI training company is seeking a Biotechnology R&D Scientist to evaluate and improve AI models by posing complex biological questions. This remote position allows candidates to choose projects and work on their own schedule, with rates starting at $40+ per hour... 
    Hourly pay
    Remote work

    DataAnnotation

    Sioux Falls, SD
    4 days ago
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines...  ...LLM-judge–based evaluations. Design test datasets and benchmarks to measure generalization... 
    Full time

    Scale Ai

    Washington DC
    1 day ago
  • $40 per hour

     ...A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated quantitative work and design problems for AI training. Enjoy the flexibility of fully remote work with a competitive pay starting at $40+ per hour. Ideal candidates... 
    Hourly pay
    Remote work

    DataAnnotation

    Washington DC
    1 day ago
  • $40 per hour

     ...DataAnnotation is seeking a Biotechnology R&D Scientist to train AI models. In this role, you will evaluate the outputs of AI chatbots and assess their logic to improve model quality. The ideal candidate should have a deep understanding of cell biology, genetics, biochemistry... 
    Hourly pay
    For contractors
    Remote work

    DataAnnotation

    New York, NY
    2 days ago
  • $60 per hour

     ...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific...  ...$60 per hour, and may work from home with flexible hours. A brief assessment test is required prior to joining. #J-18808-Ljbffr... 
    Hourly pay
    Remote work
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    1 day ago
  • $50 - $60 per hour

     ...DataAnnotation is seeking an Appellate Attorney to train AI models by evaluating their legal outputs and solving complex legal challenges. This role allows for flexible remote work, enabling you to choose projects based on your own schedule and preferences. A J.D. is mandatory... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Raleigh, NC
    1 day ago
  • $40 per hour

     ...seeking an R&D Biologist to join their team in the United States. In this remote role, you will train AI models by providing complex biology questions and evaluating chatbot responses. The ideal candidate will have an expert understanding of biology and related fields,... 
    Hourly pay
    Contract work
    Remote work
    Flexible hours

    DataAnnotation

    Oregon, WI
    1 day ago
  • $40 per hour

     ...Development Chemist to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the...  ...are not limited to: Chemistry and/or Chemical Engineering. Benefits This is a full-time or part-time... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work

    DataAnnotation

    Iowa, LA
    1 day ago
  • $40 per hour

    A growing technology company is seeking an R&D Biologist to join their team. This role involves training AI models by evaluating chatbot outputs on complex biology questions. Applicants should possess strong expertise in biology and related fields. You will work remotely... 
    Hourly pay
    Remote work

    DataAnnotation

    New York, NY
    1 day ago
  • $40 per hour

    A leading AI training firm in the United States is seeking an R&D Biologist to join their team. In this remote position, you will evaluate AI chatbots and enhance their models while ensuring the biological accuracy of their outputs. The ideal candidate should have an expert... 
    Hourly pay
    Remote work

    DataAnnotation

    Hartford, CT
    1 day ago
  • $40 per hour

    A technology company in Massachusetts is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal candidates will hold advanced qualifications in biology or biochemistry. This position allows full... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Oklahoma City, OK
    1 day ago
  • $40 per hour

     ...company in the United States is seeking an R&D Biologist to train AI models and improve their quality. This position offers remote...  ...selected projects at your own schedule. Responsibilities include evaluating the performance of AI chatbots on complex biology topics. Candidates... 
    Hourly pay
    Remote work

    DataAnnotation

    Brooklyn, NY
    2 days ago
  • $40 per hour

    A tech company specializing in AI training is looking for a Statistician to join their team. In this remote role, you'll train AI models by providing complex math problems and evaluating their outputs for quality and correctness. The ideal candidate will have strong mathematical... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Oregon, WI
    1 day ago
  • $50 - $60 per hour

     ...DataAnnotation is looking for an Appellate Attorney to help train AI models by providing complex legal problems and evaluating AI outputs. This role is ideal for General Counsel or those with similar legal expertise. Contract details include working on your own schedule... 
    Hourly pay
    Contract work
    Remote work

    DataAnnotation

    Washington DC
    5 days ago
  • $40 per hour

    A technology-focused company is seeking a Postdoctoral Researcher in Chemistry to evaluate AI chatbots based on complex chemistry questions. This remote position allows for flexible scheduling and project selection. Applicants must have a PhD in Chemistry and a strong command... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Raleigh, NC
    1 day ago
  • $40 per hour

    A technology firm specializing in AI is seeking a Biostatistician to enhance AI models by evaluating their performance and solving complex mathematical problems. This role offers flexibility as a remote position, allowing you to choose your projects and work according... 
    Hourly pay
    Remote work

    DataAnnotation

    Jackson, MS
    1 day ago
  • $40 per hour

    A healthcare technology firm is seeking medical experts to evaluate AI chatbots' performance and ensure their medical accuracy. This position allows for flexible scheduling and project selection, making it suitable for both full-time and part-time professionals. Candidates... 
    Hourly pay
    Full time
    Part time
    Remote work
    Flexible hours

    DataAnnotation

    Raleigh, NC
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to QA/Test Engineer for AI Model Evaluation. Be the first to apply!