QA/Test Engineer for AI Model Evaluation
$60 - $90 per hourSaidGig
Validate the integrity of complex, multi-step evaluation tasks used to benchmark frontier generative AI models. You will design and execute rigorous test cases, probe edge cases, and debug task environments so each task is unambiguous, correctly graded, and robust to shortcuts. Individual tasks typically represent one to two days of expert effort and span multiple technical skills. You will operate in a close feedback loop with the lab''s researchers and task authors to ensure benchmark results remain trustworthy.
Key Responsibilities- Design checks and test cases that confirm each task behaves as intended, including difficult edge cases.
- Review tasks and reference solutions in detail before finalization, identifying ambiguity, grading gaps, and missing assumptions.
- Debug task logic and verification code, actively using Python to investigate and fix failures.
- Help create simple, repeatable quality checklists and provide actionable feedback that task authors can apply quickly.
- Monitor AI agent runs for shortcuts or grading weaknesses to protect benchmark validity.
- Collaborate closely with researchers and task authors in an iterative feedback loop to improve task quality and evaluation processes.
- MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy or engineering-heavy domain.
- At least 1 year of experience in test engineering, quality assurance, or a research or software engineering role that included strong quality ownership.
- Proven ability to design test cases and quality-review processes, and to debug complex systems end-to-end.
- Working proficiency in Python and Git, with comfort navigating unfamiliar codebases and runtime environments.
- Exceptional attention to detail and clear written documentation habits.
- Prior experience with AI training, model evaluation, or quality review of AI-generated outputs is preferred.
- A perfectionist mindset, creativity in finding what others missed, and the ability to work independently on ambiguous, open-ended problems.
- Availability to engage reliably for approximately 35 hours per week.
- Status: Full-time W-2 employment with Cincinnatus LLC, acting as the employer of record for this placement.
- Placement: Opportunity to be placed at a leading AI lab as part of their extended workforce, working within the client team and enterprise workflows.
- Location: Fully remote within the United States, candidates must be able to work from the U.S.
- Schedule: Approximately 35 hours per week.
- Engagement type: Role-based employment, not a freelance or project-by-project arrangement; integration with client teams and standard enterprise processes is expected.
- Pay rate: $60.00 to $90.00 per hour, paid on an hourly basis.
- Employment, onboarding, payroll, and benefits for this role are administered by Cincinnatus LLC, the employer of record.
- Opportunities may be discovered through third-party talent platforms, but final hiring and employment administration are handled by Cincinnatus LLC.
- Cincinnatus LLC is an Equal Employment Opportunity employer and provides reasonable accommodations for qualified individuals with disabilities throughout the application process.
- ...Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced QA and Test Engineers to ensure every benchmark is reliable, reproducible, and accurately measures real AI capabilities...SuggestedFull timeContract workFor contractorsRemote workFlexible hours
- Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced QA and Test Engineers to ensure every benchmark is reliable, reproducible, and accurately measures real AI capabilities. In...SuggestedFull timeRemote work
- ...QA Test Engineer Location: Alpharetta, GA, USA Duration: 12+ Month Contract Must have strong AI experience Education: Bachelor's in Computer Science... ...are probabilistic. Evaluation strategies: golden datasets... ..., circuit breakers, model timeouts, degraded-mode behavior...SuggestedContract work
- ...Medical professionals apply clinical and workplace expertise to evaluate AI-generated content in their specialty, assess field-specific materials, and provide clear, structured feedback that improves model performance on medical tasks and language. This hourly, temporary...SuggestedHourly payTemporary workPart timeRemote workFlexible hours
$40 - $65 per hour
...scenarios that probe frontier language models, then evaluate and document model behavior so engineering teams can improve model safety,... ...written English, and iterative testing to produce high-quality... ...structure. Preferred experience with AI human data environments such as...SuggestedRemote jobHourly payFor contractorsImmediate start$60 - $80 per hour
...subject-matter expertise to a GenAI team building foundational AI models. You will create realistic marketing tasks, evaluate model outputs against structured rubrics, and advise research and engineering teams on brand strategy, growth marketing, and campaign-level reasoning...Hourly payWeekday work$100 per hour
...Role Overview Evaluate and optimize AI-generated outputs for a customer-facing project by applying... ..., clarity, and business alignment of model outputs through detailed review, prompt... ..., technical editing, and prompt engineering. Preferred experience: 3+ years delivering...Remote jobHourly payPart timeFor contractors$140 per hour
...improve how next-generation AI systems learn, reason,... ...Domain Expert you will evaluate AI outputs, create... ...and validate advanced models. This is a remote, part... ...and refine prompts that test AI reasoning, and model... ...finance, healthcare, STEM engineering, software, law, or...Remote jobHourly payPart timeFor contractorsVisa sponsorshipWork visaFree visa$60 - $80 per hour
...building foundational large language models, applying real-world... ...operations judgment to design tasks, evaluate model outputs, and guide research and engineering teams toward accurate retail reasoning... ...real retail practice. Evaluate AI model outputs against structured...Hourly payContract workWeekday work$60 - $80 per hour
...underwriting and claims judgment, and evaluate large language model outputs against structured rubrics to... ...Guide research and engineering teams to close knowledge gaps in underwriting... ...underwriting and claims practice. Evaluate AI model outputs against structured rubrics...Hourly payWeekday work$30 - $90 per hour
...contributing to the evaluation and training of next-generation AI coding tools in... ...structured product testing, reporting, and collaboration... ...alpha AI coding models in Cursor, running... ...researchers and engineers via Slack to provide... ...or developer tool QA. Interest in mentoring...Hourly payContract workPart timeRemote work$60 - $90 per hour
...serve as ground-truth references for evaluation of frontier generative AI models. You will create one-to-two day,... ...data cleaning, statistical testing, manual spot checks, and clear written... ...experience in research, research engineering, or heavy data analysis. Strong...Hourly payFull timePart timeWork experience placementFreelanceRemote work- ...in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant... ...work hours, competitive pay rates, and a requirement to complete an assessment test before joining. #J-18808-Ljbffr...Hourly payRemote workFlexible hours
$80 - $110 per hour
...Role Overview Contribute frontier research expertise to the development and evaluation of AI systems that reason about real research chemistry. This role supports model training and assessment across organic synthesis, reaction mechanism, catalysis, structural and physical...Hourly payPart timeImmediate startRemote work$80 - $110 per hour
...Contribute subject-matter expertise to the development and evaluation of next-generation AI systems that must reason about pure and applied... ...of a working research mathematician. You will help ensure models understand and produce correct, rigorous mathematics across...Hourly payPart timeImmediate startRemote work$85 per hour
...Role Overview Help evaluate and improve frontier AI coding models by completing realistic machine learning engineering tasks and assessing model-generated implementations. You will work with cutting-edge coding agents to surface bugs, failure modes, and tradeoffs in end...Hourly payRemote work$60 per hour
...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific... ...$60 per hour, and may work from home with flexible hours. A brief assessment test is required prior to joining. #J-18808-Ljbffr...Hourly payRemote workWork from homeFlexible hours$60 per hour
...and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with... ...offers a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing...Hourly payRemote workWork from homeFlexible hours$110 per hour
...Role Overview Provide clinical expertise to help train, evaluate, and shape medical AI systems by joining a Physician Expert Network. This is an... ...rolling basis. Key Responsibilities Train and evaluate AI models in medical and clinical contexts. Create tasks, case...Hourly payContract workRemote work$100 per hour
...Overview Geoscientists apply field and interpretive expertise to evaluate AI-generated geology content, design job-related questions, and... ...oil and gas, or energy transition fields Environmental or Engineering Geologists Hydrogeologists or other applied geoscientists...Full timePart timeFor contractorsRemote workFlexible hours$75 per hour
...Role Overview Evaluate and improve how medical AI systems reason about real clinical problems. In... ...clinical reasoning, then work with engineers and researchers to improve model performance. Key... ...scenarios and case vignettes that test decision-making, differential diagnosis...Contract workRemote workFlexible hours- ...Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content... ...Participants enjoy flexible hours and competitive pay, with a brief assessment test prior to joining. #J-18808-Ljbffr...Remote workFlexible hours
$60 per hour
...developing cutting-edge AI systems, while... ...advance AI development. AI models are increasingly capable... ...models on tasks like evaluating AI-generated quantitative... ...design (e.g., A/B testing, hypothesis testing, regression... ...Science, Mathematics, Engineering, or similar); a master...Hourly payFull timeRemote workFlexible hours$40 per hour
...A forward-thinking AI solutions company is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the development of cutting-edge... ...skills. Join us to directly impact the future of AI analytics and model reasoning. #J-18808-Ljbffr...Hourly payRemote workFlexible hours$40 per hour
A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative work and provide critical feedback. This role offers the flexibility of remote work, allowing you to set your own schedule while focusing on impactful projects...Hourly payRemote work$40 per hour
...A leading AI development firm is seeking experienced quantitative professionals to join their remote team. You will evaluate AI-generated quantitative analysis and solve complex problems to ensure technical accuracy. The ideal candidate should have at least 2 years of...Hourly payRemote workFlexible hours$40 per hour
A leading AI development company is seeking experienced quantitative professionals to evaluate and validate AI systems. The role is fully remote, offering flexibility in project... ...-generated work and designing problems for model training, contributing to shaping the future...Hourly payRemote work$40 per hour
...A leading AI company in the United States is seeking experienced quantitative professionals to evaluate and validate AI-generated analytical work. This fully remote position allows you to set your own schedule, with competitive hourly pay starting at $40 USD. Responsibilities...Hourly payRemote work$40 per hour
...A leading AI development firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully remote position allows for a flexible schedule, offering competitive pay starting at $40+ per hour. Ideal candidates...Hourly payRemote workFlexible hours$40 per hour
...A leading AI development firm is seeking experienced quantitative professionals to join their team remotely. The role involves evaluating AI-generated quantitative work, providing insights, and shaping the future of AI systems. Candidates should have over two years of...Hourly payRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to QA/Test Engineer for AI Model Evaluation. Be the first to apply!
- junior qa engineer United States
- software test engineer United States
- qa engineer United States
- entry level qa engineer United States
- qa automation engineer United States
- senior software quality engineer United States
- software quality engineer United States
- qa test engineer United States
- junior software test automation engineer United States
- quality assurance engineer United States



