QA/Test Engineer for AI Model Evaluation
$60 - $90 per hourSaidGig
Validate the integrity of complex, multi-step evaluation tasks used to benchmark frontier generative AI models. You will design and execute rigorous test cases, probe edge cases, and debug task environments so each task is unambiguous, correctly graded, and robust to shortcuts. Individual tasks typically represent one to two days of expert effort and span multiple technical skills. You will operate in a close feedback loop with the lab''s researchers and task authors to ensure benchmark results remain trustworthy.
Key Responsibilities- Design checks and test cases that confirm each task behaves as intended, including difficult edge cases.
- Review tasks and reference solutions in detail before finalization, identifying ambiguity, grading gaps, and missing assumptions.
- Debug task logic and verification code, actively using Python to investigate and fix failures.
- Help create simple, repeatable quality checklists and provide actionable feedback that task authors can apply quickly.
- Monitor AI agent runs for shortcuts or grading weaknesses to protect benchmark validity.
- Collaborate closely with researchers and task authors in an iterative feedback loop to improve task quality and evaluation processes.
- MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy or engineering-heavy domain.
- At least 1 year of experience in test engineering, quality assurance, or a research or software engineering role that included strong quality ownership.
- Proven ability to design test cases and quality-review processes, and to debug complex systems end-to-end.
- Working proficiency in Python and Git, with comfort navigating unfamiliar codebases and runtime environments.
- Exceptional attention to detail and clear written documentation habits.
- Prior experience with AI training, model evaluation, or quality review of AI-generated outputs is preferred.
- A perfectionist mindset, creativity in finding what others missed, and the ability to work independently on ambiguous, open-ended problems.
- Availability to engage reliably for approximately 35 hours per week.
- Status: Full-time W-2 employment with Cincinnatus LLC, acting as the employer of record for this placement.
- Placement: Opportunity to be placed at a leading AI lab as part of their extended workforce, working within the client team and enterprise workflows.
- Location: Fully remote within the United States, candidates must be able to work from the U.S.
- Schedule: Approximately 35 hours per week.
- Engagement type: Role-based employment, not a freelance or project-by-project arrangement; integration with client teams and standard enterprise processes is expected.
- Pay rate: $60.00 to $90.00 per hour, paid on an hourly basis.
- Employment, onboarding, payroll, and benefits for this role are administered by Cincinnatus LLC, the employer of record.
- Opportunities may be discovered through third-party talent platforms, but final hiring and employment administration are handled by Cincinnatus LLC.
- Cincinnatus LLC is an Equal Employment Opportunity employer and provides reasonable accommodations for qualified individuals with disabilities throughout the application process.
- ...Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced QA and Test Engineers to ensure every benchmark is reliable, reproducible, and accurately measures real AI capabilities...SuggestedFull timeContract workFor contractorsRemote workFlexible hours
- ...Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced QA and Test Engineers to ensure every benchmark is reliable, reproducible, and accurately measures real AI capabilities...SuggestedFull timeContract workFor contractorsRemote workFlexible hours
$90 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ..., and Jack Dorsey . Position: QA/Test Engineer Type: Contract Compensation:... ...Preferred Experience in AI training , model evaluation, or quality review of AI-generated...SuggestedContract workSummer workRemote work$146k - $194k
...expertise, technology, and business model of the 21st century’s most... ...is powered by Lattice OS, an AI-powered operating system that... ...skilled and experienced Test & Evaluation Manager who is passionate about... ...for a highly motivated test engineer with emphasis in developmental...SuggestedFull timeWork experience placementImmediate startRemote work$40 per hour
A biotechnology company is seeking a Biotechnology R&D Scientist to train AI models and evaluate their outputs. This role involves measuring the progress of AI chatbots with complex biology questions and ensuring their performance and correctness. Candidates should have...SuggestedHourly payRemote workFlexible hours- ...Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content... ...Participants enjoy flexible hours and competitive pay, with a brief assessment test prior to joining. #J-18808-Ljbffr...Remote workFlexible hours
$224k - $356.5k
...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our... ...performance computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in crafting...Full time$45 - $52 per hour
DescriptionKforce has a client seeking a QA Test Engineer in Boca Raton, FL to support web application testing initiatives. This role is responsible... ...to CI/CD processes and continuous testing practices* Leverage AI-assisted tools to improve testing efficiencyRequirements*...$40 per hour
...A leading AI development firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully remote position allows for a flexible schedule, offering competitive pay starting at $40+ per hour. Ideal candidates...Hourly payRemote workFlexible hours$60 per hour
...developing cutting-edge AI systems, while... ...advance AI development. AI models are increasingly capable... ...models on tasks like evaluating AI-generated quantitative... ...design (e.g., A/B testing, hypothesis testing, regression... ...Science, Mathematics, Engineering, or similar); a master...Hourly payFull timeRemote workFlexible hours$40 per hour
A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting...Hourly payRemote work$40 per hour
...A technology company in Mississippi seeks a Full Stack Engineer to improve AI models by providing coding challenges and evaluating performance. Candidates should be proficient in a programming language and fluent in English. This remote position offers flexibility in project...Hourly payContract workRemote work$40 per hour
...A forward-thinking AI solutions company is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the development of cutting-edge... ...skills. Join us to directly impact the future of AI analytics and model reasoning. #J-18808-Ljbffr...Hourly payRemote workFlexible hours$40 per hour
A leading AI training company is seeking a Biotechnology R&D Scientist to evaluate and improve AI models by posing complex biological questions. This remote position allows candidates to choose projects and work on their own schedule, with rates starting at $40+ per hour...Hourly payRemote work$208k - $300k
...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines... ...LLM-judge–based evaluations. Design test datasets and benchmarks to measure generalization...Full time$40 per hour
...A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated quantitative work and design problems for AI training. Enjoy the flexibility of fully remote work with a competitive pay starting at $40+ per hour. Ideal candidates...Hourly payRemote work$40 per hour
...DataAnnotation is seeking a Biotechnology R&D Scientist to train AI models. In this role, you will evaluate the outputs of AI chatbots and assess their logic to improve model quality. The ideal candidate should have a deep understanding of cell biology, genetics, biochemistry...Hourly payFor contractorsRemote work$60 per hour
...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific... ...$60 per hour, and may work from home with flexible hours. A brief assessment test is required prior to joining. #J-18808-Ljbffr...Hourly payRemote workWork from homeFlexible hours$50 - $60 per hour
...DataAnnotation is seeking an Appellate Attorney to train AI models by evaluating their legal outputs and solving complex legal challenges. This role allows for flexible remote work, enabling you to choose projects based on your own schedule and preferences. A J.D. is mandatory...Hourly payRemote workFlexible hours$40 per hour
...seeking an R&D Biologist to join their team in the United States. In this remote role, you will train AI models by providing complex biology questions and evaluating chatbot responses. The ideal candidate will have an expert understanding of biology and related fields,...Hourly payContract workRemote workFlexible hours$40 per hour
...Development Chemist to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the... ...are not limited to: Chemistry and/or Chemical Engineering. Benefits This is a full-time or part-time...Hourly payFull timeContract workPart timeRemote work$40 per hour
A growing technology company is seeking an R&D Biologist to join their team. This role involves training AI models by evaluating chatbot outputs on complex biology questions. Applicants should possess strong expertise in biology and related fields. You will work remotely...Hourly payRemote work$40 per hour
A leading AI training firm in the United States is seeking an R&D Biologist to join their team. In this remote position, you will evaluate AI chatbots and enhance their models while ensuring the biological accuracy of their outputs. The ideal candidate should have an expert...Hourly payRemote work$40 per hour
A technology company in Massachusetts is seeking an R&D Biologist to join their team to train AI models by evaluating chatbot outputs against complex biology questions. Ideal candidates will hold advanced qualifications in biology or biochemistry. This position allows full...Hourly payFull timePart timeRemote work$40 per hour
...company in the United States is seeking an R&D Biologist to train AI models and improve their quality. This position offers remote... ...selected projects at your own schedule. Responsibilities include evaluating the performance of AI chatbots on complex biology topics. Candidates...Hourly payRemote work$40 per hour
A tech company specializing in AI training is looking for a Statistician to join their team. In this remote role, you'll train AI models by providing complex math problems and evaluating their outputs for quality and correctness. The ideal candidate will have strong mathematical...Hourly payRemote workFlexible hours$50 - $60 per hour
...DataAnnotation is looking for an Appellate Attorney to help train AI models by providing complex legal problems and evaluating AI outputs. This role is ideal for General Counsel or those with similar legal expertise. Contract details include working on your own schedule...Hourly payContract workRemote work$40 per hour
A technology-focused company is seeking a Postdoctoral Researcher in Chemistry to evaluate AI chatbots based on complex chemistry questions. This remote position allows for flexible scheduling and project selection. Applicants must have a PhD in Chemistry and a strong command...Hourly payRemote workFlexible hours$40 per hour
A technology firm specializing in AI is seeking a Biostatistician to enhance AI models by evaluating their performance and solving complex mathematical problems. This role offers flexibility as a remote position, allowing you to choose your projects and work according...Hourly payRemote work$40 per hour
A healthcare technology firm is seeking medical experts to evaluate AI chatbots' performance and ensure their medical accuracy. This position allows for flexible scheduling and project selection, making it suitable for both full-time and part-time professionals. Candidates...Hourly payFull timePart timeRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to QA/Test Engineer for AI Model Evaluation. Be the first to apply!
- mobile qa automation engineer United States
- qa engineer remote United States
- qa engineer United States
- qa engineer python United States
- senior software test engineer United States
- qa automation engineer United States
- qa automation test engineer United States
- salesforce qa engineer United States
- qa engineer part time United States
- quality assurance engineer United States



