QA Engineer for AI Model Evaluation
$90 - $175 per hourSaidGig
Apply software quality assurance expertise to evaluate technical AI outputs and help improve how next-generation AI systems learn, reason, and perform. This remote contract opportunity supports a high-volume project and does not require prior AI experience, but it does require prior paid human-in-the-loop AI evaluation work.
Key Responsibilities- Evaluate and score AI-generated technical outputs against established quality criteria and rubrics, identifying responses that appear correct but are inaccurate or incomplete.
- Design and assess thorough functional, regression, edge-case, negative, and boundary test cases.
- Review bug reports and test documentation for reproducibility, completeness, and appropriate severity assessment.
- Identify, isolate, and document defects with precise reproduction steps using structured tracking methods.
- Provide detailed written feedback and annotations that developers can act on without further clarification.
- Work with project teams to refine evaluation guidelines and continuously improve testing standards and methods.
- Professional experience as a QA Engineer, SDET, Test Engineer, QA Analyst, or in a comparable software quality assurance role.
- Strong knowledge of quality assurance, software testing, test case design, bug tracking, regression testing, and manual and automated testing strategies.
- Hands-on experience with automation frameworks and test management tools, such as Selenium, Playwright, Cypress, Appium, Postman, Jira, TestRail, Zephyr, BrowserStack, or similar tools.
- Prior paid experience supporting AI training through human data annotation, labeling, RLHF, AI response or model evaluation, or rubric-based grading. Software QA experience alone does not meet this project requirement.
- Excellent analytical and problem-solving skills, close attention to detail, and the ability to communicate complex findings clearly in written English at B2 level or above.
- No formal degree is required. Demonstrable practical testing experience is prioritized.
- Reliable internet access and readiness to begin promptly.
- Remote contract engagement.
- $90 to $175 per hour.
Apply through the available application flow using email or a Google account. Continuing the application requires agreement to the applicable terms and privacy policies.
$60 - $90 per hour
...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Summers , and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract...SuggestedFull timeContract workSummer workRemote work$60 per hour
...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...SuggestedHourly payRemote workWork from homeFlexible hours$60 per hour
...and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with... ...offers a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing...SuggestedHourly payRemote workWork from homeFlexible hours$60 - $80 per hour
...develop advanced large language models. In this role, you will bring practical... ..., and campaign judgment to AI training data, partnering with research and engineering teams to improve model... ...quality. Develop and improve evaluation guidelines and scoring rubrics for...SuggestedHourly payWeekday work$100 - $150 per hour
...Overview Apply senior legal judgment to help a leading AI research organization improve how advanced AI models reason about real-world legal work. You will work... ...define high-quality legal tasks, standards, and evaluations. This role is designed for a practicing legal...SuggestedHourly payFull timeFreelanceInternshipLive inRelocationRelocation package$70 - $90 per hour
...kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical... ...including NKI, Pallas, or TPU. Background in compiler engineering, MLIR, or intermediate-representation lowering. Knowledge...Hourly payRemote work- ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets...For contractorsRemote work
$70 - $80 per hour
...Role Overview Apply advanced drug safety expertise to help improve AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data. This remote contract role focuses on the quality, accuracy, and regulatory alignment of complex safety reports...Hourly payContract workRemote work$70 - $90 per hour
...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness... ...NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth...Hourly payRemote work$100 per hour
...expertise to improve the performance of large language models on finance tasks. You will work with AI researchers to identify model weaknesses in areas... ...focused on advanced AI systems. Key Responsibilities Evaluate LLM performance in finance areas where models...Hourly payContract workFor contractorsFreelanceRemote work10 hours per weekFlexible hours- ...emerging trillion-dollar Voice AI economy, providing real-time... ...’s voice-native foundation models are accessed through cloud... ...looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the... ...and strong engineering and QA practices. What We're Looking...Full time
- ...computational problem solving to improve and evaluate large language models. You will design rigorous math... ...customers Accelerate frontier AI research by contributing high quality... ...mathematics at the level expected for engineering entrance exams and for graduate or PhD...Contract workFor contractorsFreelanceRemote work
- ...reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial... ...growing group of committed researchers, engineers, policy experts, and business leaders... ...for Research Engineers to build the evaluations that tell us — and the world — what Claude...Full time
$100 - $150 per hour
...will be considered for future projects evaluating how well AI systems perform real-world data... ...assess AI or human-produced analyses and models, document decisions in writing, and iterate... ...and A/B test write-ups, feature engineering, and technical reports or notebooks...Hourly payImmediate startRemote work$60 - $80 per hour
...expertise to help develop advanced large language models by bringing real-world underwriting, claims, and risk-assessment judgment to AI training and evaluation work. Key Responsibilities Partner with research and engineering teams to address knowledge gaps in...Hourly payWeekday work$36 - $72 per hour
...seekers, 1 million+ employers, and 1,600 educational institutions. Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise. Role Details Location: Onsite in Seattle, WA,...Hourly payFull timeMonday to FridayFlexible hours$60 - $80 per hour
...mathematics experts considered for future contract opportunities with AI labs and companies. This is an open application, not a posting... ...by applying mathematics knowledge to real-world tasks, model evaluation, and domain-specific feedback. Key Responsibilities...Hourly payContract workRemote work$100 per hour
...Apply consulting expertise to evaluate and improve AI-generated business content for a customer-facing project. Your judgment will help AI systems... ...executive summaries. Develop and refine large language model prompts using structured problem-solving and analytical rigor...Hourly payPart timeFor contractorsRemote work$110 per hour
...Apply to join a physician talent network supporting AI labs and companies with medical expertise. This is an open application... ...projects by bringing real-world clinical expertise to model development and evaluation. Key Responsibilities Train and evaluate AI models...Hourly payContract workRemote work$65 - $105 per hour
...Apply deep engineering judgment to help frontier AI models reason more accurately about real-world software development and engineered systems. You will... ...management team to define high-quality engineering work, evaluate model performance, and turn expert practice into...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package$60 - $90 per hour
...Role Overview Design rigorous, real-world materials science and engineering tasks that evaluate an AI model’s expert reasoning. You will create prompts, supporting data rooms, and grading criteria, then test and refine each task until it clearly distinguishes strong from...Hourly payWork at officeRemote work$65 - $105 per hour
...Help advance frontier AI models by bringing rigorous life sciences research judgment into task design, evaluation, and model improvement. You will work closely with an AI lab''s research and program management teams to ensure models can reason credibly about real scientific...Hourly payFull timeFreelanceLive inRelocationRelocation package- ...Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced QA and Test Engineers to ensure every benchmark is reliable, reproducible, and accurately measures real AI capabilities...Full timeContract workFor contractorsRemote workFlexible hours
$100 - $150 per hour
...Apply senior finance expertise to improve how frontier AI systems reason through real-world financial work. You will partner closely... ...with an AI research team to define high-quality finance tasks, evaluate model performance, and translate professional judgment into rigorous...Hourly payFull timeLive inRelocationRelocation package$70 - $110 per hour
...Role Overview Shape how advanced AI systems reason about real clinical work. In this... ...will partner with an AI research team to evaluate medical knowledge tasks, define high-quality... ...develop benchmarks that measure meaningful model improvement. Key Responsibilities Review...Hourly payFull timeLive inRelocationRelocation package- ...Apply your dermatology expertise to image-based work that helps develop and evaluate advanced AI models for clinical image interpretation. This is a non-clinical, remote opportunity with no direct patient care, focused on ensuring AI-generated medical outputs reflect real...Hourly payRemote work
$60 - $90 per hour
...Role Overview Help shape how an advanced performance-transfer model evaluates character animation, preserving an actor''s timing, emotion,... ...evaluation methods, and help build a reliable human-review process for AI-generated performance results. Key Responsibilities...Hourly payPart time$20 - $36 per hour
...Role Overview Evaluate AI-generated music in Slovak and English by listening across genres... ...labels will help improve generative music models. Key Responsibilities Compare... ...professional experience as a music producer, audio engineer, or mixing engineer Headphones or...Hourly payImmediate startRemote workFlexible hours$60 - $80 per hour
...Apply deep retail domain expertise to help build and evaluate foundational generative AI models. You will design realistic retail tasks, produce authoritative... ...practical judgment while working with research and engineering teams at a leading GenAI lab. Key Responsibilities...Hourly payWeekday work$100 - $150 per hour
...matter expertise to a GenAI research team, creating authoritative instruction specifications, golden solutions, and evaluation benchmarks so frontier AI models reason correctly about real legal work. This role centers on hands-on legal judgment, translating professional...Hourly payFull timeFreelanceInternshipLive inLocal areaRelocationRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to QA Engineer for AI Model Evaluation. Be the first to apply!
- entry level qa engineer United States
- software test engineer United States
- junior software test automation engineer United States
- qa engineer United States
- staff qa engineer United States
- senior software test automation engineer United States
- sdet qa automation engineer United States
- senior software quality engineer United States
- quality assurance engineer United States
- senior quality assurance engineer United States




