Technology AI Evaluation Expert
$60 - $75 per hourWeekday AI
Role Description
Join an advanced AI research initiative focused on improving how next-generation AI systems understand professional documents, execute complex instructions, and reason through real-world technical workflows. We are seeking experienced technology professionals to design high-quality benchmark tasks that evaluate AI performance across software engineering and data science domains.
In this role, you will create realistic, multi-step evaluation tasks based on technical documentation, code repositories, API references, architecture diagrams, and other workplace resources. Your work will help measure and improve the ability of AI models to interpret technical information, follow detailed instructions, and generate accurate, well-structured outputs.
This is a fully remote, independent contractor opportunity with flexible working hours.
Key Responsibilities
- Design AI Evaluation Tasks
- Create realistic, multi-step benchmark tasks based on professional technology workflows.
- Develop challenges using technical specifications, architecture documents, API documentation, codebases, web research, and code execution.
- Ensure each task includes a clearly defined expected output and objective evaluation criteria.
- Develop Evaluation Standards
- Write comprehensive ground-truth solutions and structured scoring rubrics.
- Design tasks that assess reasoning, technical understanding, instruction following, and output quality.
- Maintain high standards of technical accuracy, clarity, and reproducibility.
- Contribute Domain Expertise
- Apply real-world knowledge from software engineering, data science, or analytics to create authentic evaluation scenarios.
- Collaborate with research teams to improve benchmark quality and consistency.
- Continuously refine tasks based on project feedback and evolving evaluation requirements.
Qualifications
- Minimum 3 years of hands-on professional experience in one or more of the following areas:
- Software Engineering
- Data Science
- Data Analytics
- Strong understanding of technical documentation, software development workflows, and engineering best practices.
- Experience working with codebases, APIs, technical specifications, or system architecture documentation.
- Excellent analytical thinking and problem-solving skills.
- Strong written communication with the ability to create clear technical instructions and evaluation criteria.
- Ability to work independently while maintaining high standards of accuracy and consistency.
Engagement Details
- Independent contractor engagement.
- Fully remote with flexible working hours.
- Expected commitment of 15–20 hours per week.
- Projects may be extended, shortened, or concluded based on business needs and performance.
- Weekly payments processed through supported payment platforms.
Why Join
- Help shape the next generation of AI systems for technical reasoning and document understanding.
- Work on intellectually challenging projects involving real-world engineering and data science workflows.
- Apply your technical expertise to improve advanced AI evaluation benchmarks.
- Enjoy flexible remote work with meaningful impact on AI research.
Equal Opportunity Statement
We are committed to providing equal opportunities to all qualified applicants without regard to legally protected characteristics. Reasonable accommodations are available upon request.
Contract Information
- Independent contractor engagement.
- Fully remote work completed on your own schedule.
- Weekly payments are processed based on approved work completed.
- Work does not involve access to confidential or proprietary information from any employer, client, or institution.
- Please note that visa sponsorship is not available for this opportunity.
- A tech company focusing on AI research is looking for experienced Krita users for a flexible, project-based contract opportunity. This role allows you to earn while evaluating AI-generated content related to digital painting and concept art. Candidates should have at least...SuggestedRemote jobContract workFlexible hours
- ...driven Geopolitical and Event Forecasting Professionals to help evaluate AI-generated predictions about real-world events. The estimated... ...elections, central bank decisions, policy announcements, and technology/regulatory events using a provided rubric. ~Project...SuggestedContract workFreelance
$150 per hour
...Overview Join Turing as a Physics Expert (PhD, Remote) and help shape how the world’s most advanced AI models reason about physics.... ...challenges across science, technology, and industry. Role Details... ...design challenging problems, evaluate model outputs, and provide...SuggestedHourly payContract workFreelanceRemote work10 hours per weekFlexible hours$70 - $90 per hour
...content for security vulnerabilities to help AI models recognize and classify threats.... ...with a distributed team of domain experts to refine detection and reasoning approaches... ...a short interview and a questionnaire to evaluate domain expertise. If hired, onboarding...SuggestedHourly payContract workTemporary workRemote work$60 per hour
A technology company is seeking experienced quantitative professionals to contribute to the development of AI systems while enjoying fully remote work. You will evaluate AI-generated quantitative analyses, design quantitative problems for AI training, and provide feedback...SuggestedHourly payRemote workFlexible hours- ...A leading AI research firm is seeking Expert Prompt Curators to design challenging prompts for evaluating advanced AI models. The role requires advanced knowledge in diverse fields and offers flexible hours, remote work, and a competitive hourly wage. Ideal candidates...Hourly payTemporary workRemote workFlexible hours
- ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote workFlexible hours
$60 per hour
...Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Hourly payRemote workWork from homeFlexible hours$60 per hour
...Prolific is seeking Biology Experts and Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with flexible hours and offers a competitive pay rate of up to $60 per hour for reviewing model responses...Hourly payRemote workWork from homeFlexible hours$60 per hour
A technology company is seeking quantitative professionals to evaluate AI-generated analyses and contribute to the development of AI systems. This role offers the flexibility of remote work and up to $60 per hour, allowing you to choose your projects and schedule. Candidates...Remote jobHourly pay$60 per hour
A leading AI technology company is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the advancement of cutting-edge AI systems. This fully remote role allows you to set your own schedule and offers competitive pay, reaching...Remote jobHourly pay- ...Prolific is seeking Biology Experts and Life Science Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities...Hourly payRemote workFlexible hours
$30 per hour
Prolific is seeking fluent Hindi speakers to join their Expert Network, helping to train and evaluate AI models with real legal expertise. Responsibilities include analyzing and writing tasks in Hindi, judging AI’s performance, and aiding in the improvement of AI models...Remote jobHourly payFlexible hours$80 - $120 per hour
Role Description ~Evaluate AI-generated artifacts against domain-specific quality rubrics. ~Identify factual, aesthetic, and presentation errors in documents, spreadsheets, and slide decks. ~Provide clear, structured written feedback to improve AI model outputs....Part timeWork at officeRemote work$80 - $120 per hour
Role Description ~Evaluate AI-generated artifacts against domain-specific quality rubrics. ~Identify factual, aesthetic, and presentation errors in documents, spreadsheets, and slide decks. ~Provide clear, structured written feedback to improve AI model outputs....Part timeWork at officeRemote work$50 - $60 per hour
A healthcare technology company is seeking a Clinical Reviewer to enhance AI models through comprehensive evaluation. The ideal candidate must have a medical degree and fluency in English. The role allows for flexible hours and is available as both full-time and part-time...Hourly payFull timePart timeRemote workFlexible hours$50 - $60 per hour
A leading AI technology firm is seeking a Private Banker to assist in training AI models with financial expertise. This remote position... ...or PhD in a finance-related field. Responsibilities include evaluating AI outputs and improving financial reasoning. Competitive pay...Remote jobHourly payFlexible hours$50 - $60 per hour
A technology firm specializing in AI is seeking a Private Banker to provide financial expertise in training AI models. This role is fully remote and... ...analysis and reasoning. Responsibilities include evaluating AI outputs and improving accuracy in financial contexts....Remote jobHourly payFlexible hours$50 - $60 per hour
A technology firm focused on AI is seeking a Private Banker in the United States to contribute to training AI models using financial expertise.... ...remote position allows you to work on flexible projects, evaluating AI outputs related to finance. Candidates with a Master's...Remote jobHourly payFlexible hours- A leading AI technology firm is looking for a Private Banker to help train AI models in finance. This role offers flexibility and allows... ...schedule, whether part-time or full-time. Responsibilities include evaluating AI outputs and providing structured feedback on financial...Remote jobFull timePart time
$40 per hour
DataAnnotation is looking for an experienced Biology Teacher to assist in training AI models. You will evaluate complex biology questions posed to AI chatbots and measure the accuracy of their outputs. Candidates should have a strong understanding of cell biology, genetics...Remote jobHourly payFor contractors$30 per hour
...About Prolific Prolific is not just another player in the AI space – we are building the biggest pool of quality human... ...Graphic and Visual Designers to act as Domain Experts for a high-level AI evaluation project. AI models are evolving beyond simple image generation...Remote workWork from homeFlexible hours$40 per hour
A healthcare technology company in the United States seeks medical experts to evaluate AI chatbots. Responsibilities include presenting healthcare problems to AI and assessing their responses for correctness. Candidates must be fluent in English and possess a current or...Remote jobHourly pay$8 - $65 per unit
Prolific is seeking Mental Health Professionals to help train AI models by reviewing AI-generated responses and providing expertise. Participants will complete paid tasks and get paid between $8 and $65 per task. The role requires verified professional status and a solid...Remote workWork from homeFlexible hours- ...You will review AI-generated Hebrew and English responses and/or generate high-quality bilingual training content, evaluating reasoning quality and step-by-step problem‑solving while providing expert feedback that helps models produce answers that are accurate, logical...Hourly payFor contractorsRemote workFlexible hours
$80 per hour
...professionals apply their pricing, demand forecasting, inventory planning, and supply chain operations expertise to evaluate AI-generated outputs and produce expert training data. This contract role supports AI research by assessing the accuracy and relevance of model...Hourly payContract workPart timeRemote workFlexible hours$30 per hour
Prolific is looking for Fluent Urdu Speakers to join our Expert Network to help train AI models using real legal expertise. The role involves completing tasks that require one hour of uninterrupted work, with pay rates of up to $30/hr. Candidates must possess advanced Urdu...Remote jobWork from homeFlexible hours$100k - $160k
...infrastructure resiliency, contact center operations, information technology, software engineering, program management, strategic... ...Overview Pantheon Data is seeking a Flight Test & Evaluation Subject Matter Expert who will provide expertise for aircraft and systems...Work at officeLocal areaRemote work$90 - $110 per hour
...This role focuses on creating a benchmark dataset aimed at evaluating AI models for professional document understanding and instruction following specifically within the Engineering & Built Environment domain. You will engage in complex, multi-step tasks that are grounded...Hourly payRemote work- ...Senior Air Warfare Test and Evaluation / Threat M&S / HWIL / Over-the-Air Test Subject Matter Expert Belong. Connect. Grow. with... ...-end engineering and advanced technology solutions to our customers in... ...unmanned, Artificial Intelligence (AI) -enabled, or electronic...Full timePart timeWork at officeLocal areaRemote workWork from home
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Technology AI Evaluation Expert. Be the first to apply!




