Humanities Researcher for AI Model Evaluation
$40 - $90 per hourSaidGig
Help improve next-generation AI systems by creating rigorous humanities evaluations grounded in interpretation, argument, context, and analytical reasoning. This remote contractor opportunity relies on your subject-matter expertise, no prior AI experience required. Key Responsibilities
- Create original, high-difficulty question-and-answer pairs that challenge advanced AI models.
- Research, verify, and cite primary texts and reputable sources to support accurate answers and reasoning.
- Develop prompts that require deep analysis or interpretation rather than simple factual recall.
- Test questions with AI models and refine them for difficulty, clarity, and correctness.
- Produce precise, unambiguous questions and well-supported answers.
- Incorporate reviewer feedback and follow project quality standards.
- Advanced academic background or professional experience in history, literature, philosophy, linguistics, classics, art history, journalism, or another humanities discipline.
- Strong research and source-triangulation skills, analytical thinking, close reading, and analytical writing ability.
- Excellent attention to detail, written precision, and ability to document reasoning and sources clearly.
- Experience writing nuanced questions that assess understanding beyond surface-level facts.
- Self-direction and reliability when completing remote work independently.
- Experience evaluating or training AI models is helpful but not required.
- Remote, independent contractor engagement.
- Work is paid by qualifying task, and time needed may vary by experience and workflow.
- Minimum weekly task-submission requirements apply.
$40 to $90 per hour equivalent, with payment based on tasks that meet project specifications.
Availability and Application Process- Open roles are typically filled within 48 hours, so applicants should be ready to begin promptly.
- If selected, complete onboarding and begin initial tasks within 24 to 48 hours.
- ...Role Overview Apply research-grade expertise to help evaluate and improve AI reasoning across technical and humanities disciplines, including STEM, English, literature, and journalism. This remote contract role focuses on producing rigorous training data, evaluations...SuggestedHourly payContract workFor contractorsRemote work
$70 - $110 per hour
...Help advance frontier AI systems by bringing rigorous materials... ...and engineering judgment to the evaluation, design, and improvement of... ...will work closely with an AI research team to define what high quality... ...looks like in practice and ensure model outputs can withstand...SuggestedHourly payFull timeLive inRelocationRelocation package$70 - $90 per hour
...Role Overview Evaluate vulnerability reproduction and remediation tasks used to train and assess frontier AI models. You will determine whether CVE reproductions faithfully reflect... ...penetration testing, or vulnerability research. ~ Strong knowledge of CVE taxonomy and...SuggestedRemote jobHourly pay$262.5k - $299.6k
...Applied Researcher II (AI Foundations) Overview: At Capital One, we are creating... ...of AI & ML are bringing humanity and simplicity to banking.... ...textual data. Build AI foundation models through all phases of... ...from design through training, evaluation, validation, and...SuggestedFull timePart timeLocal areaFlexible hours- ...interpretable, and steerable AI systems. We want AI to... ...group of committed researchers, engineers, policy... ...Anthropic's production models undergo sophisticated post... ...the science of how we evaluate production training runs... ...Safety, and Learning from Human Preferences. Come...SuggestedFull timeWork at officeVisa sponsorshipFlexible hours
£50 per hour
...background in Psychology for completing AI training tasks. This role involves analyzing and improving AI models using psychological theories. Participants... ...reliable internet access. Join Prolific to help shape AI research in psychology and human behavior. #J-18808-Ljbffr...$50 per hour
...Overview This role focuses on improving and evaluating large language models through advanced mathematical reasoning,... ...checks. The work supports both frontier research efforts and the application of those advances into reliable AI systems for enterprise use. Key...Contract workFor contractorsFreelanceRemote work- ...reasoning and computational problem solving to improve and evaluate large language models. You will design rigorous math problems, produce clear,... ...How this work supports customers Accelerate frontier AI research by contributing high quality data and evaluation...Contract workFor contractorsFreelanceRemote work
$60 - $80 per hour
...considered for future contract opportunities with AI labs and companies. This is an open... ...Role Overview Qualified experts may support AI research by applying mathematics knowledge to real-world tasks, model evaluation, and domain-specific feedback. Key Responsibilities...Hourly payContract workRemote work$60 - $80 per hour
...Role Overview Provide high-level mathematical expertise to support AI research and product development. Mathematicians in this expert network train and evaluate mathematical models, design realistic problem tasks and deliverables, and give domain-specific feedback that...Hourly payContract workImmediate startRemote work$60 - $80 per hour
...Apply deep retail domain expertise to help build and evaluate foundational generative AI models. You will design realistic retail tasks, produce authoritative... ...reasoning, and practical judgment while working with research and engineering teams at a leading GenAI lab. Key...Hourly payWeekday work$218.7k - $249.6k
...Overview Applied Researcher 4 (AI Foundations, LLM Core and Agentic AI... ...applications of AI & ML are bringing humanity and simplicity to banking.... .... Build AI foundation models through all phases of... ...from design through training, evaluation, validation, and implementation...Full timePart timeLocal areaFlexible hours$218.7k - $249.6k
...Overview Applied Researcher 4 (AI Foundations - LLM, Optimization and... ...applications of AI & ML are bringing humanity and simplicity to banking.... .... Build AI foundation models through all phases of... ...from design through training, evaluation, validation, and implementation...Full timePart timeLocal areaFlexible hours$218.7k - $249.6k
...Overview Applied Researcher 4 (AI Foundations, Graph and Behavioral Models) At Capital One, we are creating trustworthy... ...of AI & ML are bringing humanity and simplicity to banking. We are... ..., from design through training, evaluation, validation, and implementation....Full timePart timeLocal areaFlexible hours$120 per hour
...Apply your survey research expertise to help evaluate and improve sophisticated AI-generated survey research. This role focuses on distinguishing rigorous, decision-useful surveys from designs that may introduce noise, bias, or misleading results. Key Responsibilities...Hourly payRemote work- ...advanced offensive security expertise to evaluate and validate AI-generated security analyses, exploit... ...reasoning, and vulnerability research across software, operating systems, networking... ...credentials. Why Join Work on frontier AI models addressing real-world cybersecurity...Hourly payRemote workVisa sponsorshipWork visa
$65 - $105 per hour
...Help advance frontier AI models by bringing rigorous life sciences research judgment into task design, evaluation, and model improvement. You will work closely with an AI lab''s research and program management teams to ensure models can reason credibly about real scientific...Hourly payFull timeFreelanceLive inRelocationRelocation package$65 - $105 per hour
...Help improve how frontier AI models reason about real-world life sciences research. In this senior, hands-on role, you will apply deep scientific judgment to evaluate research tasks and model outputs, define rigorous standards for correct answers, and build benchmarks...Hourly payFull timeLive inRelocationRelocation package$75 - $115 per hour
...Role Overview Help advance frontier AI models by bringing rigorous pharmaceutical research and development judgment to the evaluation, design, and improvement of domain-specific knowledge work. You will work closely with AI research and program management teams to define...Hourly payFull timeContract workLive inRelocationRelocation package$100 - $150 per hour
...Role Overview Provide senior legal subject-matter expertise to a GenAI research team, creating authoritative instruction specifications, golden solutions, and evaluation benchmarks so frontier AI models reason correctly about real legal work. This role centers on hands-on...Hourly payFull timeFreelanceInternshipLive inLocal areaRelocationRelocation package$70 - $90 per hour
...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness, CUDA-to-NKI migration fidelity, and whether implementations are well suited to...Remote jobHourly pay$20 - $60 per hour
...Help train next-generation AI systems by creating rigorous, real-world evaluations that test how well advanced models learn, reason, and perform. This remote contract opportunity... ...from any background with strong research and writing expertise. Prior AI experience...Hourly payContract workFor contractorsRemote work$100 per hour
...your finance expertise to help improve AI models across complex financial problem-solving... ...capital markets, portfolio management, research, trading, quantitative finance, investment... ...is required. Key Responsibilities Evaluate language models in finance domains where...Hourly payContract workRemote work10 hours per weekFlexible hours- ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets...For contractorsRemote work
$70 - $90 per hour
...Review GPU and accelerator kernel development tasks that support the training and evaluation of advanced AI models. This role focuses on assessing task quality, numerical correctness, completeness, fair performance benchmarking, appropriate scope, and whether kernels...Remote jobHourly pay$60 - $80 per hour
...marketing subject-matter expertise to a GenAI team building foundational AI models. You will create realistic marketing tasks, evaluate model outputs against structured rubrics, and advise research and engineering teams on brand strategy, growth marketing, and campaign-...Hourly payWeekday work$60 - $70 per hour
...alignment, and overall quality of frontier AI model outputs on complex, policy sensitive,... ...area" topics. Work through structured evaluations to identify unsafe behavior, reasoning... ...safety performance. Collaborate with AI researchers and safety teams on ongoing evaluation...Hourly payRemote work$60 - $80 per hour
...help develop advanced large language models. In this role, you will bring... ...growth, and campaign judgment to AI training data, partnering with research and engineering teams to improve model... ...quality. Develop and improve evaluation guidelines and scoring rubrics for...Hourly payWeekday work$50 per hour
...Role Overview Improve the performance of large language models on real-world finance tasks by evaluating model outputs, designing assessment rubrics, and working directly with AI researchers to shape training and evaluation. This role focuses on applying deep domain...Hourly payContract workRemote work10 hours per weekFlexible hours$100 - $150 per hour
...Overview Apply senior legal judgment to help a leading AI research organization improve how advanced AI models reason about real-world legal work. You will work... ...define high-quality legal tasks, standards, and evaluations. This role is designed for a practicing legal...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Humanities Researcher for AI Model Evaluation. Be the first to apply!
- title researcher United States
- remote researcher United States
- data collection researcher United States
- home based internet researcher United States
- music researcher United States
- design researcher United States
- researcher United States
- criminal researcher United States
- qualitative researcher United States
- online researcher United States



