AI Evaluation Researcher
Weekday
Join a pioneering AI initiative focused on developing the next generation of evaluation benchmarks for frontier AI models.
We are seeking researchers from computational STEM disciplines—as well as computationally intensive social sciences and humanities—to bring the rigor of real-world research into AI evaluation.
In this role, you will transform scientific methodologies such as experimental design, hypothesis testing, and data-driven analysis into sophisticated, multi-step benchmark tasks that challenge state-of-the-art AI systems. Working closely with AI researchers, you'll help uncover subtle reasoning errors and methodological flaws that only experienced researchers can identify.
This is a fully remote, full-time engagement requiring approximately 35 hours per week .
Requirements
Key Responsibilities
- Design complex, research-oriented benchmark tasks inspired by real-world scientific workflows, including study design, experimentation, hypothesis testing, and data analysis.
- Develop comprehensive reference solutions using Python , notebooks, and computational tools with the rigor expected in professional research.
- Define clear evaluation standards that distinguish sound scientific reasoning from plausible but incorrect conclusions.
- Review AI-generated solutions, identifying methodological weaknesses, analytical errors, and flawed reasoning that experienced researchers would recognize immediately.
- Collaborate with AI researchers and fellow domain experts to improve benchmark quality, consistency, and scientific rigor.
- Contribute to the continuous refinement of evaluation methodologies for advanced AI systems.
Required Qualifications
- Master's degree, PhD, or equivalent practical experience in a STEM discipline , computational social science, computational humanities, or another research-intensive field involving programming and data analysis.
- Minimum 1 year of experience in an active research role within academia, industry, government laboratories, or a similar research environment.
- Demonstrated experience performing computational research involving Python , data analysis, simulation, modeling, machine learning, or scientific computing.
- Strong understanding of experimental design, hypothesis testing, statistical analysis, and rigorous interpretation of research findings.
- Working knowledge of Git , integrated development environments (IDEs), and notebook platforms such as Jupyter or Google Colab .
- Experience with AI evaluation, benchmark development, AI training, or task authoring is preferred.
- Excellent analytical thinking, attention to detail, creativity, and the ability to solve complex, open-ended problems independently.
- Strong written communication skills for documenting technical methodologies and research findings.
- Ability to commit approximately 35 hours per week on a consistent basis.
Preferred Qualifications
- Experience designing reproducible computational experiments or research workflows.
- Familiarity with machine learning, large language models, or AI-assisted research tools.
- Background in benchmark design, scientific software development, or computational research infrastructure.
- Experience mentoring researchers, reviewing scientific work, or contributing to peer-reviewed publications.
Why Join
- Help shape how next-generation AI systems are evaluated using rigorous scientific methodologies.
- Collaborate with leading AI researchers working on frontier models and advanced evaluation frameworks.
- Apply your research expertise to improve AI reasoning, reliability, and scientific accuracy.
- Contribute to impactful work that advances the quality and robustness of AI systems across multiple disciplines.
- Enjoy the flexibility of a fully remote engagement while working on cutting-edge AI research initiatives.
Equal Opportunity
We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.
Contract & Engagement Details
- Independent contractor engagement.
- Fully remote with flexible working hours.
- Expected commitment of approximately 35 hours per week .
- Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
- Work does not require access to confidential or proprietary information from any current or former employer.
- Payments are issued weekly based on approved work completed.
- At this time, we are unable to support H1-B or STEM OPT candidates.
$400k
...Join a dynamic research team as a Member of Technical Staff (MTS) focused on Medical & Health... ...will play a pivotal role in advancing AI systems designed to enhance healthcare, clinical... ...emphasizes the development of robust evaluation frameworks that assess medical reasoning,...SuggestedFull timeRemote work- ...Role Overview Apply advanced offensive security expertise to evaluate and validate AI-generated security analyses, exploit development reasoning, and vulnerability research across software, operating systems, networking, cloud, and web platforms. This role focuses on...SuggestedHourly payRemote workVisa sponsorshipWork visa
$60 - $90 per hour
...Role Overview Help build next-generation agentic evaluation benchmarks for frontier AI models by turning rigorous scientific practice into challenging... ...methods and expected results at the level of a careful researcher. Define scoring and success criteria, specifying...SuggestedHourly payFreelanceImmediate startRemote work$400k
...Role Overview Define the frontier of AI-powered legal reasoning by building rigorous evaluation frameworks and benchmarks for agentic AI systems that perform... ...complex legal tasks. This role combines applied legal research, dataset curation, and collaboration with...SuggestedFull timeRemote work$200k - $400k
...intelligent machines at scale. At Scout AI, we’re developing Fury, the first robotic... ...and relentless work. We’re hiring an AI Researcher on the Fury Team to help push the frontiers... ...tests and data collection campaigns to evaluate system performance in realistic environments...SuggestedFull timeRelocation package- Role Description We're hiring our first dedicated AI Researcher to advance the core models powering Ares. You'll work alongside our VP of... ...horizons. You'll design experiments end-to-end, build the evaluation infrastructure the field doesn't yet have, and translate research...Permanent employmentFull time
$240k - $400k
...who discovers novel attack paths across AI agents, cloud infrastructure, developer platforms... ...1, you will perform advanced offensive research to harden our stack, partner with... ...and advanced persistent threat scenarios, evaluating privilege escalation, lateral movement, and...Full timeRemote work$20 - $55 per hour
...subject matter expertise to conduct literature and document research, prepare and review complex Word and PDF deliverables, and... ...-ready reports and presentations that help train and evaluate next-generation AI systems. You will contribute evidence-based insights and high...Hourly payFor contractorsRemote work$85.3k
...you eager to use artificial intelligence (AI) to unlock insights from complex, high... ...the Intelligent Systems Center, conducts research at the intersection of AI and complex systems... .... You will create, adapt, and evaluate modern AI methods, including machine learning...Interim roleRemote work$85.3k
...you eager to use artificial intelligence (AI) to unlock insights from complex, high-... ...the Intelligent Systems Center, conducts research at the intersection of AI and complex systems... ...projects by creating, adapting, and evaluating modern AI methods including machine learning...Temporary workWork experience placementInterim roleRemote workRelocation packageFlexible hours$40 per hour
A leading AI cybersecurity firm is seeking experienced cybersecurity professionals to improve AI models in the field. In this remote role, you will evaluate AI-generated security content, solve technical cybersecurity problems, and provide feedback to enhance the accuracy...Hourly payRemote workFlexible hours- ...diagnosis, software as a medical product, and AI marketplaces. Key Responsibilities ~... ...for healthcare. ~Design, build, and evaluate solutions for healthcare use cases (e.g.,... ..., semantic search), performing research, experimentation, data management, and model...Full time
- ...Responsibilities ~Interdisciplinary evaluation and advisory support assessing the feasibility... ..., and policy implications of applying AI and data-analytics tools across CARICOM IMPACS... ...related. ~Impact evaluations/applied research in governance, law enforcement, or...Full timeContract workRemote work
$50 per hour
...models. Produce clear, step-by-step solutions and work with researchers to build evaluation benchmarks covering topics from early undergraduate... ...English comprehension, and the chance to learn how to use AI to enhance your analytical workflow. Key Responsibilities...Contract workFor contractorsFreelanceRemote work$50 per hour
...thinking, and clear written explanations to improve and evaluate large language models and other AI systems. You will design challenging math problems,... ...errors and gaps. Work contributes to both frontier AI research and to projects that help enterprises deploy reliable...Contract workFor contractorsFreelanceRemote work$50 per hour
...contractor to produce clear, step-by-step solutions, annotations, and evaluation benchmarks spanning early undergraduate through PhD-level... ..., or simulations when appropriate. Collaborate with LLM researchers to align problem types and solutions with evaluation goals,...Contract workFor contractorsFreelanceRemote work- Our client is a fast-growing AI consulting firm helping enterprises deploy artificial intelligence... ...to accelerate, the firm is expanding its research team. Role Overview The AI Researcher role focuses on identifying, evaluating, and advancing emerging AI technologies for...Remote workFlexible hours
$250k - $350k
Job description AI Researcher San Carlos, CA (on-site, remote) About the Lab The 1X World Model Lab is an embodied AI research organization... ...supercharge the model’s ability in the lab and in the world. Evaluations Build the evaluation infrastructure that connects pre‑...Local areaRemote work$40 per hour
...We are looking for an Applied Mathematician to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the quality of each model. In this role you will need to hold an expert level of...Hourly payFull timeContract workPart timeRemote work$40 per hour
A technology company is seeking a Research Scientist (Biology) to enhance AI models by evaluating their responses to complex biology inquiries. This role is open to applicants throughout the United States, offering flexible scheduling and payment of $40+ hourly via PayPal...Hourly payFor contractorsRemote workFlexible hours$40 per hour
A leading tech company is seeking a Research Scientist (Biology) to join its team. This position involves training AI models and assessing their logic by evaluating the performance of AI chatbots through complex biology questions. Ideal candidates should be detail-oriented...Hourly payRemote work$40 per hour
A leading AI training company is looking for a Research Scientist (Biology) to evaluate AI models. This role involves testing chatbots with complex biology questions and requires strong expertise in biology and related fields. Candidates must be fluent in English and detail...Hourly payContract workRemote workFlexible hours$40 per hour
...DataAnnotation is seeking a Biotechnology R&D Scientist to train AI models. In this role, you will evaluate the outputs of AI chatbots and assess their logic to improve model quality. The ideal candidate should have a deep understanding of cell biology, genetics, biochemistry...Hourly payFor contractorsRemote work$40 per hour
A leading data science firm is seeking a Research Scientist (Chemistry) to evaluate AI chatbots and improve their logic and performance. This role requires an expert understanding of chemistry, and candidates can choose projects while working on their own schedule. The...Hourly payRemote work- A research-based organization is seeking a Research Scientist (Chemistry) to join their team. You will engage in training AI models, measuring their progress, and solving logic issues to enhance... ...with responsibilities focused on evaluating AI outputs on chemistry queries....Hourly payRemote work
$40 per hour
An innovative tech company in Missouri is seeking a Research Scientist (Chemistry) to train AI models and evaluate their performance. The ideal candidate will have an expert level understanding of chemistry and be detail-oriented. Responsibilities include assessing AI...Hourly payRemote workFlexible hours$40 per hour
A technology firm specializing in AI is looking for a Research Scientist (Chemistry) to join their team. This role involves training AI models by evaluating chatbot outputs on complex chemistry questions. The candidate should possess a solid understanding of chemistry...Hourly payContract workRemote work$40 per hour
A leading AI research firm seeks a Research Scientist (Chemistry) to train AI models, evaluate outputs, and enhance model quality. The role requires expertise in chemistry, with options for full-time or part-time remote work. Responsibilities include testing AI chatbots...Hourly payFull timePart timeRemote work- A tech company specializing in AI seeks a Research Scientist (Biology) to train AI models by evaluating their output and improving performance. The ideal candidate should have strong expertise in biology, particularly in cell biology and genetics, and possess a Master'...Hourly payContract workRemote workFlexible hours
$40 per hour
A leading AI training company is seeking a Research Scientist (Biology) to evaluate and train AI models in the United States. Responsibilities include assessing the performance of AI chatbots in complex biological queries. This role offers flexibility in projects and scheduling...Hourly payRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Researcher. Be the first to apply!





