Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Evaluation Researcher

Full-time

Weekday

Join a pioneering AI initiative focused on developing the next generation of evaluation benchmarks for frontier AI models.

We are seeking researchers from computational STEM disciplines—as well as computationally intensive social sciences and humanities—to bring the rigor of real-world research into AI evaluation.

In this role, you will transform scientific methodologies such as experimental design, hypothesis testing, and data-driven analysis into sophisticated, multi-step benchmark tasks that challenge state-of-the-art AI systems. Working closely with AI researchers, you'll help uncover subtle reasoning errors and methodological flaws that only experienced researchers can identify.

This is a fully remote, full-time engagement requiring approximately 35 hours per week .

Requirements

Key Responsibilities

  • Design complex, research-oriented benchmark tasks inspired by real-world scientific workflows, including study design, experimentation, hypothesis testing, and data analysis.
  • Develop comprehensive reference solutions using Python , notebooks, and computational tools with the rigor expected in professional research.
  • Define clear evaluation standards that distinguish sound scientific reasoning from plausible but incorrect conclusions.
  • Review AI-generated solutions, identifying methodological weaknesses, analytical errors, and flawed reasoning that experienced researchers would recognize immediately.
  • Collaborate with AI researchers and fellow domain experts to improve benchmark quality, consistency, and scientific rigor.
  • Contribute to the continuous refinement of evaluation methodologies for advanced AI systems.

Required Qualifications

  • Master's degree, PhD, or equivalent practical experience in a STEM discipline , computational social science, computational humanities, or another research-intensive field involving programming and data analysis.
  • Minimum 1 year of experience in an active research role within academia, industry, government laboratories, or a similar research environment.
  • Demonstrated experience performing computational research involving Python , data analysis, simulation, modeling, machine learning, or scientific computing.
  • Strong understanding of experimental design, hypothesis testing, statistical analysis, and rigorous interpretation of research findings.
  • Working knowledge of Git , integrated development environments (IDEs), and notebook platforms such as Jupyter or Google Colab .
  • Experience with AI evaluation, benchmark development, AI training, or task authoring is preferred.
  • Excellent analytical thinking, attention to detail, creativity, and the ability to solve complex, open-ended problems independently.
  • Strong written communication skills for documenting technical methodologies and research findings.
  • Ability to commit approximately 35 hours per week on a consistent basis.

Preferred Qualifications

  • Experience designing reproducible computational experiments or research workflows.
  • Familiarity with machine learning, large language models, or AI-assisted research tools.
  • Background in benchmark design, scientific software development, or computational research infrastructure.
  • Experience mentoring researchers, reviewing scientific work, or contributing to peer-reviewed publications.

Why Join

  • Help shape how next-generation AI systems are evaluated using rigorous scientific methodologies.
  • Collaborate with leading AI researchers working on frontier models and advanced evaluation frameworks.
  • Apply your research expertise to improve AI reasoning, reliability, and scientific accuracy.
  • Contribute to impactful work that advances the quality and robustness of AI systems across multiple disciplines.
  • Enjoy the flexibility of a fully remote engagement while working on cutting-edge AI research initiatives.

Equal Opportunity

We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.

Contract & Engagement Details

  • Independent contractor engagement.
  • Fully remote with flexible working hours.
  • Expected commitment of approximately 35 hours per week .
  • Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
  • Work does not require access to confidential or proprietary information from any current or former employer.
  • Payments are issued weekly based on approved work completed.
  • At this time, we are unable to support H1-B or STEM OPT candidates.
Vacancy posted 6 days ago
Similar jobs that could be interesting for youBased on the AI Evaluation Researcher in Remote vacancy
  • $400k

     ...Join a dynamic research team as a Member of Technical Staff (MTS) focused on Medical & Health...  ...will play a pivotal role in advancing AI systems designed to enhance healthcare, clinical...  ...emphasizes the development of robust evaluation frameworks that assess medical reasoning,... 
    Suggested
    Full time
    Remote work

    SaidGig

    United States
    15 hours ago
  •  ...Role Overview Apply advanced offensive security expertise to evaluate and validate AI-generated security analyses, exploit development reasoning, and vulnerability research across software, operating systems, networking, cloud, and web platforms. This role focuses on... 
    Suggested
    Hourly pay
    Remote work
    Visa sponsorship
    Work visa

    SaidGig

    United States
    1 hour ago
  • $60 - $90 per hour

     ...Role Overview Help build next-generation agentic evaluation benchmarks for frontier AI models by turning rigorous scientific practice into challenging...  ...methods and expected results at the level of a careful researcher. Define scoring and success criteria, specifying... 
    Suggested
    Hourly pay
    Freelance
    Immediate start
    Remote work

    SaidGig

    United States
    1 day ago
  • $400k

     ...Role Overview Define the frontier of AI-powered legal reasoning by building rigorous evaluation frameworks and benchmarks for agentic AI systems that perform...  ...complex legal tasks. This role combines applied legal research, dataset curation, and collaboration with... 
    Suggested
    Full time
    Remote work

    SaidGig

    United States
    15 hours ago
  • $200k - $400k

     ...intelligent machines at scale. At Scout AI, we’re developing Fury, the first robotic...  ...and relentless work. We’re hiring an AI Researcher on the Fury Team to help push the frontiers...  ...tests and data collection campaigns to evaluate system performance in realistic environments... 
    Suggested
    Full time
    Relocation package

    Scout AI

    Remote
    3 days ago
  • Role Description We're hiring our first dedicated AI Researcher to advance the core models powering Ares. You'll work alongside our VP of...  ...horizons. You'll design experiments end-to-end, build the evaluation infrastructure the field doesn't yet have, and translate research... 
    Permanent employment
    Full time

    Assail, Inc.

    Remote
    7 days ago
  • $240k - $400k

     ...who discovers novel attack paths across AI agents, cloud infrastructure, developer platforms...  ...1, you will perform advanced offensive research to harden our stack, partner with...  ...and advanced persistent threat scenarios, evaluating privilege escalation, lateral movement, and... 
    Full time
    Remote work

    SaidGig

    United States
    15 hours ago
  • $20 - $55 per hour

     ...subject matter expertise to conduct literature and document research, prepare and review complex Word and PDF deliverables, and...  ...-ready reports and presentations that help train and evaluate next-generation AI systems. You will contribute evidence-based insights and high... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    United States
    15 hours ago
  • $85.3k

     ...you eager to use artificial intelligence (AI) to unlock insights from complex, high...  ...the Intelligent Systems Center, conducts research at the intersection of AI and complex systems...  .... You will create, adapt, and evaluate modern AI methods, including machine learning... 
    Interim role
    Remote work

    Johns Hopkins Applied Physics Laboratory

    Laurel, MD
    2 days ago
  • $85.3k

     ...you eager to use artificial intelligence (AI) to unlock insights from complex, high-...  ...the Intelligent Systems Center, conducts research at the intersection of AI and complex systems...  ...projects by creating, adapting, and evaluating modern AI methods including machine learning... 
    Temporary work
    Work experience placement
    Interim role
    Remote work
    Relocation package
    Flexible hours

    The Johns Hopkins University Applied Physics Laboratory

    Laurel, MD
    4 days ago
  • $40 per hour

    A leading AI cybersecurity firm is seeking experienced cybersecurity professionals to improve AI models in the field. In this remote role, you will evaluate AI-generated security content, solve technical cybersecurity problems, and provide feedback to enhance the accuracy... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Wyoming, OH
    2 days ago
  •  ...diagnosis, software as a medical product, and AI marketplaces. Key Responsibilities ~...  ...for healthcare. ~Design, build, and evaluate solutions for healthcare use cases (e.g.,...  ..., semantic search), performing research, experimentation, data management, and model... 
    Full time

    Quantiphi

    Remote
    6 days ago
  •  ...Responsibilities ~Interdisciplinary evaluation and advisory support assessing the feasibility...  ..., and policy implications of applying AI and data-analytics tools across CARICOM IMPACS...  ...related. ~Impact evaluations/applied research in governance, law enforcement, or... 
    Full time
    Contract work
    Remote work

    DarkStar Intelligence LLC

    Remote
    2 days ago
  • $50 per hour

     ...models. Produce clear, step-by-step solutions and work with researchers to build evaluation benchmarks covering topics from early undergraduate...  ...English comprehension, and the chance to learn how to use AI to enhance your analytical workflow. Key Responsibilities... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    15 hours ago
  • $50 per hour

     ...thinking, and clear written explanations to improve and evaluate large language models and other AI systems. You will design challenging math problems,...  ...errors and gaps. Work contributes to both frontier AI research and to projects that help enterprises deploy reliable... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    15 hours ago
  • $50 per hour

     ...contractor to produce clear, step-by-step solutions, annotations, and evaluation benchmarks spanning early undergraduate through PhD-level...  ..., or simulations when appropriate. Collaborate with LLM researchers to align problem types and solutions with evaluation goals,... 
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    15 hours ago
  • Our client is a fast-growing AI consulting firm helping enterprises deploy artificial intelligence...  ...to accelerate, the firm is expanding its research team. Role Overview The AI Researcher role focuses on identifying, evaluating, and advancing emerging AI technologies for... 
    Remote work
    Flexible hours

    CT Labs

    San Francisco, CA
    1 day ago
  • $250k - $350k

    Job description AI Researcher San Carlos, CA (on-site, remote) About the Lab The 1X World Model Lab is an embodied AI research organization...  ...supercharge the model’s ability in the lab and in the world. Evaluations Build the evaluation infrastructure that connects pre‑... 
    Local area
    Remote work

    Halodi Robotics

    San Carlos, CA
    4 days ago
  • $40 per hour

     ...We are looking for an Applied Mathematician to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the quality of each model. In this role you will need to hold an expert level of... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work

    DataAnnotation

    Brooklyn, NY
    3 days ago
  • $40 per hour

    A technology company is seeking a Research Scientist (Biology) to enhance AI models by evaluating their responses to complex biology inquiries. This role is open to applicants throughout the United States, offering flexible scheduling and payment of $40+ hourly via PayPal... 
    Hourly pay
    For contractors
    Remote work
    Flexible hours

    DataAnnotation

    Oregon, WI
    2 days ago
  • $40 per hour

    A leading tech company is seeking a Research Scientist (Biology) to join its team. This position involves training AI models and assessing their logic by evaluating the performance of AI chatbots through complex biology questions. Ideal candidates should be detail-oriented... 
    Hourly pay
    Remote work

    DataAnnotation

    Florida, NY
    2 days ago
  • $40 per hour

    A leading AI training company is looking for a Research Scientist (Biology) to evaluate AI models. This role involves testing chatbots with complex biology questions and requires strong expertise in biology and related fields. Candidates must be fluent in English and detail... 
    Hourly pay
    Contract work
    Remote work
    Flexible hours

    DataAnnotation

    New York, NY
    2 days ago
  • $40 per hour

     ...DataAnnotation is seeking a Biotechnology R&D Scientist to train AI models. In this role, you will evaluate the outputs of AI chatbots and assess their logic to improve model quality. The ideal candidate should have a deep understanding of cell biology, genetics, biochemistry... 
    Hourly pay
    For contractors
    Remote work

    DataAnnotation

    New York, NY
    3 days ago
  • $40 per hour

    A leading data science firm is seeking a Research Scientist (Chemistry) to evaluate AI chatbots and improve their logic and performance. This role requires an expert understanding of chemistry, and candidates can choose projects while working on their own schedule. The... 
    Hourly pay
    Remote work

    DataAnnotation

    El Paso, TX
    2 days ago
  • A research-based organization is seeking a Research Scientist (Chemistry) to join their team. You will engage in training AI models, measuring their progress, and solving logic issues to enhance...  ...with responsibilities focused on evaluating AI outputs on chemistry queries.... 
    Hourly pay
    Remote work

    DataAnnotation

    Iowa, LA
    2 days ago
  • $40 per hour

    An innovative tech company in Missouri is seeking a Research Scientist (Chemistry) to train AI models and evaluate their performance. The ideal candidate will have an expert level understanding of chemistry and be detail-oriented. Responsibilities include assessing AI... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Louisiana, MO
    2 days ago
  • $40 per hour

    A technology firm specializing in AI is looking for a Research Scientist (Chemistry) to join their team. This role involves training AI models by evaluating chatbot outputs on complex chemistry questions. The candidate should possess a solid understanding of chemistry... 
    Hourly pay
    Contract work
    Remote work

    DataAnnotation

    New York, NY
    2 days ago
  • $40 per hour

    A leading AI research firm seeks a Research Scientist (Chemistry) to train AI models, evaluate outputs, and enhance model quality. The role requires expertise in chemistry, with options for full-time or part-time remote work. Responsibilities include testing AI chatbots... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    Washington DC
    2 days ago
  • A tech company specializing in AI seeks a Research Scientist (Biology) to train AI models by evaluating their output and improving performance. The ideal candidate should have strong expertise in biology, particularly in cell biology and genetics, and possess a Master'... 
    Hourly pay
    Contract work
    Remote work
    Flexible hours

    DataAnnotation

    Bismarck, ND
    2 days ago
  • $40 per hour

    A leading AI training company is seeking a Research Scientist (Biology) to evaluate and train AI models in the United States. Responsibilities include assessing the performance of AI chatbots in complex biological queries. This role offers flexibility in projects and scheduling... 
    Hourly pay
    Remote work

    DataAnnotation

    Raleigh, NC
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Evaluation Researcher. Be the first to apply!