Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Engineer, Code Generation & Model Evaluation

$50 - $100 per hour

SaidGig

Apply your software engineering expertise to help train and evaluate next-generation AI systems through real-world coding tasks, technical feedback, and code quality analysis. This remote contract role focuses on code generation workflows and model evaluation; prior AI experience is not required. Key Responsibilities

  • Analyze, debug, and resolve issues in codebases using Python 3, Java, Rust, C++, Go, or TypeScript.
  • Design and implement features and enhancements for code generation workflows.
  • Refactor and optimize code to improve performance, maintainability, and adaptability.
  • Create and review practical coding tasks used to assess AI model performance and accuracy.
  • Write clear feedback and annotations that support model training and assessment.
  • Collaborate with open-source contributors and technical stakeholders on technical direction and reliable outputs.
  • Document technical decisions, best practices, and solutions for transparency and knowledge sharing.
Qualifications
  • Strong competitive programming and coding problem analysis expertise.
  • Proficiency in at least one of Python 3, Java, Rust, C++, Go, or TypeScript, including solid algorithms and data structures knowledge.
  • Experience with bug fixing, feature implementation, codebase refactoring, and performance optimization.
  • A demonstrated record of open-source contributions or collaborative software project work.
  • Strong analytical skills for interpreting complex constraints and evaluating multiple solution paths.
  • Excellent written and verbal technical communication skills, close attention to code validation and output consistency, and the ability to work independently in a remote collaborative setting.
  • Commitment to delivering high-quality, well-documented code under tight deadlines.
Work Terms
  • Remote, contractor engagement.
  • Work is paid on an output basis per task that meets project specifications; completion time varies by experience and workflow.
  • Minimum submission requirements apply, including a minimum number of tasks each week.
  • Selected candidates should be ready to begin their first tasks within 24 to 48 hours after onboarding is completed.
Compensation

$50 to $100 per hour.

Application Process

Submit an application through the available email or Google sign-in option, agree to the applicable terms and privacy policies, and complete onboarding if selected.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Research Engineer, Code Generation & Model Evaluation in United States vacancy
  • $50 - $100 per hour

     ...Role Title: Research Engineer - Code Generation & Model Evaluation Role Type: Contractor Location: Remote micro1 is engaging Research Engineers to participate in a project focused on code generation and model evaluation for a customer's initiative. In this role... 
    Suggested
    For contractors
    Remote work

    micro1

    Remote
    4 hours ago
  •  ...a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to...  ...the role We're looking for Research Engineers to build the evaluations that tell us — and the world — what Claude can actually do... 
    Suggested
    Full time

    Anthropic

    Remote
    5 days ago
  •  ...Francisco, California. The Role: As a Research Engineer - Language Model Pre-Training , you'll shape our...  ...your insights into our next-generation models. You'll Work Across:...  ...Dataset collection, processing, and evaluation Architecture and methodology research... 
    Suggested
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    24 days ago
  • $136.44k - $265.11k

     ...data, and run AI agents and models directly in their workflows....  ....You’ll build the datasets, evaluations, and systems that help close...  ...the intersection of software engineering, biology, and frontier AI:...  ...engineers, scientists, and external research partners.Desire to work in a... 
    Suggested
    Work at office
    Local area
    Monday to Friday
    Shift work

    Benchling

    San Francisco, CA
    2 days ago
  •  ...solving to improve and evaluate large language models. You will design...  ...numerical results with code, and review model...  ...Accelerate frontier AI research by contributing high...  ...and annotate model generated solutions, identify...  ...level expected for engineering entrance exams and for... 
    Suggested
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    a month ago
  •  ...Job Description Job Description Research Engineer — AI Alignment & Evaluation AI Safety / Research Engineering...  ...at the intersection of frontier model evaluation, AI safety, and security...  ...research workflows. Review agent-generated work critically and identify... 
    Full time
    Work at office
    Relocation
    Visa sponsorship

    W3 Sourcing

    San Francisco, CA
    23 days ago
  • $174k - $252k

    Drive post-training research and engineering using reinforcement...  ...) to advance Gemini coding capabilities across...  ...and maintain frontier evaluation suites and automated...  ...infrastructure, reward models, and data curation...  ...workflows, code generation, or software engineering... 

    Google

    Mountain View, CA
    4 days ago
  • $224k - $356.5k

     ...Tools organization is seeking a Senior Research Engineer to join our Research team, where we build the AI coding agents, models, datasets, and evaluations at the heart of NVIDIA's strategy...  ...to rigorous evaluations for code generation or agentic systemsTrack record of shipping... 
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $70 - $80 per hour

     ...expertise to help improve AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data. This remote...  ...benefit-risk assessment methods. Hands-on experience with MedDRA coding, seriousness and causality assessments, expectedness... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    a month ago
  • $20 - $36 per hour

     ...Role Overview Evaluate generative music AI across a wide range of genres, applying your knowledge of Hungarian music and lyrics to detailed quality standards. You will work in both Hungarian and English to help assess the quality, originality, and naturalness of AI-generated... 
    Hourly pay
    For contractors
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    7 days ago
  • $70 - $90 per hour

     ...tasks that support the training and evaluation of advanced AI models. This role focuses on assessing...  ...kernel task types: specification-based generation, cross-framework translation or...  ...or TPU. Background in compiler engineering, MLIR, or intermediate-representation... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    24 days ago
  •  ...As the Manager of Model Validation & Verification (VnV) for Behavior Autonomy, you will lead an engineering and data science team responsible for evaluating, benchmarking, and validating the machine...  ...with reinforcement learning, generative AI, or distributed ML systems.... 
    Full time
    Temporary work
    Relocation package

    Zoox

    California
    4 days ago
  •  ...company based in San Francisco, California. The Role: As a Research Engineer - Model Architectures , you will be a core contributor to Zyphra’...  ...team, who will integrate your insights into our next-generation models. What We're Looking For / Requirements:... 
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    24 days ago
  • $400 per month

     ...partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured...  ...infrastructure engineering workflows and model evaluation...  ...tasks. Review model-generated implementations... 

    Mercor

    San Francisco, CA
    5 days ago
  • Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote job
    Flexible hours

    Prolific

    Charlotte, NC
    1 day ago
  •  ...Founded by a team of Stanford researchers and entrepreneurs with...  ...deep expertise in model innovation and systems engineering with a design-minded product...  ...mark on an ambitious, generational mission to change how the...  ...families, build the evaluation infrastructure to measure... 

    Sanas.AI Inc.

    Brooklyn, NY
    5 days ago
  •  ...Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities include fact-checking AI technical... 
    Remote job
    Hourly pay
    Flexible hours

    Prolific

    Jacksonville, FL
    1 day ago
  • $60 per hour

     ...Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with flexible...  ...a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing experimental... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    Dallas, TX
    1 day ago
  • $20 per hour

     ...and technical talent with leading AI research labs. Headquartered in San Francisco...  ...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths...  ...completeness of responses. Ensure model responses align with expected conversational... 
    Remote job
    Contract work
    Part time
    Summer work

    Mercor

    New York, NY
    4 days ago
  • $70 - $80 per hour

     ...expert feedback that will help train next-generation AI systems for pharmacovigilance. This...  ...fully remote and focuses on high-quality evaluation of DSURs, PSURs/PBRERs, aggregate safety...  ...-risk assessment methodology, MedDRA coding, seriousness and causality assessment, expectedness... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    United States
    2 days ago
  • $60 per hour

    Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    2 days ago
  • Research Scientist Graduate (Foundation Model, Generative AI) - 2025 Start (PhD) Join ByteDance as a Research Scientist Graduate...  ...in Computer Science or related engineering field. Highly competent in algorithms and programming; strong coding skills in Python and PyTorch.... 

    ByteDance

    Seattle, WA
    2 days ago
  • $78 per hour

    Model Evaluator Please share 2 onsite profiles for Model Evaluators. Location can be either...  ...Skills Strong understanding of LLMs, generative AI, and transformer-based architectures...  ...evaluation frameworks. Familiarity with prompt engineering, embeddings, RLHF/RLAIF, and LLM-based... 

    ClifyX

    Austin, TX
    4 days ago
  • $60 - $90 per hour

     ...Apply hands-on mechanical engineering judgment to improve how advanced AI models reason through real-...  ...You will partner with AI research and program management...  ..., and create rigorous evaluations grounded in industry practice...  ...with applicable codes and standards, including... 
    Hourly pay
    Full time
    Remote work

    SaidGig

    United States
    1 day ago
  • $100 - $150 per hour

     ...considered for future projects evaluating how well AI systems...  ...produced analyses and models, document decisions in...  ...write-ups, feature engineering, and technical reports...  ...Evaluate AI-generated or human-created work...  ...a leading technology, research, or quantitative firm,... 
    Hourly pay
    Immediate start
    Remote work

    SaidGig

    United States
    a month ago
  • $36 - $72 per hour

     ...Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise. Role Details...  ...As an AI Image Evaluator, you will help image generation models learn two things at once: what a good... 
    Hourly pay
    Full time
    Monday to Friday
    Flexible hours

    Handshake

    Seattle, WA
    12 days ago
  • $315k

    We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand...  ...Safety Level (ASL) to assign to models. Research leads on this team...  ...training infrastructure to prepare new generations of models for routine... 
    Currently hiring
    Work at office
    Immediate start
    Home office
    Visa sponsorship
    Relocation package

    Anthropic

    Seattle, WA
    5 days ago
  • $60 per hour

     ...Chemistry Experts and Chemical Engineers to join their Expert Network. Participants will evaluate AI-generated chemistry through tasks that...  ...-edge advancements in AI models. The position requires a strong...  ...or industrial experience. Researchers can earn up to $60 per hour... 
    Hourly pay

    Prolific

    Arizona City, AZ
    5 days ago
  • $305k

     ...a quickly growing group of committed researchers, engineers, policy experts, and business leaders...  ...role As a Product Manager on Claude Code's model performance team, you will drive model...  ...model behavior, prompt engineering, and evaluation methodology Are a systems thinker:... 
    Work at office
    Visa sponsorship
    Flexible hours

    United States Digital Space LLC

    Seattle, WA
    4 days ago
  • $65 - $105 per hour

     ...Apply deep engineering judgment to help frontier AI models reason more accurately about real-world...  ...work closely with an AI research and program management...  ...quality engineering work, evaluate model performance, and...  ...reasoning, subtly incorrect code, unaddressed edge cases,... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    28 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Engineer, Code Generation & Model Evaluation. Be the first to apply!