Research Engineer, Code Generation & Model Evaluation
$50 - $100 per hourSaidGig
Apply your software engineering expertise to help train and evaluate next-generation AI systems through real-world coding tasks, technical feedback, and code quality analysis. This remote contract role focuses on code generation workflows and model evaluation; prior AI experience is not required. Key Responsibilities
- Analyze, debug, and resolve issues in codebases using Python 3, Java, Rust, C++, Go, or TypeScript.
- Design and implement features and enhancements for code generation workflows.
- Refactor and optimize code to improve performance, maintainability, and adaptability.
- Create and review practical coding tasks used to assess AI model performance and accuracy.
- Write clear feedback and annotations that support model training and assessment.
- Collaborate with open-source contributors and technical stakeholders on technical direction and reliable outputs.
- Document technical decisions, best practices, and solutions for transparency and knowledge sharing.
- Strong competitive programming and coding problem analysis expertise.
- Proficiency in at least one of Python 3, Java, Rust, C++, Go, or TypeScript, including solid algorithms and data structures knowledge.
- Experience with bug fixing, feature implementation, codebase refactoring, and performance optimization.
- A demonstrated record of open-source contributions or collaborative software project work.
- Strong analytical skills for interpreting complex constraints and evaluating multiple solution paths.
- Excellent written and verbal technical communication skills, close attention to code validation and output consistency, and the ability to work independently in a remote collaborative setting.
- Commitment to delivering high-quality, well-documented code under tight deadlines.
- Remote, contractor engagement.
- Work is paid on an output basis per task that meets project specifications; completion time varies by experience and workflow.
- Minimum submission requirements apply, including a minimum number of tasks each week.
- Selected candidates should be ready to begin their first tasks within 24 to 48 hours after onboarding is completed.
$50 to $100 per hour.
Application ProcessSubmit an application through the available email or Google sign-in option, agree to the applicable terms and privacy policies, and complete onboarding if selected.
$50 - $100 per hour
...Role Title: Research Engineer - Code Generation & Model Evaluation Role Type: Contractor Location: Remote micro1 is engaging Research Engineers to participate in a project focused on code generation and model evaluation for a customer's initiative. In this role...SuggestedFor contractorsRemote work- ...a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to... ...the role We're looking for Research Engineers to build the evaluations that tell us — and the world — what Claude can actually do...SuggestedFull time
- ...Francisco, California. The Role: As a Research Engineer - Language Model Pre-Training , you'll shape our... ...your insights into our next-generation models. You'll Work Across:... ...Dataset collection, processing, and evaluation Architecture and methodology research...SuggestedWork at officeRelocation package
$136.44k - $265.11k
...data, and run AI agents and models directly in their workflows.... ....You’ll build the datasets, evaluations, and systems that help close... ...the intersection of software engineering, biology, and frontier AI:... ...engineers, scientists, and external research partners.Desire to work in a...SuggestedWork at officeLocal areaMonday to FridayShift work- ...solving to improve and evaluate large language models. You will design... ...numerical results with code, and review model... ...Accelerate frontier AI research by contributing high... ...and annotate model generated solutions, identify... ...level expected for engineering entrance exams and for...SuggestedContract workFor contractorsFreelanceRemote work
- ...Job Description Job Description Research Engineer — AI Alignment & Evaluation AI Safety / Research Engineering... ...at the intersection of frontier model evaluation, AI safety, and security... ...research workflows. Review agent-generated work critically and identify...Full timeWork at officeRelocationVisa sponsorship
$174k - $252k
Drive post-training research and engineering using reinforcement... ...) to advance Gemini coding capabilities across... ...and maintain frontier evaluation suites and automated... ...infrastructure, reward models, and data curation... ...workflows, code generation, or software engineering...$224k - $356.5k
...Tools organization is seeking a Senior Research Engineer to join our Research team, where we build the AI coding agents, models, datasets, and evaluations at the heart of NVIDIA's strategy... ...to rigorous evaluations for code generation or agentic systemsTrack record of shipping...Full timeShift work$70 - $80 per hour
...expertise to help improve AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data. This remote... ...benefit-risk assessment methods. Hands-on experience with MedDRA coding, seriousness and causality assessments, expectedness...Hourly payContract workRemote work$20 - $36 per hour
...Role Overview Evaluate generative music AI across a wide range of genres, applying your knowledge of Hungarian music and lyrics to detailed quality standards. You will work in both Hungarian and English to help assess the quality, originality, and naturalness of AI-generated...Hourly payFor contractorsImmediate startRemote workFlexible hours$70 - $90 per hour
...tasks that support the training and evaluation of advanced AI models. This role focuses on assessing... ...kernel task types: specification-based generation, cross-framework translation or... ...or TPU. Background in compiler engineering, MLIR, or intermediate-representation...Hourly payRemote work- ...As the Manager of Model Validation & Verification (VnV) for Behavior Autonomy, you will lead an engineering and data science team responsible for evaluating, benchmarking, and validating the machine... ...with reinforcement learning, generative AI, or distributed ML systems....Full timeTemporary workRelocation package
- ...company based in San Francisco, California. The Role: As a Research Engineer - Model Architectures , you will be a core contributor to Zyphra’... ...team, who will integrate your insights into our next-generation models. What We're Looking For / Requirements:...Work at officeRelocation package
$400 per month
...partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured... ...infrastructure engineering workflows and model evaluation... ...tasks. Review model-generated implementations...- Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote jobFlexible hours
- ...Founded by a team of Stanford researchers and entrepreneurs with... ...deep expertise in model innovation and systems engineering with a design-minded product... ...mark on an ambitious, generational mission to change how the... ...families, build the evaluation infrastructure to measure...
- ...Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities include fact-checking AI technical...Remote jobHourly payFlexible hours
$60 per hour
...Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with flexible... ...a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing experimental...Remote jobHourly payWork from homeFlexible hours$20 per hour
...and technical talent with leading AI research labs. Headquartered in San Francisco... ...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths... ...completeness of responses. Ensure model responses align with expected conversational...Remote jobContract workPart timeSummer work$70 - $80 per hour
...expert feedback that will help train next-generation AI systems for pharmacovigilance. This... ...fully remote and focuses on high-quality evaluation of DSURs, PSURs/PBRERs, aggregate safety... ...-risk assessment methodology, MedDRA coding, seriousness and causality assessment, expectedness...Hourly payFor contractorsRemote work$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Remote jobHourly payWork from homeFlexible hours- Research Scientist Graduate (Foundation Model, Generative AI) - 2025 Start (PhD) Join ByteDance as a Research Scientist Graduate... ...in Computer Science or related engineering field. Highly competent in algorithms and programming; strong coding skills in Python and PyTorch....
$78 per hour
Model Evaluator Please share 2 onsite profiles for Model Evaluators. Location can be either... ...Skills Strong understanding of LLMs, generative AI, and transformer-based architectures... ...evaluation frameworks. Familiarity with prompt engineering, embeddings, RLHF/RLAIF, and LLM-based...$60 - $90 per hour
...Apply hands-on mechanical engineering judgment to improve how advanced AI models reason through real-... ...You will partner with AI research and program management... ..., and create rigorous evaluations grounded in industry practice... ...with applicable codes and standards, including...Hourly payFull timeRemote work$100 - $150 per hour
...considered for future projects evaluating how well AI systems... ...produced analyses and models, document decisions in... ...write-ups, feature engineering, and technical reports... ...Evaluate AI-generated or human-created work... ...a leading technology, research, or quantitative firm,...Hourly payImmediate startRemote work$36 - $72 per hour
...Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise. Role Details... ...As an AI Image Evaluator, you will help image generation models learn two things at once: what a good...Hourly payFull timeMonday to FridayFlexible hours$315k
We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand... ...Safety Level (ASL) to assign to models. Research leads on this team... ...training infrastructure to prepare new generations of models for routine...Currently hiringWork at officeImmediate startHome officeVisa sponsorshipRelocation package$60 per hour
...Chemistry Experts and Chemical Engineers to join their Expert Network. Participants will evaluate AI-generated chemistry through tasks that... ...-edge advancements in AI models. The position requires a strong... ...or industrial experience. Researchers can earn up to $60 per hour...Hourly pay$305k
...a quickly growing group of committed researchers, engineers, policy experts, and business leaders... ...role As a Product Manager on Claude Code's model performance team, you will drive model... ...model behavior, prompt engineering, and evaluation methodology Are a systems thinker:...Work at officeVisa sponsorshipFlexible hours$65 - $105 per hour
...Apply deep engineering judgment to help frontier AI models reason more accurately about real-world... ...work closely with an AI research and program management... ...quality engineering work, evaluate model performance, and... ...reasoning, subtly incorrect code, unaddressed edge cases,...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Engineer, Code Generation & Model Evaluation. Be the first to apply!
- cyber research engineer United States
- junior machine learning research engineer United States
- engineering analyst United States
- research assistant engineering United States
- robotics research engineer United States
- research software engineer United States
- research engineer United States
- deep learning research engineer United States
- senior research engineer United States
- research programmer United States


