Research Engineer - Code Generation & Model Evaluation
$50 - $100 per hourmicro1
Role Title: Research Engineer - Code Generation & Model Evaluation
Role Type: Contractor
Location: Remote
micro1 is engaging Research Engineers to participate in a project focused on code generation and model evaluation for a customer's initiative. In this role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and perform through high-quality, real-world input. No prior experience in AI is required, your domain knowledge is what matters.
Scope of Work
- Analyze, debug, and resolve issues across diverse codebases written in Python3, Java, Rust, C++, Go, or TypeScript.
- Contribute to the design and implementation of new features and enhancements for code generation workflows.
- Refactor and optimize codebases to improve performance, maintainability, and adaptability.
- Develop and review real-world coding tasks to evaluate AI model performance and accuracy.
- Author clear, actionable feedback and annotations to support model training and assessment.
- Collaborate with other open-source contributors and technical stakeholders to set technical direction and ensure robust outputs.
- Document technical decisions, best practices, and solutions to drive project transparency and knowledge sharing.
Required Skills and Qualifications:
- Expertise in competitive programming and coding problem analysis.
- Expertise in at least one of: Python3, Java, Rust, C++, Go, or TypeScript, with solid knowledge of algorithms and data structures.
- Proven track record of open-source contributions or participation in collaborative software projects.
- Strong analytical abilities to interpret complex problem constraints and multiple solution paths.
- Exceptional written and verbal communication skills; ability to articulate technical details clearly.
- Meticulous attention to detail in code validation and output consistency.
- Experience working independently in a remote, collaborative environment.
- Commitment to producing high-quality, well-documented code under tight deadlines.
Compensation Structure
Compensation is output-based; experts are paid per task that meets the project specifications. The time required to complete work may vary depending on the expert’s experience and workflow. Minimum submission requirements apply. Experts must submit a minimum of tasks per week.
Start Timeline & Availability
We typically fill roles within 48 hours and are looking for experts ready to jump in right away. If selected, we expect you to start your first tasks within 24 to 48 hours of completing onboarding.
$50 - $100 per hour
...Apply your software engineering expertise to help train and evaluate next-generation AI systems through real-world coding tasks, technical feedback, and code quality analysis. This... ...focuses on code generation workflows and model evaluation; prior AI experience is not required...SuggestedHourly payContract workFor contractorsRemote work- ...a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to... ...the role We're looking for Research Engineers to build the evaluations that tell us — and the world — what Claude can actually do...SuggestedFull time
- ...solving to improve and evaluate large language models. You will design... ...numerical results with code, and review model... ...Accelerate frontier AI research by contributing high... ...and annotate model generated solutions, identify... ...level expected for engineering entrance exams and for...SuggestedContract workFor contractorsFreelanceRemote work
$70 - $80 per hour
...expertise to help improve AI systems through rigorous, real-world evaluation of pharmacovigilance documentation and data. This remote... ...benefit-risk assessment methods. Hands-on experience with MedDRA coding, seriousness and causality assessments, expectedness...SuggestedHourly payContract workRemote work$20 - $36 per hour
...Role Overview Evaluate generative music AI across a wide range of genres, applying your knowledge of Hungarian music and lyrics to detailed quality standards. You will work in both Hungarian and English to help assess the quality, originality, and naturalness of AI-generated...SuggestedHourly payFor contractorsImmediate startRemote workFlexible hours$70 - $90 per hour
...tasks that support the training and evaluation of advanced AI models. This role focuses on assessing... ...kernel task types: specification-based generation, cross-framework translation or... ...or TPU. Background in compiler engineering, MLIR, or intermediate-representation...Hourly payRemote work- ...As the Manager of Model Validation & Verification (VnV) for Behavior Autonomy, you will lead an engineering and data science team responsible for evaluating, benchmarking, and validating the machine... ...with reinforcement learning, generative AI, or distributed ML systems....Full timeTemporary workRelocation package
- ...Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities include fact-checking AI technical...Remote jobHourly payFlexible hours
$60 per hour
...Life Science Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to work from home with flexible... ...a competitive pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing experimental...Remote jobHourly payWork from homeFlexible hours- Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote jobFlexible hours
$20 per hour
...and technical talent with leading AI research labs. Headquartered in San Francisco... ...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths... ...completeness of responses. Ensure model responses align with expected conversational...Remote jobContract workPart timeSummer work$70 - $80 per hour
...expert feedback that will help train next-generation AI systems for pharmacovigilance. This... ...fully remote and focuses on high-quality evaluation of DSURs, PSURs/PBRERs, aggregate safety... ...-risk assessment methodology, MedDRA coding, seriousness and causality assessment, expectedness...Hourly payFor contractorsRemote work$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Remote jobHourly payWork from homeFlexible hours$60 - $90 per hour
...Apply hands-on mechanical engineering judgment to improve how advanced AI models reason through real-... ...You will partner with AI research and program management... ..., and create rigorous evaluations grounded in industry practice... ...with applicable codes and standards, including...Hourly payFull timeRemote work$100 - $150 per hour
...considered for future projects evaluating how well AI systems... ...produced analyses and models, document decisions in... ...write-ups, feature engineering, and technical reports... ...Evaluate AI-generated or human-created work... ...a leading technology, research, or quantitative firm,...Hourly payImmediate startRemote work$315k
We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand... ...Safety Level (ASL) to assign to models. Research leads on this team... ...training infrastructure to prepare new generations of models for routine...Currently hiringWork at officeImmediate startHome officeVisa sponsorshipRelocation package$90 - $175 per hour
...quality assurance expertise to evaluate technical AI outputs and help improve how next-generation AI systems learn, reason, and perform... ...experience as a QA Engineer, SDET, Test Engineer, QA Analyst... ...labeling, RLHF, AI response or model evaluation, or rubric-based grading...Hourly payContract workRemote work$65 - $105 per hour
...Apply deep engineering judgment to help frontier AI models reason more accurately about real-world... ...work closely with an AI research and program management... ...quality engineering work, evaluate model performance, and... ...reasoning, subtly incorrect code, unaddressed edge cases,...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package- ...dermatology expertise to image-based work that helps develop and evaluate advanced AI models for clinical image interpretation. This is a non-clinical,... ...with no direct patient care, focused on ensuring AI-generated medical outputs reflect real-world clinical reasoning and...Hourly payRemote work
$60 - $90 per hour
...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Summers , and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract...Full timeContract workSummer workRemote work- ...role As a Staff Research Engineer, you will join a team... ...edge challenges in the Generative AI space, with a... ...interactive video diffusion models. Within the team you’... .... Build robust evaluation frameworks and test... ...research code. Outcome-driven...Full timeWork at officeRemote workWorldwide
$159k - $296k
...makes decisions and generates trajectories for our... ...driving trucks. As a research engineer for Learnable Planner... ...Integrate cutting-edge ML models in production... ...structured and tested code. - Stay up-to-date... ...on a model including evaluation, introspection and fine...Full timeWork at officeWork from homeFlexible hours- ...Motors, through Embodied AI, seeks a Senior Engineer to measure and visualize AV model performance. You will design and implement evaluation workflows, collaborate across Data,... ...influence safety and scalability of next‑generation autonomous systems and to contribute to...Remote job
- ...Innovation Principle Engineer for a contract to hire... ...frameworks, AI gateway/model-proxy patterns, and enterprise... ...Retrieval-Augmented Generation (RAG) patterns using... ...controls, guardrails, evaluation, and observability... ..., reusable templates, code quality, API patterns,...Hourly payPermanent employmentContract workRemote work
$50 - $70 per hour
...Help improve frontier AI systems by evaluating the quality of professional work products across documents, presentations, spreadsheets... ...reasoned written feedback and ratings. Compare and rank AI generated outputs using defined evaluation criteria. Identify errors,...Hourly payRemote work$100k - $150k
...Large Language Model Specialist - Remote... ...an LLM Fine-Tuning Engineer to design, execute... ...construction, rigorous evaluation methodology, and... ...the bar through code review, design review... ...with product, research, and platform teams... ...with synthetic data generation and dataset...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship$80 - $100 per hour
...geospatial expertise to improve next-generation AI systems through practical... ...documentation, and rigorous evaluation of AI-generated solutions.... ...the use of advanced AI coding agents in technical projects.... ...geospatial data challenges for AI model development. Document...Hourly payContract workRemote work- ...Role Overview Use your investment and finance expertise to evaluate and improve AI model performance on financial reasoning, valuation, markets, and real-world investment scenarios. Key Responsibilities Assess AI model outputs on valuation, financial modeling, markets...For contractorsRemote work
$70 - $90 per hour
...Role Overview Help evaluate Neuron Kernel Interface development tasks that support the training and evaluation of advanced AI models. You will assess kernel quality, numerical correctness... ...NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth...Hourly payRemote work$100 per hour
...expertise to improve the performance of large language models on finance tasks. You will work with AI researchers to identify model weaknesses in areas such as... ...on advanced AI systems. Key Responsibilities Evaluate LLM performance in finance areas where models...Hourly payContract workFor contractorsFreelanceRemote work10 hours per weekFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Engineer - Code Generation & Model Evaluation. Be the first to apply!
- junior machine learning research engineer Remote
- research software engineer Remote
- research engineer Remote
- deep learning research engineer Remote
- senior research engineer Remote
- research programmer Remote
- anthropology research Remote
- research statistician Remote
- research and development assistant Remote
- research and development Remote



