Research Engineer - Code Generation & Model Evaluation [Remote]
AuraOne Human Data
- Remote job
Research Engineer - Code Generation & Model Evaluation is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should have written, and explain the fix so the modeling team can target the gap.
Why this role matters
Engineering model quality lives or dies on whether the generated code actually compiles, passes tests, and handles edge cases. AuraOne pairs experienced engineers with the modeling team to grade outputs the way a code reviewer would.
Responsibilities
- Run and reproduce candidate code outputs in a sandboxed environment for Research Engineer - Code Generation & Model Evaluation assignments.
- Grade research engineer code generation model evaluation engineering review solutions for correctness, style, and edge-case handling.
- Write minimal failing tests that demonstrate the bug a model output missed.
- Compare paired solutions and rank them with a written rationale tied to the rubric.
- Tag failure modes (compile error, runtime crash, off-by-one, security issue) with severity scores.
- Document recurring code-generation failures so the modeling team can target them.
- Calibrate against gold-standard reviews to keep inter-rater agreement above target.
Qualifications
- Strong day-job engineering experience — you can read, run, and debug unfamiliar code for Research Engineer - Code Generation & Model Evaluation work.
- Comfort writing concise unit tests that capture a single failure mode.
- Familiarity with a testing framework. pytest, Jest, or JUnit. Go test, RSpec, or whatever you use.
- Clear written reasoning — your review note has to convince another senior engineer.
- Reliable async availability for at least 10 hours per week.
- Prior code-review or technical-interview-grading experience is a plus.
Example tasks
- Reproduce a generated engineering solution to a coding task, run the test suite, and grade it.
- Write the smallest failing test that demonstrates a model's edge-case bug.
- Compare two paired solutions and rank them with a written rationale tied to the rubric.
- Triage a security issue surfaced by a model and document the patch the model should have produced.
Nice to have
- Open-source contributions or a public portfolio that demonstrates production-quality code.
- Experience with the target language's standard tooling, linters, and idiomatic style guides.
- Familiarity with security-review checklists (OWASP, CWE) and AppSec patterns.
Skills
- Code review
- Debugging
- Unit testing
- Software engineering judgment
- Research Engineer Code Generation Model Evaluation engineering review
- AI model evaluation
Work model
Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.
Compensation
Hourly rate confirmed after the interview process.
Application process
Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.
$50 - $100 per hour
...Apply your software engineering expertise to help train and evaluate next-generation AI systems through real-world coding tasks, technical feedback, and code quality analysis. This... ...focuses on code generation workflows and model evaluation; prior AI experience is not required...SuggestedHourly payContract workFor contractorsRemote work$50 - $100 per hour
...Role Title: Research Engineer - Code Generation & Model Evaluation Role Type: Contractor Location: Remote micro1 is engaging Research Engineers to participate in a project focused on code generation and model evaluation for a customer's initiative. In this role...SuggestedFor contractorsRemote work- ...solving to improve and evaluate large language models. You will design... ...numerical results with code, and review model... ...Accelerate frontier AI research by contributing high... ...and annotate model generated solutions, identify... ...level expected for engineering entrance exams and for...SuggestedContract workFor contractorsFreelanceRemote work
- ...Role Overview Senior Software Engineer contractor role, remote, paying $245 to $280 per hour... ...AI lab improve frontier AI systems by evaluating AI-generated code, creating challenging engineering problems, and supporting model performance on real-world software tasks....SuggestedHourly payFor contractorsRemote work
$315k
We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand... ...Safety Level (ASL) to assign to models. Research leads on this team... ...training infrastructure to prepare new generations of models for routine...SuggestedCurrently hiringWork at officeImmediate startHome officeVisa sponsorshipRelocation package$295k
...cutting-edge foundation AI models and end-to-end products that... .... Cohere is a team of researchers, engineers, designers, and more, who are... ...Paris. Join us!Why this role?Evaluation is critical to making progress... ...for creating these next-generation evaluation methods and infrastructure...Full timeWork at officeLocal areaRemote workHome office- ...Description As the Manager of Model Validation & Verification (VnV)... ...Behavior Autonomy, you will lead an engineering and data science team responsible for evaluating, benchmarking, and validating... ...Experience with reinforcement learning, generative AI, or distributed ML systems....Temporary workRelocation package
- Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote jobFlexible hours
$160k - $200k
...Posted: 2026-09-19 Category: Engineering and Sciences Subcategory: Modeling/Sim Engr Schedule: Full-Time... ...coordination, and analysis results evaluation. Recommend and document force... ...simulation scripts and interface code using government supportable software...Full timeRemote workShift work- ...Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities include fact-checking AI technical...Remote jobHourly payFlexible hours
$20 per hour
...and technical talent with leading AI research labs. Headquartered in San Francisco... ...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths... ...completeness of responses. Ensure model responses align with expected conversational...Remote jobContract workPart timeSummer work$192k - $304.75k
...now looking for a Senior Research Engineer passionate about Generative AI inference. Are you excited... ...of generative AI models, from language to images.... ...will be doing:Design and evaluate routing policies for LLM... ...source repo: design docs, code review, docs, and community...Full timeRemote work$70 - $90 per hour
...tasks that support the training and evaluation of advanced AI models. This role focuses on assessing... ...kernel task types: specification-based generation, cross-framework translation or... ...or TPU. Background in compiler engineering, MLIR, or intermediate-representation...Hourly payRemote work- OverviewFrom our Ann Arbor, MI office, we have an opening for a Research Engineer with FPGA and low-level code development (C and C++) for real-time and high-performance computing (HPC) applications. You will participate in multi-disciplinary, collaborative teams and contribute...Work at officeRemote work
$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Remote jobHourly payWork from homeFlexible hours$250k - $290k
...to help them hire. Research Engineer Location San... ...data engineering, and evaluation . This is not a traditional... ...or experimenting with models in isolation. It is... ...infrastructure, generate and curate datasets, and... ..., production-quality code Ability to debug...H1bRemote work- ...Ando Research Team Member Ando is a messaging platform... ...track, or intervene. Evaluating that judgment means... ...background, with depth in model evaluation,... ...credentials. Work samples or code that demonstrate the skills... ...hold high-bandwidth, generative technical...Work from home
- ...Research Engineer Colombia, Huila, Colombia Turing's... ...frontier AI labs to generate high-quality datasets... ...benchmarks that improve model capabilities in software... ...into better data, evaluations, and more capable models... ...areas: Coding and software engineering...Remote workShift work
$100 - $150 per hour
...Role Overview Apply your data science expertise to evaluate AI-generated slides, spreadsheets, and documents for real-world quality and usability. You will assess outputs against professional standards and deliver clear feedback that improves their accuracy and presentation...Hourly payWork at officeRemote work$60 - $90 per hour
...Apply hands-on mechanical engineering judgment to improve how advanced AI models reason through real-... ...You will partner with AI research and program management... ..., and create rigorous evaluations grounded in industry practice... ...with applicable codes and standards, including...Hourly payFull timeRemote work$100 - $150 per hour
...considered for future projects evaluating how well AI systems... ...produced analyses and models, document decisions in... ...write-ups, feature engineering, and technical reports... ...Evaluate AI-generated or human-created work... ...a leading technology, research, or quantitative firm,...Hourly payImmediate startRemote work- ...holistic approach for evaluating security posture and ecosystems... ...-conscious Innovation Engineer to join our technology... ...and scaling our generative AI capabilities. You... ..., Large Language Models (LLMs), prompt engineering... ...with Infrastructure as Code (IaC) tools like AWS...Local areaRemote workFlexible hours
$65 - $105 per hour
...Apply deep engineering judgment to help frontier AI models reason more accurately about real-world... ...work closely with an AI research and program management... ...quality engineering work, evaluate model performance, and... ...reasoning, subtly incorrect code, unaddressed edge cases,...Hourly payFull timeFreelanceInternshipLive inRelocationRelocation package$81.26k - $137.14k
...windows team, leading research related to window... ...analytical and modeling tools. You will... ...energy/market/engineering analysis and provide... ...improvement, and evaluation of the... ..., and model data generated from energy simulations... ...energy policy and codes. Knowledge in statistics...Full timeWork experience placement- ...Job Description Job Description Research Engineer – Code Generation & Model Evaluation Role Type: Contractor Location: Remote We are seeking experienced Research Engineers to contribute to a project focused on code generation, software engineering, and...For contractorsRemote work
- Python Infrastructure Engineer - Model Evaluation (AI Training) About the Role What if your Python expertise... ...depend on to train and validate next-generation models. This is a fully remote... ...fixes Collaborate with data, research, and engineering teams to support model...Hourly payOngoing contractContract workFreelanceRemote workFlexible hours
- ...Motors, through Embodied AI, seeks a Senior Engineer to measure and visualize AV model performance. You will design and implement evaluation workflows, collaborate across Data,... ...influence safety and scalability of next‑generation autonomous systems and to contribute to...Remote job
$60 - $90 per hour
...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Summers , and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract...Full timeContract workSummer workRemote work- ...role As a Staff Research Engineer, you will join a team... ...edge challenges in the Generative AI space, with a... ...interactive video diffusion models. Within the team you’... .... Build robust evaluation frameworks and test... ...research code. Outcome-driven...Full timeWork at officeRemote workWorldwide
$158k - $293k
...TR Labs, owns the engineering behind that layer: next-generation search and retrieval... ...delivers.What makes this research engineering rather... ...different ranking model, a hybrid retrieval... ...to what a coding agent hands you as... ...understandingBuild evaluation that actually discriminates...Full timeWork at officeLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Research Engineer - Code Generation & Model Evaluation [Remote]. Be the first to apply!
- junior machine learning research engineer Remote
- research engineer Remote
- research programmer Remote
- engineering business analyst Remote
- deep learning research engineer Remote
- research software engineer Remote
- senior research engineer Remote
- procurement specialist remote Remote
- physician consultant remote Remote
- remote insurance agent Remote




