Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Engineer - Code Generation & Model Evaluation [Remote]

AuraOne Human Data

Remote
  • Remote job

Research Engineer - Code Generation & Model Evaluation is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should have written, and explain the fix so the modeling team can target the gap.

Why this role matters

Engineering model quality lives or dies on whether the generated code actually compiles, passes tests, and handles edge cases. AuraOne pairs experienced engineers with the modeling team to grade outputs the way a code reviewer would.

Responsibilities

  • Run and reproduce candidate code outputs in a sandboxed environment for Research Engineer - Code Generation & Model Evaluation assignments.
  • Grade research engineer code generation model evaluation engineering review solutions for correctness, style, and edge-case handling.
  • Write minimal failing tests that demonstrate the bug a model output missed.
  • Compare paired solutions and rank them with a written rationale tied to the rubric.
  • Tag failure modes (compile error, runtime crash, off-by-one, security issue) with severity scores.
  • Document recurring code-generation failures so the modeling team can target them.
  • Calibrate against gold-standard reviews to keep inter-rater agreement above target.

Qualifications

  • Strong day-job engineering experience — you can read, run, and debug unfamiliar code for Research Engineer - Code Generation & Model Evaluation work.
  • Comfort writing concise unit tests that capture a single failure mode.
  • Familiarity with a testing framework. pytest, Jest, or JUnit. Go test, RSpec, or whatever you use.
  • Clear written reasoning — your review note has to convince another senior engineer.
  • Reliable async availability for at least 10 hours per week.
  • Prior code-review or technical-interview-grading experience is a plus.

Example tasks

  • Reproduce a generated engineering solution to a coding task, run the test suite, and grade it.
  • Write the smallest failing test that demonstrates a model's edge-case bug.
  • Compare two paired solutions and rank them with a written rationale tied to the rubric.
  • Triage a security issue surfaced by a model and document the patch the model should have produced.

Nice to have

  • Open-source contributions or a public portfolio that demonstrates production-quality code.
  • Experience with the target language's standard tooling, linters, and idiomatic style guides.
  • Familiarity with security-review checklists (OWASP, CWE) and AppSec patterns.

Skills

  • Code review
  • Debugging
  • Unit testing
  • Software engineering judgment
  • Research Engineer Code Generation Model Evaluation engineering review
  • AI model evaluation

Work model

Remote — US-eligible. Remote · Independent specialist contractor. Employment type: CONTRACTOR. Applicants must be authorized to work from US.

Compensation

Hourly rate confirmed after the interview process.

Application process

Apply through AuraOne's specialist intake for role-specific routing and review. Final project scope, schedule, and contractor terms are confirmed before placement.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Research Engineer - Code Generation & Model Evaluation [Remote] in Remote vacancy
  • $50 - $100 per hour

     ...Apply your software engineering expertise to help train and evaluate next-generation AI systems through real-world coding tasks, technical feedback, and code quality analysis. This...  ...focuses on code generation workflows and model evaluation; prior AI experience is not required... 
    Suggested
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    11 days ago
  • $50 - $100 per hour

     ...Role Title: Research Engineer - Code Generation & Model Evaluation Role Type: Contractor Location: Remote micro1 is engaging Research Engineers to participate in a project focused on code generation and model evaluation for a customer's initiative. In this role... 
    Suggested
    For contractors
    Remote work

    micro1

    Remote
    6 days ago
  •  ...solving to improve and evaluate large language models. You will design...  ...numerical results with code, and review model...  ...Accelerate frontier AI research by contributing high...  ...and annotate model generated solutions, identify...  ...level expected for engineering entrance exams and for... 
    Suggested
    Contract work
    For contractors
    Freelance
    Remote work

    SaidGig

    United States
    more than 2 months ago
  •  ...Role Overview Senior Software Engineer contractor role, remote, paying $245 to $280 per hour...  ...AI lab improve frontier AI systems by evaluating AI-generated code, creating challenging engineering problems, and supporting model performance on real-world software tasks.... 
    Suggested
    Hourly pay
    For contractors
    Remote work

    SaidGig

    United States
    29 days ago
  • $315k

    We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand...  ...Safety Level (ASL) to assign to models. Research leads on this team...  ...training infrastructure to prepare new generations of models for routine... 
    Suggested
    Currently hiring
    Work at office
    Immediate start
    Home office
    Visa sponsorship
    Relocation package

    Anthropic

    Seattle, WA
    1 day ago
  • $295k

     ...cutting-edge foundation AI models and end-to-end products that...  .... Cohere is a team of researchers, engineers, designers, and more, who are...  ...Paris. Join us!Why this role?Evaluation is critical to making progress...  ...for creating these next-generation evaluation methods and infrastructure... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    3 days ago
  •  ...Description As the Manager of Model Validation & Verification (VnV)...  ...Behavior Autonomy, you will lead an engineering and data science team responsible for evaluating, benchmarking, and validating...  ...Experience with reinforcement learning, generative AI, or distributed ML systems.... 
    Temporary work
    Relocation package

    Zoox

    Foster, CA
    12 days ago
  • Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote job
    Flexible hours

    Prolific

    Charlotte, NC
    3 days ago
  • $160k - $200k

     ...Posted: 2026-09-19 Category: Engineering and Sciences Subcategory: Modeling/Sim Engr Schedule: Full-Time...  ...coordination, and analysis results evaluation. Recommend and document force...  ...simulation scripts and interface code using government supportable software... 
    Full time
    Remote work
    Shift work

    SAIC

    Colorado Springs, CO
    17 hours ago
  •  ...Professionals in Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields and have experience in research or academia. Responsibilities include fact-checking AI technical... 
    Remote job
    Hourly pay
    Flexible hours

    Prolific

    Jacksonville, FL
    3 days ago
  • $20 per hour

     ...and technical talent with leading AI research labs. Headquartered in San Francisco...  ...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths...  ...completeness of responses. Ensure model responses align with expected conversational... 
    Remote job
    Contract work
    Part time
    Summer work

    Mercor

    New York, NY
    6 days ago
  • $192k - $304.75k

     ...now looking for a Senior Research Engineer passionate about Generative AI inference. Are you excited...  ...of generative AI models, from language to images....  ...will be doing:Design and evaluate routing policies for LLM...  ...source repo: design docs, code review, docs, and community... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $70 - $90 per hour

     ...tasks that support the training and evaluation of advanced AI models. This role focuses on assessing...  ...kernel task types: specification-based generation, cross-framework translation or...  ...or TPU. Background in compiler engineering, MLIR, or intermediate-representation... 
    Hourly pay
    Remote work

    SaidGig

    Remote
    a month ago
  • OverviewFrom our Ann Arbor, MI office, we have an opening for a Research Engineer with FPGA and low-level code development (C and C++) for real-time and high-performance computing (HPC) applications. You will participate in multi-disciplinary, collaborative teams and contribute... 
    Work at office
    Remote work

    SRI International

    Ann Arbor, MI
    2 days ago
  • $60 per hour

    Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    4 days ago
  • $250k - $290k

     ...to help them hire. Research Engineer Location San...  ...data engineering, and evaluation . This is not a traditional...  ...or experimenting with models in isolation. It is...  ...infrastructure, generate and curate datasets, and...  ..., production-quality code Ability to debug... 
    H1b
    Remote work

    Recruiting from Scratch

    San Francisco, CA
    1 day ago
  •  ...Ando Research Team Member Ando is a messaging platform...  ...track, or intervene. Evaluating that judgment means...  ...background, with depth in model evaluation,...  ...credentials. Work samples or code that demonstrate the skills...  ...hold high-bandwidth, generative technical... 
    Work from home

    ASARI S.A de CV

    San Francisco, CA
    4 days ago
  •  ...Research Engineer Colombia, Huila, Colombia Turing's...  ...frontier AI labs to generate high-quality datasets...  ...benchmarks that improve model capabilities in software...  ...into better data, evaluations, and more capable models...  ...areas: Coding and software engineering... 
    Remote work
    Shift work

    Turing

    United States
    5 days ago
  • $100 - $150 per hour

     ...Role Overview Apply your data science expertise to evaluate AI-generated slides, spreadsheets, and documents for real-world quality and usability. You will assess outputs against professional standards and deliver clear feedback that improves their accuracy and presentation... 
    Hourly pay
    Work at office
    Remote work

    SaidGig

    United States
    3 days ago
  • $60 - $90 per hour

     ...Apply hands-on mechanical engineering judgment to improve how advanced AI models reason through real-...  ...You will partner with AI research and program management...  ..., and create rigorous evaluations grounded in industry practice...  ...with applicable codes and standards, including... 
    Hourly pay
    Full time
    Remote work

    SaidGig

    United States
    7 days ago
  • $100 - $150 per hour

     ...considered for future projects evaluating how well AI systems...  ...produced analyses and models, document decisions in...  ...write-ups, feature engineering, and technical reports...  ...Evaluate AI-generated or human-created work...  ...a leading technology, research, or quantitative firm,... 
    Hourly pay
    Immediate start
    Remote work

    SaidGig

    United States
    a month ago
  •  ...holistic approach for evaluating security posture and ecosystems...  ...-conscious Innovation Engineer to join our technology...  ...and scaling our generative AI capabilities. You...  ..., Large Language Models (LLMs), prompt engineering...  ...with Infrastructure as Code (IaC) tools like AWS... 
    Local area
    Remote work
    Flexible hours

    GuidePoint Security

    United States
    4 days ago
  • $65 - $105 per hour

     ...Apply deep engineering judgment to help frontier AI models reason more accurately about real-world...  ...work closely with an AI research and program management...  ...quality engineering work, evaluate model performance, and...  ...reasoning, subtly incorrect code, unaddressed edge cases,... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    California
    a month ago
  • $81.26k - $137.14k

     ...windows team, leading research related to window...  ...analytical and modeling tools. You will...  ...energy/market/engineering analysis and provide...  ...improvement, and evaluation of the...  ..., and model data generated from energy simulations...  ...energy policy and codes. Knowledge in statistics... 
    Full time
    Work experience placement

    Lawrence Berkeley National Laboratory

    Berkeley, CA
    2 days ago
  •  ...Job Description Job Description Research Engineer – Code Generation & Model Evaluation Role Type: Contractor Location: Remote We are seeking experienced Research Engineers to contribute to a project focused on code generation, software engineering, and... 
    For contractors
    Remote work

    YO AI Labs

    New York, NY
    7 days ago
  • Python Infrastructure Engineer - Model Evaluation (AI Training) About the Role What if your Python expertise...  ...depend on to train and validate next-generation models. This is a fully remote...  ...fixes Collaborate with data, research, and engineering teams to support model... 
    Hourly pay
    Ongoing contract
    Contract work
    Freelance
    Remote work
    Flexible hours

    Alignerr

    Seattle, WA
    4 days ago
  •  ...Motors, through Embodied AI, seeks a Senior Engineer to measure and visualize AV model performance. You will design and implement evaluation workflows, collaborate across Data,...  ...influence safety and scalability of next‑generation autonomous systems and to contribute to... 
    Remote job

    General Motors

    Austin, TX
    1 day ago
  • $60 - $90 per hour

     ...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Summers , and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    a month ago
  •  ...role As a Staff Research Engineer, you will join a team...  ...edge challenges in the Generative AI space, with a...  ...interactive video diffusion models. Within the team you’...  .... Build robust evaluation frameworks and test...  ...research code. Outcome-driven... 
    Full time
    Work at office
    Remote work
    Worldwide

    Synthesia

    Remote
    16 days ago
  • $158k - $293k

     ...TR Labs, owns the engineering behind that layer: next-generation search and retrieval...  ...delivers.What makes this research engineering rather...  ...different ranking model, a hybrid retrieval...  ...to what a coding agent hands you as...  ...understandingBuild evaluation that actually discriminates... 
    Full time
    Work at office
    Local area
    Flexible hours

    Thomson Reuters

    New York, NY
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Engineer - Code Generation & Model Evaluation [Remote]. Be the first to apply!