Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer , AI Code Evaluation & Benchmarking (US candidates only)

Temporary

Turing

About Turing

Turing is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems. Turing helps leading AI labs improve the reasoning, problem-solving, and decision-making capabilities of large language models (LLMs) through high-quality human feedback, evaluation, and training data.

Role Overview

We are looking for experienced Software Engineers to help evaluate, benchmark, and improve the coding capabilities of frontier AI models. In this role, you will assess AI-generated code, validate solutions against real-world software engineering tasks, identify correctness and quality issues, and contribute to the development of high-quality evaluation datasets and benchmarks.

This position is ideal for engineers who enjoy code review, debugging, problem-solving, and applying strong software engineering judgment to complex technical scenarios. Your work will directly contribute to measuring and improving the performance of advanced AI coding systems.

What Does Day-to-Day Look Like?

  • Review and evaluate AI-generated code for correctness, efficiency, maintainability, and adherence to requirements.
  • Analyze software engineering tasks and validate whether proposed solutions meet expected outcomes.
  • Debug code, reproduce issues, and verify fixes across different programming environments.
  • Assess model-generated explanations, reasoning, and implementation approaches for technical accuracy.
  • Create, refine, and maintain evaluation datasets, benchmarks, and grading rubrics for coding tasks.
  • Identify edge cases, failure modes, and areas where AI systems struggle with software engineering problems.
  • Document findings clearly and provide structured feedback to improve evaluation quality and consistency.
  • Collaborate with project teams to establish quality standards and evaluation methodologies.

Requirements

  • Bachelor's or Master's degree in Computer Science, Software Engineering, or a related technical field.
  • 3+ years of professional software engineering experience.
  • Strong proficiency in one or more of the following languages: Python, Java, C/C++, Go, Swift, Objective-C, PHP, or SQL .
  • Strong understanding of data structures, algorithms, software design principles, and debugging methodologies.
  • Experience performing code reviews and evaluating code quality in production or large-scale codebases.
  • Ability to analyze complex technical problems and assess solution correctness with minimal supervision.
  • Familiarity with version control systems (e.g., Git) and modern software development workflows.
  • Strong written communication skills and attention to detail.
  • Experience with AI/ML data annotation, NLP, prompt engineering, model evaluation, or LLM-related projects is a plus.
  • Experience evaluating AI-generated code, benchmark creation, or software quality assessment is highly preferred.

Perks of Freelancing With Turing

  • Work in a fully remote environment.
  • Opportunity to work on cutting-edge AI projects with leading LLM companies.

Offer Details

  • Commitments Required: At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST.
  • Engagement type : Contractor assignment (no medical/paid leave)
  • Duration of contract 1 month; [expected start date is next week]
  • Location US only

Evaluation Process

  1. Online automated coding challenge for Python and Docker test (RHLF)
Vacancy posted 22 days ago
Similar jobs that could be interesting for youBased on the Software Engineer , AI Code Evaluation & Benchmarking (US candidates only) in Remote vacancy
  • $50 per hour

     ...Role Overview Evaluate, benchmark, and help improve the coding capabilities of advanced AI models by assessing AI-generated solutions against real software engineering problems. You will review code for correctness...  ...Work location, remote, candidates must be located in the... 
    Suggested
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    10 days ago
  • $80 - $100 per hour

     ...accepted countries and locations. For US applicants: This is a 1099 independent...  ...You'll Be Doing Design and build the coding benchmarks and evaluation pipelines used to test frontier AI models on real software engineering work: Design coding benchmarks that evaluate... 
    Suggested
    Remote job
    Full time
    Contract work
    For contractors

    G2i

    Miami, FL
    19 hours ago
  •  ...Overview Design and evaluate high-quality...  ...language models for code. You will work...  ...languages, assess AI-generated implementations...  ...across the software engineering lifecycle. Key...  ...training and benchmarking. Evaluate AI-generated...  ...: Remote, candidates must be based in... 
    Suggested
    For contractors
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Canada
    more than 2 months ago
  •  ...About Us Based in San Francisco...  ...for frontier AI labs and a...  ...specialize in software engineering, logical reasoning...  ...Engineering evaluator, you will create...  ...for training, benchmarking, and advancing...  ...curating code examples, providing...  ...Location: candidates must be based... 
    Suggested
    Contract work
    For contractors
    Flexible hours

    Turing

    Remote
    22 days ago
  •  ...Software Engineer III Full Time - Hybrid Cincinnati...  ...Innovate with Benchmark Gensuite as a...  ...inclusion. Join us and help make...  ...clients. The ideal candidate will perform coding, debugging,...  ...processes AI & Emerging Technology...  ...feedback Evaluate and integrate AI... 
    Suggested
    Full time
    Immediate start
    Worldwide
    Flexible hours

    Benchmark Gensuite®

    Remote
    19 hours ago
  •  ...accelerator for frontier AI labs and a trusted...  ...who specialize in software engineering, logical reasoning...  ...Day Look Like? Evaluate and refine AI-generated code to ensure that it...  ...performance benchmarks. Build agents and...  ...fit) Location: Candidates must be based in the... 
    For contractors
    Remote work
    Flexible hours

    Turing

    United States
    3 days ago
  •  ...accelerator for frontier AI labs and a...  ...specialize in software engineering, logical...  ...Day Look Like? Evaluate and refine AI-generated code to ensure that it...  ...industry performance benchmarks. Build agents...  ...) Location : Candidates must be based out of US, Canada or WEU countries... 
    Full time
    For contractors
    Remote work
    Flexible hours

    Turing

    United States
    3 days ago
  • $80 - $100 per hour

     ...Senior Python Developer (AI Evaluation & Benchmarking) An enterprise client...  ...evaluating AI-generated code, and helping improve the...  ...opportunity for experienced software engineers who enjoy solving complex...  ...the client. Qualified candidates will receive an Upwork contract... 
    Hourly pay
    Weekly pay
    Contract work
    Remote work
    10 hours per week

    Lifted, an Upwork Company™

    United States
    1 day ago
  •  ...Senior Software Engineer, AI Code Modernization   Location: Hybrid...  ...process.   Evaluate and improve AI systems...  ...simulation systems, benchmarking, or validation...  ...Like Successful candidates will demonstrate the...  ...458-5000 or sending us an email at disabilityrequest... 
    Worldwide

    Bentley Systems

    Exton, PA
    a month ago
  •  ...Senior Software Engineer — AI Coding Evaluator is a remote engineering review track for evaluating production code...  ...Run and reproduce candidate code outputs in a sandboxed environment...  ...Expert review Work model Remote — US-eligible. Remote · Independent specialist... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    a month ago
  • $139k - $204k

     ...Essential Cloud for AI™. Built for pioneers...  ...looking for a Senior Engineer for CoreWeave’s Benchmarking & Performance team....  .... You will also aid us in achieving...  ...designs and elevate coding/testing standards....  ...market rate for each candidate which can include a... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Core Weave

    Remote
    19 hours ago
  •  ...career network for the AI economy. 20 million knowledge...  ...About the Role As a Software Engineer on our Coding Pod, you will build the...  ...-scale, high-quality benchmark datasets that evaluate how models perform on real...  ...are for full-time US employees. Ownership... 
    Full time
    Freelance
    Internship
    Work at office
    Remote work
    Flexible hours

    Handshake

    Remote
    19 hours ago
  •  ...unprecedented scale. Join us to help deliver...  ...The Evaluation team builds and...  ...clear feedback for engineering and leadership,...  ...driving software performance at...  ...system design, code reviews, testing...  ...Experience leveraging AI‑assisted...  ...encourage interested candidates to review the... 
    Local area
    Work from home

    General Motors

    Brooklyn, NY
    2 days ago
  •  ...new era, we seek AI-native thinkers...  ...The Cortex Code team is building...  ...the full AI engineering lifecycle: design...  ...coding harnesses, evaluation pipelines....  ...orchestration, or software with...  ...want you to join us in building the...  ...not only one-off benchmarks. ~ Proficiency... 
    Full time

    Snowflake

    Remote
    19 hours ago
  •  ...intelligence to power the AI economy. We're...  ...Mercor's APEX benchmark family measures...  ...s platform is code —the tasks,...  ...that train and evaluate the world's...  .... As a Staff Software Engineer for Code Search...  ...embeddings and BM25 , candidate generation,...  ...that let us continuously evolve... 
    Full time
    Work at office
    Relocation package

    Mercor

    San Francisco, CA
    19 hours ago
  •  ...high-quality datasets and evaluation pipelines that improve and benchmark large language models for code generation and software engineering tasks. You will curate...  ..., evaluate and refine AI-generated solutions...  ...information. Eligibility Candidates must have the specified... 
    Full time
    For contractors
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...industry. As an AI native company, we...  ...Mainstay and help us redefine what’s...  ...full-stack senior software engineer - maximum agency,...  ...and actionable Evaluate and benchmark LLM performance...  ...engineers through code reviews, technical...  ...everyone -- from candidates to employees to... 
    Full time
    Remote work

    Mainstay Labs Inc.

    Remote
    19 hours ago
  • $65 - $105 per hour

     ...Apply deep engineering judgment to help frontier AI models reason more...  ...about real-world software development and...  ...work, evaluate model performance...  ...standards and benchmarks. Key Responsibilities...  ...subtly incorrect code, unaddressed...  ...based engagement. Candidates may discover... 
    Hourly pay
    Full time
    Freelance
    Internship
    Live in
    Relocation
    Relocation package

    SaidGig

    Bay County, FL
    21 days ago
  • $173k - $251k

     ...looking for a Senior Software Engineer, AI Platform with...  ...orchestration, evaluation, and agentic patterns...  ..., write great code, learn a huge...  ...opportunity ahead of us, but don’t rely...  ...in the “perfect” candidate and encourage you...  ...relevance tuning, benchmarking, and multi-modal... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Flexible hours
    2 days per week
    3 days per week

    Everlaw

    Remote
    19 hours ago
  • $190.8k - $267.1k

     ...As a Senior Engineer on this team,...  ...world.  Join us and help...  ...have: ~ Software development experience...  ...well-tested code. ~ Media...  ...transparency to candidates, we share...  ...country location, benchmarked against...  ...intelligence (AI). You will have...  ...to evaluate your application... 
    Full time
    For contractors
    Work experience placement
    Remote work
    Flexible hours

    Reddit

    Remote
    19 hours ago
  •  ...future of agent-centric software development. As the leader in AI code verification and governance...  ...Motor Company count on us to provide independent, explainable...  ...looking for a Software Engineer to join the team building...  ...office worth having.  Candidates need to be genuinely... 
    Full time
    Work at office
    Relocation

    Sonar

    Remote
    19 hours ago
  • $204k - $259k

     ...-of-the-art Generative AI to create a training ground...  ...Driver. The Simulator Evaluation team faces the ultimate...  ...looking for a Senior Software Engineer to build the metrics...  ...write clean, testable code that is built to last....  ...full-time position across US locations is listed below... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    19 hours ago
  • $175k - $215k

     ...-of-the-art Generative AI to create a training ground...  ...Driver. The Simulator Evaluation team faces the ultimate...  ...We are looking for a Software Engineer to build the metrics...  ...engineering and AI, writing the code that processes...  ...full-time position across US locations is listed... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    19 hours ago
  • $170k - $216k

     ...Waymo Driver. Our software allows the Waymo Driver...  ...enabling software engineers like you to develop...  ...automation and evaluation frameworks that establish...  ...in industrial AI applications involving...  ...and efficient code We Prefer MS...  ...time position across US locations is listed... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    19 hours ago
  •  ...Software Engineer, Rapid Response Assignment - (US) Location: MA / Austin TX / MD Team: Product &...  ...pipelines, automation, and AI and LLM workflows your...  ...re primarily looking for candidates based in Massachusetts, Maryland...  ...-quality, maintainable code and uphold engineering... 
    Remote job
    Full time
    Local area
    Home office
    Flexible hours

    Vulncheck

    United States
    19 hours ago
  •  ...Kubernetes-native AI infrastructure company...  ...empowers platform engineering teams to deliver...  ...and build the software that provisions, integrates...  ...Go. The right candidate has built...  ...with strong testing, code review, and operational...  ...decisions about evaluation and review connected... 
    Remote job
    Full time
    Local area

    Mirantis Inc.

    Remote
    19 hours ago
  •  ...the Role As a Senior Software Engineer at LemonEdge, you’ll...  ...to build on it. We use AI-assisted development...  ...This role is open to candidates based in the United States...  ...bar through thorough code review, design...  ...development — usage policies, evaluation of generated code,... 
    Remote job
    Full time
    Work experience placement
    Work at office
    Shift work

    Lemonedge Technology Ltd

    United States
    19 hours ago
  •  ...About Us At Pickle Robot...  ...with Physical AI. Our robots...  ...Dill Autonomy Engine: generalized...  ...Ground-Truth Evaluation: Design and...  ...beyond static benchmarks to guarantee...  ...through rigorous code reviews,...  ...Python and modern software engineering...  ...often consider candidates at different... 
    Full time
    Local area
    3 days per week

    Pickle Robot Company

    Remote
    19 hours ago
  • $160k - $200k

     ...Let’s be real, AI isn’t magic; Legion...  ...and doers to join us in shaping the future...  ...Machine Learning Engineer (Remote USA) **...  ...ID*** Senior Software Engineer, Machine...  ...structured generation, evaluation, and other...  ...evaluation frameworks, benchmarks, or datasets for AI... 
    Remote job
    Permanent employment
    Full time
    Contract work
    Local area

    Yurts

    United States
    19 hours ago
  • $60 - $100 per hour

     ...technical talent with leading AI research labs....  ...our investors include Benchmark , General Catalyst ,...  ...Dorsey . Position: Software Engineering, Data Science, and Systems...  ...Responsibilities Evaluate LLM-generated responses to coding and software engineering... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    19 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer , AI Code Evaluation & Benchmarking (US candidates only). Be the first to apply!