Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, AI Code Evaluation and Benchmarking

$50 per hour

SaidGig

Role Overview

Evaluate and benchmark the coding abilities of frontier AI models by reviewing AI-generated solutions, validating them against real-world software engineering tasks, and shaping high-quality evaluation datasets and benchmarks. This role suits experienced engineers who enjoy code review, debugging, and applying software engineering judgment to improve model correctness and reliability.

Key Responsibilities
  • Review AI-generated code for correctness, efficiency, maintainability, and compliance with task requirements.
  • Analyze software engineering tasks and validate whether proposed solutions meet expected outcomes.
  • Debug code, reproduce issues, and verify fixes across multiple programming environments.
  • Evaluate model-generated explanations and reasoning for technical accuracy and soundness.
  • Create, refine, and maintain evaluation datasets, benchmarks, and grading rubrics for coding tasks.
  • Identify edge cases, failure modes, and areas where models struggle with software engineering problems.
  • Document findings clearly and provide structured feedback to improve evaluation consistency and quality.
  • Collaborate with project teams to establish and maintain quality standards and evaluation methodologies.
Qualifications
  • Bachelor''s or Master’s degree in Computer Science, Software Engineering, or a related technical field.
  • Minimum 3 years of professional software engineering experience.
  • Strong proficiency in one or more of the following languages: Python, Java, C/C++, Go, Swift, Objective-C, PHP, or SQL.
  • Solid understanding of data structures, algorithms, software design principles, and debugging methodologies.
  • Experience performing code reviews and evaluating code quality in production or large-scale codebases.
  • Ability to analyze complex technical problems and assess solution correctness with minimal supervision.
  • Familiarity with version control systems such as Git and with modern software development workflows.
  • Strong written communication skills and attention to detail.
  • Experience with AI or ML data annotation, NLP, prompt engineering, model evaluation, or LLM-related projects is a plus.
  • Experience evaluating AI-generated code, creating benchmarks, or assessing software quality is highly preferred.
Work Terms
  • Engagement type: Contractor assignment, no medical or paid leave provided.
  • Minimum commitment: at least 4 hours per day and at least 20 hours per week, with a required daily overlap of 4 hours aligned to PST.
  • Contract length: 1 month, expected start date is next week.
  • Work location: Remote, United States only.
  • Perks: fully remote work and the opportunity to contribute to cutting-edge AI coding evaluation projects.
Eligibility
  • Applicants must be located in the United States. This role is open to US-based candidates only.
Evaluation Process

Candidates will complete an online automated coding assessment covering Python and a Docker-based test identified as RHLF as part of the selection process.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Software Engineer, AI Code Evaluation and Benchmarking in United States vacancy
  •  ...of the world’s fastest-growing AI companies, accelerating the...  ...high-quality human feedback, evaluation, and training data. Role Overview...  ...are looking for experienced Software Engineers to help evaluate, benchmark, and improve the coding capabilities of frontier AI... 
    Suggested
    Contract work
    Temporary work
    For contractors
    Freelance
    Remote work

    Turing

    Remote
    19 days ago
  •  ...Role Overview Evaluate and benchmark the coding abilities of frontier AI models by reviewing AI-generated solutions, validating them against real-world software engineering tasks, and shaping high-quality evaluation datasets and benchmarks. This role suits experienced... 
    Suggested
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    a month ago
  •  ...Surge was founded by engineers and researchers who dreamed...  ...building the next generation AI. We're building a...  ...to conducting rigorous evaluations that go beyond benchmarks. We've run a profitable...  .... The Role As a Software Engineer, Coding Evaluation & Training Data... 
    Suggested
    Full time

    Surge AI

    United States
    a month ago
  •  ...Responsibilities Review and refine AI-generated prompts, responses, and code Validate algorithms and software concepts for technical...  ...or language Support benchmarking efforts to evaluate and compare model...  ...experience in software engineering, technical research, or... 
    Suggested
    Part time
    Remote work

    Crossing Hurdles

    United States
    1 day ago
  • $80 - $100 per hour

     ...client who is a leading AI platform that enables organizations...  ...human feedback, AI evaluation, and model alignment....  ...designing programming benchmarks, evaluating AI-generated code, and helping improve the...  ...for experienced software engineers who enjoy solving complex... 
    Suggested
    Hourly pay
    Weekly pay
    Contract work
    Remote work
    10 hours per week

    Lifted, an Upwork Company™

    Texas City, TX
    a month ago
  • $60 - $100 per hour

     ...technical talent with leading AI research labs....  ...our investors include Benchmark , General Catalyst ,...  ...Dorsey . Position: Software Engineering, Data Science, and Systems...  ...Responsibilities Evaluate LLM-generated responses to coding and software engineering... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    more than 2 months ago
  •  ...new era, we seek AI-native thinkers across...  ...The Cortex Code team is building...  ...ll own the full AI engineering lifecycle: design,...  ...coding harnesses, evaluation pipelines. Build...  ...orchestration, or software with substantial state...  ...—not only one-off benchmarks. ~ Proficiency... 
    Full time

    Snowflake

    Remote
    a month ago
  • $50 - $150 per hour

     ...technical talent with leading AI research labs....  ...Francisco, our investors include Benchmark , General Catalyst , Peter...  ...Dorsey . Position: Software Engineering Expert Type: Contract...  ...Write and evaluate code across diverse programming... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    more than 2 months ago
  • $35 - $120 per hour

     ...technical talent with leading AI research labs....  ..., our investors include Benchmark , General Catalyst ,...  ...Dorsey . Position: Code-Data Eval Author — Software Engineer Type: Contract...  ...structured review. Evaluate the accuracy and depth of... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Miami, FL
    more than 2 months ago
  • $100 per hour

    Senior AI Interaction Evaluator (Codex / Claude Code) Contract | $100-$200/hour | 10-20 hrs/week | Start ASAP (through early May) Check out this Loom...  ...details! We're looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions... 
    Contract work
    Immediate start

    G2i Inc.

    Brooklyn, NY
    5 days ago
  •  ...Senior Software Engineer — AI Coding Evaluator is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    a month ago
  •  ...companies, working with frontier AI labs to accelerate model...  ...high-quality training data, evaluations, and engineering talent. About the Role...  ...looking for experienced, hands-on software engineers to help evaluate and improve AI coding models . Rather than primarily... 
    Hourly pay
    Temporary work
    For contractors
    Freelance
    Immediate start
    Remote work

    Turing

    Remote
    19 days ago
  •  ...We are looking for experienced Senior Software Engineers to contribute technical expertise to an advanced AI training and evaluation project. This opportunity is ideal for software...  ...improve AI systems by creating realistic coding challenges, evaluating technical... 

    eDataBae

    United States
    4 days ago
  • $175k - $215k

     ...state-of-the-art Generative AI to create a training ground for...  ...Waymo Driver. The Simulator Evaluation team faces the ultimate data...  ...real"? We are looking for a Software Engineer to build the metrics and pipelines...  ...and AI, writing the code that processes petabytes of simulation... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    more than 2 months ago
  • $204k - $259k

     ...state-of-the-art Generative AI to create a training ground for...  ...Waymo Driver. The Simulator Evaluation team faces the ultimate data...  ...We are looking for a Senior Software Engineer to build the metrics and systems...  ...: you write clean, testable code that is built to last. Data... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    more than 2 months ago
  • $139k - $204k

     ...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers,...  ...role  We’re looking for a Senior Engineer for CoreWeave’s Benchmarking & Performance team. You will have an...  ...cross-team designs and elevate coding/testing standards. Help ensure reproducible... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Core Weave

    Montana
    more than 2 months ago
  • $170k - $216k

     ...the Waymo Driver. Our software allows the Waymo Driver...  ...sensors, enabling software engineers like you to develop...  ...critical automation and evaluation frameworks that establish...  ...in industrial AI applications involving...  ...with robust and efficient code We Prefer MS or... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    more than 2 months ago
  •  ...organize human intelligence to power the AI economy. We partner with leading AI...  ...context that can't be captured in code alone. Today, more than 30,000 experts...  ...About the Role As a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at the core of... 
    Full time
    Work at office
    Relocation package

    Mercor

    San Francisco, CA
    more than 2 months ago
  • $405k

     ...interpretable, and steerable AI systems. We want AI to...  ...committed researchers, engineers, policy experts, and...  ...re looking for a Staff Software Engineer to set...  ...research on the Claude Code team. In this role, you...  ...systems, tooling, and evaluation infrastructure that determine... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    more than 2 months ago
  • $388k

     ...part of what’s next. AI and ML power innovation...  ...for a hands-on senior engineer to build the frameworks...  ...platform, model performance, evaluation, and vendor integration...  ...need: Experience in software, AI/ML, or platform...  ...single use case. Strong coding skills in Python and at... 
    Hourly pay
    Full time
    Immediate start
    Flexible hours

    Netflix

    Los Gatos, CA
    8 days ago
  •  ...introspect autonomous driving software performance at...  ...developers and systems engineers. Design and implement...  ...autonomy stack, including evaluation of perception,...  ...thoughtful system design, code reviews, testing, observability...  ...Experience leveraging AI‑assisted development... 
    Local area
    Work from home

    General Motors

    Denver, CO
    2 days ago
  •  ...intelligence to power the AI economy. We're a...  ...models. Mercor's APEX benchmark family measures AI's real...  ...Mercor's platform is code —the tasks, problems,...  ...solutions that train and evaluate the world's frontier...  ...models. As a Staff Software Engineer for Code Search & Retrieval... 
    Full time
    Work at office
    Relocation package

    Mercor

    San Francisco, CA
    a month ago
  • $50 - $150 per hour

    A fast-growing AI company in New York is seeking a contractor for a software engineering role focused on improving LLM performance. Responsibilities...  ...and reviewing model-generated code. Ideal candidates will have...  ...scalable applications and evaluating code quality. The position... 
    Hourly pay
    For contractors
    10 hours per week
    Flexible hours

    Turing Inc

    New York, NY
    1 day ago
  •  ...accelerator for frontier AI labs and a trusted...  ...researchers who specialize in software engineering, logical reasoning, STEM...  ...Typical Day Look Like? Evaluate and refine AI-generated code to ensure that it is...  ...against industry performance benchmarks. Build agents and... 
    Remote job
    For contractors
    Flexible hours

    Turing

    New York, NY
    3 days ago
  • Job Title: Software Engineer Job Location: Cambridge, MA Duration: 12 Months Designs...  ...potential biases. Uses existing AI tools to analyze their effectiveness in code generation, and application to...  ...models and software systems. Trains, evaluates and optimizes performance and... 
    Work experience placement
    Relocation

    LanceSoft

    Cambridge, MA
    2 days ago
  • $85 per hour

     ...technical talent with leading AI research labs. Headquartered...  ..., our investors include Benchmark , General Catalyst , Peter...  .... Position: Frontend Engineer (Coding Agent Experience) Type...  ...coding agents to complete and evaluate complex engineering tasks.... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    more than 2 months ago
  • $119.8k - $234.7k

    Overview Visual Studio Code is where millions of developers...  ...turn ideas into working software. We are looking for a Senior Software Engineer to help build its next...  ...to delegating work to AI. This is hands-on...  ...issues, and use telemetry, evaluations, and user feedback to guide... 
    Ongoing contract
    Local area
    Remote work

    Microsoft Corporation

    Redmond, WA
    3 days ago
  •  ...research accelerator for frontier AI labs and a trusted partner...  ...who specialize in coding, reasoning, STEM, multilinguality...  ...L Role Overview As a Software Engineering evaluator, you will create cutting-edge datasets for training, benchmarking, and advancing large... 
    Full time
    Temporary work
    For contractors
    Flexible hours

    Turing

    Remote
    19 days ago
  •  ...Senior Software Engineer, AI Code ModernizationLocation: Hybrid – Exton, PA/PhiladelphiaPosition SummaryBentley...  ...conversion/test/correction process.Evaluate and improve AI systems for accuracy,...  ...frameworks, simulation systems, benchmarking, or validation tooling.Experience... 
    Worldwide

    Bentley Systems

    Exton, PA
    a month ago
  •  ...development of high-quality datasets and evaluation pipelines that improve and benchmark large language models for code generation and software engineering tasks. You will curate and author reference code, evaluate and refine AI-generated solutions across multiple programming... 
    Full time
    For contractors
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, AI Code Evaluation and Benchmarking. Be the first to apply!