Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer for AI Code Evaluation and Benchmarking

$50 per hour

SaidGig

Role Overview

Evaluate and benchmark the coding abilities of frontier AI models by reviewing AI-generated solutions, validating them against real-world software engineering tasks, and shaping high-quality evaluation datasets and benchmarks. This role suits experienced engineers who enjoy code review, debugging, and applying software engineering judgment to improve model correctness and reliability.

Key Responsibilities
  • Review AI-generated code for correctness, efficiency, maintainability, and compliance with task requirements.
  • Analyze software engineering tasks and validate whether proposed solutions meet expected outcomes.
  • Debug code, reproduce issues, and verify fixes across multiple programming environments.
  • Evaluate model-generated explanations and reasoning for technical accuracy and soundness.
  • Create, refine, and maintain evaluation datasets, benchmarks, and grading rubrics for coding tasks.
  • Identify edge cases, failure modes, and areas where models struggle with software engineering problems.
  • Document findings clearly and provide structured feedback to improve evaluation consistency and quality.
  • Collaborate with project teams to establish and maintain quality standards and evaluation methodologies.
Qualifications
  • Bachelor''s or Master’s degree in Computer Science, Software Engineering, or a related technical field.
  • Minimum 3 years of professional software engineering experience.
  • Strong proficiency in one or more of the following languages: Python, Java, C/C++, Go, Swift, Objective-C, PHP, or SQL.
  • Solid understanding of data structures, algorithms, software design principles, and debugging methodologies.
  • Experience performing code reviews and evaluating code quality in production or large-scale codebases.
  • Ability to analyze complex technical problems and assess solution correctness with minimal supervision.
  • Familiarity with version control systems such as Git and with modern software development workflows.
  • Strong written communication skills and attention to detail.
  • Experience with AI or ML data annotation, NLP, prompt engineering, model evaluation, or LLM-related projects is a plus.
  • Experience evaluating AI-generated code, creating benchmarks, or assessing software quality is highly preferred.
Work Terms
  • Engagement type: Contractor assignment, no medical or paid leave provided.
  • Minimum commitment: at least 4 hours per day and at least 20 hours per week, with a required daily overlap of 4 hours aligned to PST.
  • Contract length: 1 month, expected start date is next week.
  • Work location: Remote, United States only.
  • Perks: fully remote work and the opportunity to contribute to cutting-edge AI coding evaluation projects.
Eligibility
  • Applicants must be located in the United States. This role is open to US-based candidates only.
Evaluation Process

Candidates will complete an online automated coding assessment covering Python and a Docker-based test identified as RHLF as part of the selection process.

Vacancy posted 7 days ago
Similar jobs that could be interesting for youBased on the Software Engineer for AI Code Evaluation and Benchmarking in United States vacancy
  •  ...Role Overview Evaluate and benchmark the coding abilities of frontier AI models by reviewing AI-generated solutions, validating them against real-world software engineering tasks, and shaping high-quality evaluation datasets and benchmarks. This role suits experienced... 
    Suggested
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    10 days ago
  •  ...Surge was founded by engineers and researchers who dreamed...  ...building the next generation AI. We're building a...  ...to conducting rigorous evaluations that go beyond benchmarks. We've run a profitable...  .... The Role As a Software Engineer, Coding Evaluation & Training Data... 
    Suggested
    Full time

    Surge AI

    United States
    1 day ago
  • $80 - $100 per hour

     ...employment verification for this role. What You'll Be Doing Design and build the coding benchmarks and evaluation pipelines used to test frontier AI models on real software engineering work: Design coding benchmarks that evaluate frontier models on real-world... 
    Suggested
    Remote job
    Full time
    Contract work
    For contractors

    G2i

    Miami, FL
    1 day ago
  • $40 per hour

     ...specialists with project-based AI opportunities for...  ..., focused on testing, evaluating, and improving AI...  ...dataset to evaluate AI coding agents - how well a model...  .... Not prompt engineering. Not writing code from...  ...Qualifications ~5+ years in software development. ~Core... 
    Suggested
    Part time

    Mindrift

    Remote
    3 days ago
  • $90 per hour

     ...Role Overview Software Engineers evaluate AI-generated code and technical content, providing structured, experience-based feedback to improve how AI systems understand programming tasks, system design, and engineering best practices. This is an ongoing, project-based,... 
    Suggested
    Hourly pay
    Ongoing contract
    Contract work
    Part time
    Internship
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $25 per hour

     ...Role Overview Software Engineers apply production software development experience to evaluate AI-generated code and technical content, provide structured feedback, and improve how AI models understand programming tasks, system design, and engineering best practices. This... 
    Hourly pay
    Contract work
    Part time
    Internship
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  • $80 - $100 per hour

     ...client who is a leading AI platform that enables organizations...  ...human feedback, AI evaluation, and model alignment....  ...designing programming benchmarks, evaluating AI-generated code, and helping improve the...  ...for experienced software engineers who enjoy solving complex... 
    Hourly pay
    Weekly pay
    Contract work
    Remote work
    10 hours per week

    Lifted, an Upwork Company™

    Texas City, TX
    a month ago
  • $60 - $100 per hour

     ...technical talent with leading AI research labs....  ...our investors include Benchmark , General Catalyst ,...  ...Dorsey . Position: Software Engineering, Data Science, and Systems...  ...Responsibilities Evaluate LLM-generated responses to coding and software engineering... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    1 day ago
  • $50 - $150 per hour

     ...technical talent with leading AI research labs....  ...Francisco, our investors include Benchmark , General Catalyst , Peter...  ...Dorsey . Position: Software Engineering Expert Type: Contract...  ...Write and evaluate code across diverse programming... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    1 day ago
  •  ...Senior Software Engineer, AI Code Modernization   Location: Hybrid – Exton, PA/Philadelphia...  ...conversion/test/correction process.   Evaluate and improve AI systems for accuracy,...  ...frameworks, simulation systems, benchmarking, or validation tooling. Experience... 
    Worldwide

    Bentley Systems

    Exton, PA
    7 days ago
  •  ...development of high-quality datasets and evaluation pipelines that improve and benchmark large language models for code generation and software engineering tasks. You will curate and author reference code, evaluate and refine AI-generated solutions across multiple programming... 
    Full time
    For contractors
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $225k - $250k

     ...5, we started Handshake AI and built the fastest-growing...  ...researchers to create evaluations, publish benchmarks, and push the boundary...  ...Work together with engineers, scientists, operators,...  ...the Role As a Senior Software Engineer on our Coding Pod, you'll lead the design... 
    Full time
    Work at office
    Flexible hours

    Handshake

    San Francisco, CA
    27 days ago
  • $35 - $120 per hour

     ...technical talent with leading AI research labs....  ..., our investors include Benchmark , General Catalyst ,...  ...Dorsey . Position: Code-Data Eval Author — Software Engineer Type: Contract...  ...structured review. Evaluate the accuracy and depth of... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    1 day ago
  • $70 - $120 per hour

     ...technical talent with leading AI research labs....  ...our investors include Benchmark , General Catalyst...  ...Dorsey . Position: Software Engineer Type:...  ...multiple categories, such as code review, debugging, and...  ...with AI labs to enhance evaluation datasets for AI... 
    Full time
    Summer work
    Freelance
    Immediate start

    Mercor

    Remote
    1 day ago
  • $90 per hour

     ...Software Engineers evaluate AI-generated code and technical content, provide structured feedback, and help improve how AI understands programming tasks, system design, and engineering best practices. Role Overview This is an ongoing, project-based, hourly contract role... 
    Hourly pay
    Ongoing contract
    Full time
    Contract work
    Part time
    Internship
    Remote work
    Flexible hours

    SaidGig

    United States
    3 days ago
  • $25 per hour

     ...Role Overview Software Engineers based in India will evaluate AI-generated code and technical content, provide structured feedback, and help improve how AI models understand programming tasks, system design, and engineering best practices. This is a part-time, project... 
    Hourly pay
    Part time
    For contractors
    Internship
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  •  ...Lead the design and evaluation of next-generation coding agents by creating benchmarks, measurement methodologies, datasets,...  ...model performance across diverse software engineering tasks. Develop high-...  ...researchers, engineers, and applied AI teams to design experiments... 
    Full time
    Remote work

    SaidGig

    United States
    more than 2 months ago
  •  ...Senior Software Engineer — AI Coding Evaluator is a remote engineering review track for evaluating production code, debugging traces, and developer-facing AI outputs against real-world correctness standards. Reviewers reproduce failures, write the unit test the model should... 
    Remote job
    Hourly pay
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    3 days ago
  • $175k - $215k

     ...state-of-the-art Generative AI to create a training ground for...  ...Waymo Driver. The Simulator Evaluation team faces the ultimate data...  ...real"? We are looking for a Software Engineer to build the metrics and pipelines...  ...and AI, writing the code that processes petabytes of simulation... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  •  ...organize human intelligence to power the AI economy. We partner with leading AI...  ...context that can't be captured in code alone. Today, more than 30,000 experts...  ...About the Role As a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at the core of... 
    Full time
    Work at office
    Relocation package

    Mercor

    San Francisco, CA
    1 day ago
  • $170k - $216k

     ...the Waymo Driver. Our software allows the Waymo Driver...  ...sensors, enabling software engineers like you to develop...  ...an automated system Evaluate new hardware specifications...  ...in industrial AI applications involving...  ...with robust and efficient code We prefer: MS... 
    Full time
    Remote work

    Waymo

    Remote
    1 day ago
  • $320k

     ...interpretable, and steerable AI systems. We want AI to...  ...committed researchers, engineers, policy experts, and...  ...looking for a Software Engineer to help us build...  ...and maintain new agentic coding tools for developers. The...  ...experimenting with new tools, evaluating emerging techniques,... 
    Full time
    Work experience placement
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Seattle, WA
    1 day ago
  • $180k - $240k

     ...Description Deepgram is looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team...  ...fail criteria grounded in Research benchmarks. ~Build monitoring systems that...  ...continuously. ~Help raise the bar through code reviews, technical design... 
    Full time

    Deepgram

    Remote
    2 days ago
  • $170k - $216k

     ...the Waymo Driver. Our software allows the Waymo Driver...  ...sensors, enabling software engineers like you to develop...  ...critical automation and evaluation frameworks that establish...  ...in industrial AI applications involving...  ...with robust and efficient code We Prefer MS or... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $144.7k - $221.4k

     ...About the Organization The Evaluation team builds and evolves...  ...clear feedback for engineering and leadership, and help...  ...autonomous driving software performance at interfaces...  ...thoughtful system design, code reviews, testing,...  ...Experience leveraging AI-assisted development and... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $129.4k - $198.4k

     ...About the Organization:The Evaluation team builds and evolves...  ...for autonomous vehicle software validation. Develop and...  ...insights to engineering teams and leadership, including...  ...high standards for code quality and software architecture...  ...practices. Leverage AI-assisted development... 
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, TX
    3 days ago
  • $152k - $241.5k

     ...tapping into the unlimited potential of AI to define the next era of computing....  ...impact on the world.We are seeking a Software Engineer - Scientific Evaluation to own a shared platform for classical testing, scientific benchmarking, and agentic evaluation. The portfolio... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $147k - $211k

     ...product or system development code.Participate in, or lead...  ...years of experience with software development in one or more...  ...of labeling, ratings, and evaluations.Experience with Natural Language...  ...Learning, or Generative AI.Google's software engineers develop the next-... 

    Google

    Mountain View, CA
    1 day ago
  • AI Systems Modernization DeveloperBentley Systems...  ...dedicated expert team in code modernization. This...  ...support and guide other software developers in the company...  ...: Manual evaluation of the quality of the conversion...  ...solutions for architecture, engineering, and construction -... 
    Worldwide

    Bentley Systems

    Exton, PA
    2 days ago
  • $224k - $356.5k

     ...autonomous driving, and evaluation is how we know the drive...  ...our organization develops AI drivers!We are looking for a senior engineer to own the engine that...  ...end by this person, in code and in the room, not from...  ...experience).12+ years building software, with significant time... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer for AI Code Evaluation and Benchmarking. Be the first to apply!