Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer - AI Evaluation

$60 - $100 per hour
Full-time

Mercor

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .

Position: Software Engineering, Data Science, and Systems Design Experts

Type: Contract

Compensation: $60–$100/hour

Location: Remote

Role Responsibilities

  • Evaluate LLM-generated responses to coding and software engineering queries for accuracy, reasoning, clarity, and completeness.
  • Conduct fact-checking using trusted public sources and authoritative references.
  • Conduct accuracy testing by executing code and validating outputs using appropriate tools .
  • Annotate model responses by identifying strengths, areas of improvement, and factual or conceptual inaccuracies.
  • Assess code quality, readability, algorithmic soundness, and explanation quality.
  • Ensure model responses align with expected conversational behavior and system guidelines.

Qualifications

Must-Have

  • BS, MS, or PhD in Computer Science or a closely related field .
  • Significant (3+ years) real-world experience in software engineering or related technical roles.
  • Expert in at least two relevant programming languages (e.g., Python, Java, C++, C, JavaScript, Go, Rust, Ruby, SQL, Powershell, Bash, Swift, Kotlin, R, TypeScript, HTML/CSS ).
  • Able to solve HackerRank or LeetCode Medium and Hard–level problems independently .
  • Experience contributing to well-known open-source projects, including merged pull requests.
  • Significant experience using LLMs while coding and understanding their strengths and failure modes.
  • Strong attention to detail and comfortable evaluating complex technical reasoning , identifying subtle bugs or logical flaws.

Preferred

  • Prior experience with RLHF , model evaluation, or data annotation work.
  • Track record in competitive programming.
  • Experience reviewing code in production environments.
  • Familiarity with multiple programming paradigms or ecosystems.
  • Experience explaining complex technical concepts to non-expert audiences.

Application Process (Takes 20–30 mins to complete)

  • Upload resume
  • AI interview based on your resume
  • Submit form

Resources & Support

PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.

Vacancy posted 22 hours ago
Similar jobs that could be interesting for youBased on the Software Engineer - AI Evaluation in Remote vacancy
  • $175k - $215k

     ...dynamics, and state-of-the-art Generative AI to create a training ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge: How...  ...virtual world is "real"? We are looking for a Software Engineer to build the metrics and pipelines that... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    22 hours ago
  • $170k - $216k

     ...powers the Waymo Driver. Our software allows the Waymo Driver to perceive...  ...sensors, enabling software engineers like you to develop multi-...  ...-critical automation and evaluation frameworks that establish the...  ...of experience in industrial AI applications involving the creation... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    22 hours ago
  • $170k - $216k

     ...powers the Waymo Driver. Our software allows the Waymo Driver to perceive...  ...sensors, enabling software engineers like you to develop multi-...  ...using an automated system Evaluate new hardware specifications...  ...of experience in industrial AI applications involving the creation... 
    Suggested
    Full time
    Remote work

    Waymo

    Remote
    22 hours ago
  • $120k - $250k

     ...2016 in Silicon Valley, Pony.ai has quickly become a global leader...  ..., and multi-dimensional evaluation. Design and implement high...  ...Build and optimize downstream engineering workflows for Large Language...  ...skills in C/C++, Python, and software design Strong foundation in... 
    Suggested
    Full time
    Temporary work

    Pony Ai

    Remote
    22 hours ago
  • $144.7k - $221.4k

     .... About the Organization The Evaluation team builds and evolves the evaluation...  ...into clear feedback for engineering and leadership, and help...  ...introspect autonomous driving software performance at interfaces across...  .... Experience leveraging AI-assisted development and analytics... 
    Suggested
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    2 days ago
  • $184k - $287.5k

     ...are looking for an experienced engineer to join our Planning and...  ...team to work on metrics and evaluation. In this role you will enable...  ...evaluation of our Autonomous Vehicle software.Build compelling, data driven...  ...vacancy. NVIDIA uses AI tools in its recruiting processes... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $129.4k - $198.4k

     .... About the Organization:The Evaluation team builds and evolves the evaluation...  ...used for autonomous vehicle software validation. Develop and...  ...interpretable insights to engineering teams and leadership, including...  ...best practices. Leverage AI-assisted development tools and... 
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $224k - $356.5k

     ...future of autonomous driving, and evaluation is how we know the drive is...  ...changing how our organization develops AI drivers!We are looking for a senior engineer to own the engine that makes it...  ...equivalent experience).12+ years building software, with significant time in... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $80 - $100 per hour

     ...verification for this role. What You'll Be Doing Design and build the coding benchmarks and evaluation pipelines used to test frontier AI models on real software engineering work: Design coding benchmarks that evaluate frontier models on real-world programming... 
    Remote job
    Full time
    Contract work
    For contractors

    G2i

    Miami, FL
    22 hours ago
  • $40 per hour

     ...specialists with project-based AI opportunities for leading...  ...companies, focused on testing, evaluating, and improving AI systems. Participation...  ...data labeling. Not prompt engineering. Not writing code from...  ...~5+ years in software development. ~Core stack:... 
    Part time

    Mindrift

    Remote
    3 days ago
  • $90 per hour

     ...Role Overview Software Engineers evaluate AI-generated code and technical content, providing structured, experience-based feedback to improve how AI systems understand programming tasks, system design, and engineering best practices. This is an ongoing, project-based... 
    Hourly pay
    Ongoing contract
    Contract work
    Part time
    Internship
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...development of high-quality datasets and evaluation pipelines that improve and benchmark large...  ...language models for code generation and software engineering tasks. You will curate and author reference code, evaluate and refine AI-generated solutions across multiple programming... 
    Full time
    For contractors
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...Role Overview Design and evaluate high-quality datasets and evaluations that advance...  ...corrections across multiple languages, assess AI-generated implementations for...  ...measure model capabilities across the software engineering lifecycle. Key Responsibilities Curate... 
    For contractors
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Canada
    more than 2 months ago
  •  ...Role Overview Evaluate and benchmark the coding abilities of frontier AI models by reviewing AI-generated solutions, validating them against real-world software engineering tasks, and shaping high-quality evaluation datasets and benchmarks. This role suits experienced... 
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    10 days ago
  • $25 per hour

     ...Role Overview Software Engineers apply production software development experience to evaluate AI-generated code and technical content, provide structured feedback, and improve how AI models understand programming tasks, system design, and engineering best practices.... 
    Hourly pay
    Contract work
    Part time
    Internship
    Remote work
    Flexible hours

    SaidGig

    United States
    a month ago
  • $35 - $120 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Position: Code-Data Eval Author — Software Engineer Type: Contract Compensation...  ...quality through structured review. Evaluate the accuracy and depth of AI-generated... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    22 hours ago
  • $100 per hour

     ...Role Overview Use your software engineering expertise to shape next-generation AI systems by reviewing, refining, and evaluating AI-generated technical content. In this remote, part-time contractor role you will create high-quality prompts, documentation, and rubric-based... 
    Hourly pay
    Part time
    For contractors
    Remote work

    SaidGig

    Remote
    8 days ago
  • $127k - $223k

     ...Description Waabi, founded by AI visionary Raquel Urtasun, is...  .... To learn more visit: The Evaluation Algorithms team is responsible...  ...realistic closed-loop simulation engine built with the latest in...  ...Python programming and strong software engineering fundamentals with... 
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    27 days ago
  •  ...Lead the design and evaluation of next-generation coding agents by creating benchmarks,...  ...coding model performance across diverse software engineering tasks. Develop high-quality datasets...  ...with researchers, engineers, and applied AI teams to design experiments and... 
    Full time
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $170k - $216k

     ...teams that consume map data. In this role, you will: Evaluate and launch machine learning models that automatically...  ...work on various parts of the system ~ A passion for good software engineering It's a Bonus if you have: M.S. or Ph.D in Computer... 
    Full time
    Remote work

    Waymo

    Remote
    22 hours ago
  • $180k - $240k

    Role Description Deepgram is looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team responsible for validating the quality of our speech, audio, and multilingual models before they reach customers. This team owns the evaluation and quality... 
    Full time

    Deepgram

    Remote
    2 days ago
  • $30 per hour

    A technology company is seeking a Web Platform Engineer to evaluate AI chatbots and enhance model performance. This role requires proficiency in programming languages like Python and JavaScript. You will assess AI outputs from coding challenges and writing tasks, ensuring... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Jackson, MS
    1 day ago
  • $30 - $40 per hour

    An AI training company is seeking a Web Platform Engineer to evaluate AI chatbots' outputs and improve their logic. The role allows for remote work and on-demand project selection, paying $30-$40+ per hour. Candidates should be fluent in English and have experience with... 
    Hourly pay
    For contractors
    Remote work

    DataAnnotation

    Little Rock, AR
    1 day ago
  • $238k - $302k

     ...+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in Large...  ...looking for quantitatively-minded engineers to research and propose new ways...  ...in a heavily quantitative software engineering area ~ Experience navigating... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    22 hours ago
  • $150k - $250k

     ...About Distyl AI Distyl is an applied AI technology company...  ..., we build AI systems using Evaluation-Driven Development —an approach...  ...production.   AI Evaluation Engineers focus on designing and...  ...What We Require ~2+ years of software engineering experience ~ Strong... 
    Full time
    Work at office
    Flexible hours
    3 days per week

    Distyl Ai

    Remote
    22 hours ago
  • $40 per hour

     ...A cybersecurity solutions company is seeking experienced professionals to evaluate AI-generated security content and solve technical security problems. Candidates should have over 2 years of hands-on experience in cybersecurity and coding skills, with strong writing and... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Saint Paul, MN
    1 day ago
  • $40 per hour

     ...cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content, solve technical...  ...penetration testing, red teaming, incident response, detection engineering, DFIR, malware analysis, threat intelligence, or... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    New York, NY
    1 day ago
  • $40 per hour

    A cybersecurity startup is seeking experienced cybersecurity professionals to evaluate AI-generated content and solve technical problems. This role offers the flexibility to work remotely while contributing to innovative security AI models. Candidates should have 2+ years... 
    Hourly pay
    Remote work

    DataAnnotation

    Bismarck, ND
    1 day ago
  • A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve complex cybersecurity problems. This role is fully remote, allowing candidates to choose projects and work on their own schedule. Ideal candidates should... 
    Remote work

    DataAnnotation

    Lincoln, NE
    1 day ago
  • A leading AI training firm is seeking a Web Application Developer to join their team. In this role, you will train AI models, evaluate their performance, and solve coding challenges using languages like Python or JavaScript. Candidates should possess a detail-oriented... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Annapolis, MD
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer - AI Evaluation. Be the first to apply!