Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer - Evaluation Author - AI Trainer

$35 - $120 per hour
Full-time

Mercor

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .

Position: Code-Data Eval Author — Software Engineer

Type: Contract

Compensation: $35–$120/hour

Location: Remote

Commitment: 30+ hours/week

Role Responsibilities

  • Author non-trivial coding tasks with golden solutions and automated verifiers.
  • Design rubrics and grade agent trajectories and model outputs.
  • Improve task and rubric quality through structured review.
  • Evaluate the accuracy and depth of AI-generated content to strengthen reasoning and rigor in model outputs.
  • Work independently and asynchronously to meet deadlines while improving AI model performance .

Qualifications

Must-Have

  • 5+ years of software engineering at a real product organization (big tech or venture-backed startup).
  • Strong code quality, systems design, debugging, and testing discipline.
  • Clear written communication (you write instructions others follow).

Preferred

  • Familiarity with AI coding tools and evals.

Interview Process

  • Short Mercor Technical Screen .
  • Live Code Review Session .
  • Domain Expert Interview .
  • You're paid $200 for completing all three, regardless of outcome.

Application Process (Takes 20–30 mins to complete)

  • Upload resume
  • AI interview based on your resume
  • Submit form

Resources & Support

PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.

#hiringmercor
Vacancy posted 14 hours ago
Similar jobs that could be interesting for youBased on the Software Engineer - Evaluation Author - AI Trainer in Remote vacancy
  • $170k - $216k

     ...teams that consume map data. In this role, you will: Evaluate and launch machine learning models that automatically...  ...work on various parts of the system ~ A passion for good software engineering It's a Bonus if you have: M.S. or Ph.D in Computer... 
    Suggested
    Full time
    Remote work

    Waymo

    Remote
    14 hours ago
  • $85 per hour

     ...technical talent with leading AI research labs. Headquartered...  ...Position: Materials Writer - Engineering FRQ Type: Contract...  ...Role Responsibilities Author original, free-response engineering...  ...to enhance model training and evaluation. Qualifications Must-Have... 
    Suggested
    Contract work
    Summer work
    Remote work

    Mercor

    Detroit, MI
    14 days ago
  • $100 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Dorsey . Position: Special Projects Software Engineers Type: Contract...  ...projects to enhance software solutions. Evaluate and improve AI model training and performance... 
    Suggested
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    14 hours ago
  • $90 - $120 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  .... Position: Technical Expert - Software Engineering Type: Contract Compensation...  ...test AI models on coding tasks. Evaluate model outputs in areas like algorithms... 
    Suggested
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    14 hours ago
  • $60 - $100 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ..., and Jack Dorsey . Position: Software Engineering, Data Science, and Systems Design Experts...  ...Role Responsibilities Evaluate LLM-generated responses to coding and... 
    Suggested
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    14 hours ago
  • $170k - $216k

     ...billions in simulation across 15+ U.S. states. Waymo's Release Evaluation org ensures that each version of the Waymo Driver is safe...  ...objectives under resource constraints. Collaborate with other engineers, data scientists, statisticians and the leadership team to... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    14 hours ago
  • $120k - $250k

     ...2016 in Silicon Valley, Pony.ai has quickly become a global leader...  ..., and multi-dimensional evaluation. Design and implement high...  ...Build and optimize downstream engineering workflows for Large Language...  ...skills in C/C++, Python, and software design Strong foundation in... 
    Full time
    Temporary work

    Pony Ai

    Remote
    14 hours ago
  • $170k - $216k

     ...across 15+ U.S. states. The Planner Evaluation team works on one of the key challenges in...  ...and improving the quality of the software that drives the car. We are looking for experienced data-minded software engineers and data scientists to help us improve how... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    14 hours ago
  • $170k - $216k

     ...powers the Waymo Driver. Our software allows the Waymo Driver to perceive...  ...sensors, enabling software engineers like you to develop multi-...  ...-critical automation and evaluation frameworks that establish the...  ...of experience in industrial AI applications involving the creation... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    14 hours ago
  • $170k - $216k

     ...simulator components, we create metrics and evaluation methodologies to measure realism. These...  ...this hybrid role you will report to an Engineering Manager.   You will: Develop...  ...~2+ years of industry experience in software development. ~ Software Engineering Fundamentals... 
    Full time
    Remote work

    Waymo

    Remote
    14 hours ago
  • $170k - $216k

     ...powers the Waymo Driver. Our software allows the Waymo Driver to perceive...  ...sensors, enabling software engineers like you to develop multi-...  ...using an automated system Evaluate new hardware specifications...  ...of experience in industrial AI applications involving the creation... 
    Full time
    Remote work

    Waymo

    Remote
    14 hours ago
  • $175k - $215k

     ...dynamics, and state-of-the-art Generative AI to create a training ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge: How...  ...virtual world is "real"? We are looking for a Software Engineer to build the metrics and pipelines that... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    14 hours ago
  • $170k - $216k

     ...billions in simulation across 15+ U.S. states. The Perception Evaluation team at Waymo is at the forefront of autonomous driving,...  ...safe and effective autonomous operation. We are seeking a Software Engineer play a pivotal role in shaping the future of transportation... 
    Full time
    Remote work

    Waymo

    Remote
    14 hours ago
  • $50 - $150 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ..., and Jack Dorsey . Position: Software Engineering Expert Type: Contract...  ...Role Responsibilities Write and evaluate code across diverse programming languages... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    14 hours ago
  • $80 - $120 per hour

    Role Description ~Evaluate AI-generated artifacts against domain-specific quality rubrics. ~Identify factual, aesthetic, and presentation errors in documents, spreadsheets, and slide decks. ~Provide clear, structured written feedback to improve AI outputs. ~Apply... 
    Part time
    Work at office
    Remote work

    Mercor

    Remote
    6 days ago
  • $129.4k - $198.4k

     .... About the Organization:The Evaluation team builds and evolves the evaluation...  ...used for autonomous vehicle software validation. Develop and...  ...interpretable insights to engineering teams and leadership, including...  ...best practices. Leverage AI-assisted development tools and... 
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    4 days ago
  • $224k - $356.5k

     ...future of autonomous driving, and evaluation is how we know the drive is...  ...changing how our organization develops AI drivers!We are looking for a senior engineer to own the engine that makes it...  ...equivalent experience).12+ years building software, with significant time in... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $125k - $175k

     ...starting out, NinjaTrader equips traders with award-winning software and brokerage services to navigate the world's leading...  ...the world. What you'll do:We are seeking a Sr. Software Engineer to join our Evaluation Services team. Our Evaluation Services business powers funded... 
    Work at office
    Remote work
    Worldwide
    Monday to Friday
    Flexible hours
    Shift work

    NinjaTrader Group

    Chicago, IL
    2 days ago
  • $184k - $287.5k

     ...are looking for an experienced engineer to join our Planning and...  ...team to work on metrics and evaluation. In this role you will enable...  ...evaluation of our Autonomous Vehicle software.Build compelling, data driven...  ...vacancy. NVIDIA uses AI tools in its recruiting processes... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $144.7k - $221.4k

     .... About the Organization The Evaluation team builds and evolves the evaluation...  ...into clear feedback for engineering and leadership, and help...  ...introspect autonomous driving software performance at interfaces across...  .... Experience leveraging AI-assisted development and analytics... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    2 days ago
  • $80 - $100 per hour

     ...verification for this role. What You'll Be Doing Design and build the coding benchmarks and evaluation pipelines used to test frontier AI models on real software engineering work: Design coding benchmarks that evaluate frontier models on real-world programming... 
    Remote job
    Full time
    Contract work
    For contractors

    G2i

    Miami, FL
    14 hours ago
  •  ...Summary This is a fully remote, hourly contractor role supporting AI data and language projects on a project-based, flexible hour...  ..., and other content to support AI training datasets. LLM evaluation: reviewing AI-generated responses for accuracy, reasoning quality... 
    Hourly pay
    For contractors
    Remote work
    Flexible hours

    CNTXT AI

    Brooklyn, NY
    1 day ago
  • $20 - $80 per hour

     ...Role Overview Train and evaluate next-generation AI systems by scoring model outputs, annotating real-world content, and delivering clear, actionable...  ...across fields such as finance, healthcare, STEM engineering, and more, producing the annotations and feedback loops that... 
    Hourly pay
    For contractors
    Remote work

    SaidGig

    United States
    23 days ago
  • **Job Title: AI Trainer || Image Quality Evaluator || English** **Location**: Remote | Work from Home **Employment Type:** Project-based | Contract We are looking for detail-oriented Image Quality Evaluator for a multilingual AI data Annotation and Transcription Specialists... 
    Contract work
    Remote work
    Work from home
    Monday to Friday
    Day shift

    iMerit Technology

    Louisiana
    a month ago
  • $80 - $120 per hour

    Role Description ~Evaluate AI-generated artifacts against domain-specific quality rubrics. ~Identify factual, aesthetic, and presentation errors in documents, spreadsheets, and slide decks. ~Provide clear, structured written feedback to improve AI model outputs.... 
    Part time
    Work at office
    Remote work

    Mercor

    Remote
    a month ago
  • Mercor is seeking expert Evaluators in Finance operations / audit support to review AI-generated documents, spreadsheets, and slide decks for accuracy, rigor, and domain quality. This is a remote, hourly engagement, requiring strong subject-matter expertise to grade outputs... 
    Remote job
    Hourly pay

    Mercor

    Dallas, TX
    4 days ago
  • $30 per hour

    Prolific is seeking Advanced Dutch Speakers in Chicago, IL to train AI models. You will complete AI tasks and assess AI performance in...  ...competitive rates and direct payment through PayPal. Pass the evaluation, and you can start within 15 minutes. #J-18808-Ljbffr Prolific
    Remote job
    Work from home

    Prolific

    Chicago, IL
    4 days ago
  • YO IT Consulting is seeking an AI Trainer & Evaluator for a remote contract role. You will train AI systems by evaluating outputs, annotating data, and providing actionable feedback to improve model performance. You’ll work with client teams to apply rubrics across diverse... 
    Remote job
    Contract work

    YO IT Consulting

    Atlanta, GA
    1 day ago
  • $20 per hour

    A leading AI development company in the United States is seeking detail-oriented individuals for remote opportunities in training AI...  ...include developing prompts, writing high-quality responses, and evaluating AI outputs. The ideal candidates are fluent in English with... 
    Remote job
    Hourly pay
    Flexible hours

    SupportFinity™

    Raleigh, NC
    3 days ago
  • $80 - $120 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Position: Compliance / regulatory response with financial-services AI Evaluator Type: Contract Compensation: $80–$120/hour... 
    Contract work
    Summer work
    Work at office
    Remote work

    Mercor

    Chicago, IL
    14 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer - Evaluation Author - AI Trainer. Be the first to apply!