Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer — Model Evaluation & Experimentation

Full-time

Dorado

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models.

1. Overview

A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced machine learning practitioners to act as ground-truth experts for model evaluation and experimentation. You will author complex, multi-step ML tasks — for example, taking a vague research idea like "modify how an RL reward is computed," implementing the change, running the training experiment, and analyzing the results to determine success — and verify exactly where frontier models fall short.

Each task represents one to two days of continuous, focused effort and spans multiple technical skills: implementation, experiment setup and execution, and rigorous analysis. You will work in a tight feedback loop with the lab's researchers.

This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week.

2. Key Responsibilities

  • Design tasks: Turn real ML research ideas — like tweaking how an RL reward is computed — into well-defined, multi-step tasks.

  • Run experiments: Implement changes, run training experiments, and analyze the results to show what a correct solution looks like.

  • Explore RL ideas: Build some of your tasks around reinforcement-learning basics such as reward functions and training behavior.

  • Evaluate models: See how frontier models handle your tasks, and note where and why they fall short.

  • Work as a team: Compare notes with researchers and fellow experts so tasks stay consistent, rigorous, and fair.

3. Core Qualifications

  • MSc or PhD in machine learning, computer science, or another STEM field, or equivalent practical experience in a research-heavy domain.

  • 1+ years of experience in a research or research-engineering role.

  • Hands-on experience training and evaluating ML models and running experiments end-to-end — experiment setup, execution, and analysis — not just using ML libraries superficially.

  • Strong familiarity with large language models: their capabilities, limitations, and evaluation techniques.

  • Working proficiency in Python and Git, with comfort in both scripting and notebook environments.

  • Basic understanding of reinforcement learning (reward functions, policy training) is preferred.

  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.

  • A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems.

  • Ability to engage reliably for approximately 35 hours per week.

About Cincinnatus LLC

Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives.

Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client\'s internal teams, and integration into standard enterprise workflows.

Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC.

Equal Employment Opportunity

Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

#J-18808-Ljbffr
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer — Model Evaluation & Experimentation in United States vacancy
  • $60 - $90 per hour

     ...Machine Learning Engineer — Model Evaluation & Experimentation is a remote review track for evaluating AI outputs across machine learning engineer model evaluation experimentation specialist operations workflows. Reviewers grade workflow correctness, policy adherence,... 
    Suggested
    Remote job
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    14 days ago
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure... 
    Suggested
    Full time

    Scale Ai

    Washington DC
    10 hours ago
  •  ...Author and run rigorous, multi-step machine learning evaluation tasks for a leading generative AI research...  ...results to determine where frontier models succeed or fail. Typical tasks...  ...experience in a research or research-engineering position. Hands-on experience training... 
    Suggested
    Hourly pay
    Full time
    Freelance
    Remote work

    SaidGig

    United States
    12 days ago
  • $85 per hour

     ...Role Overview Evaluate and improve frontier AI coding agents by completing realistic machine learning engineering tasks and assessing model outputs. You will perform structured technical assessments that reflect production ML workflows, helping a leading AI research lab... 
    Suggested
    Hourly pay
    Remote work

    SaidGig

    United States
    16 days ago
  • $228.7k - $343.1k

     ...enormous scale, and one bad model can mean millions in...  ..., so you critically evaluate what it produces and...  ...you did not write, learning the data, configs,...  ...and generalization. Experimentation and statistical rigor...  ...software and data engineering: production-quality Python... 
    Suggested
    Remote job
    Full time
    Local area
    Shift work

    Block

    New York, NY
    10 hours ago
  •  ...opportunity Unity’s Ads Experimentation Platform team is looking for a senior machine learning engineer to lead the evolution of how...  ...iterate on machine learning models or ads delivery pipelines is...  ...for our experimentation and evaluation roadmap. You will bridge the... 
    Full time
    Temporary work
    Work at office
    Worldwide
    Relocation package

    Unity Technologies

    Remote
    10 hours ago
  • $150k

     ...Machine Learning Engineer About the Institute of Foundation Models: We are a dedicated research lab for building, understanding, using, and risk-managing...  ...researchers to support data pipelines, experimentation, and evaluation workflows. ~ This role balances fast... 
    Visa sponsorship

    Institute of Foundation Models

    Sunnyvale, CA
    4 days ago
  •  ...WA, we are a team of engineers and technologists from...  ...Software Engineer – Evaluation, you will design and...  ..., and small language model (SLM) systems. You will...  ...work closely with machine learning and data engineering...  ...learning evaluation, experimentation, or benchmarking ~... 

    VTI Aerospace

    Washington DC
    8 days ago
  • $174.72k - $295.68k

     ...through cutting-edge R&D in AI, machine learning, and smart connectivity.We are...  ...for a full-time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development of...  ...).Conduct systematic ablation, evaluation, and visualization of model... 
    Full time

    XPENG Motors

    Santa Clara, CA
    10 hours ago
  •  ...systems enable robots to adapt, learn, and perform in the real...  ...fast, complex, and poorly modeled physics that traditional...  ...ones.We are seeking a Senior Machine Learning Engineer to lead the development of...  ...planning, process-optimization, evaluation, and synthetic-data... 
    Shift work

    Path Robotics

    Columbus, OH
    1 day ago
  •  ...Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific...  ...sciences. Responsibilities include fact-checking, critiquing experimental designs, and ensuring ethical compliance. Participants... 
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    2 days ago
  • $60 per hour

     ...advance AI development. AI models are increasingly capable of...  ...art AI models on tasks like evaluating AI-generated quantitative analysis...  ...areas like forecasting, experimental analysis, optimization, and...  ...Computer Science, Mathematics, Engineering, or similar); a master's or... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Kansas City, MO
    2 days ago
  • $213k - $263k

     ...for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge...  ...is "real"? We are seeking visionary machine learning engineers and researchers to architect the...  ...the realism of our multimodal world models. Your work will define the state of the... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    10 hours ago
  • $174.72k - $295.68k

     ...through cutting-edge R&D in AI, machine learning, and smart connectivity. We are seeking Machine Learning Engineers with strong expertise in generative modeling and large-scale deep learning systems...  ...will research, implement, and evaluate world models that learn the... 
    Full time

    XPENG

    Santa Clara, CA
    1 day ago
  • $70 per hour

     ...Role Overview Join a remote Machine Learning Engineer talent network to be considered for future contract engagements with AI labs and companies...  .... Members contribute to advancing AI by training and evaluating models, designing real-world tasks and deliverables, and... 
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    14 days ago
  • $107.25k - $164.45k

     ..., leveraging modern approaches in machine learning and artificial intelligence. We have...  ...We are seeking Machine Learning Engineers (Applied Research & Model Development) to join our team and...  ...scalable ML solutions. - Contribute to experimental design and analysis, including... 
    Hourly pay
    Full time

    PathAI

    Remote
    1 day ago
  • $224k - $356.5k

    We are looking for outstanding Machine Learning Engineers to join our Physical AI teams. As the pioneers...  ...state-of-the-art multimodal models and diffusion techniques to simulate...  ...Establish a strong mentality for KPI evaluation and validation to ensure the quality... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $204k - $259k

     ...The Driver Understanding and Evaluation (DUE) team at Waymo is...  ...the Waymo Driver.  The DUE Machine Learning team will build and operate...  ...and advanced machine learning models to deliver training and evaluation...  ...researchers and software engineers who are passionate about developing... 
    Full time

    Waymo

    Remote
    10 hours ago
  •  ...Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields...  ...fact-checking AI technical claims and assessing experimental logic. The role includes flexible work hours, competitive... 
    Remote job
    Hourly pay
    Flexible hours

    Prolific

    Jacksonville, FL
    5 days ago
  • $60 per hour

     ...Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to...  ...pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing experimental designs. Successful candidates should have a strong... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    Dallas, TX
    10 hours ago
  • $170k - $216k

     ...15+ U.S. states. The DUE Machine Learning team will build and operate...  ...tools, improve and speed up the evaluation and onboard developer...  ...and advanced machine learning models to deliver training and evaluation...  ...researchers and software engineers who are passionate about... 
    Full time

    Waymo

    Remote
    10 hours ago
  • $238k - $302k

     ...Foundations team is to develop machine learning solutions addressing open...  ...from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a...  ...to a Senior Staff Software Engineer.   You will: Work with... 
    Full time
    Remote work

    Waymo

    Remote
    10 hours ago
  • $60 per hour

     ...is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards...  ...technical claims from public databases, and critiquing experimental designs. Successful candidates will earn competitive pay rates... 
    Remote job
    Hourly pay
    Work from home
    Flexible hours

    Prolific

    San Jose, CA
    1 day ago
  •  ...Lead the design and evaluation of next-generation coding...  ...of coding models. Key Responsibilities...  ...across diverse software engineering tasks. Develop...  ...support large-scale experimentation, data generation, review...  ...engineering, machine learning, AI research, evaluation... 
    Full time
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $30 - $90 per hour

     ...refine high-performance backend APIs in Rust to support experimental developer tools and cutting-edge AI research. You...  ...GraphQL services, collaborate closely with research engineers, and evaluate AI-powered coding models to improve developer workflows. You will join micro1... 
    Hourly pay
    Contract work
    For contractors
    Remote work
    Free visa

    SaidGig

    United States
    a month ago
  •  ...providing Information Technology, Engineering Services, Program Management, and...  ...Solerity is seeking Mid to Senior Machine Learning Engineers and AI Model Developers to support an upcoming...  ...This effort focuses on developing, evaluating, and integrating machine learning... 
    Full time
    For contractors
    Remote work
    Flexible hours

    Solerity

    Remote
    10 hours ago
  • $117.7k - $221.4k

     ...the intersection of machine learning, data infrastructure...  ...across perception and evaluation workflows.Our goal...  ...not only on stronger models, but also on better...  ...model reflects how Cola engineers think: build durable...  ...support both rapid experimentation and production... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    10 hours ago
  • $100 per hour

     ...knowledge to help train and evaluate next-generation AI systems by...  ..., clarity, and relevance of model outputs through rubric-based...  ...experience in Data Science, Machine Learning, Applied AI, Statistics,...  ...desirable. Background in prompt engineering, AI output evaluation, fact... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    Indiana
    9 days ago
  • $170k - $216k

     ...team builds the system which learns the spatial-temporal representation...  ...set of sensors, enabling engineers like you to (1) develop...  ...real-world data, to (2) develop models and model training at scale,...  ...experience ~3+ years experience in Machine Learning and/or Computer... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    10 hours ago
  •  ...build cutting-edge foundation AI models and end-to-end products that...  ...is a team of researchers, engineers, designers, and more, who are...  ...Paris. Join us!Why this role?Evaluation is critical to making progress...  ...credit.Education & learning stipend for conferences, courses... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    10 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer — Model Evaluation & Experimentation. Be the first to apply!