Machine Learning Engineer — Model Evaluation & Experimentation
Dorado
Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models.
1. Overview
A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced machine learning practitioners to act as ground-truth experts for model evaluation and experimentation. You will author complex, multi-step ML tasks — for example, taking a vague research idea like "modify how an RL reward is computed," implementing the change, running the training experiment, and analyzing the results to determine success — and verify exactly where frontier models fall short.
Each task represents one to two days of continuous, focused effort and spans multiple technical skills: implementation, experiment setup and execution, and rigorous analysis. You will work in a tight feedback loop with the lab's researchers.
This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week.
2. Key Responsibilities
Design tasks: Turn real ML research ideas — like tweaking how an RL reward is computed — into well-defined, multi-step tasks.
Run experiments: Implement changes, run training experiments, and analyze the results to show what a correct solution looks like.
Explore RL ideas: Build some of your tasks around reinforcement-learning basics such as reward functions and training behavior.
Evaluate models: See how frontier models handle your tasks, and note where and why they fall short.
Work as a team: Compare notes with researchers and fellow experts so tasks stay consistent, rigorous, and fair.
3. Core Qualifications
MSc or PhD in machine learning, computer science, or another STEM field, or equivalent practical experience in a research-heavy domain.
1+ years of experience in a research or research-engineering role.
Hands-on experience training and evaluating ML models and running experiments end-to-end — experiment setup, execution, and analysis — not just using ML libraries superficially.
Strong familiarity with large language models: their capabilities, limitations, and evaluation techniques.
Working proficiency in Python and Git, with comfort in both scripting and notebook environments.
Basic understanding of reinforcement learning (reward functions, policy training) is preferred.
Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.
A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems.
Ability to engage reliably for approximately 35 hours per week.
About Cincinnatus LLC
Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives.
Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client\'s internal teams, and integration into standard enterprise workflows.
Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC.
Equal Employment Opportunity
Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.
Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.
#J-18808-Ljbffr$60 - $90 per hour
...Machine Learning Engineer — Model Evaluation & Experimentation is a remote review track for evaluating AI outputs across machine learning engineer model evaluation experimentation specialist operations workflows. Reviewers grade workflow correctness, policy adherence,...SuggestedRemote jobFor contractorsWork experience placement10 hours per week$208k - $300k
...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure...SuggestedFull time- ...Author and run rigorous, multi-step machine learning evaluation tasks for a leading generative AI research... ...results to determine where frontier models succeed or fail. Typical tasks... ...experience in a research or research-engineering position. Hands-on experience training...SuggestedHourly payFull timeFreelanceRemote work
$85 per hour
...Role Overview Evaluate and improve frontier AI coding agents by completing realistic machine learning engineering tasks and assessing model outputs. You will perform structured technical assessments that reflect production ML workflows, helping a leading AI research lab...SuggestedHourly payRemote work$228.7k - $343.1k
...enormous scale, and one bad model can mean millions in... ..., so you critically evaluate what it produces and... ...you did not write, learning the data, configs,... ...and generalization. Experimentation and statistical rigor... ...software and data engineering: production-quality Python...SuggestedRemote jobFull timeLocal areaShift work- ...opportunity Unity’s Ads Experimentation Platform team is looking for a senior machine learning engineer to lead the evolution of how... ...iterate on machine learning models or ads delivery pipelines is... ...for our experimentation and evaluation roadmap. You will bridge the...Full timeTemporary workWork at officeWorldwideRelocation package
$150k
...Machine Learning Engineer About the Institute of Foundation Models: We are a dedicated research lab for building, understanding, using, and risk-managing... ...researchers to support data pipelines, experimentation, and evaluation workflows. ~ This role balances fast...Visa sponsorship- ...WA, we are a team of engineers and technologists from... ...Software Engineer – Evaluation, you will design and... ..., and small language model (SLM) systems. You will... ...work closely with machine learning and data engineering... ...learning evaluation, experimentation, or benchmarking ~...
$174.72k - $295.68k
...through cutting-edge R&D in AI, machine learning, and smart connectivity.We are... ...for a full-time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development of... ...).Conduct systematic ablation, evaluation, and visualization of model...Full time- ...systems enable robots to adapt, learn, and perform in the real... ...fast, complex, and poorly modeled physics that traditional... ...ones.We are seeking a Senior Machine Learning Engineer to lead the development of... ...planning, process-optimization, evaluation, and synthetic-data...Shift work
- ...Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific... ...sciences. Responsibilities include fact-checking, critiquing experimental designs, and ensuring ethical compliance. Participants...Remote workFlexible hours
$60 per hour
...advance AI development. AI models are increasingly capable of... ...art AI models on tasks like evaluating AI-generated quantitative analysis... ...areas like forecasting, experimental analysis, optimization, and... ...Computer Science, Mathematics, Engineering, or similar); a master's or...Hourly payFull timeRemote workFlexible hours$213k - $263k
...for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge... ...is "real"? We are seeking visionary machine learning engineers and researchers to architect the... ...the realism of our multimodal world models. Your work will define the state of the...Full timeRemote work$174.72k - $295.68k
...through cutting-edge R&D in AI, machine learning, and smart connectivity. We are seeking Machine Learning Engineers with strong expertise in generative modeling and large-scale deep learning systems... ...will research, implement, and evaluate world models that learn the...Full time$70 per hour
...Role Overview Join a remote Machine Learning Engineer talent network to be considered for future contract engagements with AI labs and companies... .... Members contribute to advancing AI by training and evaluating models, designing real-world tasks and deliverables, and...Hourly payContract workFor contractorsRemote work$107.25k - $164.45k
..., leveraging modern approaches in machine learning and artificial intelligence. We have... ...We are seeking Machine Learning Engineers (Applied Research & Model Development) to join our team and... ...scalable ML solutions. - Contribute to experimental design and analysis, including...Hourly payFull time$224k - $356.5k
We are looking for outstanding Machine Learning Engineers to join our Physical AI teams. As the pioneers... ...state-of-the-art multimodal models and diffusion techniques to simulate... ...Establish a strong mentality for KPI evaluation and validation to ensure the quality...Full time$204k - $259k
...The Driver Understanding and Evaluation (DUE) team at Waymo is... ...the Waymo Driver. The DUE Machine Learning team will build and operate... ...and advanced machine learning models to deliver training and evaluation... ...researchers and software engineers who are passionate about developing...Full time- ...Jacksonville, Florida, to join our Expert Network for evaluating AI-generated science models. Candidates should hold a BS, MS, or PhD in relevant fields... ...fact-checking AI technical claims and assessing experimental logic. The role includes flexible work hours, competitive...Remote jobHourly payFlexible hours
$60 per hour
...Professionals to join their Expert Network to evaluate AI-generated science. This role allows you to... ...pay rate of up to $60 per hour for reviewing model responses, validating technical claims, and critiquing experimental designs. Successful candidates should have a strong...Remote jobHourly payWork from homeFlexible hours$170k - $216k
...15+ U.S. states. The DUE Machine Learning team will build and operate... ...tools, improve and speed up the evaluation and onboard developer... ...and advanced machine learning models to deliver training and evaluation... ...researchers and software engineers who are passionate about...Full time$238k - $302k
...Foundations team is to develop machine learning solutions addressing open... ...from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a... ...to a Senior Staff Software Engineer. You will: Work with...Full timeRemote work$60 per hour
...is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards... ...technical claims from public databases, and critiquing experimental designs. Successful candidates will earn competitive pay rates...Remote jobHourly payWork from homeFlexible hours- ...Lead the design and evaluation of next-generation coding... ...of coding models. Key Responsibilities... ...across diverse software engineering tasks. Develop... ...support large-scale experimentation, data generation, review... ...engineering, machine learning, AI research, evaluation...Full timeRemote work
$30 - $90 per hour
...refine high-performance backend APIs in Rust to support experimental developer tools and cutting-edge AI research. You... ...GraphQL services, collaborate closely with research engineers, and evaluate AI-powered coding models to improve developer workflows. You will join micro1...Hourly payContract workFor contractorsRemote workFree visa- ...providing Information Technology, Engineering Services, Program Management, and... ...Solerity is seeking Mid to Senior Machine Learning Engineers and AI Model Developers to support an upcoming... ...This effort focuses on developing, evaluating, and integrating machine learning...Full timeFor contractorsRemote workFlexible hours
$117.7k - $221.4k
...the intersection of machine learning, data infrastructure... ...across perception and evaluation workflows.Our goal... ...not only on stronger models, but also on better... ...model reflects how Cola engineers think: build durable... ...support both rapid experimentation and production...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours$100 per hour
...knowledge to help train and evaluate next-generation AI systems by... ..., clarity, and relevance of model outputs through rubric-based... ...experience in Data Science, Machine Learning, Applied AI, Statistics,... ...desirable. Background in prompt engineering, AI output evaluation, fact...Hourly payContract workPart timeFor contractorsRemote work$170k - $216k
...team builds the system which learns the spatial-temporal representation... ...set of sensors, enabling engineers like you to (1) develop... ...real-world data, to (2) develop models and model training at scale,... ...experience ~3+ years experience in Machine Learning and/or Computer...Full timeRemote work- ...build cutting-edge foundation AI models and end-to-end products that... ...is a team of researchers, engineers, designers, and more, who are... ...Paris. Join us!Why this role?Evaluation is critical to making progress... ...credit.Education & learning stipend for conferences, courses...Full timeWork at officeLocal areaRemote workHome office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Engineer — Model Evaluation & Experimentation. Be the first to apply!
- junior machine learning engineer United States
- graduate machine learning engineer United States
- senior ml engineer United States
- data scientist machine learning engineer United States
- machine learning engineer United States
- lead machine learning engineer United States
- machine learning ai engineer United States
- entry level machine learning engineer United States
- machine learning software engineer United States
- ai ml engineer United States



