Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Engineer - Model Evaluation Expert

$60 - $90 per hour
Full-time

Mercor

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .

Position: Machine Learning Engineer — Model Evaluation & Experimentation

Type: Contract

Compensation: $60–$90/hour

Location: Remote

Commitment: 35 hours/week

Role Responsibilities

  • Design tasks by transforming real ML research ideas into well-defined, multi-step tasks.
  • Run experiments by implementing changes, executing training experiments, and analyzing results to define correct solutions.
  • Explore reinforcement learning concepts such as reward functions and training behavior in task development.
  • Evaluate frontier models' performance on tasks and identify areas of improvement.
  • Collaborate with researchers to ensure tasks are consistent, rigorous, and fair.

Qualifications

Must-Have

  • MSc or PhD in machine learning , computer science , or a related STEM field.
  • 1+ years in a research or research-engineering role.
  • Experience in training and evaluating ML models and conducting end-to-end experiments.
  • Proficiency in Python and Git .
  • Strong understanding of large language models and their evaluation.

Preferred

  • Basic knowledge of reinforcement learning .
  • Experience in AI training, model evaluation, or benchmark/task authoring.

Compensation & Legal

  • Hourly contractor
  • Paid weekly

Application Process (Takes 20–30 mins to complete)

  • Upload resume
  • AI interview based on your resume
  • Submit form

Resources & Support

PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.

Vacancy posted 12 hours ago
Similar jobs that could be interesting for youBased on the ML Engineer - Model Evaluation Expert in Remote vacancy
  • $189.4k - $300.6k

     ...behavior across real-world scenarios.The Evaluation Foundations team—part of Embodied AI’s...  ...in autonomous vehicle development. We engineer high-performance tools that identify top-performing models and partner with data-intensive ML teams to drive rapid innovation.In this... 
    Suggested
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    3 days ago
  • $281k - $356k

     ...improve and speed up the evaluation and onboard developer...  ...journeys. It will combine expert human judgements and...  ...advanced machine learning models to deliver training and...  ...and software engineers who are passionate about...  ...in Python and standard ML frameworks (e.g., JAX,... 
    Suggested
    Full time

    Waymo

    Remote
    21 hours ago
  •  ...seeking a highly skilled Machine Learning Engineer who specializes in leveraging Large Language Models (LLMs) for automated evaluation and quality assessment. In this role, you will...  ...~5+ years of experience in ML engineering, NLP, or AI/ML automation. ~... 
    Suggested
    Work at office
    Remote work
    Flexible hours

    Grid Dynamics Holdings

    United States
    4 days ago
  • $189.4k - $300.6k

     ...behavior across real-world scenarios. The Evaluation Foundations team—part of Embodied AI’s...  ...in autonomous vehicle development. We engineer high-performance tools that identify top-performing models and partner with data-intensive ML teams to drive rapid innovation. In... 
    Suggested
    Full time
    Relocation package
    Flexible hours

    General Motors

    Remote
    1 day ago
  • $213k - $263k

     ...ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge:...  ...are seeking visionary machine learning engineers and researchers to architect the...  ...measure the realism of our multimodal world models. Your work will define the state of the... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    21 hours ago
  • $204k - $259k

     ...Driver Understanding and Evaluation (DUE) team at Waymo is...  .... It will combine expert human judgements and advanced...  ...machine learning models to deliver training and...  ...researchers and software engineers who are passionate about...  ...implement scalable and robust ML pipelines for training,... 
    Full time

    Waymo

    Remote
    21 hours ago
  • $170k - $216k

     ...improve and speed up the evaluation and onboard developer...  ...journeys. It will combine expert human judgements and...  ...advanced machine learning models to deliver training and...  ...and software engineers who are passionate about...  ...distributed systems covering the ML lifecycle, supporting... 
    Full time

    Waymo

    Remote
    21 hours ago
  • $238k - $302k

     ...from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid...  ...to a Senior Staff Software Engineer.   You will: Work with...  ...~5+ years of experience in ML engineering and applied Deep Learning... 
    Full time
    Remote work

    Waymo

    Remote
    21 hours ago
  •  ...Innodata is expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will...  ...optimize the technical foundations that power model improvement for foundation model builders... 
    Full time

    Innodata

    Remote
    21 days ago
  • $152k - $241.5k

     ...working for us! We believe open-weight models are foundational to American AI...  ...scientific scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find,...  .... We are looking for an Evaluation/ML-Systems Engineer to own how we... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...step machine learning evaluation tasks for a leading generative...  ...where frontier models succeed or fail....  ...tasks that translate real ML research ideas into well...  ...and fellow experts. Qualifications...  ...research or research-engineering position. Hands-on... 
    Hourly pay
    Full time
    Freelance
    Remote work

    SaidGig

    United States
    12 days ago
  • $85 per hour

     ...Role Overview Evaluate and improve frontier AI coding agents by completing realistic machine learning engineering tasks and assessing model outputs. You will perform structured technical assessments that reflect production ML workflows, helping a leading AI research lab... 
    Hourly pay
    Remote work

    SaidGig

    United States
    16 days ago
  •  ...the most advanced AI models. 1. Overview A leading...  ...generation of agentic evaluation benchmarks for...  ...to act as ground-truth experts for model evaluation and...  ...author complex, multi-step ML tasks — for example,...  ...research or research-engineering role. ~ Hands-on... 
    Full time
    Contract work
    Part time
    Freelance
    Remote work

    Dorado

    United States
    1 day ago
  • $177.3k - $212.8k

     ...teams use them to train various models, mapping teams use them to...  ...latest developments in AI and ML for autonomous driving, 3D reconstruction...  ...a customer-centric manner. - Evaluate and make recommendations...  ...consensus. Mentors and guides engineers within the group. - Bachelor’... 
    Full time
    Work at office
    Immediate start
    Relocation

    Torc Robotics

    Remote
    12 hours ago
  •  ...efficiency. About the Role We’re hiring an ML Engineer to join Kodex’s Verifications / Threat...  ...workflows into production-grade models, pipelines, and decision-support tooling...  ...-to-end (feature generation, training, evaluation, batch/streaming inference, backfills, and... 
    Remote job
    Full time
    Flexible hours

    Kodex

    United States
    21 hours ago
  • $199.2k - $298.8k

     ...teams use it to train various models and simulation teams use it for...  ...latest developments in AI and ML for autonomous driving....  ...a customer-centric manner.  Evaluate and make recommendations regarding...  ...consensus.  Mentors and guides engineers within the group.   Bachelor... 
    Full time
    Immediate start
    Relocation

    Company

    Remote
    21 hours ago
  • $213k - $263k

     ...are currently focusing on include reinforcement learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid work schedule and reports to a Principal Research Scientist.... 
    Full time
    Temporary work
    Remote work

    Waymo

    New York, NY
    21 hours ago
  • $195k - $300k

     ...Built by a world-class team: Engineers, designers, and operators from...  ...— it's the foundation. As an ML Engineer, you'll be working at...  ...intersection of cutting-edge model development and real-world legal...  ...data curation and fine-tuning to evaluation and production deployment —... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours

    Eve

    Remote
    21 hours ago
  •  ...Machine Learning Engineer (Llama AI Platform) Location: Remote (Preferred...  ...by open-source large language models, agentic workflows, and...  ...workflows. Fine-tune and evaluate LLM performance for business use...  ...tuning open-source LLMs. ML Engineering and MLOps... 
    Full time
    Remote work

    Performacentric

    Indianapolis, IN
    21 hours ago
  • $153.2k - $183.3k

     ...Team   As a Machine Learning Engineer II – Road & Lane, you will help develop next‑generation models that estimate road surfaces,...  ...Implement scalable training and evaluation pipelines for lane perception...  ...Hands‑on experience developing ML models for perception tasks... 
    Remote job
    Full time

    Torc Robotics

    Remote
    21 hours ago
  •  ...science. We build foundational understanding of models to advance the frontier of intelligence. About the role: As a ML Engineer, you’ll build and operate the...  ...work on the systems that support training and evaluating large models, scaling experimental pipelines... 
    Full time
    Internship

    Tilde Research

    Remote
    21 hours ago
  • $153.2k - $183.3k

     ...the Team:   As a Machine Learning Engineer II – Learned Behaviors, you will help develop and deploy behavior models that power decision-making for...  ...~ Implement production-quality ML code to support model training, evaluation, and inference within the autonomy... 
    Full time

    Torc Robotics

    Remote
    21 hours ago
  •  ...About the Role We’re looking for an Applied ML Engineer to design, evaluate, and scale recommendation and ranking systems that power how content...  ...techniques, grounding them in real user behavior, and shipping models that perform reliably at scale. The work is hands-on,... 
    Full time

    Darwin

    Palo Alto, CA
    21 hours ago
  • $200k - $300k

     ...preparation configurations. As a Senior ML Engineer, Manipulation, you will own the learning...  ...parallel jaw, multi-finger) Implement and evaluate modern policy architectures (diffusion policies, transformer-based action models, action chunking) and adapt them to Chef's... 
    Full time
    Flexible hours

    Chef Robotics

    Remote
    21 hours ago
  • $213k - $263k

     ...are currently focusing on include reinforcement learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid work schedule and you will report to a Research Director of AI... 
    Full time
    Remote work

    Waymo

    Remote
    21 hours ago
  •  ...Responsibilities: Own evaluation pipelines — design, build, and...  ...keep our speech and multimodal models honest in production. Harness...  ...Required Qualifications: Expert-level PyTorch. Proven software engineer who loves ML; comfortable writing production... 
    Full time
    Contract work
    Flexible hours
    Shift work

    Sesame, L.l.c.

    San Francisco, CA
    21 hours ago
  • $200k - $300k

     ...ingredients — it will come from foundation models that generalize across thousands of food...  ...Food Foundation Model. As a Senior ML Engineer, Foundation Models, you will work at the...  ...for the Food Foundation Model — evaluating tradeoffs across generalization, sample... 
    Full time
    Flexible hours

    Chef Robotics

    Remote
    21 hours ago
  • $174k - $253k

     ...implement robust agents and LLM-powered journeys that evaluate, gate, and improve engineering artifacts.Engineer and refine skills and context provided...  ...experience.5 years of experience with Large Language Modeling (LLM).3 years of experience with generative AI Agents.Preferred... 

    Google

    San Jose, CA
    3 days ago
  • $100.4k - $180.7k

    Posting TitleML and Optimization Engineer.LocationCO - Golden.Position...  ...scientists, engineers, and experts are accelerating energy...  ...high‑performance computing, AI/ML, modeling and simulation, and...  ...end development and red‑team evaluation. NLR collaborates across DOE... 
    Full time
    Fixed term contract
    Live in
    Local area
    Remote work
    Relocation
    Shift work

    National Renewable Energy Laboratory

    Golden, CO
    2 days ago
  • $175k - $215k

     ...learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. In this hybrid role, you will report to a...  ...industrial or research setting developing recipes for ML models We prefer: - Track record of... 
    Full time
    Remote work

    Waymo

    New York, NY
    12 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Engineer - Model Evaluation Expert. Be the first to apply!