Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Engineer - Model Evaluation Expert

$60 - $90 per hour
Full-time

Mercor

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .

Position: Machine Learning Engineer — Model Evaluation & Experimentation

Type: Contract

Compensation: $60–$90/hour

Location: Remote

Commitment: 35 hours/week

Role Responsibilities

  • Design tasks by transforming real ML research ideas into well-defined, multi-step tasks.
  • Run experiments by implementing changes, executing training experiments, and analyzing results to define correct solutions.
  • Explore reinforcement learning concepts such as reward functions and training behavior in task development.
  • Evaluate frontier models' performance on tasks and identify areas of improvement.
  • Collaborate with researchers to ensure tasks are consistent, rigorous, and fair.

Qualifications

Must-Have

  • MSc or PhD in machine learning , computer science , or a related STEM field.
  • 1+ years in a research or research-engineering role.
  • Experience in training and evaluating ML models and conducting end-to-end experiments.
  • Proficiency in Python and Git .
  • Strong understanding of large language models and their evaluation.

Preferred

  • Basic knowledge of reinforcement learning .
  • Experience in AI training, model evaluation, or benchmark/task authoring.

Compensation & Legal

  • Hourly contractor
  • Paid weekly

Application Process (Takes 20–30 mins to complete)

  • Upload resume
  • AI interview based on your resume
  • Submit form

Resources & Support

PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the ML Engineer - Model Evaluation Expert in Remote vacancy
  • $60 - $90 per hour

     ...Summers , and Jack Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type: Contract...  ...Responsibilities Design tasks by transforming real ML research ideas into well-defined, multi-step tasks. Run... 
    Suggested
    Hourly pay
    Weekly pay
    Full time
    Contract work
    For contractors
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    a month ago
  • $60 - $90 per hour

     ...Dorsey . Position: Machine Learning Engineer — Model Evaluation & Experimentation Type:...  ...Design tasks by turning real ML research ideas into well-defined, multi...  ...Collaborate with researchers and experts to maintain task consistency and rigor... 
    Suggested
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    Remote
    a month ago
  • $144.7k - $261.3k

    Job DescriptionAbout the TeamThe Evaluation Foundations team, part of Embodied AI's Scaling...  ...for autonomous vehicle development. We engineer high-performance tools that identify top-performing models and partner with data-intensive ML teams to drive rapid innovation.Why... 
    Suggested
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Warren, MI
    4 days ago
  • General Motors is seeking a Senior Engineer for the Embodied AI Scaling Foundations team to measure and visualize AV model performance at scale. You will design and implement evaluation and introspection tools used by GM AV teams, collaborating across Data, Infra, and... 
    Suggested
    Remote job

    General Motors

    Seattle, WA
    4 days ago
  • $213k - $263k

     ...ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge:...  ...are seeking visionary machine learning engineers and researchers to architect the...  ...measure the realism of our multimodal world models. Your work will define the state of the... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    more than 2 months ago
  •  ...looking for a Senior Machine Learning Engineer to join our Search & Intelligence organization...  ...-native experiences, agentic systems, models, evaluation frameworks, and data platforms that...  .... You will own projects across the ML lifecycle, from problem definition and... 
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    5 days ago
  • $238k - $302k

     ...from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid...  ...to a Senior Staff Software Engineer.   You will: Work with...  ...~5+ years of experience in ML engineering and applied Deep Learning... 
    Full time
    Remote work

    Waymo

    Remote
    more than 2 months ago
  • $204k - $259k

     ...Driver Understanding and Evaluation (DUE) team at Waymo is...  .... It will combine expert human judgements and advanced...  ...machine learning models to deliver training and...  ...researchers and software engineers who are passionate about...  ...implement scalable and robust ML pipelines for training,... 
    Full time

    Waymo

    Remote
    more than 2 months ago
  • $141.02k - $204.53k

     ...Responsibilities As the Senior AI/ML Engineer - Validation & Evaluation within AI Validation & Monitoring (...  ...evidence gaps or changes to the model, data, workflow, population, interface...  ...product development. Ability to conduct expert reviews using established usability... 
    Full time
    Work at office
    Remote work
    Flexible hours
    Weekend work

    Mayo Clinic

    Rochester, MN
    17 days ago
  • $200k - $300k

     ...ingredients — it will come from foundation models that generalize across thousands of food...  ...Food Foundation Model. As a Senior ML Engineer, Foundation Models, you will work at the...  ...for the Food Foundation Model — evaluating tradeoffs across generalization, sample... 
    Full time
    Flexible hours

    Chef Robotics

    Remote
    more than 2 months ago
  •  ...Responsibilities: Own evaluation pipelines — design, build, and...  ...keep our speech and multimodal models honest in production. Harness...  ...Required Qualifications: Expert-level PyTorch. Proven software engineer who loves ML; comfortable writing production... 
    Full time
    Contract work
    Flexible hours
    Shift work

    Sesame, L.l.c.

    San Francisco, CA
    more than 2 months ago
  •  ...Machine Learning Engineer (Llama AI Platform) Location: Remote (Preferred...  ...by open-source large language models, agentic workflows, and...  ...workflows. Fine-tune and evaluate LLM performance for business use...  ...tuning open-source LLMs. ML Engineering and MLOps... 
    Full time
    Remote work

    Performacentric

    Dallas, TX
    more than 2 months ago
  •  ...specialize in physics-informed ML and enterprise AI solutions...  ...Role: We are seeking an ML Engineer who builds ML systems directly...  ...multiple systems — and ends with a model running on a schedule inside...  ...and cleaning, modeling, evaluation, deployment into the client environment... 
    Full time
    Remote work
    Work visa
    Flexible hours

    Azx Inc

    Seattle, WA
    22 days ago
  • $177.3k - $212.8k

     ...teams use them to train various models, mapping teams use them to...  ...latest developments in AI and ML for autonomous driving, 3D reconstruction...  ...a customer-centric manner. Evaluate and make recommendations...  ...consensus. Mentors and guides engineers within the group. Bachelor’... 
    Full time
    Work at office
    Immediate start
    Relocation

    Torc Robotics

    Remote
    a month ago
  • $195k - $300k

     ...Built by a world-class team: Engineers, designers, and operators from...  ...— it's the foundation. As an ML Engineer, you'll be working at...  ...intersection of cutting-edge model development and real-world legal...  ...data curation and fine-tuning to evaluation and production deployment —... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours

    Eve

    Remote
    more than 2 months ago
  • $213k - $263k

     ...are currently focusing on include reinforcement learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid work schedule and you will report to a Research Director of AI... 
    Full time
    Remote work

    Waymo

    Remote
    more than 2 months ago
  •  ...science. We build foundational understanding of models to advance the frontier of intelligence. About the role: As a ML Engineer, you’ll build and operate the...  ...work on the systems that support training and evaluating large models, scaling experimental pipelines... 
    Full time
    Internship

    Tilde Research

    Remote
    more than 2 months ago
  • $140k - $300k

     ...of R1 and its AI innovation engine, R37 Lab , bringing Phare’s...  ...do it. The Role As a ML Engineer, you’ll lead the development...  ...Research Scientists to make models production-ready with clear...  ...decisions by defining the right evaluation datasets, metrics, and... 
    Full time
    Work at office
    Local area
    Flexible hours

    R37 Lab, R1 Rcm

    Remote
    12 days ago
  • $199.2k - $298.8k

     ...teams use it to train various models and simulation teams use it for...  ...latest developments in AI and ML for autonomous driving....  ...a customer-centric manner.  Evaluate and make recommendations regarding...  ...consensus.  Mentors and guides engineers within the group.   Bachelor... 
    Full time
    Immediate start
    Relocation

    Company

    Remote
    more than 2 months ago
  •  ...that carry our data from an assay recording to a reproducible model, an evaluated prediction, and a usable recommendation for the next...  ...controls Qualifications • Strong production software engineering experience in Python and modern machine-learning or data systems... 
    Full time
    Work at office
    Shift work

    Monarch

    Remote
    11 days ago
  • $175k - $200k

     ...Senior Machine Learning Engineer Truveta is the world...  ...clinician to be an expert, and help families make...  ...fine-tuned foundation models that continuously learn...  ...engineering. You’ll blend ML craftsmanship with platform...  .... Understand evaluation deeply — fluent in ML validation... 
    Full time
    For contractors
    Visa sponsorship
    Work visa
    Flexible hours

    Truveta

    Remote
    23 days ago
  • $100.4k - $180.7k

    Posting TitleML and Optimization Engineer.LocationCO - Golden.Position...  ...scientists, engineers, and experts are accelerating energy...  ...high‑performance computing, AI/ML, modeling and simulation, and...  ...end development and red‑team evaluation. NLR collaborates across DOE... 
    Full time
    Fixed term contract
    Live in
    Local area
    Remote work
    Relocation
    Shift work

    National Renewable Energy Laboratory

    Golden, CO
    5 days ago
  • $143.2k - $243.4k

     ...DescriptionIt all started when engineer Fred Luddy wrote code that...  ...looks like. The role As a Senior ML Engineer, you build core components...  ...of what you ship: tests, evaluation, observability, and the metrics...  ...wider system. The data and model plumbing that keeps the graph... 
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    1 day ago
  • $170k - $245k

     ...impact.Learn more at .As a Senior ML Engineer, you'll be building and...  ...measure, designing time-series models and the fallback strategies...  ...production workloads. Design, evaluate, and productionize time-series...  ...systems in production. Expert-level Python: you write typed... 
    Remote work

    CloudBolt Software

    Rockville, MD
    4 days ago
  • $85 per hour

     ...Summers , and Jack Dorsey . Position: ML Engineer (Coding Agent Experience) Type:...  ...frontier AI coding agents to complete and evaluate complex machine learning and AI engineering tasks. Review model-generated implementations involving model training... 
    Full time
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    a month ago
  • $175k - $215k

     ...learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. In this hybrid role, you will report to a...  ...industrial or research setting developing recipes for ML models We prefer: Track record of... 
    Full time
    Remote work

    Waymo

    New York, NY
    a month ago
  •  ...work spans cutting-edge LLMs / ML, large-scale data systems,...  ...joining a team of exceptional engineers, analysts, and investors working...  ...inform how you think about modeling, signals, and product design....  ...model training to inference, evaluation, and production operations.... 
    Full time
    Work at office
    Local area

    Versant Limited

    Remote
    a month ago
  • $251k - $310k

     ...from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid...  .... Partner effectively with engineering and research teams across...  ....g., xprof), and debugging of ML models. You have: PhD... 
    Full time
    Temporary work
    Remote work

    Waymo

    New York, NY
    a month ago
  • $220k - $247k

     ...As a  Senior Staff Machine Learning Engineer , you will operate at the company level—...  ...You will lead the design of large-scale ML systems and shared platforms that power all...  ...design of  scalable ML platforms (training, evaluation, inference, safety) used across multiple... 
    Full time
    Work at office
    Immediate start
    Flexible hours
    3 days per week

    Typeface

    Remote
    more than 2 months ago
  • $132.8k - $250.8k

     ...Information Systems, Data Engineering, or a related field, or...  ...database design/modeling, including designing or...  ...architectures. Understanding of AI/ML capabilities, with...  ..., a technical expert, a culture builder…or all...  ...data product ecosystem. Evaluate and implement metadata... 
    Immediate start
    Relocation package
    Flexible hours

    Ford

    Redford, MI
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Engineer - Model Evaluation Expert. Be the first to apply!