ML Engineer - Model Evaluation Expert
$60 - $90 per hourMercor
About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .
Position: Machine Learning Engineer — Model Evaluation & Experimentation
Type: Contract
Compensation: $60–$90/hour
Location: Remote
Commitment: 35 hours/week
Role Responsibilities
- Design tasks by transforming real ML research ideas into well-defined, multi-step tasks.
- Run experiments by implementing changes, executing training experiments, and analyzing results to define correct solutions.
- Explore reinforcement learning concepts such as reward functions and training behavior in task development.
- Evaluate frontier models' performance on tasks and identify areas of improvement.
- Collaborate with researchers to ensure tasks are consistent, rigorous, and fair.
Qualifications
Must-Have
- MSc or PhD in machine learning , computer science , or a related STEM field.
- 1+ years in a research or research-engineering role.
- Experience in training and evaluating ML models and conducting end-to-end experiments.
- Proficiency in Python and Git .
- Strong understanding of large language models and their evaluation.
Preferred
- Basic knowledge of reinforcement learning .
- Experience in AI training, model evaluation, or benchmark/task authoring.
Compensation & Legal
- Hourly contractor
- Paid weekly
Application Process (Takes 20–30 mins to complete)
- Upload resume
- AI interview based on your resume
- Submit form
Resources & Support
- For details about the interview process and platform information, please check:
- For any help or support, reach out to: View email address on jobs.jobcopilot.com
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
$189.4k - $300.6k
...behavior across real-world scenarios.The Evaluation Foundations team—part of Embodied AI’s... ...in autonomous vehicle development. We engineer high-performance tools that identify top-performing models and partner with data-intensive ML teams to drive rapid innovation.In this...SuggestedFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$281k - $356k
...improve and speed up the evaluation and onboard developer... ...journeys. It will combine expert human judgements and... ...advanced machine learning models to deliver training and... ...and software engineers who are passionate about... ...in Python and standard ML frameworks (e.g., JAX,...SuggestedFull time- ...seeking a highly skilled Machine Learning Engineer who specializes in leveraging Large Language Models (LLMs) for automated evaluation and quality assessment. In this role, you will... ...~5+ years of experience in ML engineering, NLP, or AI/ML automation. ~...SuggestedWork at officeRemote workFlexible hours
$189.4k - $300.6k
...behavior across real-world scenarios. The Evaluation Foundations team—part of Embodied AI’s... ...in autonomous vehicle development. We engineer high-performance tools that identify top-performing models and partner with data-intensive ML teams to drive rapid innovation. In...SuggestedFull timeRelocation packageFlexible hours$213k - $263k
...ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge:... ...are seeking visionary machine learning engineers and researchers to architect the... ...measure the realism of our multimodal world models. Your work will define the state of the...SuggestedFull timeRemote work$204k - $259k
...Driver Understanding and Evaluation (DUE) team at Waymo is... .... It will combine expert human judgements and advanced... ...machine learning models to deliver training and... ...researchers and software engineers who are passionate about... ...implement scalable and robust ML pipelines for training,...Full time$170k - $216k
...improve and speed up the evaluation and onboard developer... ...journeys. It will combine expert human judgements and... ...advanced machine learning models to deliver training and... ...and software engineers who are passionate about... ...distributed systems covering the ML lifecycle, supporting...Full time$238k - $302k
...from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid... ...to a Senior Staff Software Engineer. You will: Work with... ...~5+ years of experience in ML engineering and applied Deep Learning...Full timeRemote work- ...Innodata is expanding its team of technical experts in LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you will... ...optimize the technical foundations that power model improvement for foundation model builders...Full time
$152k - $241.5k
...working for us! We believe open-weight models are foundational to American AI... ...scientific scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find,... .... We are looking for an Evaluation/ML-Systems Engineer to own how we...Full timeRemote work- ...step machine learning evaluation tasks for a leading generative... ...where frontier models succeed or fail.... ...tasks that translate real ML research ideas into well... ...and fellow experts. Qualifications... ...research or research-engineering position. Hands-on...Hourly payFull timeFreelanceRemote work
$85 per hour
...Role Overview Evaluate and improve frontier AI coding agents by completing realistic machine learning engineering tasks and assessing model outputs. You will perform structured technical assessments that reflect production ML workflows, helping a leading AI research lab...Hourly payRemote work- ...the most advanced AI models. 1. Overview A leading... ...generation of agentic evaluation benchmarks for... ...to act as ground-truth experts for model evaluation and... ...author complex, multi-step ML tasks — for example,... ...research or research-engineering role. ~ Hands-on...Full timeContract workPart timeFreelanceRemote work
$177.3k - $212.8k
...teams use them to train various models, mapping teams use them to... ...latest developments in AI and ML for autonomous driving, 3D reconstruction... ...a customer-centric manner. - Evaluate and make recommendations... ...consensus. Mentors and guides engineers within the group. - Bachelor’...Full timeWork at officeImmediate startRelocation- ...efficiency. About the Role We’re hiring an ML Engineer to join Kodex’s Verifications / Threat... ...workflows into production-grade models, pipelines, and decision-support tooling... ...-to-end (feature generation, training, evaluation, batch/streaming inference, backfills, and...Remote jobFull timeFlexible hours
$199.2k - $298.8k
...teams use it to train various models and simulation teams use it for... ...latest developments in AI and ML for autonomous driving.... ...a customer-centric manner. Evaluate and make recommendations regarding... ...consensus. Mentors and guides engineers within the group. Bachelor...Full timeImmediate startRelocation$213k - $263k
...are currently focusing on include reinforcement learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid work schedule and reports to a Principal Research Scientist....Full timeTemporary workRemote work$195k - $300k
...Built by a world-class team: Engineers, designers, and operators from... ...— it's the foundation. As an ML Engineer, you'll be working at... ...intersection of cutting-edge model development and real-world legal... ...data curation and fine-tuning to evaluation and production deployment —...Full timeTemporary workWork at officeLocal areaFlexible hours- ...Machine Learning Engineer (Llama AI Platform) Location: Remote (Preferred... ...by open-source large language models, agentic workflows, and... ...workflows. Fine-tune and evaluate LLM performance for business use... ...tuning open-source LLMs. ML Engineering and MLOps...Full timeRemote work
$153.2k - $183.3k
...Team As a Machine Learning Engineer II – Road & Lane, you will help develop next‑generation models that estimate road surfaces,... ...Implement scalable training and evaluation pipelines for lane perception... ...Hands‑on experience developing ML models for perception tasks...Remote jobFull time- ...science. We build foundational understanding of models to advance the frontier of intelligence. About the role: As a ML Engineer, you’ll build and operate the... ...work on the systems that support training and evaluating large models, scaling experimental pipelines...Full timeInternship
$153.2k - $183.3k
...the Team: As a Machine Learning Engineer II – Learned Behaviors, you will help develop and deploy behavior models that power decision-making for... ...~ Implement production-quality ML code to support model training, evaluation, and inference within the autonomy...Full time- ...About the Role We’re looking for an Applied ML Engineer to design, evaluate, and scale recommendation and ranking systems that power how content... ...techniques, grounding them in real user behavior, and shipping models that perform reliably at scale. The work is hands-on,...Full time
$200k - $300k
...preparation configurations. As a Senior ML Engineer, Manipulation, you will own the learning... ...parallel jaw, multi-finger) Implement and evaluate modern policy architectures (diffusion policies, transformer-based action models, action chunking) and adapt them to Chef's...Full timeFlexible hours$213k - $263k
...are currently focusing on include reinforcement learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid work schedule and you will report to a Research Director of AI...Full timeRemote work- ...Responsibilities: Own evaluation pipelines — design, build, and... ...keep our speech and multimodal models honest in production. Harness... ...Required Qualifications: Expert-level PyTorch. Proven software engineer who loves ML; comfortable writing production...Full timeContract workFlexible hoursShift work
$200k - $300k
...ingredients — it will come from foundation models that generalize across thousands of food... ...Food Foundation Model. As a Senior ML Engineer, Foundation Models, you will work at the... ...for the Food Foundation Model — evaluating tradeoffs across generalization, sample...Full timeFlexible hours$174k - $253k
...implement robust agents and LLM-powered journeys that evaluate, gate, and improve engineering artifacts.Engineer and refine skills and context provided... ...experience.5 years of experience with Large Language Modeling (LLM).3 years of experience with generative AI Agents.Preferred...$100.4k - $180.7k
Posting TitleML and Optimization Engineer.LocationCO - Golden.Position... ...scientists, engineers, and experts are accelerating energy... ...high‑performance computing, AI/ML, modeling and simulation, and... ...end development and red‑team evaluation. NLR collaborates across DOE...Full timeFixed term contractLive inLocal areaRemote workRelocationShift work$175k - $215k
...learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. In this hybrid role, you will report to a... ...industrial or research setting developing recipes for ML models We prefer: - Track record of...Full timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Engineer - Model Evaluation Expert. Be the first to apply!
- graduate machine learning engineer Remote
- senior ml engineer Remote
- data scientist machine learning engineer Remote
- machine learning engineer Remote
- machine learning ai engineer Remote
- machine learning software engineer Remote
- ai ml engineer Remote
- computer vision machine learning engineer Remote
- junior machine learning research engineer Remote
- machine learning researcher Remote




