Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Engineer - Model Evaluation

$60 - $90 per hour

Mercor

Job Description

Job Description

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .

Position: Machine Learning Engineer — Model Evaluation & Experimentation
Type: Contract
Compensation: $60–$90/hour
Location: Remote
Commitment: 35 hours/week

Role Responsibilities

  • Design tasks by turning real ML research ideas into well-defined, multi-step tasks.
  • Run experiments by implementing changes, running training experiments, and analyzing results.
  • Explore reinforcement learning ideas, focusing on reward functions and training behavior.
  • Evaluate models to identify where frontier models fall short.
  • Collaborate with researchers and experts to maintain task consistency and rigor.

Qualifications

Must-Have

  • MSc or PhD in machine learning , computer science , or a related STEM field.
  • 1+ years of experience in a research or research-engineering role.
  • Experience in training and evaluating ML models and running experiments end-to-end.
  • Strong familiarity with large language models and their evaluation techniques.
  • Proficiency in Python and Git .

Preferred

  • Basic understanding of reinforcement learning .
  • Experience in AI training , model evaluation, or benchmark/task authoring.

Application Process (Takes 20–30 mins to complete)

  • Upload resume
  • AI interview based on your resume
  • Submit form

Resources & Support

PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the ML Engineer - Model Evaluation in New York, NY vacancy
  • $175k - $215k

     ...learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. In this hybrid role, you will report to a...  ...industrial or research setting developing recipes for ML models We prefer: - Track record of... 
    Suggested
    Full time
    Remote work

    Waymo

    New York, NY
    19 hours ago
  • $228.7k - $343.1k

     ...at enormous scale, and one bad model can mean millions in credit...  ...validate at scale, so you critically evaluate what it produces and own the...  ...in parallel. Reason about ML systems end to end — how...  ...accuracy. Solid software and data engineering: production-quality Python,... 
    Suggested
    Remote job
    Full time
    Local area
    Shift work

    Block

    New York, NY
    19 hours ago
  •  ...company. We build cutting-edge foundation AI models and end-to-end products that are...  ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all...  ...Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    4 days ago
  •  ...for an Insurance Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting judgments, with placement at the client lab as part of the extended... 
    Suggested
    Weekday work

    Mercor

    New York, NY
    2 days ago
  • Alignerr is seeking a Senior Python Infrastructure Engineer to design and build data pipelines, evaluation harnesses, and annotation tooling powering AI systems...  ...role focuses on production-grade Python systems for model evaluation at scale. You will collaborate with research... 
    Suggested
    Remote job
    Hourly pay
    Contract work

    Alignerr

    New York, NY
    5 hours ago
  • $400 per month

     ...lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic data engineering workflows and model evaluation. Spots are limited and filling quickly... 

    Mercor

    New York, NY
    1 day ago
  • We are seeking an expert to evaluate and improve our AI models through comprehensive testing and analysis. You will be responsible for designing...  ...outputs for bias, fairness, and accuracy Collaborate with ML engineers to implement improvements Document findings and... 

    MERIT Beauty

    New York, NY
    4 days ago
  • Senior Research Scientist, Model Evaluation Cohere | Posted Mar 2 | Full-time | New York | Negotiable | Unknown Why this role? Evaluation...  ...the capabilities you care about. You have strong software engineering skills. We value and celebrate diversity and strive to create... 
    Full time
    Work at office
    Remote work
    Flexible hours

    SupportFinity™

    New York, NY
    4 days ago
  • Mercor is seeking a Generalist who can operate in English and Punjabi. This contract, remote position focuses on evaluating AI outputs and supporting model evaluation tasks. You will conduct fact-checking, assess reasoning, clarity, tone and completeness, and provide actionable... 
    Remote job
    Contract work

    United States Digital Space LLC

    New York, NY
    2 days ago
  • Dorado is seeking an experienced Investment Banking SME to support the development, evaluation, and improvement of advanced AI models in finance. You will assess AI-generated analyses for accuracy, reasoned judgments, and data integrity, guiding model refinements and prompts... 
    Remote job

    Dorado

    New York, NY
    3 days ago
  • Cohere is seeking a Senior Research Engineer, Model Evaluation, to create next‑generation evaluation methods and scalable infrastructure. You will develop benchmarks, datasets, and environments to measure frontier model capabilities, and you will push the state‑of‑the‑... 

    cohere

    New York, NY
    1 day ago
  • $175k - $280k

     ...New York is seeking an expert in optimizing machine learning models to turbocharge their serving layer, integrating LLM, speech, and...  ...significant experience in systems programming and performance engineering, aiming to improve high-throughput, low-latency serving. Join... 

    SESAME

    New York, NY
    4 days ago
  • Dorado is seeking a Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems, probe model reasoning at the frontier, and help identify where models fail under rigorous scientific scrutiny.... 
    Remote job

    Dorado

    New York, NY
    4 days ago
  •  ...opportunities through their expert network. Qualified candidates should hold a MS or PhD in a relevant field and have experience in evaluating complex biology content. Strong communication skills and proficient English are essential for success in this role. #J-18808-... 
    Immediate start

    SME Careers

    New York, NY
    3 days ago
  • Mercor is seeking experienced Musicians to evaluate generative musical AI models in partnership with a leading AI lab. In this role, you will assess...  ...Ideal candidates have 3+ years as a music producer/audio engineer, a college degree in music, and native or near-native... 

    Mercor

    New York, NY
    2 days ago
  •  ...seeking Insurance domain SMEs to join a cutting-edge AI training program. You will evaluate AI model outputs against real underwriting practice and rubrics, guiding research and engineering teams to close knowledge gaps in underwriting, claims, and risk assessment. The... 

    Obsidian

    New York, NY
    4 days ago
  •  ...company. We build cutting‑edge foundation AI models and end‑to‑end products that are...  ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all...  ...Germany and Paris. Join us! Why this role? Evaluation is critical to making progress in scaling... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    cohere

    New York, NY
    3 days ago
  • Take a new model and get it running — correctly — on our ASIC in record time. When a frontier...  ...we're looking for 5+ years in systems or ML systems, with real depth in at least one...  ...or LLM-driven tooling that did real engineering work, not demos. #J-18808-Ljbffr General... 
    Live in

    General Compute Inc.

    New York, NY
    3 days ago
  • $152k - $241.5k

     ...working for us! We believe open-weight models are foundational to American AI...  ...scientific scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find,...  .... We are looking for an Evaluation/ML-Systems Engineer to own how we... 
    Full time
    Remote work

    Nvidia

    New York, NY
    1 day ago
  • SME Careers is seeking a remote Kotlin Engineer to review AI-generated responses and create...  ...optimizing AI performance, and ensuring model accuracy. The ideal candidate has a...  ...an expert network and requires critical evaluation of technical concepts. #J-18808-Ljbffr... 
    Remote job

    SME Careers

    New York, NY
    4 days ago
  • $85 per hour

     ...Jack Dorsey . Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience) Type: Contract...  ...Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving cloud platforms... 
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    2 days ago
  • $60 - $90 per hour

     ...Help shape how AI systems evaluate and preserve the timing, emotion, micro-expressions, and physical nuance of human and character performances...  ...-time role supports the development of performance-transfer models by defining high-quality evaluation standards and helping... 
    Hourly pay
    Part time

    SaidGig

    New York, NY
    16 days ago
  • Mercor partners with a leading AI research lab to support a Frontier Code Agents project, focusing on realistic data engineering workflows and model evaluation. Contributors help evaluate and improve frontier AI coding models through structured technical assessments and... 

    Mercor

    New York, NY
    1 day ago
  • SME Careers is looking for a remote R Engineer to review AI-generated responses and create high-quality R and data-analysis content. The position requires a strong background in R programming, applied statistics, and excellent writing skills to document analyses. Candidates... 
    Remote job

    SME Careers

    New York, NY
    4 days ago
  •  ...building extraction agents, evaluating accuracy, deploying to production...  ...architecture changes, prompt engineering, fine-tuning, or rule-based...  ...~2+ years building ML/AI systems in production ~...  ...infrastructure glue, not just model training scripts ~ Practical... 
    Full time

    Triomics

    New York, NY
    19 hours ago
  • $213k - $263k

     ...are currently focusing on include reinforcement learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid work schedule and reports to a Principal Research Scientist.... 
    Full time
    Temporary work
    Remote work

    Waymo

    New York, NY
    19 hours ago
  • $100k - $250k

     ...This innovative field blends AI, engineering, and materials science,...  .... The opportunity As an ML Engineer at Radical AI, you will...  ...also data ingestion, enrichment, model exporting and serving, and...  ...infrastructure. Continually evaluate model performance and maintain... 
    Full time

    Radical Ai

    New York, NY
    19 hours ago
  • $20 per hour

    Feedinkoo is looking for a Web Developer/Designer to enhance AI models by evaluating design work, including interfaces and user experiences. This role involves reviewing AI‑generated visuals and providing feedback to improve users' experience with AI tools. Working remotely... 
    Remote job

    Feedinkoo

    New York, NY
    19 hours ago
  • $80 per hour

     ...Summers , and Jack Dorsey . Position: Data Engineer (Coding Agent Experience) Type: Contract Compensation...  ...Use frontier AI coding agents to complete and evaluate complex data engineering tasks. Review model-generated implementations involving ETL pipelines... 
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    8 days ago
  •  ...workflow orchestration layer for computer vision models. Job Description We are seeking a Staff ML Engineer with a passion for building application-layer AI...  ...OpenAI, along with pgvector Langchain and evaluation frameworks in Langsmith Google Cloud, Docker... 
    Full time
    Work at office
    Work from home
    Flexible hours
    3 days per week

    Spara

    New York, NY
    19 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Engineer - Model Evaluation. Be the first to apply!