Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer for AI Model Evaluation

$85 per hour

SaidGig

Role Overview

Evaluate and improve frontier AI coding models by completing structured technical assessments that mirror realistic machine learning engineering workflows, model training and inference systems, MLOps, and LLM application scenarios.

Key Responsibilities
  • Use frontier AI coding agents to complete and evaluate complex ML and AI engineering tasks.
  • Review model-generated implementations across model training, inference systems, deployment infrastructure, and LLM applications.
  • Identify bugs, edge cases, performance regressions, and failure modes in model outputs and implementations.
  • Compare outputs from multiple frontier models and assess their relative strengths, weaknesses, and tradeoffs.
  • Apply professional engineering judgment to realistic ML engineering scenarios, documenting findings and recommendations.
Qualifications
  • At least 2 years of professional machine learning engineering experience.
  • Experience building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated machine learning implementations and reason about technical tradeoffs.
  • Experience deploying ML systems to production is preferred.
Work Terms
  • Location: Remote.
  • Employment type: hourly.
  • Sprint-based engagement, with work organized into 12-24 hour stretches based on client requirements.
  • Spots are limited and are filled on a first-come, first-serve basis.
Compensation
  • $400 per accepted task.
  • Typical tasks take approximately 2-3 hours after ramp-up.
  • Compensation is tied to accepted work.
  • Hourly rate (metadata): $85 per hour.
Eligibility

This role is intended for engineers with 2+ years of ML engineering experience who regularly use AI coding agents and can assess model-generated ML solutions. Preference is given to candidates with experience deploying ML systems to production. Compensation is contingent on accepted deliverables.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer for AI Model Evaluation in United States vacancy
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure... 
    Suggested
    Full time

    Scale Ai

    Washington DC
    1 day ago
  • $60 - $90 per hour

     ...Machine Learning Engineer — Model Evaluation & Experimentation is a remote review track for evaluating AI outputs across machine learning engineer model evaluation experimentation specialist operations workflows. Reviewers grade workflow correctness, policy adherence,... 
    Suggested
    Remote job
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  • $70 per hour

     ...Role Overview Join a distributed talent network to provide expert Machine Learning engineering support on contract projects with AI labs and companies. Contributors help train and evaluate models, design real-world tasks and deliverables, and give domain-specific feedback... 
    Suggested
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    5 days ago
  • $228.7k - $343.1k

     ...enormous scale, and one bad model can mean millions in...  ...to models applies to AI. We build the tooling that...  ..., so you critically evaluate what it produces and own...  ...you did not write, learning the data, configs, and...  ...Solid software and data engineering: production-quality Python... 
    Suggested
    Remote job
    Full time
    Local area
    Shift work

    Block

    New York, NY
    1 day ago
  • $213k - $263k

     ...state-of-the-art Generative AI to create a training ground...  ...Driver. The Simulator Evaluation team faces the ultimate data...  ...We are seeking visionary machine learning engineers and researchers to architect...  ...realism of our multimodal world models. Your work will define the... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $60 - $90 per hour

     ...Drive research-grade data analyses and create evaluation tasks that reveal where frontier generative AI models fail. You will author realistic, multi-skill analysis...  ...1 year of experience in a research, research-engineering, or intensive data-analysis role. Strong, hands... 
    Hourly pay
    Full time
    Freelance
    Remote work

    SaidGig

    United States
    7 days ago
  •  ...Role Overview Medical professionals evaluate AI-generated medical content and use their clinical and field experience to improve model outputs. No prior AI experience is required. In this role you will assess model responses, create realistic prompts that reflect clinical... 
    Hourly pay
    Part time
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    6 days ago
  • $40 - $65 per hour

     ...Contribute domain expertise to evaluate and harden frontier large language models by crafting adversarial multi-turn...  ...improving how next-generation AI systems learn, reason, and behave, and does...  ...evaluations, annotation, or prompt engineering. Preferred: deep familiarity... 
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Frontier County, NE
    1 day ago
  • $238k - $302k

     ...The mission of the Waymo AI Foundations team is to develop machine learning solutions addressing...  ...demonstration, generative modeling, Bayesian inference,...  ...hierarchical learning, and robust evaluation. This role follows a...  ...Senior Staff Software Engineer.   You will:... 
    Full time
    Remote work

    Waymo

    Remote
    1 day ago
  • $80 per hour

     ...Role Overview Help evaluate and improve frontier AI coding models by using AI coding agents to complete realistic data engineering tasks, then assess the outputs for correctness, scalability, and failure modes. Work centers on end-to-end data engineering workflows including... 
    Hourly pay
    Work at office
    Remote work

    SaidGig

    United States
    2 days ago
  • $204k - $259k

     ...Driver Understanding and Evaluation (DUE) team at Waymo is...  ...Driver.  The DUE Machine Learning team will build and operate...  ...machine learning models to deliver training and...  ...and software engineers who are passionate about...  ...modeling and generative AI into robust, production... 
    Full time

    Waymo

    Remote
    1 day ago
  • $120 per hour

     ...Government Quantitative Professionals evaluate AI model outputs and provide structured, domain-...  ...by the research project. Optionally learn new evaluation techniques and tools as...  ...techniques to practical problems in science, engineering, business, or security, and sharing... 
    Hourly pay
    Temporary work
    Part time
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    6 days ago
  •  ...providing Information Technology, Engineering Services, Program Management, and...  ...Solerity is seeking Mid to Senior Machine Learning Engineers and AI Model Developers to support an upcoming...  ...This effort focuses on developing, evaluating, and integrating machine learning... 
    Full time
    For contractors
    Remote work
    Flexible hours

    Solerity

    Remote
    1 day ago
  • $251k - $310k

     ...S. states. The DUE Machine Learning team will build and operate...  ...and speed up the evaluation and onboard developer...  ...advanced machine learning models to deliver training...  ...researchers and software engineers who are passionate...  ...for evaluating complex AI systems. ~ Track record... 
    Full time

    Waymo

    Remote
    1 day ago
  • $39 per hour

     ...Music Audio Expert, where you will play a crucial role in evaluating generative musical AI models in collaboration with a leading AI lab. This position...  ...quality, and production/mix quality, using audio-engineering terminology. Annotate songs in detail, including genre... 
    Hourly pay
    Part time
    Immediate start
    10 hours per week

    SaidGig

    United States
    25 days ago
  • $60 - $80 per hour

     ...foundational Large Language Models, by designing realistic marketing...  ...accurate solutions, and evaluating model outputs with rigorous,...  ...places you within a leading AI lab''s extended workforce while...  ...Responsibilities Guide research and engineering teams to close knowledge gaps... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    6 days ago
  • $30 - $90 per hour

     ...Role Overview Build and maintain backend services in Go while testing and evaluating alpha-stage AI coding models. This part-time, remote contract role combines hands-on Go development with structured model evaluation, bug reporting, and real-time collaboration to improve... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    2 days ago
  • $75 per hour

     ...Role Overview Lead the clinical evaluation of medical AI by designing and executing assessments that probe clinical reasoning and decision-making...  ...teams to create realistic clinical problems, find where models fail or lack medical knowledge, and shape improvements so AI... 
    Contract work
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  • $60 - $150 per hour

     ...Role Overview Provide legal subject-matter expertise to improve and evaluate AI systems, by designing realistic legal tasks, reviewing model outputs, and giving domain-specific feedback that advances frontier AI research. This is an open application to join a Law Expert... 
    Hourly pay
    Contract work
    Immediate start
    Remote work

    SaidGig

    United States
    4 days ago
  • $85 per hour

     ...environmental assessment, GIS analysis, and renewable energy siting to evaluate AI-generated geospatial outputs and develop expert training data...  ..., GIS Analyst, Interconnection Specialist, Solar Design Engineer, or equivalent positions in the energy industry. Hands-on... 
    Hourly pay
    Contract work
    Part time
    Work at office
    Remote work
    Flexible hours

    SaidGig

    United States
    6 days ago
  •  ...team of experienced sellers, engineers, and researchers. Many of us worked...  ...the role Lightfield's AI/ML team builds the experiences...  ...Pioneer the training of new models that leverage both historical...  ...strong understanding of deep learning AI/ML frameworks or cloud services... 
    Full time

    Lightfield

    Remote
    1 day ago
  • $40 per hour

     ...data annotation company is seeking professionals in quantitative fields to enhance AI development. This fully remote role allows individuals to set flexible schedules while evaluating AI-generated analyses and solving complex quantitative problems. Candidates should have... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Oregon, WI
    3 days ago
  • $40 per hour

    A leading AI development firm in Michigan is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the evolution of AI models. Candidates should have a background in data science, statistics, or similar fields, with at least... 
    Hourly pay
    Remote work

    DataAnnotation

    Lansing, MI
    3 days ago
  • $40 per hour

     ...A forward-thinking AI team is seeking quantitative professionals to evaluate and improve cutting-edge AI systems. The role involves working on AI-generated analyses and providing critical feedback for model enhancement. Candidates should have a strong quantitative background... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Helena, MT
    9 hours ago
  • $60 per hour

     ...contribute to developing cutting-edge AI systems, while enjoying the...  ...advance AI development. AI models are increasingly capable of...  ...-art AI models on tasks like evaluating AI-generated quantitative...  ...Computer Science, Mathematics, Engineering, or similar); a master's or... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Little Rock, AR
    4 days ago
  • $110 per hour

     ...Role Overview Join a Physician Expert Network to provide clinical expertise to AI research teams and companies. Members contribute medical knowledge to train and evaluate AI models, design realistic clinical tasks, and give domain-specific feedback that advances medical... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    10 days ago
  • $40 per hour

    A data science team is seeking experienced quantitative professionals to evaluate AI-generated work and contribute to the development of cutting-edge AI systems. This fully remote position offers flexible scheduling and competitive hourly pay starting at $40+. Ideal candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Madison, WI
    3 days ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals for a remote role. Candidates will evaluate AI-generated quantitative work, solve complex problems, and provide valuable feedback. The ideal candidate has 2+ years of experience in a quantitative... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Topeka, KS
    3 days ago
  • $40 per hour

     ...An innovative AI development company is seeking experienced quantitative professionals to contribute to AI advancements. This fully remote role involves evaluating AI-generated analyses and ensuring they are technically accurate and valid in real-world scenarios. Candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Santa Fe, NM
    9 hours ago
  •  ...Physicians apply clinical judgment and frontline medical experience to evaluate AI-generated medical content, ensuring clinical accuracy, sound...  ...planning. Assess clarity, relevance, and safety of model outputs in realistic care scenarios. Provide detailed, constructive... 
    Full time
    For contractors
    Private practice
    Remote work
    Flexible hours

    SaidGig

    United States
    6 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer for AI Model Evaluation. Be the first to apply!