Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer for AI Model Evaluation

$85 per hour

SaidGig

Role Overview

Help evaluate and improve frontier AI coding models by completing structured technical assessments that simulate realistic machine learning engineering workflows. This role focuses on using and critiquing coding agents to surface bugs, failure modes, and deployment tradeoffs, with limited spots available on a first come, first serve basis.

Key Responsibilities
  • Use frontier AI coding agents to complete and evaluate complex ML and AI engineering tasks.
  • Review model-generated implementations related to model training, inference systems, MLOps, and LLM applications.
  • Identify bugs, edge cases, performance problems, and failure modes in model outputs and implementations.
  • Compare outputs from multiple frontier models, assessing their relative strengths and weaknesses.
  • Apply professional engineering judgment to realistic ML engineering scenarios and recommend actionable improvements.
Qualifications
  • At least 2 years of professional machine learning engineering experience.
  • Experience building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated ML implementations and reason about technical tradeoffs.
  • Experience deploying ML systems to production is preferred.
Work Terms
  • Remote engagement.
  • Employment type, hourly.
  • Sprint-based project work, with assignments running in 12 to 24 hour stretches depending on client requirements.
  • Spots are limited and are filled on a first come, first serve basis.
  • Task availability and placement depend on client needs.
Compensation
  • Listed hourly rate: $85 per hour.
  • Alternate payment detail: $400 per accepted task, with typical tasks taking approximately 2 to 3 hours after ramp-up.
  • Compensation is tied to accepted work.
Eligibility
  • Ideal candidates have 2 or more years of professional ML engineering experience and relevant production deployment experience.
  • Regular practical experience with AI coding agents and with evaluating model-generated engineering outputs is required.
Vacancy posted 16 days ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer for AI Model Evaluation in United States vacancy
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure... 
    Suggested
    Full time

    Scale Ai

    Washington DC
    1 day ago
  •  ...Author and run rigorous, multi-step machine learning evaluation tasks for a leading generative AI research team. You will take...  ...to determine where frontier models succeed or fail. Typical tasks...  ...experience in a research or research-engineering position. Hands-on... 
    Suggested
    Hourly pay
    Full time
    Freelance
    Remote work

    SaidGig

    United States
    12 days ago
  • $228.7k - $343.1k

     ...enormous scale, and one bad model can mean millions in...  ...to models applies to AI. We build the tooling that...  ..., so you critically evaluate what it produces and own...  ...you did not write, learning the data, configs, and...  ...Solid software and data engineering: production-quality Python... 
    Suggested
    Remote job
    Full time
    Local area
    Shift work

    Block

    New York, NY
    1 day ago
  • $174.72k - $295.68k

     ..., integrating advanced AI and autonomous driving...  ...cutting-edge R&D in AI, machine learning, and smart connectivity...  ...-time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development...  ...systematic ablation, evaluation, and visualization of... 
    Suggested
    Full time

    XPENG Motors

    Santa Clara, CA
    6 hours ago
  • $60 - $90 per hour

     ...Machine Learning Engineer — Model Evaluation & Experimentation is a remote review track for evaluating AI outputs across machine learning engineer model evaluation experimentation specialist operations workflows. Reviewers grade workflow correctness, policy adherence,... 
    Suggested
    Remote job
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    14 days ago
  •  ...embodied intelligence. Our AI-driven systems enable robots to adapt, learn, and perform in the...  ..., complex, and poorly modeled physics that...  ...are seeking a Senior Machine Learning Engineer to lead the development...  ...process-optimization, evaluation, and synthetic-data workflows... 
    Shift work

    Path Robotics

    Columbus, OH
    1 day ago
  • $213k - $263k

     ...state-of-the-art Generative AI to create a training ground...  ...Driver. The Simulator Evaluation team faces the ultimate data...  ...We are seeking visionary machine learning engineers and researchers to architect...  ...realism of our multimodal world models. Your work will define the... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $150k

     ...Machine Learning Engineer About the Institute of Foundation Models: We are a dedicated research lab for building, understanding,...  ...nurture the next generation of AI builders, and drive transformative...  ..., experimentation, and evaluation workflows. ~ This role... 
    Visa sponsorship

    Institute of Foundation Models

    Sunnyvale, CA
    4 days ago
  • $70 per hour

     ...Role Overview Join a remote Machine Learning Engineer talent network to be considered for future contract engagements with AI labs and companies. This is an open application...  ...contribute to advancing AI by training and evaluating models, designing real-world tasks and... 
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    14 days ago
  • $224k - $356.5k

    We are looking for outstanding Machine Learning Engineers to join our Physical AI teams. As the pioneers of the GPU—the...  ...state-of-the-art multimodal models and diffusion techniques to simulate...  ...Establish a strong mentality for KPI evaluation and validation to ensure the... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $204k - $259k

     ...Driver Understanding and Evaluation (DUE) team at Waymo is...  ...Driver.  The DUE Machine Learning team will build and operate...  ...machine learning models to deliver training and...  ...and software engineers who are passionate about...  ...modeling and generative AI into robust, production... 
    Full time

    Waymo

    Remote
    1 day ago
  • $238k - $302k

     ...The mission of the Waymo AI Foundations team is to develop machine learning solutions addressing...  ...demonstration, generative modeling, Bayesian inference,...  ...hierarchical learning, and robust evaluation. This role follows a...  ...Senior Staff Software Engineer.   You will:... 
    Full time
    Remote work

    Waymo

    Remote
    1 day ago
  •  ...providing Information Technology, Engineering Services, Program Management, and...  ...Solerity is seeking Mid to Senior Machine Learning Engineers and AI Model Developers to support an upcoming...  ...This effort focuses on developing, evaluating, and integrating machine learning... 
    Full time
    For contractors
    Remote work
    Flexible hours

    Solerity

    Remote
    1 day ago
  • $251k - $310k

     ...S. states. The DUE Machine Learning team will build and operate...  ...and speed up the evaluation and onboard developer...  ...advanced machine learning models to deliver training...  ...researchers and software engineers who are passionate...  ...for evaluating complex AI systems. ~ Track record... 
    Full time

    Waymo

    Remote
    1 day ago
  • $100 per hour

     ...knowledge to help train and evaluate next-generation AI systems by reviewing,...  ...clarity, and relevance of model outputs through rubric-based...  ...in Data Science, Machine Learning, Applied AI, Statistics,...  ...Background in prompt engineering, AI output evaluation, fact... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    Indiana
    9 days ago
  •  ...VTI Aerospace builds AI-powered perception and...  ...WA, we are a team of engineers and technologists from...  ...Senior Software Engineer – Evaluation, you will design and...  ...), and small language model (SLM) systems. You...  ...will work closely with machine learning and data engineering teams... 

    VTI Aerospace

    Seattle, WA
    8 days ago
  •  ...team of experienced sellers, engineers, and researchers. Many of us worked...  ...the role Lightfield's AI/ML team builds the experiences...  ...Pioneer the training of new models that leverage both historical...  ...strong understanding of deep learning AI/ML frameworks or cloud services... 
    Full time

    Lightfield

    Remote
    1 day ago
  • $80 per hour

     ...Role Overview Evaluate and improve frontier AI coding agents by applying professional data engineering judgment to realistic data infrastructure and pipeline scenarios. You will use and assess model-generated implementations for ETL, data warehouses, analytics platforms... 
    Hourly pay
    Remote work

    SaidGig

    United States
    16 days ago
  • $224k - $356.5k

     ...tapping into the unlimited potential of AI to define the next era of computing....  .... As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful...  ...developing or assessing contemporary machine learning and deep learning systems.Hands... 
    Full time

    Nvidia

    Santa Clara, CA
    6 hours ago
  • $40 per hour

    A biotechnology company is seeking a Biotechnology R&D Scientist to train AI models and evaluate their outputs. This role involves measuring the progress of AI chatbots with complex biology questions and ensuring their performance and correctness. Candidates should have... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Wisconsin
    2 days ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    2 days ago
  • $40 per hour

    A forward-thinking AI development firm seeks experienced quantitative professionals to evaluate AI-generated work, applying their skills in statistical analysis, predictive modeling, and technical writing. This fully remote opportunity offers a flexible schedule and projects... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Sioux Falls, SD
    2 days ago
  • $60 per hour

     ...A leading data analysis firm is seeking experienced quantitative professionals to join their remote team. In this role, you'll evaluate AI-generated quantitative analysis, design problem-solving tasks for AI training, and provide insightful feedback on AI systems. Candidates... 
    Remote work

    DataAnnotation

    Nevada, IA
    2 days ago
  • $60 per hour

     ...A leading AI development company is seeking quantitative professionals to evaluate AI-generated analyses and develop solutions in various quantitative fields. This fully remote position allows for flexible scheduling and competitive pay up to $60/hour. Candidates should... 
    Remote work
    Flexible hours

    DataAnnotation

    Providence, RI
    2 days ago
  • $20 per hour

     ...DataAnnotation is committed to creating quality AI. Join our team to help train AI chatbots while...  ...chatbots. You will develop complex prompts to test AI models, write high-quality responses to demonstrate excellence, and evaluate different model outputs based on accuracy and... 
    Hourly pay
    Full time
    Contract work
    Part time
    For contractors
    Self employment
    Freelance
    Remote work

    DataAnnotation

    Wyoming, OH
    2 days ago
  • $60 per hour

     ...A tech company focused on AI is seeking quantitative professionals to evaluate AI-generated analytical work and help advance AI development. Responsibilities include statistical analysis, predictive modeling, and providing feedback to shape AI systems. The role offers... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Helena, MT
    3 days ago
  • $60 per hour

     ...contribute to developing cutting-edge AI systems, while enjoying the...  ...advance AI development. AI models are increasingly capable of...  ...-art AI models on tasks like evaluating AI-generated quantitative...  ...Computer Science, Mathematics, Engineering, or similar); a master's or... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Kansas City, MO
    2 days ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting... 
    Hourly pay
    Remote work

    DataAnnotation

    Juneau, AK
    2 days ago
  • $60 per hour

    A leading data science company is seeking quantitative professionals to evaluate AI-generated analyses and design quantitative problems crucial for advancing AI systems. This fully remote role allows you to work flexibly and choose your projects, with competitive hourly... 
    Hourly pay
    Remote work

    DataAnnotation

    Columbia, SC
    2 days ago
  • $60 per hour

    A leading AI development firm is seeking experienced quantitative professionals to contribute to AI systems. The role involves evaluating AI-generated work and solving technical problems with a flexible remote schedule. Candidates should have a background in data science... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Nashville, TN
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer for AI Model Evaluation. Be the first to apply!