Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer for Model Evaluation

SaidGig

Role Overview

Author and run rigorous, multi-step machine learning evaluation tasks for a leading generative AI research team. You will take high-level research ideas, implement changes to training or evaluation procedures, execute experiments, and analyze results to determine where frontier models succeed or fail. Typical tasks require one to two days of continuous, focused work and combine implementation, experiment execution, and careful analysis. You will collaborate closely with the lab''s researchers in a fast feedback loop.

Key Responsibilities
  • Design tasks that translate real ML research ideas into well-defined, multi-step experiments, for example modifying how an RL reward is computed and specifying success criteria.
  • Implement changes to code and training pipelines required by each task.
  • Set up, run, and monitor training experiments end-to-end, ensuring reproducibility and correctness.
  • Analyze experiment outputs to demonstrate what a correct solution looks like and to identify model failure modes.
  • Build tasks that explore reinforcement learning concepts such as reward functions and training behavior.
  • Evaluate how frontier models handle your tasks, documenting where and why they fall short.
  • Share findings and align task design, rigour, and fairness with researchers and fellow experts.
Qualifications
  • MSc or PhD in machine learning, computer science, or another STEM field, or equivalent practical experience in a research-heavy role.
  • At least 1 year of experience in a research or research-engineering position.
  • Hands-on experience training and evaluating ML models and running experiments end-to-end, including setup, execution, and analysis.
  • Strong familiarity with large language models, including their capabilities, limitations, and common evaluation techniques.
  • Working proficiency in Python and Git, comfortable in both scripting and notebook environments.
  • Basic understanding of reinforcement learning concepts such as reward functions and policy training is preferred.
  • Prior experience in AI training, model evaluation, or benchmark or task authoring is preferred.
  • High attention to detail, creative task design, strong written communication, and the ability to work independently on ambiguous, open-ended problems.
  • Availability to engage reliably for approximately 35 hours per week.
Work Terms
  • Employment type: W-2 employment through Cincinnatus LLC, placed to work as part of a leading AI lab''s extended workforce.
  • Work location: fully remote within the United States.
  • Time commitment: approximately 35 hours per week, full-time role-based assignment integrated into the client team''s workflows.
  • Role structure: this is a structured, role-based position rather than a project-based or freelance engagement; it involves close collaboration with the client team.
Compensation
  • Hourly W-2 pay range: $60.00 to $90.00 per hour.
Eligibility and Application Process
  • Cincinnatus LLC serves as the employer of record and administers employment, onboarding, payroll, and benefits for this role.
  • To apply, submit your application through the job listing or application portal; applications and subsequent employment administration will be handled by Cincinnatus LLC.
  • Equal Employment Opportunity: Cincinnatus is an equal opportunity employer and provides reasonable accommodations for qualified applicants with disabilities.
Vacancy posted 12 days ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer for Model Evaluation in United States vacancy
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure... 
    Suggested
    Full time

    Scale Ai

    Washington DC
    1 day ago
  • $85 per hour

     ...Role Overview Evaluate and improve frontier AI coding agents by completing realistic machine learning engineering tasks and assessing model outputs. You will perform structured technical assessments that reflect production ML workflows, helping a leading AI research lab... 
    Suggested
    Hourly pay
    Remote work

    SaidGig

    United States
    16 days ago
  • $60 - $90 per hour

     ...Machine Learning Engineer — Model Evaluation & Experimentation is a remote review track for evaluating AI outputs across machine learning engineer model evaluation experimentation specialist operations workflows. Reviewers grade workflow correctness, policy adherence,... 
    Suggested
    Remote job
    For contractors
    Work experience placement
    10 hours per week

    AuraOne Human Data

    Remote
    14 days ago
  • $228.7k - $343.1k

     ...enormous scale, and one bad model can mean millions in credit losses...  ...at scale, so you critically evaluate what it produces and own the...  ...codebases you did not write, learning the data, configs, and...  .... Solid software and data engineering: production-quality Python, SQL... 
    Suggested
    Remote job
    Full time
    Local area
    Shift work

    Block

    New York, NY
    1 day ago
  • $174.72k - $295.68k

     ...through cutting-edge R&D in AI, machine learning, and smart connectivity.We are...  ...for a full-time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development of...  ...).Conduct systematic ablation, evaluation, and visualization of model... 
    Suggested
    Full time

    XPENG Motors

    Santa Clara, CA
    5 days ago
  •  ...systems enable robots to adapt, learn, and perform in the real...  ...fast, complex, and poorly modeled physics that traditional...  ...ones.We are seeking a Senior Machine Learning Engineer to lead the development of...  ...planning, process-optimization, evaluation, and synthetic-data... 
    Shift work

    Path Robotics

    Columbus, OH
    1 day ago
  • $213k - $263k

     ...for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge...  ...is "real"? We are seeking visionary machine learning engineers and researchers to architect the...  ...the realism of our multimodal world models. Your work will define the state of the... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $150k

     ...Machine Learning Engineer About the Institute of Foundation Models: We are a dedicated research lab for building, understanding, using, and risk-managing foundation...  ...support data pipelines, experimentation, and evaluation workflows. ~ This role balances fast-... 
    Visa sponsorship

    Institute of Foundation Models

    Sunnyvale, CA
    4 days ago
  • $70 per hour

     ...Role Overview Join a remote Machine Learning Engineer talent network to be considered for future contract engagements with AI labs and companies...  .... Members contribute to advancing AI by training and evaluating models, designing real-world tasks and deliverables, and... 
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    14 days ago
  • $224k - $356.5k

    We are looking for outstanding Machine Learning Engineers to join our Physical AI teams. As the pioneers...  ...state-of-the-art multimodal models and diffusion techniques to simulate...  ...Establish a strong mentality for KPI evaluation and validation to ensure the quality... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $204k - $259k

     ...The Driver Understanding and Evaluation (DUE) team at Waymo is...  ...the Waymo Driver.  The DUE Machine Learning team will build and operate...  ...and advanced machine learning models to deliver training and evaluation...  ...researchers and software engineers who are passionate about developing... 
    Full time

    Waymo

    Remote
    1 day ago
  • $170k - $216k

     ...15+ U.S. states. The DUE Machine Learning team will build and operate...  ...tools, improve and speed up the evaluation and onboard developer...  ...and advanced machine learning models to deliver training and evaluation...  ...researchers and software engineers who are passionate about... 
    Full time

    Waymo

    Remote
    1 day ago
  • $238k - $302k

     ...Foundations team is to develop machine learning solutions addressing open...  ...from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a...  ...to a Senior Staff Software Engineer.   You will: Work with... 
    Full time
    Remote work

    Waymo

    Remote
    1 day ago
  •  ...providing Information Technology, Engineering Services, Program Management, and...  ...Solerity is seeking Mid to Senior Machine Learning Engineers and AI Model Developers to support an upcoming...  ...This effort focuses on developing, evaluating, and integrating machine learning... 
    Full time
    For contractors
    Remote work
    Flexible hours

    Solerity

    Remote
    1 day ago
  • $100 per hour

     ...knowledge to help train and evaluate next-generation AI systems by...  ..., clarity, and relevance of model outputs through rubric-based...  ...experience in Data Science, Machine Learning, Applied AI, Statistics,...  ...desirable. Background in prompt engineering, AI output evaluation, fact... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    Indiana
    9 days ago
  •  ...Seattle, WA, we are a team of engineers and technologists from...  ...Senior Software Engineer – Evaluation, you will design and implement...  ...recognition (ASR), and small language model (SLM) systems. You will...  ...You will work closely with machine learning and data engineering teams... 

    VTI Aerospace

    Seattle, WA
    8 days ago
  •  ...customized by a team of experienced sellers, engineers, and researchers. Many of us worked on...  ...execs Pioneer the training of new models that leverage both historical data and synthetic...  ...You have a strong understanding of deep learning AI/ML frameworks or cloud services... 
    Full time

    Lightfield

    Remote
    1 day ago
  • $170k - $216k

     ...team builds the system which learns the spatial-temporal representation...  ...set of sensors, enabling engineers like you to (1) develop...  ...real-world data, to (2) develop models and model training at scale,...  ...experience ~3+ years experience in Machine Learning and/or Computer... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  •  ...build cutting-edge foundation AI models and end-to-end products that...  ...is a team of researchers, engineers, designers, and more, who are...  ...Paris. Join us!Why this role?Evaluation is critical to making progress...  ...credit.Education & learning stipend for conferences, courses... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    5 days ago
  • $80 per hour

     ...Role Overview Evaluate and improve frontier AI coding agents by applying professional data engineering judgment to realistic data infrastructure and pipeline scenarios. You will use and assess model-generated implementations for ETL, data warehouses, analytics platforms... 
    Hourly pay
    Remote work

    SaidGig

    United States
    16 days ago
  • $298k - $368k

     ...team builds the system which learns the spatial-temporal representation...  ...set of sensors, enabling engineers like you to (1) develop...  ...real-world data, to (2) develop models and model training at scale,...  ...~7+ years of experience in Machine Learning, with a focus on large... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $40 per hour

    A biotechnology company is seeking a Biotechnology R&D Scientist to train AI models and evaluate their outputs. This role involves measuring the progress of AI chatbots with complex biology questions and ensuring their performance and correctness. Candidates should have... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Wisconsin
    2 days ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    2 days ago
  • $224k - $356.5k

     ...high-performance computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in...  ...typically 12+ years) developing or assessing contemporary machine learning and deep learning systems.Hands-on experience with... 
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $272k - $431.25k

     ...re generating it! Our world model team is pushing the boundaries...  ...Manager to lead world-model evaluation and benchmarking across...  ...Strong research background in machine learning, computer vision, multimodal...  ...Computer Science, Electrical Engineering, Robotics, Machine Learning,... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $40 per hour

    A forward-thinking AI development firm seeks experienced quantitative professionals to evaluate AI-generated work, applying their skills in statistical analysis, predictive modeling, and technical writing. This fully remote opportunity offers a flexible schedule and projects... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Sioux Falls, SD
    2 days ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting... 
    Hourly pay
    Remote work

    DataAnnotation

    Juneau, AK
    2 days ago
  • $60 per hour

     ...professionals to help advance AI development. AI models are increasingly capable of performing...  ...-of-the-art AI models on tasks like evaluating AI-generated quantitative analysis,...  ...Statistics, Computer Science, Mathematics, Engineering, or similar); a master's or PhD is a... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Kansas City, MO
    2 days ago
  • $60 per hour

     ...A leading data analysis firm is seeking experienced quantitative professionals to join their remote team. In this role, you'll evaluate AI-generated quantitative analysis, design problem-solving tasks for AI training, and provide insightful feedback on AI systems. Candidates... 
    Remote work

    DataAnnotation

    Nevada, IA
    2 days ago
  • $20 per hour

     ...contractors to join our team and teach AI chatbots. You will develop complex prompts to test AI models, write high-quality responses to demonstrate excellence, and evaluate different model outputs based on accuracy and style guidelines. This role is ideal for professionals... 
    Hourly pay
    Full time
    Contract work
    Part time
    For contractors
    Self employment
    Freelance
    Remote work

    DataAnnotation

    Wyoming, OH
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer for Model Evaluation. Be the first to apply!