Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

DevOps Engineer for AI Model Evaluation

$85 per hour

SaidGig

Role Overview

Contribute to evaluations that improve frontier AI coding models by completing and assessing realistic infrastructure engineering tasks. You will work with model-generated code and implementations that touch cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation, applying professional engineering judgment to identify failure modes and reliability issues.

Key Responsibilities
  • Use frontier AI coding agents to complete complex infrastructure engineering tasks, then evaluate the results.
  • Review model-generated implementations for cloud platforms, Kubernetes, CI/CD, observability, and infrastructure-as-code.
  • Identify bugs, edge cases, reliability problems, and failure modes in model outputs.
  • Compare and contrast outputs from multiple frontier models to assess strengths and weaknesses.
  • Apply professional engineering judgment to realistic scenarios to judge correctness, safety, and operational readiness.
Qualifications
  • At least 2 years of professional experience in DevOps, SRE, or Cloud Engineering.
  • Hands-on experience with one or more cloud providers, such as AWS, Azure, or GCP.
  • Familiarity with Kubernetes, Terraform, CI/CD pipelines, and observability tooling.
  • Regular use of AI coding agents, for example Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.
Work Terms
  • Remote engagement.
  • Hourly employment model.
  • Work is organized in sprint windows that run in 12 to 24 hour stretches, scheduled per client requirement.
  • Typical tasks require about 2 to 3 hours of work after an initial ramp-up.
  • Spots are limited and are filled on a first come, first serve basis.
  • Project participation and assignment depend on client project availability.
Compensation
  • Hourly rate stated: $85 per hour.
  • Alternate task-based pay: $400 per accepted task.
  • Payment is tied to accepted work.
Eligibility
  • Meets the qualifications listed above, including the 2+ years of relevant professional experience.
  • Capability to use and evaluate outputs from AI coding agents.
  • No additional work-authorization details were specified, applicants should ensure they are eligible to work remotely under the engagement terms.
Vacancy posted 16 days ago
Similar jobs that could be interesting for youBased on the DevOps Engineer for AI Model Evaluation in United States vacancy
  • $400 per month

     ...partnering with a leading AI research lab to support a...  .... Contributors help evaluate and improve frontier AI coding models through structured technical...  ...realistic infrastructure engineering workflows and model evaluation...  ...2+ years of professional DevOps, SRE, or Cloud... 
    Suggested

    Mercor Inc

    Doral, FL
    2 days ago
  • $85 per hour

     ...technical talent with leading AI research labs. Headquartered...  ...Dorsey . Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience)...  ...coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving... 
    Suggested
    Contract work
    Summer work
    Remote work

    Mercor

    New York, NY
    5 days ago
  • $85 per hour

     ...technical talent with leading AI research labs. Headquartered...  ...Dorsey . Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience)...  ...coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving... 
    Suggested
    Contract work
    Summer work
    Remote work

    Mercor

    Miami, FL
    5 days ago
  • $70 - $150 per hour

     ...Apply to join a talent network for future remote DevOps and platform engineering contract opportunities with AI research labs and companies. Selected experts...  ...this domain. Key Responsibilities Train and evaluate AI models for DevOps and platform engineering use cases.... 
    Suggested
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    14 days ago
  • $208k - $300k

     ...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure... 
    Suggested
    Full time

    Scale Ai

    Washington DC
    1 day ago
  • $40 per hour

    A biotechnology company is seeking a Biotechnology R&D Scientist to train AI models and evaluate their outputs. This role involves measuring the progress of AI chatbots with complex biology questions and ensuring their performance and correctness. Candidates should have... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Wisconsin
    2 days ago
  •  ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant... 
    Remote work
    Flexible hours

    Prolific

    Charlotte, NC
    2 days ago
  • $224k - $356.5k

     ...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our...  ...performance computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in crafting... 
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $40 per hour

    A forward-thinking AI development firm seeks experienced quantitative professionals to evaluate AI-generated work, applying their skills in statistical analysis, predictive modeling, and technical writing. This fully remote opportunity offers a flexible schedule and projects... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Sioux Falls, SD
    2 days ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting... 
    Hourly pay
    Remote work

    DataAnnotation

    Juneau, AK
    2 days ago
  • $60 per hour

     ...contribute to developing cutting-edge AI systems, while enjoying the...  ...advance AI development. AI models are increasingly capable of...  ...-art AI models on tasks like evaluating AI-generated quantitative...  ...Computer Science, Mathematics, Engineering, or similar); a master's or... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Kansas City, MO
    2 days ago
  • $40 per hour

     ...A forward-thinking AI solutions company is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the development of cutting-edge...  ...skills. Join us to directly impact the future of AI analytics and model reasoning. #J-18808-Ljbffr... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Lincoln, NE
    2 days ago
  • $40 per hour

     ...A leading AI development firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully remote position allows for a flexible schedule, offering competitive pay starting at $40+ per hour. Ideal candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Bismarck, ND
    2 days ago
  • $40 per hour

    A leading AI training company is seeking a Biotechnology R&D Scientist to evaluate and improve AI models by posing complex biological questions. This remote position allows candidates to choose projects and work on their own schedule, with rates starting at $40+ per hour... 
    Hourly pay
    Remote work

    DataAnnotation

    Sioux Falls, SD
    5 days ago
  • $30 - $90 per hour

     ...maintain backend services in Go while evaluating and training alpha-stage AI coding tools. This contract role...  ...combines hands-on Go development, backend engineering, and structured testing of AI-...  .... Test and evaluate alpha AI models using Cursor, conducted over multiple... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    a month ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals to work remotely. In this role, you'll evaluate AI-generated quantitative work and solve technical problems while providing feedback to shape AI systems. Qualifications include 2+ years... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Salt Lake City, UT
    2 days ago
  • $40 per hour

     ...A leading AI development company is seeking experienced quantitative professionals to evaluate and validate AI systems. The role is fully remote, offering flexibility in project...  ...-generated work and designing problems for model training, contributing to shaping the future... 
    Hourly pay
    Remote work

    DataAnnotation

    Columbia, SC
    2 days ago
  • $40 per hour

     ...A leading AI company in the United States is seeking experienced quantitative professionals to evaluate and validate AI-generated analytical work. This fully remote position allows you to set your own schedule, with competitive hourly pay starting at $40 USD. Responsibilities... 
    Hourly pay
    Remote work

    DataAnnotation

    Jackson, MS
    2 days ago
  • $40 per hour

     ...A leading AI development firm is seeking experienced quantitative professionals to join their remote team. You will evaluate AI-generated quantitative analysis and solve complex problems to ensure technical accuracy. The ideal candidate should have at least 2 years of... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Denver, CO
    6 hours ago
  • $40 per hour

    A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative work and provide critical feedback. This role offers the flexibility of remote work, allowing you to set your own schedule while focusing on impactful projects... 
    Hourly pay
    Remote work

    DataAnnotation

    Wisconsin
    2 days ago
  • $40 per hour

     ...A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative analysis and provide impactful feedback. This fully remote role allows for flexible scheduling and competitive pay starting at $40 per hour. Candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Honolulu, HI
    3 days ago
  • $40 per hour

     ...A leading AI development firm is seeking experienced quantitative professionals to join their team remotely. The role involves evaluating AI-generated quantitative work, providing insights, and shaping the future of AI systems. Candidates should have over two years of... 
    Hourly pay
    Remote work

    DataAnnotation

    Indiana, PA
    2 days ago
  •  ...technology company is seeking a Director of Finance to enhance AI models relevant to finance. This role allows for flexible remote work...  ...expertise in financial reasoning. Responsibilities include evaluating AI performance and providing structured feedback. This is an independent... 
    Hourly pay
    Full time
    Part time
    For contractors
    Remote work
    Flexible hours

    DataAnnotation

    Maine
    2 days ago
  • $40 per hour

     ...A forward-thinking analytics company is seeking quantitative professionals to evaluate AI-generated work, ensuring accuracy in statistical analysis and predictive modeling. This fully remote role offers a flexible schedule and competitive pay starting at $40 per hour.... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Saint Paul, MN
    2 days ago
  • $40 per hour

     ...A tech-focused company is seeking experienced quantitative professionals to evaluate AI-generated analysis and provide technical feedback. This remote role allows you to choose your schedule and work from various countries including the US. Ideal candidates will have a... 
    Hourly pay
    Remote work

    DataAnnotation

    Providence, RI
    5 days ago
  • $50 - $60 per hour

     ...DataAnnotation is seeking an Appellate Attorney to train AI models by evaluating their legal outputs and solving complex legal challenges. This role allows for flexible remote work, enabling you to choose projects based on your own schedule and preferences. A J.D. is mandatory... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Raleigh, NC
    2 days ago
  •  ...Opportunity Deepgram is looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team responsible for validating the...  ...production before they reach customers. Partner with DevOps/Infra to stand up ephemeral test environments and results... 
    Full time

    Deepgram

    Remote
    18 days ago
  • $40 per hour

     ...company in the United States is seeking an R&D Biologist to train AI models and improve their quality. This position offers remote...  ...selected projects at your own schedule. Responsibilities include evaluating the performance of AI chatbots on complex biology topics. Candidates... 
    Hourly pay
    Remote work

    DataAnnotation

    Brooklyn, NY
    3 days ago
  • $30 per hour

    A technology company is seeking a Web Platform Engineer to evaluate AI chatbots and enhance model performance. This role requires proficiency in programming languages like Python and JavaScript. You will assess AI outputs from coding challenges and writing tasks, ensuring... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Jackson, MS
    2 days ago
  • $40 per hour

     ...Development Chemist to join our team to train AI models. You will measure the progress of these AI chatbots, evaluate their logic, and solve problems to improve the...  ...are not limited to: Chemistry and/or Chemical Engineering. Benefits This is a full-time or part-time... 
    Hourly pay
    Full time
    Contract work
    Part time
    Remote work

    DataAnnotation

    Iowa, LA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to DevOps Engineer for AI Model Evaluation. Be the first to apply!