Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

DevOps Engineer for AI Model Evaluation

$85 per hour

SaidGig

Role Overview

This role supports a Frontier Code Agents initiative at a leading AI research lab, focused on evaluating and improving advanced AI coding models by performing structured technical assessments of realistic infrastructure engineering workflows and model outputs.

Key Responsibilities
  • Use frontier AI coding agents to perform and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations that involve cloud platforms, Kubernetes, CI/CD systems, observability, and infrastructure automation.
  • Identify bugs, edge cases, reliability problems, and likely failure modes in model outputs.
  • Compare outputs from multiple frontier models, assessing relative strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure and reliability engineering scenarios.
Qualifications
  • Minimum 2 years of professional experience in DevOps, SRE, or Cloud Engineering.
  • Hands-on experience with one or more cloud platforms, such as AWS, Azure, or GCP.
  • Familiarity with Kubernetes, Terraform, CI/CD pipelines, and observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.
  • Experience supporting production-scale systems is preferred.
Work Terms
  • Remote role.
  • Hourly engagement, sprint based, with project sprints running in 12 to 24 hour stretches depending on client requirements.
  • Spots are limited and are filled on a first come, first serve basis.
Compensation
  • $400 paid per accepted task.
  • Typical tasks take approximately 2 to 3 hours of work after ramp-up.
  • Compensation is tied to accepted work.
  • Metadata indicates an hourly reference rate of $85 per hour.
Eligibility
  • Role is open to remote applicants. No specific work authorization or sponsorship information was provided in the source; applicants should confirm they are eligible to work in their location.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the DevOps Engineer for AI Model Evaluation in United States vacancy
  • $85 per hour

     ...technical talent with leading AI research labs. Headquartered...  ...Dorsey . Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience)...  ...coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving... 
    Suggested
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    1 day ago
  • $85 per hour

     ...technical talent with leading AI research labs. Headquartered...  ...Dorsey . Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience)...  ...coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving... 
    Suggested
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    16 days ago
  • $70 - $150 per hour

     ...Role Overview Contribute DevOps and platform engineering expertise to AI research and product teams through hands-on work training and evaluating models, designing and maintaining infrastructure, and producing real-world tasks and deliverables. This is an open application... 
    Suggested
    Hourly pay
    Contract work
    For contractors
    Remote work

    SaidGig

    United States
    10 days ago
  •  ...Role Overview Medical professionals evaluate AI-generated medical content and use their clinical and field experience to improve model outputs. No prior AI experience is required. In this role you will assess model responses, create realistic prompts that reflect clinical... 
    Suggested
    Hourly pay
    Part time
    Immediate start
    Remote work
    Flexible hours

    SaidGig

    United States
    6 days ago
  • $40 - $65 per hour

     ...Overview Contribute domain expertise to evaluate and harden frontier large language models by crafting adversarial multi-turn...  ...on improving how next-generation AI systems learn, reason, and behave,...  ..., annotation, or prompt engineering. Preferred: deep familiarity with... 
    Suggested
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Frontier County, NE
    1 day ago
  • $39 per hour

     ...Music Audio Expert, where you will play a crucial role in evaluating generative musical AI models in collaboration with a leading AI lab. This position...  ...quality, and production/mix quality, using audio-engineering terminology. Annotate songs in detail, including genre... 
    Hourly pay
    Part time
    Immediate start
    10 hours per week

    SaidGig

    United States
    25 days ago
  • $60 - $80 per hour

     ...foundational Large Language Models, by designing realistic marketing...  ...accurate solutions, and evaluating model outputs with rigorous,...  ...places you within a leading AI lab''s extended workforce while...  ...Responsibilities Guide research and engineering teams to close knowledge gaps... 
    Hourly pay
    Weekday work

    SaidGig

    United States
    6 days ago
  • $30 - $90 per hour

     ...Role Overview Build and maintain backend services in Go while testing and evaluating alpha-stage AI coding models. This part-time, remote contract role combines hands-on Go development with structured model evaluation, bug reporting, and real-time collaboration to improve... 
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    2 days ago
  • $80 - $110 per hour

     ...Role Overview Work directly with a frontier Generative AI research team to design difficult, real-world machine learning and natural language tasks, produce reference solutions, and evaluate model outputs to reveal reasoning and capability gaps. You will create executable... 
    Hourly pay
    Part time
    Freelance
    Remote work

    SaidGig

    United States
    3 days ago
  • $75 per hour

     ...Role Overview Lead the clinical evaluation of medical AI by designing and executing assessments that probe clinical reasoning and decision-making...  ...teams to create realistic clinical problems, find where models fail or lack medical knowledge, and shape improvements so AI... 
    Contract work
    Remote work
    Flexible hours

    SaidGig

    United States
    1 day ago
  • $60 - $150 per hour

     ...Role Overview Provide legal subject-matter expertise to improve and evaluate AI systems, by designing realistic legal tasks, reviewing model outputs, and giving domain-specific feedback that advances frontier AI research. This is an open application to join a Law Expert... 
    Hourly pay
    Contract work
    Immediate start
    Remote work

    SaidGig

    United States
    4 days ago
  • $85 per hour

     ...environmental assessment, GIS analysis, and renewable energy siting to evaluate AI-generated geospatial outputs and develop expert training data...  ..., GIS Analyst, Interconnection Specialist, Solar Design Engineer, or equivalent positions in the energy industry. Hands-on... 
    Hourly pay
    Contract work
    Part time
    Work at office
    Remote work
    Flexible hours

    SaidGig

    United States
    6 days ago
  • $60 per hour

     ...contribute to developing cutting-edge AI systems, while enjoying the...  ...advance AI development. AI models are increasingly capable of...  ...-art AI models on tasks like evaluating AI-generated quantitative...  ...Computer Science, Mathematics, Engineering, or similar); a master's or... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Little Rock, AR
    4 days ago
  • $40 per hour

    A leading AI development firm in Michigan is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the evolution of AI models. Candidates should have a background in data science, statistics, or similar fields, with at least... 
    Hourly pay
    Remote work

    DataAnnotation

    Lansing, MI
    3 days ago
  • $40 per hour

     ...data annotation company is seeking professionals in quantitative fields to enhance AI development. This fully remote role allows individuals to set flexible schedules while evaluating AI-generated analyses and solving complex quantitative problems. Candidates should have... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Oregon, WI
    3 days ago
  • $40 per hour

     ...A forward-thinking AI team is seeking quantitative professionals to evaluate and improve cutting-edge AI systems. The role involves working on AI-generated analyses and providing critical feedback for model enhancement. Candidates should have a strong quantitative background... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Helena, MT
    8 hours ago
  • $40 per hour

     ...An innovative AI development company is seeking experienced quantitative professionals to contribute to AI advancements. This fully remote role involves evaluating AI-generated analyses and ensuring they are technically accurate and valid in real-world scenarios. Candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Santa Fe, NM
    8 hours ago
  • $40 per hour

    A data science team is seeking experienced quantitative professionals to evaluate AI-generated work and contribute to the development of cutting-edge AI systems. This fully remote position offers flexible scheduling and competitive hourly pay starting at $40+. Ideal candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Madison, WI
    3 days ago
  • $60 - $90 per hour

     ...Drive research-grade data analyses and create evaluation tasks that reveal where frontier generative AI models fail. You will author realistic, multi-skill analysis...  ...1 year of experience in a research, research-engineering, or intensive data-analysis role. Strong, hands... 
    Hourly pay
    Full time
    Freelance
    Remote work

    SaidGig

    United States
    7 days ago
  • $85 per hour

     ...Role Overview Evaluate and improve frontier AI coding models by completing structured technical assessments that mirror realistic machine learning engineering workflows, model training and inference systems, MLOps, and LLM application scenarios. Key Responsibilities... 
    Hourly pay
    Remote work

    SaidGig

    United States
    2 days ago
  • $80 - $110 per hour

     ...expertise at the frontier of physics to help develop and evaluate next-generation AI systems for scientific reasoning. You will apply active research...  ...physics problems used to train and evaluate advanced AI models. Evaluate model outputs on tasks including analytical... 
    Hourly pay
    Part time
    Immediate start
    Remote work

    SaidGig

    United States
    7 days ago
  • $40 per hour

    A leading AI development company is seeking experienced quantitative professionals for a remote role. Candidates will evaluate AI-generated quantitative work, solve complex problems, and provide valuable feedback. The ideal candidate has 2+ years of experience in a quantitative... 
    Hourly pay
    Full time
    Remote work
    Flexible hours

    DataAnnotation

    Topeka, KS
    3 days ago
  • $110 per hour

     ...Role Overview Join a Physician Expert Network to provide clinical expertise to AI research teams and companies. Members contribute medical knowledge to train and evaluate AI models, design realistic clinical tasks, and give domain-specific feedback that advances medical... 
    Hourly pay
    Contract work
    Remote work

    SaidGig

    United States
    10 days ago
  •  ...Physicians apply clinical judgment and frontline medical experience to evaluate AI-generated medical content, ensuring clinical accuracy, sound...  ...planning. Assess clarity, relevance, and safety of model outputs in realistic care scenarios. Provide detailed, constructive... 
    Full time
    For contractors
    Private practice
    Remote work
    Flexible hours

    SaidGig

    United States
    6 days ago
  • $40 per hour

     ...A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated analysis and solve complex technical problems. Ideal candidates will have 2+ years in quantitative roles, knowledge of statistical methods, and experience with analytical... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Oklahoma City, OK
    3 days ago
  • $40 per hour

     ...A leading AI development firm is looking for experienced quantitative professionals to evaluate AI-generated work and design problems for AI training. This fully remote position allows for a flexible schedule, offering competitive pay starting at $40+ per hour. Ideal candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Bismarck, ND
    3 days ago
  • $40 per hour

     ...A leading AI development firm is seeking experienced quantitative professionals to join their team remotely. The role involves evaluating AI-generated quantitative work, providing insights, and shaping the future of AI systems. Candidates should have over two years of... 
    Hourly pay
    Remote work

    DataAnnotation

    Indiana, PA
    3 days ago
  • $40 per hour

     ...A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative analysis and provide impactful feedback. This fully remote role allows for flexible scheduling and competitive pay starting at $40 per hour. Candidates... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Honolulu, HI
    8 hours ago
  • $40 per hour

     ...A forward-thinking AI solutions company is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the development of cutting-edge...  ...skills. Join us to directly impact the future of AI analytics and model reasoning. #J-18808-Ljbffr... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Lincoln, NE
    3 days ago
  • $40 per hour

    A leading AI development firm is seeking experienced quantitative professionals to evaluate AI-generated quantitative work and provide critical feedback. This role offers the flexibility of remote work, allowing you to set your own schedule while focusing on impactful projects... 
    Hourly pay
    Remote work

    DataAnnotation

    Wisconsin
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to DevOps Engineer for AI Model Evaluation. Be the first to apply!