Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

DevOps Engineer - AI Model Evaluator

$85 per hour

Mercor

Job Description

Job Description

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .

Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience)
Type: Contract
Compensation: $85/hour
Location: Remote

Role Responsibilities

  • Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks.
  • Review model-generated implementations involving cloud platforms , Kubernetes , CI/CD systems , observability , and infrastructure automation .
  • Identify bugs, edge cases, reliability issues, and failure modes in model outputs.
  • Compare outputs from multiple frontier models to assess their strengths and weaknesses.
  • Apply professional engineering judgment to realistic infrastructure engineering scenarios.

Qualifications

Must-Have

  • 2+ years of professional DevOps , SRE , or Cloud Engineering experience.
  • Experience with AWS , Azure , GCP , Kubernetes , Terraform , CI/CD pipelines , or observability tooling.
  • Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
  • Ability to evaluate model-generated infrastructure and reliability engineering solutions.

Preferred

  • Experience supporting production-scale systems.

Compensation & Legal

  • $400 per accepted task
  • Compensation tied to accepted work.

Application Process (Takes 20–30 mins to complete)

  • Upload resume
  • AI interview based on your resume
  • Submit form

Resources & Support

PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the DevOps Engineer - AI Model Evaluator in San Francisco, CA vacancy
  • $85 per hour

     ...technical talent with leading AI research labs. Headquartered...  ...Dorsey . Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience)...  ...coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving... 
    Suggested
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    1 day ago
  • $50 - $75 per hour

    A leading tech company based in Australia is seeking an AI Model Evaluator on a contract basis. The role involves evaluating AI-generated responses, writing prompts, and providing justifications based on specific criteria. Ideal candidates will hold a Master's degree in... 
    Suggested
    Hourly pay
    Contract work

    Mercor

    San Francisco, CA
    4 days ago
  • $218.5k - $288k

     ...Scientist specializing in Small Language Models and AI Training, you will lead research and...  .... You will work closely with research, engineering, and product teams to advance model training...  ...language models.Design, implement, and evaluate model training experiments to improve... 
    Suggested
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    2 days ago
  • $180k - $225k

    As a Software Engineer on the ML Infrastructure team, you will design...  ...engineers to integrate and optimize models for production and research...  ...to ensure a fair and thorough evaluation of all applicants.About Us:At...  ...is to develop reliable AI systems for the world's most important... 
    Suggested
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  •  ...Description Job Description We are looking for a Senior Software Engineer, Infrastructure with ideally 5+ years of experience to own and...  ...build, release, and runtime foundations at a fast-scaling voice AI company. You'll design and automate deployment pipelines across... 
    Suggested

    NxT Level

    San Francisco, CA
    more than 2 months ago
  • $150k - $240k

     ...DevOps Engineer Title of Role: DevOps Engineer Location: San Francisco, on-site or remote Company Stage of Funding: Series B Office...  ...software development space, particularly within API SDKs and AI-driven solutions. With a strong commitment to innovation and... 
    Work at office
    Remote work

    Recruiting from Scratch

    San Francisco, CA
    4 days ago
  • $150k - $200k

     ...Job Description Job Description About the Role Join a fast-moving, venture-backed AI messaging infrastructure startup as a Founding Engineer on the macOS DevOps team . This is a mid-level (Member of Technical Staff) position based on-site in San Francisco, CA... 
    Remote work

    Clera

    San Francisco, CA
    1 day ago
  • $130k - $196.5k

     ...service environment creation for engineering teams.Lead and collaborate...  ...quality.Mentor engineers on DevOps best practices, helping teams...  ...infrastructure technologies by evaluating options, building proofs of...  ...processing tools, analytics, and AI/ML workloads across LiveRamp... 
    Full time
    Work from home
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    1 day ago
  • $176.6k - $239k

     ...Solutions Architect, you will partner some of the world’s leading AI model providers companies, AWS Sales, and several other AWS teams to...  ...areas (e.g. software development, cloud computing, systems engineering, infrastructure, security, networking, data & analytics)... 
    Local area
    Worldwide
    Flexible hours

    AmazonWebServices

    San Francisco, CA
    1 day ago
  •  ...or Kubernetes-based environments. Experience building AI-powered solutions, MCP Servers, Agentic AI systems, or...  ...skills. Preferred Qualifications Experience with DevOps and Site Reliability Engineering (SRE) practices. Strong production support, incident... 

    Eitacies Inc

    San Francisco, CA
    a month ago
  •  ...critical inference for the world's most dynamic AI companies, like Cursor, Notion,...  ...the frontier of AI to bring cutting-edge models into production. We're growing quickly and...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    22 hours ago
  •  ...We are rebuilding biotech for the AI era. When a breakthrough is delayed, the world...  ...structured data, and run AI agents and models directly in their workflows. Over 200,000...  ...-saving therapeutics. As a full-stack engineer on the team, you’ll focus on building the... 
    Full time
    Work at office
    Local area
    Monday to Friday
    Shift work

    Benchling

    San Francisco, CA
    22 hours ago
  •  ...enterprise and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to...  ...via model inference.  About the Role We are looking for an engineer who wants to take the world's largest and most capable AI... 
    Full time

    OpenAI

    San Francisco, CA
    22 hours ago
  •  ...About the Team We’re hiring software engineers to make OpenAI’s Model Performance teams more productive. These teams work on the systems, tooling,...  ...with forward progress About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that... 
    Full time

    OpenAI

    San Francisco, CA
    22 hours ago
  •  ...is taking the hard out of hardware, by developing the first AI Hardware Engineer. Our goal is to democratize the ability to create bleeding edge...  ...that means everything around it has to just work. As a DevOps Engineer, you'll work on the full-stack systems that power... 
    Local area
    Shift work

    Flux Defunct

    San Francisco, CA
    1 day ago
  •  ...Summary: This position is responsible for: Build DevOps solutions across applications, network, deployment & scaling to...  ...~ Experience scaling different geographies. High level of AI usage in Day-to-Day activity. (Good to have) Experience... 
    Work experience placement
    Work at office

    Texas State Library and Archives Commision

    San Francisco, CA
    3 days ago
  • Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should... 
    Weekday work

    Mercor Inc

    San Francisco, CA
    22 hours ago
  • A pioneering technology company in San Francisco is seeking a DevOps Engineer to join the founding team. This role requires building and maintaining scalable cloud infrastructure, own CI/CD systems, and implement observability tooling. Ideal candidates have 4-10+ years... 

    Fabrion

    San Francisco, CA
    2 days ago
  •  ...DevOps Engineer Encord is the universal data layer for AI that helps 300+ AI teams train and run models on the right data. Our platform indexes, curates, annotates, and evaluates data across the full AI lifecycle, from development through production. We're looking... 
    Work at office
    Flexible hours

    Encord

    San Francisco, CA
    1 day ago
  • Vapi Inc. in San Francisco is seeking a dedicated engineer to enhance their deploy pipeline and implement progressive delivery systems. The...  ...on customer satisfaction for internal engineers. Join a dynamic environment driving innovation in voice AI! #J-18808-Ljbffr VAPI

    VAPI

    San Francisco, CA
    1 day ago
  •  ...investors, we're building the category-defining AI workflow automation platform that...  ...Role Plenful is hiring a Senior DevOps Engineer to join our engineering team and help build...  ...and New York. R&D roles follow a hybrid model, with two days per week in our San... 
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    2 days ago
  • Obsidian is hiring expert Evaluators in Investment analysis / valuation / credit to review AI-generated work products for accuracy and quality. This remote, hourly position requires deep subject-matter expertise and professional fluency in English to provide structured... 
    Hourly pay
    Work at office
    Remote work

    Obsidian

    San Francisco, CA
    2 days ago
  • Obsidian is seeking a Spanish Audio Generalist Evaluator Expert to contribute to a high-impact audio AI research project. You will handle transcription, annotation...  ...to help train and benchmark advanced language models. The ideal candidate should have strong writing skills... 
    Part time
    10 hours per week

    Obsidian

    San Francisco, CA
    3 days ago
  • Welo Data is seeking Data Labeling Associates in California to evaluate AI outputs and ensure cultural context and safety in Arabic datasets. This role requires professional-level proficiency in Portuguese (Brazil), a bachelor's degree, and at least 2 years of experience... 

    Welo Data

    San Francisco, CA
    4 days ago
  • $172.43k - $230.95k

     ...intelligence. As the only vertically integrated AI infrastructure company built from the...  ....About This Role:The Senior Software Engineer for the AI Model Lifecycle team will play a crucial...  ...management: versioning, lineage, evaluation, and reproducible fine-tuning at scale... 
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  • A leading data and AI company in San Francisco is seeking a Staff Engineer to design and implement systems for their AI/ML Model Serving platform. You will collaborate with product, infrastructure, and research teams to ensure high-performance system delivery. The ideal... 

    Jobleads-US

    San Francisco, CA
    2 days ago
  • Obsidian is hiring expert Evaluators in Real estate, hospitality, and events to review AI-generated work for accuracy, rigor, and domain quality. This remote position requires deep expertise and involves grading outputs like documents and presentations. Applicants must... 
    Remote job
    Work at office

    Obsidian

    San Francisco, CA
    2 days ago
  • Synthires is seeking a PhD-level expert to contribute to advanced AI research and evaluation projects in San Francisco. The role centers on applying deep domain knowledge to design problems, evaluate AI responses, and craft high-quality reference solutions that push the... 
    Part time

    Synthires

    San Francisco, CA
    3 days ago
  •  ...About the job Staff DevOps Engineer At Cube, we're redefining how organizations deliver...  ...and analytics across teams, tools, and AI agents. Our mission is to enable Agentic...  ...While Cube Cloud is our primary delivery model, larger enterprise customers can run Cube... 
    Remote work

    Cube Dev, Inc

    San Francisco, CA
    4 days ago
  • Obsidian is looking for expert Evaluators in Finance operations/audit support to review AI-generated work products for accuracy and quality. This remote hourly position requires a minimum of 5 years in finance and fluency in English. Your role will involve evaluating outputs... 
    Remote job
    Hourly pay
    Work at office

    Obsidian

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to DevOps Engineer - AI Model Evaluator. Be the first to apply!