Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Go Developer for AI Model Evaluation

$30 - $90 per hour

SaidGig

Role Overview

Build and maintain backend services in Go while evaluating and training alpha-stage AI coding tools. This contract role combines hands-on Go development, backend engineering, and structured testing of AI-assisted coding workflows to improve next-generation developer tools.

Key Responsibilities
  • Develop REST and GraphQL endpoints within a modern Go stack.
  • Implement rigorous data validation and comprehensive error handling across services.
  • Perform database migrations and apply performance optimizations.
  • Test and evaluate alpha AI models using Cursor, conducted over multiple 4-day testing bursts with 5+ hour sessions per day.
  • Identify, document, and report incidents, bugs, and edge cases with detailed bug traces and annotated screenshots.
  • Collaborate with the research team via Slack, providing real-time feedback and thoughtful technical insights.
  • Complete structured surveys after each testing burst to capture overall impressions and model usability.
Qualifications
  • Required : Minimum 5 years of professional experience as a Golang developer, strong command of Go, REST APIs, GraphQL, and backend development best practices.
  • Proven experience with AI-powered coding tools, ideally experience using Cursor.
  • Experience in data validation, robust error handling, and complex debugging.
  • Exceptional written and verbal communication skills, with the ability to document findings clearly without disclosing confidential information.
  • Enthusiasm for hands-on coding, rapid prototyping, and technical exploration.
  • Preferred : Public open-source contributions such as notable GitHub repositories, prior experience in technical product evaluations or developer tool QA, and interest in mentoring or sharing insights with research teams.
Work Terms
  • Contract role, fully remote.
  • Part-time engagement organized around multiple testing bursts, each burst lasting 4 days with 5 or more hours of testing per day.
  • Work involves highly confidential, alpha-stage model evaluation, requiring strict protection of sensitive information.
  • Collaboration and feedback will occur primarily via Slack and structured survey forms after testing bursts.
Compensation

Hourly rate: $30 to $90 per hour.

Eligibility
  • This position is offered as a contract engagement, candidates must be able to work as independent contractors.
  • Candidates must be able to work remotely and commit to the scheduled testing bursts and confidentiality requirements.
Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Go Developer for AI Model Evaluation in United States vacancy
  • $30 - $90 per hour

     ...Role Overview Work as a remote Go developer contributing to the evaluation and training of next-generation AI coding tools in confidential alpha stages. This part-time,...  ...performance. Test and evaluate alpha AI coding models in Cursor, running focused testing sessions.... 
    Suggested
    Hourly pay
    Contract work
    Part time
    Remote work

    SaidGig

    United States
    4 days ago
  • $100 - $130 per hour

     ...to help train next-generation AI systems by creating realistic...  ...facing project that shapes how models learn, reason, and perform. No...  ...into high-quality training data, evaluations, and feedback for frontier...  ...and implement code solutions in Go, Python, and TypeScript to... 
    Suggested
    Hourly pay
    For contractors
    Remote work

    SaidGig

    Indiana
    14 days ago
  •  ...Opportunity Deepgram is looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team responsible for validating the...  ...experience in a language such as Python , Rust , Go , or similar. ~ Experience designing and building automated... 
    Suggested
    Full time

    Deepgram

    Remote
    18 days ago
  • $105 per hour

     ...engineering skills to help refine and evaluate Large Language Models for high-stakes business communications...  ...intersection of web development and AI research, supporting model evaluation...  ...for high-stakes business contexts. Develop domain-specific prompts and assess LLM... 
    Suggested
    Part time
    Work experience placement
    Remote work
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  • $100 per hour

     ...software engineering expertise to shape next-generation AI systems by reviewing, refining, and evaluating AI-generated technical content. In this remote, part-...  ...rubric-based assessments that directly influence how models learn, reason, and produce engineering deliverables.... 
    Suggested
    Hourly pay
    Part time
    For contractors
    Remote work

    SaidGig

    Remote
    9 days ago
  •  ...Lead the design and evaluation of next-generation coding agents by creating...  ...and improvement of coding models. Key Responsibilities Design...  ...engineering tasks. Develop high-quality datasets, golden...  ...researchers, engineers, and applied AI teams to design experiments and... 
    Full time
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $30 - $90 per hour

     ...performance backend APIs in Rust to support experimental developer tools and cutting-edge AI research. You will implement scalable REST and...  ...collaborate closely with research engineers, and evaluate AI-powered coding models to improve developer workflows. You will join... 
    Hourly pay
    Contract work
    For contractors
    Remote work
    Free visa

    SaidGig

    United States
    a month ago
  • $40 per hour

     ...connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems....  ...evaluate AI coding agents - how well a model handles real-world developer tasks. ~Creating challenging tasks and... 
    Permanent employment
    Temporary work
    Part time

    Mindrift

    Remote
    5 days ago
  • $85 per hour

     ...Role Overview Evaluate and improve frontier AI coding models by using AI coding agents to complete and assess realistic infrastructure engineering workflows. Work will focus on reviewing model-generated infrastructure solutions across cloud platforms, container orchestration... 
    Hourly pay
    Remote work

    SaidGig

    United States
    16 days ago
  • $100 - $130 per hour

     ...documentation that train next-generation AI systems. This contract role centers on using Go as the primary language, with...  ...and technical feedback that help models learn how software is designed,...  ...to support AI training data and evaluations. Review, analyze, and provide... 
    Remote job
    Hourly pay
    Contract work
    For contractors

    SaidGig

    Remote
    13 days ago
  • $85 per hour

     ...connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our...  ...~Use frontier AI coding agents to complete and evaluate complex engineering tasks. ~Review model-generated mobile application code for correctness, quality... 
    Contract work
    Part time
    Summer work
    Remote work

    Mercor

    Remote
    a month ago
  •  ...We are hiring a Senior Solutions Engineer to help shape and scale our AI Data Solutions, working with leading AI labs, frontier model developers, and enterprise AI teams on complex data, evaluation, and model development workflows. This is a senior, customer-facing role... 
    Full time
    For contractors

    L10n People Ltd

    Remote
    7 days ago
  • $40 per hour

     ...We are looking for experienced cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content, solve technical cybersecurity problems, and provide feedback to improve how AI systems reason about real... 
    Hourly pay
    Full time
    Part time
    Remote work

    DataAnnotation

    New York, NY
    2 days ago
  • $400 per month

     ...About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic infrastructure engineering... 

    Mercor Inc

    Doral, FL
    2 days ago
  • $40 per hour

    A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve technical problems in a flexible remote role. Ideal candidates should have over 2 years in cybersecurity, coding experience, strong analytical and writing... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Sioux Falls, SD
    5 days ago
  • $238k - $302k

     ...in simulation across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in Large Language Models...  ...deployed in the Waymo Driver. You will: Develop novel metrics and sampling techniques to measure the... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  •  ...Opportunity Deepgram is looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team responsible for validating the...  ...experience in a language such as Python , Rust , Go , or similar. ~ Experience designing and building automated... 
    Full time

    Deepgram

    Remote
    8 days ago
  • $30 per hour

    A technology company is seeking a Web Platform Engineer to evaluate AI chatbots and enhance model performance. This role requires proficiency in programming languages like Python and JavaScript. You will assess AI outputs from coding challenges and writing tasks, ensuring... 
    Hourly pay
    Remote work
    Flexible hours

    DataAnnotation

    Jackson, MS
    2 days ago
  • $45 - $65 per hour

     ...live Codeforces contest as part of the application process for AI-data work. This role focuses on competitive programming problem...  ...solutions under time constraints while collaborating remotely with the evaluation team. Key Responsibilities Analyze competitive programming... 
    Hourly pay
    Remote work

    SaidGig

    United States
    17 days ago
  •  ...leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated cybersecurity content and solve technical security problems. You will play a significant role in training AI models, providing critical feedback, and improving system accuracy. This... 
    Remote job
    Flexible hours

    DataAnnotation

    New York, NY
    3 days ago
  • $40 per hour

    A leading cybersecurity solutions provider is seeking experienced cybersecurity professionals for a remote position. You will evaluate AI-generated security content, solve technical problems, and provide essential feedback to improve AI systems. The ideal candidate will... 
    Remote job
    Hourly pay

    DataAnnotation

    Helena, MT
    5 days ago
  • $40 per hour

    A cybersecurity company is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. This remote position offers the flexibility to choose projects and work on your own schedule, with projects starting at $40 per hour. Candidates... 
    Remote job
    Hourly pay

    DataAnnotation

    Columbia, SC
    3 days ago
  • A leading cybersecurity firm is seeking experienced cybersecurity professionals for a remote role to help train AI models. Candidates will evaluate AI-generated security content, solve technical cybersecurity problems, and provide valuable feedback for the improvement of... 
    Remote job
    Flexible hours

    DataAnnotation

    Santa Fe, NM
    4 days ago
  • $40 per hour

    A leading AI security solutions provider is seeking experienced cybersecurity professionals to evaluate AI-generated security content and solve real-world technical problems. In this remote role, candidates will require over 2 years of cybersecurity experience, fluency... 
    Remote job
    Hourly pay

    DataAnnotation

    Brooklyn, NY
    3 days ago
  •  ...needless overhead of meetings. Our AI assistant captures, summarizes,...  ...ROLE OVERVIEW We're hiring a Model Performance Engineer to own the...  ...so an AI Engineer can go more quickly from dataset to deployed...  ...that gets 1.3x speedup with % Evaluate serving frameworks (vLLM vs... 
    Full time
    Remote work

    Fathom

    Remote
    1 day ago
  • $40 per hour

    A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve technical cybersecurity problems. You will enhance how AI systems handle real-world threats while working remotely on an hourly project basis starting at... 
    Remote job
    Hourly pay

    DataAnnotation

    California, MO
    3 days ago
  • $50 - $150 per hour

     ...reference solutions that test AI systems on complex software engineering...  ...help train next-generation models by supplying high-quality, real...  ...Learning environments that evaluate an AI model''s ability to solve...  ...languages: Python3, Java, Rust, Go, C++, or TypeScript. Analyze... 
    Hourly pay
    For contractors
    Immediate start
    Remote work

    SaidGig

    Indiana
    15 days ago
  • $220k - $320k

     ...trains and hosts specialized language models for companies that need frontier-quality AI at a fraction of the cost. The...  ...-to-end: distillation, training, evaluation, and planet-scale hosting. We...  ...our performance standards before going to production Experiment with... 
    Work at office

    Inference

    San Francisco, CA
    3 days ago
  • $180k - $225k

     ...integrate and optimize models for production and research...  ...design and scalability.Develop monitoring and...  ...languages (e.g., Python, Go, Rust, C++).Experience...  ...ensure a fair and thorough evaluation of all applicants.About...  ...is to develop reliable AI systems for the world's... 
    Full time

    Scale AI

    San Francisco, CA
    5 days ago
  •  ...We are rebuilding biotech for the AI era. When a breakthrough is delayed, the world...  ...structured data, and run AI agents and models directly in their workflows. Over 200,00...  ..., figuring out new patterns for how we develop and go to market. We’ll win if we stay curious... 
    Full time
    Work at office
    Local area
    Monday to Friday
    Shift work

    Benchling

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Go Developer for AI Model Evaluation. Be the first to apply!