Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer for AI Model Evaluation

$50 - $75 per hour

SaidGig

Contribute to the advancement of AI models by evaluating AI-generated solutions in a cutting-edge project focused on improving how language models understand, reason about, and generate production-quality software. This role offers a unique opportunity to leverage your software engineering expertise to shape the future of AI-assisted software development. Key Responsibilities

  • Become familiar with assigned open-source codebases and their engineering practices.
  • Review code changes, technical discussions, and related software artifacts.
  • Evaluate AI-generated code for correctness, quality, maintainability, and engineering best practices.
  • Provide structured evaluations with clear technical reasoning.
  • Collaborate with fellow technical reviewers to ensure high-quality and consistent evaluations.
Qualifications
  • Years of Experience: 3+ years of professional software engineering experience or equivalent open-source experience.
  • Technical Expertise: Proficiency in JavaScript/TypeScript (Node.js, React, TypeScript), Go (backend services, distributed systems, REST/gRPC APIs, cloud-native development), and Python (backend development, open-source libraries, DevOps/infrastructure automation, testing frameworks).
  • Software Engineering: Experience with large production codebases, ability to quickly understand unfamiliar repositories and software architectures, experience reviewing code for readability and maintainability, strong engineering judgment, and understanding of software engineering best practices.
  • GitHub & Open Source: Strong understanding of Git and GitHub workflows (branching, Issues, Pull Requests, Code Reviews), experience navigating open-source repositories, familiarity with open-source development workflows, and familiarity with SWE-Bench or SWE-Bench Pro is a plus.
  • Communication: Strong English communication skills.
Work Terms
  • Employment Type: Hourly
  • Expected Commitment: At least 10 hours
  • Project Duration: Approximately 1 week
  • Expected Timeline: July 8, 14
  • Location: Candidates must be located in the US/Canada for real-time collaboration with the project team.
Compensation
  • Hourly Rate: $50 - $75
Eligibility
  • This is not a software development role; the focus is on evaluating AI-generated solutions.
Vacancy posted 28 days ago
Similar jobs that could be interesting for youBased on the Software Engineer for AI Model Evaluation in Canada vacancy
  • $70 - $126 per hour

     ...technical domain expertise to advance agentic AI workflows by working alongside...  ...plan, implement, debug, and refine real software engineering tasks. Your contributions will help train...  ...behind decisions and changes. Evaluate AI-generated code for correctness, hallucinations... 
    Suggested
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    3 days ago
  • $40 - $65 per hour

     ...Overview Contribute domain expertise to evaluate and harden frontier large language models by crafting adversarial multi-turn...  ...on improving how next-generation AI systems learn, reason, and behave,...  ..., annotation, or prompt engineering. Preferred: deep familiarity with... 
    Suggested
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    3 days ago
  • $50 per hour

     ...Overview This role develops high-quality datasets and evaluation workflows to train and benchmark large language models on software engineering tasks. You will curate and author code examples, correct and improve AI-generated solutions across multiple languages, design... 
    Suggested
    For contractors
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Canada
    28 days ago
  • $70 per hour

     ...This role focuses on evaluating AI models in the Underwriting domain through a benchmark dataset project. Experts will create complex, grounded tasks that include a clear ground-truth output and an objective rubric, contributing to the advancement of visual document understanding... 
    Suggested
    Hourly pay
    Remote work

    SaidGig

    Canada
    a month ago
  • $75 per hour

     ...Role Overview Author complex, grounded evaluation tasks and objective scoring rubrics that measure AI models on visual document understanding and instruction-following within the Electrical and Electronics domain. Your work will feed a benchmark dataset used to evaluate... 
    Suggested
    Hourly pay
    Remote work

    SaidGig

    Canada
    23 days ago
  • $75 per hour

     ...Join a project focused on evaluating AI models in the architecture domain, specifically in visual document understanding and instruction-following. This role involves authoring complex, grounded tasks that include a clear ground-truth output and an objective rubric.... 
    Hourly pay
    Remote work

    SaidGig

    Canada
    21 days ago
  • $60 per hour

     ...This role focuses on evaluating AI models for visual document understanding and instruction-following within the Marketing Analytics domain. You will work on a benchmark dataset project that involves creating and assessing complex tasks related to campaign dashboards,... 
    Hourly pay
    Remote work

    SaidGig

    Canada
    a month ago
  • $100 per hour

     ...Join a project focused on evaluating AI models for visual document understanding and instruction-following specifically within the Surveying & GIS domain. In this role, you will author complex, grounded tasks that include clear ground-truth outputs and objective rubrics... 
    Hourly pay
    Remote work

    SaidGig

    Canada
    25 days ago
  • $90 per hour

     ...This role involves contributing to a benchmark dataset project that evaluates AI models focused on visual document understanding and instruction-following specifically within the Telecom and Network domain. As an expert, you will be responsible for authoring complex, grounded... 
    Hourly pay
    Remote work

    SaidGig

    Canada
    26 days ago
  • $75 per hour

     ...This role involves contributing to a benchmark dataset project that evaluates AI models focused on visual document understanding and instruction-following specifically within the Energy and Utilities domain. As an expert, you will author complex, grounded tasks that include... 
    Hourly pay
    Remote work

    SaidGig

    Canada
    24 days ago
  • $70 per hour

     ...This role focuses on evaluating AI models for visual document understanding and instruction-following specifically within the Pharma and Clinical Trials domain. You will contribute to a benchmark dataset project by authoring complex, grounded tasks that include a clear... 
    Hourly pay
    Remote work

    SaidGig

    Canada
    a month ago
  • $60 per hour

     ...This role focuses on evaluating AI models for visual document understanding and instruction-following within the Logistics & Supply Chain domain. You will contribute to a benchmark dataset project by authoring complex, grounded tasks that include clear ground-truth outputs... 
    Hourly pay
    Remote work

    SaidGig

    Canada
    27 days ago
  • $20 - $40 per hour

     ...that will be used to train next-generation AI systems. This role contributes...  ...world legal reasoning, helping improve how models learn and perform. No prior AI experience...  ...contributions will be used for training and evaluating AI systems. Materials must be original... 
    Remote job
    Hourly pay
    Contract work
    For contractors

    SaidGig

    Canada
    3 days ago
  • $80 per hour

     ...Role Overview Design and author grounded evaluation tasks and objective rubrics that test AI models on visual document understanding and instruction-following...  ...Subject-matter expertise in Mechanical engineering and CAD, demonstrated by work or domain knowledge... 
    Hourly pay
    Remote work

    SaidGig

    Canada
    17 days ago
  • $75 per hour

     ...This role focuses on evaluating AI models in the Construction Estimating domain through a benchmark dataset project. Experts will be responsible for authoring complex, grounded tasks that include a clear ground-truth output and an objective rubric. Key Responsibilities... 
    Hourly pay
    Remote work

    SaidGig

    Canada
    25 days ago
  • $20 - $60 per hour

     ...realistic business scenarios, create and evaluate Office Open XML documents, and provide detailed...  ...helps train and improve large language models. This contractor role focuses on...  ...complex document-related prompts and assessing AI outputs. Create, open, and review... 
    Remote job
    Hourly pay
    For contractors
    Work at office
    Free visa

    SaidGig

    Canada
    3 days ago
  • $80 - $160 per hour

     ...Role Overview Model stochastic bacterial population dynamics to derive asymptotic growth rates, analyze effects of growth-rate switching...  ...clear, reproducible methodological documentation to train and evaluate AI systems. Key Responsibilities Analyze and model... 
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    1 day ago
  • $60 - $85 per hour

     ..., and optimize GPU-based software and kernels used to train large language models, with a strong focus on runtime...  ...for LLM training and evaluation. About the company micro1 is an AI data lab that produces...  ...finance, healthcare, and STEM engineering, contribute real-world... 
    Remote job
    Hourly pay
    For contractors
    Immediate start

    SaidGig

    Canada
    2 days ago
  • $80 - $160 per hour

     ...assignments. Your contributions will be used to train and evaluate next generation AI systems, providing high quality, domain-grounded data and judgments...  ..., and producing annotated explanations suitable for model training. Compensation ~ Pay range, $80.00 to $160.00... 
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    1 day ago
  • $80 - $160 per hour

     ...many-body systems, with a focus on problems such as the PXP model, Rydberg blockade, quantum many-body scars, and constrained...  ...analyses that support research-level benchmarks and help train and evaluate advanced AI systems. Key Responsibilities Implement and analyze... 
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    1 day ago
  • $60 per hour

     ...Role Overview Develop and author evaluation tasks for a benchmark dataset that measures AI models on visual document understanding and instruction following within the Real Estate Appraisal domain. Tasks must be complex, grounded in real appraisal practice, and paired... 
    Hourly pay
    Remote work

    SaidGig

    Canada
    24 days ago
  • $40 - $50 per hour

     ...quality inputs for next-generation AI systems. You will analyze and...  ...and troubleshooting scenarios so models learn from accurate, practical engineering judgment. No prior AI experience is...  ...reports, and operational procedures. Evaluate electrical troubleshooting... 
    Remote job
    Hourly pay
    For contractors
    Local area

    SaidGig

    Canada
    3 days ago
  • $170 - $190 per hour

     ...-person role in San Francisco supports a healthcare AI partner by ensuring annotated data and model outputs meet clinical standards. Key Responsibilities...  ...model interpretation of medical context. Model evaluation and feedback: assess AI recommendations or clinical... 
    Hourly pay

    SaidGig

    Canada
    3 days ago
  •  ...LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you...  ...technical foundations that power model improvement for foundation model builders...  ...across model comparisons Software / Data Engineering Proficiency... 
    Full time

    Innodata

    Canada
    10 days ago
  • $50 per hour

     ...Role Overview Author and deliver benchmark tasks that evaluate AI models on visual document understanding and instruction-following within the front-end engineering domain. Work will focus on UI screenshots, design specifications, and component diagrams. This opening... 
    Hourly pay
    Remote work

    SaidGig

    Canada
    17 days ago
  •  ...BeyondTrust is seeking a Software Engineer (Backend) for the team that develops our Password Safe...  ...Demonstrated experience using agentic AI as a fundamental tool integrated into daily...  ..., and performant code Propose and evaluate technical solutions as part of research... 
    Full time
    Remote work

    Beyondtrust

    Canada
    6 days ago
  • $100 - $150 per hour

     ...Role Overview Provide experienced litigation expertise to evaluate, refine, and improve AI-generated legal analysis and written advocacy. In this...  ...expert judgment on ambiguous or complex edge cases to inform model calibration on nuanced legal issues. Draft or edit... 
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    2 days ago
  •  ...Description We're seeking a talented and motivated full-time Software Engineer, AI Enablement to help Tailscale's engineering organization...  ..., Codex , OpenCode , and Pi more effectively. Evaluate new AI models, agents, and tools, and make clear recommendations on... 
    Full time

    Tailscale

    Canada
    9 days ago
  • $20 - $60 per hour

     ...document expertise to a project that trains next-generation AI systems. You will design realistic Fortune 500 style scenarios and interact iteratively with an advanced language model to create, edit, and evaluate Office Open XML files, with a focus on .pptx deliverables.... 
    Remote job
    Hourly pay
    Contract work
    For contractors
    Work at office

    SaidGig

    Canada
    3 days ago
  • $50 - $90 per hour

     ...validate mental-health safety and crisis-prevention frameworks used to train next-generation AI systems. This contractor role focuses on shaping how digital tools detect, evaluate, and respond to self-harm, eating-disorder risk, emotional dependency, and suicide-related... 
    Remote job
    Hourly pay
    For contractors

    SaidGig

    Canada
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer for AI Model Evaluation. Be the first to apply!