Software Engineer for AI Model Evaluation
$50 - $75 per hourSaidGig
Contribute to the advancement of AI models by evaluating AI-generated solutions in a cutting-edge project focused on improving how language models understand, reason about, and generate production-quality software. This role offers a unique opportunity to leverage your software engineering expertise to shape the future of AI-assisted software development. Key Responsibilities
- Become familiar with assigned open-source codebases and their engineering practices.
- Review code changes, technical discussions, and related software artifacts.
- Evaluate AI-generated code for correctness, quality, maintainability, and engineering best practices.
- Provide structured evaluations with clear technical reasoning.
- Collaborate with fellow technical reviewers to ensure high-quality and consistent evaluations.
- Years of Experience: 3+ years of professional software engineering experience or equivalent open-source experience.
- Technical Expertise: Proficiency in JavaScript/TypeScript (Node.js, React, TypeScript), Go (backend services, distributed systems, REST/gRPC APIs, cloud-native development), and Python (backend development, open-source libraries, DevOps/infrastructure automation, testing frameworks).
- Software Engineering: Experience with large production codebases, ability to quickly understand unfamiliar repositories and software architectures, experience reviewing code for readability and maintainability, strong engineering judgment, and understanding of software engineering best practices.
- GitHub & Open Source: Strong understanding of Git and GitHub workflows (branching, Issues, Pull Requests, Code Reviews), experience navigating open-source repositories, familiarity with open-source development workflows, and familiarity with SWE-Bench or SWE-Bench Pro is a plus.
- Communication: Strong English communication skills.
- Employment Type: Hourly
- Expected Commitment: At least 10 hours
- Project Duration: Approximately 1 week
- Expected Timeline: July 8, 14
- Location: Candidates must be located in the US/Canada for real-time collaboration with the project team.
- Hourly Rate: $50 - $75
- This is not a software development role; the focus is on evaluating AI-generated solutions.
$70 - $126 per hour
...technical domain expertise to advance agentic AI workflows by working alongside... ...plan, implement, debug, and refine real software engineering tasks. Your contributions will help train... ...behind decisions and changes. Evaluate AI-generated code for correctness, hallucinations...SuggestedRemote jobHourly payFor contractors$40 - $65 per hour
...Overview Contribute domain expertise to evaluate and harden frontier large language models by crafting adversarial multi-turn... ...on improving how next-generation AI systems learn, reason, and behave,... ..., annotation, or prompt engineering. Preferred: deep familiarity with...SuggestedRemote jobHourly payFor contractors$50 per hour
...Overview This role develops high-quality datasets and evaluation workflows to train and benchmark large language models on software engineering tasks. You will curate and author code examples, correct and improve AI-generated solutions across multiple languages, design...SuggestedFor contractorsRemote work10 hours per weekFlexible hours$70 per hour
...This role focuses on evaluating AI models in the Underwriting domain through a benchmark dataset project. Experts will create complex, grounded tasks that include a clear ground-truth output and an objective rubric, contributing to the advancement of visual document understanding...SuggestedHourly payRemote work$75 per hour
...Role Overview Author complex, grounded evaluation tasks and objective scoring rubrics that measure AI models on visual document understanding and instruction-following within the Electrical and Electronics domain. Your work will feed a benchmark dataset used to evaluate...SuggestedHourly payRemote work$75 per hour
...Join a project focused on evaluating AI models in the architecture domain, specifically in visual document understanding and instruction-following. This role involves authoring complex, grounded tasks that include a clear ground-truth output and an objective rubric....Hourly payRemote work$60 per hour
...This role focuses on evaluating AI models for visual document understanding and instruction-following within the Marketing Analytics domain. You will work on a benchmark dataset project that involves creating and assessing complex tasks related to campaign dashboards,...Hourly payRemote work$100 per hour
...Join a project focused on evaluating AI models for visual document understanding and instruction-following specifically within the Surveying & GIS domain. In this role, you will author complex, grounded tasks that include clear ground-truth outputs and objective rubrics...Hourly payRemote work$90 per hour
...This role involves contributing to a benchmark dataset project that evaluates AI models focused on visual document understanding and instruction-following specifically within the Telecom and Network domain. As an expert, you will be responsible for authoring complex, grounded...Hourly payRemote work$75 per hour
...This role involves contributing to a benchmark dataset project that evaluates AI models focused on visual document understanding and instruction-following specifically within the Energy and Utilities domain. As an expert, you will author complex, grounded tasks that include...Hourly payRemote work$70 per hour
...This role focuses on evaluating AI models for visual document understanding and instruction-following specifically within the Pharma and Clinical Trials domain. You will contribute to a benchmark dataset project by authoring complex, grounded tasks that include a clear...Hourly payRemote work$60 per hour
...This role focuses on evaluating AI models for visual document understanding and instruction-following within the Logistics & Supply Chain domain. You will contribute to a benchmark dataset project by authoring complex, grounded tasks that include clear ground-truth outputs...Hourly payRemote work$20 - $40 per hour
...that will be used to train next-generation AI systems. This role contributes... ...world legal reasoning, helping improve how models learn and perform. No prior AI experience... ...contributions will be used for training and evaluating AI systems. Materials must be original...Remote jobHourly payContract workFor contractors$80 per hour
...Role Overview Design and author grounded evaluation tasks and objective rubrics that test AI models on visual document understanding and instruction-following... ...Subject-matter expertise in Mechanical engineering and CAD, demonstrated by work or domain knowledge...Hourly payRemote work$75 per hour
...This role focuses on evaluating AI models in the Construction Estimating domain through a benchmark dataset project. Experts will be responsible for authoring complex, grounded tasks that include a clear ground-truth output and an objective rubric. Key Responsibilities...Hourly payRemote work$20 - $60 per hour
...realistic business scenarios, create and evaluate Office Open XML documents, and provide detailed... ...helps train and improve large language models. This contractor role focuses on... ...complex document-related prompts and assessing AI outputs. Create, open, and review...Remote jobHourly payFor contractorsWork at officeFree visa$80 - $160 per hour
...Role Overview Model stochastic bacterial population dynamics to derive asymptotic growth rates, analyze effects of growth-rate switching... ...clear, reproducible methodological documentation to train and evaluate AI systems. Key Responsibilities Analyze and model...Remote jobHourly payFor contractors$60 - $85 per hour
..., and optimize GPU-based software and kernels used to train large language models, with a strong focus on runtime... ...for LLM training and evaluation. About the company micro1 is an AI data lab that produces... ...finance, healthcare, and STEM engineering, contribute real-world...Remote jobHourly payFor contractorsImmediate start$80 - $160 per hour
...assignments. Your contributions will be used to train and evaluate next generation AI systems, providing high quality, domain-grounded data and judgments... ..., and producing annotated explanations suitable for model training. Compensation ~ Pay range, $80.00 to $160.00...Remote jobHourly payFor contractors$80 - $160 per hour
...many-body systems, with a focus on problems such as the PXP model, Rydberg blockade, quantum many-body scars, and constrained... ...analyses that support research-level benchmarks and help train and evaluate advanced AI systems. Key Responsibilities Implement and analyze...Remote jobHourly payFor contractors$60 per hour
...Role Overview Develop and author evaluation tasks for a benchmark dataset that measures AI models on visual document understanding and instruction following within the Real Estate Appraisal domain. Tasks must be complex, grounded in real appraisal practice, and paired...Hourly payRemote work$40 - $50 per hour
...quality inputs for next-generation AI systems. You will analyze and... ...and troubleshooting scenarios so models learn from accurate, practical engineering judgment. No prior AI experience is... ...reports, and operational procedures. Evaluate electrical troubleshooting...Remote jobHourly payFor contractorsLocal area$170 - $190 per hour
...-person role in San Francisco supports a healthcare AI partner by ensuring annotated data and model outputs meet clinical standards. Key Responsibilities... ...model interpretation of medical context. Model evaluation and feedback: assess AI recommendations or clinical...Hourly pay- ...LLM training, post-training, and evaluation systems. As an AI/ML Research Engineer, LLM Training & Evaluation , you... ...technical foundations that power model improvement for foundation model builders... ...across model comparisons Software / Data Engineering Proficiency...Full time
$50 per hour
...Role Overview Author and deliver benchmark tasks that evaluate AI models on visual document understanding and instruction-following within the front-end engineering domain. Work will focus on UI screenshots, design specifications, and component diagrams. This opening...Hourly payRemote work- ...BeyondTrust is seeking a Software Engineer (Backend) for the team that develops our Password Safe... ...Demonstrated experience using agentic AI as a fundamental tool integrated into daily... ..., and performant code Propose and evaluate technical solutions as part of research...Full timeRemote work
$100 - $150 per hour
...Role Overview Provide experienced litigation expertise to evaluate, refine, and improve AI-generated legal analysis and written advocacy. In this... ...expert judgment on ambiguous or complex edge cases to inform model calibration on nuanced legal issues. Draft or edit...Remote jobHourly payFor contractors- ...Description We're seeking a talented and motivated full-time Software Engineer, AI Enablement to help Tailscale's engineering organization... ..., Codex , OpenCode , and Pi more effectively. Evaluate new AI models, agents, and tools, and make clear recommendations on...Full time
$20 - $60 per hour
...document expertise to a project that trains next-generation AI systems. You will design realistic Fortune 500 style scenarios and interact iteratively with an advanced language model to create, edit, and evaluate Office Open XML files, with a focus on .pptx deliverables....Remote jobHourly payContract workFor contractorsWork at office$50 - $90 per hour
...validate mental-health safety and crisis-prevention frameworks used to train next-generation AI systems. This contractor role focuses on shaping how digital tools detect, evaluate, and respond to self-harm, eating-disorder risk, emotional dependency, and suicide-related...Remote jobHourly payFor contractors
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer for AI Model Evaluation. Be the first to apply!
- network software engineer Canada
- senior robotics software engineer Canada
- entry level software engineer remote Canada
- cybersecurity software engineer Canada
- agile software developer Canada
- financial software developer Canada
- software engineer Canada
- software engineer healthcare Canada
- startup software engineer Canada
- intel software engineer Canada




