Software Engineer for AI Model Evaluation
$220kSaidGig
This role focuses on advancing the evaluation and development of frontier coding agents, blending AI research, software engineering, and model evaluation. You will design benchmarks, methodologies, and data systems that define how next-generation coding models are assessed and enhanced. Key Responsibilities
- Design and own evaluation frameworks for coding agents, including benchmark specifications, scoring methodologies, rubrics, and quality standards.
- Lead end-to-end research initiatives aimed at measuring and improving coding model performance across various software engineering tasks.
- Develop high-quality datasets, golden examples, and evaluation protocols for reliable assessment of frontier coding systems.
- Analyze model behavior and failure modes, identifying systematic weaknesses and translating findings into actionable improvements for training and evaluation.
- Build tooling and infrastructure to support large-scale experimentation, data generation, review workflows, and evaluation pipelines.
- Establish best practices for coding-agent assessment, ensuring methodological rigor, reproducibility, and measurement quality.
- Collaborate closely with researchers, engineers, and applied AI teams to design experiments and evaluate emerging model capabilities.
- Contribute to technical reports, benchmark studies, and client-facing research initiatives that communicate model performance and insights.
- Strong software engineering background with expertise in Python, C++, or comparable programming languages.
- 3+ years of experience in software engineering, machine learning, AI research, evaluation, or related technical disciplines.
- Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies.
- Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems.
- Proven ability to build tooling, automate workflows, and improve technical processes through systematic experimentation.
- Strong analytical skills with the ability to investigate model behavior and derive insights from complex technical systems.
- Excellent written and verbal communication skills, with the ability to clearly articulate technical findings to diverse audiences.
- Comfortable operating in fast-moving research environments with significant ambiguity and evolving priorities.
- Experience working on frontier AI systems, coding agents, or model evaluation research.
- Deep interest in understanding how data, evaluations, and feedback mechanisms influence model capabilities.
- Track record of independently driving ambiguous technical or research projects from conception to execution.
- Experience designing benchmarks or datasets for machine learning systems at scale.
- Familiarity with agentic workflows, tool use, reinforcement learning, or post-training methodologies.
- Publications, open-source contributions, or demonstrated technical leadership in AI research.
Full-time position with remote work flexibility.
CompensationAnnual salary range of $220, 000 - $500, 000.
EligibilityOpen to candidates with the required skills and experience, regardless of location.
$40 per hour
...specialists with project-based AI opportunities for... ..., focused on testing, evaluating, and improving AI... ...coding agents - how well a model handles real-world... ...labeling. Not prompt engineering. Not writing code... ...Qualifications ~5+ years in software development. ~Core...SuggestedPermanent employmentTemporary workPart time$30 per hour
A technology company is seeking a Web Platform Engineer to evaluate AI chatbots and enhance model performance. This role requires proficiency in programming languages like Python and JavaScript. You will assess AI outputs from coding challenges and writing tasks, ensuring...SuggestedHourly payRemote workFlexible hours$105 per hour
...leverage their technical skills to contribute to AI research projects, focusing on enhancing the capabilities of Large Language Models (LLMs) in business communication and... ...Responsibilities Develop domain-specific prompts and evaluate LLM responses. Conduct independent...SuggestedRemote workFlexible hours$40 per hour
A cybersecurity company is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. This remote position offers the flexibility to choose projects and work on your own schedule, with projects starting at $40 per hour. Candidates...SuggestedHourly payRemote work$238k - $302k
...across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in... ...for quantitatively-minded engineers to research and propose new ways... ...experience in a heavily quantitative software engineering area ~ Experience navigating...SuggestedFull timeRemote work- A leading cybersecurity platform is seeking experienced professionals to evaluate AI-generated security content and solve technical cybersecurity issues. This role offers the flexibility of full-time or part-time remote work, allowing you to choose projects and set your...Full timePart timeRemote work
$40 per hour
A cybersecurity solutions provider is seeking experienced cybersecurity professionals for a REMOTE position. You will evaluate AI-generated security content, solve technical problems, and contribute to cybersecurity tools using your expertise. Candidates should have 2+...Hourly payRemote work$40 per hour
A technology company specializing in cybersecurity is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. This remote position is suitable for candidates with 2+ years in cybersecurity and a background in penetration...Hourly payRemote workFlexible hours$40 per hour
A cybersecurity-focused company is looking for experienced professionals to evaluate AI-generated security content and provide feedback to improve AI systems' understanding of threats. This role, which can be full-time or part-time, allows for flexible project selection...Hourly payFull timePart timeRemote workFlexible hours$40 per hour
...professionals to join their remote team. In this role, you will evaluate AI-generated security content, design solutions to cybersecurity problems, and provide essential feedback for improving AI models. Candidates should have over 2 years of hands-on experience in cybersecurity...Hourly payRemote workFlexible hours$40 per hour
A leading cybersecurity solutions provider is seeking experienced cybersecurity professionals for a remote position. You will evaluate AI-generated security content, solve technical problems, and provide essential feedback to improve AI systems. The ideal candidate will...Hourly payRemote work$40 per hour
...cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content, solve technical... ...penetration testing, red teaming, incident response, detection engineering, DFIR, malware analysis, threat intelligence, or...Hourly payFull timePart timeRemote work$20 - $60 per hour
...expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and perform through high-quality... ...in computational, simulation, or systems engineering to inform advanced AI benchmarking and evaluation processes. Analyze and provide...Hourly payContract workRemote work$204k - $259k
...The core challenge within Model Lifecycle is accelerating Waymo... ...role, you will report to an engineering manager. You will:... ...efficient model training and evaluation. Develop infrastructure to... ...Passionate about data-centric AI and autonomous driving applications...Full timeTemporary workRemote work$405k
...interpretable, and steerable AI systems. We want AI to... ...committed researchers, engineers, policy experts, and... ...re looking for a Staff Software Engineer to set... ...systems, tooling, and evaluation infrastructure that determine... ...frameworks that measure model capabilities across diverse...Full timeWork at officeVisa sponsorshipFlexible hours$40 per hour
A cybersecurity firm is seeking experienced cybersecurity professionals to join their team in a remote capacity. You will evaluate AI-generated security content and solve technical cybersecurity problems. The ideal candidate will have a minimum of 2 years hands-on experience...Remote jobHourly payFlexible hours$40 per hour
A leading AI-focused cybersecurity firm is looking for experienced cybersecurity professionals to evaluate AI-generated content and solve technical security problems. In this flexible role, you can work remotely and choose your projects. Ideal candidates will have 2+ years...Remote jobHourly payFlexible hours$30 - $90 per hour
...collaborating with cutting-edge AI research. As an... ...a high-caliber engineering team. Key Responsibilities... ...test new AI-powered models in Cursor, providing actionable... ...designing or evaluating experimental tooling and... ...AI advancements in software development. Work Terms...Hourly payContract workRemote work$30 - $40 per hour
An AI training company is seeking a Web Platform Engineer to evaluate AI chatbots' outputs and improve their logic. The role allows for remote work and on-demand project selection, paying $30-$40+ per hour. Candidates should be fluent in English and have experience with...Remote jobHourly payFor contractors$40 per hour
A leading AI security solutions provider is seeking experienced cybersecurity professionals to evaluate AI-generated security content and solve real-world technical problems. In this remote role, candidates will require over 2 years of cybersecurity experience, fluency...Remote jobHourly pay$40 per hour
A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve technical cybersecurity problems. You will enhance how AI systems handle real-world threats while working remotely on an hourly project basis starting at...Remote jobHourly pay$40 per hour
A leading AI training firm is seeking experienced cybersecurity professionals to evaluate AI-generated security content and solve technical problems. This role is remote, allowing you to choose your projects and work schedule. Candidates should have over 2 years of hands...Remote jobHourly pay- A cybersecurity solutions company is looking for experienced cybersecurity professionals to help train AI models. You will work remotely to evaluate AI-generated security content, solve technical problems, and provide feedback to improve AI systems. Ideal candidates have...Remote jobFlexible hours
$40 per hour
...cybersecurity innovation company is seeking experienced professionals to evaluate AI-generated security content and solve technical cybersecurity... .... In this remote position, you will work with advanced AI models and contribute to improving cybersecurity tools. The ideal...Remote jobHourly payFlexible hours$40 per hour
...leading cybersecurity firm is looking for experienced cybersecurity professionals to evaluate AI-generated content and solve technical problems. The role involves working with advanced AI models, providing feedback, and contributing to the cybersecurity industry's future....Remote jobHourly payFlexible hours$40 per hour
A cybersecurity firm is looking for experienced professionals to join its team. This remote role involves evaluating AI-generated security content and solving technical cybersecurity problems. Candidates should have over 2 years of hands-on experience in cybersecurity and...Remote jobHourly payFlexible hours$40 per hour
A leading cybersecurity firm is seeking qualified professionals to evaluate AI-generated security content and solve technical problems. This remote role requires hands-on experience in cybersecurity, including penetration testing or related areas. Candidates should possess...Remote jobHourly payFlexible hours$40 per hour
A leading cybersecurity firm is seeking experienced professionals to join their team in evaluating AI-generated security content. You will solve technical problems and provide feedback to enhance AI capabilities related to real-world threats. The ideal candidate has over...Remote jobHourly pay$40 per hour
A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve technical problems in a flexible remote role. Ideal candidates should have over 2 years in cybersecurity, coding experience, strong analytical and writing...Remote jobHourly payFlexible hours- ...leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated cybersecurity content and solve technical security problems. You will play a significant role in training AI models, providing critical feedback, and improving system accuracy. This...Remote jobFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer for AI Model Evaluation. Be the first to apply!
- software engineer full time United States
- software system engineer United States
- consulting software engineer United States
- software engineer travel United States
- software engineer mainframe United States
- software developer trainee United States
- real time software engineer United States
- network software engineer United States
- senior software engineer remote United States
- entry level software engineer remote United States




