Software Engineer for AI Model Evaluation
SaidGig
Lead the design and evaluation of next-generation coding agents by creating benchmarks, measurement methodologies, datasets, and the tooling that enables rigorous, large-scale assessment and improvement of coding models. Key Responsibilities
- Design and own evaluation frameworks for coding agents, including benchmark specifications, scoring methodologies, rubrics, and quality standards.
- Lead end-to-end research initiatives that measure and improve coding model performance across diverse software engineering tasks.
- Develop high-quality datasets, golden examples, and evaluation protocols to enable reliable assessment of frontier coding systems.
- Analyze model behavior and failure modes, identify systematic weaknesses, and translate findings into actionable improvements for training and evaluation.
- Build tooling and infrastructure to support large-scale experimentation, data generation, review workflows, and evaluation pipelines.
- Establish and document best practices for coding-agent assessment, ensuring methodological rigor, reproducibility, and measurement quality.
- Collaborate with researchers, engineers, and applied AI teams to design experiments and evaluate emerging model capabilities.
- Contribute to technical reports, benchmark studies, and client-facing research deliverables that communicate model performance and insights.
- Required skills: LLMs, coding, evaluation, AI evaluation, ML systems.
- Strong software engineering background with expertise in Python, C++, or comparable programming languages.
- Minimum 3 years of experience in software engineering, machine learning, AI research, evaluation, or related technical disciplines.
- Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies.
- Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems.
- Proven ability to build tooling, automate workflows, and improve technical processes through systematic experimentation.
- Strong analytical skills, with the ability to investigate model behavior and derive insights from complex technical systems.
- Excellent written and verbal communication skills, including the ability to clearly articulate technical findings to diverse audiences.
- Comfortable operating in fast-moving research environments with significant ambiguity and evolving priorities.
- Preferred experience: working on frontier AI systems, coding agents, or model evaluation research; designing benchmarks or datasets for machine learning at scale; familiarity with agentic workflows, tool use, reinforcement learning, or post-training methodologies.
- Preferred evidence of impact: publications, open-source contributions, or demonstrated technical leadership.
- Employment type: Full-time.
- Location: Remote.
- Salary range: $400, 000 to $800, 000 per year.
This is a full-time remote position. The listing does not specify work authorization or visa sponsorship details, candidates should ensure they are able to work in a remote capacity under their own authorization.
$40 per hour
...specialists with project-based AI opportunities for... ..., focused on testing, evaluating, and improving AI... ...coding agents - how well a model handles real-world... ...labeling. Not prompt engineering. Not writing code... ...Qualifications ~5+ years in software development. ~Core...SuggestedPermanent employmentTemporary workPart time$180k - $240k
Role Description Deepgram is looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team responsible for validating the quality of our speech, audio, and multilingual models before they reach customers. This team owns the evaluation and quality...SuggestedFull time$100 per hour
...Role Overview Help train and refine advanced AI systems by applying deep software engineering expertise to evaluate, edit, and produce high‑quality technical content... ...part‑time contractor role focuses on improving how models learn and reason by providing precise, domain‑...SuggestedRemote jobHourly payContract workPart timeFor contractors$85 per hour
...and technical talent with leading AI research labs. Headquartered in San... ...and Jack Dorsey. Position: iOS Engineer (Coding Agent Experience) Type:... ...AI coding agents to complete and evaluate complex engineering tasks. ~Review model-generated mobile application code...SuggestedContract workPart timeSummer workRemote work$40 per hour
A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve technical problems in a flexible remote role. Ideal candidates should have over 2 years in cybersecurity, coding experience, strong analytical and writing...SuggestedHourly payRemote workFlexible hours- ...The Opportunity Deepgram is looking for a Senior Software Engineer - Model Evaluation & AI Systems to join the team responsible for validating the quality of our speech, audio, and multilingual models before they reach customers. This team owns the evaluation and...Full time
$40 per hour
A leading AI training firm is seeking experienced cybersecurity professionals to evaluate AI-generated security content and solve technical problems. This role is remote, allowing you to choose your projects and work schedule. Candidates should have over 2 years of hands...Hourly payRemote work$40 per hour
...cybersecurity professionals to join our team to help train AI models. In this role, you will evaluate AI-generated security content, solve technical... ...penetration testing, red teaming, incident response, detection engineering, DFIR, malware analysis, threat intelligence, or...Hourly payFull timePart timeRemote work$40 per hour
A cybersecurity firm is seeking experienced cybersecurity professionals to join their team in a remote capacity. You will evaluate AI-generated security content and solve technical cybersecurity problems. The ideal candidate will have a minimum of 2 years hands-on experience...Hourly payRemote workFlexible hours- A leading cybersecurity platform is seeking experienced professionals to evaluate AI-generated security content and solve technical cybersecurity issues. This role offers the flexibility of full-time or part-time remote work, allowing you to choose projects and set your...Full timePart timeRemote work
$40 per hour
...professionals to join their remote team. In this role, you will evaluate AI-generated security content, design solutions to cybersecurity problems, and provide essential feedback for improving AI models. Candidates should have over 2 years of hands-on experience in cybersecurity...Hourly payRemote workFlexible hours$40 per hour
A leading cybersecurity solutions provider is seeking experienced cybersecurity professionals for a remote position. You will evaluate AI-generated security content, solve technical problems, and provide essential feedback to improve AI systems. The ideal candidate will...Hourly payRemote work$40 per hour
A leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated security content and solve technical cybersecurity problems. You will enhance how AI systems handle real-world threats while working remotely on an hourly project basis starting at...Hourly payRemote work$30 per hour
A technology company is seeking a Web Platform Engineer to evaluate AI chatbots and enhance model performance. This role requires proficiency in programming languages like Python and JavaScript. You will assess AI outputs from coding challenges and writing tasks, ensuring...Hourly payRemote workFlexible hours$30 - $40 per hour
An AI training company is seeking a Web Platform Engineer to evaluate AI chatbots' outputs and improve their logic. The role allows for remote work and on-demand project selection, paying $30-$40+ per hour. Candidates should be fluent in English and have experience with...Hourly payFor contractorsRemote work$238k - $302k
...across 15+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in... ...for quantitatively-minded engineers to research and propose new ways... ...experience in a heavily quantitative software engineering area ~ Experience navigating...Full timeRemote work$172.43k - $230.95k
.... As the only vertically integrated AI infrastructure company built from the... ...Crusoe.About This Role:The Senior Software Engineer for the AI Model Lifecycle team will play a crucial role... ...management: versioning, lineage, evaluation, and reproducible fine-tuning at scale...Temporary work$204k - $259k
...The core challenge within Model Lifecycle is accelerating Waymo... ...role, you will report to an engineering manager. You will:... ...efficient model training and evaluation. Develop infrastructure to... ...Passionate about data-centric AI and autonomous driving applications...Full timeRemote work$405k
...interpretable, and steerable AI systems. We want AI to... ...committed researchers, engineers, policy experts, and... ...re looking for a Staff Software Engineer to set... ...systems, tooling, and evaluation infrastructure that determine... ...frameworks that measure model capabilities across diverse...Full timeWork at officeVisa sponsorshipFlexible hours$40 per hour
A cybersecurity company is seeking experienced professionals to evaluate AI-generated security content and solve technical problems. This remote position offers the flexibility to choose projects and work on your own schedule, with projects starting at $40 per hour. Candidates...Remote jobHourly pay$40 per hour
A leading AI security solutions provider is seeking experienced cybersecurity professionals to evaluate AI-generated security content and solve real-world technical problems. In this remote role, candidates will require over 2 years of cybersecurity experience, fluency...Remote jobHourly pay- A leading cybersecurity firm is seeking experienced cybersecurity professionals for a remote role to help train AI models. Candidates will evaluate AI-generated security content, solve technical cybersecurity problems, and provide valuable feedback for the improvement of...Remote jobFlexible hours
- ...leading cybersecurity firm is seeking experienced professionals to evaluate AI-generated cybersecurity content and solve technical security problems. You will play a significant role in training AI models, providing critical feedback, and improving system accuracy. This...Remote jobFlexible hours
$400 per month
...Mercor is partnering with a leading AI research lab to support a Frontier... ...project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments... ...on realistic infrastructure engineering workflows and model evaluation. Spots...$145k - $200k
...builds the world’s leading software for data-driven decisions and... ....The Role We are a software engineering team with expertise in enabling ML models in production. We deploy AI models to run in variety of... ...and the ability to quickly evaluate and integrate new models and...Full timeWork experience placementWork at officeRemote workWork from homeRelocation package$86.8k - $198k
Model and Simulation Software EngineerThe Opportunity: You will play a critical role... ..., and integrating AI‑enabled models that support... ...DoW. Using core software engineering principles, you’ll build scalable... ...generation M&S systems by evaluating new frameworks, enhancing...Full timeContract workPart timeWork at officeLocal areaRemote work$220k - $320k
...hosts specialized language models for companies that need frontier-quality AI at a fraction of the... ..., training, evaluation, and planet-scale hosting... ...funded ten‑person team of engineers who work in‑person in downtown... ...and run their own software companies. We are high‑...Work at office$152k - $241.5k
...tapping into the unlimited potential of AI to define the next era of computing.... ...on the world.We are seeking a Software Engineer - Scientific Evaluation to own a shared platform for classical... ...packages, PyTorch integrations, scientific models, and AI agents. This hands-on role...Full time$120k - $200k
...in Silicon Valley, Pony.ai has quickly become a global... ...algorithms and evaluation metrics to drive core AI... ...and optimize downstream engineering workflows for Large Language Models (LLMs), programmatically... ...skills in C/C++, Python, and software designStrong foundation...Full timeTemporary work$144.7k - $221.4k
...the Organization The Evaluation team builds and evolves... ...into clear feedback for engineering and leadership, and... ...introspect autonomous driving software performance at... ...prediction, and planning models. Build and maintain... ...Experience leveraging AI-assisted development and...Full timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer for AI Model Evaluation. Be the first to apply!
- agile software developer United States
- software developer internship no experience United States
- intermediate software engineer United States
- software engineer staff United States
- experienced software developer United States
- software engineer co-op United States
- work from home software developer United States
- software developer no experience United States
- software developer fintech United States
- software data engineer United States







