Machine Learning Engineer for AI Model Evaluation
$85 per hourSaidGig
Evaluate and improve frontier AI coding models by completing structured technical assessments that mirror realistic machine learning engineering workflows, model training and inference systems, MLOps, and LLM application scenarios.
Key Responsibilities- Use frontier AI coding agents to complete and evaluate complex ML and AI engineering tasks.
- Review model-generated implementations across model training, inference systems, deployment infrastructure, and LLM applications.
- Identify bugs, edge cases, performance regressions, and failure modes in model outputs and implementations.
- Compare outputs from multiple frontier models and assess their relative strengths, weaknesses, and tradeoffs.
- Apply professional engineering judgment to realistic ML engineering scenarios, documenting findings and recommendations.
- At least 2 years of professional machine learning engineering experience.
- Experience building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products.
- Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
- Ability to evaluate model-generated machine learning implementations and reason about technical tradeoffs.
- Experience deploying ML systems to production is preferred.
- Location: Remote.
- Employment type: hourly.
- Sprint-based engagement, with work organized into 12-24 hour stretches based on client requirements.
- Spots are limited and are filled on a first-come, first-serve basis.
- $400 per accepted task.
- Typical tasks take approximately 2-3 hours after ramp-up.
- Compensation is tied to accepted work.
- Hourly rate (metadata): $85 per hour.
This role is intended for engineers with 2+ years of ML engineering experience who regularly use AI coding agents and can assess model-generated ML solutions. Preference is given to candidates with experience deploying ML systems to production. Compensation is contingent on accepted deliverables.
$208k - $300k
...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure...SuggestedFull time$60 - $90 per hour
...Machine Learning Engineer — Model Evaluation & Experimentation is a remote review track for evaluating AI outputs across machine learning engineer model evaluation experimentation specialist operations workflows. Reviewers grade workflow correctness, policy adherence,...SuggestedRemote jobFor contractorsWork experience placement10 hours per week$70 per hour
...Role Overview Join a distributed talent network to provide expert Machine Learning engineering support on contract projects with AI labs and companies. Contributors help train and evaluate models, design real-world tasks and deliverables, and give domain-specific feedback...SuggestedHourly payContract workRemote work$228.7k - $343.1k
...enormous scale, and one bad model can mean millions in... ...to models applies to AI. We build the tooling that... ..., so you critically evaluate what it produces and own... ...you did not write, learning the data, configs, and... ...Solid software and data engineering: production-quality Python...SuggestedRemote jobFull timeLocal areaShift work$213k - $263k
...state-of-the-art Generative AI to create a training ground... ...Driver. The Simulator Evaluation team faces the ultimate data... ...We are seeking visionary machine learning engineers and researchers to architect... ...realism of our multimodal world models. Your work will define the...SuggestedFull timeRemote work$60 - $90 per hour
...Drive research-grade data analyses and create evaluation tasks that reveal where frontier generative AI models fail. You will author realistic, multi-skill analysis... ...1 year of experience in a research, research-engineering, or intensive data-analysis role. Strong, hands...Hourly payFull timeFreelanceRemote work- ...Role Overview Medical professionals evaluate AI-generated medical content and use their clinical and field experience to improve model outputs. No prior AI experience is required. In this role you will assess model responses, create realistic prompts that reflect clinical...Hourly payPart timeImmediate startRemote workFlexible hours
$40 - $65 per hour
...Contribute domain expertise to evaluate and harden frontier large language models by crafting adversarial multi-turn... ...improving how next-generation AI systems learn, reason, and behave, and does... ...evaluations, annotation, or prompt engineering. Preferred: deep familiarity...Remote jobHourly payFor contractors$238k - $302k
...The mission of the Waymo AI Foundations team is to develop machine learning solutions addressing... ...demonstration, generative modeling, Bayesian inference,... ...hierarchical learning, and robust evaluation. This role follows a... ...Senior Staff Software Engineer. You will:...Full timeRemote work$80 per hour
...Role Overview Help evaluate and improve frontier AI coding models by using AI coding agents to complete realistic data engineering tasks, then assess the outputs for correctness, scalability, and failure modes. Work centers on end-to-end data engineering workflows including...Hourly payWork at officeRemote work$204k - $259k
...Driver Understanding and Evaluation (DUE) team at Waymo is... ...Driver. The DUE Machine Learning team will build and operate... ...machine learning models to deliver training and... ...and software engineers who are passionate about... ...modeling and generative AI into robust, production...Full time$120 per hour
...Government Quantitative Professionals evaluate AI model outputs and provide structured, domain-... ...by the research project. Optionally learn new evaluation techniques and tools as... ...techniques to practical problems in science, engineering, business, or security, and sharing...Hourly payTemporary workPart timeImmediate startRemote workFlexible hours- ...providing Information Technology, Engineering Services, Program Management, and... ...Solerity is seeking Mid to Senior Machine Learning Engineers and AI Model Developers to support an upcoming... ...This effort focuses on developing, evaluating, and integrating machine learning...Full timeFor contractorsRemote workFlexible hours
$251k - $310k
...S. states. The DUE Machine Learning team will build and operate... ...and speed up the evaluation and onboard developer... ...advanced machine learning models to deliver training... ...researchers and software engineers who are passionate... ...for evaluating complex AI systems. ~ Track record...Full time$39 per hour
...Music Audio Expert, where you will play a crucial role in evaluating generative musical AI models in collaboration with a leading AI lab. This position... ...quality, and production/mix quality, using audio-engineering terminology. Annotate songs in detail, including genre...Hourly payPart timeImmediate start10 hours per week$60 - $80 per hour
...foundational Large Language Models, by designing realistic marketing... ...accurate solutions, and evaluating model outputs with rigorous,... ...places you within a leading AI lab''s extended workforce while... ...Responsibilities Guide research and engineering teams to close knowledge gaps...Hourly payWeekday work$30 - $90 per hour
...Role Overview Build and maintain backend services in Go while testing and evaluating alpha-stage AI coding models. This part-time, remote contract role combines hands-on Go development with structured model evaluation, bug reporting, and real-time collaboration to improve...Hourly payContract workPart timeFor contractorsRemote work$75 per hour
...Role Overview Lead the clinical evaluation of medical AI by designing and executing assessments that probe clinical reasoning and decision-making... ...teams to create realistic clinical problems, find where models fail or lack medical knowledge, and shape improvements so AI...Contract workRemote workFlexible hours$60 - $150 per hour
...Role Overview Provide legal subject-matter expertise to improve and evaluate AI systems, by designing realistic legal tasks, reviewing model outputs, and giving domain-specific feedback that advances frontier AI research. This is an open application to join a Law Expert...Hourly payContract workImmediate startRemote work$85 per hour
...environmental assessment, GIS analysis, and renewable energy siting to evaluate AI-generated geospatial outputs and develop expert training data... ..., GIS Analyst, Interconnection Specialist, Solar Design Engineer, or equivalent positions in the energy industry. Hands-on...Hourly payContract workPart timeWork at officeRemote workFlexible hours- ...team of experienced sellers, engineers, and researchers. Many of us worked... ...the role Lightfield's AI/ML team builds the experiences... ...Pioneer the training of new models that leverage both historical... ...strong understanding of deep learning AI/ML frameworks or cloud services...Full time
$40 per hour
...data annotation company is seeking professionals in quantitative fields to enhance AI development. This fully remote role allows individuals to set flexible schedules while evaluating AI-generated analyses and solving complex quantitative problems. Candidates should have...Hourly payRemote workFlexible hours$40 per hour
A leading AI development firm in Michigan is seeking experienced quantitative professionals to evaluate AI-generated analyses and contribute to the evolution of AI models. Candidates should have a background in data science, statistics, or similar fields, with at least...Hourly payRemote work$40 per hour
...A forward-thinking AI team is seeking quantitative professionals to evaluate and improve cutting-edge AI systems. The role involves working on AI-generated analyses and providing critical feedback for model enhancement. Candidates should have a strong quantitative background...Hourly payRemote workFlexible hours$60 per hour
...contribute to developing cutting-edge AI systems, while enjoying the... ...advance AI development. AI models are increasingly capable of... ...-art AI models on tasks like evaluating AI-generated quantitative... ...Computer Science, Mathematics, Engineering, or similar); a master's or...Hourly payFull timeRemote workFlexible hours$110 per hour
...Role Overview Join a Physician Expert Network to provide clinical expertise to AI research teams and companies. Members contribute medical knowledge to train and evaluate AI models, design realistic clinical tasks, and give domain-specific feedback that advances medical...Hourly payContract workRemote work$40 per hour
A data science team is seeking experienced quantitative professionals to evaluate AI-generated work and contribute to the development of cutting-edge AI systems. This fully remote position offers flexible scheduling and competitive hourly pay starting at $40+. Ideal candidates...Hourly payRemote workFlexible hours$40 per hour
A leading AI development company is seeking experienced quantitative professionals for a remote role. Candidates will evaluate AI-generated quantitative work, solve complex problems, and provide valuable feedback. The ideal candidate has 2+ years of experience in a quantitative...Hourly payFull timeRemote workFlexible hours$40 per hour
...An innovative AI development company is seeking experienced quantitative professionals to contribute to AI advancements. This fully remote role involves evaluating AI-generated analyses and ensuring they are technically accurate and valid in real-world scenarios. Candidates...Hourly payRemote workFlexible hours- ...Physicians apply clinical judgment and frontline medical experience to evaluate AI-generated medical content, ensuring clinical accuracy, sound... ...planning. Assess clarity, relevance, and safety of model outputs in realistic care scenarios. Provide detailed, constructive...Full timeFor contractorsPrivate practiceRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Engineer for AI Model Evaluation. Be the first to apply!
- computer vision machine learning engineer United States
- junior machine learning research engineer United States
- machine learning software engineer United States
- junior machine learning engineer United States
- machine learning ai engineer United States
- data scientist machine learning engineer United States
- machine learning engineer United States
- graduate machine learning engineer United States
- ai ml engineer United States
- senior ml engineer United States



