Machine Learning Engineer for AI Model Evaluation
$85 per hourSaidGig
Help evaluate and improve frontier AI coding models by completing structured technical assessments that simulate realistic machine learning engineering workflows. This role focuses on using and critiquing coding agents to surface bugs, failure modes, and deployment tradeoffs, with limited spots available on a first come, first serve basis.
Key Responsibilities- Use frontier AI coding agents to complete and evaluate complex ML and AI engineering tasks.
- Review model-generated implementations related to model training, inference systems, MLOps, and LLM applications.
- Identify bugs, edge cases, performance problems, and failure modes in model outputs and implementations.
- Compare outputs from multiple frontier models, assessing their relative strengths and weaknesses.
- Apply professional engineering judgment to realistic ML engineering scenarios and recommend actionable improvements.
- At least 2 years of professional machine learning engineering experience.
- Experience building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products.
- Regular use of AI coding agents such as Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or similar tools.
- Ability to evaluate model-generated ML implementations and reason about technical tradeoffs.
- Experience deploying ML systems to production is preferred.
- Remote engagement.
- Employment type, hourly.
- Sprint-based project work, with assignments running in 12 to 24 hour stretches depending on client requirements.
- Spots are limited and are filled on a first come, first serve basis.
- Task availability and placement depend on client needs.
- Listed hourly rate: $85 per hour.
- Alternate payment detail: $400 per accepted task, with typical tasks taking approximately 2 to 3 hours after ramp-up.
- Compensation is tied to accepted work.
- Ideal candidates have 2 or more years of professional ML engineering experience and relevant production deployment experience.
- Regular practical experience with AI coding agents and with evaluating model-generated engineering outputs is required.
$208k - $300k
...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure...SuggestedFull time- ...Author and run rigorous, multi-step machine learning evaluation tasks for a leading generative AI research team. You will take... ...to determine where frontier models succeed or fail. Typical tasks... ...experience in a research or research-engineering position. Hands-on...SuggestedHourly payFull timeFreelanceRemote work
$228.7k - $343.1k
...enormous scale, and one bad model can mean millions in... ...to models applies to AI. We build the tooling that... ..., so you critically evaluate what it produces and own... ...you did not write, learning the data, configs, and... ...Solid software and data engineering: production-quality Python...SuggestedRemote jobFull timeLocal areaShift work$174.72k - $295.68k
..., integrating advanced AI and autonomous driving... ...cutting-edge R&D in AI, machine learning, and smart connectivity... ...-time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development... ...systematic ablation, evaluation, and visualization of...SuggestedFull time$60 - $90 per hour
...Machine Learning Engineer — Model Evaluation & Experimentation is a remote review track for evaluating AI outputs across machine learning engineer model evaluation experimentation specialist operations workflows. Reviewers grade workflow correctness, policy adherence,...SuggestedRemote jobFor contractorsWork experience placement10 hours per week- ...embodied intelligence. Our AI-driven systems enable robots to adapt, learn, and perform in the... ..., complex, and poorly modeled physics that... ...are seeking a Senior Machine Learning Engineer to lead the development... ...process-optimization, evaluation, and synthetic-data workflows...Shift work
$213k - $263k
...state-of-the-art Generative AI to create a training ground... ...Driver. The Simulator Evaluation team faces the ultimate data... ...We are seeking visionary machine learning engineers and researchers to architect... ...realism of our multimodal world models. Your work will define the...Full timeRemote work$150k
...Machine Learning Engineer About the Institute of Foundation Models: We are a dedicated research lab for building, understanding,... ...nurture the next generation of AI builders, and drive transformative... ..., experimentation, and evaluation workflows. ~ This role...Visa sponsorship$70 per hour
...Role Overview Join a remote Machine Learning Engineer talent network to be considered for future contract engagements with AI labs and companies. This is an open application... ...contribute to advancing AI by training and evaluating models, designing real-world tasks and...Hourly payContract workFor contractorsRemote work$224k - $356.5k
We are looking for outstanding Machine Learning Engineers to join our Physical AI teams. As the pioneers of the GPU—the... ...state-of-the-art multimodal models and diffusion techniques to simulate... ...Establish a strong mentality for KPI evaluation and validation to ensure the...Full time$204k - $259k
...Driver Understanding and Evaluation (DUE) team at Waymo is... ...Driver. The DUE Machine Learning team will build and operate... ...machine learning models to deliver training and... ...and software engineers who are passionate about... ...modeling and generative AI into robust, production...Full time$238k - $302k
...The mission of the Waymo AI Foundations team is to develop machine learning solutions addressing... ...demonstration, generative modeling, Bayesian inference,... ...hierarchical learning, and robust evaluation. This role follows a... ...Senior Staff Software Engineer. You will:...Full timeRemote work- ...providing Information Technology, Engineering Services, Program Management, and... ...Solerity is seeking Mid to Senior Machine Learning Engineers and AI Model Developers to support an upcoming... ...This effort focuses on developing, evaluating, and integrating machine learning...Full timeFor contractorsRemote workFlexible hours
$251k - $310k
...S. states. The DUE Machine Learning team will build and operate... ...and speed up the evaluation and onboard developer... ...advanced machine learning models to deliver training... ...researchers and software engineers who are passionate... ...for evaluating complex AI systems. ~ Track record...Full time$100 per hour
...knowledge to help train and evaluate next-generation AI systems by reviewing,... ...clarity, and relevance of model outputs through rubric-based... ...in Data Science, Machine Learning, Applied AI, Statistics,... ...Background in prompt engineering, AI output evaluation, fact...Hourly payContract workPart timeFor contractorsRemote work- ...VTI Aerospace builds AI-powered perception and... ...WA, we are a team of engineers and technologists from... ...Senior Software Engineer – Evaluation, you will design and... ...), and small language model (SLM) systems. You... ...will work closely with machine learning and data engineering teams...
- ...team of experienced sellers, engineers, and researchers. Many of us worked... ...the role Lightfield's AI/ML team builds the experiences... ...Pioneer the training of new models that leverage both historical... ...strong understanding of deep learning AI/ML frameworks or cloud services...Full time
$80 per hour
...Role Overview Evaluate and improve frontier AI coding agents by applying professional data engineering judgment to realistic data infrastructure and pipeline scenarios. You will use and assess model-generated implementations for ETL, data warehouses, analytics platforms...Hourly payRemote work$224k - $356.5k
...tapping into the unlimited potential of AI to define the next era of computing.... .... As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful... ...developing or assessing contemporary machine learning and deep learning systems.Hands...Full time$40 per hour
A biotechnology company is seeking a Biotechnology R&D Scientist to train AI models and evaluate their outputs. This role involves measuring the progress of AI chatbots with complex biology questions and ensuring their performance and correctness. Candidates should have...Hourly payRemote workFlexible hours- ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote workFlexible hours
$40 per hour
A forward-thinking AI development firm seeks experienced quantitative professionals to evaluate AI-generated work, applying their skills in statistical analysis, predictive modeling, and technical writing. This fully remote opportunity offers a flexible schedule and projects...Hourly payRemote workFlexible hours$60 per hour
...A leading data analysis firm is seeking experienced quantitative professionals to join their remote team. In this role, you'll evaluate AI-generated quantitative analysis, design problem-solving tasks for AI training, and provide insightful feedback on AI systems. Candidates...Remote work$60 per hour
...A leading AI development company is seeking quantitative professionals to evaluate AI-generated analyses and develop solutions in various quantitative fields. This fully remote position allows for flexible scheduling and competitive pay up to $60/hour. Candidates should...Remote workFlexible hours$20 per hour
...DataAnnotation is committed to creating quality AI. Join our team to help train AI chatbots while... ...chatbots. You will develop complex prompts to test AI models, write high-quality responses to demonstrate excellence, and evaluate different model outputs based on accuracy and...Hourly payFull timeContract workPart timeFor contractorsSelf employmentFreelanceRemote work$60 per hour
...A tech company focused on AI is seeking quantitative professionals to evaluate AI-generated analytical work and help advance AI development. Responsibilities include statistical analysis, predictive modeling, and providing feedback to shape AI systems. The role offers...Hourly payRemote workFlexible hours$60 per hour
...contribute to developing cutting-edge AI systems, while enjoying the... ...advance AI development. AI models are increasingly capable of... ...-art AI models on tasks like evaluating AI-generated quantitative... ...Computer Science, Mathematics, Engineering, or similar); a master's or...Hourly payFull timeRemote workFlexible hours$40 per hour
A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting...Hourly payRemote work$60 per hour
A leading data science company is seeking quantitative professionals to evaluate AI-generated analyses and design quantitative problems crucial for advancing AI systems. This fully remote role allows you to work flexibly and choose your projects, with competitive hourly...Hourly payRemote work$60 per hour
A leading AI development firm is seeking experienced quantitative professionals to contribute to AI systems. The role involves evaluating AI-generated work and solving technical problems with a flexible remote schedule. Candidates should have a background in data science...Hourly payRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Engineer for AI Model Evaluation. Be the first to apply!
- junior machine learning engineer United States
- graduate machine learning engineer United States
- senior ml engineer United States
- data scientist machine learning engineer United States
- machine learning engineer United States
- lead machine learning engineer United States
- machine learning ai engineer United States
- entry level machine learning engineer United States
- machine learning software engineer United States
- ai ml engineer United States



