Machine Learning Engineer for Model Evaluation
SaidGig
Author and run rigorous, multi-step machine learning evaluation tasks for a leading generative AI research team. You will take high-level research ideas, implement changes to training or evaluation procedures, execute experiments, and analyze results to determine where frontier models succeed or fail. Typical tasks require one to two days of continuous, focused work and combine implementation, experiment execution, and careful analysis. You will collaborate closely with the lab''s researchers in a fast feedback loop.
Key Responsibilities- Design tasks that translate real ML research ideas into well-defined, multi-step experiments, for example modifying how an RL reward is computed and specifying success criteria.
- Implement changes to code and training pipelines required by each task.
- Set up, run, and monitor training experiments end-to-end, ensuring reproducibility and correctness.
- Analyze experiment outputs to demonstrate what a correct solution looks like and to identify model failure modes.
- Build tasks that explore reinforcement learning concepts such as reward functions and training behavior.
- Evaluate how frontier models handle your tasks, documenting where and why they fall short.
- Share findings and align task design, rigour, and fairness with researchers and fellow experts.
- MSc or PhD in machine learning, computer science, or another STEM field, or equivalent practical experience in a research-heavy role.
- At least 1 year of experience in a research or research-engineering position.
- Hands-on experience training and evaluating ML models and running experiments end-to-end, including setup, execution, and analysis.
- Strong familiarity with large language models, including their capabilities, limitations, and common evaluation techniques.
- Working proficiency in Python and Git, comfortable in both scripting and notebook environments.
- Basic understanding of reinforcement learning concepts such as reward functions and policy training is preferred.
- Prior experience in AI training, model evaluation, or benchmark or task authoring is preferred.
- High attention to detail, creative task design, strong written communication, and the ability to work independently on ambiguous, open-ended problems.
- Availability to engage reliably for approximately 35 hours per week.
- Employment type: W-2 employment through Cincinnatus LLC, placed to work as part of a leading AI lab''s extended workforce.
- Work location: fully remote within the United States.
- Time commitment: approximately 35 hours per week, full-time role-based assignment integrated into the client team''s workflows.
- Role structure: this is a structured, role-based position rather than a project-based or freelance engagement; it involves close collaboration with the client team.
- Hourly W-2 pay range: $60.00 to $90.00 per hour.
- Cincinnatus LLC serves as the employer of record and administers employment, onboarding, payroll, and benefits for this role.
- To apply, submit your application through the job listing or application portal; applications and subsequent employment administration will be handled by Cincinnatus LLC.
- Equal Employment Opportunity: Cincinnatus is an equal opportunity employer and provides reasonable accommodations for qualified applicants with disabilities.
$208k - $300k
...Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems—including LLMs, agentic models, and multimodal pipelines—into mission-critical government environments. We build evaluation frameworks that ensure...SuggestedFull time$85 per hour
...Role Overview Evaluate and improve frontier AI coding agents by completing realistic machine learning engineering tasks and assessing model outputs. You will perform structured technical assessments that reflect production ML workflows, helping a leading AI research lab...SuggestedHourly payRemote work$60 - $90 per hour
...Machine Learning Engineer — Model Evaluation & Experimentation is a remote review track for evaluating AI outputs across machine learning engineer model evaluation experimentation specialist operations workflows. Reviewers grade workflow correctness, policy adherence,...SuggestedRemote jobFor contractorsWork experience placement10 hours per week$228.7k - $343.1k
...enormous scale, and one bad model can mean millions in credit losses... ...at scale, so you critically evaluate what it produces and own the... ...codebases you did not write, learning the data, configs, and... .... Solid software and data engineering: production-quality Python, SQL...SuggestedRemote jobFull timeLocal areaShift work$174.72k - $295.68k
...through cutting-edge R&D in AI, machine learning, and smart connectivity.We are... ...for a full-time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic development of... ...).Conduct systematic ablation, evaluation, and visualization of model...SuggestedFull time- ...systems enable robots to adapt, learn, and perform in the real... ...fast, complex, and poorly modeled physics that traditional... ...ones.We are seeking a Senior Machine Learning Engineer to lead the development of... ...planning, process-optimization, evaluation, and synthetic-data...Shift work
$213k - $263k
...for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge... ...is "real"? We are seeking visionary machine learning engineers and researchers to architect the... ...the realism of our multimodal world models. Your work will define the state of the...Full timeRemote work$150k
...Machine Learning Engineer About the Institute of Foundation Models: We are a dedicated research lab for building, understanding, using, and risk-managing foundation... ...support data pipelines, experimentation, and evaluation workflows. ~ This role balances fast-...Visa sponsorship$70 per hour
...Role Overview Join a remote Machine Learning Engineer talent network to be considered for future contract engagements with AI labs and companies... .... Members contribute to advancing AI by training and evaluating models, designing real-world tasks and deliverables, and...Hourly payContract workFor contractorsRemote work$224k - $356.5k
We are looking for outstanding Machine Learning Engineers to join our Physical AI teams. As the pioneers... ...state-of-the-art multimodal models and diffusion techniques to simulate... ...Establish a strong mentality for KPI evaluation and validation to ensure the quality...Full time$204k - $259k
...The Driver Understanding and Evaluation (DUE) team at Waymo is... ...the Waymo Driver. The DUE Machine Learning team will build and operate... ...and advanced machine learning models to deliver training and evaluation... ...researchers and software engineers who are passionate about developing...Full time$170k - $216k
...15+ U.S. states. The DUE Machine Learning team will build and operate... ...tools, improve and speed up the evaluation and onboard developer... ...and advanced machine learning models to deliver training and evaluation... ...researchers and software engineers who are passionate about...Full time$238k - $302k
...Foundations team is to develop machine learning solutions addressing open... ...from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a... ...to a Senior Staff Software Engineer. You will: Work with...Full timeRemote work- ...providing Information Technology, Engineering Services, Program Management, and... ...Solerity is seeking Mid to Senior Machine Learning Engineers and AI Model Developers to support an upcoming... ...This effort focuses on developing, evaluating, and integrating machine learning...Full timeFor contractorsRemote workFlexible hours
$100 per hour
...knowledge to help train and evaluate next-generation AI systems by... ..., clarity, and relevance of model outputs through rubric-based... ...experience in Data Science, Machine Learning, Applied AI, Statistics,... ...desirable. Background in prompt engineering, AI output evaluation, fact...Hourly payContract workPart timeFor contractorsRemote work- ...Seattle, WA, we are a team of engineers and technologists from... ...Senior Software Engineer – Evaluation, you will design and implement... ...recognition (ASR), and small language model (SLM) systems. You will... ...You will work closely with machine learning and data engineering teams...
- ...customized by a team of experienced sellers, engineers, and researchers. Many of us worked on... ...execs Pioneer the training of new models that leverage both historical data and synthetic... ...You have a strong understanding of deep learning AI/ML frameworks or cloud services...Full time
$170k - $216k
...team builds the system which learns the spatial-temporal representation... ...set of sensors, enabling engineers like you to (1) develop... ...real-world data, to (2) develop models and model training at scale,... ...experience ~3+ years experience in Machine Learning and/or Computer...Full timeRemote work- ...build cutting-edge foundation AI models and end-to-end products that... ...is a team of researchers, engineers, designers, and more, who are... ...Paris. Join us!Why this role?Evaluation is critical to making progress... ...credit.Education & learning stipend for conferences, courses...Full timeWork at officeLocal areaRemote workHome office
$80 per hour
...Role Overview Evaluate and improve frontier AI coding agents by applying professional data engineering judgment to realistic data infrastructure and pipeline scenarios. You will use and assess model-generated implementations for ETL, data warehouses, analytics platforms...Hourly payRemote work$298k - $368k
...team builds the system which learns the spatial-temporal representation... ...set of sensors, enabling engineers like you to (1) develop... ...real-world data, to (2) develop models and model training at scale,... ...~7+ years of experience in Machine Learning, with a focus on large...Full timeRemote work$40 per hour
A biotechnology company is seeking a Biotechnology R&D Scientist to train AI models and evaluate their outputs. This role involves measuring the progress of AI chatbots with complex biology questions and ensuring their performance and correctness. Candidates should have...Hourly payRemote workFlexible hours- ...Prolific is seeking Biology Experts and Life Science Professionals to join an expert network that evaluates and trains AI models. This role involves reviewing AI-generated scientific content for accuracy and validation, requiring candidates with a BS, MS, or PhD in relevant...Remote workFlexible hours
$224k - $356.5k
...high-performance computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in... ...typically 12+ years) developing or assessing contemporary machine learning and deep learning systems.Hands-on experience with...Full time$272k - $431.25k
...re generating it! Our world model team is pushing the boundaries... ...Manager to lead world-model evaluation and benchmarking across... ...Strong research background in machine learning, computer vision, multimodal... ...Computer Science, Electrical Engineering, Robotics, Machine Learning,...Full time$40 per hour
A forward-thinking AI development firm seeks experienced quantitative professionals to evaluate AI-generated work, applying their skills in statistical analysis, predictive modeling, and technical writing. This fully remote opportunity offers a flexible schedule and projects...Hourly payRemote workFlexible hours$40 per hour
A leading AI development company is seeking experienced quantitative professionals to evaluate AI-generated analyses and design quantitative problems for AI training. This fully remote role offers flexibility in project selection and scheduling, with competitive pay starting...Hourly payRemote work$60 per hour
...professionals to help advance AI development. AI models are increasingly capable of performing... ...-of-the-art AI models on tasks like evaluating AI-generated quantitative analysis,... ...Statistics, Computer Science, Mathematics, Engineering, or similar); a master's or PhD is a...Hourly payFull timeRemote workFlexible hours$60 per hour
...A leading data analysis firm is seeking experienced quantitative professionals to join their remote team. In this role, you'll evaluate AI-generated quantitative analysis, design problem-solving tasks for AI training, and provide insightful feedback on AI systems. Candidates...Remote work$20 per hour
...contractors to join our team and teach AI chatbots. You will develop complex prompts to test AI models, write high-quality responses to demonstrate excellence, and evaluate different model outputs based on accuracy and style guidelines. This role is ideal for professionals...Hourly payFull timeContract workPart timeFor contractorsSelf employmentFreelanceRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Engineer for Model Evaluation. Be the first to apply!
- junior machine learning engineer United States
- graduate machine learning engineer United States
- senior ml engineer United States
- data scientist machine learning engineer United States
- machine learning engineer United States
- lead machine learning engineer United States
- machine learning ai engineer United States
- entry level machine learning engineer United States
- machine learning software engineer United States
- ai ml engineer United States



