ML Engineer - Model Evaluation
$60 - $90 per hourMercor
Job Description
Job Description
About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .
Position: Machine Learning Engineer — Model Evaluation & Experimentation
Type: Contract
Compensation: $60–$90/hour
Location: Remote
Commitment: 35 hours/week
Role Responsibilities
- Design tasks by turning real ML research ideas into well-defined, multi-step tasks.
- Run experiments by implementing changes, running training experiments, and analyzing results.
- Explore reinforcement learning ideas, focusing on reward functions and training behavior.
- Evaluate models to identify where frontier models fall short.
- Collaborate with researchers and experts to maintain task consistency and rigor.
Qualifications
Must-Have
- MSc or PhD in machine learning , computer science , or a related STEM field.
- 1+ years of experience in a research or research-engineering role.
- Experience in training and evaluating ML models and running experiments end-to-end.
- Strong familiarity with large language models and their evaluation techniques.
- Proficiency in Python and Git .
Preferred
- Basic understanding of reinforcement learning .
- Experience in AI training , model evaluation, or benchmark/task authoring.
Application Process (Takes 20–30 mins to complete)
- Upload resume
- AI interview based on your resume
- Submit form
Resources & Support
- For details about the interview process and platform information, please check:
- For any help or support, reach out to: View email address on ziprecruiter.com
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
- ...company. We build cutting-edge foundation AI models and end-to-end products that are... ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all... ...Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling...SuggestedFull timeWork at officeLocal areaRemote workHome office
- ...for an Insurance Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting judgments, with placement at the client lab as part of the extended...SuggestedWeekday work
$80 per hour
...Summers , and Jack Dorsey . Position: Data Engineer (Coding Agent Experience) Type: Contract Compensation... ...Use frontier AI coding agents to complete and evaluate complex data engineering tasks. Review model-generated implementations involving ETL pipelines...SuggestedContract workSummer workRemote work- Mercor is seeking a Generalist who can operate in English and Punjabi. This contract, remote position focuses on evaluating AI outputs and supporting model evaluation tasks. You will conduct fact-checking, assess reasoning, clarity, tone and completeness, and provide actionable...SuggestedRemote jobContract work
- Senior Research Scientist, Model Evaluation Cohere | Posted Mar 2 | Full-time | New York | Negotiable | Unknown Why this role? Evaluation... ...the capabilities you care about. You have strong software engineering skills. We value and celebrate diversity and strive to create...SuggestedFull timeWork at officeRemote workFlexible hours
- Cincinnatus LLC is hiring for a Marketing SME to support GenAI model evaluation and brand/growth tasks. The role centers on applying rigorous marketing judgment to AI training data and guiding cross-functional teams to improve model outputs. Ideal candidates have extensive...
- Dorado is seeking an experienced Investment Banking SME to support the development, evaluation, and improvement of advanced AI models in finance. You will assess AI-generated analyses for accuracy, reasoned judgments, and data integrity, guiding model refinements and prompts...Remote job
- ...is seeking a Generalist to perform English and Tamil language evaluation tasks on a remote contract basis. You will conduct rigorous fact... ...data that highlights strengths, weaknesses, and inaccuracies in model responses. The role requires strong English writing, Tamil...Remote jobContract work
- Cohere is seeking a Senior Research Engineer, Model Evaluation, to create next‑generation evaluation methods and scalable infrastructure. You will develop benchmarks, datasets, and environments to measure frontier model capabilities, and you will push the state‑of‑the‑...
- Screenshot Understanding Model Evaluator is a remote evaluation track for reviewing screenshot understanding model evaluation prompts and responses against AuraOne's quality rubric. Reviewers compare paired outputs, label edge cases, and write the kind of structured feedback...Hourly payFor contractorsRemote work10 hours per week
$175k - $280k
...New York is seeking an expert in optimizing machine learning models to turbocharge their serving layer, integrating LLM, speech, and... ...significant experience in systems programming and performance engineering, aiming to improve high-throughput, low-latency serving. Join...- ...seeking Insurance domain SMEs to join a cutting-edge AI training program. You will evaluate AI model outputs against real underwriting practice and rubrics, guiding research and engineering teams to close knowledge gaps in underwriting, claims, and risk assessment. The...
- Dorado is seeking a Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems, probe model reasoning at the frontier, and help identify where models fail under rigorous scientific scrutiny....Remote job
- Mercor is seeking experienced Musicians to evaluate generative musical AI models in partnership with a leading AI lab. In this role, you will assess... ...Ideal candidates have 3+ years as a music producer/audio engineer, a college degree in music, and native or near-native...
- ...company. We build cutting‑edge foundation AI models and end‑to‑end products that are... ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all... ...Germany and Paris. Join us! Why this role? Evaluation is critical to making progress in scaling...Full timeWork at officeLocal areaRemote workHome office
$40 per hour
Feedinkoo is seeking a Frontend Developer to join our team in the United States. The role involves training AI models by evaluating their logic and solving problems to enhance their quality. This independent contract position allows for flexible scheduling and offers hourly...Remote jobHourly payContract workFlexible hours- SME Careers is seeking a remote Kotlin Engineer to review AI-generated responses and create... ...optimizing AI performance, and ensuring model accuracy. The ideal candidate has a... ...an expert network and requires critical evaluation of technical concepts. #J-18808-Ljbffr...Remote job
- Obsidian is seeking contributors to evaluate frontier AI coding models through structured technical assessments. You will work on realistic ML engineering workflows, model evaluation, and production-ready deployment considerations. The role emphasizes hands-on use of frontier...
$85 per hour
...Jack Dorsey . Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience) Type: Contract... ...Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving cloud platforms...Contract workSummer workRemote work$60 - $90 per hour
...Role Overview Help shape how advanced performance-transfer models preserve an actor’s timing, emotion, and subtle physical choices... ...character-animation expertise to define quality standards, strengthen evaluation methods, and guide the reviewers who assess model performance....Hourly payPart time$40 per hour
Feedinkoo is seeking a Software Developer to join their remote team focused on training AI models. The role involves providing coding challenges to AI chatbots and evaluating their outputs for quality and performance. Applicants should be fluent in English and detail-oriented...Remote jobHourly payFor contractorsFlexible hours- SME Careers is looking for a remote R Engineer to review AI-generated responses and create high-quality R and data-analysis content. The position requires a strong background in R programming, applied statistics, and excellent writing skills to document analyses. Candidates...Remote job
- Block Market-Based PayBlock takes a market-based approach to pay, and pay may vary depending on your location. U.S. locations are categorized into one of four zones based on a cost of labor index for that geographic area. The successful candidate's starting pay will be...
$20 per hour
Feedinkoo is looking for a Web Developer/Designer to enhance AI models by evaluating design work, including interfaces and user experiences. This role involves reviewing AI‑generated visuals and providing feedback to improve users' experience with AI tools. Working remotely...Remote job- ...providing independent assurance and evaluating the company's risk management... .... You will be deploying your engineering, data analytics and data... ...in frameworks for auditing models, including criteria like robustness... ...with LLMs and traditional ML models and at least 5 years...
$184.9k - $250.2k
We are seeking a Machine Learning Engineer to work directly alongside our research scientists to train, evaluate, and deploy the models that make our robots move, perceive, and act in the real world. This is a hands-on ML role: you will train policies, debug convergence...InternshipFlexible hours$77k - $202k
...Bachelor's Degree- At least 4 years of experience in software engineering or data engineeringWhat Sets You Apart- Master's Degree in Computer... ...assets, or collaborating closely with team members. We evaluate these factors thoughtfully to establish a secure and trusted workplace...Full timeH1b- Mindrift, powered by Toloka, is launching a Management Consulting domain to translate real-world consulting engagements into structured learning environments for advanced AI systems. We are assembling a team of strategy consultants from top-tier firms who can convert authentic...Remote job
- In-House Fit Model (Women's Size 8 / Medium)About the RoleWe are looking for a reliable and detail-oriented Women's Fit Model to support... ..., technical design, merchandising, and sales teams to help evaluate garment fit, comfort, and overall wearability. This position plays...Full time
- ...of applied intelligence from model optimization to productized AI... ...and deploy production‑grade ML systems with end‑to‑end ownership... ...experience in ML engineering. Strong programming skills in... ...pipelines for model training and evaluation. Familiarity with FastAPI, OpenAI...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Engineer - Model Evaluation. Be the first to apply!
- data scientist machine learning engineer New York, NY
- machine learning ai engineer New York, NY
- computer vision machine learning engineer New York, NY
- machine learning engineer New York, NY
- ai ml engineer New York, NY
- machine learning software engineer New York, NY
- junior machine learning engineer New York, NY
- junior machine learning research engineer New York, NY
- senior ml engineer New York, NY
- internship machine learning New York, NY


