ML Engineer - Model Evaluation
$60 - $90 per hourMercor
Job Description
Job Description
About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .
Position: Machine Learning Engineer — Model Evaluation & Experimentation
Type: Contract
Compensation: $60–$90/hour
Location: Remote
Commitment: 35 hours/week
Role Responsibilities
- Design tasks by turning real ML research ideas into well-defined, multi-step tasks.
- Run experiments by implementing changes, running training experiments, and analyzing results.
- Explore reinforcement learning ideas, focusing on reward functions and training behavior.
- Evaluate models to identify where frontier models fall short.
- Collaborate with researchers and experts to maintain task consistency and rigor.
Qualifications
Must-Have
- MSc or PhD in machine learning , computer science , or a related STEM field.
- 1+ years of experience in a research or research-engineering role.
- Experience in training and evaluating ML models and running experiments end-to-end.
- Strong familiarity with large language models and their evaluation techniques.
- Proficiency in Python and Git .
Preferred
- Basic understanding of reinforcement learning .
- Experience in AI training , model evaluation, or benchmark/task authoring.
Application Process (Takes 20–30 mins to complete)
- Upload resume
- AI interview based on your resume
- Submit form
Resources & Support
- For details about the interview process and platform information, please check:
- For any help or support, reach out to: View email address on us.fitly.work
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.
$175k - $215k
...learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. In this hybrid role, you will report to a... ...industrial or research setting developing recipes for ML models We prefer: - Track record of...SuggestedFull timeRemote work$228.7k - $343.1k
...at enormous scale, and one bad model can mean millions in credit... ...validate at scale, so you critically evaluate what it produces and own the... ...in parallel. Reason about ML systems end to end — how... ...accuracy. Solid software and data engineering: production-quality Python,...SuggestedRemote jobFull timeLocal areaShift work- ...company. We build cutting-edge foundation AI models and end-to-end products that are... ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all... ...Germany and Paris. Join us!Why this role?Evaluation is critical to making progress in scaling...SuggestedFull timeWork at officeLocal areaRemote workHome office
- ...for an Insurance Subject‑Matter Expert to join a leading AI lab's GenAI team in New York. This W-2 role involves evaluating insurance tasks and guiding model development for high‑quality underwriting judgments, with placement at the client lab as part of the extended...SuggestedWeekday work
- Alignerr is seeking a Senior Python Infrastructure Engineer to design and build data pipelines, evaluation harnesses, and annotation tooling powering AI systems... ...role focuses on production-grade Python systems for model evaluation at scale. You will collaborate with research...SuggestedRemote jobHourly payContract work
$400 per month
...lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic data engineering workflows and model evaluation. Spots are limited and filling quickly...- We are seeking an expert to evaluate and improve our AI models through comprehensive testing and analysis. You will be responsible for designing... ...outputs for bias, fairness, and accuracy Collaborate with ML engineers to implement improvements Document findings and...
- Senior Research Scientist, Model Evaluation Cohere | Posted Mar 2 | Full-time | New York | Negotiable | Unknown Why this role? Evaluation... ...the capabilities you care about. You have strong software engineering skills. We value and celebrate diversity and strive to create...Full timeWork at officeRemote workFlexible hours
- Mercor is seeking a Generalist who can operate in English and Punjabi. This contract, remote position focuses on evaluating AI outputs and supporting model evaluation tasks. You will conduct fact-checking, assess reasoning, clarity, tone and completeness, and provide actionable...Remote jobContract work
- Dorado is seeking an experienced Investment Banking SME to support the development, evaluation, and improvement of advanced AI models in finance. You will assess AI-generated analyses for accuracy, reasoned judgments, and data integrity, guiding model refinements and prompts...Remote job
- Cohere is seeking a Senior Research Engineer, Model Evaluation, to create next‑generation evaluation methods and scalable infrastructure. You will develop benchmarks, datasets, and environments to measure frontier model capabilities, and you will push the state‑of‑the‑...
$175k - $280k
...New York is seeking an expert in optimizing machine learning models to turbocharge their serving layer, integrating LLM, speech, and... ...significant experience in systems programming and performance engineering, aiming to improve high-throughput, low-latency serving. Join...- Dorado is seeking a Physics Specialist to contribute deep scientific expertise to AI model evaluation. You will craft and assess challenging physics problems, probe model reasoning at the frontier, and help identify where models fail under rigorous scientific scrutiny....Remote job
- ...opportunities through their expert network. Qualified candidates should hold a MS or PhD in a relevant field and have experience in evaluating complex biology content. Strong communication skills and proficient English are essential for success in this role. #J-18808-...Immediate start
- Mercor is seeking experienced Musicians to evaluate generative musical AI models in partnership with a leading AI lab. In this role, you will assess... ...Ideal candidates have 3+ years as a music producer/audio engineer, a college degree in music, and native or near-native...
- ...seeking Insurance domain SMEs to join a cutting-edge AI training program. You will evaluate AI model outputs against real underwriting practice and rubrics, guiding research and engineering teams to close knowledge gaps in underwriting, claims, and risk assessment. The...
- ...company. We build cutting‑edge foundation AI models and end ‑to‑end products that are... ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all... ...Germany and Paris. Join us! Why this role? Evaluation is critical to making progress in scaling...Full timeWork at officeLocal areaRemote workHome office
- Take a new model and get it running — correctly — on our ASIC in record time. When a frontier... ...we're looking for 5+ years in systems or ML systems, with real depth in at least one... ...or LLM-driven tooling that did real engineering work, not demos. #J-18808-Ljbffr General...Live in
$152k - $241.5k
...working for us! We believe open-weight models are foundational to American AI... ...scientific scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find,... .... We are looking for an Evaluation/ML-Systems Engineer to own how we...Full timeRemote work- SME Careers is seeking a remote Kotlin Engineer to review AI-generated responses and create... ...optimizing AI performance, and ensuring model accuracy. The ideal candidate has a... ...an expert network and requires critical evaluation of technical concepts. #J-18808-Ljbffr...Remote job
$85 per hour
...Jack Dorsey . Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience) Type: Contract... ...Use frontier AI coding agents to complete and evaluate complex infrastructure engineering tasks. Review model-generated implementations involving cloud platforms...Contract workSummer workRemote work$60 - $90 per hour
...Help shape how AI systems evaluate and preserve the timing, emotion, micro-expressions, and physical nuance of human and character performances... ...-time role supports the development of performance-transfer models by defining high-quality evaluation standards and helping...Hourly payPart time- Mercor partners with a leading AI research lab to support a Frontier Code Agents project, focusing on realistic data engineering workflows and model evaluation. Contributors help evaluate and improve frontier AI coding models through structured technical assessments and...
- SME Careers is looking for a remote R Engineer to review AI-generated responses and create high-quality R and data-analysis content. The position requires a strong background in R programming, applied statistics, and excellent writing skills to document analyses. Candidates...Remote job
- ...building extraction agents, evaluating accuracy, deploying to production... ...architecture changes, prompt engineering, fine-tuning, or rule-based... ...~2+ years building ML/AI systems in production ~... ...infrastructure glue, not just model training scripts ~ Practical...Full time
$213k - $263k
...are currently focusing on include reinforcement learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role follows a hybrid work schedule and reports to a Principal Research Scientist....Full timeTemporary workRemote work$100k - $250k
...This innovative field blends AI, engineering, and materials science,... .... The opportunity As an ML Engineer at Radical AI, you will... ...also data ingestion, enrichment, model exporting and serving, and... ...infrastructure. Continually evaluate model performance and maintain...Full time$20 per hour
Feedinkoo is looking for a Web Developer/Designer to enhance AI models by evaluating design work, including interfaces and user experiences. This role involves reviewing AI‑generated visuals and providing feedback to improve users' experience with AI tools. Working remotely...Remote job$80 per hour
...Summers , and Jack Dorsey . Position: Data Engineer (Coding Agent Experience) Type: Contract Compensation... ...Use frontier AI coding agents to complete and evaluate complex data engineering tasks. Review model-generated implementations involving ETL pipelines...Contract workSummer workRemote work- ...workflow orchestration layer for computer vision models. Job Description We are seeking a Staff ML Engineer with a passion for building application-layer AI... ...OpenAI, along with pgvector Langchain and evaluation frameworks in Langsmith Google Cloud, Docker...Full timeWork at officeWork from homeFlexible hours3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Engineer - Model Evaluation. Be the first to apply!
- data scientist machine learning engineer New York, NY
- machine learning ai engineer New York, NY
- computer vision machine learning engineer New York, NY
- machine learning engineer New York, NY
- ai ml engineer New York, NY
- machine learning software engineer New York, NY
- entry level machine learning engineer New York, NY
- junior machine learning research engineer New York, NY
- senior ml engineer New York, NY
- internship machine learning New York, NY


