ML Evaluation Specialist, Human Data
$144.6k - $263.8kApple Inc.
Cupertino, California, United States Machine Learning and AI At Apple, we don’t just build products — we build experiences fueled by world-class data. The Human-centered AI team within Apple Services Engineering is looking for an ML Evaluation Specialist, Human Data to join our Data Quality and Operations division to spearhead complex, multi-stakeholder operations that specialize in data collection, curation, annotation, and human evaluation efforts across Apple Music, App Store, TV+, Podcasts, and Books. In this role, you will own the operational strategy and continuous improvement of large-scale, multilingual human data programs, from designing onboarding scaffolds that progressively build annotator calibration, to analyzing annotator behavior patterns to identify where automation can offload low‑judgment decisions, to enforcing quality frameworks that close the loop between annotator struggle and task redesign. You will identify where human judgment is essential and where it could be better directed, then build the scaffolding, automation, and feedback systems that let annotators focus their cognitive energy where it matters most. Because this work cuts across engineering, data science, research, procurement, and legal, a critical part of the role is serving as the connective tissue between teams who each own a piece of this space, aligning on shared standards, surfacing gaps, and ensuring that insights from the annotation layer inform upstream decisions about task design and tooling. You will bring a point of view on human data best practices and translate it into scalable, human‑centered approaches that make generative AI features safer and more reliable. The ideal candidate brings a rare combination of technical depth and program execution skills. You are comfortable designing and deploying sophisticated data pipelines in the morning, and then seamlessly transitioning to present comprehensive quality rectification strategies to stakeholders in the afternoon. You care deeply about data quality and human alignment, have a creative and systematic approach to finding and fixing problems, and find motivation in wide‑ranging work whose impact shows up in everyday Apple experiences. Responsibilities Lead the end‑to‑end execution of human data collection programs for multilingual, multimodal, and multi‑turn AI features, from intake and scoping, to delivery and retrospective Estimate and maintain project timelines, capacity needs, and cost while anticipating and proactively resolving bottlenecks to ensure timely execution Design and own the measurement framework for human data collection initiatives, defining key indicators such as spend, speed, inter‑rater reliability, and volume, and building reporting systems that surface actionable insights to stakeholders Build and implement human data quality frameworks, including developing statistical process controls and behavioral signals to proactively detect and remediate quality degradation throughout the annotation lifecycle Apply human‑centered AI principles to influence data collection task design, identifying and reducing sources of cognitive burden that impact human performance, consistency, and wellbeing Systematically analyze data collection workflows to identify inefficiencies and scalability gaps, then design and implement innovative workflows (e.g. agent‑ and machine‑in‑the‑loop) and build the supporting data pipelines and ETL services to improve data quality and diversity while optimizing cost and lead time Partner with Legal, Privacy, and New Product Security to design and implement compliant data collection and user study programs, ensuring all human data workflows adhere to Apple’s privacy, governance, and regulatory standards Design and own the data collection lifecycle with external and internal workforces, including building onboarding and calibration programs, performance frameworks, and quality audits that ensure reliable, high‑quality deliverables Act as the connective tissue across engineering, product, legal, security, procurement, and vendor teams, ensuring alignment, clear communication, and follow‑through Minimum Qualifications Bachelor’s degree or higher in Cognitive Science, Linguistics, or a related field that includes an experimental or empirical component 4+ years of experience defining and leading cross‑team human data programs for AI/ML, including annotation operations, quality frameworks, and evaluation strategies, within an NLP/NLU or generative AI environment Proficiency in programming and data languages (Python, R, SQL) to process, analyze, query large datasets, extract insights, automate tasks, and monitor program performance Hands‑on experience designing and managing 0→1 human‑in‑the‑loop data collection, annotation, and evaluation initiatives, including driving and incorporating agentic workflows to improve quality and scalability Experience working with diverse data types (speech, text, multimodal) across multiple languages Expertise in end‑to‑end data annotation quality management, including the ability to develop statistical process controls and data quality metrics Familiarity with privacy‑preserving data handling practices and compliance frameworks Demonstrated success optimizing data pipelines and workflows to improve quality, reduce lead time, and scale operations Experience working cross‑functionally with engineering, data science, legal, privacy, and third‑party suppliers Preferred Qualifications Master’s degree or higher in Cognitive Science, Linguistics, or a related field that includes an experimental or empirical component 2+ years of experience owning data strategy for frontier AI development and evaluation, with experience in human alignment methodologies and agentic GenAI systems Experience managing external vendor or workforce partners at scale Familiarity with AI Safety and Responsible AI principles, including experience applying them to data collection or annotation workflows Strong organizational skills and execution‑oriented mindset; ability to balance attention to detail with big‑picture thinking in an environment where program scope and priorities evolve quickly Excellent written and verbal communication skills; able to translate technical concepts for non‑technical stakeholders At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $144,600 and $263,800, and your base pay will depend on your skills, qualifications, experience, and location. Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits. Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program. Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant. At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong. Learn about accessibility in Apple’s workplace. Learn about reasonable accommodations for job applicants. Apple accepts applications to this posting on an ongoing basis. #J-18808-Ljbffr Apple Inc.
- Apple Inc. is seeking a Machine Learning Evaluation Specialist to lead human data collection for AI features, ensuring high-quality data and effective collaboration across teams. This role requires technical expertise and program management skills to enhance data processes...Data
- ...focused on recovering accurate 3D human body and hand motion from... ...video. Human demonstration data is the fuel for robot learning... ...annotation tooling, model training and evaluation, and production deployment at... ...-world, production-oriented ML pipelines. ~ Strong problem-...DataFull time
$124k - $137.89k
...organisation. We combine world-class human craft with next generation... ...-edge media intelligence and data solutions, world-class... ...What does an API Integration Specialist do at WPP Production?The API... ...partners and internal teams to evaluate integration approaches and refine...DataWork at office3 days per week$224k - $356.5k
...seeking a highly analytical Product Evaluations Lead to own the evaluation,... ...the intersection of rigorous data science, GenAI product... ...with a focus on measuring AI/ML or GenAI systems. Master's or... ...quality evaluation datasets and human-evaluation strategies.Ability...DataFull time$101.6k - $177.8k
...seeking a Go-To-Market (GTM) Specialist to define, build and lead the... ...and operationalize GenAI and ML workloads with specific focus... ...customer engagements and through Geo Data & AI specialists, to prospect... ...workshops/ PoVs/ PoCs for evaluation. You will solicit and help current...DataLocal areaWorldwideFlexible hours$207k - $300k
...implementation of solutions in specialized ML areas, optimize ML... ...of model optimization and data processing strategies.Minimum... ...duplicating and responding to the human voice), reinforcement learning... ...e.g., model deployment, model evaluation, data processing, debugging, fine...Data$162.7k - $220.2k
...provider of choice for customers? Join the Data & AI team as a Go-To-Market Specialist focused on Search and AI!This is an... ...in search technologies and AI/ML to create best-in-class field enablement... ..., finance, engineering, human resources, or related field, or PMP...DataLocal areaWorldwideFlexible hoursDay shift- ...computing experiences—from AI and data centers, to PCs, gaming and... ...comes from bold ideas, human ingenuity and a shared passion... ...career. THE ROLEWe are hiring AI / ML Platform Engineers to build the... ...distributed inference, batch evaluation, and large-scale agent rollout...Data
$130k - $220k
Santa Clara, CAData Engineering - ML Infrastructure /Full-time /HybridFinding the right data is central to improving... ...valuable moments for training and evaluation, then turn those methods into reliable... ...team but do not replace human judgment. Final hiring decisions...DataFull time$176.6k - $239k
...generative AI (GenAI)? AWS Worldwide Specialists Org (WWSO) is responsible for... ...) on AWS, model performance evaluations, develop demos and proof-of-... ...(Amazon EC2, Lustre), ML frameworks PyTorch, JAX, orchestration... ..., security, networking, data & analytics) experience- 3+...DataLocal areaWorldwideFlexible hours- ...surgery smarter, safer, and more human. Every day, our work helps... ...scientists, health economists, and data scientists.HFRE conducts... ...in Python or C++ required for ML performance modelingDevelop Snowflake... ...sources from research and evaluations to inform product design...DataWork at officeLocal areaWorldwideFlexible hours
$189.4k - $300.6k
...teams are redefining mobility. Through a human-centered design process, we create... ...behavior across real-world scenarios.The Evaluation Foundations team—part of Embodied AI’s... ...top-performing models and partner with data-intensive ML teams to drive rapid innovation.In this...DataFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours- ...design and implementation of automated model evaluation: standing up LLM-as-a-judge in existing... ...model regressions before they reach human evaluation or the live on population. This... ...practical knowledge of Python, including data-pipeline fluency (JSON/YAML, REST APIs).Hands...Data
$144.7k - $221.4k
...experiences. About the Organization The Evaluation team builds and evolves the... ...approaches that enable data-driven decisions across AV development... ...develop new statistical and ML methods to quantify... ...validation efforts, integrating human-in-the-loop review where needed...DataFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$151k - $204.3k
...leveraging the most comprehensive GenAI/ML platform available? Come join us!At Amazon... ...agent frameworks, prompt engineering, model evaluation, fine-tuning) is required. A strong... ...responsibilitiesWorking with customers' development, data science, and AI engineering teams to...DataLocal areaWorldwideFlexible hoursDay shift- ...revolutionize the future of data centers. nEye’s MEMS-based silicon... ...-oriented Purchasing Specialist to join our Operations team.... ...supplier selection. Request and evaluate supplier quotations for pricing... ...recruitment team but do not replace human judgment. Final hiring...Data
$153.6k - $207.8k
...sales expertise to help customers evaluate, design, and adopt AWS as... ...platform for Generative AI and ML workloads? Do you enjoy building... ...solutions? Join the GenAI/ML Specialists Solution Architect team as a... ...architects, product leaders, data scientists, and executives at...DataLocal areaWorldwideFlexible hours- ...enterprise. Design-in risk management, evaluation, compliance, and other... ...across four pillars: Data Substrate (lineage at ingest),... ..., and Activation (risk-tiered human-in-the-loop). This role exists... ...Experience building or operating ML/LLM evaluation infrastructure...DataFull timeWorldwideShift work
- .... About the Organization: The Evaluation team builds and evolves the evaluation... ...approaches that enable data-driven decisions across AV... ...validation efforts, integrating human-in-the-loop whereappropriate.... ...+yearsapplied experience indata analysis, ML evaluation, or autonomy...DataLocal areaWork from home
$147.4k - $272.1k
ML Engineer - Automated Evaluation and Adversarial Design Cupertino, California, United States Software and Services... ...alignment between automated and human evaluation methods on an ongoing... ...leveraging automation to scale evaluation data generation and analysis Experience...DataRelocationShift work$170k - $216k
...scalable machine learning and data systems, simulation workflow and... ...tools, improve and speed up the evaluation and onboard developer journeys. It will combine expert human judgements and advanced machine... ...systems covering the ML lifecycle, supporting planet-scale...DataFull time$83.35k - $101.62k
...EVALUATIONS SPECIALIST, SENIOR San Jose/Evergreen Community College District Close/First Review... ...calculate awarded degree and certificate data and statistical information as... ...excellence; fosters cultural, racial and human understanding; provides positive roles...DataFull timeWork at officeMonday to FridayFlexible hours$160k - $250k
...The Role: CrowdStrike's Data Platform team is... ...security researchers, AI/ML teams, and customer-facing... ...lakehouse architectures Evaluate and integrate emerging... ...ontology/data modeling specialists Foster a team culture of... ...or activity in a local human rights commission, status...DataWork experience placementWork at officeLocal area- Apple Inc. is seeking an expert in evaluating machine learning and deep learning models, including foundation models... ...and a deep understanding of statistical methods, data quality, and model robustness, collaborating with ML engineers, data scientists, and ML infrastructure...Data
$150.4k - $277.6k
...Engineer - Multimodal for Human Understanding Sunnyvale... ..., implementing, and evaluating novel algorithms and models... ...and productizing ML systems capable of human... ...engineers, software engineers, data scientists, human-... ...designers, and domain specialists—working in an...DataWorldwideRelocation$184.7k - $324.8k
.../ Engineer, Foundation Model Evaluation Cupertino, California, United... ...with model training, training data, and product teams to ensure evaluation... ...in Python and experience with ML frameworks (PyTorch, JAX, or... ...teams Familiarity with human evaluation methodology and experience...DataRelocation$57.47 - $60.37 per hour
...Per Diem Imaging Specialist El Camino Health Medical Network (ECHMN... ...and maintains follow-up data in the system, facilitating accurate... ...policies, including human resources, safety, HIPAA, and... ...analyzing business opportunities and evaluating ROI. Proficient in...DataDaily paidWork at officeFlexible hoursDay shift$238k - $302k
...across 15+ U.S. states. The Waymo ML Frameworks & Efficiency team... ..., including training and evaluation. They are geared towards both... ...that can scale across compute, data, and environments to improve model... ...and alignment with human drivers. You Will:...DataFull timeRemote work$106.4k - $195.1k
...space. About the Role As a Sr. FinOps Specialist at Workday, you will be the connective... ...wide initiatives, transforming raw cloud data into the operational excellence that... ...to achieve a desired business outcome by evaluating, diagnosing, and evolving existing business...DataFull timeWork experience placementWork at officeRemote workHome officeFlexible hours$152k - $241.5k
...Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find,... ...evidence first. We are looking for an Evaluation/ML-Systems Engineer to own how we measure the... ..., including experiment tracking and data pipelines.Ways to Stand Out from the Crowd...DataFull timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Evaluation Specialist, Human Data. Be the first to apply!
- junk removal specialist Cupertino, CA
- continuous improvement specialist Cupertino, CA
- loss prevention specialist Cupertino, CA
- hospitality specialist Cupertino, CA
- fabrication specialist Cupertino, CA
- process specialist Cupertino, CA
- credit specialist Cupertino, CA
- health specialist Cupertino, CA
- reporting specialist Cupertino, CA
- esports specialist Cupertino, CA

