Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
$184.7k - $324.8kApple Inc.
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Santa Clara, California, United States Machine Learning and AI
We are the Foundation Model Inference team within Cloud OS and AI Inference organization. We are on a mission to build the most highly performant, secure and private inference stack that powers Siri AI, Apple Intelligence and Apps that are powered with the largest foundation models.Our systems serve billions of queries daily across Siri AI, Apple Intelligence, Apple Search, Apple Music, Apple TV, App Store, iMessage, Photos, Camera, Spotlight & Safari, at remarkably low latency with every ounce of compute extracted from the hardware beneath them. We optimize language, vision, and speech models with billions of parameters using state-of-the-art techniques and ship them at Apple scale.This is a rare opportunity to directly shape how AI reaches billions of people worldwide.
Description
You will work at the intersection of research and production, partnering closely with the Foundation Model Research team and our external partners to bring cutting-edge model architectures from prototype to planetary-scale deployment. You will own hard problems in inference efficiency, hardware/software codesign, systems architecture, and tooling, and help set the technical direction for the engineers around you.This role sits within CloudOS and Private Cloud Compute (PCC) - Apple's purpose-built, privacy-preserving cloud infrastructure for AI workloads. PCC represents a first-of-its-kind approach to running foundation models in the cloud with verifiable privacy guarantees, and CloudOS is the systems foundation that makes it possible. You will be building and optimizing inference systems on top of this infrastructure, working closely with platform and security teams to deliver both performance and trust at scale.
Responsibilities
- Partner with the Foundation Model Research team and our external partners to optimize inference for the latest model architectures across language, vision, and speech.
- Design and ship production-grade inference systems serving millions of customers in real time.
- Build profiling tools and simulators to identify and resolve performance bottlenecks across different hardware configurations and use cases.
- Drive technical decisions on high-throughput, low-latency serving at supercomputing scale.
- Mentor and grow engineers across the organization.
Minimum Qualifications
- 5+ years of experience leading complex, ambiguous technical projects from end to end.
- Hands-on experience with LLM inference stacks.
- Working knowledge of GPU or TPU programming concepts.
- Proficiency with PyTorch, JAX, or TensorFlow.
- Experience building and operating high-throughput services at large distributed scale.
- Proficiency deploying applications on cloud platforms (AWS, GCP, or equivalent) using Kubernetes and Docker.
- BS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field.
Preferred Qualifications
- Experience building productions systems in Go or Python.
- Strong knowledge of deep learning architectures including Transformers, encoder/decoder models, and multimodal variants.
- Experience with inference optimization frameworks such as TensorRT-LLM, vLLM, SGLang, TGI, or Nvidia Triton Server.
- Experience authoring custom CUDA kernels using CUDA C++ or OpenAI Triton.
- MS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field.
At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $184,700 and $324,800, and your base pay will depend on your skills, qualifications, experience, and location.
Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits
Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.
Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant
At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.
Learn about accessibility in Apple’s workplace
Learn about reasonable accommodations for job applicants
Apple accepts applications to this posting on an ongoing basis.
#J-18808-Ljbffr- ...Apple Inc. is seeking a Sr. Machine Learning Engineer for the Foundation Models Inference team in Santa Clara, CA. You will collaborate with research and external partners to bring cutting-edge model architectures from prototype to planetary-scale deployment, owning hard...CloudSeniorFoundation
- ...leading training and inference speeds; over 10... ...based hyperscale cloud inference... ...with the leading model labs, global enterprises... ...state-of-the-art foundation models and... ...Cerebras' Wafer-Scale Engine (WSE). We build the... ...intersection of machine learning frameworks, compiler...CloudSeniorFoundation
- ...We are Foundation Model Inference Team, within AI, Search & Knowledge Platform... ...use cases.Mentor and guide engineers in the organization. Minimum... ...running applications on Cloud (AWS / Azure or equivalent... ...Artificial Intelligence, Machine Learning, Information Retrieval,...CloudSeniorFoundation
- ...Cerebras Systems is seeking an engineering leader to build and scale the Inference Model Scaling organization. You will define... ...team enabling latest foundation models on Cerebras hardware, leading... ...Work across compiler, runtime, cloud, hardware, product management, and...CloudFoundation
$229.9k - $262.4k
...Overview Sr. Lead AI Engineer (Inference Optimization, FM hosting,... ...industry leader in using machine learning to create real-time... ...The Intelligent Foundations and Experiences (... ...customers. Our AI models and platforms empower... ...AI solutions on cloud platforms (e.g. AWS...CloudSeniorFoundationFull timePart timeLocal area$193.3k - $261.5k
...software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators. Join us to optimize the latest models to run really fast on the Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you will onboard and...CloudSeniorInternshipLocal areaFlexible hours- ...specific focus on large-scale model inference on AMD Instinct™ and Radeon™... ...it. You will influence engineering roadmaps, represent AMD in open... ...relationships with key OSS maintainers, foundation working groups, and... ...Hugging Face, Red Hat, and cloud providers on joint...CloudSeniorFoundationRemote work
$117.7k - $221.4k
...sits at the intersection of machine learning, data infrastructure, and... ...not only on stronger models, but also on better infrastructure... ...model reflects how Cola engineers think: build durable... ...processing, featurization, and inference foundations that power scalable world...FoundationFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours- ...leading training and inference speeds and allows... ...include top model labs, global enterprises... ...of product, engineering, sales, and marketing... ...semiconductor, HPC, cloud infrastructure, or... ...language models, foundation model training, and... ...that supports learning, growth, and collaboration...CloudSeniorFoundation
$184k - $287.5k
...network stack for distributed inference which will be used to orchestrate wide area networks/cloud Networks. Enabling a workload... ...Computer Science, Electrical Engineering, Software Engineer, or... ...orchestration frameworks with strong foundation Network topology and...CloudSeniorFoundationFull timeWork experience placement$192k - $279k
...Product Manager, TPU Inference Location:... ...Strong business foundation with the ability to... ...closely with creative engineers, designers, marketers... ...is to help Google Cloud customers, from AI... ...serve their AI/ML models efficiently and at... ...equity + benefits Learn more about...CloudSeniorFoundationTemporary workWorldwideShift work- ...intelligence to every moving machine on the planet.... ...Seoul; and Tokyo. Learn more at applied.co... ...for a performance engineer who specializes in... ...-throughput batch inference sweeping petabytes... ...and performance models for our workloads,... ...machine learning foundations, and the ability...FoundationFull timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift
$229k - $343k
...themselves, live in the moment, learn about the world, and have... ...generative AI, including foundational models, efficient infrastructure,... ...on-device and server-side inference. Our team creates... ...Spectacles.We’re looking for a Machine Learning Engineer to join Snap Inc!What you’...Full timeLive inWork at officeLocal areaWorldwide$195.2k - $361.2k
...combines the best of local and cloud intelligence — private,... .... Small, efficient models run directly on the user's machine (AI PC, edge, on-prem,... ...actually own. You optimize inference engines (llama.cpp, vLLM) for... ...it helps usWhat you’ll learn / grow intoCuriosity is...CloudSeniorFull timeInternshipLocal areaImmediate startShift work- Databricks in Mountain View is seeking an Engineering Manager to lead the Foundation Model Inference (FMAPI) organization. You will build and guide teams responsible for large-scale inference workloads, partnering with product to deliver reliable, scalable AI infrastructure...Foundation
- ...customer commitment, model launch, and infrastructure... ...strategy for our Inference Service organization.... ...working directly with Engineering, Product, Infrastructure... ...operations experience in cloud infrastructure, large‑... ...through continuous learning, growth and support of...CloudSenior
$272k - $431.25k
...and FlashDreams are the foundation for a new generation of interactive world-model systems. With this... ...for fidelity, real-time inference performance in world models... ...source inferencing engine. We are on a mission to... ...passionate about developing cloud services we want to...CloudSeniorFoundationFull time$230k - $265k
...veteran scientists and engineers. As a Senior Machine Learning Engineer, you’ll... ...post-training, and inference strategies for... ...language and speech models using PyTorch and/or... ...deployment workflows in a cloud environment.... ...large language or foundation models, with production...CloudSeniorFoundationPermanent employment- ...combines the best of local and cloud intelligence - private,... .... Small, efficient models run directly on the user's machine (AI PC, edge, on-prem,... ...actually own. You optimize inference engines (llama.cpp, vLLM) for... ...it helps us What you’ll learn / grow into The internals...CloudSeniorLocal areaShift work
$159.5k - $236.5k
...trade. Our beliefs are the foundation for how we conduct business... ...design, develop, and implement machine learning models and algorithms to solve... ...data scientists, software engineers, and product teams to enhance... ...-learn.Familiarity with cloud platforms (AWS, Azure, GCP)...CloudSeniorFoundationFull timeWork at officeLocal areaImmediate startFlexible hours$148.7k - $240.53k
...it’s needed. This model supports real-time... ...Networks is seeking a Sr. Product Manager to... ...strategy for PAN-OS, the core operating... ...physical, virtual, and cloud-delivered form... ...with cross-functional engineering and central... ...security into the foundation of the product lifecycle...CloudSeniorFoundationFull timeWork at officeShift work- ...our proprietary models, and a... ...and continuously learn and adapt.Moveworks... ...on the Forbes Cloud 100 and AI 50 lists... ...’ Reasoning Engine and natural language... ...looking for a Machine Learning... ...distributed training and inference pipeline for... ...as a strong foundation for our hundreds...CloudSeniorFoundationWork at officeRemote workFlexible hours
$224k - $356.5k
Intelligent machines powered by artificial... ...that can learn, reason, and interact... ...provides the foundation for machines to... ...Perception Engineer to help design... ...centric foundation models, and multi-sensor... ...of radar point cloud data at every... ...optimizing training or inference pipelines...CloudSeniorFoundationFull time$203.5k - $299.3k
...Role We are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash grows New... ...causal systems: uplift models, heterogeneous treatment effect... ...practical experience with causal inference, econometrics, experimentation,...FoundationHourly payFull timeWork at officeLocal areaRemote workFlexible hours$92k - $135k
...is The Essential Cloud for AI™. Built for... ...CRWV) in March 2025. Learn more at What... ...'ll Do: Join the Inference team to ship production... ..., and cost for model serving on our GPU... ...from experienced engineers. About the role... ...practical experience. Foundations in data structures...CloudFoundationPermanent employmentFull timeTemporary workCasual workInternshipWork at officeFlexible hours$151.8k - $265.35k
...is seeking SeniorMachine Learning Engineers for our GenAI Services area... ...and develop efficient inference pipelines, optimize models for latency and through at... ...PhD in Computer Science, Machine Learning, or a related field... ...Adobe Firefly, Creative Cloud, Adobe Experience...CloudSeniorFull timeTemporary workLocal areaWorldwide$119.25k - $150.85k
...drivers at unprecedented scale. Within GM AV, the Model Deployment & Inference Solutions team deploys machine learning models from training frameworks (e.g., PyTorch)... ...on the Cadillac Escalade IQ, and we’re hiring engineers to help deliver the next generation of safe,...Full timeInternshipLocal areaWork from homeRelocation packageFlexible hours$193.3k - $261.5k
The Product: AWS Machine Learning accelerators are at... ...best-in-class ML inference performance at the lowest cost in cloud. Trainium will deliver... ...including silicon engineering, hardware design... ...neural net models on our custom-built... ...performance.You: As a Sr. Machine Learning...CloudSeniorInternshipLocal areaWork from homeRelocationFlexible hours$182.5k - $260.5k
...networking for the cloud and AI era. We secure... ..., its Zero Trust Engine, and the powerful NewEdge... ...at Netskope to learn more. Follow us on... ...a Senior Staff Machine Learning Scientist, you own the inference and optimization layer... ...-tune and evaluate models, push latency and throughput...CloudSenior- ...shape the global strategy for scaled-out AI inference. You will architect high-throughput, low-latency distributed pipelines and model serving strategies for massive scale and... ...versioning, and automated scaling in enterprise and cloud environments, overseeing cross-functional...CloudSenior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference. Be the first to apply!
- senior ml engineer Santa Clara, CA
- computer vision machine learning engineer Santa Clara, CA
- machine learning engineer Santa Clara, CA
- machine learning software engineer Santa Clara, CA
- senior cloud security engineer Santa Clara, CA
- senior cloud solutions architect Santa Clara, CA
- senior cloud data engineer Santa Clara, CA
- cloud operations engineer Santa Clara, CA
- cloud engineering manager Santa Clara, CA
- informatica cloud developer Santa Clara, CA


