Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference

$184.7k - $324.8k

Apple Oakbrook

Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference Santa Clara, California, United States Machine Learning and AI We are the Foundation Model Inference team within Cloud OS and AI Inference organization. We are on a mission to build the most highly performant, secure and private inference stack that powers Siri AI, Apple Intelligence and Apps that are powered with the largest foundation models.Our systems serve billions of queries daily across Siri AI, Apple Intelligence, Apple Search, Apple Music, Apple TV, App Store, iMessage, Photos, Camera, Spotlight & Safari, at remarkably low latency with every ounce of compute extracted from the hardware beneath them. We optimize language, vision, and speech models with billions of parameters using state-of-the-art techniques and ship them at Apple scale.This is a rare opportunity to directly shape how AI reaches billions of people worldwide. Description You will work at the intersection of research and production, partnering closely with the Foundation Model Research team and our external partners to bring cutting-edge model architectures from prototype to planetary-scale deployment. You will own hard problems in inference efficiency, hardware/software codesign, systems architecture, and tooling, and help set the technical direction for the engineers around you.This role sits within CloudOS and Private Cloud Compute (PCC) - Apple's purpose-built, privacy-preserving cloud infrastructure for AI workloads. PCC represents a first-of-its-kind approach to running foundation models in the cloud with verifiable privacy guarantees, and CloudOS is the systems foundation that makes it possible. You will be building and optimizing inference systems on top of this infrastructure, working closely with platform and security teams to deliver both performance and trust at scale. Responsibilities Partner with the Foundation Model Research team and our external partners to optimize inference for the latest model architectures across language, vision, and speech. Design and ship production-grade inference systems serving millions of customers in real time. Build profiling tools and simulators to identify and resolve performance bottlenecks across different hardware configurations and use cases. Drive technical decisions on high-throughput, low-latency serving at supercomputing scale. Mentor and grow engineers across the organization. Minimum Qualifications 5+ years of experience leading complex, ambiguous technical projects from end to end. Hands-on experience with LLM inference stacks. Working knowledge of GPU or TPU programming concepts. Proficiency with PyTorch, JAX, or TensorFlow. Experience building and operating high-throughput services at large distributed scale. Proficiency deploying applications on cloud platforms (AWS, GCP, or equivalent) using Kubernetes and Docker. BS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field. Preferred Qualifications Experience building productions systems in Go or Python. Strong knowledge of deep learning architectures including Transformers, encoder/decoder models, and multimodal variants. Experience with inference optimization frameworks such as TensorRT-LLM, vLLM, SGLang, TGI, or Nvidia Triton Server. Experience authoring custom CUDA kernels using CUDA C++ or OpenAI Triton. MS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field. At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $184,700 and $324,800, and your base pay will depend on your skills, qualifications, experience, and location. Apple employees also have the opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan. You’ll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program. Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong. Learn about accessibility in Apple’s workplace Learn about reasonable accommodations for job applicants Apple accepts applications to this posting on an ongoing basis. #J-18808-Ljbffr Apple

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference in Santa Clara, CA vacancy
  • Apple Inc. is seeking a Sr. Machine Learning Engineer for the Foundation Models Inference team in Santa Clara, CA. You will collaborate with research and external partners to bring cutting-edge model architectures from prototype to planetary-scale deployment, owning hard... 
    Cloud
    Senior
    Foundation

    Apple

    Santa Clara, CA
    3 days ago
  •  ...leading training and inference speeds; over 10...  ...based hyperscale cloud inference...  ...with the leading model labs, global enterprises...  ...state-of-the-art foundation models and...  ...Cerebras' Wafer-Scale Engine (WSE). We build the...  ...intersection of machine learning frameworks, compiler... 
    Cloud
    Senior
    Foundation

    Cerebras Systems

    Sunnyvale, CA
    5 days ago
  • Cerebras Systems is seeking an engineering leader to build and scale the Inference Model Scaling organization. You will define...  ...team enabling latest foundation models on Cerebras hardware, leading...  ...Work across compiler, runtime, cloud, hardware, product management, and... 
    Cloud
    Foundation

    Cerebras

    Sunnyvale, CA
    3 days ago
  • $193.3k - $261.5k

     ...software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators. Join us to optimize the latest models to run really fast on the Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you will onboard and... 
    Cloud
    Senior
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  •  ...specific focus on large-scale model inference on AMD Instinct™ and Radeon™...  ...it. You will influence engineering roadmaps, represent AMD in open...  ...relationships with key OSS maintainers, foundation working groups, and...  ...Hugging Face, Red Hat, and cloud providers on joint... 
    Cloud
    Senior
    Foundation
    Remote work

    AMD

    Santa Clara, CA
    5 days ago
  • $117.7k - $221.4k

     ...sits at the intersection of machine learning, data infrastructure, and...  ...not only on stronger models, but also on better infrastructure...  ...model reflects how Cola engineers think: build durable...  ...processing, featurization, and inference foundations that power scalable world... 
    Foundation
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $190.2k - $345.65k

     ...image, video, and 3D models built on each...  ...hiring a Senior Staff Machine Learning Engineer to architect and lead...  ...Strong data-engineering foundations — large-scale batch...  ...embedding models and the inference paths that produce...  ...CI/CD, and a major cloud (AWS or Azure).... 
    Cloud
    Senior
    Foundation
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    2 days ago
  • $184k - $287.5k

     ...network stack for distributed inference which will be used to orchestrate wide area networks/cloud Networks. Enabling a workload...  ...Computer Science, Electrical Engineering, Software Engineer, or...  ...orchestration frameworks with strong foundation Network topology and... 
    Cloud
    Senior
    Foundation
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    2 days ago
  • $215k - $285k

     ...intelligence to every moving machine on the planet....  ...Seoul; and Tokyo. Learn more at applied.co...  ...for a performance engineer who specializes in...  ...-throughput batch inference sweeping petabytes...  ...and performance models for our workloads,...  ...machine learning foundations, and the ability... 
    Foundation
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    NLP PEOPLE

    Sunnyvale, CA
    3 days ago
  •  ...The Role We are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash grows New...  ...production causal systems: uplift models, heterogeneous treatment effect...  ...experience with causal inference, econometrics, experimentation,... 
    Foundation
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    DoorDash

    Sunnyvale, CA
    3 days ago
  • $184.7k - $324.8k

    Staff / Sr. Machine Learning Engineer, AI, Search & Knowledge Platforms...  ...large language models, advanced NLP, and entity...  ...corpora for Apple's foundation models, directly...  ...Experience working in a cloud-native environment...  ...Server, or similar inference frameworks. Experience... 
    Cloud
    Senior
    Foundation
    Relocation

    Apple

    Cupertino, CA
    4 days ago
  •  ...-leading training and inference speeds; over 10 times...  ...GPU-based hyperscale cloud inference services. This...  ...works with the leading model labs, global...  ...RoleWe're hiring a Staff Engineer to help lead, drive, and...  ...cloud components with machine learning services. We are often... 
    Cloud
    Senior

    Cerebras Systems

    Sunnyvale, CA
    5 days ago
  •  ...-leading training and inference speeds; over 10 times...  ...GPU-based hyperscale cloud inference services. This...  ...works with the leading model labs, global...  ...working directly with Engineering, Product, Infrastructure...  ...work through continuous learning, growth and support of... 
    Cloud
    Senior
    Remote work

    Cerebras Systems

    Sunnyvale, CA
    5 days ago
  • $229k - $343k

     ...themselves, live in the moment, learn about the world, and have...  ...generative AI, including foundational models, efficient infrastructure,...  ...on-device and server-side inference. Our team creates...  ...Spectacles.We’re looking for a Machine Learning Engineer to join Snap Inc!What you’... 
    Full time
    Live in
    Work at office
    Local area
    Worldwide

    Snap

    Palo Alto, CA
    6 days ago
  • GMI Cloud is a fast-growing AI infrastructure company backed by Headline VC and...  ...services from GPU compute service to AI model inference API solutions. As an NVIDIA Reference...  .... About this role We are hiring a Machine Learning Engineer, LLM Optimization to build a world-leading... 
    Cloud
    Worldwide

    GMI Cloud

    Mountain View, CA
    22 hours ago
  • $272k - $431.25k

     ...and FlashDreams are the foundation for a new generation of interactive world-model systems. With this...  ...for fidelity, real-time inference performance in world models...  ...source inferencing engine. We are on a mission to...  ...passionate about developing cloud services we want to... 
    Cloud
    Senior
    Foundation
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $230k - $265k

     ...veteran scientists and engineers. As a Senior Machine Learning Engineer, you’ll...  ...post-training, and inference strategies for...  ...language and speech models using PyTorch and/or...  ...deployment workflows in a cloud environment....  ...large language or foundation models, with production... 
    Cloud
    Senior
    Foundation
    Permanent employment

    Otter.ai

    Mountain View, CA
    2 days ago
  • $195.2k - $361.2k

    ## Sr. Inference Optimization Engineer (local / edge runtime)Applylocations: US,...  ...the best of local and cloud intelligence — private...  ...design. Small, efficient models run directly on the user's machine (AI PC, edge, on-prem,...  ...helps us# What you’ll learn / grow into*Curiosity... 
    Cloud
    Senior
    Internship
    Local area
    Shift work

    Intel

    Santa Clara, CA
    22 hours ago
  • $184k - $287.5k

     ...to change how the AI inference industry works. With agent...  ...of vast NCP (NVIDIA Cloud Partner)...  ...executives, distinguished engineers and datacenter developers...  ...for someone with fast-learning technical agility and...  ..., or TensorRT-LLM for model optimization and serving... 
    Cloud
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

    Intelligent machines powered by artificial...  ...that can learn, reason, and interact...  ...provides the foundation for machines to...  ...Perception Engineer to help design...  ...centric foundation models, and multi-sensor...  ...of radar point cloud data at every...  ...optimizing training or inference pipelines... 
    Cloud
    Senior
    Foundation
    Full time

    Nvidia

    Santa Clara, CA
    6 days ago
  • $148.7k - $240.53k

     ...it’s needed. This model supports real-time...  ...Networks is seeking a Sr. Product Manager to...  ...strategy for PAN-OS, the core operating...  ...physical, virtual, and cloud-delivered form...  ...with cross-functional engineering and central...  ...security into the foundation of the product lifecycle... 
    Cloud
    Senior
    Foundation
    Full time
    Work at office
    Shift work

    Palo Alto Networks

    Santa Clara, CA
    6 days ago
  •  ...shape the global strategy for scaled-out AI inference. You will architect high-throughput, low-latency distributed pipelines and model serving strategies for massive scale and...  ...versioning, and automated scaling in enterprise and cloud environments, overseeing cross-functional... 
    Cloud
    Senior

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $92k - $135k

     ...is The Essential Cloud for AI™. Built for...  ...CRWV) in March 2025. Learn more at  What...  ...'ll Do: Join the Inference team to ship production...  ..., and cost for model serving on our GPU...  ...from experienced engineers. About the role...  ...experience. Foundations in data structures... 
    Cloud
    Foundation
    Permanent employment
    Full time
    Temporary work
    Casual work
    Internship
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    2 days ago
  •  ...our proprietary models, and a...  ...and continuously learn and adapt.Moveworks...  ...on the Forbes Cloud 100 and AI 50 lists...  ...’ Reasoning Engine and natural language...  ...looking for a Machine Learning...  ...distributed training and inference pipeline for...  ...as a strong foundation for our hundreds... 
    Cloud
    Senior
    Foundation
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    5 days ago
  • $151.8k - $265.35k

     ...is seeking SeniorMachine Learning Engineers for our GenAI Services area...  ...and develop efficient inference pipelines, optimize models for latency and through at...  ...PhD in Computer Science, Machine Learning, or a related field...  ...Adobe Firefly, Creative Cloud, Adobe Experience... 
    Cloud
    Senior
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    2 days ago
  • $124.5k - $272k

     ...networking for the cloud and AI era. We secure...  ..., its Zero Trust Engine, and the powerful NewEdge...  ...at Netskope to learn more. Follow us on...  ...a Senior Staff Machine Learning Scientist, you own the inference and optimization layer...  ...-tune and evaluate models, push latency and throughput... 
    Cloud
    Senior

    Netskope

    Santa Clara, CA
    3 days ago
  • $195.2k - $262.2k

     ...leading a new era in cloud infrastructure for the...  ...enterprises from data and model training through to...  ...infrastructure. Built by engineers, for engineers. From...  ...GPU orchestration to inference optimization, we own...  ...in computer science, machine learning, ML systems, computer... 
    Cloud
    Senior
    Temporary work
    Immediate start
    Remote work

    Nebius

    Palo Alto, CA
    22 hours ago
  • $193.3k - $261.5k

    The Product: AWS Machine Learning accelerators are at...  ...best-in-class ML inference performance at the lowest cost in cloud. Trainium will deliver...  ...including silicon engineering, hardware design...  ...neural net models on our custom-built...  ...performance.You: As a Sr. Machine Learning... 
    Cloud
    Senior
    Internship
    Local area
    Work from home
    Relocation
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  •  ...industry-leading training and inference speeds; over 10 times...  ...GPU-based hyperscale cloud inference services....  ...works with the leading model labs, global enterprises...  ...'re hiring a Principal Engineer for our Inference Cloud...  ...work through continuous learning, growth and support of... 
    Cloud

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  •  ...and integrate large language models (LLMs) and other state-of-...  ...experiences. Knowledge and passion in machine learning algorithms, Gen AI, LLMs,...  ...or PyTorch. Experience with cloud platforms (AWS) and...  ...strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT... 
    Cloud
    Full time
    Work experience placement

    Eightfold

    Santa Clara, CA
    8 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference. Be the first to apply!