ML Engineer: Distributed Inference & Model Serving
ByteDance
A leading technology company in Seattle is seeking a Machine Learning Engineer for Model Serving Infrastructure. The ideal candidate will have at least 5 years of experience and strong programming skills in C/C++/CUDA. You will design and implement distributed inference infrastructure and collaborate with product teams. This position offers a competitive salary and benefits in a dynamic work environment. #J-18808-Ljbffr ByteDance
$175k - $280k
Sesame, located in Bellevue, Washington, is seeking a talented engineer to join our team focused on revolutionizing the way computers interact... ...with humans. The role involves optimizing machine learning models for a new consumer product category, working with state-of-the-...Suggested$175k - $280k
...Responsibilities Turbocharge our serving layer, consisting of a variety of LLM, speech, and vision models. Partner with ML infrastructure and training engineers to build a fast, cost‑effective,... ...and custom kernels to speed up inference. Find ways to reduce model...SuggestedContract workFlexible hours$177.69k - $416.1k
Machine Learning Engineer-Model Serving Infrastructure Machine Learning Engineer-Model Serving Infrastructure... ...for the design and implementation of distributed inference infrastructure for feeds, ads and... ...-Design, High Performance Computing, ML Hardware Acceleration (e.g., GPU/RDMA...SuggestedFull timeTemporary workLocal area$209k - $313k
...digital services.Snap Engineering teams build fun and technically... ...ll do:Design and build models that quantify causal... ...of causal inference and modern approaches to... ...and leveraging causal ML in production systemsPreferred... ...reinforce our values, and serve our community,...SuggestedFull timeLive inWork at officeLocal area$200k - $300k
Member of Technical Staff — Model Optimization and Inference (New Grad) Seattle,... ...stack from model weights to serving infrastructure: quantization... ...posting is aimed at early-career engineers finishing or recently... ...For BS, MS, or PhD in CS, ML, or a related field — completed...SuggestedInternshipH1bWork at officeVisa sponsorship- Apple Inc. seeks a Sr. Machine Learning Engineer for Foundation Models Inference in Cloud OS & Inference to advance private, scalable AI inference across... ...scale. You will own high-throughput, low-latency serving, mentor engineers, and collaborate with security teams...
$175k - $308.5k
...it than Apple. The Foundation Model Services team builds the frameworks... ..., Safari, Siri, and more — serving millions of queries at... ...researchers to prototype and develop inference for cutting-edge model... ...year+ industry experience in ML technologies (LLMs, Machine Learning...Relocation$229k - $343k
...services.Snap’s Generative ML Platform team builds... ...including foundational models, efficient... ...device and server-side inference. Our team creates intuitive... ...for a Machine Learning Engineer to join Snap Inc!What you... ...technology and products that serve millions of SnapchattersBuild...Full timeLive inWork at officeLocal areaWorldwide- ...building machine learning models and systems to protect... ...stability3. Context engineering and tool collaboration... ....2. Business value: Serving content safety for billions... ...understanding of distributed computing framework &... ...for training/finetuning/inference; Being familiar with...Flexible hoursShift work
$184.7k - $324.8k
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference Santa Clara, California, United... ...foundation models.Our systems serve billions of queries daily across Siri... ...high‑throughput services at large distributed scale. Proficiency deploying...WorldwideRelocation- ...architectures, combining rigorous engineering with learning systems... ...are seeking a Staff ML Systems Engineer to architect and build the distributed infrastructure that... ...data processing, model training, evaluation,... ...learning training and inference systems. Familiarity...Local area
$182.8k - $247.3k
...supporting Foundation Model Providers (FMP) on... ..., and distributed computing at extraordinary... ...train, fine-tune, and serve state-of-the-art... ...) along with ML expertise to build... ...relationships with customer engineering teams• Dive deep... ..., training/inference lifecycles, and optimization...Flexible hours$150.4k - $277.6k
...seeking research engineers to build systems for... ...stack, including model training, harness... ...experience. 2+ years of ML engineering... ...maintaining training and inference systems -... ...infrastructure, model serving, or evaluation... ...on experience with distributed ML systems - CI/CD...Relocation$180k
...motivated, and focused on engineering excellence. This organization... ...engineer on the Imagine Model Team, you will develop cutting... ..., modeling, training, inference serving, and product integration, covering... ...working with large-scale distributed machine learning systems....Temporary work$200.8k - $251k
...seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a...Full time- ...to take your software engineering career to the next... ...experience, with emphasis on ML systems.Hands-on... ...working with distributed systems concepts, microservices... ...to ML model serving frameworks (e.g., TorchServe... ...TensorFlow Serving, Triton Inference Server)Familiarity...
$151.8k - $265.35k
...Senior Machine Learning Engineers for our GenAI... ...and develop efficient inference pipelines, optimize models for latency and through... ...of products that serve individual and enterprise... ..., and mentor other ML engineers.Job... ...expertise in Kubernetes, distributed systems, and MLOps...Full timeTemporary workLocal areaWorldwide$184.5k
...This Senior Machine Learning Engineer role is part of the Distribution & Supply team which sits... ...deploy, and scale robust models that directly improve the... ...customer problems into clear ML‑driven solutions,... ...workloads or large‑scale batch inference, including clear, well‑versioned...Full time$150.75k - $241.2k
...Machine Learning Engineer at Axon, you’ll help... ...talented ML engineers and scientists... ...machine learning models across Axon devices... ...scale evaluation, inference optimization, data... ...machine learning, distributed systems, cloud infrastructure... ...communities we serve.Studies have shown...Work experience placementWork at officeRemote work$168.1k - $227.4k
...over dozens of chained inferences, coupled to real... ...systems problem as a modeling problem - agent performance... ...infrastructure, and the serving stack co-designed... ...Senior Machine Learning Engineer to build and own core... ...architectures, and large-scale distributed systems, serving...InternshipFlexible hoursDay shift$202.16k - $368.22k
...Infrastructures team oversees the distributed training,... ...framework, high-performance inference, and heterogeneous... ...for AI foundation models. Responsibilities Design... ...science, mathematics, engineering, or a related field,... ...Dola and Dreamnia — and serves enterprise customers...Temporary workInternshipLocal area$130k - $260k
...production-grade ML systems powering customer... ...pipelines, model training, evaluation, deployment, serving, monitoring, and continuous... ...systems.Establish engineering patterns and best... ..., deployment, inference, monitoring, and... ...understanding of distributed systems, APIs and...Full timeTemporary workPart timeWork experience placement$206.4k - $379.1k
...image, video, and 3D models built on each customer... ...Principal Machine Learning Engineer to serve as the technical lead... .... You will set the inference architecture and technical... .... Director, ML Engineering and ML Engineering... ...in Kubernetes, distributed systems, and MLOps platforms...Full timeTemporary workLocal areaWorldwide$197.3k - $313.7k
...for a Staff Machine Learning Engineer with deep expertise in model training and finetuning to join our ML team. You'll design, train, and... ...at Slack ship models that serve millions of users daily. This... ...with model optimization for inference (quantization, pruning, speculative...Full time- ...Design and build ML platform orchestration... ...FinOps. Build online model-serving lifecycle orchestration... ...model and image distribution, deployment, upgrades... ...computer science, software engineering, artificial... ...model distribution, or inference performance analysis....Full timeInternship
$200k - $250k
...Machine Learning Engineering within the Advanced... ...annotation pipelines, ML Infrastructure and... ...state-of-art models into robust, autonomous... ...infrastructure to serve computer vision... ..., including distributed training infrastructure... ..., and low-latency inference services. Ensure high...Temporary workWork at officeLocal area- ...Responsibilities The Data-AML-Engine Orchestration team... ...that powers online model serving across ByteDance... ...infrastructure with production ML workloads.You will... ...including model and image distribution, deployment, upgrades,... ...distribution, or inference performance analysis....Internship
- ...Inc. in Seattle is seeking a Software Engineer focused on Machine Learning and AI to build production-grade systems that serve real-time model inferences at scale. You will collaborate with... ...shipping production software, 5+ years ML/AI experience, and a BS in CS or...
- Haus Analytics in Seattle is seeking a senior ML engineer to drive high-impact projects on the cMMM space, blending ML, causal inference, and scalable production code. You will... ...cross-functional teams to deliver trustworthy models and scalable processes, while mentoring...Flexible hours
- ...Own the architecture for inference serving, GPU scheduling, Kubernetes operators... ...evaluation systems for model, prompt, and agent changes, including... .... Apply physics-informed ML and enterprise AI expertise... ...evaluation practices, mentor engineers, and maintain architectural...Full timeRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Engineer: Distributed Inference & Model Serving. Be the first to apply!
- ai ml engineer Seattle, WA
- graduate machine learning engineer Seattle, WA
- senior ml engineer Seattle, WA
- machine learning ai engineer Seattle, WA
- computer vision machine learning engineer Seattle, WA
- data scientist machine learning engineer Seattle, WA
- machine learning engineer Seattle, WA
- machine learning software engineer Seattle, WA
- artificial intelligence - machine learning intern Seattle, WA
- internship machine learning Seattle, WA



