ML Engineer: Distributed Inference & Model Serving
ByteDance
A leading technology company in Seattle is seeking a Machine Learning Engineer for Model Serving Infrastructure. The ideal candidate will have at least 5 years of experience and strong programming skills in C/C++/CUDA. You will design and implement distributed inference infrastructure and collaborate with product teams. This position offers a competitive salary and benefits in a dynamic work environment. #J-18808-Ljbffr ByteDance
- ...Inc. in the San Francisco Bay Area is looking for a Staff/Sr. Machine Learning Engineer to optimize inference for large foundation models and deliver production-grade solutions that serve millions of users in real time. You will collaborate closely with research and product...Suggested
$175k - $280k
Sesame, located in Bellevue, Washington, is seeking a talented engineer to join our team focused on revolutionizing the way computers interact... ...with humans. The role involves optimizing machine learning models for a new consumer product category, working with state-of-the-...Suggested$175k - $280k
...Responsibilities Turbocharge our serving layer, consisting of a variety of LLM, speech, and vision models. Partner with ML infrastructure and training engineers to build a fast, cost‑effective,... ...and custom kernels to speed up inference. Find ways to reduce model...SuggestedContract workFlexible hours$177.69k - $416.1k
Machine Learning Engineer-Model Serving Infrastructure Machine Learning Engineer-Model Serving Infrastructure... ...for the design and implementation of distributed inference infrastructure for feeds, ads and... ...-Design, High Performance Computing, ML Hardware Acceleration (e.g., GPU/RDMA...SuggestedFull timeTemporary workLocal area$200k - $300k
...developing foundation models designed for it from the... ...from model weights to serving infrastructure: quantization... ...aimed at early-career engineers finishing or recently... ...to end-to-end inference optimization across our... ...BS, MS, or PhD in CS, ML, or a related field — completed...SuggestedFull timeInternshipH1bWork at officeVisa sponsorship$229k - $343k
...services.Snap’s Generative ML Platform team builds... ...including foundational models, efficient... ...device and server-side inference. Our team creates intuitive... ...for a Machine Learning Engineer to join Snap Inc!What you... ...technology and products that serve millions of SnapchattersBuild...Full timeLive inWork at officeLocal areaWorldwide- ...building machine learning models and systems to protect... ...stability3. Context engineering and tool collaboration... ....2. Business value: Serving content safety for billions... ...understanding of distributed computing framework &... ...for training/finetuning/inference; Being familiar with...Flexible hoursShift work
$182.8k - $247.3k
...supporting Foundation Model Providers (FMP) on... ..., and distributed computing at extraordinary... ...train, fine-tune, and serve state-of-the-art... ...) along with ML expertise to build... ...relationships with customer engineering teams• Dive deep... ..., training/inference lifecycles, and optimization...Flexible hours$180k
...motivated, and focused on engineering excellence. This organization... ...engineer on the Imagine Model Team, you will develop cutting... ..., modeling, training, inference serving, and product integration, covering... ...working with large-scale distributed machine learning systems....Temporary work$184.5k
...This Senior Machine Learning Engineer role is part of the Distribution & Supply team which sits... ...deploy, and scale robust models that directly improve the... ...customer problems into clear ML‑driven solutions,... ...workloads or large‑scale batch inference, including clear, well‑versioned...Full time$151.8k - $265.35k
...SeniorMachine Learning Engineers for our GenAI... ...and develop efficient inference pipelines, optimize models for latency and through... ...of products that serve individual and enterprise... ..., and mentor other ML engineers.Job... ...expertise in Kubernetes, distributed systems, and MLOps platforms...Full timeTemporary workLocal areaWorldwide$200.8k - $251k
...seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a...Full time$202.16k - $368.22k
...Infrastructures team oversees the distributed training,... ...framework, high-performance inference, and heterogeneous... ...for AI foundation models. Responsibilities Design... ...science, mathematics, engineering, or a related field,... ...Dola and Dreamnia — and serves enterprise customers...Temporary workInternshipLocal area$197.3k - $313.7k
...for a Staff Machine Learning Engineer with deep expertise in model training and finetuning to join our ML team. You'll design, train, and... ...at Slack ship models that serve millions of users daily. This... ...with model optimization for inference (quantization, pruning, speculative...Full time$183.7k - $248.6k
...Learning Infrastructure Engineer to join our Vector Ads... ...that brings ML models from training into production... ...intersection of ML systems and distributed infrastructure,... ...infrastructure that serves ML models in real-time... ...model versioning, and inference optimization What...Work at officeRemote workWorldwideRelocation package$200k - $250k
...Machine Learning Engineering within the Advanced... ...annotation pipelines, ML Infrastructure and... ...state-of-art models into robust, autonomous... ...infrastructure to serve computer vision... ..., including distributed training infrastructure... ..., and low-latency inference services. Ensure high...Temporary workWork at officeLocal area- Sift is seeking an experienced Machine Learning Engineer to bridge data science and distributed systems in a production setting. You will build end-to-end ML pipelines, train merchant-specific models, and serve predictions at scale with low latency. You will work on an...
$100k - $150k
...ML Performance Engineer - Remote Bright Vision Technologies... ...across training and inference workloads for large... ...kernel optimization to distributed system tuning,... ...of GPU architecture, model parallelism, memory management... ...decoding for LLM serving. Drive compiler-level...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship- ...Responsibilities The Data-AML-Engine Orchestration team... ...that powers online model serving across ByteDance... ...infrastructure with production ML workloads. You will... ...model and image distribution, deployment, upgrades,... ...model distribution, or inference performance analysis....Internship
$200k - $345k
...updates on our news and engineering blogs and join us as we... ...the future of AI and ML at Whatnot. You’ll design... ...-hosted large language model applications across the... ...-latency, large model serving to distributed training & high-throughput GPU inference. What you'll do: Own the...Work experience placementWork at officeLocal areaRemote workWork from homeHome officeFlexible hours$176.76k - $232k
...efficiency.Core responsibilities As a Senior AI/ML Engineer, you will lead the delivery of scalable... ...engineering challenges from setting up model training and fine-tuning to architectures and system design for serving AI/ML inference solutions in production. You will help...Permanent employmentFull timeContract workPart timeWork visa$169.8k - $355.4k
...Senior Principal AI Agent / ML Software Engineer is a Senior Staff-level,... ...autonomous workflows, scalable inference infrastructure, and... ...ideal candidate combines deep distributed systems experience with... ...execution, inference systems, model serving, AI workflow orchestration...Temporary workFlexible hours$146.83k - $192.72k
...Core responsibilities As an AI/ML Engineer, you will contribute to the... ...AI/ML systems across the full model lifecycle from dataset... ...architecture, training strategies, and serving infrastructure, working with... ...training pipelines using distributed training frameworks (PyTorch...Permanent employmentFull timePart timeWork visa$185.6k - $255k
...segment without waiting on an engineering queue. The bet underneath... ...you.The RoleAt Amperity, ML Engineers work in small,... ..., and predictive models at scale.Improve model inference latency to deliver predictions... ..., model registry, and serving.About YouYou're a ML engineer...Work at officeLocal areaRemote work$141.9k - $190.3k
...global organization of engineers, product developers,... ...consumer media touch points serving millions of people... ..., advertising, and distribution businesses for years to... ...help build services and models enabling efficient ad... ...testable software and apply ML solutions where needed...Work experience placementWork at office$141.9k - $190.3k
...global organization of engineers, product developers,... ...consumer media touch points serving millions of people... ..., advertising, and distribution businesses for years to... ...help build services and models enabling efficient ad... ...testable software and apply ML solutions where needed...Full timeWork experience placementWork at office$168.1k - $227.4k
...Senior Machine Learning Engineer with expertise in agentic system, production ML systems, and scalable deploymentarchitectures... ..., machine learning model development, model validation and serving• Research and implement... ...architecture, training/inference lifecycles, and...InternshipWorldwideFlexible hours$202.16k - $368.22k
...resume. Building a next-generation big model as a service platform to serve hundreds of LLMs based applications.... ...offline training/finetuning, online inference, model management, and resource... ...ideally with experience with scalable ML systems and real-world deployment, feedback...Temporary workInternshipLocal area$143.7k - $194.4k
What does it take to distribute billions of security credentials every day to every host across... ...of this role.As a Software Development Engineer on the AWS Credentials Distribution... ...tier-0 credential distribution system that serves every service team across Amazon's...InternshipFlexible hours$172k
...Senior AI/ML Engineer Chicago, IL, USA; New York, NY, USA; San Francisco... ...foundational transformer models that convert behavioral and... ...for training, serving, and monitoring large-scale... ...proficiency in Python, SQL, and distributed computing and model training...Full timeWork at officeLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Engineer: Distributed Inference & Model Serving. Be the first to apply!
- graduate machine learning engineer Seattle, WA
- senior ml engineer Seattle, WA
- data scientist machine learning engineer Seattle, WA
- machine learning engineer Seattle, WA
- machine learning ai engineer Seattle, WA
- machine learning software engineer Seattle, WA
- ai ml engineer Seattle, WA
- computer vision machine learning engineer Seattle, WA
- machine learning researcher Seattle, WA
- machine learning part time Seattle, WA


