ML Engineer: Distributed Inference & Model Serving
ByteDance
A leading technology company in Seattle is seeking a Machine Learning Engineer for Model Serving Infrastructure. The ideal candidate will have at least 5 years of experience and strong programming skills in C/C++/CUDA. You will design and implement distributed inference infrastructure and collaborate with product teams. This position offers a competitive salary and benefits in a dynamic work environment. #J-18808-Ljbffr ByteDance
- ...Inc. in the San Francisco Bay Area is looking for a Staff/Sr. Machine Learning Engineer to optimize inference for large foundation models and deliver production-grade solutions that serve millions of users in real time. You will collaborate closely with research and product...Suggested
$175k - $280k
Sesame, located in Bellevue, Washington, is seeking a talented engineer to join our team focused on revolutionizing the way computers interact... ...with humans. The role involves optimizing machine learning models for a new consumer product category, working with state-of-the-...Suggested$175k - $280k
...Responsibilities Turbocharge our serving layer, consisting of a variety of LLM, speech, and vision models. Partner with ML infrastructure and training engineers to build a fast, cost‑effective,... ...and custom kernels to speed up inference. Find ways to reduce model...SuggestedContract workFlexible hours$177.69k - $416.1k
Machine Learning Engineer-Model Serving Infrastructure Machine Learning Engineer-Model Serving Infrastructure... ...for the design and implementation of distributed inference infrastructure for feeds, ads and... ...-Design, High Performance Computing, ML Hardware Acceleration (e.g., GPU/RDMA...SuggestedFull timeTemporary workLocal area$200k - $300k
Member of Technical Staff — Model Optimization and Inference (New Grad) Seattle,... ...stack from model weights to serving infrastructure: quantization... ...posting is aimed at early-career engineers finishing or recently... ...For BS, MS, or PhD in CS, ML, or a related field — completed...SuggestedInternshipH1bWork at officeVisa sponsorship$229k - $343k
...services.Snap’s Generative ML Platform team builds... ...including foundational models, efficient... ...device and server-side inference. Our team creates intuitive... ...for a Machine Learning Engineer to join Snap Inc!What you... ...technology and products that serve millions of SnapchattersBuild...Full timeLive inWork at officeLocal areaWorldwide- ...building machine learning models and systems to protect... ...stability3. Context engineering and tool collaboration... ....2. Business value: Serving content safety for billions... ...understanding of distributed computing framework &... ...for training/finetuning/inference; Being familiar with...Flexible hoursShift work
$182.8k - $247.3k
...supporting Foundation Model Providers (FMP) on... ..., and distributed computing at extraordinary... ...train, fine-tune, and serve state-of-the-art... ...) along with ML expertise to build... ...relationships with customer engineering teams• Dive deep... ..., training/inference lifecycles, and optimization...Flexible hours$200.8k - $251k
...seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a...Full time$184.5k
...This Senior Machine Learning Engineer role is part of the Distribution & Supply team which sits... ...deploy, and scale robust models that directly improve the... ...customer problems into clear ML‑driven solutions,... ...workloads or large‑scale batch inference, including clear, well‑versioned...Full time$151.8k - $265.35k
...SeniorMachine Learning Engineers for our GenAI... ...and develop efficient inference pipelines, optimize models for latency and through... ...of products that serve individual and enterprise... ..., and mentor other ML engineers.Job... ...expertise in Kubernetes, distributed systems, and MLOps platforms...Full timeTemporary workLocal areaWorldwide$180k
...motivated, and focused on engineering excellence. This organization... ...engineer on the Imagine Model Team, you will develop cutting... ..., modeling, training, inference serving, and product integration, covering... ...working with large‑scale distributed machine learning systems....Temporary work- ...architectures, combining rigorous engineering with learning systems... ...are seeking a Staff ML Systems Engineer to architect and build the distributed infrastructure that... ...data processing, model training, evaluation,... ...learning training and inference systems. Familiarity...Local area
$202.16k - $368.22k
...Infrastructures team oversees the distributed training,... ...framework, high-performance inference, and heterogeneous... ...for AI foundation models. Responsibilities Design... ...science, mathematics, engineering, or a related field,... ...Dola and Dreamnia — and serves enterprise customers...Temporary workInternshipLocal area$197.3k - $313.7k
...for a Staff Machine Learning Engineer with deep expertise in model training and finetuning to join our ML team. You'll design, train, and... ...at Slack ship models that serve millions of users daily. This... ...with model optimization for inference (quantization, pruning, speculative...Full time$106.9k - $160.4k
...enterprise, we are seeking a skilled ML Engineer to design, build, and... ...transition from experimental models to trusted, production-grade AI... ...cases. Model Deployment & Serving Operationalize and deploy batch and real-time inference solutions using cloud-native services...Full timeTemporary work$156.75k - $250.8k
...seasoned Machine Learning Engineer to join a new team... ...between a promising model and a production... ...infrastructure, the inference and serving stack, the retrieval... ...maintaining large-scale distributed platforms in production... ...building systems across the ML lifecycle: data...Work experience placementWork at officeRemote work$200k - $250k
...Machine Learning Engineering within the Advanced... ...annotation pipelines, ML Infrastructure and... ...state-of-art models into robust, autonomous... ...infrastructure to serve computer vision... ..., including distributed training infrastructure... ..., and low-latency inference services. Ensure high...Temporary workWork at officeLocal area- Sift is seeking an experienced Machine Learning Engineer to bridge data science and distributed systems in a production setting. You will build end-to-end ML pipelines, train merchant-specific models, and serve predictions at scale with low latency. You will work on an...
- ...Responsibilities The Data-AML-Engine Orchestration team... ...that powers online model serving across ByteDance... ...infrastructure with production ML workloads.You will... ...including model and image distribution, deployment, upgrades,... ...distribution, or inference performance analysis....Internship
$100k - $150k
...ML Performance Engineer - Remote Bright Vision Technologies... ...across training and inference workloads for large... ...kernel optimization to distributed system tuning,... ...of GPU architecture, model parallelism, memory management... ...decoding for LLM serving. Drive compiler-level...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship- ...multi-modality foundation model to drive the next... ...Optimization & Deployment Engineer, you will focus on bringing... ...You will optimize the ML models, write custom... ...build highly concurrent inference code to ensure real-... ...Radar). Experience with distributed training pipelines and...Temporary workRelocation package
- ...updates on our news and engineering blogs and join us as we... ...the future of AI and ML at Whatnot. You'll design... ...-hosted large language model applications across the... ...-latency, large model serving to distributed training & high-throughput GPU inference. What you'll do: Own...Work experience placementWork at officeLocal areaRemote workWork from homeHome officeFlexible hours
$176.76k - $232k
...efficiency.Core responsibilities As a Senior AI/ML Engineer, you will lead the delivery of scalable... ...engineering challenges from setting up model training and fine-tuning to architectures and system design for serving AI/ML inference solutions in production. You will help...Permanent employmentFull timeContract workPart timeWork visa$146.83k - $192.72k
...Core responsibilities As an AI/ML Engineer, you will contribute to the... ...AI/ML systems across the full model lifecycle from dataset... ...architecture, training strategies, and serving infrastructure, working with... ...training pipelines using distributed training frameworks (PyTorch...Permanent employmentFull timePart timeWork visa$169.8k - $355.4k
...Senior Principal AI Agent / ML Software Engineer is a Senior Staff-level,... ...autonomous workflows, scalable inference infrastructure, and... ...ideal candidate combines deep distributed systems experience with... ...execution, inference systems, model serving, AI workflow orchestration...Temporary workFlexible hours$168.1k - $227.4k
...Senior Machine Learning Engineer with expertise in agentic system, production ML systems, and scalable deploymentarchitectures... ..., machine learning model development, model validation and serving• Research and implement... ...architecture, training/inference lifecycles, and...InternshipWorldwideFlexible hours$141.9k - $190.3k
...global organization of engineers, product developers,... ...consumer media touch points serving millions of people... ..., advertising, and distribution businesses for years to... ...help build services and models enabling efficient ad... ...testable software and apply ML solutions where needed...Full timeWork experience placementWork at office$141.9k - $190.3k
...global organization of engineers, product developers,... ...consumer media touch points serving millions of people... ..., advertising, and distribution businesses for years to... ...help build services and models enabling efficient ad... ...testable software and apply ML solutions where needed...Work experience placementWork at office$185.6k - $255k
...segment without waiting on an engineering queue. The bet underneath... ...you.The RoleAt Amperity, ML Engineers work in small,... ..., and predictive models at scale.Improve model inference latency to deliver predictions... ..., model registry, and serving.About YouYou're a ML engineer...Work at officeLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Engineer: Distributed Inference & Model Serving. Be the first to apply!
- ai ml engineer Seattle, WA
- senior ml engineer Seattle, WA
- computer vision machine learning engineer Seattle, WA
- machine learning engineer Seattle, WA
- machine learning ai engineer Seattle, WA
- data scientist machine learning engineer Seattle, WA
- graduate machine learning engineer Seattle, WA
- machine learning software engineer Seattle, WA
- machine learning scientist Seattle, WA
- senior research scientist - machine learning Seattle, WA


