Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Engineer: Distributed Inference & Model Serving

ByteDance

A leading technology company in Seattle is seeking a Machine Learning Engineer for Model Serving Infrastructure. The ideal candidate will have at least 5 years of experience and strong programming skills in C/C++/CUDA. You will design and implement distributed inference infrastructure and collaborate with product teams. This position offers a competitive salary and benefits in a dynamic work environment. #J-18808-Ljbffr ByteDance

Vacancy posted 17 hours ago
Similar jobs that could be interesting for youBased on the ML Engineer: Distributed Inference & Model Serving in Seattle, WA vacancy
  • $175k - $280k

    Sesame, located in Bellevue, Washington, is seeking a talented engineer to join our team focused on revolutionizing the way computers interact...  ...with humans. The role involves optimizing machine learning models for a new consumer product category, working with state-of-the-... 
    Suggested

    SESAME

    Bellevue, WA
    1 day ago
  • $175k - $280k

     ...Responsibilities Turbocharge our serving layer, consisting of a variety of LLM, speech, and vision models. Partner with ML infrastructure and training engineers to build a fast, cost‑effective,...  ...and custom kernels to speed up inference. Find ways to reduce model... 
    Suggested
    Contract work
    Flexible hours

    SESAME

    Bellevue, WA
    2 days ago
  • $177.69k - $416.1k

    Machine Learning Engineer-Model Serving Infrastructure Machine Learning Engineer-Model Serving Infrastructure...  ...for the design and implementation of distributed inference infrastructure for feeds, ads and...  ...-Design, High Performance Computing, ML Hardware Acceleration (e.g., GPU/RDMA... 
    Suggested
    Full time
    Temporary work
    Local area

    ByteDance

    Seattle, WA
    17 hours ago
  • $209k - $313k

     ...digital services.Snap Engineering teams build fun and technically...  ...ll do:Design and build models that quantify causal...  ...of causal inference and modern approaches to...  ...and leveraging causal ML in production systemsPreferred...  ...reinforce our values, and serve our community,... 
    Suggested
    Full time
    Live in
    Work at office
    Local area

    Snap

    Seattle, WA
    4 days ago
  • $200k - $300k

    Member of Technical Staff — Model Optimization and Inference (New Grad) Seattle,...  ...stack from model weights to serving infrastructure: quantization...  ...posting is aimed at early-career engineers finishing or recently...  ...For BS, MS, or PhD in CS, ML, or a related field — completed... 
    Suggested
    Internship
    H1b
    Work at office
    Visa sponsorship

    Nuance Labs

    Seattle, WA
    2 days ago
  • Apple Inc. seeks a Sr. Machine Learning Engineer for Foundation Models Inference in Cloud OS & Inference to advance private, scalable AI inference across...  ...scale. You will own high-throughput, low-latency serving, mentor engineers, and collaborate with security teams... 

    Apple

    Seattle, WA
    1 day ago
  • $175k - $308.5k

     ...it than Apple. The Foundation Model Services team builds the frameworks...  ..., Safari, Siri, and more — serving millions of queries at...  ...researchers to prototype and develop inference for cutting-edge model...  ...year+ industry experience in ML technologies (LLMs, Machine Learning... 
    Relocation

    Apple Inc.

    Seattle, WA
    3 days ago
  • $229k - $343k

     ...services.Snap’s Generative ML Platform team builds...  ...including foundational models, efficient...  ...device and server-side inference. Our team creates intuitive...  ...for a Machine Learning Engineer to join Snap Inc!What you...  ...technology and products that serve millions of SnapchattersBuild... 
    Full time
    Live in
    Work at office
    Local area
    Worldwide

    Snap

    Seattle, WA
    4 days ago
  •  ...building machine learning models and systems to protect...  ...stability3. Context engineering and tool collaboration...  ....2. Business value: Serving content safety for billions...  ...understanding of distributed computing framework &...  ...for training/finetuning/inference; Being familiar with... 
    Flexible hours
    Shift work

    TikTok

    Seattle, WA
    2 days ago
  • $184.7k - $324.8k

    Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference Santa Clara, California, United...  ...foundation models.Our systems serve billions of queries daily across Siri...  ...high‑throughput services at large distributed scale. Proficiency deploying... 
    Worldwide
    Relocation

    Apple

    Seattle, WA
    1 day ago
  •  ...architectures, combining rigorous engineering with learning systems...  ...are seeking a Staff ML Systems Engineer to architect and build the distributed infrastructure that...  ...data processing, model training, evaluation,...  ...learning training and inference systems. Familiarity... 
    Local area

    FieldAI

    Seattle, WA
    12 days ago
  • $182.8k - $247.3k

     ...supporting Foundation Model Providers (FMP) on...  ..., and distributed computing at extraordinary...  ...train, fine-tune, and serve state-of-the-art...  ...) along with ML expertise to build...  ...relationships with customer engineering teams• Dive deep...  ..., training/inference lifecycles, and optimization... 
    Flexible hours

    AmazonWebServices

    Seattle, WA
    17 hours ago
  • $150.4k - $277.6k

     ...seeking research engineers to build systems for...  ...stack, including model training, harness...  ...experience. 2+ years of ML engineering...  ...maintaining training and inference systems -...  ...infrastructure, model serving, or evaluation...  ...on experience with distributed ML systems - CI/CD... 
    Relocation

    Apple

    Seattle, WA
    17 hours ago
  • $180k

     ...motivated, and focused on engineering excellence. This organization...  ...engineer on the Imagine Model Team, you will develop cutting...  ..., modeling, training, inference serving, and product integration, covering...  ...working with large-scale distributed machine learning systems.... 
    Temporary work

    SpaceXAI

    Seattle, WA
    5 days ago
  • $200.8k - $251k

     ...seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a... 
    Full time

    Scale AI

    Seattle, WA
    1 day ago
  •  ...to take your software engineering career to the next...  ...experience, with emphasis on ML systems.Hands-on...  ...working with distributed systems concepts, microservices...  ...to ML model serving frameworks (e.g., TorchServe...  ...TensorFlow Serving, Triton Inference Server)Familiarity... 

    JP Morgan Chase

    Seattle, WA
    2 days ago
  • $151.8k - $265.35k

     ...Senior Machine Learning Engineers for our GenAI...  ...and develop efficient inference pipelines, optimize models for latency and through...  ...of products that serve individual and enterprise...  ..., and mentor other ML engineers.Job...  ...expertise in Kubernetes, distributed systems, and MLOps... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    Seattle, WA
    1 day ago
  • $184.5k

     ...This Senior Machine Learning Engineer role is part of the Distribution & Supply team which sits...  ...deploy, and scale robust models that directly improve the...  ...customer problems into clear ML‑driven solutions,...  ...workloads or large‑scale batch inference, including clear, well‑versioned... 
    Full time

    Expedia

    Seattle, WA
    3 days ago
  • $150.75k - $241.2k

     ...Machine Learning Engineer at Axon, you’ll help...  ...talented ML engineers and scientists...  ...machine learning models across Axon devices...  ...scale evaluation, inference optimization, data...  ...machine learning, distributed systems, cloud infrastructure...  ...communities we serve.Studies have shown... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Seattle, WA
    1 day ago
  • $168.1k - $227.4k

     ...over dozens of chained inferences, coupled to real...  ...systems problem as a modeling problem - agent performance...  ...infrastructure, and the serving stack co-designed...  ...Senior Machine Learning Engineer to build and own core...  ...architectures, and large-scale distributed systems, serving... 
    Internship
    Flexible hours
    Day shift

    Amazon

    Bellevue, WA
    2 days ago
  • $202.16k - $368.22k

     ...Infrastructures team oversees the distributed training,...  ...framework, high-performance inference, and heterogeneous...  ...for AI foundation models. Responsibilities Design...  ...science, mathematics, engineering, or a related field,...  ...Dola and Dreamnia — and serves enterprise customers... 
    Temporary work
    Internship
    Local area

    ByteDance

    Seattle, WA
    17 hours ago
  • $130k - $260k

     ...production-grade ML systems powering customer...  ...pipelines, model training, evaluation, deployment, serving, monitoring, and continuous...  ...systems.Establish engineering patterns and best...  ..., deployment, inference, monitoring, and...  ...understanding of distributed systems, APIs and... 
    Full time
    Temporary work
    Part time
    Work experience placement

    Walmart

    Bellevue, WA
    3 days ago
  • $206.4k - $379.1k

     ...image, video, and 3D models built on each customer...  ...Principal Machine Learning Engineer to serve as the technical lead...  .... You will set the inference architecture and technical...  .... Director, ML Engineering and ML Engineering...  ...in Kubernetes, distributed systems, and MLOps platforms... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    Seattle, WA
    1 day ago
  • $197.3k - $313.7k

     ...for a Staff Machine Learning Engineer with deep expertise in model training and finetuning to join our ML team. You'll design, train, and...  ...at Slack ship models that serve millions of users daily. This...  ...with model optimization for inference (quantization, pruning, speculative... 
    Full time

    Salesforce

    Seattle, WA
    4 days ago
  •  ...Design and build ML platform orchestration...  ...FinOps. Build online model-serving lifecycle orchestration...  ...model and image distribution, deployment, upgrades...  ...computer science, software engineering, artificial...  ...model distribution, or inference performance analysis.... 
    Full time
    Internship

    ByteDance

    Seattle, WA
    1 day ago
  • $200k - $250k

     ...Machine Learning Engineering within the Advanced...  ...annotation pipelines, ML Infrastructure and...  ...state-of-art models into robust, autonomous...  ...infrastructure to serve computer vision...  ..., including distributed training infrastructure...  ..., and low-latency inference services. Ensure high... 
    Temporary work
    Work at office
    Local area

    Metropolis Corp

    Seattle, WA
    17 hours ago
  •  ...Responsibilities The Data-AML-Engine Orchestration team...  ...that powers online model serving across ByteDance...  ...infrastructure with production ML workloads.You will...  ...including model and image distribution, deployment, upgrades,...  ...distribution, or inference performance analysis.... 
    Internship

    ByteDance

    Seattle, WA
    4 days ago
  •  ...Inc. in Seattle is seeking a Software Engineer focused on Machine Learning and AI to build production-grade systems that serve real-time model inferences at scale. You will collaborate with...  ...shipping production software, 5+ years ML/AI experience, and a BS in CS or... 

    Apple

    Seattle, WA
    17 hours ago
  • Haus Analytics in Seattle is seeking a senior ML engineer to drive high-impact projects on the cMMM space, blending ML, causal inference, and scalable production code. You will...  ...cross-functional teams to deliver trustworthy models and scalable processes, while mentoring... 
    Flexible hours

    Haus Analytics

    Seattle, WA
    3 days ago
  •  ...Own the architecture for inference serving, GPU scheduling, Kubernetes operators...  ...evaluation systems for model, prompt, and agent changes, including...  .... Apply physics-informed ML and enterprise AI expertise...  ...evaluation practices, mentor engineers, and maintain architectural... 
    Full time
    Remote work
    Flexible hours

    AZX

    Seattle, WA
    12 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Engineer: Distributed Inference & Model Serving. Be the first to apply!