Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior ML Serving Engineer for LLMs & Inference

Alldus

A tech company in AI/ML is seeking a Senior Software Engineer specializing in ML Serving to build robust infrastructure for ML models. The ideal candidate has 5+ years of experience in software engineering, with a focus on ML serving. Proficiency in Python and knowledge of various serving frameworks are essential. This full-time role is located in San Jose, California and offers a competitive salary.#J-18808-Ljbffr

Vacancy posted 6 hours ago
Similar jobs that could be interesting for youBased on the Senior ML Serving Engineer for LLMs & Inference in San Jose, CA vacancy
  • $184k - $287.5k

     ...computing. An era where our GPU serves as the intelligence behind...  ...systems at scale. We seek a Senior ML Engineer to compose and deliver next-...  ...large language models (LLMs), vision-language models (VLMs...  ...quantization, and real-time inference optimization for production... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $193.3k - $261.5k

     ...team is looking for a Senior Software Development Engineer to own the design...  ...implementation of our inference data plane. We build...  ...data movement, and serving integration.Our work...  ...kernels for a custom ML accelerator...  ...development tools - uses LLMs or code-generation agents... 
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    17 hours ago
  • $174.72k - $295.68k

     ...tuning, PTQ, QAT, on-vehicle inference and related fields.Key ResponsibilitiesDevelop...  ...QAT orchestration workflows.Serve as the primary interface with...  ...programming and software engineering skills.Ability to work...  ....Experience deploying LLMs on resource-constrained or heterogeneous... 
    Senior
    Full time

    XPENG Motors

    Santa Clara, CA
    1 day ago
  • $151.8k - $265.35k

     ...verticals. We are hiring a Senior Machine Learning Engineer to build the...  ...including finetuned LLMs, image and video generation...  ..., all while ensuring served quality matches the...  ...SLAs. Run production ML operationally - on-call...  ...of production ML or inference services at scale.... 
    Senior
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    1 day ago
  • $201.3k - $352.3k

     ...DescriptionIt all started when engineer Fred Luddy wrote code that...  ...behind modern deep learning, LLMs, and agent architectures. Hands...  ...— model integration, APIs, serving infrastructure, and the application...  ...to LLM fine-tuning or inference optimization in production. Why... 
    Senior
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    2 days ago
  •  ...Description: We are looking for a senior ML infrastructure engineer to build and evolve the systems that...  ...infrastructure, deployment workflows, model serving, platform tooling, automation, and...  ...a core focus, and experience with inference deployment is a strong plus. We care... 
    Senior

    Maxinsights

    Santa Clara, CA
    28 days ago
  • $195k - $230k

     ...RoleWe are looking for a Senior Machine Learning Engineer to help evolve our...  ...and ranking systems serving tens of millions of users...  ...training online inference A/B experimentation metric...  ...ApplicationsApply LLMs and foundation models...  ...large-scale data and ML systems (e.g., Spark,... 
    Senior
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    2 days ago
  •  ...Apple Inc. is seeking a Sr. Machine Learning Engineer for the Foundation Models Inference team in Santa Clara, CA. You will collaborate with research and external partners to bring cutting-edge model architectures from prototype to planetary-scale deployment, owning hard... 
    Senior

    Apple Inc.

    Santa Clara, CA
    6 hours ago
  • $148.75k - $361k

     ...Learning, Experimentation, and Inference Platform that powers the...  ...a talented and experienced Senior Software Engineer, MLOps/DevOps, to join the Advertising...  ...platforms that accelerate ML experimentation and...  ...for critical ML training and serving infrastructurePartner with data... 
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  • $229.5k - $360k

     ...—it depends on a robust, flexible ML platform built for experimentation...  ...technology. Our work blends innovation, engineering excellence, and a deep commitment...  ...: feature store, real-time inference services, vector DBs, etc., that serve millions of transactions per secondRun... 
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    1 day ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software...  ...inference recipes for LLMs. A recipe defines which operators...  ...toolingExperience with ML accelerators with a basic...  ...to inference serving frameworks (vLLM, TRT-LLM... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    We are now looking for a Senior Deep Learning Architect, LLM Inference!NVIDIA is at the forefront...  ...Large Language Models (LLMs). If you're passionate about...  ...terms like disaggregated serving, data parallel attention,...  ....Collaborate with engineers from AI startup companies... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

    We are now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA is...  ...team building Generative AI inference platform to make design and...  ...develop open source software to serve inference of trained AI...  ...task costs for self-hosted LLMs.Build and evolve Dynamo’s... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...the world's most advanced LLMs, our proprietary models,...  ...with Moveworks’ Reasoning Engine and natural language...  ...help build cutting edge ML infrastructure for building and serving LLM’s at Moveworks. This...  ...distributed training and inference pipeline for large language... 
    Senior
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    3 days ago
  • $184.7k - $324.8k

     ...Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference Santa Clara, California, United States Machine Learning and AI...  ...powered with the largest foundation models.Our systems serve billions of queries daily across Siri AI, Apple Intelligence... 
    Senior
    Worldwide
    Relocation

    Apple Inc.

    Santa Clara, CA
    6 hours ago
  •  ...drive technical strategy for ML model development, training pipelines, and inference systems on SambaNova’s RDU and...  ...and long-context modeling. Serve as the senior technical voice in design reviews...  ...principal and senior ML engineers through collaboration, design... 
    Senior
    Full time
    Temporary work
    Flexible hours

    SambaNova Systems

    San Jose, CA
    3 days ago
  •  ...innovative AI startup is seeking a Founding ML Infrastructure Engineer to take charge of deploying and...  ...responsible for building and managing a full ML serving stack, working closely with product...  ...ML infrastructure, particularly with LLMs, and will be proficient in relevant... 

    Realmlabs

    Sunnyvale, CA
    14 hours ago
  • $184.7k - $324.8k

     ...experienced Machine Learning Engineer to build, operate, and...  ...expertise in model serving, deployment pipelines,...  ...distributed systems, and ML platform infrastructure...  ...human annotations or LLMs Design and implement...  ...Qualifications Familiarity with inference optimization techniques... 
    Senior
    Work experience placement
    Relocation

    Apple Inc.

    Cupertino, CA
    6 hours ago
  • $150k - $350k

     ...edge generative AI to assist engineers in RTL design, simulation,...  ...We are seeking an ML Systems Engineer to optimize...  ...efficiency of large language model inference powering our agentic AI...  ...deploying and optimizing LLMs in production: model serving, batching strategies,... 

    ChipAgents

    San Jose, CA
    3 days ago
  • $246.5k

     ...core of this is our Machine Learning and Inference Platform that powers the entire...  ...technical leader with deep experience in ML serving, high-performance computing, and industry...  ...frameworks - someone excited to mentor engineers, innovate at scale, and shape the future... 
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    1 day ago
  • $193.3k - $261.5k

     ...Inferentia and Trainium ML accelerators. This...  ...unparalleled ML inference and training...  ...software boundary, our engineers build systematic...  ...mentorship. Our senior members enjoy one-...  ...Machine learning and LLMs, their...  ...offline inference serving with vLLM, SGLang,... 
    Senior
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    17 hours ago
  •  ...: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The...  ...with other software (ML and compilers) and hardware...  ...with inference servers/model serving frameworks (such as TensorRT...  ...fundamentalsExperience deploying ML workloads (LLMs, VLMs, NLP, etc.) on... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    2 days ago
  • $223.61k - $268.33k

     ...fundamentally different from conventional ML engineering. The data spans on-chain activity, fiat...  ...pipelines, training workflows, model serving, decision integrations, monitoring,...  ...consistency between offline training and online inference. Ensure models and decision systems... 
    Senior
    Full time

    OKX

    San Jose, CA
    3 days ago
  • $206.4k - $379.1k

     ...Principal Machine Learning Engineer to serve as the technical lead...  ...role — it is the senior-most hands-on engineering...  .... You will set the inference architecture and technical...  ...model pipelines — LLMs, diffusion and transformer...  .... Director, ML Engineering and ML Engineering... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    2 days ago
  • $184.7k - $324.8k

     ...world? We truly believe it can! We are the ML Data Team of the Intelligent System...  ...looking for an exceptional software & data engineer who is passionate about Apple products and...  ...of industry experience as a data engineer serving various ML applications (vision domain preferred... 
    Senior
    Temporary work
    Relocation

    Apple Inc.

    Cupertino, CA
    4 days ago
  • $189k - $301k

     ...Conductor in San Jose, CA is seeking a seasoned engineer to lead co-design efforts for optimizing AI model inference performance. The role requires a deep understanding...  ..., covering everything from model definition to serving. The ideal candidate should have extensive... 
    Senior

    Conductor

    San Jose, CA
    6 hours ago
  • $119.25k - $150.85k

     ...the Model Deployment & Inference Solutions team deploys...  ...mission is to build the ML deployment platform that...  ...IQ, and we’re hiring engineers to help deliver the next...  ...work with and learn from senior engineers on real...  ...or distributed training/serving infrastructure. Familiarity... 
    Full time
    Internship
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

    We are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who are mindful...  ...language and multimodal model inference as part of NVIDIA Inference Microservices...  ...-LLM, NVIDIA’s open-source inference serving library.Profile and analyze bottlenecks... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $174.72k - $295.68k

     ...a full-time Machine Learning Engineer - AI Foundation, with deep knowledge...  ...establishing a state-of-art ML infrastructure for training...  ...and accelerating model training/inference. Our mission is to solve the...  ...Responsibilities:Optimize transformer-based LLMs for low-latency and high-... 
    Senior
    Full time

    XPENG Motors

    Santa Clara, CA
    2 days ago
  •  ...Apple Inc. is looking for an experienced Machine Learning Engineer in Cupertino, California, to develop and scale systems...  ...millions of users. In this role, you will handle model serving, deployment pipelines, and ML infrastructure, collaborating closely with... 
    Senior

    Apple Inc.

    Cupertino, CA
    6 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior ML Serving Engineer for LLMs & Inference. Be the first to apply!