Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior AI Model Serving Engineer — Low-Latency Inference

Jobleads-US

A leading data and AI company in San Francisco is seeking a Senior Engineer to enhance their Model Serving platform. This role requires expertise in building large-scale distributed systems and collaboration across teams to optimize performance and reliability. Ideal candidates will have a strong foundation in algorithms and system design, along with a passion for mentoring others. The position offers a competitive salary and generous benefits. #J-18808-Ljbffr Jobleads-US

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior AI Model Serving Engineer — Low-Latency Inference in San Francisco, CA vacancy
  • A leading data and AI company in San Francisco is seeking a Staff Engineer to design and implement systems for their AI/ML Model Serving platform. You will collaborate with product, infrastructure, and research teams to ensure high-performance system delivery. The ideal... 
    Senior

    Jobleads-US

    San Francisco, CA
    2 days ago
  • $166k - $225k

     ...the world's best data and AI infrastructure platform so...  ...their business. Databricks’ Model Serving product provides...  ...models. It offers real-time, low-latency inference, governance, monitoring, and...  ...and cost efficiency.As a Senior Engineer, you’ll play a critical role... 
    Senior
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  • $325k

    A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate has over 5 years of software engineering experience, strong familiarity with ML architectures, and experience... 
    Senior

    Jobleads-US

    San Francisco, CA
    4 days ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability...  ...in production-grade serving infrastructure, be fluent in Python... 
    Senior

    MakerMaker.AI

    San Francisco, CA
    1 day ago
  • $220k - $320k

     ...Help us make inference blazingly fast. If you...  ...specialized language models for companies that...  ...frontier-quality AI at a fraction of...  ...ten‑person team of engineers who work in‑person...  ...with the goal of serving models faster and...  ...inference performance: latency, throughput, cost... 
    Senior
    Work at office

    SOLANA FOUNDATION

    San Francisco, CA
    4 days ago
  • $190k - $265k

     ...about enabling data and AI teams to solve the...  .... Founded by engineers — and customer-obsessed...  ...data apps, AI agents, model training, model serving, and Vector Search....  ...real-time and batch inference, powering model inference...  ...reliability, latency, and efficiency of distributed... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    16 hours ago
  • $300 per month

     ...vertically integrated AI infrastructure...  ...our Production Engineering team ensures...  ...looking for a Senior Production...  ...large language models to help us build...  ...compute-intensive, latency-sensitive...  ...with a focus on serving and scaling LLM...  ...pipelines and inference servicesDefine,... 
    Senior
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  •  ...Responsibilities As a senior Machine Learning Systems Engineer on the Search...  ...implement scalable search serving infrastructure,...  ...of high-throughput, low-latency search systems that...  ...relevance quality.ML Model Development & ServingBuild...  ...with Rovo and AI platform teams to evolve... 
    Senior
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    3 days ago
  •  ...San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting...  ...physical observations. You will design techniques to improve latency and throughput, optimize the inference stack to... 
    Senior

    Causal Labs

    San Francisco, CA
    2 days ago
  • $218.4k - $273k

     ...bottleneck in Physical AI. This position will...  ...Robotics data and model evaluation. We...  ...and experienced Senior Machine Learning Engineer, Computer Vision to...  ...deployment on edge devices (low power, low latency) and/or large-scale...  ...Leadership: Serve as the subject... 
    Senior
    Full time

    Scale Ai

    San Francisco, CA
    16 hours ago
  • $250k

     ...Ready to architect AI infrastructure...  ...a serverless inference platform, beginning...  ...expanding into low-latency, real-time inference and custom model hosting. This is...  ...chance to join as a Senior Inference Platform Engineer at an early...  ...latest models, serving frameworks, and... 
    Senior
    Full time
    San Francisco, CA
    more than 2 months ago
  • MakerMaker in San Francisco is seeking a Senior ML systems engineer to build and operate production inference systems for large models. You will own performance, profiling, and...  ...optimizations to ensure high throughput and low latency in production. You will collaborate with... 
    Senior

    MakerMaker

    San Francisco, CA
    2 days ago
  • $166.6k - $208.3k

     ...growing fast in scope and in stakes. Models increasingly drive real-time...  ...through deployment, real-time inference, observability, and retraining....  ...burden themselves — and we serve low-latency, highly available scores to the decision engine that depends on them. The platform... 
    Senior

    Mercury

    San Francisco, CA
    3 days ago
  •  ...is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in...  ...infrastructure that handles billions of inference requests, optimizing for latency, throughput, cost, and...  ...responsibilities spanning scalable services, model serving, load balancing,... 
    Senior

    Inception

    San Francisco, CA
    4 days ago
  • $250k - $300k

     ...vertically integrated AI infrastructure...  ...large language models run faster,...  ...owning the inference stack end to end...  ...deep into the serving code when the defaults...  ...patterns, latency targets, and cost...  ...with customer engineering teams to tailor...  ...profiling, and low-level... 
    Senior
    Temporary work

    Crusoe

    San Francisco, CA
    16 hours ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like...  ...bring cutting-edge models into production. We...  ...build the platform engineers turn to to ship AI...  ...distributed systems, model serving, and developer...  ...record of owning low‑latency, reliable backend... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    16 hours ago
  •  ...About the Team Our Inference team brings OpenAI’s most capable research...  ...access our start-of-the-art AI models, allowing them to do things...  ...We are looking for an engineer who wants to take the world'...  ...them for use in a high-volume, low-latency, and high-availability... 
    Full time

    OpenAI

    San Francisco, CA
    16 hours ago
  • $300 per month

     ...vertically integrated AI infrastructure...  ...Crusoe, our Production Engineering team ensures the...  ...’re looking for a Senior Production...  ...compute-intensive, latency-sensitive workloads...  ...services with a focus on serving and scaling LLM...  ...training and inference clustersAutomate observability... 
    Senior
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  •  ...A tech startup in AI model serving located in San Francisco is seeking a qualified candidate to architect scalable inference systems. The role focuses on optimizing model serving performance...  ...in Python and PyTorch, along with low-level systems knowledge. This in-person... 

    Reducto, Inc.

    San Francisco, CA
    20 hours ago
  •  ...is a leader in foundational AI models for image and video. Based in...  ...presence in San Francisco, we seek engineers who can bridge research...  ...scalable APIs and optimize GPU inference, collaborating with...  ...performance, and production ML serving, offering a hybrid setup with... 
    Senior

    Black Forest Labs

    San Francisco, CA
    4 days ago
  • $192k - $260k

     ...world's best data and AI infrastructure platform...  ...business. Foundation Model Serving is the API Product for...  ...serving frontier AI model inference for open source models...  .... We’re looking for engineers who have owned high...  ...enable high-throughput, low-latency inference on GPU... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  •  ...worldwide.We’re a team of engineers, clinicians, and...  ...Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be...  ...Responsibilities• GPU & Model Performance & Optimization...  ...models to real-time onboard inference—while serving as a core contributor to... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    2 days ago
  • $160k - $230k

     ...the RoleAt Together.ai, we are building state...  ...and scalable inference for large language models (LLMs). Our mission...  ...Frameworks and Optimization Engineer to design, develop,...  ...role will focus on low-latency, high-throughput inference...  ...high-performance serving.Apply CUDA graph... 
    Full time

    Together AI

    San Francisco, CA
    2 days ago
  • $206.3k - $388k

     ...Principal ML Engineer to architect...  ...multimodal foundation models (image, video...  .... This is a senior individual...  ...Scale up inference throughput across...  ..., index, and serve billions of...  ..., throughput/latency tradeoffs)...  ...stack, from low-level systems...  ..., powered by AI and driven by... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    4 days ago
  • $182k - $242k

     ...Essential Cloud for AI™. Built for...  ...internal and customer engineering teams, offering valuable...  ...role, you will: Serve as the primary...  ...contributions to open-source inference frameworks...  ...publications/talks on latency, optimization, or advanced model-server architectures... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    22 days ago
  • $298k - $368k

     ...a diverse set of sensors, enabling engineers like you to (1) develop methods for...  ...scale real-world data, to (2) develop models and model training at scale, to (3)...  ...foundation models). ~ Proven expertise in low-latency on-device inference techniques and a deep understanding... 
    Senior
    Full time
    Remote work

    Waymo

    San Francisco, CA
    16 hours ago
  • Xcede is looking for a Member of Technical Staff focused on AI Safety to lead red-teaming efforts and ensure the robustness of next...  ...should have deep expertise in LLM safety, strong software engineering skills, and relevant academic qualifications in AI or related fields... 
    Senior

    Xcede

    San Francisco, CA
    3 days ago
  • $180k - $237.5k

     ...17, we’re delivering low-cost and large-scale...  ...already have.   Senior Software Engineer – Site Controller, Energy...  ...that ensure low-latency, reliable data flow between...  ..."Pack Manager" to serve as a universal translator...  ...information, and inferences drawn from your PI. We... 
    Senior
    Full time
    Local area

    Redwood Materials

    San Francisco, CA
    16 hours ago
  •  ...HP IQ, HP’s AI innovation lab, seeks a firmware engineer to design and develop firmware for low‑power ARM‑based MCUs, coordinating with EE, RF, Security, and Cloud teams to deliver secure, high‑performance systems. You will bring up hardware in the lab, support remote... 
    Senior
    Remote work

    Hpiq

    San Francisco, CA
    20 hours ago
  • A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and...  ...Ideal candidates should have strong software engineering skills and experience with ML inference... 
    Senior

    Gimlet Labs

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior AI Model Serving Engineer — Low-Latency Inference. Be the first to apply!