Senior AI Model Serving Engineer — Low-Latency Inference
Jobleads-US
A leading data and AI company in San Francisco is seeking a Senior Engineer to enhance their Model Serving platform. This role requires expertise in building large-scale distributed systems and collaboration across teams to optimize performance and reliability. Ideal candidates will have a strong foundation in algorithms and system design, along with a passion for mentoring others. The position offers a competitive salary and generous benefits. #J-18808-Ljbffr Jobleads-US
- A leading data and AI company in San Francisco is seeking a Staff Engineer to design and implement systems for their AI/ML Model Serving platform. You will collaborate with product, infrastructure, and research teams to ensure high-performance system delivery. The ideal...Senior
$166k - $225k
...the world's best data and AI infrastructure platform so... ...their business. Databricks’ Model Serving product provides... ...models. It offers real-time, low-latency inference, governance, monitoring, and... ...and cost efficiency.As a Senior Engineer, you’ll play a critical role...SeniorLocal areaWorldwide$325k
A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate has over 5 years of software engineering experience, strong familiarity with ML architectures, and experience...Senior- MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability... ...in production-grade serving infrastructure, be fluent in Python...Senior
$220k - $320k
...Help us make inference blazingly fast. If you... ...specialized language models for companies that... ...frontier-quality AI at a fraction of... ...ten‑person team of engineers who work in‑person... ...with the goal of serving models faster and... ...inference performance: latency, throughput, cost...SeniorWork at office$190k - $265k
...about enabling data and AI teams to solve the... .... Founded by engineers — and customer-obsessed... ...data apps, AI agents, model training, model serving, and Vector Search.... ...real-time and batch inference, powering model inference... ...reliability, latency, and efficiency of distributed...Local areaWorldwide$300 per month
...vertically integrated AI infrastructure... ...our Production Engineering team ensures... ...looking for a Senior Production... ...large language models to help us build... ...compute-intensive, latency-sensitive... ...with a focus on serving and scaling LLM... ...pipelines and inference servicesDefine,...SeniorTemporary work- ...Responsibilities As a senior Machine Learning Systems Engineer on the Search... ...implement scalable search serving infrastructure,... ...of high-throughput, low-latency search systems that... ...relevance quality.ML Model Development & ServingBuild... ...with Rovo and AI platform teams to evolve...SeniorWork at officeLocal area
- ...San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting... ...physical observations. You will design techniques to improve latency and throughput, optimize the inference stack to...Senior
$218.4k - $273k
...bottleneck in Physical AI. This position will... ...Robotics data and model evaluation. We... ...and experienced Senior Machine Learning Engineer, Computer Vision to... ...deployment on edge devices (low power, low latency) and/or large-scale... ...Leadership: Serve as the subject...SeniorFull time$250k
...Ready to architect AI infrastructure... ...a serverless inference platform, beginning... ...expanding into low-latency, real-time inference and custom model hosting. This is... ...chance to join as a Senior Inference Platform Engineer at an early... ...latest models, serving frameworks, and...SeniorFull time- MakerMaker in San Francisco is seeking a Senior ML systems engineer to build and operate production inference systems for large models. You will own performance, profiling, and... ...optimizations to ensure high throughput and low latency in production. You will collaborate with...Senior
$166.6k - $208.3k
...growing fast in scope and in stakes. Models increasingly drive real-time... ...through deployment, real-time inference, observability, and retraining.... ...burden themselves — and we serve low-latency, highly available scores to the decision engine that depends on them. The platform...Senior- ...is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in... ...infrastructure that handles billions of inference requests, optimizing for latency, throughput, cost, and... ...responsibilities spanning scalable services, model serving, load balancing,...Senior
$250k - $300k
...vertically integrated AI infrastructure... ...large language models run faster,... ...owning the inference stack end to end... ...deep into the serving code when the defaults... ...patterns, latency targets, and cost... ...with customer engineering teams to tailor... ...profiling, and low-level...SeniorTemporary work- ...powers mission-critical inference for the world's most dynamic AI companies, like... ...bring cutting-edge models into production. We... ...build the platform engineers turn to to ship AI... ...distributed systems, model serving, and developer... ...record of owning low‑latency, reliable backend...Full timeFlexible hours
- ...About the Team Our Inference team brings OpenAI’s most capable research... ...access our start-of-the-art AI models, allowing them to do things... ...We are looking for an engineer who wants to take the world'... ...them for use in a high-volume, low-latency, and high-availability...Full time
$300 per month
...vertically integrated AI infrastructure... ...Crusoe, our Production Engineering team ensures the... ...’re looking for a Senior Production... ...compute-intensive, latency-sensitive workloads... ...services with a focus on serving and scaling LLM... ...training and inference clustersAutomate observability...SeniorTemporary work- ...A tech startup in AI model serving located in San Francisco is seeking a qualified candidate to architect scalable inference systems. The role focuses on optimizing model serving performance... ...in Python and PyTorch, along with low-level systems knowledge. This in-person...
- ...is a leader in foundational AI models for image and video. Based in... ...presence in San Francisco, we seek engineers who can bridge research... ...scalable APIs and optimize GPU inference, collaborating with... ...performance, and production ML serving, offering a hybrid setup with...Senior
$192k - $260k
...world's best data and AI infrastructure platform... ...business. Foundation Model Serving is the API Product for... ...serving frontier AI model inference for open source models... .... We’re looking for engineers who have owned high... ...enable high-throughput, low-latency inference on GPU...Local areaWorldwide- ...worldwide.We’re a team of engineers, clinicians, and... ...Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be... ...Responsibilities• GPU & Model Performance & Optimization... ...models to real-time onboard inference—while serving as a core contributor to...SeniorLocal areaWorldwideFlexible hours
$160k - $230k
...the RoleAt Together.ai, we are building state... ...and scalable inference for large language models (LLMs). Our mission... ...Frameworks and Optimization Engineer to design, develop,... ...role will focus on low-latency, high-throughput inference... ...high-performance serving.Apply CUDA graph...Full time$206.3k - $388k
...Principal ML Engineer to architect... ...multimodal foundation models (image, video... .... This is a senior individual... ...Scale up inference throughput across... ..., index, and serve billions of... ..., throughput/latency tradeoffs)... ...stack, from low-level systems... ..., powered by AI and driven by...Full timeTemporary workLocal areaWorldwide$182k - $242k
...Essential Cloud for AI™. Built for... ...internal and customer engineering teams, offering valuable... ...role, you will: Serve as the primary... ...contributions to open-source inference frameworks... ...publications/talks on latency, optimization, or advanced model-server architectures...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$298k - $368k
...a diverse set of sensors, enabling engineers like you to (1) develop methods for... ...scale real-world data, to (2) develop models and model training at scale, to (3)... ...foundation models). ~ Proven expertise in low-latency on-device inference techniques and a deep understanding...SeniorFull timeRemote work- Xcede is looking for a Member of Technical Staff focused on AI Safety to lead red-teaming efforts and ensure the robustness of next... ...should have deep expertise in LLM safety, strong software engineering skills, and relevant academic qualifications in AI or related fields...Senior
$180k - $237.5k
...17, we’re delivering low-cost and large-scale... ...already have. Senior Software Engineer – Site Controller, Energy... ...that ensure low-latency, reliable data flow between... ..."Pack Manager" to serve as a universal translator... ...information, and inferences drawn from your PI. We...SeniorFull timeLocal area- ...HP IQ, HP’s AI innovation lab, seeks a firmware engineer to design and develop firmware for low‑power ARM‑based MCUs, coordinating with EE, RF, Security, and Cloud teams to deliver secure, high‑performance systems. You will bring up hardware in the lab, support remote...SeniorRemote work
- A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and... ...Ideal candidates should have strong software engineering skills and experience with ML inference...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Model Serving Engineer — Low-Latency Inference. Be the first to apply!
- ai ml engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- ai developer San Francisco, CA
- ai engineer San Francisco, CA
- ai research engineer San Francisco, CA
- ai prompt engineer San Francisco, CA
- senior ai engineer San Francisco, CA
- ai engineer remote San Francisco, CA
- senior lead project manager San Francisco, CA
- senior robotics software engineer San Francisco, CA





