Senior AI Model Serving Engineer — Low-Latency Inference
Jobleads-US
A leading data and AI company in San Francisco is seeking a Senior Engineer to enhance their Model Serving platform. This role requires expertise in building large-scale distributed systems and collaboration across teams to optimize performance and reliability. Ideal candidates will have a strong foundation in algorithms and system design, along with a passion for mentoring others. The position offers a competitive salary and generous benefits. #J-18808-Ljbffr Jobleads-US
- A leading data and AI company in San Francisco is seeking a Staff Engineer to design and implement systems for their AI/ML Model Serving platform. You will collaborate with product, infrastructure, and research teams to ensure high-performance system delivery. The ideal...Senior
$166k - $225k
...the world's best data and AI infrastructure platform so... ...their business. Databricks’ Model Serving product provides... ...models. It offers real-time, low-latency inference, governance, monitoring, and... ...and cost efficiency.As a Senior Engineer, you’ll play a critical role...SeniorLocal areaWorldwide$325k
A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate has over 5 years of software engineering experience, strong familiarity with ML architectures, and experience...Senior- MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability... ...in production-grade serving infrastructure, be fluent in Python...Senior
$190.8k - $267.1k
...partners closely with Ads Engineering teams to improve... ...domains including ad serving, auctions, targeting,... ...Experience operating high-QPS, low-latency services where latency... ...machine learning inference or recommendation systems... ...intelligence (AI). You will have the opportunity...SeniorFor contractorsWork experience placementRemote workFlexible hours$190k - $265k
...about enabling data and AI teams to solve the... .... Founded by engineers — and customer-obsessed... ...data apps, AI agents, model training, model serving, and Vector Search.... ...real-time and batch inference, powering model inference... ...reliability, latency, and efficiency of distributed...Local areaWorldwide$220k - $320k
Help us make inference blazingly fast. If you... ...specialized language models for companies that... ...frontier-quality AI at a fraction of... ...ten‑person team of engineers who work in‑person... ...with the goal of serving models faster and... ...inference performance: latency, throughput, cost...SeniorWork at office$193.4k - $290k
...frontier agentic AI, an enterprise-grade... ...a Staff Software Engineer on the Model Infrastructure... ...managementReliability and latency... ...high availability, low latency, and operational... ...excellence for AI inference.Design and... ...infrastructure, LLM serving, or machine learning...Senior- ...San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting... ...physical observations. You will design techniques to improve latency and throughput, optimize the inference stack to...Senior
$300 per month
...vertically integrated AI infrastructure... ...our Production Engineering team ensures... ...looking for a Senior Production... ...large language models to help us build... ...compute-intensive, latency-sensitive... ...with a focus on serving and scaling LLM... ...pipelines and inference servicesDefine,...SeniorTemporary work- ...Responsibilities As a senior Machine Learning Systems Engineer on the Search... ...implement scalable search serving infrastructure,... ...of high-throughput, low-latency search systems that... ...relevance quality.ML Model Development & ServingBuild... ...with Rovo and AI platform teams to evolve...SeniorWork at officeLocal area
$250k
...Ready to architect AI infrastructure... ...a serverless inference platform, beginning... ...expanding into low-latency, real-time inference and custom model hosting. This is... ...chance to join as a Senior Inference Platform Engineer at an early... ...latest models, serving frameworks, and...SeniorFull time- ...is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in... ...infrastructure that handles billions of inference requests, optimizing for latency, throughput, cost, and... ...responsibilities spanning scalable services, model serving, load balancing,...Senior
- ...powers mission-critical inference for the world's most dynamic AI companies, like... ...bring cutting-edge models into production. We... ...build the platform engineers turn to to ship AI... ...distributed systems, model serving, and developer... ...record of owning low‑latency, reliable backend...Full timeFlexible hours
- ...organization develops AI-native silicon and... ...of frontier models. By co-designing chips... ...within the inference engine that executes complex... ...layers of the cluster serving software stack,... ...optimizing for throughput, latency, utilization, and... ...retaining the low-level control...Full time
- ...About the Team Our Inference team brings OpenAI’s most capable research... ...access our start-of-the-art AI models, allowing them to do things... ...We are looking for an engineer who wants to take the world'... ...them for use in a high-volume, low-latency, and high-availability...Full time
$300 per month
...vertically integrated AI infrastructure... ...Crusoe, our Production Engineering team ensures the... ...’re looking for a Senior Production... ...compute-intensive, latency-sensitive workloads... ...services with a focus on serving and scaling LLM... ...training and inference clustersAutomate observability...SeniorTemporary work- ...is a leader in foundational AI models for image and video. Based in... ...presence in San Francisco, we seek engineers who can bridge research... ...scalable APIs and optimize GPU inference, collaborating with... ...performance, and production ML serving, offering a hybrid setup with...Senior
$192k - $260k
...world's best data and AI infrastructure platform... ...business. Foundation Model Serving is the API Product for... ...serving frontier AI model inference for open source models... .... We’re looking for engineers who have owned high... ...enable high-throughput, low-latency inference on GPU...Local areaWorldwide- ...growing fast in scope and in stakes. Models increasingly drive real-time... ...through deployment, real-time inference, observability, and retraining.... ...burden themselves. We also serve low-latency, highly available scores to the decision engine that depends on them. The platform...Senior
- ...worldwide.We’re a team of engineers, clinicians, and... ...Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be... ...Responsibilities• GPU & Model Performance & Optimization... ...models to real-time onboard inference—while serving as a core contributor to...SeniorLocal areaWorldwideFlexible hours
- ...Sciforium is an AI infrastructure... ...multimodal AI models and a proprietary... ...-efficiency serving platform. Backed... ...support from AMD engineers the team is... ...to market. As a senior technical leader... ...and distributed inference systems. Develop... ...and ensure low-latency, scalable inference...Full timeWork at officeFlexible hours
$298k - $368k
...a diverse set of sensors, enabling engineers like you to (1) develop methods for... ...scale real-world data, to (2) develop models and model training at scale, to (3)... ...foundation models). ~ Proven expertise in low-latency on-device inference techniques and a deep understanding...SeniorFull timeRemote work- Xcede is looking for a Member of Technical Staff focused on AI Safety to lead red-teaming efforts and ensure the robustness of next... ...should have deep expertise in LLM safety, strong software engineering skills, and relevant academic qualifications in AI or related fields...Senior
$206.3k - $388k
...Principal ML Engineer to architect... ...multimodal foundation models (image, video... .... This is a senior individual... ...Scale up inference throughput across... ..., index, and serve billions of... ..., throughput/latency tradeoffs)... ...stack, from low-level systems... ..., powered by AI and driven by...Full timeTemporary workLocal areaWorldwide$180k - $300k
...creating the generative models that power how... ...becoming usable APIs Inference is slower than it... ...-performance APIs serving millions of requests... ...Optimize inference latency and throughput... ...includes: Real-time or low-latency inference... ...edge of generative AI. #J-18808-Ljbffr Black...Remote workWorldwide2 days per week- A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and... ...Ideal candidates should have strong software engineering skills and experience with ML inference...Senior
$220k
...Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience...Senior$156k - $195k
Ironclad is the leading AI contracting platform... ...Ironclad is hiring a Senior Finance Systems & AI Engineer to own the data... ...and maintain reliable, low-manual-touch integrations... ...pipelines that serve as the analytics layer... ...platform end-to-end: model architecture, formula...SeniorFull timeContract work$188k - $275k
...The Essential Cloud for AI™. Built for pioneers by... ...'ll Do: The Field Engineering organization at CoreWeave... ...can train and inference on at scale, spanning infrastructure... ...role, you will: Serve as the primary... ...Define and operationalize models for managing customer bare...SeniorPermanent employmentFull timeContract workTemporary workCasual workWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Model Serving Engineer — Low-Latency Inference. Be the first to apply!
- ai research engineer San Francisco, CA
- senior ai engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- ai engineer remote San Francisco, CA
- ai developer San Francisco, CA
- ai prompt engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- ai engineer San Francisco, CA
- senior technical analyst San Francisco, CA
- senior associate attorney San Francisco, CA




