Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff ML Performance Engineer High-Throughput Inference

$180k - $250k

fal

A tech company in San Francisco is looking for a Staff Software Engineer to enhance model performance for generative media models. The ideal candidate will design advanced model serving architectures, develop performance tools, and collaborate with the Applied ML team. Offering a competitive salary of $180,000 - $250,000 with equity and comprehensive health benefits, this position provides substantial growth opportunities. #J-18808-Ljbffr fal

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Staff ML Performance Engineer High-Throughput Inference in San Francisco, CA vacancy
  • Reactor in San Francisco is seeking an ML Inference Engineer to maximize performance of generative media models and push ultra-low-latency, high-throughput inference. You will craft an in-house runtime, implement optimizations with PyTorch tools, and collaborate with partner... 
    Performance

    Reactor

    San Francisco, CA
    3 days ago
  •  ...to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern inference frameworks and a solid... 
    Performance

    Reflection AI

    San Francisco, CA
    1 day ago
  • $180k - $270k

     ...professionals to elevate productivity and performance through note-taking solutions,...  ...building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational...  ...intersection between the core ML training team and the backend... 
    Performance
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    1 day ago
  •  ...technology company in San Francisco is seeking a Founding Engineer specializing in ML Inference. This highly technical role requires expertise in the ML...  ...infrastructure stack and aims to optimize generative media performance. The ideal candidate will drive innovations in real... 
    Performance
    Relocation package

    Reactor.am

    San Francisco, CA
    1 day ago
  • OpenAI in San Francisco seeks an experienced Software Engineer to help bring inference workloads to AWS Trainium and build the software stack to...  ..., compilers, and model execution. You will develop high‑performance kernels, improve compiler support, and ensure scalable... 
    Performance

    Slope

    San Francisco, CA
    3 days ago
  • Jaide Health is seeking an engineer for their Model...  ...focuses on building reliable ML systems while enhancing core performance metrics across model execution...  ...5 years of experience in high-performance coding, plus...  ...and insights into the LLM inference ecosystem. A commitment... 
    Performance
    Remote job

    Jaide Health

    San Francisco, CA
    4 days ago
  •  ...About the role: As a ML Engineer, you’ll build and operate...  ...that push the boundaries of performance, interpretability, and...  ...might work on: Optimize inference and training throughput for novel model...  ...architectures Build and maintain high-performance distributed... 
    Performance
    Full time
    Internship

    Tilde Research

    San Francisco, CA
    1 day ago
  • $160k - $230k

     ...Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role...  ...about AI inference, PyTorch, and developing high-performance systems, we want to hear from... 
    Performance
    Full time

    Together Ai

    San Francisco, CA
    1 day ago
  • $170.1k - $258.3k

     ...export, kernel development, and performance engineering so that every cycle on our...  ...AI Kernels team builds high‑performance GPU kernels and...  ...the heart of our on‑vehicle ML inference for ADAS and autonomous...  ...consistently meeting strict latency, throughput, and reliability targets.... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    6 hours ago
  • $128.7k - $261.3k

     ...kernel development, and performance engineering so that every cycle on...  ...the compiler that turns high‑level models into fast, reliable inference across GPUs powering GM...  ..., and effortless for ML engineers across the AV...  ...simultaneously optimize compilation throughput, model fidelity, and on... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    2 days ago
  •  ...the OpportunityAn ultra-high-growth artificial...  ...Learning Infrastructure Engineer to help architect the...  ...next-generation model performance. Backed by top-tier venture...  ...and sustain high-throughput execution and training...  ...distributed training or inference engines.Practical experience... 
    Performance
    Full time
    Work at office
    Flexible hours

    Objective Paradigm

    San Francisco, CA
    12 hours ago
  • $155k - $180k

     ...(not only product and engineering), so Roboflow employs...  ...center of all of this is inference — one of our most...  ...ability to keep quality high and cut releases on a...  ...computer vision and ML models to our users....  ...contributors. ~ Level-up your performance with AI agents.... 
    Performance
    Full time
    Second job
    Remote work
    Work from home
    Relocation package
    Flexible hours
    Night shift

    Roboflow

    San Francisco, CA
    1 day ago
  •  ...new Machine Learning Engineer opportunities posted...  ...optimize end-to-end ML pipelines...  ...observability, and performance across all AI systems...  ...environments to ensure high performance and low...  ...tuning training and inference end-to-end for high throughput. Design data control... 
    Performance
    Flexible hours

    AI Chopping Block, Inc.

    San Francisco, CA
    23 hours ago
  • $200k - $300k

     ...ML Infrastructure EngineerThe company...  ...ML Infrastructure Engineer, you'll own critical training and inference infrastructure powering...  ...latency, throughput, reliability, and...  ...executionExperience building highly scalable...  ...infrastructure and performance bottlenecksBachelor... 
    Performance

    Recruiting from Scratch

    San Francisco, CA
    5 days ago
  • $250k

    Title : ML Inference Engineer Location : San Francisco, CA Salary : $250k base + equity An AI Unicorn...  ..., Python, and PyTorch. This is a highly autonomous role with significant ownership...  ...across inference systems and model performance in production. This role is hybrid in... 
    Performance
    Full time

    Oscar Technology

    San Francisco, CA
    2 days ago
  •  ...researchers and model engineers to translate ideas...  ...is a hands‑on, high‑leverage role at...  ...the intersection of ML, software...  ...Will Own training/inference infrastructure: Design...  ...friction. Optimize performance: Profile and improve...  ...utilization, throughput, and distributed synchronization... 
    Performance
    Full time

    Monograph

    San Francisco, CA
    3 days ago
  • Position: Senior ML Performance Engineer Location: SF Bay Area (US) or Toronto...  ...infrastructure company is building a high-performance, portable...  ...testing platform for LLM inference workloads across GPU...  ...and test suites (latency, throughput, memory utilization, power... 
    Performance
    Full time

    Amadeus Search

    San Francisco, CA
    3 days ago
  •  ...Associates Limited (US) in San Francisco offers a hybrid Senior ML Inference Engineer role. You will enhance AI-native infrastructure powering...  ..., with ownership across inference systems and production performance. The ideal candidate has 3+ years of experience, strong Python... 
    Performance

    Oscar Technology

    San Francisco, CA
    2 days ago
  • Together AI is building the best inference infrastructure for voice applications. We seek a Staff ML Engineer to own the model serving stack and optimize latency and throughput for real-time voice workloads....  ...foundational role on a small, high-impact team. #J-18808-Ljbffr... 

    Together

    San Francisco, CA
    1 day ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability. The ideal candidate will have 3+ years of experience in production... 
    Performance

    MakerMaker.AI

    San Francisco, CA
    3 days ago
  • $300k - $400k

     ...Description   As a Principal AI/ML Engineer in our AdTech team, you...  ...ensure our ML systems are highly performant, scalable, and reliable....  ...and training to real-time inference , for our real-time bidding...  ...handle low-latency, high-throughput requirements.   ~ Technical... 
    Performance
    Full time

    Zeta Global

    San Francisco, CA
    1 day ago
  •  ...seeking a Member of Technical Staff to design and optimize inference systems. The role involves...  ...allocation and improving execution performance across various components....  ...should have strong software engineering skills and experience with ML inference systems, particularly... 
    Performance

    Gimlet Labs

    San Francisco, CA
    4 days ago
  •  ...looking for a Member of Technical Staff focused on ML systems and inference in San Francisco. You will design...  ...ensuring fast, predictable, and scalable performance. Key responsibilities include...  ...strong foundations in software engineering, experience with ML inference systems... 
    Performance

    Gimlet Labs, Inc.

    San Francisco, CA
    1 day ago
  • $200k - $260k

     ...is building the best inference infrastructure for voice...  ...looking for a Senior ML Engineer to drive the model serving...  ...— pushing latency and throughput to the frontier. You'...  ...hire on a small, high-impact team. Voice inference...  ...Optimize inference performance for voice models (STT,... 
    Performance
    Full time

    Together Ai

    San Francisco, CA
    1 day ago
  •  ...and a proprietary, high-efficiency serving...  ...support from AMD engineers the team is scaling...  ...and deployment + inference optimization ....  ...architectures Improve throughput, memory efficiency...  ...to enhance performance Sandbox Environments...  ...to-end production ML systems with... 
    Performance
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    1 day ago
  •  ...models and a proprietary, high-efficiency serving...  ...hands-on support from AMD engineers the team is scaling...  ...About the role As an ML Engineer at Sciforium,...  ...architecting and shipping performance-critical systems....  ...TensorFlow, or JAX, or Ray and inference engines like TensorRT,... 
    Performance
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    1 day ago
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 
    Performance

    Causal Labs

    San Francisco, CA
    4 days ago
  • $148.7k - $199.4k

     ...Senior Machine Learning Engineer - ESPNReq ID:10150610...  ...operating distributed data and ML infrastructure that supports high‑throughput, low‑latency data...  ...services such as inference inputs, feature APIs, and...  ...improve system availability, performance, and cost efficiency.4)... 
    Performance
    Full time
    Worldwide

    Hulu

    San Francisco, CA
    3 days ago
  •  ...new initiatives to building highly-scalable enterprise-grade AI...  ...developers and enterprise users.ML performance, quality, and systems acumen...  ...7+ years of software engineering experience building and operating...  ...and improving latency, throughput, saturation, error rates, availability... 
    Performance
    Work at office

    Atlassian

    San Francisco, CA
    2 days ago
  •  ...Francisco is seeking a GPU kernel engineer with deep CUDA expertise to accelerate our training and inference workloads. You will write, optimize, and maintain high-performance kernels close to the metal....  ...engineers to squeeze throughput and reduce latency, with the... 
    Performance
    Work at office

    TypeSafe AI

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff ML Performance Engineer High-Throughput Inference. Be the first to apply!