Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

High-Performance ML Inference Engineer

Reactor

Reactor in San Francisco is seeking an ML Inference Engineer to maximize performance of generative media models and push ultra-low-latency, high-throughput inference. You will craft an in-house runtime, implement optimizations with PyTorch tools, and collaborate with partner teams to integrate external models. The role requires deep expertise in PyTorch, TensorRT, CUDA, and model optimization techniques, with a focus on delivering cutting-edge inference capabilities at scale. #J-18808-Ljbffr Reactor

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the High-Performance ML Inference Engineer in San Francisco, CA vacancy
  •  ...to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern inference frameworks and a solid... 
    Performance

    Reflection AI

    San Francisco, CA
    1 day ago
  •  ...technology company in San Francisco is seeking a Founding Engineer specializing in ML Inference. This highly technical role requires expertise in the ML...  ...infrastructure stack and aims to optimize generative media performance. The ideal candidate will drive innovations in real... 
    Performance
    Relocation package

    Reactor.am

    San Francisco, CA
    1 day ago
  • $180k - $250k

     ...company in San Francisco is looking for a Staff Software Engineer to enhance model performance for generative media models. The ideal candidate will design...  ...performance tools, and collaborate with the Applied ML team. Offering a competitive salary of $180,000 - $250,00... 
    Performance

    fal

    San Francisco, CA
    4 days ago
  • OpenAI in San Francisco seeks an experienced Software Engineer to help bring inference workloads to AWS Trainium and build the software stack to...  ..., compilers, and model execution. You will develop high‑performance kernels, improve compiler support, and ensure scalable... 
    Performance

    Slope

    San Francisco, CA
    3 days ago
  • Jaide Health is seeking an engineer for their Model...  ...focuses on building reliable ML systems while enhancing core performance metrics across model execution...  ...5 years of experience in high-performance coding, plus...  ...and insights into the LLM inference ecosystem. A commitment... 
    Performance
    Remote job

    Jaide Health

    San Francisco, CA
    4 days ago
  • $160k - $230k

     ...Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role...  ...about AI inference, PyTorch, and developing high-performance systems, we want to hear from... 
    Performance
    Full time

    Together Ai

    San Francisco, CA
    1 day ago
  • $180k - $270k

     ...professionals to elevate productivity and performance through note-taking solutions,...  ...building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational...  ...intersection between the core ML training team and the backend... 
    Performance
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    1 day ago
  • $155k - $180k

     ...(not only product and engineering), so Roboflow employs...  ...center of all of this is inference — one of our most...  ...ability to keep quality high and cut releases on a...  ...computer vision and ML models to our users....  ...contributors. ~ Level-up your performance with AI agents.... 
    Performance
    Full time
    Second job
    Remote work
    Work from home
    Relocation package
    Flexible hours
    Night shift

    Roboflow

    San Francisco, CA
    1 day ago
  • $250k

    Title : ML Inference Engineer Location : San Francisco, CA Salary : $250k base + equity An AI Unicorn...  ..., Python, and PyTorch. This is a highly autonomous role with significant ownership...  ...across inference systems and model performance in production. This role is hybrid in... 
    Performance
    Full time

    Oscar Technology

    San Francisco, CA
    2 days ago
  •  ...Associates Limited (US) in San Francisco offers a hybrid Senior ML Inference Engineer role. You will enhance AI-native infrastructure powering...  ..., with ownership across inference systems and production performance. The ideal candidate has 3+ years of experience, strong Python... 
    Performance

    Oscar Technology

    San Francisco, CA
    2 days ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability. The ideal candidate will have 3+ years of experience in production... 
    Performance

    MakerMaker.AI

    San Francisco, CA
    3 days ago
  •  ...Technical Staff to design and optimize inference systems. The role involves...  ...allocation and improving execution performance across various components. Ideal candidates...  ...should have strong software engineering skills and experience with ML inference systems, particularly in... 
    Performance

    Gimlet Labs

    San Francisco, CA
    4 days ago
  •  ...Member of Technical Staff focused on ML systems and inference in San Francisco. You will design...  ...ensuring fast, predictable, and scalable performance. Key responsibilities include...  ...have strong foundations in software engineering, experience with ML inference systems... 
    Performance

    Gimlet Labs, Inc.

    San Francisco, CA
    1 day ago
  •  ...About the role: As a ML Engineer, you’ll build and operate the...  ...that push the boundaries of performance, interpretability, and control...  ...might work on: Optimize inference and training throughput for...  ...architectures Build and maintain high-performance distributed... 
    Performance
    Full time
    Internship

    Tilde Research

    San Francisco, CA
    1 day ago
  •  ...models and a proprietary, high-efficiency serving...  ...hands-on support from AMD engineers the team is scaling...  ...About the role As an ML Engineer at Sciforium,...  ...architecting and shipping performance-critical systems....  ...TensorFlow, or JAX, or Ray and inference engines like TensorRT,... 
    Performance
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    1 day ago
  • $170.1k - $258.3k

     ...model export, kernel development, and performance engineering so that every cycle on our accelerators...  ...the TeamThe AI Kernels team builds high‑performance GPU kernels and custom libraries...  ...sit at the heart of our on‑vehicle ML inference for ADAS and autonomous driving. We... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    4 hours ago
  • $128.7k - $261.3k

     ...export, kernel development, and performance engineering so that every cycle on our accelerators...  ...We own the compiler that turns high‑level models into fast, reliable inference across GPUs powering GM’s next‑...  ..., reliable, and effortless for ML engineers across the AV... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    2 days ago
  •  ...About the OpportunityAn ultra-high-growth artificial...  ...Machine Learning Infrastructure Engineer to help architect the compute...  ...powering next-generation model performance. Backed by top-tier venture capital...  ...established distributed training or inference engines.Practical experience... 
    Performance
    Full time
    Work at office
    Flexible hours

    Objective Paradigm

    San Francisco, CA
    10 hours ago
  • $197.3k - $313.7k

     ...Staff Machine Learning Engineer with deep expertise in...  ...finetuning to join our ML team. You'll design, train...  ...!) user base.Produce high-quality results by leading...  ...model optimization for inference (quantization, pruning,...  ..., assessment of job performance, discipline, termination... 
    Performance
    Full time

    Salesforce

    San Francisco, CA
    3 days ago
  •  ...2 new Machine Learning Engineer opportunities posted on...  ...and optimize end-to-end ML pipelines encompassing...  ...quality, observability, and performance across all AI systems....  ...environments to ensure high performance and low...  ...for tuning training and inference end-to-end for high... 
    Performance
    Flexible hours

    AI Chopping Block, Inc.

    San Francisco, CA
    22 hours ago
  • $200k - $300k

     ...ML Infrastructure EngineerThe company is building...  ...As an ML Infrastructure Engineer, you'll own critical training and inference infrastructure powering...  ...executionExperience building highly scalable infrastructure...  ...infrastructure and performance bottlenecksBachelor's degree... 
    Performance

    Recruiting from Scratch

    San Francisco, CA
    5 days ago
  • $180k - $240k

     ...Senior Machine Learning Engineer - this is a fully...  ...join our small but mighty ML team building production...  ...systems that handle real, high-stakes conversations at...  ...research into high-performing systems and models that...  ...problems — low latency inference, hallucination reduction... 
    Performance
    Remote work
    Flexible hours

    Career Renew

    San Francisco, CA
    2 days ago
  •  ...with researchers and model engineers to translate ideas into...  .... This is a hands‑on, high‑leverage role at the intersection of ML, software engineering,...  ...You Will Own training/inference infrastructure: Design,...  ...minimal friction. Optimize performance: Profile and improve... 
    Performance
    Full time

    Monograph

    San Francisco, CA
    3 days ago
  • Position: Senior ML Performance Engineer Location: SF Bay Area (US) or Toronto (Canada) - Hybrid Employment...  ...infrastructure company is building a high-performance, portable compiler...  ...performance testing platform for LLM inference workloads across GPU clusters Define... 
    Performance
    Full time

    Amadeus Search

    San Francisco, CA
    3 days ago
  •  ...flight data, interpret human performance, and deliver feedback that...  ...and more effective. As an AI/ML Engineer, you’ll design and deploy the...  ...airline operations to high-performance fighter jet sorties...  ...backend systems for scalable inference across schools, airlines, and... 
    Performance
    Full time
    Weekend work

    Navi Ai

    San Francisco, CA
    1 day ago
  • $189.6k - $237k

    Scale’s ML platform (RLXF) team builds our internal...  ...model training and inference. The platform has been...  ...software engineering skills, proficient in...  ...qualifications, interview performance, and relevant education...  ...products provide the high-quality data and full-... 
    Performance
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  • Together AI is building the best inference infrastructure for voice applications. We seek a Staff ML Engineer to own the model serving stack and optimize latency and throughput...  ...'s platform. This is a foundational role on a small, high-impact team. #J-18808-Ljbffr Together

    Together

    San Francisco, CA
    1 day ago
  •  ...Talent Network for Principal AI/ML Engineer Roles at VC-Backed Startups...  ...AI-powered systems at high-growth startups. By joining...  ...Develop scalable training and inference pipelines for AI-powered applications...  ..., and optimization for performance ~ Strong knowledge of... 
    Performance
    Full time

    Signal Fire Inc

    San Francisco, CA
    1 day ago
  • $300k - $400k

     ...Role Description   As a Principal AI/ML Engineer in our AdTech team, you will be a key...  ...science teams to ensure our ML systems are highly performant, scalable, and reliable. You will also...  ...ingestion and training to real-time inference , for our real-time bidding,... 
    Performance
    Full time

    Zeta Global

    San Francisco, CA
    1 day ago
  • $209k - $313k

     ...and other digital services.Snap Engineering teams build fun and...  ...interpretabilityConduct code reviews, maintain high engineering standards, and...  ...understanding of causal inference and modern approaches to estimating...  ...tests) and leveraging causal ML in production systemsPreferred... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to High-Performance ML Inference Engineer. Be the first to apply!