Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Scale Low-Latency ML Inference Engineer (GPU/CUDA)

AI Breaking Wire

AI Breaking Wire seeks a Machine Learning Engineer to join our Inference Infrastructure team. You will build and optimize the high-throughput, low-latency distributed systems that power models like GPT-4 and Sora for millions of users worldwide. The role emphasizes C++ and Python proficiency, GPU optimization with CUDA and Triton, and collaboration with research teams to productionize new architectures. This hybrid position is based in San Francisco. #J-18808-Ljbffr AI Breaking Wire

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Scale Low-Latency ML Inference Engineer (GPU/CUDA) in San Francisco, CA vacancy
  • $170.1k - $258.3k

     ...real vehicles at scale. We pioneer new...  ...performance engineering so that every cycle...  ...‑performance GPU kernels and...  ...our on‑vehicle ML inference for ADAS and autonomous...  ...meeting strict latency, throughput,...  ...and iterate on CUDA-based kernels...  ...basedExperience with low latency or real... 
    Suggested
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    3 days ago
  • A media technology company in San Francisco is seeking a Founding Engineer specializing in ML Inference. This highly technical role requires expertise in the ML infrastructure stack and aims to optimize generative media performance. The ideal candidate will drive innovations... 
    Suggested
    Relocation package

    Reactor.am

    San Francisco, CA
    5 days ago
  •  ...San Francisco is seeking an ML Inference Engineer to maximize performance of generative...  ...media models and push ultra-low-latency, high-throughput inference....  ...in PyTorch, TensorRT, CUDA, and model optimization...  ...edge inference capabilities at scale. #J-18808-Ljbffr Reactor.am
    Suggested

    Reactor.am

    San Francisco, CA
    20 hours ago
  •  ...company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on... 
    Suggested

    Reflection AI

    San Francisco, CA
    4 days ago
  •  ...Machine Learning Engineer to join our Inference Infrastructure team...  ...high-throughput, low-latency distributed...  ...Architect, build, and scale low-latency...  ...models. Optimize GPU utilization, memory...  ...distributed systems or ML infrastructure....  ...experience with CUDA, Triton, or deep... 
    Suggested
    Worldwide
    Flexible hours

    AI Breaking Wire

    San Francisco, CA
    5 days ago
  • $180k - $270k

     ...deploying high-throughput, ultra-low-latency inference engines for large language models or...  ...a deep understanding of GPU architectures (NVIDIA Ampere...  ...between the core ML training team and the backend...  ...naturalness or ASR accuracy. Large-Scale Distributed Systems:... 
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    1 day ago
  • Oscar is hiring a Senior Machine Learning Inference Engineer for a full-time role in the San Francisco Bay Area. You will focus on improving...  ...ideal candidate has 3+ years of professional experience, deep GPU infrastructure knowledge, and strong Python and PyTorch skills... 
    Full time

    Oscar

    San Francisco, CA
    3 days ago
  •  ...deploy, and maintain large distributed ML training and inference clusters Develop efficient, scalable...  ...-to-end pipelines to manage petabyte-scale datasets and model training...  ...model scales Analyze, profile and debug low-level GPU operations to optimize performance Stay... 

    Kindredventures

    San Francisco, CA
    20 hours ago
  •  ...documentation platform scaling to thousands of...  ...We are hiring two ML Engineers / Researchers to help...  ...improving quality, cost, latency, and control. You...  ...Optimize models for low-latency, real-time inference at production scale...  ...and inference GPU-based model serving... 
    Full time

    Knowtex

    San Francisco, CA
    1 day ago
  • $160k - $230k

     ...seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing...  ...and effectively at scale. If you are passionate about...  ...Excellent understanding of low-level operating systems...  ...~ Preferred: Knowledge of CUDA/Triton programming. ~ Nice... 
    Full time

    Together Ai

    San Francisco, CA
    1 day ago
  • Sciforium is seeking a GPU Kernel Engineer to push performance on...  ...custom GPU kernels, from low-level development to integrating ops in ML frameworks used for large-scale training and inference. Ideal candidates have...  ...Python skills, and deep CUDA/ROCm expertise.... 

    Sciforium

    San Francisco, CA
    3 days ago
  •  ...ML Infrastructure Engineer, Model InferenceAs an ML Infrastructure...  ...Engineer, Model Inference at Abridge, you'...  ...-performance and low-latency.Collaborate with...  ...product teams to scale backend...  ...workflows and enhance GPU utilization for ML...  ...cluster management and CUDA... 
    Hourly pay
    Full time
    Flexible hours

    Abridge

    San Francisco, CA
    3 days ago
  • $128.7k - $261.3k

     ...real vehicles at scale. We pioneer new...  ...performance engineering so that every cycle...  ...fast, reliable inference across GPUs...  ..., systems, and GPU engineers who enjoy...  ...MLIR/ONNX and CUDA/TensorRT...  ...effortless for ML engineers across...  ...and on‑vehicle latency. Along the way,... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    5 days ago
  • $160k - $230k

     ...efficient and scalable inference for large language...  ...Frameworks and Optimization Engineer to design, develop,...  ...language models at scale. This role will focus on low-latency, high-throughput inference, GPU/accelerator...  ...performance serving.Apply CUDA graph optimizations,... 
    Full time

    Together AI

    San Francisco, CA
    2 days ago
  •  ...Learning Infrastructure Engineer to help architect the...  ..., this team is scaling rapidly to solve complex...  ...flexibility, and ultra-low latency serving.Establish comprehensive...  ...across extensive GPU clusters.Build...  ...distributed training or inference engines.Practical experience... 
    Full time
    Work at office
    Flexible hours

    Objective Paradigm

    San Francisco, CA
    3 days ago
  • David Joseph & Company seeks a talented ML infrastructure engineer to own the distributed training and inference backbone for a large foundation model. You...  ...clusters, build data pipelines for petabyte-scale datasets, and squeeze GPU performance across model scales. You will... 
    Relocation package

    David Joseph & Company

    San Francisco, CA
    3 days ago
  •  ...-on support from AMD engineers the team is scaling rapidly to build the...  ...About the role As an ML Engineer at...  ...accuracy, reliability, latency, and task completion....  ..., or JAX, or Ray and inference engines like TensorRT...  ...Familiarity with low-level performance considerations... 
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    1 day ago
  •  ...complex documents at scale. We have a...  ...fast-growing team of engineers in San Francisco powering...  ...models Low-latency OCR models for targeted...  ...pipelines Optimize inference, batching, and quantization on GPU Productionize models...  ...+ years in applied ML or research, or... 
    Full time
    Work at office
    Visa sponsorship
    Relocation package

    The Pulse

    San Francisco, CA
    1 day ago
  • An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate... 

    Jobleads-US

    San Francisco, CA
    4 days ago
  • $180k - $280k

     ...societal shift on the scale of the agricultural...  ...-2024, we've been engineering the foundation for...  ...'re looking for a GPU kernel engineer with deep, low-level CUDA expertise to make our training and inference faster and more efficient...  ...and minimum latency out of our hardware... 
    Work at office
    Visa sponsorship
    Shift work

    TypeSafe AI

    San Francisco, CA
    20 hours ago
  • $220k - $320k

    inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques... 

    inference.net

    San Francisco, CA
    4 days ago
  •  ...leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will...  ...computation efficiency. Ideal candidates have 1-5 years of CUDA development experience and a strong understanding of GPU architecture... 

    Baseten

    San Francisco, CA
    3 days ago
  • $100k - $120k

    Coda Robotics is scaling the compute...  ...As training and inference workloads grow,...  ...innovations to reduce latency, memory usage,...  ...and optimize low‑level compute...  ...and system engineers focused on performance...  ...AVX/ARM NEON), GPU (CUDA/ROCm), and...  ...into distributed ML frameworks (e.g... 

    Coda Robotics

    San Francisco, CA
    1 day ago
  • $200k - $400k

     ...Infrastructure Engineer San Francisco...  ...building large-scale foundation models...  ...distributed training and inference backbone for a...  ...of GPUs at a low level across...  ...large distributed ML training and...  ...debug low-level GPU operations to optimize...  ...and debugging (CUDA, JAX) Why Join... 
    Full time
    Visa sponsorship
    Relocation package

    David Joseph & Company

    San Francisco, CA
    14 days ago
  •  ...Zensors builds the engine that powers our...  ...Engineer in ML Runtime & Optimization...  ...training and inference of computer...  ...throughput and minimize latency. Building...  ...(e.g., PyTorch, CUDA, TensorRT, and NVIDIA...  ...of GPU hardware performance...  ...systems or cloud-scale inference serving... 

    Zensors

    San Francisco, CA
    24 days ago
  • Jaide Health is seeking an engineer for their Model Efficiency team...  ...focuses on building reliable ML systems while enhancing core...  ...performance techniques such as GPU/CUDA optimizations and collaborate...  ...and insights into the LLM inference ecosystem. A commitment to diversity... 
    Remote job

    Jaide Health

    San Francisco, CA
    2 days ago
  • $151.8k - $265.35k

     ...Machine Learning Engineer to build the...  ...them as services, scale those services to...  ...to meet target latency and throughput budgets...  ...production. Own GPU capacity and...  ...Run production ML operationally -...  ...production ML or inference services at scale...  ...; custom CUDA a plus. Strong,... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    6 days ago
  •  ...Staff Machine Learning Engineer, you’ll own AI-driven products...  ...something reliable at scale, drawing on real...  ...package, deploy, and operate low-latency, high-concurrency inference (Triton, vLLM, GPU-backed serving) that stays...  ...shipping and operating ML-driven functionality.... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Primer.ai

    San Francisco, CA
    1 day ago
  • $220k - $280k

     ...building the best inference infrastructure for...  ...with best-in-class latency and reliability....  ...looking for a Staff ML Engineer to drive the model...  .... You'll profile GPU utilization, design...  ...voice models at scale — design the serving...  ...optimization — CUDA kernels, memory hierarchies... 
    Full time

    Together Ai

    San Francisco, CA
    1 day ago
  • $250k

    Title : ML Inference Engineer Location : San Francisco, CA Salary : $250k base + equity An AI Unicorn...  ...and a strong understanding of GPU infrastructure, Python, and PyTorch. This...  ...Experience : Building AI applications at scale from the ground up Strong understanding... 
    Full time

    Oscar Technology

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Scale Low-Latency ML Inference Engineer (GPU/CUDA). Be the first to apply!