Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

On-Device LLM Inference Engineer — Real-Time, Rust & GPU

Mirai Labs

Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack for real-time local intelligence. You will primarily work on uzu, our inference engine, and focus on supporting new modalities and a variety of features. The ideal candidates should have a deep understanding of how computers and modern language models work, and experience in writing high-performance GPU kernels or Rust systems programming. We welcome applications from talented students and early-career engineers. #J-18808-Ljbffr Mirai Labs

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the On-Device LLM Inference Engineer — Real-Time, Rust & GPU in San Francisco, CA vacancy
  • $160k - $230k

     ...efficient and scalable inference for large language models...  ...and Optimization Engineer to design, develop, and...  ...-throughput inference, GPU/accelerator optimizations...  ...to shape the future of LLM inference infrastructure...  ...salary range for this full-time position is: $160,000 -... 
    Suggested
    Full time

    Together AI

    San Francisco, CA
    4 days ago
  • $170k - $245k

     ...progress of AI applications out into the real world.With Anyscale, we’re building the...  ...to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and...  ...consistent. As the market data changes over time, the target salary for this role may be... 
    Suggested
    Work at office

    Anyscale

    San Francisco, CA
    2 days ago
  •  ...worldwide.We’re a team of engineers, clinicians, and innovators...  ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you...  ...and scalable medical device products and infrastructures...  ...from foundational models to real-time onboard inference—while serving as a core contributor... 
    Suggested
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    4 days ago
  • Redwood Materials, Inc. is hiring an Embedded Software Engineer to develop real‑time firmware for power electronics in energy storage systems. You...  ...with 2+ years shipping firmware, expertise in C/C++ (and Rust), and experience with Cortex‑M/R, CAN/Ethernet, and safety... 
    Suggested

    Redwood Materials, Inc.

    San Francisco, CA
    3 days ago
  • $220k

    Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience... 
    Suggested

    Perplexity

    San Francisco, CA
    5 days ago
  • $300k

     ...GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its limits — not in theory, but in production systems handling real-time speech...  ...with attention kernels, decoding paths, or LLM-style runtimes Comfort profiling with low-level GPU tooling... 
    Relocation
    Visa sponsorship
    Free visa

    techire ai

    San Francisco, CA
    1 day ago
  •  ...AI Systems Engineer - Codex Core Agents...  ...Employment Type Full time Department...  ...agents operating in real development...  ...model behavior, inference/runtime issues,...  ...across layers: Rust systems code, Python...  ...experience with LLM applications, coding...  ...optimization, GPU systems, benchmarking... 
    Full time
    Work at office
    Local area
    Relocation package
    Flexible hours

    Slope

    San Francisco, CA
    4 days ago
  •  ...seeking an expert in high‑performance LLM serving systems and inference optimization. In this role, you will push...  ...and optimizing major inference engines such as SGLang, vLLM, or TensorRT. Deep knowledge of state‑of‑the‑art GPU architectures, and effectively exploit... 

    NEAR.AI

    San Francisco, CA
    4 days ago
  •  ...shipping excellence. We seek engineers with strong intrinsic...  ...to help scale AI inference. You’ll leverage your knowledge...  ...systems to optimize GPU performance at the bleeding edge of AI. Full-Time On-site at either our...  ...(virtual, 30 minutes) LLM-assisted coding... 
    Full time
    Work at office

    Vast.ai

    San Francisco, CA
    3 days ago
  • $173.5k - $331.05k

     ...for a senior, hands-on engineer to own and evolve the cross-platform GPU rendering platform at the...  ...platform that underpins real-time video processing across...  ...performance, stability, and device compatibility across...  ...driver / hardware vendors ML inference integration (e.g.,... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    5 hours ago
  • Kindredventures is hiring software engineers to own the path from trained model to customer value, building production systems that deliver predictions in real time across cloud, VPC, on‑prem, and restricted networks. You will create the surfaces customers touch, from APIs... 

    Kindredventures

    San Francisco, CA
    5 days ago
  •  ...About the Team OpenAI’s Inference team powers the...  ...small, fast-moving team of engineers focused on delivering a...  ...infrastructure for serving real-time audio, image, and other...  ...improvements including GPU utilization, tensor...  ...tooling like vLLM, TensorRT-LLM, or custom model... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $300k

     ...committed researchers, engineers, policy experts,...  ...role Our Inference team is responsible...  ...management systems LLM inference...  ..., GCP) Python or Rust Representative projects...  ...performance based on real-world production workloads...  ...package for full-time employees includes... 
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic...  ...us and help build the platform engineers turn to to ship AI products....  ...hardware. We believe that as LLM and multi-modal workloads scale...  ...foundational engineers to lead our GPU Networking efforts, making... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $300k

     ...committed researchers, engineers, policy experts,...  ...The Cloud Inference team scales and optimizes...  ...remediation based on real-world production workloads...  ...familiarity with LLM inference...  ...Proficiency in Python or Rust   The annual...  ...at least 25% of the time. However, some roles... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  •  ...powers mission-critical inference for the world's most...  ...help build the platform engineers turn to to ship AI...  ...throughput, and improve GPU efficiency via profiling...  ...Build large-scale, real-time infrastructure for multi...  ...profiling across host-device boundaries (e.g.... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...Fluidstack Production Engineering TeamExamples of key exciting...  ...of GPUs legible in real time: build the observability...  ...down to individual device and link. At this scale...  ...so every new site and GPU generation lands cleanly...  ...fluent with AI tooling. LLM APIs, MCP servers, and... 
    Local area

    Fluidstack

    San Francisco, CA
    2 days ago
  • $160k - $190k

     ...Manufacturing Co in San Francisco is looking for a full-time Senior Robotics Software Engineer to develop advanced robotic control systems. The successful...  ...in robotic control development. A strong grasp of Rust or C++ and hands-on experience with industrial robotic arms... 
    Full time

    Dormont Manufacturing Company

    San Francisco, CA
    1 day ago
  • $298k - $368k

     ...research to address real-world problems and...  ...sensors, enabling engineers like you to (1) develop...  ...: Design VLM/LLM model architecture...  ...performance for on-device use cases (memory,...  ...-latency on-device inference techniques and a deep...  ...for this full-time position across US... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • Socket.dev is seeking an experienced Software Engineer to lead the development of software and services for safe, reliable operation...  ...The role emphasizes ownership of major development projects, real-time data handling, and integration with customer systems, with opportunities... 

    Socket.dev

    San Francisco, CA
    5 days ago
  • $190k - $265k

     ...business. Founded by engineers — and customer-obsessed...  ...research, product, and real-world enterprise use...  ...language models across real-time and batch inference, powering model...  ...impact you will have:Build LLM infrastructure...  ...ML infrastructure, or GPU orchestrationFamiliarity... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  •  ...Building a new kind of platform for real-time generative media, enabling...  ...unicorn founders and senior engineers with deep expertise in 3D,...  ...for a Founding Engineer, ML Inference with deep expertise in high-performance...  ...Working knowledge of GPU hardware (NVIDIA) and the... 
    Relocation
    Visa sponsorship
    Relocation package

    Reactor.am

    San Francisco, CA
    2 days ago
  • $200k - $300k

     ...technology — from AI research to systems engineering to product design — creating a tight...  ...Role Overview We’re looking for a Real-time Graphics Engineer who can develop efficient...  ...graphics API stack, efficiently utilizing GPU hardware (both desktop and mobile), code... 
    Full time

    World Labs

    San Francisco, CA
    1 day ago
  • Aurelius Systems, Inc is seeking a Perception Engineer to join our software team in San Francisco. You will design, train, and deploy...  ...vision and sensor-based models, manage data pipelines, and enable real-time detection and tracking for our laser-based defense system. The... 

    Aurelius Systems, Inc

    San Francisco, CA
    4 days ago
  • $220k - $320k

    inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques... 

    inference.net

    San Francisco, CA
    1 day ago
  •  ...Machines, with offices in Emeryville and Santa Clara, CA, is building the Matter Compiler™ platform for real-time micro-manufacturing. We seek a Mechanical Engineer to shape the foundation of this multi-material, multi-process system, from design concepts through deployment... 

    Atomic Machines

    Emeryville, CA
    19 hours ago
  • $190k - $240k

     ...Computer Vision Full-time Hybrid Simbe is...  ...You will work on real robots operating in...  ...robot software, edge inference, and field...  ...CUDA, and embedded GPU workflows. Advance...  ...and mentorship for engineers working at the intersection...  ...record of adoption of LLM assisted and... 
    Full time

    SOSV SF & NY (Fka IndieBio)

    San Francisco, CA
    2 days ago
  • $165k - $310k

    Senior Research Engineer, LLM Training & Post-Training New York...  ..., and production inference, with security, observability...  ...AI’s platform and real‑world customer workloads...  ...training across multi‑GPU environments by improving...  ...priorities evolve over time. Master’s degree, PhD,... 
    For contractors
    For subcontractor
    Work at office
    Remote work
    Work from home
    Flexible hours
    2 days per week

    Lightning AI

    San Francisco, CA
    2 days ago
  • $300k

     ...model training, or inference.  Our client operates...  ...high-performance GPU clusters powering...  ...into low-latency, real-time inference and custom...  ...operate inference engines such as vLLM, SGLang, and TensorRT-LLM across multiple model...  ...in Python, Go, Rust, or a comparable language... 
    Permanent employment
    Worldwide
    San Francisco, CA
    more than 2 months ago
  • $170k - $250k

     ...This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The role...  ...Have: Top 0.1% C++ ability, elite Rust considered for the right candidate...  ..., or systems debugging under real production conditions No visa sponsorship... 
    Full time
    Visa sponsorship
    Flexible hours
    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to On-Device LLM Inference Engineer — Real-Time, Rust & GPU. Be the first to apply!