Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

On-Device LLM Inference Engineer — Real-Time, Rust & GPU

Mirai Labs

Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack for real-time local intelligence. You will primarily work on uzu, our inference engine, and focus on supporting new modalities and a variety of features. The ideal candidates should have a deep understanding of how computers and modern language models work, and experience in writing high-performance GPU kernels or Rust systems programming. We welcome applications from talented students and early-career engineers. #J-18808-Ljbffr Mirai Labs

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the On-Device LLM Inference Engineer — Real-Time, Rust & GPU in San Francisco, CA vacancy
  • $160k - $230k

     ...efficient and scalable inference for large language models...  ...and Optimization Engineer to design, develop, and...  ...-throughput inference, GPU/accelerator optimizations...  ...to shape the future of LLM inference infrastructure...  ...salary range for this full-time position is: $160,000 -... 
    Suggested
    Full time

    Together AI

    San Francisco, CA
    1 day ago
  • $170k - $245k

     ...progress of AI applications out into the real world.With Anyscale, we’re building the...  ...to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and...  ...consistent. As the market data changes over time, the target salary for this role may be... 
    Suggested
    Work at office

    Anyscale

    San Francisco, CA
    4 days ago
  •  ...worldwide.We’re a team of engineers, clinicians, and innovators...  ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you...  ...and scalable medical device products and infrastructures...  ...from foundational models to real-time onboard inference—while serving as a core contributor... 
    Suggested
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    1 day ago
  • Redwood Materials, Inc. is hiring an Embedded Software Engineer to develop real‑time firmware for power electronics in energy storage systems. You...  ...with 2+ years shipping firmware, expertise in C/C++ (and Rust), and experience with Cortex‑M/R, CAN/Ethernet, and safety... 
    Suggested

    Redwood Materials, Inc.

    San Francisco, CA
    5 days ago
  • Redwood Materials is seeking an Embedded Software Engineer to design real‑time firmware for power conversion units. You will work at the intersection...  ...candidates have 2+ years in firmware, strong C/C++ (and Rust), and hands‑on experience with Cortex‑M/R, peripherals, and... 
    Suggested

    Redwood Materials

    San Francisco, CA
    5 days ago
  • Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will leverage your...  ...record in high-performance systems. This role is full-time and on-site, offering equity and a fast-paced startup... 
    Full time

    Vast.ai

    San Francisco, CA
    1 day ago
  • $220k

     ...Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience... 

    Perplexity

    San Francisco, CA
    2 days ago
  • Parafin, Inc. is seeking a Software Engineer to lead the evolution of its ML Platform within...  ...experimentation, training, evaluation, inference, and retraining powering underwriting...  ...platform end-to-end, supporting batch and real-time underwriting infrastructure in... 

    Parafin, Inc.

    San Francisco, CA
    1 day ago
  • Lumicity is seeking an Embedded Software Engineer to develop the foundational software...  ...build high‑performance, safety‑critical, real‑time embedded systems that enable reliable, deterministic...  ...creating RTOS-based control software, device drivers, BSPs for ARM64 platforms, and... 

    Lumicity

    San Francisco, CA
    5 days ago
  •  ...platforms. You will design and integrate control systems, working on real hardware alongside a small, dedicated team. Applicants should...  ...a strong background in robotics with hands-on experience in real-time control system design. The position offers competitive salary, meaningful... 
    Relocation package

    Industrial Next (YC W22)

    San Francisco, CA
    2 days ago
  •  ...research lab is seeking a Senior Software Engineer focused on building next-generation...  ...expertise in Python and experience with modern LLM frameworks. This position involves...  ...cross-functional teams to create scalable real-time interfaces. Join a well-funded startup at... 

    Jack & Jill

    San Francisco, CA
    1 day ago
  • $300k

    GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its limits — not in theory, but in production systems handling real-time speech...  ...with attention kernels, decoding paths, or LLM-style runtimes Comfort profiling with low-level GPU tooling... 
    Relocation
    Visa sponsorship
    Free visa

    Techire Ai

    San Francisco, CA
    5 days ago
  •  ...shipping excellence. We seek engineers with strong intrinsic...  ...to help scale AI inference. You’ll leverage your knowledge...  ...systems to optimize GPU performance at the bleeding edge of AI. Full-Time On-site at either our...  ...(virtual, 30 minutes) LLM-assisted coding... 
    Full time
    Work at office

    Vast.ai Inc.

    San Francisco, CA
    6 days ago
  • Exploration Technology Group is seeking a hands-on Vision Systems Engineer to own the detection, tracking, and target discrimination software for a space-based IR sensing program. You will develop real-time pipelines on embedded hardware, integrating with simulations and... 

    Exploration Technology Group

    San Francisco, CA
    2 days ago
  • Lykos Controls in the San Francisco Bay Area is seeking a software engineer to design and develop MES applications, enabling real-time production visibility and quality management. You will integrate MES with ERP and PLCs, implement real-time data collection, and create... 

    Lykos Controls

    San Francisco, CA
    2 days ago
  • $137.5k - $227.5k

     ...company in San Francisco is seeking an Embedded Software Engineer to drive firmware development for power electronics. You will...  ...in firmware engineering, along with expert proficiency in Rust or C. This full-time role offers competitive compensation between $137,500 and... 
    Full time

    Redwood Materials

    San Francisco, CA
    5 days ago
  • $173.5k - $331.05k

     ...for a senior, hands-on engineer to own and evolve the cross-platform GPU rendering platform at the...  ...platform that underpins real-time video processing across...  ...performance, stability, and device compatibility across...  ...driver / hardware vendors ML inference integration (e.g.,... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    2 days ago
  • Kindredventures is hiring software engineers to own the path from trained model to customer value, building production systems that deliver predictions in real time across cloud, VPC, on‑prem, and restricted networks. You will create the surfaces customers touch, from APIs... 

    Kindredventures

    San Francisco, CA
    2 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 

    Vast.ai Inc.

    San Francisco, CA
    1 day ago
  • Aurelius Systems is seeking a Perception Engineer to join our software team and advance end-to-end perception and sensor fusion for our...  ...with hardware engineers to integrate sensors and optimize real-time detection and tracking in challenging environments. #J-18808-... 

    Aurelius Systems

    San Francisco, CA
    1 day ago
  • SR2 | Socially Responsible Recruitment | Certified B Corporation™ seeks a Senior/Staff Computer Vision Engineer to advance autonomous robotic platforms with real-time perception capabilities. You will influence sensor fusion, SLAM, and ML-based CV, collaborating across... 

    SR2 | Socially Responsible Recruitment | Certified B Corpora...

    San Francisco, CA
    5 days ago
  •  ...technology company in San Francisco is seeking a Senior Software Engineer for Backend (Systems / Infrastructure). You will architect and...  ...scalability as demand grows. This role involves optimizing APIs, managing GPU workloads, and collaborating with cross-functional teams. Ideal... 

    Vizcom

    San Francisco, CA
    1 day ago
  • $300k

     ...committed researchers, engineers, policy experts,...  ...role Our Inference team is responsible...  ...management systems LLM inference...  ..., GCP) Python or Rust Representative projects...  ...performance based on real-world production workloads...  ...package for full-time employees includes... 
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    17 hours ago
  •  ...About the Team OpenAI’s Inference team powers the...  ...small, fast-moving team of engineers focused on delivering a...  ...infrastructure for serving real-time audio, image, and other...  ...improvements including GPU utilization, tensor...  ...tooling like vLLM, TensorRT-LLM, or custom model... 
    Full time

    OpenAI

    San Francisco, CA
    17 hours ago
  • NextGenEnergyJobs in San Francisco is seeking an Embedded Software Engineer Intern for Fall 2026 to help develop bare-metal firmware for...  ...on Cortex-R and Cortex-M microcontrollers, and participate in real-time control and BMS tasks. You will gain hands-on experience... 
    Internship

    NextGenEnergyJobs

    San Francisco, CA
    5 days ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic...  ...us and help build the platform engineers turn to to ship AI products....  ...hardware. We believe that as LLM and multi-modal workloads scale...  ...foundational engineers to lead our GPU Networking efforts, making... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    17 hours ago
  • $300k

     ...committed researchers, engineers, policy experts,...  ...The Cloud Inference team scales and optimizes...  ...remediation based on real-world production workloads...  ...familiarity with LLM inference...  ...Proficiency in Python or Rust   The annual...  ...at least 25% of the time. However, some roles... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    17 hours ago
  • CHAOS Industries is seeking exceptional backend and cloud engineers to design and operate the secure cloud platform behind advanced radar...  ..., and systems engineers to ingest, process, and deliver real-time sensor data with low latency and high security. This high-ownership... 

    CHAOS Industries

    San Francisco, CA
    5 days ago
  •  ...powers mission-critical inference for the world's most...  ...help build the platform engineers turn to to ship AI...  ...throughput, and improve GPU efficiency via profiling...  ...Build large-scale, real-time infrastructure for multi...  ...profiling across host-device boundaries (e.g.... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    17 hours ago
  •  ...out into multiple AI inference requests running in real time. Behind that sits a large GPU fleet spread across several...  ...Today, our inference engineers and researchers build...  ...-level code in Go, Rust or C++. You’ve supported...  ..., SGLang, or TensorRT-LLM. Slurm or other HPC... 
    Shift work

    Neura Market

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to On-Device LLM Inference Engineer — Real-Time, Rust & GPU. Be the first to apply!