On-Device LLM Inference Engineer — Real-Time, Rust & GPU
Mirai Labs
Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack for real-time local intelligence. You will primarily work on uzu, our inference engine, and focus on supporting new modalities and a variety of features. The ideal candidates should have a deep understanding of how computers and modern language models work, and experience in writing high-performance GPU kernels or Rust systems programming. We welcome applications from talented students and early-career engineers. #J-18808-Ljbffr Mirai Labs
$160k - $230k
...efficient and scalable inference for large language models... ...and Optimization Engineer to design, develop, and... ...-throughput inference, GPU/accelerator optimizations... ...to shape the future of LLM inference infrastructure... ...salary range for this full-time position is: $160,000 -...SuggestedFull time$170k - $245k
...progress of AI applications out into the real world.With Anyscale, we’re building the... ...to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and... ...consistent. As the market data changes over time, the target salary for this role may be...SuggestedWork at office- ...worldwide.We’re a team of engineers, clinicians, and innovators... ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you... ...and scalable medical device products and infrastructures... ...from foundational models to real-time onboard inference—while serving as a core contributor...SuggestedLocal areaWorldwideFlexible hours
- Redwood Materials, Inc. is hiring an Embedded Software Engineer to develop real‑time firmware for power electronics in energy storage systems. You... ...with 2+ years shipping firmware, expertise in C/C++ (and Rust), and experience with Cortex‑M/R, CAN/Ethernet, and safety...Suggested
$220k
Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience...Suggested$300k
GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its limits — not in theory, but in production systems handling real-time speech... ...with attention kernels, decoding paths, or LLM-style runtimes Comfort profiling with low-level GPU tooling...RelocationVisa sponsorshipFree visa- ...research lab is seeking a Senior Software Engineer focused on building next-generation... ...expertise in Python and experience with modern LLM frameworks. This position involves... ...cross-functional teams to create scalable real-time interfaces. Join a well-funded startup at...
- ...shipping excellence. We seek engineers with strong intrinsic... ...to help scale AI inference. You’ll leverage your knowledge... ...systems to optimize GPU performance at the bleeding edge of AI. Full-Time On-site at either our... ...(virtual, 30 minutes) LLM-assisted coding...Full timeWork at office
- Lykos Controls in the San Francisco Bay Area is seeking a software engineer to design and develop MES applications, enabling real-time production visibility and quality management. You will integrate MES with ERP and PLCs, implement real-time data collection, and create...
- Kindredventures is hiring software engineers to own the path from trained model to customer value, building production systems that deliver predictions in real time across cloud, VPC, on‑prem, and restricted networks. You will create the surfaces customers touch, from APIs...
$173.5k - $331.05k
...for a senior, hands-on engineer to own and evolve the cross-platform GPU rendering platform at the... ...platform that underpins real-time video processing across... ...performance, stability, and device compatibility across... ...driver / hardware vendors ML inference integration (e.g.,...Full timeTemporary workLocal areaWorldwide- Aurelius Systems is seeking a Perception Engineer to join our software team and advance end-to-end perception and sensor fusion for our... ...with hardware engineers to integrate sensors and optimize real-time detection and tracking in challenging environments. #J-18808-...
- ...About the Team OpenAI’s Inference team powers the... ...small, fast-moving team of engineers focused on delivering a... ...infrastructure for serving real-time audio, image, and other... ...improvements including GPU utilization, tensor... ...tooling like vLLM, TensorRT-LLM, or custom model...Full time
$300k
...committed researchers, engineers, policy experts,... ...role Our Inference team is responsible... ...management systems LLM inference... ..., GCP) Python or Rust Representative projects... ...performance based on real-world production workloads... ...package for full-time employees includes...Full timeWork at officeWorldwideVisa sponsorshipFlexible hours- ...Baseten powers mission-critical inference for the world's most dynamic... ...us and help build the platform engineers turn to to ship AI products.... ...hardware. We believe that as LLM and multi-modal workloads scale... ...foundational engineers to lead our GPU Networking efforts, making...Full timeFlexible hours
$300k
...committed researchers, engineers, policy experts,... ...The Cloud Inference team scales and optimizes... ...remediation based on real-world production workloads... ...familiarity with LLM inference... ...Proficiency in Python or Rust The annual... ...at least 25% of the time. However, some roles...Full timeWork at officeVisa sponsorshipFlexible hours- ...powers mission-critical inference for the world's most... ...help build the platform engineers turn to to ship AI... ...throughput, and improve GPU efficiency via profiling... ...Build large-scale, real-time infrastructure for multi... ...profiling across host-device boundaries (e.g....Full timeFlexible hours
- AI Systems Engineer - Codex Core Agents Location... ...Type Full time Department... ...agents operating in real development environments... ...model behavior, inference/runtime issues,... ...across layers: Rust systems code,... ...experience with LLM applications,... ...inference optimization, GPU systems,...Full timeWork at officeLocal areaRelocation packageFlexible hours
- Atomic Machines is seeking a Staff Mechanical Engineer to help shape the Matter Compiler platform from conception through deployment.... ...engineers to push the capabilities of precision micromachining and real-time manufacturing, with hands-on prototyping and rigorous...
$160k - $190k
...Manufacturing Co in San Francisco is looking for a full-time Senior Robotics Software Engineer to develop advanced robotic control systems. The successful... ...in robotic control development. A strong grasp of Rust or C++ and hands-on experience with industrial robotic arms...Full time$298k - $368k
...research to address real-world problems and... ...sensors, enabling engineers like you to (1) develop... ...: Design VLM/LLM model architecture... ...performance for on-device use cases (memory,... ...-latency on-device inference techniques and a deep... ...for this full-time position across US...Full timeRemote work- Socket.dev is seeking an experienced Software Engineer to lead the development of software and services for safe, reliable operation... ...The role emphasizes ownership of major development projects, real-time data handling, and integration with customer systems, with opportunities...
- ...where AI agents hang out with real people in group chats, DMs, voice... ...Staff is the title we use for engineers who own hard problems end to... ...messages Serving low-latency LLM responses for group chats with... ...training, fine-tuning, evaluation, inference, or RAG at scale High-...
- An innovative robotics company in San Francisco is seeking an experienced engineer to join their team. The role involves designing and implementing real-time, performance-critical components for the robotics software platform. Candidates should have a Bachelor's degree...
- ...robotics firm in San Francisco is seeking a Robotics Software Engineer to develop high-level application software for autonomous systems... ...strong programming skills in C++ and Python, experience with real-time systems, and familiarity with ROS2. This full-time position...Full time
- Point One Navigation is seeking a Staff Computer Vision Engineer to own the complete lifecycle of spatial AI and visual navigation features. You will drive architecture decisions, optimize for edge devices, and ensure robust performance across devices and applications....
$180k - $280k
...quietly rethinking the LLM stack from first principles... ...model designed for real-world reliability, decision... ...mid-2024, we've been engineering the foundation for what... ...We're looking for a GPU kernel engineer with deep... ...make our training and inference faster and more efficient...Work at officeVisa sponsorshipShift work- ...Computer Lab in San Francisco is seeking a hands-on Perception Engineer to build the computer vision and ML system at the core of our... ...perception models that detect people, gestures, and scenes, enabling real-time understanding on edge hardware. You will collaborate with the...
- MENFEM is seeking its first engineering hire in San Francisco to build core infrastructure for real-time content generation. This role offers a unique opportunity to work directly with founders and define technical architecture from day one. Ideal candidates will have...
- ...Building a new kind of platform for real-time generative media, enabling... ...unicorn founders and senior engineers with deep expertise in 3D,... ...for a Founding Engineer, ML Inference with deep expertise in high-performance... ...Working knowledge of GPU hardware (NVIDIA) and the...RelocationVisa sponsorshipRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to On-Device LLM Inference Engineer — Real-Time, Rust & GPU. Be the first to apply!

