Member of Technical Staff, ML Inference Engineering
Sanas
Member of Technical Staff, ML Inference Engineering Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers and entrepreneurs with deep industry experience, Sanas has developed the world's first real-time speech AI platform capable of accent translation, noise cancellation, speech enhancement, cross-language communication, and more. Sanas makes conversations clearer, more inclusive, and more effective, removing barriers that prevent people from being understood, regardless of accent, background noise, or native language. Sanas is currently one of the fastest growing startups in Silicon Valley, growing from $16M to $50M ARR in 2025. The company's core business is profitable and is on track to end 2026 with >$120M ARR. Our team combines deep expertise in model innovation and systems engineering with a design-minded product engineering culture to build and ship cutting-edge AI models and experiences — entirely in-house. Sanas is a 130 person team, established in 2020. In this short span, we've successfully secured over $100 million in funding. Our innovation has been supported by the industry's leading investors, including Insight Partners, Google Ventures, Quadrille Capital, General Catalyst, Quiet Capital, and other influential investors. Our reputation is further solidified by collaborations with numerous Fortune 100 companies. With Sanas, you're not just adopting a product; you're investing in the future of communication. If you’re looking to have a significant role in roadmapping and driving technical directions, if you’re looking to deploy challenging and big ideas without much overhead or slowness, if you're looking to leave your mark on an ambitious, generational mission to change how the worlds thinks about speech + AI, then Sanas is a well-suited place for you. Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable. We're looking for a deeply hands-on, senior engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them. Performance Optimization Optimize system and GPU performance for high-throughput AI workloads across multi-node training and inference Analyze and improve latency, throughput, memory usage, and compute efficiency Profile system performance to detect and resolve GPU- and kernel-level bottlenecks Implement low-level optimizations using CUDA, Triton, and other performance tooling Improve support for mixed precision, quantization, and model graph optimization Build and maintain performance benchmarking and monitoring infrastructure Scale inference and training systems across multi-GPU, multi-node environments Inference Systems & Reliability Own and evolve our inference engine, enabling reliability and performance at scale Develop and optimize runtime inference services for large-scale AI applications Implement robust, fault-tolerant systems for data ingestion and processing Requirements Must-have: 5+ years of experience writing high-quality, high-performance code Familiarity with NVIDIA GPU architecture and CUDA Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling A research-leaning or systems background in LLM, Speech-to-Text, Text-to-Speech, or Speech-to-Speech inference, with work you can point to A record of shipping research or systems that other people build on, whether in a lab or in industry Nice-to-have: Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes Experience maintaining or contributing to open-source ML projects Experience managing machine learning workloads on Kubernetes clusters Experience with InfiniBand or RoCE networking Experience with bare-metal provisioning and lifecycle management Experience operating large-scale AI training or inference clusters Experience with hardware health monitoring and predictive failure detection Experience with distributed storage systems #J-18808-Ljbffr Sanas
- Member of Technical Staff — Kernel / Compiler / Communication RadixArk is seeking a deeply technical engineer who pushes the limits of performance for frontier... ...to scaling training and inference across thousands of GPUs... ...and runtime stacks for ML systems Improve communication...SuggestedFlexible hours
- Member of Technical Staff — Kernel / Compiler / Communication About the Role RadixArk... ...to scaling training and inference across thousands of GPUs,... ...deeply technical role for engineers who enjoy working close to... ...Contributions to kernel/compiler/ML systems open source...SuggestedFlexible hours
- ...Role RadixArk is seeking a Member of Technical Staff — Diffusion Model to advance... ...thinking with strong engineering execution—from designing novel... ...5+ years of experience in ML research or applied ML engineering... ...to scale training and inference Translate research ideas into...SuggestedFlexible hours
- RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on the core systems that... ...GPUs. This role sits at the intersection of systems engineering, ML infrastructure, and performance optimization. Your work...SuggestedWorldwideFlexible hours
$180k - $250k
Member of Technical Staff -- TPU Systems (JAX / XLA / PALLAS) About the Role RadixArk is looking for a TPU Systems Engineer to build high-performance inference and training systems using JAX, XLA, and Pallas. You... ...experience building production ML systems with JAX, XLA, or...SuggestedFull timeFlexible hours- ...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who... ...looking for an engineer to help with low precision RL training and inference. RESPONSIBILITIES: Design and optimize our inference stack for...
- Perplexity is seeking experienced ML engineers to design, build, and optimize the recommendation systems that power core experiences on... ...these systems learn and improve with usage. Help shape the technical direction of ranking, recommendations, and personalization at...
$209k - $313k
...other digital services.Snap Engineering teams build fun and technically sophisticated products... ...understanding of causal inference and modern approaches to... ...tests) and leveraging causal ML in production... ...approach and expect our team members to work in an office 4+ days...Full timeLive inWork at officeLocal area$250k - $350k
About the RoleWe are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment...Work at office3 days per week- ...The Role RadixArk is seeking a Member of Technical Staff — Training to build and scale the... ...role sits at the intersection of ML, systems, and performance engineering. Your work will directly impact... ...Experience working on training / inference correctness or other precision-...Flexible hours
- Member of Technical Staff — Accelerator Systems About the Role RadixArk is seeking... ...systems. Most performance engineering assumes a single vendor's... ...in systems, performance, or ML infrastructure engineering... ...Experience with distributed inference systems (SGLang, vLLM) or training...Flexible hours
$148.5k - $223.9k
Senior Member of Technical Staff - AI ResearchSkip to main content#Senior Member... ...future of Salesforce.*ware engineers, product managers and solution... ...skills.** *Has deep ML knowledge with meaningful implementation... ...training, evaluation, and inference pipelines** *Infrastructure...Work at office- About The Role RadixArk is hiring a Member of Technical Staff — CI Engineer to own the infrastructure that keeps... ...the fastest‑growing open‑source LLM inference engines. When CI is green and fast,... ...experience in CI contexts Familiarity with ML inference workloads (model loading,...Flexible hoursNight shift
$180k
Member of Technical Staff - Multimodal Understanding About xAI xAI’s mission is... ...motivated, and focused on engineering excellence. This organization... ...pre‑training, post‑training, inference, data processing, and... ...optimizing large‑scale distributed ML systems (training/inference...Temporary work- RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems... ...role sits at the intersection of ML, systems, and performance engineering. Your work will directly impact... ...infrastructure for AI training and inference and partner with frontier AI...Flexible hours
- Member of Technical Staff, LLM Post-Training, Applied Sanas is pioneering the future... ...innovation and systems engineering with a design-minded product... ...‑tuning, alignment, and inference optimization — into models... ...Proficiency with the open‑source ML ecosystem (Hugging Face,...
$229k - $343k
...services.Snap’s Generative ML Platform team builds... ...device and server-side inference. Our team creates... ...for a Machine Learning Engineer to join Snap Inc!What you... ...Qualifications:Bachelor's degree in technical field such as computer... ...and expect our team members to work in an office 4+...Full timeLive inWork at officeLocal areaWorldwide$190k - $250k
...time Location Type Hybrid Department AI We are looking for an AI Inference engineer to join our growing team. Our current stack is Python, Rust,... ...LLM inference optimizations Qualifications Experience with ML systems and deep learning frameworks (e.g. PyTorch, TensorFlow...Full time- Bright Vision Technologies is seeking a Machine Learning Infrastructure Engineer to design, build, and operate high-performance inference platforms for serving large ML models in production. This remote U.S. role focuses on the systems engineering side of AI deployment...Remote job
- ...The Role RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech) to make LLM inference and training dramatically... ...a high-performance inference engine that serves trillions of tokens... ...contributions to open-source AI/ML projects. About RadixArk RadixArk...Flexible hours
$120k - $200k
...curated datasets, or full-cycle data engineering, Abaka AI provides the foundation for... ...AI systems. About the Role As a Member of Technical Staff, Platform, you'll build full-stack product... ...-facing products at scale. - ML and Deep Learning knowledge and experience...Full timeFlexible hours$300k - $350k
Member of Technical Staff Level 1 - Engineering Bellevue | Hybrid NTT DATA AIVista, Inc., a wholly owned subsidiary of NTT DATA, is based in Silicon Valley... ...Large Language Models (LLMs) AI infrastructure and inference systems Data processing and knowledge pipelines Agent...Work experience placementLocal areaFlexible hours- ...-latency, high-throughput AI for multi-node GPU workloads. As a Senior Engineer, you will shape core infrastructure and architecture decisions, lead performance optimizations, and own the inference engine to scale research and production workloads. This role is in the...
$180k
...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who... ...generation models. Background in distributed systems, real-time inference serving, Kubernetes, observability tools, or large-scale data...Temporary workWorldwide$180k
...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who... ...SKILLS AND EXPERIENCE: Experience with real-time systems, inference serving, or multi-modal data processing at scale. Familiarity...Temporary workWorldwide$200k - $420k
Member of Technical Staff, Software Engineering At River, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are... ...stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs...Local areaVisa sponsorshipRelocation package$180k
Member of Technical Staff - Pre-Training About xAI xAI’s mission is to create AI systems that can accurately understand the universe and aid humanity... .... Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who...Temporary work- ...looking for We hire deeply technical staff working in world models, video... ...runs, or working on inference, we hire people who can take... ...to-shoulder across research, engineering, and product. Strong work here... ...multimodal learning, large-scale ML systems, or adjacent areas....
- ...design custom ASICs alongside evolving ML workloads, and enable a new era of... ...and Intel. What You’ll Do As a Founding Member of the Technical Staff on the RTL Design team at Architect, you... ...), refresh management, and ECC/RAS engines. Architect the memory hierarchy : including...
- ...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who... ...to a 15-30 minutes phone interview, during which a member of our team will ask technical questions. If you clear the phone interview, you will...Temporary workH1bRelocationWork visa
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff, ML Inference Engineering. Be the first to apply!
- operations support technician Palo Alto, CA
- product support technician Palo Alto, CA
- systems support technician Palo Alto, CA
- user support analyst Palo Alto, CA
- technical support specialist Palo Alto, CA
- help desk assistant Palo Alto, CA
- personal computer support technician Palo Alto, CA
- mri tech aide Palo Alto, CA
- desktop support analyst Palo Alto, CA
- tech assistant Palo Alto, CA

