Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff, ML Inference Engineering

Sanas

Member of Technical Staff, ML Inference Engineering Sanas is pioneering the future of human communication. Founded by a team of Stanford researchers and entrepreneurs with deep industry experience, Sanas has developed the world's first real-time speech AI platform capable of accent translation, noise cancellation, speech enhancement, cross-language communication, and more. Sanas makes conversations clearer, more inclusive, and more effective, removing barriers that prevent people from being understood, regardless of accent, background noise, or native language. Sanas is currently one of the fastest growing startups in Silicon Valley, growing from $16M to $50M ARR in 2025. The company's core business is profitable and is on track to end 2026 with >$120M ARR. Our team combines deep expertise in model innovation and systems engineering with a design-minded product engineering culture to build and ship cutting-edge AI models and experiences — entirely in-house. Sanas is a 130 person team, established in 2020. In this short span, we've successfully secured over $100 million in funding. Our innovation has been supported by the industry's leading investors, including Insight Partners, Google Ventures, Quadrille Capital, General Catalyst, Quiet Capital, and other influential investors. Our reputation is further solidified by collaborations with numerous Fortune 100 companies. With Sanas, you're not just adopting a product; you're investing in the future of communication. If you’re looking to have a significant role in roadmapping and driving technical directions, if you’re looking to deploy challenging and big ideas without much overhead or slowness, if you're looking to leave your mark on an ambitious, generational mission to change how the worlds thinks about speech + AI, then Sanas is a well-suited place for you. Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable. We're looking for a deeply hands-on, senior engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them. Performance Optimization Optimize system and GPU performance for high-throughput AI workloads across multi-node training and inference Analyze and improve latency, throughput, memory usage, and compute efficiency Profile system performance to detect and resolve GPU- and kernel-level bottlenecks Implement low-level optimizations using CUDA, Triton, and other performance tooling Improve support for mixed precision, quantization, and model graph optimization Build and maintain performance benchmarking and monitoring infrastructure Scale inference and training systems across multi-GPU, multi-node environments Inference Systems & Reliability Own and evolve our inference engine, enabling reliability and performance at scale Develop and optimize runtime inference services for large-scale AI applications Implement robust, fault-tolerant systems for data ingestion and processing Requirements Must-have: 5+ years of experience writing high-quality, high-performance code Familiarity with NVIDIA GPU architecture and CUDA Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling A research-leaning or systems background in LLM, Speech-to-Text, Text-to-Speech, or Speech-to-Speech inference, with work you can point to A record of shipping research or systems that other people build on, whether in a lab or in industry Nice-to-have: Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes Experience maintaining or contributing to open-source ML projects Experience managing machine learning workloads on Kubernetes clusters Experience with InfiniBand or RoCE networking Experience with bare-metal provisioning and lifecycle management Experience operating large-scale AI training or inference clusters Experience with hardware health monitoring and predictive failure detection Experience with distributed storage systems #J-18808-Ljbffr Sanas

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff, ML Inference Engineering in Palo Alto, CA vacancy
  • Member of Technical Staff — Kernel / Compiler / Communication RadixArk is seeking a deeply technical engineer who pushes the limits of performance for frontier...  ...to scaling training and inference across thousands of GPUs...  ...and runtime stacks for ML systems Improve communication... 
    Suggested
    Flexible hours

    RadixArk

    Palo Alto, CA
    5 days ago
  • Member of Technical Staff — Kernel / Compiler / Communication About the Role RadixArk...  ...to scaling training and inference across thousands of GPUs,...  ...deeply technical role for engineers who enjoy working close to...  ...Contributions to kernel/compiler/ML systems open source... 
    Suggested
    Flexible hours

    RadixArk

    Palo Alto, CA
    4 days ago
  •  ...Role RadixArk is seeking a Member of Technical Staff — Diffusion Model to advance...  ...thinking with strong engineering execution—from designing novel...  ...5+ years of experience in ML research or applied ML engineering...  ...to scale training and inference Translate research ideas into... 
    Suggested
    Flexible hours

    RadixArk

    Palo Alto, CA
    5 days ago
  • RadixArk is seeking a Member of Technical Staff — Inference to push the limits of large-scale AI inference. You will work on the core systems that...  ...GPUs. This role sits at the intersection of systems engineering, ML infrastructure, and performance optimization. Your work... 
    Suggested
    Worldwide
    Flexible hours

    Dormont Manufacturing Co

    Palo Alto, CA
    2 days ago
  • $180k - $250k

    Member of Technical Staff -- TPU Systems (JAX / XLA / PALLAS) About the Role RadixArk is looking for a TPU Systems Engineer to build high-performance inference and training systems using JAX, XLA, and Pallas. You...  ...experience building production ML systems with JAX, XLA, or... 
    Suggested
    Full time
    Flexible hours

    RadixArk

    Palo Alto, CA
    3 days ago
  •  ...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who...  ...looking for an engineer to help with low precision RL training and inference. RESPONSIBILITIES: Design and optimize our inference stack for... 

    Pantera Capital

    Palo Alto, CA
    5 days ago
  • Perplexity is seeking experienced ML engineers to design, build, and optimize the recommendation systems that power core experiences on...  ...these systems learn and improve with usage. Help shape the technical direction of ranking, recommendations, and personalization at... 

    Pantera Capital

    Palo Alto, CA
    2 days ago
  • $209k - $313k

     ...other digital services.Snap Engineering teams build fun and technically sophisticated products...  ...understanding of causal inference and modern approaches to...  ...tests) and leveraging causal ML in production...  ...approach and expect our team members to work in an office 4+ days... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    Palo Alto, CA
    3 days ago
  • $250k - $350k

    About the RoleWe are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment... 
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    1 day ago
  •  ...The Role RadixArk is seeking a Member of Technical Staff — Training to build and scale the...  ...role sits at the intersection of ML, systems, and performance engineering. Your work will directly impact...  ...Experience working on training / inference correctness or other precision-... 
    Flexible hours

    RadixArk

    Palo Alto, CA
    2 days ago
  • Member of Technical Staff — Accelerator Systems About the Role RadixArk is seeking...  ...systems. Most performance engineering assumes a single vendor's...  ...in systems, performance, or ML infrastructure engineering...  ...Experience with distributed inference systems (SGLang, vLLM) or training... 
    Flexible hours

    RadixArk

    Palo Alto, CA
    1 day ago
  • $148.5k - $223.9k

    Senior Member of Technical Staff - AI ResearchSkip to main content#Senior Member...  ...future of Salesforce.*ware engineers, product managers and solution...  ...skills.** *Has deep ML knowledge with meaningful implementation...  ...training, evaluation, and inference pipelines** *Infrastructure... 
    Work at office

    Salesforce, Inc.

    Palo Alto, CA
    4 days ago
  • About The Role RadixArk is hiring a Member of Technical Staff — CI Engineer to own the infrastructure that keeps...  ...the fastest‑growing open‑source LLM inference engines. When CI is green and fast,...  ...experience in CI contexts Familiarity with ML inference workloads (model loading,... 
    Flexible hours
    Night shift

    RadixArk

    Palo Alto, CA
    5 days ago
  • $180k

    Member of Technical Staff - Multimodal Understanding About xAI xAI’s mission is...  ...motivated, and focused on engineering excellence. This organization...  ...pre‑training, post‑training, inference, data processing, and...  ...optimizing large‑scale distributed ML systems (training/inference... 
    Temporary work

    xAI

    Palo Alto, CA
    4 days ago
  • RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems...  ...role sits at the intersection of ML, systems, and performance engineering. Your work will directly impact...  ...infrastructure for AI training and inference and partner with frontier AI... 
    Flexible hours

    RadixArk

    Palo Alto, CA
    5 days ago
  • Member of Technical Staff, LLM Post-Training, Applied Sanas is pioneering the future...  ...innovation and systems engineering with a design-minded product...  ...‑tuning, alignment, and inference optimization — into models...  ...Proficiency with the open‑source ML ecosystem (Hugging Face,... 

    Sanas

    Palo Alto, CA
    2 days ago
  • $229k - $343k

     ...services.Snap’s Generative ML Platform team builds...  ...device and server-side inference. Our team creates...  ...for a Machine Learning Engineer to join Snap Inc!What you...  ...Qualifications:Bachelor's degree in technical field such as computer...  ...and expect our team members to work in an office 4+... 
    Full time
    Live in
    Work at office
    Local area
    Worldwide

    Snap

    Palo Alto, CA
    3 days ago
  • $190k - $250k

     ...time Location Type Hybrid Department AI We are looking for an AI Inference engineer to join our growing team. Our current stack is Python, Rust,...  ...LLM inference optimizations Qualifications Experience with ML systems and deep learning frameworks (e.g. PyTorch, TensorFlow... 
    Full time

    Kindredventures

    Palo Alto, CA
    4 days ago
  • Bright Vision Technologies is seeking a Machine Learning Infrastructure Engineer to design, build, and operate high-performance inference platforms for serving large ML models in production. This remote U.S. role focuses on the systems engineering side of AI deployment... 
    Remote job

    Bright Vision Technologies

    Mountain View, CA
    3 days ago
  •  ...The Role RadixArk is seeking a Member of Technical Staff, Developer Technology (DevTech) to make LLM inference and training dramatically...  ...a high-performance inference engine that serves trillions of tokens...  ...contributions to open-source AI/ML projects. About RadixArk RadixArk... 
    Flexible hours

    RadixArk

    Palo Alto, CA
    4 days ago
  • $120k - $200k

     ...curated datasets, or full-cycle data engineering, Abaka AI provides the foundation for...  ...AI systems. About the Role As a Member of Technical Staff, Platform, you'll build full-stack product...  ...-facing products at scale. - ML and Deep Learning knowledge and experience... 
    Full time
    Flexible hours

    Embedding VC

    Mountain View, CA
    4 days ago
  • $300k - $350k

    Member of Technical Staff Level 1 - Engineering Bellevue | Hybrid NTT DATA AIVista, Inc., a wholly owned subsidiary of NTT DATA, is based in Silicon Valley...  ...Large Language Models (LLMs) AI infrastructure and inference systems Data processing and knowledge pipelines Agent... 
    Work experience placement
    Local area
    Flexible hours

    NTT DATA AIVista

    Palo Alto, CA
    2 days ago
  •  ...-latency, high-throughput AI for multi-node GPU workloads. As a Senior Engineer, you will shape core infrastructure and architecture decisions, lead performance optimizations, and own the inference engine to scale research and production workloads. This role is in the... 

    Sanas

    Palo Alto, CA
    2 days ago
  • $180k

     ...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who...  ...generation models. Background in distributed systems, real-time inference serving, Kubernetes, observability tools, or large-scale data... 
    Temporary work
    Worldwide

    SpaceXAI

    Palo Alto, CA
    3 days ago
  • $180k

     ...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who...  ...SKILLS AND EXPERIENCE: Experience with real-time systems, inference serving, or multi-modal data processing at scale. Familiarity... 
    Temporary work
    Worldwide

    SpaceXAI

    Palo Alto, CA
    3 days ago
  • $200k - $420k

    Member of Technical Staff, Software Engineering At River, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are...  ...stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs... 
    Local area
    Visa sponsorship
    Relocation package

    River AI Inc.

    Palo Alto, CA
    3 days ago
  • $180k

    Member of Technical Staff - Pre-Training About xAI xAI’s mission is to create AI systems that can accurately understand the universe and aid humanity...  .... Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who... 
    Temporary work

    Xai

    Palo Alto, CA
    4 days ago
  •  ...looking for We hire deeply technical staff working in world models, video...  ...runs, or working on inference, we hire people who can take...  ...to-shoulder across research, engineering, and product. Strong work here...  ...multimodal learning, large-scale ML systems, or adjacent areas.... 

    Doist

    Palo Alto, CA
    2 days ago
  •  ...design custom ASICs alongside evolving ML workloads, and enable a new era of...  ...and Intel. What You’ll Do As a Founding Member of the Technical Staff on the RTL Design team at Architect, you...  ...), refresh management, and ECC/RAS engines. Architect the memory hierarchy : including... 

    Kindredventures

    Palo Alto, CA
    2 days ago
  •  ...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who...  ...to a 15-30 minutes phone interview, during which a member of our team will ask technical questions. If you clear the phone interview, you will... 
    Temporary work
    H1b
    Relocation
    Work visa

    xAI

    Palo Alto, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff, ML Inference Engineering. Be the first to apply!