Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff / Principal Machine Learning Engineer, Inference - USA

Jobleads-US

About Inworld

Inworld is a research lab and inference provider focused on realtime AI for consumer-facing applications. We build first-party speech models, serve LLMs, and run the inference behind modular APIs designed for high-volume, realtime workloads.

Hundreds of millions of users interact with Inworld powered apps every day and we serve over 10 trillion LLM tokens per month. Our models and infrastructure support consumer applications across companions, healthcare, fitness, education, media, and more. Our work spans model research, realtime inference, large-scale serving infrastructure, and the APIs developers use to bring these capabilities into production.

We’ve raised more than $125M from Lightspeed Venture Partners, Section 32, Kleiner Perkins, Microsoft’s M12 venture fund, Founders Fund, Meta, Stanford, and others. Our technology has powered experiences from companies including NVIDIA, Microsoft Xbox, Niantic, Logitech Streamlabs, Wishroll, Little Umbrella, and Bible Chat. Inworld has also been recognized by CB Insights as one of the 100 most promising AI companies globally and named one of LinkedIn’s Top 10 Startups in the USA.

Who We're Looking For

A year ago, reliably working agentic systems and sub-second multimodal inference at scale barely existed. Nobody has a decade of experience here. So we're not screening for a resume template — we're looking for strong people from varied backgrounds who learn fast, thrive in ambiguity, and can show us what they've built, broken, and understood.

Experience We Find Useful

You don't need all of this. But you need enough to make a case.

  • Inference Optimization. Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM.

  • Model Acceleration . Hands-on experience with quantization, distillation, caching strategies , continuous batching, paged attention, and speculative decoding.

  • High-Performance Systems. Proficiency in C++, CUDA, Rust, or highly optimized Python. You know how to profile code and squeeze every ounce of performance out of NVIDIA GPUs.

  • Distributed Systems & Scaling. Experience with Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference, and reliably handling thousands of concurrent connections.

  • Public work. Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups.

  • Full-cycle ownership. You can take a model from the research team, containerize it, optimize its serving, and ensure it runs reliably in production.

  • Background. PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems.

Who Thrives Here

  • You don’t need a roadmap to start walking; you’re comfortable picking a direction and building the map as you go.

  • You believe engineering isn't finished until it’s shipped and stable. You have a bias for impact over purely theoretical optimizations.

  • You don't just ship code; you obsess over the why. You’re the first to question an architecture if you think there’s a better way to solve the core latency or throughput problem.

  • You aren't satisfied with "the PM said so." You thrive on deep context and want to understand the fundamental logic behind every decision we make.

What Working Here Is Like

We hand you unclear problems and expect you to make them clear. We value engineers who say "I don't know yet" and then design the benchmark or prototype that finds out. We treat performance, latency, and reliability as first-class product features, not a box to check before launch. Impact comes before everything else, though we support sharing work and open-source contributions that move the field forward. Your work should be visible. Flat structure, fast iterations, minimal process theater.

We believe in the power of in-person collaboration to solve the hardest problems and foster a strong team culture. We offer relocation assistance and look forward to you joining us in our Mountain View office.

The base salary range for this full-time position is $270,000 - $500,000+ bonus + equity + benefits.

Inworld Jobs Privacy

#J-18808-Ljbffr Jobleads-US
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff / Principal Machine Learning Engineer, Inference - USA in Mountain View, CA vacancy
  • $280k - $350k

     ...its class. Watch the launch video Staff / Principal Software Engineer - USA Mountain View, California, USA...  ...categories like health, fitness, learning, therapy, companions, customer experience...  ...-art models, optimizing realtime inference, and creating best-in-class APIs... 
    Suggested
    Full time
    Work at office
    Relocation

    Jobleads-US

    Mountain View, CA
    3 days ago
  • $174.08k - $229.04k

     ...users around the world and 300 billion ideas saved, Pinterest Machine Learning engineers help build personalized experiences to help Pinners create...  ...our various ML teams. Please only apply once within the USA or Canada as multiple applications may delay our recruitment... 
    Suggested
    Work at office
    Local area

    Pinterest

    Palo Alto, CA
    3 days ago
  • $130k - $260k

     ...: WalmartBusiness Segment: Home OfficeRole summary: As a Staff Machine Learning Engineer at Walmart, you will lead the design and deployment of scalable...  ...intelligence, machine learning, measurement, and causal inference to redefine retail experiences, optimize operations, and... 
    Suggested
    Full time
    Temporary work
    Part time
    Work experience placement

    Walmart

    Sunnyvale, CA
    3 days ago
  • $130k - $260k

     ...systems.Establish engineering patterns and best...  ...ranking, deep learning, representation learning...  ...other advanced machine learning...  ...evaluation, deployment, inference, monitoring, and...  ...StockSummaryThe Principal, Machine Learning...  ...mentoring technical staff and promoting... 
    Suggested
    Full time
    Temporary work
    Part time
    Work experience placement

    Walmart

    Sunnyvale, CA
    2 days ago
  • $209k - $313k

    Machine Learning Engineer Snap Inc is a technology company. We believe the camera presents the greatest opportunity to improve the way people...  ...Knowledge, Skills & Abilities: Strong understanding of causal inference and modern approaches to estimating treatment effects (e.g... 
    Suggested
    Live in
    Work at office
    Local area

    Snapchat

    Palo Alto, CA
    1 day ago
  • $169k - $338k

    (USA) Principal, Machine Learning Engineer Sunnyvale, CA (USA) Principal, Machine Learning Engineer 811 11 Th Ave Sunnyvale, CA 94089-4731...  ...solutions with organizational goals while mentoring technical staff and promoting engineering excellence. The successful... 
    Permanent employment
    Full time
    Temporary work
    Part time
    Work experience placement
    Local area
    Home office
    Flexible hours

    Jobleads-US

    Sunnyvale, CA
    1 day ago
  • $270k

     ...research lab of top researchers and engineers, building the world’s top-...  ...like health, fitness, learning, therapy, companions, customer...  ...models, optimizing realtime inference, and creating best-in-class APIs...  ...LinkedIn’s Top 10 Startups in the USA. Who We're Looking For A... 
    Full time
    Work at office
    Relocation package

    Inworld AI

    Mountain View, CA
    a month ago
  • ## Staff Machine Learning EngineerApply: Mountain View, CA, USA: Full time: Posted Yesterday: JOBREQ-2616280## **The opportunity...  ...a Staff Machine Learning Engineer to lead this effort. You will own...  ...in forecasting, causal inference, calibration or agentic systems... 
    Full time
    Work at office
    Worldwide

    Unity Enterprise

    Mountain View, CA
    3 days ago
  • $119.25k - $150.85k

     ...unprecedented scale. Within GM AV, the Model Deployment & Inference Solutions team deploys machine learning models from training frameworks (e.g., PyTorch) onto...  ...2028 on the Cadillac Escalade IQ, and we’re hiring engineers to help deliver the next generation of safe,... 
    Full time
    Internship
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $188.5k - $282.7k

     ...SAGE, Rubrik's Semantic AI Governance Engine, which is the first system designed to...  ...Engineering High-Performance Model Serving and Inference Infrastructure (25% of time)Designing...  ...(or higher) in Computer Science, Machine Learning, Computer Engineering, Statistics, or a... 
    Permanent employment

    Rubrik

    Palo Alto, CA
    2 days ago
  • $193.3k - $261.5k

     ...a passionate, talented, and inventive Machine Learning Engineer with a strong deep learning background,...  ...with other employees, supervisors, and staff; adhere to standards of excellence despite...  .... Learn more about our benefits at .USA, CA, Sunnyvale - 193,300.00 - 261,500.0... 
    Internship
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    3 days ago
  •  ...where we have a legal entity. Responsibilities Senior Machine Learning Engineer — Agentic Search & Query IntelligenceAtlassian is seeking...  ...improvements.Balance search quality with latency, inference cost, and reliability, building systems that operate effectively... 
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    4 days ago
  • $202.5k - $274k

     ...to make that possible.Job OverviewIntuit is looking for a Staff Machine Learning Engineer to own the data and platform layer beneath our consumer...  ...data path, training and evaluation frameworks, real-time inference serving, and the handoff into our decision engine, across... 
    Worldwide

    Intuit

    Mountain View, CA
    17 hours ago
  • $229k - $343k

     ...themselves, live in the moment, learn about the world, and have fun...  ...other digital services.Snap Engineering teams build fun and...  ...forefront.We're looking for a Machine Learning Engineering Manager...  ...applied ML, large-scale data and inference pipelines, LLM-powered workflows... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    Palo Alto, CA
    4 days ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ...next-generation AI and machine learning capabilities for the da...  ...robotic surgical platform. As a Staff Machine Learning Engineer, you...  ...structured modeling and inference approaches using statistical... 
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    1 day ago
  •  ...Position Summary We are looking for an experienced Machine Learning Engineer to lead the development of prompt injection and prompt safety...  ...systems. Build hybrid deployment pipelines that split safety inference between on-device (mobile, XR/AR) and cloud, optimizing... 
    Local area

    Trilyon, Inc.

    Mountain View, CA
    3 days ago
  • $175k - $230k

     ...Are Atoms is building the machines that power the next era of progress...  ...environments, operate them, learn from them, and improve them...  ...at scale. We are roboticists, engineers, operators, and builders. We believe...  ...bonus. Benefits Summary (USA Full‑Time Exempt Employees) ~... 
    Full time
    Temporary work
    Work at office
    Flexible hours

    Atoms

    Mountain View, CA
    2 days ago
  •  ...goal is to build the foundations to democratize AI and Machine Learning for Atlassian’s teams, customers, and ecosystem. We...  ...About This RoleAs a Senior ML System Engineer on the AI & ML Platform’s Inference team, you will design and optimize large-scale model... 
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    2 days ago
  • $172.2k - $283.9k

     ...Unity Technologies SF is hiring a Staff Machine Learning Engineer to build production-grade AI for game experiences with a focus on computer vision...  ...scenarios from cloud GPUs to efficient on-device inference Partner with research scientists to translate novel CV... 
    Work at office

    Unity Technologies SF

    Mountain View, CA
    3 days ago
  •  ...so the models that power these experiences must be deployed and accelerated entirely within that runtime. As our Principal Engineer for On-Device AI Inference & Systems, you will be the foremost engineering authority on taking state-of-the-art multi-modal models (... 

    Unity

    Mountain View, CA
    2 days ago
  •  ...and accelerated entirely within that runtime. As a Senior Machine Learning Engineer for On-Device & Mobile AI, you will take state-of-the-art...  ...the optimization and deployment of significant parts of the inference stack — from a trained checkpoint leaving research,... 

    Unity

    Mountain View, CA
    3 days ago
  • $270k

     ...Inworld in Mountain View is seeking engineers to advance real-time, agentic AI systems at scale. You will own end-to-end deployment from research to production, optimizing latency and throughput on multi-GPU infrastructure. Relocation assistance is offered; base salary... 
    Relocation package

    Jobleads-US

    Mountain View, CA
    1 day ago
  •  ...NVIDIA Corporation is seeking a Senior Machine Learning Applications and Compiler Engineer to advance end-to-end inference optimization across NVIDIA platforms. You will work at the intersection of large-scale systems, compilers, and deep learning to map neural network... 

    Jobleads-US

    Santa Clara, CA
    2 days ago
  • $195k - $230k

     ..., visit  About the Role We are looking for a Senior Machine Learning Engineer to help evolve our large-scale recommendation systems and...  ...metrics. Own systems from offline training → online inference → A/B experimentation → metric analysis . Identify and... 
    Full time
    Local area
    Work from home

    NewsBreak

    Mountain View, CA
    a month ago
  • $117k - $234k

     ...Data Scientist leads the development of advanced machine learning solutions that power Search relevance,...  ....Optimize model performance, scalability, and inference efficiency.Collaborate with Product Managers, Engineers, and Applied Scientists to define AI roadmaps.... 
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    17 hours ago
  •  ...researchers, data scientists, and engineers, tackling the most fundamental...  ...performance computing in deep learning, driving impactful discoveries...  ...performance for the machine learning software stacks, especially at training and inference, and support the team to develop... 
    Work experience placement
    Visa sponsorship

    Institute of Foundation Models

    Sunnyvale, CA
    4 days ago
  • $150k - $230k

     ...information, visit  About the Role We are looking for a hands-on Machine Learning Engineer to drive the post-training of our large language models,...  ...as Hugging Face TRL/Accelerate, DeepSpeed or FSDP, and inference engines like vLLM. Solid understanding of tokenization,... 
    Full time
    Local area
    Work from home

    NewsBreak

    Mountain View, CA
    a month ago
  • $172.5k - $313.7k

    ## Staff / Principal Data Scientist - SlackApply: Office Tech-Flexible: California - San Francisco...  ...’ll partner with Product, Design, and Engineering (PDE), as well as Go-To-Market (GTM)...  ...statistics, experimentation, causal inference, ML, and analytical problem-solving.*... 
    Full time
    Work at office
    Flexible hours

    Jobleads-US

    Palo Alto, CA
    1 day ago
  • $150k

     ...researchers, data scientists, and engineers, tackling the most...  ...performance computing in deep learning, driving impactful discoveries...  ...architectures. Training & Inference Integration – Connect training...  ...Experience with large-scale machine learning workloads (strong ML... 
    Full time
    Visa sponsorship
    Flexible hours

    Institute of Foundation Models

    Sunnyvale, CA
    1 day ago
  •  ...how teams work, discover, and create.We’re seeking a Principal Machine Learning Systems Engineer (P60) to lead technical directions of GenAI Products &...  ...production.Experience with large-scale model training, inference pipelines, or search/retrieval systems.ð¡ SkillsProficiency... 
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff / Principal Machine Learning Engineer, Inference - USA. Be the first to apply!