Staff Machine Learning Engineer - Vision-Language Foundation Models
$251k - $310kWaymo
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to mobility while saving thousands of lives now lost to traffic crashes. The Waymo Driver powers Waymo’s fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states.
The Team & Mission: In the Oracle Perception team, our mission is to build the ultimate cognitive engine for autonomous driving. We are pioneering the use of large multimodal foundation models (e.g., Gemini) to build a powerful offboard reasoning and data flywheel system. We are moving beyond traditional perception to true scene understanding and driving actions—building offboard models that can comprehend complex driving problems, predict object/scene dynamics, and deduce driving paths with logical rationale.
Our core focus is advancing the VLM foundation itself. By pushing the boundaries of multimodal pre-training and state-of-the-art post-training (SFT, RL) , we are creating models capable of rich, reasoning-based autolabeling at a massive scale. This closed-loop data engine directly powers the training and evolution of Waymo's real-time onboard models. If you are passionate about defining VLM training recipes, scaling laws, and unlocking complex reasoning via RL, this is your opportunity to redefine the foundation of autonomous driving.
In this hybrid role, you will report to a Senior Staff Technical Lead Manager.
You Will:
- Drive Pre-training & Domain Adaptation: Lead the technical strategy for curating and constructing massive-scale, high-quality multimodal pre-training datasets. Define data mixture strategies to instill deep, Waymo-specific driving intuition and physics-grounded understanding into foundation models without catastrophic forgetting.
- Lead Post-Training & Reasoning Enhancement: Design and implement state-of-the-art fine-tuning (SFT) and Reinforcement Learning (RLHF/RLAIF, DPO/GRPO/PPO) pipelines. Drastically improve the model’s instruction-following and complex reasoning capabilities (e.g., Chain-of-Thought, spatial-temporal reasoning, and driving rationale prediction).
- Pioneer the VLM Data Flywheel: Architect the highly scalable inference and evaluation pipelines that leverage these trained Gemini-class models to autonomously source, sample, and autolabel critical edge cases, directly accelerating the onboard perception models.
- Define Training Recipes & Scaling Laws: Conduct rigorous ablation studies to optimize model architectures, token budgets, and loss functions. Establish best practices for scaling multimodal training efficiently on large GPU/TPU clusters.
- Drive Cross-Functional AI Strategy: Act as the principal technical visionary across ML Infra, Perception, Behavior, and AI Foundation teams. Drive consensus on the data flywheel architecture and embed VLM reasoning capabilities seamlessly into the broader autonomous vehicle stack.
- Provide Staff-Level Technical Leadership: Own the long-term technical roadmap for foundation model development. Mentor senior engineers, lead rigorous design reviews, and establish standard-setting engineering practices from advanced prototyping to production deployment.
You Have:
- Master’s degree in Computer Science, AI, ML, or a related technical field.
- 8+ years of hands-on experience designing, training, and scaling deep learning models, with at least 3+ years focused deeply on training Large Language Models (LLMs) or Vision-Language Models (VLMs) .
- Proven expertise in the full lifecycle of Foundation Models: from pre-training data curation (interleaved formats, tokenization) and distributed training to advanced post-training techniques.
- Expert-level understanding of training infrastructure and distributed paradigms (e.g., FSDP, Megatron, JAX/Pax) required for training massive models reliably.
- Expert-level software engineering fundamentals using Python, PyTorch, or JAX, with a track record of building reliable, highly scalable ML systems.
- Proven ability to operate with high ambiguity, define technical roadmaps, and drive complex, multi-quarter technical initiatives across multiple teams in a fast-paced environment.
We Prefer:
- PhD in Computer Science, Artificial Intelligence, or a related field.
- Strong publication record in top-tier AI venues (e.g., NeurIPS, ICML, ICLR, CVPR) focusing on foundation models, large-scale training, reinforcement learning, or reasoning.
- Deep experience with advanced Reinforcement Learning paradigms applied to language or vision tasks ( focusing on improving System 2 thinking, logical deduction, and model alignment ).
- Demonstrated experience in Data Engineering for Foundation Models at the scale of billions/trillions of tokens (e.g., deduplication, quality filtering, synthetic data generation).
- Familiarity with the systemic challenges of multimodal perception in robotics or autonomous driving (e.g., 3D scene understanding, trajectory prediction).
- A proven track record of Staff-level impact: influencing product direction, pioneering zero-to-one ML architectures, and multiplying team efficiency through technical leadership.
The expected base salary range for this full-time position across US locations is listed below. Actual starting pay will be based on job-related factors, including exact work location, experience, relevant training and education, and skill level. Your recruiter can share more about the specific salary range for the role location or, if the role can be performed remote, the specific salary range for your preferred location, during the hiring process.
Waymo employees are also eligible to participate in Waymo’s discretionary annual bonus program, equity incentive plan, and generous Company benefits program, subject to eligibility requirements.
Salary Range
$251,000—$310,000 USD
$204k - $259k
...serving as the foundation for training and... ...advanced ML and engineering team that leverages... ...-art computer vision, deep learning, and generative... ...report to a Senior Staff Technical Lead... .../ multimodal models (e.g., Gemini) to... ...strategies for Vision-Language Models (VLMs) to...LanguageFoundationFull timeRemote work$311.85k - $370k
...The role As a Senior Machine Learning Engineer on Wayve's Measurement team in AI Evaluation... ..., you will build the computer vision and scene understanding models Wayve uses to measure the performance... ...our on-vehicle models and Wayve Foundation Models into offline models that...FoundationFull timeWork at officeWork from home$311.85k - $370k
...advanced AI software and foundation models enable vehicles to perceive... ...driving systems. Our vision is to create autonomy that... ...of excellence, constantly learning and evolving as we pave the... ...The role As a Senior Machine Learning Engineer on Wayve's Measurement...FoundationFull timeWork at officeWork from home$244.14k - $413.16k
...cutting-edge R&D in AI, machine learning, and smart... ...developing large-scale Vision-Language-Action (VLA) models and World Models to handle... ...driving. As a Senior Staff Machine Learning Engineer, you will architect the... ...laws for driving foundation models, overseeing data...LanguageFoundationFull timeOverseas$174.72k - $295.68k
...through cutting-edge R&D in AI, machine learning, and smart connectivity.We... ...-time Machine Learning Engineer / Research Scientist to drive the modeling and algorithmic... ...of XPENG’s next-generation Vision-Language-Action (VLA) Foundation Model — the core brain that...LanguageFoundationFull time$174.72k - $295.68k
...-edge R&D in AI, machine learning, and smart connectivity... ...Machine Learning Engineers with strong... ...expertise in generative modeling and large-scale... ...in computer vision, generative AI, and... ...improve Vision-Language-Action (VLA) driving... ...with multimodal foundation models and video...LanguageFoundationFull time- ...We are Foundation Model Inference Team, within AI, Search & Knowledge... ...billions of parameter language and vision and speech models using state... ...cases.Mentor and guide engineers in the organization.... ...Artificial Intelligence, Machine Learning, Information Retrieval, Data...LanguageFoundation
$165k - $185k
...Silicon Valley focuses on Foundation Models, Big Data Visual... ...Explainable AI (XAI), Natural Language Processing, Computer Vision & Mixed Reality, Cloud... ...Science, AI System Engineering, Time-series Analysis.... ...engineering in core AI and machine learning fields to enable...LanguageFoundationWork experience placementLocal areaWorldwide- ...About the Institute of Foundation Models We are a dedicated... ...scientists, and engineers, tackling the most fundamental... ...computing in deep learning, driving impactful... ...Scientist in the Vision Language Model (VLM) team,... ...research experience in Machine Learning, Computer...LanguageFoundation
- ...worldwide.We’re a team of engineers, clinicians, and... ...(Computer Vision), you will develop the perception models that let our Embodied... ...flagging, active learning, human-in-the-loop... ...action (VA) / vision-language-action (VLA)... ...research stage.Solid foundations in linear algebra,...LanguageFoundationLocal areaWorldwideFlexible hours
- ...AI software and foundation models enable vehicles to... ...systems. Our vision is to create autonomy... ..., constantly learning and evolving as we... ...As an ML Engineer within the Application... ...for success as a Machine Learning Engineer... ...and other relevant languages (e.g. C++ and CUDA...LanguageFoundationFull timeWork at officeWork from home
$150k
...the Institute of Foundation Models We are a dedicated... ...scientists, and engineers, tackling the most... ...computing in deep learning, driving impactful... ...The Role As a Machine Learning Engineer... ...infrastructure, Natural Language Processing or Computer Vision. ~2 years of...LanguageFoundationFull timeWorldwideVisa sponsorship$160k - $225k
...expand our product and engineering teams, bringing our vision of intelligent,... .... As an early Machine Learning Engineer at MAI,... ...stack , from the foundational data platforms that... ...experience. Bring Models to Life: You will... ...with Large Language Models, agentic frameworks...LanguageFoundationFull time$204k - $259k
...the system which learns the spatial-... ...sensors, enabling engineers like you to (1)... ...to (2) develop models and model training... ...cutting-edge VLM foundation models. Conduct... ...in large language models, vision-language models,... ...of experience in Machine Learning, with a...LanguageFoundationFull timeRemote work- ...Labor for dull, dirty, and dangerous work. The team develops vision-language models and world models to enable safe, real-world robot... ...modal perception, planning, and action policies, advancing foundation models, and delivering production-grade solutions on RoboForce...LanguageFoundationWork at office
$184.7k - $324.8k
...Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference Santa Clara, California, United States Machine Learning and AI... ...extracted from the hardware beneath them. We optimize language, vision, and speech models with billions of parameters using...LanguageFoundationWorldwideRelocation$184k - $287.5k
...application is built. We are seeking a senior vision language model engineer to design and build agentic data and... ...background in modern deep learning, including transformer‑based architectures... ...modeling, and multimodal VLM/VLA or foundation models.Excellent experience training...LanguageFoundationFull time- ...monsters and make this vision a reality. We are... ...seamlessly connect to foundational models - whether via APIs or... ...effective. We are seeking Staff and Principal level Machine Learning Engineers to solve these... ...algorithms in natural language processing, speech processing...LanguageFoundationFull timeWork at officeRelocation packageFlexible hoursShift work
$407.33k
...The role As a Principal Engineer on the Model Foundations team you will build the geometric vision and 3D foundation models that underpin... ...of large-scale deep learning, geometric computer vision, and... ...up for success as a Principal Machine Learning Engineer, Geometric...FoundationFull timeWork at officeWork from home$130k - $260k
.... Role Overview The vision of the Documents and Vision... ...of business. As a Staff Machine Learning Engineer, you will serve as a technical... ...implement machine learning models, services, and components... ...Statistical Modeling ~ Strong foundation in advanced machine...FoundationHourly payFull timeWork experience placementLocal area- # Machine Learning Engineer, Multimodal AIGoogle DeepMind## Job Description###... ...machine learning models capable of understanding vision, audio, text, and video... ...optimize state-of-the-art foundation models for real-world... ...vision, natural language processing, or multimodal...LanguageFoundationFlexible hours
$204k - $259k
...Design, train, and deploy machine learning models to automate the creation... ...ML techniques, including Vision-Language Models (VLMs) and other Generative... ...Perception and Waymo AI Foundations, to adapt cutting-edge... ...-time Job function Engineering and Information...LanguageFoundationFull timeRemote work- ...Vice President, AI & Machine Learning Engineering About the Company A highly... ..., and natural language interfaces. The ideal candidate... ...understanding of LLMs and foundation models, and experience in building... ...in NLP, computer vision, or advanced AI research...LanguageFoundation
- ....We’re a team of engineers, clinicians, and... ...platforms. As a Staff AI/ML Architect,... ...which a high-level model interprets sensory... ...prototyping and learning while working toward... ...and multimodal, vision-language, LLM-based... ...including vision foundation models (VFM), vision...LanguageFoundationContract workLocal areaWorldwideFlexible hours
$175k - $296k
...We are looking for a full-time Machine Learning Engineer, with deep knowledge and strong enthusiasm... ...for training very large foundation model and accelerating model training/inference... ...in training large scale vision or language models Previous experience in the...LanguageFoundationFull time- ...for you? Constant learning, skill growth,... ...process knowledge. Our vision is to infuse... ...businesses operate. Large Language Models (LLMs) hold... ...the landscape of Machine Learning across... ...nimble and versatile engineering team to empower... ...within the AI Foundations and Research team...LanguageFoundationPermanent employmentFull timeWorldwideFlexible hours
- ...About the Institute of Foundation Models We are a dedicated... ...data scientists, and engineers, tackling the most... ...performance computing in deep learning, driving impactful... ...models to unlock machine intelligence beyond... ...individuals who share our vision and are eager to push...FoundationVisa sponsorship
$164k - $282k
...-house robotics, computer vision, and AI — at a scale and speed... ...Planning and Control Engineer to develop robust, high-performance... ...planning, control, learning, and foundation models, and evaluate their... ...approaches, including vision-language-action (VLA) models, world...LanguageFoundationTemporary work$222.72k - $389.75k
...experience and abilities, we’ll explore your foundational skills and how you collaborate with AI.... ...process here.We are looking for a Staff Machine Learning Engineer to lead the technical vision for our Ads Conversion Core Modeling team, building the state-of-the-art...FoundationWork at officeLocal areaRelocationRelocation package$150.4k - $277.6k
Applied Machine Learning Research Engineer - Multimodal for Human Understanding Sunnyvale, California, United States... ...incredible potential of multimodal foundation and large language models, and many applications in the computer vision and machine learning domain that...LanguageFoundationWorldwideRelocation
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Machine Learning Engineer - Vision-Language Foundation Models. Be the first to apply!
- assistant engineer Mountain View, CA
- staff engineer Mountain View, CA
- software engineer staff Mountain View, CA
- senior staff engineer Mountain View, CA
- senior staff systems engineer Mountain View, CA
- technology administrator Mountain View, CA
- engineering aide Mountain View, CA
- ai ml engineer Mountain View, CA
- senior ml engineer Mountain View, CA
- computer vision machine learning engineer Mountain View, CA



