Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Engineering Manager, ML Training Infrastructure

$140k - $260k
Full-time

Woven by Toyota

Woven by Toyota is enabling Toyota’s once-in-a-century transformation into a mobility company. Inspired by a legacy of innovating for the benefit of others, our mission is to challenge the current state of mobility through human-centric innovation — expanding what “mobility” means and how it serves society.

Our work centers on four pillars: AD/ADAS, our autonomous driving and advanced driver assist technologies; Arene, our software development platform for software-defined vehicles; Woven City, a test course for mobility; and Cloud & AI, the digital infrastructure powering our collaborative foundation. Business-critical functions empower these teams to execute, and together, we’re working toward one bold goal: a world with zero accidents and enhanced well-being for all.

About the Team

Enterprise AI is a platform that provides end-to-end machine learning tooling experience to support and accelerate machine learning development, including autonomous driving and other related projects. Our platform serves customers as a standardized machine learning platform within Woven by Toyota as the larger Toyota Group companies.

The Enterprise AI ML Training Infrastructure team builds and operates the infrastructure that enables engineers and researchers to train, evaluate, and iterate on machine learning models at scale. Our customers include teams working on AD/ADAS and Woven City, where large-scale machine learning workloads, simulation, and data processing are critical to developing the next generation of mobility technologies.

The team operates at the intersection of machine learning, distributed systems, cloud infrastructure, and developer platforms. Our systems include large-scale GPU compute environments, Kubernetes-based cluster infrastructure, and workflow and pipeline orchestration technologies such as Ray, Airflow, and Temporal.

We are looking for an Engineering Manager to lead a team to help shape the technical direction, engineering culture, and long-term evolution of our ML training infrastructure platform.

WHO ARE WE LOOKING FOR?

As the Engineering Manager, ML Training Infrastructure, you will lead the team responsible for building and operating the infrastructure that powers machine learning training workloads across Woven by Toyota.

You will combine engineering leadership with strong technical judgment. You will work closely with engineering, research, and product stakeholders to understand customer needs, translate them into scalable technical solutions, and ensure that our platform is reliable, efficient, secure, and easy to use.

This role is particularly suited to an engineering leader who enjoys working on complex infrastructure problems and is comfortable operating across Kubernetes, distributed systems, GPU infrastructure, cloud platforms, and machine learning workloads.

You will also play an important role in building a strong engineering culture, developing engineers, establishing effective engineering practices, and partnering with our customers to continuously improve the platform.

What You Will Enable

In this role, you will help build the infrastructure that enables teams across Woven by Toyota to train and operate increasingly large and sophisticated machine learning workloads.

Your work will directly support engineers and researchers working on AD/ADAS, Woven City, and other mobility initiatives by making compute resources easier to access, workloads more reliable, and ML development faster and more scalable.

You will have the opportunity to shape both the technical architecture of a critical ML infrastructure platform and the engineering culture of the team building it.

\n RESPONSIBILITIES
  • Lead engineers and foster a collaborative and innovative environment.

  • Manage and nurture a group of around 7–8 engineers, offering technical vision, mentorship, professional growth, and effective project delivery.

  • Establish and articulate the technical vision and future roadmap for machine learning training infrastructure.

  • Direct the architecture, construction, and operation of resilient systems supporting massive ML model training workloads.

  • Construct and maintain Kubernetes-driven compute clusters and container ecosystems tailored for machine learning workloads.

  • Oversee and optimize GPU fleets, addressing cluster scheduling, resource utilization, system stability, and cost performance.

  • Advance distributed model training capabilities and workload management leveraging frameworks like Ray and adjacent tools.

  • Enhance workflow orchestration platforms by deploying and supporting solutions such as Airflow, Temporal, and similar pipeline frameworks.

  • Elevate the system resilience, monitoring insights, elasticity, and operational productivity of our ML training stack.

  • Work closely alongside AD/ADAS, Woven City, and internal platform partners to gather feature requests and ship infrastructure solutions that accelerate development.

  • Partner with cross-functional engineering groups to specify platform interfaces, APIs, system architectures, and operational standards.

  • Translate high-level organizational objectives and user requirements into concrete, executable engineering milestones with product leads.

  • Institute rigorous software practices across implementation, code reviews, automated testing, continuous deployment, and live site management.

  • Engage in production support workflows, lead post-incident reviews, and champion system refinements based on operational findings.

  • Balance immediate client feature requests against long-term architectural stability, technical hygiene, and system maintainability.

  • Foster an inclusive, highly collaborative, and excellence-driven engineering culture throughout Woven by Toyota.

MINIMUM QUALIFICATIONS
  • 8+ years of experience in software engineering, with at least 2 years in a leadership role.

  • Proven track record mentoring and managing high-performing software or platform teams.

  • Demonstrated expertise in architecting, deploying, and maintaining large-scale mission-critical infrastructure.

  • Deep technical mastery of Kubernetes orchestration alongside containerized runtime systems.

  • Hands-on exposure to cloud platform architectures and highly resilient distributed systems.

  • Background engineering platform foundation layers designed for machine learning model development or data pipeline intensive workloads.

  • Practical understanding of massive GPU cluster management, hardware scheduling algorithms, and capacity efficiency tuning.

  • Operational background overseeing high-availability environments, leading post-incident reviews, and advancing monitoring capabilities.

  • Excellent interpersonal abilities to align technical roadmaps across diverse engineering partners and leadership stakeholders.

NICE TO HAVES
  • Experience working with Japanese customers or clients.

  • Experience with Golang development (especially for cloud services)

  • Experience in the automotive industry or knowledge of automotive software development (e.g., ASPICE, V-model).

  • Experience working in environments with a focus on security and safety.

  • English/Japanese bilingual ability

\n

The base pay for this position ranges from $140,000 - $260,000 a year.

Your base salary is one part of your total compensation. We offer a base salary, short term and long term incentives, and a comprehensive benefits package. The total compensation offered to an employee will be dependent upon the individual's skills, experience, qualifications, location, and level.

WHAT WE OFFER

We are committed to creating a modern work environment that supports our employees and their loved ones. We offer many options of the best programs to allow you to do your most meaningful work and to help you shape the future of mobility.

・Excellent health, wellness, dental and vision coverage

・A rewarding 401k program

・Flexible vacation policy

・Family planning and care benefits

Our Commitment

・We are an equal opportunity employer and value diversity.

・Any information we receive from you will be used only in the hiring and onboarding process. Please see our privacy notice for more details.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Engineering Manager, ML Training Infrastructure in Palo Alto, CA vacancy
  •  ...Head Of Engineering At Poesis Whoever builds the leading...  ...portfolios, and manage risk. Our founders managed...  ...) and led enterprise ML at Goldman Sachs and Amazon...  ...live capital is the training ground for an...  ...integrity. Designing the infrastructure that enables... 
    Training
    Full time
    Work at office
    Visa sponsorship
    Work visa
    Relocation package
    3 days per week

    Poesis LLC

    Menlo Park, CA
    13 hours ago
  • $193.93k - $352.29k

     ...What is not fungible is the infrastructure that decides whether an autonomous...  ...inside Nuro's own engineering organization, under the same...  ...that loop runs against the training pipelines behind the driving...  ...helped. ~ Experience with ML training or research infrastructure... 
    Training
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    a month ago
  • $230k - $260k

     ...a Principal Machine Learning Engineer, you willoperateat the company...  ...the design of large-scale ML systems and shared platforms...  ...design ofscalable ML platforms(training, evaluation, inference, safety...  ...generationDefine and evolveshared ML infrastructure and platformsfor training,... 
    Training
    Work at office
    Immediate start
    3 days per week

    Typeface

    Palo Alto, CA
    1 day ago
  • $190k - $260k

     ...as the speed at which we can train it. Every improvement to our...  ...scale world models – depends on infrastructure that turns thousands of hours...  .... We are looking for engineers who make model training fast:...  ...streamable formats Partner with ML teams to scale new architectures... 
    Training
    Temporary work
    Work at office
    Visa sponsorship
    Flexible hours

    Kodiak

    Mountain View, CA
    23 days ago
  • $255k - $300k

     ...giving a growing number of engineers and employees an AI...  ...! As Engineering Manager, you will lead the...  ...in developing critical infrastructure for Agentic applications...  ...expertise in applied ML at scale. You'll partner...  ...include education, training, experience, location,... 
    Training
    Full time
    Work at office
    Flexible hours
    Shift work

    Robinhood

    Menlo Park, CA
    1 day ago
  •  ...Implements machine learning (ML) models for production...  ...workflows. Creates infrastructure and frameworks to...  ...Development Leads, Product Management, Operations, and...  ...with design criteria of trained models and/or systems....  ...Machine Learning, Computer Engineering, Mathematics, Physics,... 
    Training
    Shift work

    Oracle

    Santa Clara, CA
    2 days ago
  • $207k - $300k

     ...Engineering Manager, Applied AI, Cloud Support Intelligence Google Cloud's mission is...  ...combining cutting-edge technology, infrastructure, and talent. AI/ML software engineers in Cloud bridge...  ..., and relevant education or training. US: $207000 - $300000 (USD) + 20%... 
    Training
    Worldwide

    Google

    Sunnyvale, CA
    1 day ago
  •  ...worldwide to plan, track, and manage their work effectively...  ...Machine Learning Engineer to join our Search &...  ..., models, and infrastructure that accelerate AI feature...  ...mentorship to emerging ML engineers. Beyond these...  ...haveExperience fine-tuning, post-training, and optimizing state-... 
    Training
    Work at office
    Local area
    Worldwide

    Atlassian

    Mountain View, CA
    19 hours ago
  • $250k - $320k

     ...Staff Infrastructure Engineer We are partnered with a Stealth AI Lab (backed by top-tier investors and...  ...optimize the inference platform, GPU‑based training clusters, and data processing pipelines...  ...5+ years of experience in Software / ML Infrastructure Engineering. Deep... 
    Training
    Full time
    Immediate start

    Strativ Group

    Menlo Park, CA
    3 days ago
  • $160k - $225k

     ...world's best data and AI infrastructure platform so our...  ...improve their business. Training and customizing state‑...  ...Runtime (AIR) is our managed platform for large‑scale...  ...As a Senior Software Engineer for AI Runtime, you...  ...performance computing, or ML systems. ~... 
    Training
    Local area

    Databricks

    Mountain View, CA
    1 day ago
  • $115k - $150k

     ...technologies. As a Senior Project Manager, you will lead complex,...  ...that bring together engineering, design, construction, and client...  ...deliver critical manufacturing infrastructure. Recognized as one of the...  ...day one, you'll receive the training and support needed to leverage... 
    Training
    For contractors
    Currently hiring
    Work at office
    Flexible hours

    SSOE, Inc.

    Santa Clara, CA
    13 hours ago
  • $182k - $242k

     ...enterprises, CoreWeave combines superior infrastructure performance with deep technical...  ...You'll Do CoreWeave is looking for an Engineering Manager to lead a team building and operating...  ...systems that power high-performance AI and ML workloads. You will lead engineers... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    16 days ago
  •  ...Principal Machine Learning Systems Engineer (P60) to lead technical...  ...with reliable, high-performance infrastructure.Working at AtlassianAtlassians...  ...What You'll Doð Design and Build ML SystemsArchitect and implement scalable systems for training, fine-tuning, and serving... 
    Training
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    19 hours ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering...  ..., networking, cloud infrastructure, Kubernetes, databases, capacity management, continuous delivery,...  ...Security, Networking, and AI/ML organizations to make...  ..., inference systems, training environments, or high-... 
    Training
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  •  ...advanced AI, data, and engineering capabilities. Our...  ...runtime, model serving, GPU infrastructure, distributed systems,...  ..., long-context management, conversation memory,...  ...user experience. 6. Training & Research Infrastructure...  ...systems architecture, ML infrastructure, distributed... 
    Training

    Accellor

    Mountain View, CA
    10 days ago
  • $145k

     ...dynamic workplace environments. As a CBRE Engineering Ops Manager, you will be embedded onsite at the...  ...of critical building systems and infrastructure. You will serve as a key partner to client...  ...' growth and development through training programs, certifications, and mentorship... 
    Training
    Contract work
    Work at office
    Local area
    Visa sponsorship
    Flexible hours

    CBRE

    Sunnyvale, CA
    4 days ago
  • $112.7k - $169.1k

     ...world's leading game engine. Recommendation and ranking...  ...full pipeline from training data to deployed model...  ...large‑scale data and ML systems, whether through...  ..., or experimentation infrastructure. Experience...  ...website or directly to managers. Unity does not accept... 
    Training
    Internship
    Work at office
    Worldwide
    Relocation package
    Shift work

    Unity South APAC (SEA, ANZ, IND Subcont.)

    Mountain View, CA
    2 days ago
  • $181.1k - $245k

     ...Photonics & PcbA Test Manager Aeva's mission is to bring the...  ...while managing a team of test engineers and technicians. The...  ...manufacturing and can drive test infrastructure, DFT, yield improvement, and...  ..., relevant education or training, and market conditions. These... 
    Training
    Contract work
    Overseas
    Flexible hours

    Aeva, Inc

    Mountain View, CA
    2 days ago
  • $160.36k - $240.54k

     ...autonomous driving technology. In an ML-first system, the overall...  ...and diversity of its training and evaluation data. The team...  ...a scalable and reliable data infrastructure. This infrastructure is designed...  ...closely with system engineers to thoroughly validate the autonomous... 
    Training
    Work experience placement

    Kindredventures

    Mountain View, CA
    3 days ago
  • $150k - $230k

     ...researchers and veteran systems engineers who share a vision for...  ...increasingly complex, traditional infrastructure struggles to meet the demands...  ...-performance distributed GPU training. You'll work at the...  ...networking (RDMA, InfiniBand) ML framework or runtime internals... 
    Training

    Clockwork.io

    Palo Alto, CA
    11 days ago
  •  ...industry-leading total spend management platform for businesses...  ...The Impact of a Principal Engineer, AI/ML Architecture at Coupa: Coupa...  ..., designing how we train, evaluate, and serve models...  ...partnership evaluations with AI infrastructure and model providers. Architect... 
    Training

    Coupa Software, Inc.

    Foster, CA
    5 days ago
  • $120k - $132k

     ...automation, and other critical infrastructure. This leader will...  ...developing a high-performing Engineering team while ensuring that the...  ...Capital Projects & Vendor Management Lead and coordinate capital...  ...Engineering team through training, cross-training, and hands-... 
    Training
    Temporary work
    For contractors

    Ensemble Hospitality

    Menlo Park, CA
    8 days ago
  • $170k - $200k

     ...Description Job Description Senior Engineering Manager (LAND DEVELOPMENT) - Santa Clara, CA...  ..., utility systems, and stormwater infrastructure in accordance with jurisdictional standards...  ...and business planning. May lead training initiatives on various technical... 
    Training
    Contract work
    Work at office

    Kier & Wright

    Santa Clara, CA
    16 days ago
  • $240k - $290k

     ...Sonatus is looking for an experienced Senior Engineering Manager to build and lead our AI Validation...  ...end-to-end validation strategy for AI/ML models, LLMs, RAG pipelines, and...  ...workflows—from data pipelines and model training through cloud services and in-vehicle deployment... 
    Training
    Work at office
    Worldwide
    Flexible hours
    Shift work

    Sonatus

    Sunnyvale, CA
    15 days ago
  • $236.7k - $309.03k

     ...experimentation, quality, and production systems. You’ll have the opportunity to work across a number of areas including model post training, personalization/recommendation, proactive and other agentic experiences. You’ll tackle ambiguous problems end to end: defining... 
    Training
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    1 day ago
  •  ...team leading the equipment lifecycle management, service and engineering function for a variety of equipment...  ...equipment management practices, infrastructure, and technical capabilities that will...  ...systems. Maintain equipment training programs and ensure laboratory personnel... 
    Training
    Relocation
    Flexible hours
    3 days per week

    DELFI Diagnostics, Inc.

    Palo Alto, CA
    20 days ago
  • $150k - $224k

     ...Design and build data processing infrastructure for model training and feature serving, optimizing for performance...  ...tooling for data infrastructure used across ML teams Required Qualifications Strong software engineering fundamentals, with experience building... 
    Training
    Full time

    AppLovin

    Palo Alto, CA
    6 days ago
  • $148.32k - $203.94k

     ...visit: Job Summary We are seeking a hands-on Principal Infrastructure Hardware Engineer to architect, design, and deliver system platforms...  ...by factors such as location, experience, education, and training. In addition to base salary, this role is eligible for... 
    Training

    SiTime Corporation

    Santa Clara, CA
    17 days ago
  • $139.9k - $274.8k

     ...apply for the Principal Design Engineer role at Microsoft ....  ...Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) is the team...  ...operations, globalization, and manageability solutions. Our focus is on...  ...Intelligence (AI) / Machine Learning (ML) SoCs Working knowledge of... 
    Work at office
    Local area
    Worldwide

    Microsoft Corporation

    Mountain View, CA
    1 day ago
  • $180k

     ...motivated, and focused on engineering excellence. This organization...  ...ABOUT THE ROLE: As an ML Infrastructure Engineer, you will play a pivotal...  ...compute infrastructure, training frameworks, and experimentation...  ....g., Slurm), configuration management (Puppet/Ansible), or... 
    Training
    Temporary work
    Work experience placement

    SpaceXAI

    Palo Alto, CA
    10 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Engineering Manager, ML Training Infrastructure. Be the first to apply!