Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Engineering Manager, ML Training Infrastructure

$140k - $260k

Woven by Toyota

Job Description

Job Description

Woven by Toyota is enabling Toyota’s once-in-a-century transformation into a mobility company. Inspired by a legacy of innovating for the benefit of others, our mission is to challenge the current state of mobility through human-centric innovation — expanding what “mobility” means and how it serves society.

Our work centers on four pillars: AD/ADAS, our autonomous driving and advanced driver assist technologies; Arene, our software development platform for software-defined vehicles; Woven City, a test course for mobility; and Cloud & AI, the digital infrastructure powering our collaborative foundation. Business-critical functions empower these teams to execute, and together, we’re working toward one bold goal: a world with zero accidents and enhanced well-being for all.

About the Team

Enterprise AI is a platform that provides end-to-end machine learning tooling experience to support and accelerate machine learning development, including autonomous driving and other related projects. Our platform serves customers as a standardized machine learning platform within Woven by Toyota as the larger Toyota Group companies.

The Enterprise AI ML Training Infrastructure team builds and operates the infrastructure that enables engineers and researchers to train, evaluate, and iterate on machine learning models at scale. Our customers include teams working on AD/ADAS and Woven City, where large-scale machine learning workloads, simulation, and data processing are critical to developing the next generation of mobility technologies.

The team operates at the intersection of machine learning, distributed systems, cloud infrastructure, and developer platforms. Our systems include large-scale GPU compute environments, Kubernetes-based cluster infrastructure, and workflow and pipeline orchestration technologies such as Ray, Airflow, and Temporal.

We are looking for an Engineering Manager to lead a team to help shape the technical direction, engineering culture, and long-term evolution of our ML training infrastructure platform.

 

WHO ARE WE LOOKING FOR?

As the Engineering Manager, ML Training Infrastructure, you will lead the team responsible for building and operating the infrastructure that powers machine learning training workloads across Woven by Toyota.

You will combine engineering leadership with strong technical judgment. You will work closely with engineering, research, and product stakeholders to understand customer needs, translate them into scalable technical solutions, and ensure that our platform is reliable, efficient, secure, and easy to use.

This role is particularly suited to an engineering leader who enjoys working on complex infrastructure problems and is comfortable operating across Kubernetes, distributed systems, GPU infrastructure, cloud platforms, and machine learning workloads.

You will also play an important role in building a strong engineering culture, developing engineers, establishing effective engineering practices, and partnering with our customers to continuously improve the platform.

 

What You Will Enable

In this role, you will help build the infrastructure that enables teams across Woven by Toyota to train and operate increasingly large and sophisticated machine learning workloads.

Your work will directly support engineers and researchers working on AD/ADAS, Woven City, and other mobility initiatives by making compute resources easier to access, workloads more reliable, and ML development faster and more scalable.

You will have the opportunity to shape both the technical architecture of a critical ML infrastructure platform and the engineering culture of the team building it.

RESPONSIBILITIES

  • Lead engineers and foster a collaborative and innovative environment.

  • Manage and nurture a group of around 7–8 engineers, offering technical vision, mentorship, professional growth, and effective project delivery.

  • Establish and articulate the technical vision and future roadmap for machine learning training infrastructure.

  • Direct the architecture, construction, and operation of resilient systems supporting massive ML model training workloads.

  • Construct and maintain Kubernetes-driven compute clusters and container ecosystems tailored for machine learning workloads.

  • Oversee and optimize GPU fleets, addressing cluster scheduling, resource utilization, system stability, and cost performance.

  • Advance distributed model training capabilities and workload management leveraging frameworks like Ray and adjacent tools.

  • Enhance workflow orchestration platforms by deploying and supporting solutions such as Airflow, Temporal, and similar pipeline frameworks.

  • Elevate the system resilience, monitoring insights, elasticity, and operational productivity of our ML training stack.

  • Work closely alongside AD/ADAS, Woven City, and internal platform partners to gather feature requests and ship infrastructure solutions that accelerate development.

  • Partner with cross-functional engineering groups to specify platform interfaces, APIs, system architectures, and operational standards.

  • Translate high-level organizational objectives and user requirements into concrete, executable engineering milestones with product leads.

  • Institute rigorous software practices across implementation, code reviews, automated testing, continuous deployment, and live site management.

  • Engage in production support workflows, lead post-incident reviews, and champion system refinements based on operational findings.

  • Balance immediate client feature requests against long-term architectural stability, technical hygiene, and system maintainability.

  • Foster an inclusive, highly collaborative, and excellence-driven engineering culture throughout Woven by Toyota.

MINIMUM QUALIFICATIONS

  • 8+ years of experience in software engineering, with at least 2 years in a leadership role.

  • Proven track record mentoring and managing high-performing software or platform teams.

  • Demonstrated expertise in architecting, deploying, and maintaining large-scale mission-critical infrastructure.

  • Deep technical mastery of Kubernetes orchestration alongside containerized runtime systems.

  • Hands-on exposure to cloud platform architectures and highly resilient distributed systems.

  • Background engineering platform foundation layers designed for machine learning model development or data pipeline intensive workloads.

  • Practical understanding of massive GPU cluster management, hardware scheduling algorithms, and capacity efficiency tuning.

  • Operational background overseeing high-availability environments, leading post-incident reviews, and advancing monitoring capabilities.

  • Excellent interpersonal abilities to align technical roadmaps across diverse engineering partners and leadership stakeholders.

NICE TO HAVES

  • Experience working with Japanese customers or clients.

  • Experience with Golang development (especially for cloud services)

  • Experience in the automotive industry or knowledge of automotive software development (e.g., ASPICE, V-model).

  • Experience working in environments with a focus on security and safety.

  • English/Japanese bilingual ability

The base pay for this position ranges from $140,000 - $260,000 a year.

Your base salary is one part of your total compensation. We offer a base salary, short term and long term incentives, and a comprehensive benefits package. The total compensation offered to an employee will be dependent upon the individual's skills, experience, qualifications, location, and level.

WHAT WE OFFER

We are committed to creating a modern work environment that supports our employees and their loved ones. We offer many options of the best programs to allow you to do your most meaningful work and to help you shape the future of mobility.

・Excellent health, wellness, dental and vision coverage

・A rewarding 401k program

・Flexible vacation policy

・Family planning and care benefits

Our Commitment

・We are an equal opportunity employer and value diversity.

・Any information we receive from you will be used only in the hiring and onboarding process. Please see our privacy notice for more details.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Engineering Manager, ML Training Infrastructure in Palo Alto, CA vacancy
  • $140k - $260k

     ...and Cloud & AI, the digital infrastructure powering our collaborative foundation...  .... The Enterprise AI ML Training Infrastructure team builds...  ...infrastructure that enables engineers and researchers to train,...  ...looking for an Engineering Manager to lead a team to help shape... 
    Training
    Full time
    Temporary work
    Work at office
    Immediate start
    Flexible hours

    Woven by Toyota

    Palo Alto, CA
    3 days ago
  •  ...Head Of Engineering At Poesis Whoever builds the leading...  ...portfolios, and manage risk. Our founders managed...  ...) and led enterprise ML at Goldman Sachs and Amazon...  ...live capital is the training ground for an...  ...integrity. Designing the infrastructure that enables... 
    Training
    Full time
    Work at office
    Visa sponsorship
    Work visa
    Relocation package
    3 days per week

    Poesis LLC

    Menlo Park, CA
    18 hours ago
  • $193.93k - $352.29k

     ...What is not fungible is the infrastructure that decides whether an autonomous...  ...inside Nuro's own engineering organization, under the same...  ...that loop runs against the training pipelines behind the driving...  ...helped. ~ Experience with ML training or research infrastructure... 
    Training
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    a month ago
  • $230k - $260k

     ...a Principal Machine Learning Engineer, you willoperateat the company...  ...the design of large-scale ML systems and shared platforms...  ...design ofscalable ML platforms(training, evaluation, inference, safety...  ...generationDefine and evolveshared ML infrastructure and platformsfor training,... 
    Training
    Work at office
    Immediate start
    3 days per week

    Typeface

    Palo Alto, CA
    2 days ago
  • $190k - $260k

     ...as the speed at which we can train it. Every improvement to our...  ...scale world models – depends on infrastructure that turns thousands of hours...  .... We are looking for engineers who make model training fast:...  ...streamable formats Partner with ML teams to scale new architectures... 
    Training
    Temporary work
    Work at office
    Visa sponsorship
    Flexible hours

    Kodiak

    Mountain View, CA
    24 days ago
  • $182k - $242k

     ...enterprises, CoreWeave combines superior infrastructure performance with deep technical...  ...You'll Do CoreWeave is looking for an Engineering Manager to lead a team building and operating...  ...systems that power high-performance AI and ML workloads. You will lead engineers... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    28 days ago
  • $255k - $300k

     ...giving a growing number of engineers and employees an AI...  ...! As Engineering Manager, you will lead the...  ...in developing critical infrastructure for Agentic applications...  ...expertise in applied ML at scale. You'll partner...  ...include education, training, experience, location,... 
    Training
    Full time
    Work at office
    Flexible hours
    Shift work

    Robinhood

    Menlo Park, CA
    1 day ago
  •  ...Implements machine learning (ML) models for production...  ...workflows. Creates infrastructure and frameworks to...  ...Development Leads, Product Management, Operations, and...  ...with design criteria of trained models and/or systems....  ...Machine Learning, Computer Engineering, Mathematics, Physics,... 
    Training
    Shift work

    Oracle

    Santa Clara, CA
    3 days ago
  • $207k - $300k

     ...Engineering Manager, Applied AI, Cloud Support Intelligence Google Cloud's mission is...  ...combining cutting-edge technology, infrastructure, and talent. AI/ML software engineers in Cloud bridge...  ..., and relevant education or training. US: $207000 - $300000 (USD) + 20%... 
    Training
    Worldwide

    Google

    Sunnyvale, CA
    2 days ago
  • $148.32k - $203.94k

     ...visit: Job Summary We are seeking a hands-on Principal Infrastructure Hardware Engineer to architect, design, and deliver system platforms...  ...by factors such as location, experience, education, and training. In addition to base salary, this role is eligible for... 
    Training

    SiTime Corporation

    Santa Clara, CA
    28 days ago
  • $150k - $230k

     ...fulfill our mission: building the infrastructure layer for content intelligence....  ...looking for a hands-on Machine Learning Engineer to drive the post-training of our large language models, with...  ...scripts. Strong data engineering for ML. You can independently design data... 
    Training
    Full time
    Local area
    Work from home

    NewsBreak

    Mountain View, CA
    28 days ago
  •  ...worldwide to plan, track, and manage their work effectively...  ...Machine Learning Engineer to join our Search &...  ..., models, and infrastructure that accelerate AI feature...  ...mentorship to emerging ML engineers. Beyond these...  ...haveExperience fine-tuning, post-training, and optimizing state-... 
    Training
    Work at office
    Local area
    Worldwide

    Atlassian

    Mountain View, CA
    23 hours ago
  • $150k

     ...highly motivated, and focused on engineering excellence. This...  ...scenarios. We own the full training pipeline: massive data curation...  ...quality, and experimentation infrastructure to measure and improve performance...  ..., efficient code for AI/ML systems. Hands-on experience... 
    Training
    Temporary work

    SpaceXAI

    Palo Alto, CA
    28 days ago
  • $250k - $320k

     ...Staff Infrastructure Engineer We are partnered with a Stealth AI Lab (backed by top-tier investors and...  ...optimize the inference platform, GPU‑based training clusters, and data processing pipelines...  ...5+ years of experience in Software / ML Infrastructure Engineering. Deep... 
    Training
    Full time
    Immediate start

    Strativ Group

    Menlo Park, CA
    4 days ago
  • $198k - $326k

     ...responsible for scaling LinkedIn's AI model training, feature engineering and serving with hundreds of billions...  ...of user queries.Model Training Infrastructure: As an engineer on the AI Training...  ...agility (experiment with hundreds of new ML models per quarter using thousands of... 
    Training
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Sunnyvale, CA
    2 days ago
  • $160k - $225k

     ...world's best data and AI infrastructure platform so our...  ...improve their business. Training and customizing state‑...  ...Runtime (AIR) is our managed platform for large‑scale...  ...As a Senior Software Engineer for AI Runtime, you...  ...performance computing, or ML systems. ~... 
    Training
    Local area

    Databricks

    Mountain View, CA
    2 days ago
  •  ...Principal Machine Learning Systems Engineer (P60) to lead technical...  ...with reliable, high-performance infrastructure.Working at AtlassianAtlassians...  ...What You'll Doð Design and Build ML SystemsArchitect and implement scalable systems for training, fine-tuning, and serving... 
    Training
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    23 hours ago
  • $115k - $150k

     ...technologies. As a Senior Project Manager, you will lead complex,...  ...that bring together engineering, design, construction, and client...  ...deliver critical manufacturing infrastructure. Recognized as one of the...  ...day one, you'll receive the training and support needed to leverage... 
    Training
    For contractors
    Currently hiring
    Work at office
    Flexible hours

    SSOE, Inc.

    Santa Clara, CA
    18 hours ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering...  ..., networking, cloud infrastructure, Kubernetes, databases, capacity management, continuous delivery,...  ...Security, Networking, and AI/ML organizations to make...  ..., inference systems, training environments, or high-... 
    Training
    Full time

    NVIDIA

    Santa Clara, CA
    2 days ago
  •  ...advanced AI, data, and engineering capabilities. Our...  ...runtime, model serving, GPU infrastructure, distributed systems,...  ..., long-context management, conversation memory,...  ...user experience. 6. Training & Research Infrastructure...  ...systems architecture, ML infrastructure, distributed... 
    Training

    Accellor

    Mountain View, CA
    11 days ago
  • $145k

     ...dynamic workplace environments. As a CBRE Engineering Ops Manager, you will be embedded onsite at the...  ...of critical building systems and infrastructure. You will serve as a key partner to client...  ...' growth and development through training programs, certifications, and mentorship... 
    Training
    Contract work
    Work at office
    Local area
    Visa sponsorship
    Flexible hours

    CBRE

    Sunnyvale, CA
    5 days ago
  • $181.1k - $245k

     ...Photonics & PcbA Test Manager Aeva's mission is to bring the...  ...while managing a team of test engineers and technicians. The...  ...manufacturing and can drive test infrastructure, DFT, yield improvement, and...  ..., relevant education or training, and market conditions. These... 
    Training
    Contract work
    Overseas
    Flexible hours

    Aeva, Inc

    Mountain View, CA
    3 days ago
  • $112.7k - $169.1k

     ...world's leading game engine. Recommendation and ranking...  ...full pipeline from training data to deployed model...  ...large‑scale data and ML systems, whether through...  ..., or experimentation infrastructure. Experience...  ...website or directly to managers. Unity does not accept... 
    Training
    Internship
    Work at office
    Worldwide
    Relocation package
    Shift work

    Unity South APAC (SEA, ANZ, IND Subcont.)

    Mountain View, CA
    3 days ago
  • $118k - $160k

    BKF is a multi‑service infrastructure consulting firm providing civil engineering, construction management, environmental, planning, and surveying services across California...  ...with design teams and provide cross‑team training and technical support Prepare proposals including... 
    Training
    For subcontractor
    Work at office
    Local area
    Flexible hours

    BKF Engineers

    Redwood City, CA
    2 days ago
  • $250k - $280k

     ...For a small team, Afero engineers collectively do a little bit...  ...from cloud applications and infrastructure to mobile development on multiple...  ...at all levels of management. This is a unique opportunity...  ...experience, relevant education, training or certifications, and other... 
    Training
    Full time
    Work experience placement
    Casual work
    Work at office
    Immediate start
    3 days per week

    Afero

    Los Altos, CA
    a month ago
  •  ...seeking a senior Principal Engineer to join our Platform team and...  ...development, Kubernetes, cloud infrastructure, and production support. In...  ...including feature engineering, model training, experimentation, evaluation, or production ML systems. Experience with on-... 
    Training
    Summer work
    Work at office
    Flexible hours

    Shakudo

    Menlo Park, CA
    8 days ago
  • $150k - $230k

     ...researchers and veteran systems engineers who share a vision for...  ...increasingly complex, traditional infrastructure struggles to meet the demands...  ...-performance distributed GPU training. You'll work at the...  ...networking (RDMA, InfiniBand) ML framework or runtime internals... 
    Training

    Clockwork.io

    Palo Alto, CA
    12 days ago
  •  ...industry-leading total spend management platform for businesses...  ...The Impact of a Principal Engineer, AI/ML Architecture at Coupa: Coupa...  ..., designing how we train, evaluate, and serve models...  ...partnership evaluations with AI infrastructure and model providers. Architect... 
    Training

    Coupa Software, Inc.

    Foster, CA
    6 days ago
  • $160.36k - $240.54k

     ...autonomous driving technology. In an ML-first system, the overall...  ...and diversity of its training and evaluation data.  The team...  ...a scalable and reliable data infrastructure. This infrastructure is designed...  ...closely with system engineers to thoroughly validate the autonomous... 
    Training
    Work experience placement
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    28 days ago
  • $120k - $132k

     ...automation, and other critical infrastructure. This leader will...  ...developing a high-performing Engineering team while ensuring that the...  ...Capital Projects & Vendor Management Lead and coordinate capital...  ...Engineering team through training, cross-training, and hands-... 
    Training
    Temporary work
    For contractors

    Ensemble Hospitality

    Menlo Park, CA
    9 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Engineering Manager, ML Training Infrastructure. Be the first to apply!