Engineering Manager, ML Training Infrastructure
$140k - $260kWoven by Toyota
Woven by Toyota is enabling Toyota’s once-in-a-century transformation into a mobility company. Inspired by a legacy of innovating for the benefit of others, our mission is to challenge the current state of mobility through human-centric innovation — expanding what “mobility” means and how it serves society.
Our work centers on four pillars: AD/ADAS, our autonomous driving and advanced driver assist technologies; Arene, our software development platform for software-defined vehicles; Woven City, a test course for mobility; and Cloud & AI, the digital infrastructure powering our collaborative foundation. Business-critical functions empower these teams to execute, and together, we’re working toward one bold goal: a world with zero accidents and enhanced well-being for all.
About the Team
Enterprise AI is a platform that provides end-to-end machine learning tooling experience to support and accelerate machine learning development, including autonomous driving and other related projects. Our platform serves customers as a standardized machine learning platform within Woven by Toyota as the larger Toyota Group companies.
The Enterprise AI ML Training Infrastructure team builds and operates the infrastructure that enables engineers and researchers to train, evaluate, and iterate on machine learning models at scale. Our customers include teams working on AD/ADAS and Woven City, where large-scale machine learning workloads, simulation, and data processing are critical to developing the next generation of mobility technologies.
The team operates at the intersection of machine learning, distributed systems, cloud infrastructure, and developer platforms. Our systems include large-scale GPU compute environments, Kubernetes-based cluster infrastructure, and workflow and pipeline orchestration technologies such as Ray, Airflow, and Temporal.
We are looking for an Engineering Manager to lead a team to help shape the technical direction, engineering culture, and long-term evolution of our ML training infrastructure platform.
WHO ARE WE LOOKING FOR?
As the Engineering Manager, ML Training Infrastructure, you will lead the team responsible for building and operating the infrastructure that powers machine learning training workloads across Woven by Toyota.
You will combine engineering leadership with strong technical judgment. You will work closely with engineering, research, and product stakeholders to understand customer needs, translate them into scalable technical solutions, and ensure that our platform is reliable, efficient, secure, and easy to use.
This role is particularly suited to an engineering leader who enjoys working on complex infrastructure problems and is comfortable operating across Kubernetes, distributed systems, GPU infrastructure, cloud platforms, and machine learning workloads.
You will also play an important role in building a strong engineering culture, developing engineers, establishing effective engineering practices, and partnering with our customers to continuously improve the platform.
What You Will Enable
In this role, you will help build the infrastructure that enables teams across Woven by Toyota to train and operate increasingly large and sophisticated machine learning workloads.
Your work will directly support engineers and researchers working on AD/ADAS, Woven City, and other mobility initiatives by making compute resources easier to access, workloads more reliable, and ML development faster and more scalable.
You will have the opportunity to shape both the technical architecture of a critical ML infrastructure platform and the engineering culture of the team building it.
\n RESPONSIBILITIES-
Lead engineers and foster a collaborative and innovative environment.
-
Manage and nurture a group of around 7–8 engineers, offering technical vision, mentorship, professional growth, and effective project delivery.
-
Establish and articulate the technical vision and future roadmap for machine learning training infrastructure.
-
Direct the architecture, construction, and operation of resilient systems supporting massive ML model training workloads.
-
Construct and maintain Kubernetes-driven compute clusters and container ecosystems tailored for machine learning workloads.
-
Oversee and optimize GPU fleets, addressing cluster scheduling, resource utilization, system stability, and cost performance.
-
Advance distributed model training capabilities and workload management leveraging frameworks like Ray and adjacent tools.
-
Enhance workflow orchestration platforms by deploying and supporting solutions such as Airflow, Temporal, and similar pipeline frameworks.
-
Elevate the system resilience, monitoring insights, elasticity, and operational productivity of our ML training stack.
-
Work closely alongside AD/ADAS, Woven City, and internal platform partners to gather feature requests and ship infrastructure solutions that accelerate development.
-
Partner with cross-functional engineering groups to specify platform interfaces, APIs, system architectures, and operational standards.
-
Translate high-level organizational objectives and user requirements into concrete, executable engineering milestones with product leads.
-
Institute rigorous software practices across implementation, code reviews, automated testing, continuous deployment, and live site management.
-
Engage in production support workflows, lead post-incident reviews, and champion system refinements based on operational findings.
-
Balance immediate client feature requests against long-term architectural stability, technical hygiene, and system maintainability.
-
Foster an inclusive, highly collaborative, and excellence-driven engineering culture throughout Woven by Toyota.
-
8+ years of experience in software engineering, with at least 2 years in a leadership role.
-
Proven track record mentoring and managing high-performing software or platform teams.
-
Demonstrated expertise in architecting, deploying, and maintaining large-scale mission-critical infrastructure.
-
Deep technical mastery of Kubernetes orchestration alongside containerized runtime systems.
-
Hands-on exposure to cloud platform architectures and highly resilient distributed systems.
-
Background engineering platform foundation layers designed for machine learning model development or data pipeline intensive workloads.
-
Practical understanding of massive GPU cluster management, hardware scheduling algorithms, and capacity efficiency tuning.
-
Operational background overseeing high-availability environments, leading post-incident reviews, and advancing monitoring capabilities.
-
Excellent interpersonal abilities to align technical roadmaps across diverse engineering partners and leadership stakeholders.
-
Experience working with Japanese customers or clients.
-
Experience with Golang development (especially for cloud services)
-
Experience in the automotive industry or knowledge of automotive software development (e.g., ASPICE, V-model).
-
Experience working in environments with a focus on security and safety.
-
English/Japanese bilingual ability
The base pay for this position ranges from $140,000 - $260,000 a year.
Your base salary is one part of your total compensation. We offer a base salary, short term and long term incentives, and a comprehensive benefits package. The total compensation offered to an employee will be dependent upon the individual's skills, experience, qualifications, location, and level.WHAT WE OFFER
We are committed to creating a modern work environment that supports our employees and their loved ones. We offer many options of the best programs to allow you to do your most meaningful work and to help you shape the future of mobility.
・Excellent health, wellness, dental and vision coverage
・A rewarding 401k program
・Flexible vacation policy
・Family planning and care benefits
Our Commitment
・We are an equal opportunity employer and value diversity.
・Any information we receive from you will be used only in the hiring and onboarding process. Please see our privacy notice for more details.
- ...Head Of Engineering At Poesis Whoever builds the leading... ...portfolios, and manage risk. Our founders managed... ...) and led enterprise ML at Goldman Sachs and Amazon... ...live capital is the training ground for an... ...integrity. Designing the infrastructure that enables...TrainingFull timeWork at officeVisa sponsorshipWork visaRelocation package3 days per week
$193.93k - $352.29k
...What is not fungible is the infrastructure that decides whether an autonomous... ...inside Nuro's own engineering organization, under the same... ...that loop runs against the training pipelines behind the driving... ...helped. ~ Experience with ML training or research infrastructure...TrainingImmediate startFlexible hours$230k - $260k
...a Principal Machine Learning Engineer, you willoperateat the company... ...the design of large-scale ML systems and shared platforms... ...design ofscalable ML platforms(training, evaluation, inference, safety... ...generationDefine and evolveshared ML infrastructure and platformsfor training,...TrainingWork at officeImmediate start3 days per week$190k - $260k
...as the speed at which we can train it. Every improvement to our... ...scale world models – depends on infrastructure that turns thousands of hours... .... We are looking for engineers who make model training fast:... ...streamable formats Partner with ML teams to scale new architectures...TrainingTemporary workWork at officeVisa sponsorshipFlexible hours$255k - $300k
...giving a growing number of engineers and employees an AI... ...! As Engineering Manager, you will lead the... ...in developing critical infrastructure for Agentic applications... ...expertise in applied ML at scale. You'll partner... ...include education, training, experience, location,...TrainingFull timeWork at officeFlexible hoursShift work- ...Implements machine learning (ML) models for production... ...workflows. Creates infrastructure and frameworks to... ...Development Leads, Product Management, Operations, and... ...with design criteria of trained models and/or systems.... ...Machine Learning, Computer Engineering, Mathematics, Physics,...TrainingShift work
$207k - $300k
...Engineering Manager, Applied AI, Cloud Support Intelligence Google Cloud's mission is... ...combining cutting-edge technology, infrastructure, and talent. AI/ML software engineers in Cloud bridge... ..., and relevant education or training. US: $207000 - $300000 (USD) + 20%...TrainingWorldwide- ...worldwide to plan, track, and manage their work effectively... ...Machine Learning Engineer to join our Search &... ..., models, and infrastructure that accelerate AI feature... ...mentorship to emerging ML engineers. Beyond these... ...haveExperience fine-tuning, post-training, and optimizing state-...TrainingWork at officeLocal areaWorldwide
$250k - $320k
...Staff Infrastructure Engineer We are partnered with a Stealth AI Lab (backed by top-tier investors and... ...optimize the inference platform, GPU‑based training clusters, and data processing pipelines... ...5+ years of experience in Software / ML Infrastructure Engineering. Deep...TrainingFull timeImmediate start$160k - $225k
...world's best data and AI infrastructure platform so our... ...improve their business. Training and customizing state‑... ...Runtime (AIR) is our managed platform for large‑scale... ...As a Senior Software Engineer for AI Runtime, you... ...performance computing, or ML systems. ~...TrainingLocal area$115k - $150k
...technologies. As a Senior Project Manager, you will lead complex,... ...that bring together engineering, design, construction, and client... ...deliver critical manufacturing infrastructure. Recognized as one of the... ...day one, you'll receive the training and support needed to leverage...TrainingFor contractorsCurrently hiringWork at officeFlexible hours$182k - $242k
...enterprises, CoreWeave combines superior infrastructure performance with deep technical... ...You'll Do CoreWeave is looking for an Engineering Manager to lead a team building and operating... ...systems that power high-performance AI and ML workloads. You will lead engineers...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...Principal Machine Learning Systems Engineer (P60) to lead technical... ...with reliable, high-performance infrastructure.Working at AtlassianAtlassians... ...What You'll Doð Design and Build ML SystemsArchitect and implement scalable systems for training, fine-tuning, and serving...TrainingWork at officeLocal area
$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering... ..., networking, cloud infrastructure, Kubernetes, databases, capacity management, continuous delivery,... ...Security, Networking, and AI/ML organizations to make... ..., inference systems, training environments, or high-...TrainingFull time- ...advanced AI, data, and engineering capabilities. Our... ...runtime, model serving, GPU infrastructure, distributed systems,... ..., long-context management, conversation memory,... ...user experience. 6. Training & Research Infrastructure... ...systems architecture, ML infrastructure, distributed...Training
$145k
...dynamic workplace environments. As a CBRE Engineering Ops Manager, you will be embedded onsite at the... ...of critical building systems and infrastructure. You will serve as a key partner to client... ...' growth and development through training programs, certifications, and mentorship...TrainingContract workWork at officeLocal areaVisa sponsorshipFlexible hours$112.7k - $169.1k
...world's leading game engine. Recommendation and ranking... ...full pipeline from training data to deployed model... ...large‑scale data and ML systems, whether through... ..., or experimentation infrastructure. Experience... ...website or directly to managers. Unity does not accept...TrainingInternshipWork at officeWorldwideRelocation packageShift work$181.1k - $245k
...Photonics & PcbA Test Manager Aeva's mission is to bring the... ...while managing a team of test engineers and technicians. The... ...manufacturing and can drive test infrastructure, DFT, yield improvement, and... ..., relevant education or training, and market conditions. These...TrainingContract workOverseasFlexible hours$160.36k - $240.54k
...autonomous driving technology. In an ML-first system, the overall... ...and diversity of its training and evaluation data. The team... ...a scalable and reliable data infrastructure. This infrastructure is designed... ...closely with system engineers to thoroughly validate the autonomous...TrainingWork experience placement$150k - $230k
...researchers and veteran systems engineers who share a vision for... ...increasingly complex, traditional infrastructure struggles to meet the demands... ...-performance distributed GPU training. You'll work at the... ...networking (RDMA, InfiniBand) ML framework or runtime internals...Training- ...industry-leading total spend management platform for businesses... ...The Impact of a Principal Engineer, AI/ML Architecture at Coupa: Coupa... ..., designing how we train, evaluate, and serve models... ...partnership evaluations with AI infrastructure and model providers. Architect...Training
$120k - $132k
...automation, and other critical infrastructure. This leader will... ...developing a high-performing Engineering team while ensuring that the... ...Capital Projects & Vendor Management Lead and coordinate capital... ...Engineering team through training, cross-training, and hands-...TrainingTemporary workFor contractors$170k - $200k
...Description Job Description Senior Engineering Manager (LAND DEVELOPMENT) - Santa Clara, CA... ..., utility systems, and stormwater infrastructure in accordance with jurisdictional standards... ...and business planning. May lead training initiatives on various technical...TrainingContract workWork at office$240k - $290k
...Sonatus is looking for an experienced Senior Engineering Manager to build and lead our AI Validation... ...end-to-end validation strategy for AI/ML models, LLMs, RAG pipelines, and... ...workflows—from data pipelines and model training through cloud services and in-vehicle deployment...TrainingWork at officeWorldwideFlexible hoursShift work$236.7k - $309.03k
...experimentation, quality, and production systems. You’ll have the opportunity to work across a number of areas including model post training, personalization/recommendation, proactive and other agentic experiences. You’ll tackle ambiguous problems end to end: defining...TrainingWork at officeLocal area- ...team leading the equipment lifecycle management, service and engineering function for a variety of equipment... ...equipment management practices, infrastructure, and technical capabilities that will... ...systems. Maintain equipment training programs and ensure laboratory personnel...TrainingRelocationFlexible hours3 days per week
$150k - $224k
...Design and build data processing infrastructure for model training and feature serving, optimizing for performance... ...tooling for data infrastructure used across ML teams Required Qualifications Strong software engineering fundamentals, with experience building...TrainingFull time$148.32k - $203.94k
...visit: Job Summary We are seeking a hands-on Principal Infrastructure Hardware Engineer to architect, design, and deliver system platforms... ...by factors such as location, experience, education, and training. In addition to base salary, this role is eligible for...Training$139.9k - $274.8k
...apply for the Principal Design Engineer role at Microsoft .... ...Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) is the team... ...operations, globalization, and manageability solutions. Our focus is on... ...Intelligence (AI) / Machine Learning (ML) SoCs Working knowledge of...Work at officeLocal areaWorldwide$180k
...motivated, and focused on engineering excellence. This organization... ...ABOUT THE ROLE: As an ML Infrastructure Engineer, you will play a pivotal... ...compute infrastructure, training frameworks, and experimentation... ....g., Slurm), configuration management (Puppet/Ansible), or...TrainingTemporary workWork experience placement
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Engineering Manager, ML Training Infrastructure. Be the first to apply!
- infrastructure engineer Palo Alto, CA
- remote infrastructure engineer Palo Alto, CA
- senior infrastructure engineer Palo Alto, CA
- infrastructure developer Palo Alto, CA
- IT operations director Palo Alto, CA
- internship machine learning Palo Alto, CA
- machine learning Palo Alto, CA
- artificial intelligence - machine learning intern Palo Alto, CA
- machine learning scientist Palo Alto, CA
- machine learning research scientist Palo Alto, CA


