Tech Lead Manager, ML Accelerator Fleet Efficiency
$207k - $300kDrive technical strategy, roadmaps, and adoption for large-scale ML infrastructure development.Innovate next directions for infrastructure over a 12-month time horizon given a rapidly changing technology landscape.Exercise sound engineering judgment to guide sustainable engineering choices for ML systems at scale.Seek additional opportunities to drive efficiencies in ML workloads using scaling, idle suspend, and improving these capabilities with existing and novel technologies.Lead a team of ~10 engineers to develop solutions that drive the efficiency of ML workloads.Minimum qualifications:Bachelor’s degree or equivalent practical experience.8 years of experience in software development.5 years of experience with one or more of the following: Speech/audio (e.g., technology duplicating and responding to the human voice), reinforcement learning (e.g., sequential decision making), ML infrastructure, or specialization in another ML field.5 years of experience leading ML design and optimizing ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).3 years of experience in a technical leadership role.2 years of experience in a people management or team leadership role.Preferred qualifications:Master’s degree or PhD in Engineering, Computer Science, or a related technical field.3 years of experience working in a complex, matrixed organization involving cross-functional, or cross-business projects.Experience with ML development, modeling, optimization, and infrastructure.Experience with TPUs, TPU system design, and GPUs.Experience with low-level programming.Expertise in ML compilers and runtimes.Like Google's own ambitions, the work of a Software Engineer goes beyond just Search. Software Engineering Managers have not only the technical expertise to take on and provide technical leadership to major projects, but also manage a team of Engineers. You not only optimize your own code but make sure Engineers are able to optimize theirs. As a Software Engineering Manager you manage your project goals, contribute to product strategy and help develop your team. Teams work all across the company, in areas such as information retrieval, artificial intelligence, natural language processing, distributed computing, large-scale system design, networking, security, data compression, user interface design; the list goes on and is growing every day. Operating with scale and speed, our exceptional software engineers are just getting started -- and as a manager, you guide the way.With technical and leadership expertise, you manage engineers across multiple teams and locations, a large product budget and oversee the deployment of large-scale projects across multiple sites internationally.Our team drives machine learning computational efficiency. We manage software to optimize the utilization of hundreds of thousands of Google Accelerator Units globally. We build the software abstraction layer between ML models and physical TPU/GPU hardware.Our mission is to eliminate resource waste across Google’s accelerator fleet, maximizing the physical utility of compute clusters while maintaining peak developer velocity and seamless runtime execution. We strive to provide a cohesive, highly efficient, and transparent runtime environment that enables ML teams to focus entirely on modeling and research rather than physical infrastructure constraints.Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.Individual pay is determined by factors including job-related skills, experience, and relevant education or training. US: $207000 - $300000 (USD) + 20% bonus target + equity + benefitsLearn more about benefits at Google.Bachelor’s degree or equivalent practical experience.8 years of experience in software development.5 years of experience with one or more of the following: Speech/audio (e.g., technology duplicating and responding to the human voice), reinforcement learning (e.g., sequential decision making), ML infrastructure, or specialization in another ML field.5 years of experience leading ML design and optimizing ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).3 years of experience in a technical leadership role.2 years of experience in a people management or team leadership role.
$212k - $340k
...will bring a safer, more efficient, and more accessible... ...Aurora, visit aurora.tech or follow us on LinkedIn... ...Senior Staff Tech Lead Manager to join the Aurora Services... ...Aurora's suite of fleet management tools.The Aurora... ...coding workflows to accelerate developer velocity as...FleetWork at officeLocal areaRemote work3 days per week$152k - $241.5k
We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI... ...make their ML models more efficient leading to significant productivity improvements... ...infrastructure Proactively monitor fleet wide utilization patterns, analyze existing...FleetFull timeRemote work$207k - $300k
...on Google Cloud TPUs. Manage up to 6 engineers to drive... .... Collaborate with the ML research, ML... ....5 years of experience leading ML design and optimizing... ...sites internationally.As a Tech Lead Manager, you will... ...GPU Kernels.Google Cloud accelerates every organization’s ability...Suggested$262k - $364k
...personalized search capabilities.Lead a small team that works on... ...leading technical project strategy, ML design, and optimizing... ...technical expertise you will manage project priorities, deadlines,... ...of user context and ultimately efficiency of applying that context in producing...Suggested$207k - $300k
On-board emerging co-accelerators into Google's ML accelerator families to enable... ...improved performance and efficiency.Collaborate with internal... ...challenges related to host and management software stacks.Provide... ...technical leadership role leading project teams and setting...SuggestedWorldwide$189k - $274k
...will bring a safer, more efficient, and more accessible... ...from Aurora, visit aurora.tech or follow us on... ...focusing on Deep Learning Acceleration at Aurora, you will play... ...experience in optimizing DL/ML workloads at the... ...empathy and our ability to lead effectively. As a result...Work at officeLocal area3 days per week$192k - $278k
Manage a power design team responsible for power architecture... ...and system power integration.Lead an experienced team, define... ...improve the power efficiency of the TPU designs.Contribute... ...to shape the future of AI/ML hardware acceleration. You will have an opportunity...Worldwide$235.03k - $352.29k
...from robotaxis and logistics fleets to personal vehicles.With years... ...Fidelity, T. Rowe Price, and other leading investorsAbout the RoleNuro... ...driving technology. In an ML-first system, the overall system... ...suite. Our tools must handle the efficient processing and annotation of...FleetImmediate startFlexible hours- ...from robotaxis and logistics fleets to personal vehicles.With years... ...Fidelity, T. Rowe Price, and other leading investorsAbout the... ...stack spans both heuristic and ML-based approaches, covering everything... ...intuition for where they accelerate engineering work and where they...FleetTemporary workImmediate startFlexible hours
$262k - $364k
...above requirements and lead a multi-year technical... ...industry leading utilization efficiency.Drive the roadmap and partner with Product Management on the definition and... ...for deploying Google’s ML hardware. You will help... ...of Google’s Fleet planning and optimization...FleetTemporary workWorldwide$207k - $300k
Lead the software architecture, design, development and testing of Arm/... ...for Servers that's in the Google fleet. This is our compute servers that... ...Colossus, our Machine Learning (ML) headnodes to powers Gemini.Google Cloud accelerates every organization’s ability to digitally...Fleet$165.2k - $223.6k
...real workloads on your platform efficiently- Develop and improve the virtual... ...significant pieces of the stackNo ML background needed. You'll learn the ML accelerator domain on the job.This role can... ..., code reviews, source control management, build processes, testing, and operations...Local areaFlexible hours$207k - $300k
...application, model, and distributed fleet infrastructure layers to... ...workloads using standard ML profilers (e.g., PyTorch profiler... ....Familiarity with GPU/TPU/accelerator performance concepts (e.g. memory... ...faster, cheaper, and more efficient. You will analyze the entire...Fleet$138k - $197k
...infrastructure and architecture for on-device AI/ML accelerators.Implement, model, analyze, and... ...development, testing, and simulation.Lead the end-to-end delivery of sophisticated... ..., delivering unparalleled performance, efficiency, and integration.In this role, you will...Worldwide- ...deliver industry-leading training and inference... ...that combine GPU-accelerated prefill with ultra... ...a new accelerator fleet, and drive... ...latency, and capacity efficiency. This is a hands-on... ...checking, capacity-management, and failure-recovery... ...open-source ML systems project.Experience...Fleet
$160.36k - $240.54k
...from robotaxis and logistics fleets to personal vehicles.With... ...T. Rowe Price, and other leading investorsAbout the... ...data processing to join our ML Infrastructure team. In this... ...workload scheduling, and efficient feature management to accelerate the Nuro Driver development...FleetImmediate startFlexible hours$207k - $300k
Identify and maintain ML training and serving benchmarks... ...models) to train efficiently on a very large-scale... ...solutions at Google fleet-wide scale.Minimum... ...models to exploit ML accelerator architecture strengths... ..., as well as industry leading open-source models, to...Fleet$262k - $364k
...for complex, long-running loops.Lead the development of low-latency... ...tool selection, OAuth/identity managing, and API integration for short-... ...team of Software Engineers and Tech Leads, driving resource allocation... ...technical project strategy, ML design, and optimizing industry...$175k - $287k
...the largest privately managed compute infrastructures... ...petabytes daily, and the AI/ML infrastructure driving... ...reliability and efficiency, GPU compute optimization for AI/ML, and fleet health at scale — all at... ...tools and workflows to accelerate engineering productivitySuggested...FleetFor contractorsWork experience placementWork at officeFlexible hours$207k - $300k
...models.Translate AI/ML research into... ...-scale TPU/GPU fleet utilization.... ...of experience leading ML design and optimizing... ...and enhance efficiency.Demonstrated... ...you will manage project priorities... ...Research, and Cloud) accelerates scientific... ...bringing AI to tech, healthcare, finance...Fleet$272k - $431.25k
...outstanding opportunity to lead and build the... ...best platform for ML/AI infrastructure?... ...DSX Kubernetes Fleet team within NVIDIA... ..., lifecycle management, and deployment safety... ...to our massive GPU accelerated container platform... ...native applications.Efficiently multitask across...FleetFull timeWork experience placement$122.6k - $185k
...inference? Want to do industry leading work delivering... ...and scalability in AI/ML and HPC workloads.You are... ...Engineers, TPMs, Managers, Principals) and groups... ...future/new designs for AWS Accelerated server solutions for AWS... ...launching hardware in the fleet. Located out of Seattle...FleetLocal areaFlexible hours$193.3k - $261.5k
...these custom-designed accelerator SoCs for use by AWS internal... ....As part of the ML accelerator modeling team... ...modeling of the ML and management regions of our chips, and... ...- 5+ years of leading design or architecture... ...Experience as a mentor, tech lead or leading an engineering...InternshipLocal areaFlexible hours$174k - $252k
Direct full-stack Software (SW) role, focusing on ML compiler and inference infrastructure. Design and develop ML compiler stack... ...Python and C++.Experience with machine learning (ML) hardware accelerators, ML compiler stacks (e.g., JAX, PyTorch, XLA), and kernel development...$174k - $252k
...equivalent practical experience.Experience in software development using Python and C++.Experience with Machine Learning (ML) hardware accelerators and ML inference software.Preferred qualifications:PhD in Computer Engineering, Computer Science, or a related field.2 years...$147k - $210k
Improve the training and serving efficiency of Gemini models on hardware accelerators (TPUs and GPUs), spanning model configurations, execution runtimes, and dedicated compiler passes.Profile large-scale distributed workloads to diagnose compute, memory, and communication...- ...and trust our employees to manage their schedules responsibly.... ...learning workloads fast and cost-efficient in the datacenter. This role... ...the gap between what our fleet of accelerators is theoretically capable of... ...intersection of accelerators, ML frameworks, and large-scale...FleetFull timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift
$168.4k - $193.5k
...code reviews, source control management, build processes, testing,... ...seeking experience as a mentor, tech lead, or engineering team lead.... ...with machine learning accelerator hardware and/or software.... ...engineers can run real workloads efficiently. We work directly with software...Full timeInternship$142.8k - $274.8k
..., globalization, and manageability solutions. Our focus... ...on smart growth, high efficiency, and delivering a trusted... ...to enable industry-leading AI training and... ...seeking a Principal AI Accelerator Tools Development Engineer... ...validation, and fleet readiness for both current...FleetOngoing contractWork at officeLocal areaWorldwide3 days per week$182k - $242k
...confidence. Trusted by leading AI labs, startups,... ...expertise to accelerate breakthroughs and... ...time, whether a GPU fleet, a fabric, or a... ...GPU utilization and efficiency, interconnect (NVLink... ..., product managers, and executives, but... ...infrastructure or ML systems rather than...FleetPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Tech Lead Manager, ML Accelerator Fleet Efficiency. Be the first to apply!
- technical leader Sunnyvale, CA
- technical lead Sunnyvale, CA
- certification manager Sunnyvale, CA
- senior manager tax Sunnyvale, CA
- ranch manager Sunnyvale, CA
- lean manager Sunnyvale, CA
- employment manager Sunnyvale, CA
- e-learning manager Sunnyvale, CA
- refrigeration manager Sunnyvale, CA
- broadcast manager Sunnyvale, CA


