Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Tech Lead Manager, ML Accelerator Fleet Efficiency

$207k - $300k

Google

Drive technical strategy, roadmaps, and adoption for large-scale ML infrastructure development.Innovate next directions for infrastructure over a 12-month time horizon given a rapidly changing technology landscape.Exercise sound engineering judgment to guide sustainable engineering choices for ML systems at scale.Seek additional opportunities to drive efficiencies in ML workloads using scaling, idle suspend, and improving these capabilities with existing and novel technologies.Lead a team of ~10 engineers to develop solutions that drive the efficiency of ML workloads.Minimum qualifications:Bachelor’s degree or equivalent practical experience.8 years of experience in software development.5 years of experience with one or more of the following: Speech/audio (e.g., technology duplicating and responding to the human voice), reinforcement learning (e.g., sequential decision making), ML infrastructure, or specialization in another ML field.5 years of experience leading ML design and optimizing ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).3 years of experience in a technical leadership role.2 years of experience in a people management or team leadership role.Preferred qualifications:Master’s degree or PhD in Engineering, Computer Science, or a related technical field.3 years of experience working in a complex, matrixed organization involving cross-functional, or cross-business projects.Experience with ML development, modeling, optimization, and infrastructure.Experience with TPUs, TPU system design, and GPUs.Experience with low-level programming.Expertise in ML compilers and runtimes.Like Google's own ambitions, the work of a Software Engineer goes beyond just Search. Software Engineering Managers have not only the technical expertise to take on and provide technical leadership to major projects, but also manage a team of Engineers. You not only optimize your own code but make sure Engineers are able to optimize theirs. As a Software Engineering Manager you manage your project goals, contribute to product strategy and help develop your team. Teams work all across the company, in areas such as information retrieval, artificial intelligence, natural language processing, distributed computing, large-scale system design, networking, security, data compression, user interface design; the list goes on and is growing every day. Operating with scale and speed, our exceptional software engineers are just getting started -- and as a manager, you guide the way.With technical and leadership expertise, you manage engineers across multiple teams and locations, a large product budget and oversee the deployment of large-scale projects across multiple sites internationally.Our team drives machine learning computational efficiency. We manage software to optimize the utilization of hundreds of thousands of Google Accelerator Units globally. We build the software abstraction layer between ML models and physical TPU/GPU hardware.Our mission is to eliminate resource waste across Google’s accelerator fleet, maximizing the physical utility of compute clusters while maintaining peak developer velocity and seamless runtime execution. We strive to provide a cohesive, highly efficient, and transparent runtime environment that enables ML teams to focus entirely on modeling and research rather than physical infrastructure constraints.Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.Individual pay is determined by factors including job-related skills, experience, and relevant education or training. US: $207000 - $300000 (USD) + 20% bonus target + equity + benefitsLearn more about benefits at Google.Bachelor’s degree or equivalent practical experience.8 years of experience in software development.5 years of experience with one or more of the following: Speech/audio (e.g., technology duplicating and responding to the human voice), reinforcement learning (e.g., sequential decision making), ML infrastructure, or specialization in another ML field.5 years of experience leading ML design and optimizing ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).3 years of experience in a technical leadership role.2 years of experience in a people management or team leadership role.

Vacancy posted 17 hours ago
Similar jobs that could be interesting for youBased on the Tech Lead Manager, ML Accelerator Fleet Efficiency in Sunnyvale, CA vacancy
  • $212k - $340k

     ...will bring a safer, more efficient, and more accessible...  ...Aurora, visit aurora.tech or follow us on LinkedIn...  ...Senior Staff Tech Lead Manager to join the Aurora Services...  ...Aurora's suite of fleet management tools.The Aurora...  ...coding workflows to accelerate developer velocity as... 
    Fleet
    Work at office
    Local area
    Remote work
    3 days per week

    Aurora Innovation

    Mountain View, CA
    2 days ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI...  ...make their ML models more efficient leading to significant productivity improvements...  ...infrastructure Proactively monitor fleet wide utilization patterns, analyze existing... 
    Fleet
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $207k - $300k

     ...on Google Cloud TPUs. Manage up to 6 engineers to drive...  .... Collaborate with the ML research, ML...  ....5 years of experience leading ML design and optimizing...  ...sites internationally.As a Tech Lead Manager, you will...  ...GPU Kernels.Google Cloud accelerates every organization’s ability... 
    Suggested

    Google

    Sunnyvale, CA
    3 days ago
  • $262k - $364k

     ...personalized search capabilities.Lead a small team that works on...  ...leading technical project strategy, ML design, and optimizing...  ...technical expertise you will manage project priorities, deadlines,...  ...of user context and ultimately efficiency of applying that context in producing... 
    Suggested

    Google

    Mountain View, CA
    2 days ago
  • $207k - $300k

    On-board emerging co-accelerators into Google's ML accelerator families to enable...  ...improved performance and efficiency.Collaborate with internal...  ...challenges related to host and management software stacks.Provide...  ...technical leadership role leading project teams and setting... 
    Suggested
    Worldwide

    Google

    Sunnyvale, CA
    18 hours ago
  • $189k - $274k

     ...will bring a safer, more efficient, and more accessible...  ...from Aurora, visit aurora.tech or follow us on...  ...focusing on Deep Learning Acceleration at Aurora, you will play...  ...experience in optimizing DL/ML workloads at the...  ...empathy and our ability to lead effectively. As a result... 
    Work at office
    Local area
    3 days per week

    Aurora Innovation

    Mountain View, CA
    4 days ago
  • $192k - $278k

    Manage a power design team responsible for power architecture...  ...and system power integration.Lead an experienced team, define...  ...improve the power efficiency of the TPU designs.Contribute...  ...to shape the future of AI/ML hardware acceleration. You will have an opportunity... 
    Worldwide

    Google

    Sunnyvale, CA
    2 days ago
  • $235.03k - $352.29k

     ...from robotaxis and logistics fleets to personal vehicles.With years...  ...Fidelity, T. Rowe Price, and other leading investorsAbout the RoleNuro...  ...driving technology. In an ML-first system, the overall system...  ...suite. Our tools must handle the efficient processing and annotation of... 
    Fleet
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    1 day ago
  •  ...from robotaxis and logistics fleets to personal vehicles.With years...  ...Fidelity, T. Rowe Price, and other leading investorsAbout the...  ...stack spans both heuristic and ML-based approaches, covering everything...  ...intuition for where they accelerate engineering work and where they... 
    Fleet
    Temporary work
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    3 days ago
  • $262k - $364k

     ...above requirements and lead a multi-year technical...  ...industry leading utilization efficiency.Drive the roadmap and partner with Product Management on the definition and...  ...for deploying Google’s ML hardware. You will help...  ...of Google’s Fleet planning and optimization... 
    Fleet
    Temporary work
    Worldwide

    Google

    Sunnyvale, CA
    1 day ago
  • $207k - $300k

    Lead the software architecture, design, development and testing of Arm/...  ...for Servers that's in the Google fleet. This is our compute servers that...  ...Colossus, our Machine Learning (ML) headnodes to powers Gemini.Google Cloud accelerates every organization’s ability to digitally... 
    Fleet

    Google

    Sunnyvale, CA
    4 days ago
  • $165.2k - $223.6k

     ...real workloads on your platform efficiently- Develop and improve the virtual...  ...significant pieces of the stackNo ML background needed. You'll learn the ML accelerator domain on the job.This role can...  ..., code reviews, source control management, build processes, testing, and operations... 
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    18 hours ago
  • $207k - $300k

     ...application, model, and distributed fleet infrastructure layers to...  ...workloads using standard ML profilers (e.g., PyTorch profiler...  ....Familiarity with GPU/TPU/accelerator performance concepts (e.g. memory...  ...faster, cheaper, and more efficient. You will analyze the entire... 
    Fleet

    Google

    Mountain View, CA
    4 days ago
  • $138k - $197k

     ...infrastructure and architecture for on-device AI/ML accelerators.Implement, model, analyze, and...  ...development, testing, and simulation.Lead the end-to-end delivery of sophisticated...  ..., delivering unparalleled performance, efficiency, and integration.In this role, you will... 
    Worldwide

    Google

    Mountain View, CA
    2 days ago
  •  ...deliver industry-leading training and inference...  ...that combine GPU-accelerated prefill with ultra...  ...a new accelerator fleet, and drive...  ...latency, and capacity efficiency. This is a hands-on...  ...checking, capacity-management, and failure-recovery...  ...open-source ML systems project.Experience... 
    Fleet

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $160.36k - $240.54k

     ...from robotaxis and logistics fleets to personal vehicles.With...  ...T. Rowe Price, and other leading investorsAbout the...  ...data processing to join our ML Infrastructure team. In this...  ...workload scheduling, and efficient feature management to accelerate the Nuro Driver development... 
    Fleet
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    4 days ago
  • $207k - $300k

    Identify and maintain ML training and serving benchmarks...  ...models) to train efficiently on a very large-scale...  ...solutions at Google fleet-wide scale.Minimum...  ...models to exploit ML accelerator architecture strengths...  ..., as well as industry leading open-source models, to... 
    Fleet

    Google

    Sunnyvale, CA
    4 days ago
  • $262k - $364k

     ...for complex, long-running loops.Lead the development of low-latency...  ...tool selection, OAuth/identity managing, and API integration for short-...  ...team of Software Engineers and Tech Leads, driving resource allocation...  ...technical project strategy, ML design, and optimizing industry... 

    Google

    Mountain View, CA
    2 days ago
  • $175k - $287k

     ...the largest privately managed compute infrastructures...  ...petabytes daily, and the AI/ML infrastructure driving...  ...reliability and efficiency, GPU compute optimization for AI/ML, and fleet health at scale — all at...  ...tools and workflows to accelerate engineering productivitySuggested... 
    Fleet
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    3 days ago
  • $207k - $300k

     ...models.Translate AI/ML research into...  ...-scale TPU/GPU fleet utilization....  ...of experience leading ML design and optimizing...  ...and enhance efficiency.Demonstrated...  ...you will manage project priorities...  ...Research, and Cloud) accelerates scientific...  ...bringing AI to tech, healthcare, finance... 
    Fleet

    Google

    Sunnyvale, CA
    1 day ago
  • $272k - $431.25k

     ...outstanding opportunity to lead and build the...  ...best platform for ML/AI infrastructure?...  ...DSX Kubernetes Fleet team within NVIDIA...  ..., lifecycle management, and deployment safety...  ...to our massive GPU accelerated container platform...  ...native applications.Efficiently multitask across... 
    Fleet
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    2 days ago
  • $122.6k - $185k

     ...inference? Want to do industry leading work delivering...  ...and scalability in AI/ML and HPC workloads.You are...  ...Engineers, TPMs, Managers, Principals) and groups...  ...future/new designs for AWS Accelerated server solutions for AWS...  ...launching hardware in the fleet. Located out of Seattle... 
    Fleet
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $193.3k - $261.5k

     ...these custom-designed accelerator SoCs for use by AWS internal...  ....As part of the ML accelerator modeling team...  ...modeling of the ML and management regions of our chips, and...  ...- 5+ years of leading design or architecture...  ...Experience as a mentor, tech lead or leading an engineering... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $174k - $252k

    Direct full-stack Software (SW) role, focusing on ML compiler and inference infrastructure. Design and develop ML compiler stack...  ...Python and C++.Experience with machine learning (ML) hardware accelerators, ML compiler stacks (e.g., JAX, PyTorch, XLA), and kernel development... 

    Google

    Mountain View, CA
    2 days ago
  • $174k - $252k

     ...equivalent practical experience.Experience in software development using Python and C++.Experience with Machine Learning (ML) hardware accelerators and ML inference software.Preferred qualifications:PhD in Computer Engineering, Computer Science, or a related field.2 years... 

    Google

    Mountain View, CA
    2 days ago
  • $147k - $210k

    Improve the training and serving efficiency of Gemini models on hardware accelerators (TPUs and GPUs), spanning model configurations, execution runtimes, and dedicated compiler passes.Profile large-scale distributed workloads to diagnose compute, memory, and communication... 

    Google

    Mountain View, CA
    2 days ago
  •  ...and trust our employees to manage their schedules responsibly....  ...learning workloads fast and cost-efficient in the datacenter. This role...  ...the gap between what our fleet of accelerators is theoretically capable of...  ...intersection of accelerators, ML frameworks, and large-scale... 
    Fleet
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Decisive Point

    Sunnyvale, CA
    2 days ago
  • $168.4k - $193.5k

     ...code reviews, source control management, build processes, testing,...  ...seeking experience as a mentor, tech lead, or engineering team lead....  ...with machine learning accelerator hardware and/or software....  ...engineers can run real workloads efficiently. We work directly with software... 
    Full time
    Internship

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    4 days ago
  • $142.8k - $274.8k

     ..., globalization, and manageability solutions. Our focus...  ...on smart growth, high efficiency, and delivering a trusted...  ...to enable industry-leading AI training and...  ...seeking a Principal AI Accelerator Tools Development Engineer...  ...validation, and fleet readiness for both current... 
    Fleet
    Ongoing contract
    Work at office
    Local area
    Worldwide
    3 days per week

    Microsoft

    Mountain View, CA
    3 days ago
  • $182k - $242k

     ...confidence. Trusted by leading AI labs, startups,...  ...expertise to accelerate breakthroughs and...  ...time, whether a GPU fleet, a fabric, or a...  ...GPU utilization and efficiency, interconnect (NVLink...  ..., product managers, and executives, but...  ...infrastructure or ML systems rather than... 
    Fleet
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    11 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Tech Lead Manager, ML Accelerator Fleet Efficiency. Be the first to apply!