Tech Lead, TPU AI Infrastructure
$207k - $300kDefine the technical roadmap and architecture for inventory and fleet management, focusing on Temporal orchestration, automated node re-bootstrapping, and self-service remediation.Own architectural and design decisions for the platform, collaborating with leadership and cross-functional teams to prioritize efforts, resolve roadblocks, and manage technical debt.Establish and maintain machine health Service Level Objectives (SLOs), driving continuous improvements in fleet-wide observability and environmental monitoring.Foster a culture of engineering excellence by establishing shared responsibility models with tenant teams and enforcing robust system-level access guardrails.Coach and mentor engineers on the team, guiding their technical development and helping them grow their impact.Minimum qualifications:Bachelor's degree or equivalent practical experience.8 years of experience with software development in one or more programming languages (e.g., Python, C, C++, Java, JavaScript).5 years of experience building and developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage, or hardware architecture. 5 years of experience testing, and launching software products.3 years of experience with software design and architecture.Preferred qualifications:Experience with Temporal orchestration workflows or similar workflow engines.Expertise in Kubernetes, ArgoCD, and container orchestration.Familiarity with hardware/software integration, hardware lifecycle, or bare-metal provisioning (PXE).Excellent communication skills with the ability to manage stakeholder relationships and drive cross-functional alignment.Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.Our team is building the future of at-scale computing by delivering a groundbreaking, on-premises AI supercomputer—the next evolution of Google's AI accelerator hardware. This is a rare opportunity to join a specialized team creating novel solutions for customers with sophisticated security and data locality requirements. You will be at the heart of building the next generation of AI hardware and the critical software that brings it to life, making a Google-scale impact on the future of AI.We manage complex infrastructure, integrating with Kubernetes, Temporal, and custom orchestration tools to ensure our compute nodes meet rigorous health standards. We are evolving our infrastructure toward automated node repair, self-healing frameworks, and advanced fleet-wide observability to reduce manual operational overhead and build robust platform boundaries.Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.Individual pay is determined by factors including job-related skills, experience, and relevant education or training. US: $207000 - $300000 (USD) + 20% bonus target + equity + benefitsLearn more about benefits at Google.Bachelor's degree or equivalent practical experience.8 years of experience with software development in one or more programming languages (e.g., Python, C, C++, Java, JavaScript).5 years of experience building and developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage, or hardware architecture. 5 years of experience testing, and launching software products.3 years of experience with software design and architecture.
$207k - $300k
Lead the technical strategy and architecture of next-generation control planes and software infrastructure for highly-available ML network fabrics.Design and implement agentic SDLC, leveraging... ...and hardware across the full stack.The AI and Infrastructure team is redefining...SuggestedWorldwide$207k - $300k
...across single-host and multi-host TPU.Design and implement key... ...sequential decision making), ML infrastructure, or specialization in another... ...field.5 years of experience leading ML design and optimizing ML... ...multiple sites internationally.As a Tech Lead Manager, you will use...Suggested$207k - $300k
Lead and manage the team responsible for TPU development and Silicon Validation by setting clear priorities, providing coaching and performance feedback... ...our cutting-edge chip and hardware systems.The AI and Infrastructure team is redefining what’s possible. We empower...SuggestedWorldwide$147k - $210k
...-silicon, root-cause performance analysis, and TPU mapping optimization.Drive full-stack hardware-... ...emerging software abstractions, such as Compound AI and multi-step agentic systems, execute across Google’s AI infrastructure. You will collaborate with various product area...Suggested$207k - $300k
...optimizing Generative AI performance across heterogeneous... ...decision making), ML infrastructure, or specialization in... ....Track record of leading and delivering ML projects... ...(e.g., GPU / Pixel TPU / NPUs / CPU) on Android... ...mission of bringing AI to tech, healthcare, finance,...SuggestedShift work$174k - $252k
....3 years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with... ...infrastructure. Contribute to projects enabling AI networking and high-performance fabrics for TPU and GPU workloads.Our Platforms Infrastructure Engineering...$163k - $236k
Lead advanced SoC integration projects from concept to final delivery... ...work to shape the future of AI/ML hardware acceleration. You... ...have an opportunity to drive TPU (Tensor Processing Unit) technology... ...project schedules.The AI and Infrastructure team is redefining what’s...Worldwide$116k - $165k
...for digital designs within the TPU.Write performant, and power-... ...ll work to shape the future of AI/ML hardware acceleration. You... ...collaborative environment.The AI and Infrastructure team is redefining what’s... ...shaping the future of world-leading hyperscale computing, with key...Worldwide$138k - $197k
...ll work to shape the future of AI/ML hardware acceleration. You... ...opportunity to drive cutting-edge TPU (Tensor Processing Unit)... ...testing lifecycle.The AI and Infrastructure team is redefining what’s possible... ...shaping the future of world-leading hyperscale computing, with key...Worldwide$147k - $210k
...sources of issues and the impact on hardware, network, or service operations and quality.Implement GenAI solutions, utilize ML infrastructure, and contribute to data preparation, optimization, and performance enhancements.Minimum qualifications:Bachelor’s degree or equivalent...$174k - $252k
...compiler parallelization features and optimization techniques for TPU backend necessary for large-scale workloads.Contribute to... ...including Python and C++.3 years of experience with Machine Learning infrastructure, ML execution frameworks (e.g., TensorFlow, JAX, PyTorch), or...$147k - $210k
...system development code.Participate in, or lead design reviews with peers and... ...the architecture built by the Technical Infrastructure team to keep it running. From developing... ...best and fastest experience possible.The AI and Infrastructure team is redefining what...Worldwide$207k - $300k
...new ML models and products on Google new TPU hardware, enabling larger models (giant... ...(e.g., sequential decision making), ML infrastructure, or specialization in another ML field.5... ...will work on Gemini, as well as industry leading open-source models, to understand model...$207k - $300k
...experience building and developing large-scale infrastructure, distributed systems or networks, or... ...in a technical leadership role leading project teams and setting technical direction... ...continue to push technology forward.The AI and Infrastructure team is redefining what...Worldwide$207k - $300k
...experience building and developing large-scale infrastructure, distributed systems or networks, or... ...in a technical leadership role leading project teams and setting technical direction... ...make every business successful through AI by combining cutting-edge technology, infrastructure...$147k - $210k
...stack as we continue to push technology forward.The Secure Network Infrastructure (SecNI) team aims to secure Google's network infrastructure... ...be responsible for securing various network devices, including AI/ML and optical fabrics, and are working on projects such as secure...Remote work$207k - $300k
...roadmaps, and adoption for large-scale ML infrastructure development.Innovate next directions for... ...capabilities with existing and novel technologies.Lead a team of ~10 engineers to develop... ...and infrastructure.Experience with TPUs, TPU system design, and GPUs.Experience with...$147k - $210k
Participate in, or lead design reviews with peers and stakeholders to decide amongst available... ...experience with developing large-scale infrastructure, distributed systems or networks, or... ...platforms, and Google Cloud infrastructure.The AI and Infrastructure team is redefining...Worldwide$207k - $300k
...product ordering and platform engineering.Lead complex, cross-functional projects from... ...experience building and developing large-scale infrastructure, distributed systems, or large-scale... ...network infrastructure.As we scale into the AI era, demand financial analysis (FA) is...Worldwide$207k - $300k
...reductions across organizational boundaries.Lead the development of software that improves... ....Experience integrating generative AI tools or LLM interfaces into workflows.Preferred... ...to push technology forward.The AI and Infrastructure team is redefining what’s possible. We...Worldwide$174k - $252k
...system development code.Participate in, or lead design reviews with peers and... ...years of experience developing large-scale infrastructure, distributed systems or networks, or experience... ...maintain, and enhance software solutions.The AI and Infrastructure team is redefining...Worldwide$240k - $333k
Lead, mentor and manage a team of RTL Design and DV Engineers developing DRAM subsystems... ...integration.As Technical Manager for TPU DRAM, you will develop and validate... ...manufacturers and third party IP providers.The AI and Infrastructure team is redefining what’s possible. We...Worldwide$147k - $210k
...system development code.Participate in, or lead design reviews with peers and... ...of experience with developing large-scale infrastructure, distributed systems or networks, or experience... ...maintain, and enhance software solutions. The AI and Infrastructure team is redefining...Worldwide$262k - $364k
...development.7 years of experience building and developing large-scale infrastructure, distributed systems or networks, or experience with compute... ....5 years of experience in a technical leadership role leading project teams and setting technical direction.3 years of...$147k - $210k
...system development code.Participate in, or lead design reviews with peers and... ...of experience with developing large-scale infrastructure, distributed systems or networks, or experience... ...platforms or frameworks.Experience with AI Algorithms.Preferred qualifications:Master...$207k - $300k
...maintain and improve switch software.Work on projects enabling AI networking infrastructure, enabling high performing fabrics for Tensor Processing... ...to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development...Worldwide$207k - $300k
...experience building and developing large-scale infrastructure, distributed systems or networks, or... ...in a technical leadership role leading project teams and setting technical direction... ...maintain, and enhance software solutions.The AI and Infrastructure team is redefining...Worldwide$262k - $364k
Technical Leadership and Mentorship: Lead and coach a distributed engineering team, fostering... ...roadmaps bridging RDMA, storage, and AI/ML, staying ahead of training and... ...with Storage Systems or Machine Learning Infrastructure.Google's software engineers develop the next...Remote workWorldwide$159k - $231k
...correlation studies, and root-cause analyses. Lead the triage and resolution of power... ...Power Integrity Engineer within Platforms Infrastructure Engineering, you will play a pivotal role... ...will directly shape our next generation AI hardware for AL/ML, server, networking, and...Worldwide$126k - $180k
...scale, multi-node environments.Experience using or integrating AI/LLMs (such as Gemini) into daily engineering or troubleshooting... ...security, reliability and scalability, running the full stack from infrastructure to applications to devices and hardware. Our teams are...Internship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Tech Lead, TPU AI Infrastructure. Be the first to apply!
