Principal Nexus Capacity Engineer: AI-Driven Fleet Planning
Jobleads-US
Google LLC is seeking a Principal Engineer, Nexus Capacity Engineer to drive planet-scale capacity delivery across data centers. You will shape the architectural path, lead a full-stack team, and partner with hardware, software, supply chain, and product management to optimize capacity usage.
This role requires a deep background in large-scale capacity planning, IaaS/PaaS, and AI/ML infrastructure, with the ability to influence without direct authority and deliver transformative solutions.
#J-18808-Ljbffr Jobleads-US$307k - $427k
...Principal Engineer, Nexus Capacity Engineer Share Principal Engineer, Nexus Capacity... ...large-scale capacity planning, IaaS/PaaS solutions, or fleet management systems.... ...understanding of modern AI/ML infrastructure... ...ability to integrate AI-driven solutions. Deep understanding...PrincipalFleetTemporary work- ...Technical Program Manager to lead capacity planning and fleet strategy for our Inference... ...You will work closely with Engineering, Product, Infrastructure,... ...utilization of our AI inference fleet. You will... ...own capacity modeling, data-driven forecasting, and coordination...Fleet
$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is... ...Kubernetes, databases, capacity management,... ...environments.As a Principal SRE, you will shape... ...direction of NVIDIA’s AI Platform Runtime... ...next-generation AI-driven products and... ...budgets, capacity planning, fault tolerance,...PrincipalFull time$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements team as... ...operate these systems reliably at fleet scale. In this role, you... ...orchestration, and SW-driven serviceability. You will drive... ...increasingly known as “the AI computing company.” We're looking...PrincipalFleetFull timeRemote workShift work$152k - $241.5k
...harness the power of AI to deliver groundbreaking... ...excellence.Conduct capacity management and planning to meet ongoing... ...lifecycle management, fleet reliability/auto-healing... ...observability or data-driven operations (AIOps/ML-... ...Ruby.Mentored other engineers and influenced technical...FleetFull time- ...the world's largest AI chip, 56 times... ...by the Wafer-Scale Engine (WSE). This team will... ...labs.As a Principal SRE, you will define... ...scaling our inference fleet through self-service... ...observability, capacity orchestration, rollout... ...reliable capacity planning, workload...PrincipalFleetShift work
- ...potential of generative AI to power the... ...tackling challenges and are driven by execution. Ready to... ...infrastructure layer that every engineering team and customer... ...across colo server fleets, on-premises lab clusters... ...environments. Lead capacity planning and hardware lifecycle...Fleet
- ...potential of generative AI to power the... ...tackling challenges and are driven by execution. Ready to... ...infrastructure underpinning our engineering organization must be... ...lifecycle automation, fleet auto-remediation, and... .... Own FinOps and capacity planning across cloud,...FleetRemote work
- ...of Generative AI cloud at AWS? Do... ...The AWS Hardware Engineering team creates... ...organizational, planning, and communication... ...of AI driven automation and... ...seemingly infinite capacity at the lowest possible... ..., Managers, Principals) and groups (... ...decisions, and fleet health thereafter...FleetInternshipLocal area
- ...Management, R&D, and roadmap planning. What You’ll Be Doing... ...with customer architects and engineering teams to define suitable PCIe... ...partner for next-generation AI, HPC, storage, networking, and... ...firmware bring-up, or software-driven debug of PCIe/CXL subsystems....Principal
$140k - $215k
...world’s most advanced AI-native platform. We work... ...to work for a mission-driven company leveraging AI to... ...sensor platform.As a Senior Engineer on this team, you will... ...position in cross-org planning and architecture... ...knowledge of failure modes at fleet scaleChampion...FleetFull timeWork experience placementWork at officeLocal area$272k - $431.25k
...cloud environments. We are looking for Principal Software Engineers to help shape the technical direction... ...the crowd:Experience with GPU clusters, AI/ML infrastructure, Kubernetes... ...VMaaS, managed Kubernetes, or multi-cloud fleet operations.Experience building internal...PrincipalFleetFull time$272k - $431.25k
...the unlimited potential of AI to define the next era of... ...world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide... ...operate safely at rack and fleet scale. Build open source... ...specifications, and execution plans that align teams across...PrincipalFleetFull timeRemote workShift work$307k - $427k
...for information engineering consisting of data... ...100x more capacity for explosive AI/ML demand growth... ...and integrate AI-driven solutions to improve... ....Google Global Fleet (GGF) is responsible... ...and capacity planning. Its mission is... ...its mission.As a Principal Software...PrincipalFleet$272k - $431.25k
...into the unlimited potential of AI to define the next era of... ...will work with a diverse team of engineers in mapping, perception, reconstruction... ...pipelines that transform fleet data into reliable map... ...perception, localization, simulation, planning, and infrastructure teams to...PrincipalFleetFull timeWorldwide$96.8k - $306.4k
...Lead Principal Platform Software Engineer Seattle, WA, United States United... ...placement, and fleet management at cloud scale... ...translates into decisions and plans that can have a major... ...innovative data-driven techniques to address... ...care. And with AI embedded across our products...PrincipalFleetTemporary workFlexible hoursShift work- ...concurrency, caching, sharding, fault isolation, capacity, and graceful degradation.Set the long-... ...fragmented systems, mentor senior engineers, and align technical and product leaders... ..., availability, multi-tenancy, capacity planning, observability, and cost efficiency....PrincipalWork at officeLocal area
$195k - $240k
...seeking a Senior Advanced Packaging Engineer to lead the development and... ...activities, including DOE planning, failure analysis, and reliability... ...computing.computing, AI/ML hardware or mobile products... ...semiconductor solution provider driven by its Purpose, To Make Our Lives...PrincipalHourly pay$195k - $285k
...potential of generative AI to power the... ...tackling challenges and are driven by execution. Ready... ...Matrix is looking for a Principal Systems Hardware Engineer to own system... ...delivery/liquid cooling; fleet-scale RAS; signal and... ...our inclusive rewards plan empowers our people to...PrincipalFleet$142.8k - $274.8k
...Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft... ...manufacturing, improving the planning process, quality, delivery,... .... We are looking for a Principal Hardware Engineer - Networking... ...platforms for Microsoft Azure AI infrastructure. This role...PrincipalOngoing contractWork at officeLocal areaWorldwide3 days per week$255k - $351k
...Staff/Principal AI Transformation Engineer San Jose, CA About the Company DiDi... ...into a shared-mobility fleet will generate immense social... ...reimagined and entirely driven by an AI-Native architecture... ...(Perception, Prediction, Planning & Control, and Simulation)...PrincipalFleetFull time- ...Jose SummaryJoin a world-class engineering group at Celestica focused on... ...advanced systems, including AI servers, switches,... ...strategy, and join department planning. Guide design engineers in complete... ...networking solutions enabling AI-driven growth.Built on a legacy of...PrincipalTemporary workWork at officeLocal areaImmediate startWorldwideShift work
$272k - $431.25k
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global... ...and optimize the speed and cost efficiency of AI development and testing systems.Leading software development...PrincipalFull timeWork experience placementWorldwide$175k - $265k
...potential of generative AI to power the... ...tackling challenges and are driven by execution. Ready to... ...infrastructure layer that every engineering team and customer... ...across colo server fleets, on-premises lab clusters... ...environments.Lead capacity planning and hardware lifecycle...Fleet- ...generation computing experiences—from AI and data centers, to PCs,... ...latency for perception/planning (ms-scale)Guide architectural... ...subsystemsCloud (training, simulation, fleet learning)Provide architectural... ...credibly with customer’s engineering leaders, AI architectsTrack record...PrincipalFleet
$142k - $194k
...require a BS or MS in Computer Science, Engineering, Robotics, Life Sciences, or a related field... ...automation, scientific workflows, AI-enabled systems, or hardware-software products... ...quality assurance, scripting, software test plans, user experience development, code review...PrincipalFull time$224k - $356.5k
...time. You will work with engineers running AI clouds at scale on the problems... ...for new NVIDIA platforms, capacity, services, and use cases... ...GPU or Network Operators.Driven improved fleet health or unit economics... ...by improving the factory planning function. Are you creative...FleetFull timeWork experience placement$174k - $252k
...technologies.Google's software engineers develop the next-generation... ...software to optimize the design, planning, deployment, and... ...for Google data centers. Our fleet consumes as much power as the... ...Cloud. Additionally, you will use AI to automate rack placement at...Fleet- ...and operate low-latency event-driven services end to end,... ...rollback strategies. Drive engineering standards, code reviews, testing... ..., performance optimization, capacity planning, autoscaling, and cloud cost... ...Foundational understanding of AI/ML technologies and...PrincipalFull timeContract workWork at office3 days per week
$285k - $335k
...vertically integrated AI infrastructure company... ...run their models. As a Principal Product Manager on our... ...foundation: the fleet-wide design, the fabric... ...senior infrastructure engineers and architects, and your... ...including cost per unit and capacity planning. ~ Analytical...PrincipalFleetTemporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Nexus Capacity Engineer: AI-Driven Fleet Planning. Be the first to apply!
- chief engineer Sunnyvale, CA
- senior chief engineer Sunnyvale, CA
- principal infrastructure engineer Sunnyvale, CA
- principal developer Sunnyvale, CA
- general engineer Sunnyvale, CA
- director software engineering Sunnyvale, CA
- engineering director Sunnyvale, CA
- director data engineering Sunnyvale, CA
- principal cloud engineer Sunnyvale, CA
- hotel chief engineer Sunnyvale, CA



