HOME Cloud & Infrastructure Engineering Principal Engineer, Resource Optimization and Fleet Logic
  • Google
  • Thornton, CO
  • Full-Time
  • <72 Hours
Google VERIFIED EMPLOYER

Principal Engineer, Resource Optimization and Fleet Logic.

Cloud & Infrastructure Engineering Full-Time

Principal Engineer, Resource Optimization and Fleet Logic: our view in 3 lines...

  • The Role:This role is for a principal engineer building global capacity planning and fleet optimisation architecture for Google Cloud’s large-scale compute infrastructure.
  • The Person:The person will lead the design of global fleet planning, model and forecast capacity across compute and storage resources, and integrate AI/ML and mathematical optimisation into planning workflows.
  • Requirements:The role requires a bachelor’s degree in Computer Science or similar technical field, 15 years of software engineering experience, experience with large-scale capacity planning and IaaS/PaaS solutions, and preferred experience with a master’s degree or PhD.

About the role

Minimum qualifications:

  • Bachelor's degree in Computer Science or similar technical field, or equivalent practical experience.
  • 15 years of experience as a software engineer.
  • Experience delivering large-scale capacity planning, IaaS/PaaS solutions, or fleet management systems.

Preferred qualifications:

  • Master's degree or PhD in Computer Science or a field related (e.g., Networking or Security Systems).
  • Experience architecting, leading, and delivering large-scale capacity planning, combinatorial optimization, fleet management, or distributed infrastructure transformations from concept to deployment.
  • Deep understanding of modern AI/ML infrastructure demands (TPU/GPU topologies, accelerators) with the ability to integrate AI-driven solutions.
  • Ability to influence and lead without direct authority, building strong cross-organizational relationships across disparate teams (e.g., Hardware, Software, Supply Chain, and Product Management).
  • Exceptional communication skills, and ability to articulate complex mathematical, economic, and architectural concepts to engineering leaders and executive business stakeholders.

About the job:

Google runs one of the largest computational fleets in the world, with compute, storage, networking and dedicated accelerators spread across all continents. As Principal Engineer, you will own the architectural outlook and end-to-end technical strategy for our global capacity planning ecosystem. You will serve as the primary technical authority for a challenge, knitting together complex planning constraints, optimization scenarios, and machine deployment workflows to get optimal capacity online in the shortest time.

As the Principal Engineer, Resource Optimization and Fleet Logic, you will pioneer the technological vision for Google’s planetary-scale capacity planning and fleet optimization architecture. You will define the multi-year architectural path to orchestrate data center capacity planning, placement, supply planning, and fleet reconfiguration for ML, compute & storage demand —ultimately empowering our largest, most demanding internal and external AI/ML customers.

Google Cloud accelerates every organization’s ability to digitally transform its business and industry. We deliver enterprise-grade solutions that leverage Google’s cutting-edge technology, and tools that help developers build more sustainably. Customers in more than 200 countries and territories turn to Google Cloud as their trusted partner to enable growth and solve their most critical business problems.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $307000 - $427000 (USD) + 30% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities:

  • Lead the design and evolution of next generation global fleet planning, defining a multi-year engineering roadmap to deliver 10x more data center capacity with ML, compute & storage resources.
  • Leverage deep technical knowledge across the data center stack, spanning ML hardware (TPUs/GPUs), general compute, power, cooling, physical space, supply chain workflows, and network topology to model, forecast, and dynamically reconfigure fleet resources.
  • Partner across AI & Infrastructure organizations to define a unified capacity management product suite meeting AI training and inference for internal and external customers.
  • Integrate cutting-edge AI/ML, operations research, and advanced mathematical optimization into capacity planning workflows to compress capacity cycles, eliminate stranded capacity, and maximize data center utilization.
  • Collaborate with academic institutions and industry pioneers to anticipate technology trends.
Published September 15, 2026
Location Thornton, CO
Job Type Full-Time