Stanford CS153 Frontier Systems | The Discipline of Delivering Value per Gigawatt

Original source
Artwork for Stanford CS153 Frontier Systems | The Discipline of Delivering Value per Gigawatt

Summary

This episode is a deep dive into the engineering and operational discipline behind Google’s frontier AI infrastructure. Amin Vahdat frames the central challenge as delivering maximum capability and user value per gigawatt, emphasizing that reliability, utilization, goodput, and workload fit matter more than raw infrastructure spend. He walks through why system balance across flops, HBM, SRAM, network, storage, and CPU is essential, and why classic architectural ideas like Amdahl’s law still apply to modern distributed AI systems. The conversation also covers optical circuit switching, TPU topology, hardware depreciation, fast-changing inference demand from coding agents, and why Google is splitting TPUs into specialized inference and training variants. A recurring theme is that compute remains bottlenecked by energy, power procurement, and multi-year infrastructure lead times, even if algorithmic efficiency improves.

Notes

Hosts

Amin VahdatAnjney MidhaMichael Abbott

Topics

AI Infrastructure StrategyGigawatt-Scale Capacity PlanningInference Versus TrainingEnergy And Data Center PowerHyperscale Supply ChainsData Center PowerDistributed TrainingAI Infrastructure

Mentioned

Norm JupySundar PichaiAmin BadatAmeen BadatJeff DeanJensen HuangDemis HassabisEric SchmidtRich Sutton