
Google's AI Infrastructure Chief, Amin Vahdat, on the Physics & Economics of Frontier AI
At 100,000-accelerator scale, something fails multiple times an hour, which is why Google's Chief Technologist for AI Infrastructure Amin Vahdat thinks FLOPS is a vanity metric. The metric that matters is what he calls goodput: the useful work a workload actually delivers through real failures. Amin walks through the calculus that split the TPU line into 8i and 8t for the first time, why the TPU's core primitives haven't changed since v1, and how Google and DeepMind co-design in the same rooms, intercepting chip architectures mid-flight before tape-out. He explains why long-horizon agents are sending demand for CPUs and storage through the roof alongside accelerators, how optical circuit switches reroute light to a spare rack in milliseconds, and why Google would rather wait on a utility than build its own gigawatt. We also cover orbital data centers and the multi-megawatt rack of 2036. Hosted by Sonya Huang, Sequoia Capital 0:00 – Introduction 1:47 – What makes a data center an AI data center 5:30 – Goodput, not FLOPS: holding yourself accountable when something fails every hour 11:52 – Doubling token capacity every six months, and where the gains actually come from 15:32 – The TPU bet: from a contrarian call in 2013 to splitting 8i and 8t 23:30 – The case for and against co-design 26:11 – Shoulder to shoulder with DeepMind: intercepting chips mid-flight 34:16 – Long-horizon agents change the shape of the data center 37:50 – Optical circuit switching and the state of networking 42:35 – Power is the binding constraint: utilities, gigawatts, and sizing a data center 49:23 – Training vs. serving clusters, seven-year-old TPUs, and open standards 58:29 – Orbital data centers and the supercomputer of 2036