The GPU Myth: State of AI Compute 2026 | Stephen Balaban
Original source
Guest
Stephen Balaban is the co-founder and CTO of Lambda, an AI infrastructure and GPU cloud company.
Summary
Stephen Balaban makes the case that AI compute is still massively underbuilt and structurally non-commoditized because delivering it requires control over land, utility power, data-center construction, HPC architecture, software orchestration, and capital markets. He says demand keeps expanding as LLMs move from assistants to code generation and “software out,” so even a 10x efficiency gain would mostly become 10x more tokens consumed. The conversation goes deep on the physical stack: direct-to-chip liquid cooling, dry coolers, NVLink inside racks, InfiniBand/Ethernet between racks, RDMA, and the economics of depreciation, utilization, and GPU usable life. Balaban also details Lambda’s origin story, from face recognition and iPhone convnet experiments to DreamScope, workstation clusters, and ultimately a vertically integrated neocloud. He closes with a forward-looking thesis on “neural software,” self-assembling software, and a future where many people may effectively need at least one GPU of compute in daily life.