The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman
Original source
Guest
Andrew Feldman is the co-founder, CEO, and president of Cerebras Systems.
Summary
Andrew Feldman says the next AI bottleneck is inference, not training, and that the right metric is tokens per second per user because latency shapes chat, reasoning, and multi-step agent workflows. He walks through the chip stack—from CPUs and GPUs to TPUs, Trainium, and ASICs—and argues that HBM, CoWoS packaging, and 3nm capacity are the real supply-chain constraints, which Cerebras sidesteps with a 5nm wafer-scale design built around SRAM and redundancy. Feldman also frames reasoning, verification, guardrails, and multimodal/video as compute-expanding trends that favor faster systems. The conversation closes on megawatt-scale infrastructure, including a large OpenAI deal, AWS’s disaggregated prefill/decode setup, and Feldman’s belief that faster AI will eventually reshape SaaS by making custom tools easier to generate than to buy.