
Summary
This episode is a technical first reaction to OpenAI’s Jalapeño reveal at Hot Chips. The hosts argue that OpenAI optimized for user-facing latency and energy per request, not just peak throughput, and that the right way to judge the chip is on Pareto curves and delivered performance per watt. They highlight the architecture’s NUMA-like local HBM slices, the focus on KV-cache locality, and a scale-up system built around Broadcom networking to support 128-chip and 2048-chip domains. A major theme is that AI-assisted EDA may have helped OpenAI go from first RTL to tapeout in about nine months, which they see as a wake-up call for the rest of the accelerator industry. They also debate whether OpenAI should monetize the chip externally or keep it internal as a strategic advantage.