SemiAnalysis Weekly

Ep. 030 - Long Live the Short King: Why 4-HI HBM Wins (Memory) | Myron Xie, Jordan Nanos

Original source
Artwork for Ep. 030 - Long Live the Short King: Why 4-HI HBM Wins (Memory) | Myron Xie, Jordan Nanos

Summary

Jordan Nanos and Myron Xie explain why the industry is revisiting the tradeoff between HBM bandwidth and capacity, and why 4-high stacks may win economically in a memory-constrained market. They say Nvidia’s Rubin Ultra roadmap appears to have moved from an early 1TB/package concept with four compute dies and 16-high HBM4E stacks to a 192GB HBM4 design with two compute dies and 8-high stacks. The core reason, they argue, is supply: DRAM wafer capacity is tight, HBM demand from accelerators is rising, and shipping less memory per package can maximize the number of accelerators delivered. They emphasize that once a system has enough memory for weights and KV cache, extra capacity has diminishing returns for inference, while bandwidth remains the real limiter. The episode also forecasts more custom SKUs, a potential shift of bottlenecks toward logic wafers and substrates, and a memory crunch that may not ease this decade.

Notes

Hosts

Myron XieJordan Nanos