SemiAnalysis

Ep. 030 - Rubin Ultra Drops From 1 TB to 192GB of HBM (Memory)

Original source
Artwork for Ep. 030 - Rubin Ultra Drops From 1 TB to 192GB of HBM (Memory)

Summary

The episode’s core thesis is that the memory market has shifted enough that taller HBM stacks no longer make the most sense for many frontier AI systems. Instead of maximizing bytes per package, vendors and customers are optimizing for total shipped accelerators, bandwidth per dollar, and scarce supply-chain resources. Nvidia’s Rubin Ultra is presented as the key case study, moving from an early terabyte-class concept to a much smaller 192 GB/package configuration built around 8-high HBM4. The speakers argue this is driven by DRAM scarcity, limited wafer additions, and the fact that inference and post-training now dominate demand. They also say model scaling trends and memory-skewed SKUs could fragment accelerator offerings further, with future systems right-sizing memory by workload rather than defaulting to the largest possible stack.

Notes