SemiAnalysis

Ep. 034 - The Fight for Fast Tokens, TPU v7, Vera Rubin, and Engrams (AI Supply Chain, InferenceX)

Original source
Artwork for Ep. 034 -  The Fight for Fast Tokens, TPU v7, Vera Rubin, and Engrams (AI Supply Chain, InferenceX)

Summary

SemiAnalysis uses this episode to connect model architecture, memory hierarchy, and chip economics into one inference-stack story. The first half digs into engrams as learned n-gram embeddings stored in a hash table so models can push some meaning-building earlier and offload more of the inference footprint to DRAM, SSD, or CPU-pinned host memory. The second half introduces AgentX, a system-level benchmark for real agentic coding traffic, and uses it to estimate profitability across different accelerators and serving stacks. They then compare TPU v7, arguing it is already highly cost-competitive and will benefit from broader external availability and a native torch backend. The Vera Rubin discussion centers on memory bandwidth, disaggregated serving, and lookup-table quantization, with the main takeaway that inference economics are increasingly determined by how well vendors can convert bandwidth and memory capacity into realized tokens per dollar.

Notes