Ep. 034 - Engrams: How DeepSeek Offloads KV Cache to DRAM and SSD (Core Research) | Jordan Nanos, Cam Quilici, Alec Ibarra, Bryan Shan

Original source
Artwork for Ep. 034 - Engrams: How DeepSeek Offloads KV Cache to DRAM and SSD (Core Research) | Jordan Nanos, Cam Quilici, Alec Ibarra, Bryan Shan

Guests

Cam QuiliciSemiAnalysis frontend developer

Cam Quilici is a frontend developer at SemiAnalysis working on AI infrastructure and inference tooling.

Bryan ShanSemiAnalysis

Bryan Shan appears to be associated with SemiAnalysis and discussed AI infrastructure and accelerator performance on the podcast.

Alec IbarraSemiAnalysis / InferenceX

Alec Ibarra appears to be a member of the SemiAnalysis / InferenceX team working on AI infrastructure research and benchmarking.

Summary

Jordan Nanos, Cam Quilici, Bryan Shan, and Alec Ibarra center the episode on inference-system design: Engrams, DRAM/SSD offload, and the KV-cache bottlenecks that arise in agentic coding workloads. They describe Engrams as learned n-gram-style token embeddings that can be stored in a hash table and retrieved at inference time, with the core motivation being to reduce HBM pressure and move more of the memory burden to cheaper tiers like DRAM. The team then walks through Agent X, a replay benchmark built from real internal agentic coding traces, to measure cache reuse, prefill/decode behavior, and offloading under realistic traffic. They also discuss profitability for open-source inference on current high-end GPUs, arguing that under their TCO and pricing assumptions GB300-class serving can be highly lucrative. Later segments compare TPU v7 against Nvidia hardware, frame Vera Rubin as an incremental but important bandwidth-driven upgrade, and conclude that ultra-fast token generation only matters when it shortens the full tool-using workflow rather than just the model’s raw decode time.

Notes