Ep. 034 - Engrams: How DeepSeek Offloads KV Cache to DRAM and SSD (Core Research) | Jordan Nanos, Cam Quilici, Alec Ibarra, Bryan Shan
Original source
Guests
Cam Quilici is a frontend developer at SemiAnalysis working on AI infrastructure and inference tooling.
Bryan Shan appears to be associated with SemiAnalysis and discussed AI infrastructure and accelerator performance on the podcast.
Alec Ibarra appears to be a member of the SemiAnalysis / InferenceX team working on AI infrastructure research and benchmarking.
Summary
Jordan Nanos, Cam Quilici, Bryan Shan, and Alec Ibarra center the episode on inference-system design: Engrams, DRAM/SSD offload, and the KV-cache bottlenecks that arise in agentic coding workloads. They describe Engrams as learned n-gram-style token embeddings that can be stored in a hash table and retrieved at inference time, with the core motivation being to reduce HBM pressure and move more of the memory burden to cheaper tiers like DRAM. The team then walks through Agent X, a replay benchmark built from real internal agentic coding traces, to measure cache reuse, prefill/decode behavior, and offloading under realistic traffic. They also discuss profitability for open-source inference on current high-end GPUs, arguing that under their TCO and pricing assumptions GB300-class serving can be highly lucrative. Later segments compare TPU v7 against Nvidia hardware, frame Vera Rubin as an incremental but important bandwidth-driven upgrade, and conclude that ultra-fast token generation only matters when it shortens the full tool-using workflow rather than just the model’s raw decode time.