Topic

KV Cache

Artwork for WEKA's Val Bercovici: KV Cache, DeepSeek V4, HBF, SLC vs QLC NAND, CXL, NVLink, Tokenomics
Semi DopedVal Bercovici

WEKA's Val Bercovici: KV Cache, DeepSeek V4, HBF, SLC vs QLC NAND, CXL, NVLink, Tokenomics

Val Bercovici argues AI inference is becoming a memory-and-token-economics problem, with KV cache compression, network-attached memory, and NAND tiering reshaping system design. He also makes a blunt case that CXL is losing relevance, HBF will likely need SLC-like endurance buffering, and SaaS economics are being upended by token OPEX.