WEKA's Val Bercovici: KV Cache, DeepSeek V4, HBF, SLC vs QLC NAND, CXL, NVLink, Tokenomics
Original source
Guest
Chief AI Officer at WEKA.
Summary
Val Bercovici of WEKA and Vikram Sekar dig into how AI infrastructure is being rewritten by longer contexts, agent swarms, and exploding token usage. Val frames WEKA as an AI data and memory infrastructure company, arguing that inference is fundamentally memory-bound and that high-bandwidth fabrics can make remote memory paths competitive with or even faster than DRAM in some architectures. The conversation goes deep on KV cache compression, DeepSeek-style attention mechanisms, and why lower per-token memory use often increases total demand via longer contexts and more agents. They also debate NAND tiering, SLC versus TLC/QLC, high-bandwidth flash (HBF), and why Val thinks CXL has lost momentum relative to NVLink-style scale-up and RDMA-backed designs. The episode closes with a broader thesis: software is becoming compiled/run/inferenced, token OPEX is replacing traditional SaaS costs, and companies will increasingly need to own more of the token stack.