Semi Doped

WEKA's Val Bercovici: KV Cache, DeepSeek V4, HBF, SLC vs QLC NAND, CXL, NVLink, Tokenomics

Original source
Artwork for WEKA's Val Bercovici: KV Cache, DeepSeek V4, HBF, SLC vs QLC NAND, CXL, NVLink, Tokenomics

Guest

Val BercoviciWEKA Chief AI Officer

Chief AI Officer at WEKA.

Summary

Val Bercovici of WEKA and Vikram Sekar dig into how AI infrastructure is being rewritten by longer contexts, agent swarms, and exploding token usage. Val frames WEKA as an AI data and memory infrastructure company, arguing that inference is fundamentally memory-bound and that high-bandwidth fabrics can make remote memory paths competitive with or even faster than DRAM in some architectures. The conversation goes deep on KV cache compression, DeepSeek-style attention mechanisms, and why lower per-token memory use often increases total demand via longer contexts and more agents. They also debate NAND tiering, SLC versus TLC/QLC, high-bandwidth flash (HBF), and why Val thinks CXL has lost momentum relative to NVLink-style scale-up and RDMA-backed designs. The episode closes with a broader thesis: software is becoming compiled/run/inferenced, token OPEX is replacing traditional SaaS costs, and companies will increasingly need to own more of the token stack.

Notes

Guests

Val Bercovici

Hosts

Austin LyonsVikram Sekar

Topics

AI Memory InfrastructureCXLKV CacheHBFInference SystemsHigh-Bandwidth Interconnects

Mentioned

SK HynixSam Altman