SemiAnalysis Weekly//Cam Quilici, Kimbo Chen, Bryan Shan
Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX) | Kimbo Chen, Cam Quilici, Bryan Shan, Jordan Nanos
DeepSeek V4’s big leap is 1M context via aggressive sparse-attention and KV-cache compression, paired with a mega-MOE/mega-kernel approach to speed expert computation. The episode also compares day-zero support across Nvidia, Huawei Ascend, and AMD, highlighting how early access, kernel fusion, and tooling maturity shape real inference performance.