Person

Bryan Shan

SemiAnalysis

Artwork for Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX) | Kimbo Chen, Cam Quilici, Bryan Shan, Jordan Nanos
SemiAnalysis WeeklyCam Quilici, Kimbo Chen, Bryan Shan

Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX) | Kimbo Chen, Cam Quilici, Bryan Shan, Jordan Nanos

DeepSeek V4’s big leap is 1M context via aggressive sparse-attention and KV-cache compression, paired with a mega-MOE/mega-kernel approach to speed expert computation. The episode also compares day-zero support across Nvidia, Huawei Ascend, and AMD, highlighting how early access, kernel fusion, and tooling maturity shape real inference performance.