EPISODE · Jul 1, 2026 · 34 MIN
Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX) | Kimbo Chen, Cam Quilici, Bryan Shan, Jordan Nanos
from SemiAnalysis Weekly · host Jordan Nanos, Doug O'Laughlin
DeepSeek V4 claims a 100x KVcache reduction versus a standard MoE model, hitting 1M context length through compressed sparse attention and heavily compressed attention. Kimbo (@Kimbochen), Cam Quilici (@noslawextratost), Bryan Shan join Jordan Nanos (@JordanNanos) to break down what changed from V3, why the new MHC dimension tripped up NVIDIA on day zero, and how Mega MoE fuses compute and communication into a single kernel. The vLLM versus SGLang NDA access gap and the Huawei day zero optimization guide circulating on Twitter.The crew walks through the InferenceX article on going from day zero to day 43 support and what that grind actually looks like across different hardware. Subscribe for weekly semiconductor and AI infrastructure analysis from the SemiAnalysis team.Referenced:DeepSeekV4 1.6T Day 0 to Day 43 Performance Over Time - GB300 NVL72, Huawei, MI355X, B200: https://newsletter.semianalysis.com/p/deepseekv4-16t-day-0-to-day-43-performanceChapters:(00:00) DeepSeek V4 vs V3 changes(01:00) Sparse attention and KV cache reduction(03:04) Day zero runtime support challenges(05:34) What Mega MoE actually is(08:38) Downsides of fusing kernels(10:25) MegaKernel benchmark claims(12:59) AMD FP4 optimization gains(15:14) Compounding step by step improvements(17:59) vLLM versus SGLang competition(19:34) Open source vs vendor libraries
Embed this episode
NOW PLAYING
Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX) | Kimbo Chen, Cam Quilici, Bryan Shan, Jordan Nanos
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.