Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX) | Kimbo Chen, Cam Quilici, Bryan Shan, Jordan Nanos episode artwork

EPISODE · Jul 1, 2026 · 34 MIN

Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX) | Kimbo Chen, Cam Quilici, Bryan Shan, Jordan Nanos

from SemiAnalysis Weekly · host Jordan Nanos, Doug O'Laughlin

DeepSeek V4 claims a 100x KVcache reduction versus a standard MoE model, hitting 1M context length through compressed sparse attention and heavily compressed attention. Kimbo (@Kimbochen), Cam Quilici (@noslawextratost), Bryan Shan join Jordan Nanos (@JordanNanos) to break down what changed from V3, why the new MHC dimension tripped up NVIDIA on day zero, and how Mega MoE fuses compute and communication into a single kernel. The vLLM versus SGLang NDA access gap and the Huawei day zero optimization guide circulating on Twitter.The crew walks through the InferenceX article on going from day zero to day 43 support and what that grind actually looks like across different hardware. Subscribe for weekly semiconductor and AI infrastructure analysis from the SemiAnalysis team.Referenced:DeepSeekV4 1.6T Day 0 to Day 43 Performance Over Time - GB300 NVL72, Huawei, MI355X, B200: https://newsletter.semianalysis.com/p/deepseekv4-16t-day-0-to-day-43-performanceChapters:(00:00) DeepSeek V4 vs V3 changes(01:00) Sparse attention and KV cache reduction(03:04) Day zero runtime support challenges(05:34) What Mega MoE actually is(08:38) Downsides of fusing kernels(10:25) MegaKernel benchmark claims(12:59) AMD FP4 optimization gains(15:14) Compounding step by step improvements(17:59) vLLM versus SGLang competition(19:34) Open source vs vendor libraries

Episode metadata supplied by the publisher feed · Published Jul 1, 2026

Embed this episode

NOW PLAYING

Ep. 017 - DeepSeek V4 and Huawei Ascend NPU Performance (InferenceX) | Kimbo Chen, Cam Quilici, Bryan Shan, Jordan Nanos

0:00 34:56

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of SemiAnalysis Weekly?

This episode is 34 minutes long.

When was this SemiAnalysis Weekly episode published?

This episode was published on July 1, 2026.

Can I download this SemiAnalysis Weekly episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!