WEKA's Val Bercovici: KV Cache, DeepSeek V4, HBF, SLC vs QLC NAND, CXL, NVLink, Tokenomics episode artwork

EPISODE · Jul 10, 2026 · 1H 4M

WEKA's Val Bercovici: KV Cache, DeepSeek V4, HBF, SLC vs QLC NAND, CXL, NVLink, Tokenomics

from Semi Doped · host Vikram Sekar and Austin Lyons

Vik welcomes Val Bercovici from Weka to discuss the rapidly evolving landscape of AI memory and storage. Val explains how Weka's architecture leverages high-bandwidth networks to make storage faster than motherboard DRAM. They dive into KV cache optimizations, the future of NAND flash tiers, and the role of CXL in AI inference. The episode concludes with a look at predictive memory offloading and the AI flywheel.Chapters:0:00 Welcome Val Bercovici, Weka1:59 Memory situation and model routing3:50 KV cache offloading to CMX6:10 Network faster than motherboard13:10 Weka as AI memory infrastructure14:45 Inference market is different16:06 Memory hierarchy and KV cache19:40 KV cache optimizations and demand25:20 DeepSeek's cache read pricing34:49 NAND flash tiers: SLC vs QLC43:01 High Bandwidth Flash (HBF)49:59 CXL versus other interconnectsFollow Chipstrat:Newsletter: https://www.chipstrat.comX: https://x.com/chipstratFollow Vik:Newsletter: https://www.viksnewsletter.com/X: https://x.com/vikramskrFollow Semi Doped:Get more of Austin and Vik daily, free!Sign up: https://daily.semidoped.com/

Episode metadata supplied by the publisher feed · Published Jul 10, 2026

Embed this episode

Vik welcomes Val Bercovici from Weka to discuss the rapidly evolving landscape of AI memory and storage. Val explains how Weka's architecture leverages high-bandwidth networks to make storage faster than motherboard DRAM. They dive into KV cache optimizations, the future of NAND flash tiers, and the role of CXL in AI inference. The episode concludes with a look at predictive memory offloading and the AI flywheel. Chapters: 0:00 Welcome Val Bercovici, Weka 1:59 Memory situation and model rout...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

WEKA's Val Bercovici: KV Cache, DeepSeek V4, HBF, SLC vs QLC NAND, CXL, NVLink, Tokenomics

0:00 1:04:11

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Semi Doped?

This episode is 1 hour and 4 minutes long.

When was this Semi Doped episode published?

This episode was published on July 10, 2026.

Can I download this Semi Doped episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!