Episode 170: DeepSeek V3 China’s open source 671B parameter LLM episode artwork

EPISODE · Dec 26, 2024 · 25 MIN

Episode 170: DeepSeek V3 China’s open source 671B parameter LLM

from A Cast of Pods · host Jose Acierto

The document details DeepSeek-V3, a 671B-parameter Mixture-of-Expert large language model. It covers the model's architecture, including Multi-Head Latent Attention and an innovative auxiliary-loss-free load balancing strategy for DeepSeekMoE. The training process, encompassing pre-training on 14.8 trillion tokens and post-training using supervised fine-tuning and reinforcement learning, is described. Extensive evaluations demonstrate DeepSeek-V3's strong performance across various benchmarks, surpassing many open-source and achieving results comparable to leading closed-source models. Finally, the document explores infrastructure optimizations, including an FP8 mixed-precision framework, and suggests improvements for future AI hardware design. DOWNLOAD HERE:DeepSeek-V3 Documentation: GitHub Deepseek-V3 Download : GitHub

Episode metadata supplied by the publisher feed · Published Dec 26, 2024

Embed this episode

NOW PLAYING

Episode 170: DeepSeek V3 China’s open source 671B parameter LLM

0:00 25:32

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of A Cast of Pods?

This episode is 25 minutes long.

When was this A Cast of Pods episode published?

This episode was published on December 26, 2024.

Can I download this A Cast of Pods episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!