MoE Giants: Decoding the 670 Billion Parameter Showdown Between DeepSeek V3 and Mistral Large episode artwork

EPISODE · Dec 25, 2025 · 30 MIN

MoE Giants: Decoding the 670 Billion Parameter Showdown Between DeepSeek V3 and Mistral Large

from Neural intel Pod · host Neuralintel.org

Neural Intel Podcast EpisodeMoE Giants: Decoding the 670 Billion Parameter Showdown Between DeepSeek V3 and Mistral LargeThis week on Neural Intel, we dive deep into the architectural blueprints of two colossal Mixture-of-Experts (MoE) models: DeepSeek V3 (673B/671B) and Mistral 3 Large (675B/673B). We explore the configurations that define these massive language models, noting their shared traits, such as an embedding dimension of 7,168 and a vocabulary size of 129K. Both architectures employ a FeedForward (SwiGLU) module, and the initial three blocks use a dense FFN with a hidden size of 18,432 instead of the MoE layer.The core of the discussion focuses on how each model utilizes its MoE layer, both of which contain 128 experts. We contrast the resource allocation and expert frequency: DeepSeek V3/R1 is configured to activate one shared expert plus six additional experts per token (1 shared + 6 experts active per token), resulting in only 37B active parameters per inference step. In contrast, Mistral 3 Large activates one shared expert plus four additional experts per token (1 shared + 4 experts active per token), leading to 39B active parameters per inference step.We also analyze other crucial architectural differences visible in their configuration files, including the intermediate hidden layer dimensions—2,048 for DeepSeek V3/R1 versus 4,096 for Mistral 3 Large. Join us as we dissect how these subtle parameter choices—affecting multi-head latent attention, expert distribution, and shared experts—impact overall efficiency and performance in the race to build the most capable and resourceful large language models.

Episode metadata supplied by the publisher feed · Published Dec 25, 2025

Embed this episode

NOW PLAYING

MoE Giants: Decoding the 670 Billion Parameter Showdown Between DeepSeek V3 and Mistral Large

0:00 30:18

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Neural intel Pod?

This episode is 30 minutes long.

When was this Neural intel Pod episode published?

This episode was published on December 25, 2025.

Can I download this Neural intel Pod episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!