What DeepSeek's Training Data Reveals About Model Voice episode artwork

EPISODE · Jul 29, 2026 · 26 MIN

What DeepSeek's Training Data Reveals About Model Voice

from My Weird Prompts

When Daniel ran his model evaluation for podcast script writing, DeepSeek V4 Pro won not on benchmarks but on feel — its dialogue simply sounded more authentic. This episode traces why: DeepSeek's training corpus is 60% English and 30% Chinese, but that Chinese portion is heavily weighted toward narrative literature — Tang dynasty chuanqi tales, modern WeChat fiction, and other voice-driven storytelling. By contrast, Kimi from Moonshot AI trains on conversational social media like Zhihu and Weibo, while Qwen deliberately filters cultural bias through data augmentation. We explore the ASR analogy of alingual models, the concrete differences in how each model structures dialogue, and why subjective reasoning style may become the real differentiator as model capabilities converge. Episode #755716 — open it directly at myweirdprompts.com/755716

Episode metadata supplied by the publisher feed · Published Jul 29, 2026

Embed this episode

NOW PLAYING

What DeepSeek's Training Data Reveals About Model Voice

0:00 26:23

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of My Weird Prompts?

This episode is 26 minutes long.

When was this My Weird Prompts episode published?

This episode was published on July 29, 2026.

Can I download this My Weird Prompts episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!