EPISODE · Jul 28, 2026
Kimi K3 Goes Trillion-Scale With Sparse MoE Efficiency
from AI Post Transformers
This episode examines Kimi K3, Moonshot AI's open-weight frontier model boasting 2.8 trillion total parameters with only 104 billion active per token, a million-token context window, and native multimodal training from the ground up. The discussion traces the architectural lineage behind the model's claimed 2.5x scaling-efficiency gain over its predecessor Kimi K2, connecting its Mixture-of-Experts design back to Shazeer's 2017 sparsely-gated MoE work and contrasting its hybrid attention approach with the limits of standard residual connections from the 2015 ResNet paper. It also unpacks the systems-engineering side of running a model this large, particularly how Expert Parallelism turns token routing into a datacenter networking problem once hundreds of experts are sharded across GPUs. Listeners get a clear breakdown of why native multimodal training tends to be more stable than bolting a pretrained vision encoder onto a text-only model after the fact. The episode sets up a deeper dive into which of K3's four credited innovations — Kimi Delta Attention, Attention Residuals, Stable LatentMoE, and refined training recipes — is actually doing the heavy lifting. Sources: 1. Kimi K3 Goes Trillion-Scale With Sparse MoE Efficiency https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf 2. Deep Residual Learning for Image Recognition — Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun, 2015 (CVPR 2016) https://scholar.google.com/scholar?q=Deep+Residual+Learning+for+Image+Recognition 3. DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging — Matteo Pagliardini, Amirkeivan Mohtashami, Francois Fleuret, Martin Jaggi, 2024 https://scholar.google.com/scholar?q=DenseFormer%3A+Enhancing+Information+Flow+in+Transformers+via+Depth+Weighted+Averaging 4. Value Residual Learning for Alleviating Attention Concentration in Transformers — Zhanchao Zhou, Tianyi Wu, Zhiyun Jiang, Zhenzhong Lan (and collaborators), 2024 https://scholar.google.com/scholar?q=Value+Residual+Learning+for+Alleviating+Attention+Concentration+in+Transformers 5. Hyper-Connections — Defa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng, Yunyao Mao, Banggu Wu, Qiyang Min, Xun Zhou (ByteDance Seed), 2024 https://scholar.google.com/scholar?q=Hyper-Connections 6. Attention Residuals — Kimi Team, 2026 https://scholar.google.com/scholar?q=Attention+Residuals 7. Kimi Linear: An Expressive, Efficient Attention Architecture — Kimi Team et al., 2025 (arXiv:2510.26692) https://scholar.google.com/scholar?q=Kimi+Linear%3A+An+Expressive%2C+Efficient+Attention+Architecture 8. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI, 2024 https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model 9. ReplaySSM: Cache SSM Inputs, Not State — Dao AI Lab, 2026 (blog) https://scholar.google.com/scholar?q=ReplaySSM%3A+Cache+SSM+Inputs%2C+Not+State Interactive Visualization: Kimi K3 Goes Trillion-Scale With Sparse MoE Efficiency
Embed this episode
NOW PLAYING
Kimi K3 Goes Trillion-Scale With Sparse MoE Efficiency
No transcript for this episode yet
Similar Episodes
No similar episodes found.