EPISODE · Jun 1, 2026 · 52 MIN
2026-05-31 · 特辑|从 LLM 到 VLA:会看会说会动的模型,差在哪个字母
from The Long View
一期把当下最容易混淆的几个模型范式一次讲清的技术特辑:大语言模型 LLM、视觉语言模型 VLM、视觉动作模型 VA、视觉语言动作模型 VLA,以及世界模型。从最底层的 token 与 Transformer 讲起,一路深入到 2026 年最前沿的 VLA 是怎么把连续动作塞进一个语言模型里的——离散化、扩散、流匹配三条路线,双系统快慢分离,以及联合训练为什么不可或缺。再讲清世界模型这个唯一向前预测世界的范式,和 VLA 的三种互补关系,以及 LeCun 的世界模型派与规模派之间真正的分歧。给已经在这个领域、但想把这几个概念边界彻底厘清的人听。
Embed this episode
Ready to play
2026-05-31 · 特辑|从 LLM 到 VLA:会看会说会动的模型,差在哪个字母
0:00
52:28
1×
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.
Frequently Asked Questions
How long is this episode of The Long View?
This episode is 52 minutes long.
When was this The Long View episode published?
This episode was published on June 1, 2026.
Can I download this The Long View episode?
Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!