EPISODE · Jul 23, 2026 · 16 MIN
Video-Action Models for Robot Learning
from Embodied AI 101 · host Shaoqing Tan
Introduces Video-Action Models (VAMs) that leverage pretrained internet-scale video models such as Cosmos-Predict2 as backbones instead of VLMs, paired with a flow-matching action decoder. Claims approximately 10x sample efficiency gains over standard vision-language-action models on real-world pick-and-place tasks.
Embed this episode
NOW PLAYING
Video-Action Models for Robot Learning
0:00
16:50
1×
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
Frequently Asked Questions
How long is this episode of Embodied AI 101?
This episode is 16 minutes long.
When was this Embodied AI 101 episode published?
This episode was published on July 23, 2026.
Can I download this Embodied AI 101 episode?
Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!