EPISODE · Jul 9, 2026 · 35 MIN
mimic-video: Video-Action Models for Robot Learning
from Embodied AI 101 · host Shaoqing Tan
Introduces Video-Action Models (VAMs) that replace standard VLMs in VLAs with a pretrained internet-scale video backbone (Cosmos-Predict2) combined with a flow-matching action decoder. Achieves approximately 10× better sample efficiency on real-world pick-and-place tasks.
Embed this episode
NOW PLAYING
mimic-video: Video-Action Models for Robot Learning
0:00
35:37
1×
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
Frequently Asked Questions
How long is this episode of Embodied AI 101?
This episode is 35 minutes long.
When was this Embodied AI 101 episode published?
This episode was published on July 9, 2026.
Can I download this Embodied AI 101 episode?
Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!