EPISODE · May 19, 2026 · 42 MIN
CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models
from Embodied AI 101 · host Shaoqing Tan
Closed-loop framework coupling Vision-Language Models with Video Generation Models at step-level granularity. Mitigates long-horizon drift and mid-clip errors in goal-directed video reasoning for robotic planning.
Embed this episode
NOW PLAYING
CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models
0:00
42:10
1×
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
Frequently Asked Questions
How long is this episode of Embodied AI 101?
This episode is 42 minutes long.
When was this Embodied AI 101 episode published?
This episode was published on May 19, 2026.
Can I download this Embodied AI 101 episode?
Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!