CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models episode artwork

EPISODE · May 19, 2026 · 42 MIN

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models

from Embodied AI 101 · host Shaoqing Tan

Closed-loop framework coupling Vision-Language Models with Video Generation Models at step-level granularity. Mitigates long-horizon drift and mid-clip errors in goal-directed video reasoning for robotic planning.

Episode metadata supplied by the publisher feed · Published May 19, 2026

Embed this episode

NOW PLAYING

CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models

0:00 42:10

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Embodied AI 101?

This episode is 42 minutes long.

When was this Embodied AI 101 episode published?

This episode was published on May 19, 2026.

Can I download this Embodied AI 101 episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!