EPISODE · Jun 2, 2026 · 28 MIN
Generative Depth Supervision for Embodied Vision-Language Models
from Embodied AI 101 · host Shaoqing Tan
Vision-language model that adds generative depth prediction during pre-training for physical grounding; achieves SOTA on embodied benchiments and transfers directly to real-robot tasks.
Embed this episode
NOW PLAYING
Generative Depth Supervision for Embodied Vision-Language Models
0:00
28:36
1×
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
Frequently Asked Questions
How long is this episode of Embodied AI 101?
This episode is 28 minutes long.
When was this Embodied AI 101 episode published?
This episode was published on June 2, 2026.
Can I download this Embodied AI 101 episode?
Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!