Qwen-VLA: A Generalist Vision–Language–Action Robot Model episode artwork

EPISODE · May 29, 2026 · 35 MIN

Qwen-VLA: A Generalist Vision–Language–Action Robot Model

from Embodied AI 101 · host Shaoqing Tan

A single generalist VLA built on Qwen3.5-4B + 1.15B DiT flow-matching action decoder that unifies manipulation, navigation, and trajectory prediction across 11 embodiments via text-described embodiment prompts. Trained in four stages and outperforms task-specific specialists on real ALOHA and sim benchmarks without per-task fine-tuning.

Episode metadata supplied by the publisher feed · Published May 29, 2026

Embed this episode

NOW PLAYING

Qwen-VLA: A Generalist Vision–Language–Action Robot Model

0:00 35:31

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Embodied AI 101?

This episode is 35 minutes long.

When was this Embodied AI 101 episode published?

This episode was published on May 29, 2026.

Can I download this Embodied AI 101 episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!