Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator episode artwork

EPISODE · Apr 15, 2026 · 22 MIN

Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 39 | cs.CV, cs.AI Authors: Luozheng Qin, Jia Gong, Qian Qiao, Tianjiao Li, Li Xu, Haoyu Pan, Chao Qu, Zhiyu Tan, Hao Li Title: Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator Arxiv: http://arxiv.org/abs/2604.08121v1 Abstract: Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher computational costs than understanding, particularly for video. This imbalance motivates us to invert the conventional paradigm: rather than extending understanding-centric MLLMs to support generation, we propose Uni-ViGU, a framework that unifies video generation and understanding by extending a video generator as the foundation. We introduce a unified flow method that performs continuous flow matching for video and discrete flow matching for text within a single process, enabling coherent multimodal generation. We further propose a modality-driven MoE-based framework that augments Transformer blocks with lightweight layers for text generation while preserving generative priors. To repurpose generation knowledge for understanding, we design a bidirectional training mechanism with two stages: Knowledge Recall reconstructs input prompts to leverage learned text-video correspondences, while Capability Refinement fine-tunes on detailed captions to establish discriminative shared representations. Experiments demonstrate that Uni-ViGU achieves competitive performance on both video generation and understanding, validating generation-centric architectures as a scalable path toward unified multimodal intelligence. Project Page and Code: https://fr0zencrane.github.io/uni-vigu-page/.

Episode metadata supplied by the publisher feed · Published Apr 15, 2026

Embed this episode

NOW PLAYING

Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator

0:00 22:44

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 22 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on April 15, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!