VINO: A Unified Visual Generator with Interleaved OmniModal Context episode artwork

EPISODE · Jan 7, 2026 · 23 MIN

VINO: A Unified Visual Generator with Interleaved OmniModal Context

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 22 | cs.CV Authors: Junyi Chen, Tong He, Zhoujie Fu, Pengfei Wan, Kun Gai, Weicai Ye Title: VINO: A Unified Visual Generator with Interleaved OmniModal Context Arxiv: http://arxiv.org/abs/2601.02358v1 Abstract: We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independent modules for each modality, VINO uses a shared diffusion backbone that conditions on text, images and videos, enabling a broad range of visual creation and editing tasks under one model. Specifically, VINO couples a vision-language model (VLM) with a Multimodal Diffusion Transformer (MMDiT), where multimodal inputs are encoded as interleaved conditioning tokens, and then used to guide the diffusion process. This design supports multi-reference grounding, long-form instruction following, and coherent identity preservation across static and dynamic content, while avoiding modality-specific architectural components. To train such a unified system, we introduce a multi-stage training pipeline that progressively expands a video generation base model into a unified, multi-task generator capable of both image and video input and output. Across diverse generation and editing benchmarks, VINO demonstrates strong visual quality, faithful instruction following, improved reference and attribute preservation, and more controllable multi-identity edits. Our results highlight a practical path toward scalable unified visual generation, and the promise of interleaved, in-context computation as a foundation for general-purpose visual creation.

Episode metadata supplied by the publisher feed · Published Jan 7, 2026

Embed this episode

NOW PLAYING

VINO: A Unified Visual Generator with Interleaved OmniModal Context

0:00 23:53

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 23 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on January 7, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!