MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer episode artwork

EPISODE · Sep 23, 2025 · 25 MIN

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 37 | cs.CV, cs.CL, cs.LG Authors: Yanghao Li, Rui Qian, Bowen Pan, Haotian Zhang, Haoshuo Huang, Bowen Zhang, Jialing Tong, Haoxuan You, Xianzhi Du, Zhe Gan, Hyunjik Kim, Chao Jia, Zhenbang Wang, Yinfei Yang, Mingfei Gao, Zi-Yi Dou, Wenze Hu, Chang Gao, Dongxu Li, Philipp Dufter, Zirui Wang, Guoli Yin, Zhengdong Zhang, Chen Chen, Yang Zhao, Ruoming Pang, Zhifeng Chen Title: MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer Arxiv: http://arxiv.org/abs/2509.16197v1 Abstract: Unified multimodal Large Language Models (LLMs) that can both understand and generate visual content hold immense potential. However, existing open-source models often suffer from a performance trade-off between these capabilities. We present Manzano, a simple and scalable unified framework that substantially reduces this tension by coupling a hybrid image tokenizer with a well-curated training recipe. A single shared vision encoder feeds two lightweight adapters that produce continuous embeddings for image-to-text understanding and discrete tokens for text-to-image generation within a common semantic space. A unified autoregressive LLM predicts high-level semantics in the form of text and image tokens, with an auxiliary diffusion decoder subsequently translating the image tokens into pixels. The architecture, together with a unified training recipe over understanding and generation data, enables scalable joint learning of both capabilities. Manzano achieves state-of-the-art results among unified models, and is competitive with specialist models, particularly on text-rich evaluation. Our studies show minimal task conflicts and consistent gains from scaling model size, validating our design choice of a hybrid tokenizer.

Episode metadata supplied by the publisher feed · Published Sep 23, 2025

Embed this episode

NOW PLAYING

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

0:00 25:40

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 25 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on September 23, 2025.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!