InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation episode artwork

EPISODE · May 19, 2026 · 24 MIN

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 31 | cs.CV Authors: Yang Yue, Fangyun Wei, Tianyu He, Jinjing Zhao, Zanlin Ni, Zeyu Liu, Jiayi Guo, Lei Shi, Yue Dong, Li Chen, Ji Li, Gao Huang, Dong Chen Title: InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation Arxiv: http://arxiv.org/abs/2605.14333v1 Abstract: Text and faces are among the most perceptually salient and practically important patterns in visual generation, yet they remain challenging for autoregressive generators built on discrete tokenization. A central bottleneck is the tokenizer: aggressive downsampling and quantization often discard the fine-grained structures needed to preserve readable glyphs and distinctive facial features. We attribute this gap to standard discrete-tokenizer objectives being weakly aligned with text legibility and facial fidelity, as these objectives typically optimize generic reconstruction while compressing diverse content uniformly. To address this, we propose InsightTok, a simple yet effective discrete visual tokenization framework that enhances text and face fidelity through localized, content-aware perceptual losses. With a compact 16k codebook and a 16x downsampling rate, InsightTok significantly outperforms prior tokenizers in text and face reconstruction without compromising general reconstruction quality. These gains consistently transfer to autoregressive image generation in InsightAR, producing images with clearer text and more faithful facial details. Overall, our results highlight the potential of specialized supervision in tokenizer training for advancing discrete image generation.

Episode metadata supplied by the publisher feed · Published May 19, 2026

Embed this episode

NOW PLAYING

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation

0:00 24:22

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 24 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on May 19, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!