Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding? episode artwork

EPISODE · Jun 13, 2026 · 20 MIN

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 71 | cs.CV, cs.AI, cs.CL Authors: Jiaqi Tang, Jianmin Chen, Youyang Zhai, Wei Wei, Runtao Liu, Mengjie Zhao, Xiangyu Wu, Qingfa Xiao, Qifeng Chen Title: Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding? Arxiv: http://arxiv.org/abs/2606.08063v1 Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in visual understanding, yet their performance degrades significantly under real-world visual corruptions. While existing robustness enhancement approaches exist, they are limited: black-box feature alignment lacks interpretability, and white-box text-based reasoning cannot restore lost pixel-level details. This work investigates a fundamental research question: Can MLLMs recover corrupted visual content by themselves? To address this, we propose Robust-U1, a novel framework that equips MLLMs with explicit visual self-recovery capability for robust understanding. The approach comprises three core stages: supervised fine-tuning for initial reconstruction, reinforcement learning with dual rewards (pixel-level SSIM and semantic-level CLIP similarity) for aligning high visual quality, and multimodal reasoning that jointly considers both the corrupted input and the recovered image. Extensive experiments demonstrate that Robust-U1 achieves state-of-the-art robustness on the real-world corruption benchmark and maintains superior performance under adversarial corruptions on general VQA benchmarks. Analysis confirms that high-quality visual recovery directly enhances reasoning performance, establishing self-recovery as a critical mechanism for robust visual understanding. The source code is available at https://github.com/jqtangust/Robust-U1.

Episode metadata supplied by the publisher feed · Published Jun 13, 2026

Embed this episode

NOW PLAYING

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

0:00 20:28

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 20 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on June 13, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!