Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval episode artwork

EPISODE · Aug 8, 2026 · 20 MIN

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 26 | cs.CV Authors: Zelong Sun, Jun Wang, Kaicheng Yang, Tiancheng Gu, Ziyong Feng, Zhiwu Lu Title: Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval Arxiv: http://arxiv.org/abs/2608.06060v1 Abstract: Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.

Episode metadata supplied by the publisher feed · Published Aug 8, 2026

Embed this episode

NOW PLAYING

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

0:00 20:35

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 20 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on August 8, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!