Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text episode artwork

EPISODE · Jul 25, 2026 · 20 MIN

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 34 | cs.CV Authors: Xu Wang, Kaixiang Yao, Miao Pan, Xiaohe Zhou, Xuanyu Liu, Wenqi Zhang, Xuhong Zhang Title: Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Arxiv: http://arxiv.org/abs/2607.21072v1 Abstract: Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-generation models. This makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments directly in pixel space. We propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics. ProVisE also includes an Agentic builder that constructs and validates task-specific protocols for new benchmarks. We further introduce SpatialGen-Bench, a curated diagnostic benchmark of 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Results show that image-generation models are competitive when spatial answers can be externalized directly in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning and establish a metric-compatible testbed for studying spatial cognition in image-generation models.

Episode metadata supplied by the publisher feed · Published Jul 25, 2026

Embed this episode

NOW PLAYING

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

0:00 20:39

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 20 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on July 25, 2026.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!