TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning episode artwork

EPISODE · Oct 9, 2025 · 27 MIN

TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning

from Daily Paper Cast · host Jingwen Liang, Gengyu Wang

🤗 Upvotes: 59 | cs.AI, cs.CL, cs.LG Authors: Jiaru Zou, Soumya Roy, Vinay Kumar Verma, Ziyi Wang, David Wipf, Pan Lu, Sumit Negi, James Zou, Jingrui He Title: TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning Arxiv: http://arxiv.org/abs/2510.06217v1 Abstract: Process Reward Models (PRMs) have recently emerged as a powerful framework for enhancing the reasoning capabilities of large reasoning models (LRMs), particularly in the context of test-time scaling (TTS). However, their potential for supervising LRMs on tabular reasoning domains remains underexplored. Through detailed empirical analyses, we identify that existing PRMs, though widely adopted for supervising text-only reasoning steps, struggle with table-specific operations such as sub-table retrieval and schema interaction, leading to critical performance bottlenecks. To address this limitation, we propose TaTToo, a novel table-grounded PRM framework that (i) reasons explicitly over tabular reasoning steps and (ii) integrates tool-based verification to provide precise reward supervision. Concretely, we first design a scalable data curation pipeline that constructs over 60k high-quality step-level annotations by integrating table verification rationales with tool-based executions. Building on the collected data, we train TaTToo with a dual-stage paradigm: cold-start supervised fine-tuning to capture tool-use reasoning patterns, followed by reinforcement learning with tool-grounded reward shaping to align our model with table-based verification. We provide a comprehensive evaluation of the policy improvement induced by our newly designed PRM. Across 5 challenging tabular reasoning benchmarks covering numerical reasoning, fact-checking, and data analysis, TaTToo improves downstream policy LRMs by 30.9% at inference, surpasses strong PRM baselines such as Qwen-2.5-Math-PRM-72B with only 8B parameters, and demonstrates strong generalizability across diverse TTS strategies.

Episode metadata supplied by the publisher feed · Published Oct 9, 2025

Embed this episode

NOW PLAYING

TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning

0:00 27:21

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Daily Paper Cast?

This episode is 27 minutes long.

When was this Daily Paper Cast episode published?

This episode was published on October 9, 2025.

Can I download this Daily Paper Cast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!