【第75期】cDPO:通过发掘critical tokens去修正回答 episode artwork

EPISODE · Dec 14, 2024 · 12 MIN

【第75期】cDPO:通过发掘critical tokens去修正回答

from Seventy3

Seventy3: 用NotebookLM将论文生成播客,让大家跟着AI一起进步。今天的主题是:Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM’s Reasoning CapabilitySummaryThis research paper introduces cDPO, a novel approach to improve the reasoning capabilities of Large Language Models (LLMs). cDPO identifies "critical tokens"—tokens crucial to correct or incorrect reasoning—using contrastive estimation by comparing models trained on correct and incorrect reasoning trajectories. This allows for token-level reward adjustments during preference optimization, enhancing accuracy. Experiments on GSM8K and MATH500 benchmarks using Llama-3 and DeepSeek-math models demonstrate cDPO's superior performance over existing methods. The paper also explores the impact of various hyperparameters and offers an in-depth comparison with related techniques in contrastive estimation and reinforcement learning. The findings suggest that focusing on critical tokens significantly improves LLM reasoning accuracy.原文链接:https://arxiv.org/abs/2411.19943前往小宇宙评论区与主播互动

Episode metadata supplied by the publisher feed · Published Dec 14, 2024

Embed this episode

NOW PLAYING

【第75期】cDPO:通过发掘critical tokens去修正回答

0:00 12:51

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Seventy3?

This episode is 12 minutes long.

When was this Seventy3 episode published?

This episode was published on December 14, 2024.

Can I download this Seventy3 episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!