Large Language Models Can Self-Improve in Long-context Reasoning episode artwork

EPISODE · Nov 22, 2024 · 11 MIN

Large Language Models Can Self-Improve in Long-context Reasoning

from Artificial Discourse · host Kenpachi

This research paper investigates the potential for large language models (LLMs) to self-improve in long-context reasoning, which involves processing and understanding complex information spread across long stretches of text. The authors propose a novel approach called SEALONG that leverages the LLMs' ability to generate multiple outputs for a given question and then scores these outputs using a method called Minimum Bayes Risk (MBR). The MBR approach prioritizes outputs that align better with each other, thereby filtering out outputs that might be incorrect or hallucinatory. SEALONG then uses these high-scoring outputs for further training, either through supervised fine-tuning or preference optimization. The authors demonstrate through extensive experiments that SEALONG significantly improves the long-context reasoning performance of LLMs without requiring expert model annotations or human labeling.

Episode metadata supplied by the publisher feed · Published Nov 22, 2024

Embed this episode

NOW PLAYING

Large Language Models Can Self-Improve in Long-context Reasoning

0:00 11:49

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Artificial Discourse?

This episode is 11 minutes long.

When was this Artificial Discourse episode published?

This episode was published on November 22, 2024.

Can I download this Artificial Discourse episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!