Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL episode artwork

EPISODE · May 29, 2025 · 25 MIN

Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL

from Best AI papers explained · host Enoch H. Kang

This paper introduces Planning with a Natural Language Critic (PNLC), a novel approach for improving the planning capabilities of large language models (LLMs) in complex interactive tasks without relying on computationally expensive reinforcement learning (RL) fine-tuning or extensive inference-time search. PNLC trains a lightweight, goal-conditioned value function offline that predicts the likelihood of various future outcomes based on a proposed thought or strategy by the LLM agent. During inference, this value function acts as a natural language critic, providing the LLM with feedback on the potential positive and negative results of its thoughts, enabling the LLM to refine its reasoning and actions effectively and efficiently. Experiments on interactive tasks like web shopping, social deduction, and persuasion demonstrate that PNLC outperforms existing RL and prompting methods in both performance and efficiency, scaling to larger LLMs.

Episode metadata supplied by the publisher feed · Published May 29, 2025

Embed this episode

NOW PLAYING

Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL

0:00 25:11

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Best AI papers explained?

This episode is 25 minutes long.

When was this Best AI papers explained episode published?

This episode was published on May 29, 2025.

Can I download this Best AI papers explained episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!