A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement Learning episode artwork

EPISODE · Jun 2, 2025 · 25 MIN

A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement Learning

from Best AI papers explained · host Enoch H. Kang

This academic paper proposes a local data attribution framework for online reinforcement learning (RL). The framework uses influence functions to identify which training data records negatively impact the RL agent's learning within each training round. By filtering out these harmful records, the proposed method, called Influence-guided Intervention and Filtering (IIF), demonstrates improved performance and sample efficiency in standard RL tasks and also shows promise in reducing toxicity in Reinforcement Learning from Human Feedback (RLHF) for large language models. The paper analyzes the characteristics of influential records and the impact of different filtering levels on learning.

Episode metadata supplied by the publisher feed · Published Jun 2, 2025

Embed this episode

NOW PLAYING

A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement Learning

0:00 25:48

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Best AI papers explained?

This episode is 25 minutes long.

When was this Best AI papers explained episode published?

This episode was published on June 2, 2025.

Can I download this Best AI papers explained episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!