What’s In My Human Feedback? Learning Interpretable Descriptions of Preference Data episode artwork

EPISODE · Dec 19, 2025 · 16 MIN

What’s In My Human Feedback? Learning Interpretable Descriptions of Preference Data

from Best AI papers explained · host Enoch H. Kang

This paper introduces a method for automatically decoding hidden preferences from language model training data. By utilizing sparse autoencoders, the method translates complex text embeddings into a small set of interpretable features that explain why human annotators prefer one response over another. The research reveals that feedback datasets often contain conflicting signals, such as Reddit users favoring informal jokes while other groups disfavor them. Notably, the authors demonstrate that What’s In My Human Feedback? (WIMHF) can identify misaligned or unsafe preferences, such as a bias against model refusals in certain benchmarks. These discovered features allow developers to curate safer datasets by flipping harmful labels and to personalize model behavior based on specific user stylistic choices. Ultimately, the work provides a human-centered diagnostic tool to make the black-box process of model alignment more transparent and controllable.

Episode metadata supplied by the publisher feed · Published Dec 19, 2025

Embed this episode

NOW PLAYING

What’s In My Human Feedback? Learning Interpretable Descriptions of Preference Data

0:00 16:14

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Best AI papers explained?

This episode is 16 minutes long.

When was this Best AI papers explained episode published?

This episode was published on December 19, 2025.

Can I download this Best AI papers explained episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!