EPISODE · Dec 19, 2025 · 16 MIN
What’s In My Human Feedback? Learning Interpretable Descriptions of Preference Data
from Best AI papers explained · host Enoch H. Kang
This paper introduces a method for automatically decoding hidden preferences from language model training data. By utilizing sparse autoencoders, the method translates complex text embeddings into a small set of interpretable features that explain why human annotators prefer one response over another. The research reveals that feedback datasets often contain conflicting signals, such as Reddit users favoring informal jokes while other groups disfavor them. Notably, the authors demonstrate that What’s In My Human Feedback? (WIMHF) can identify misaligned or unsafe preferences, such as a bias against model refusals in certain benchmarks. These discovered features allow developers to curate safer datasets by flipping harmful labels and to personalize model behavior based on specific user stylistic choices. Ultimately, the work provides a human-centered diagnostic tool to make the black-box process of model alignment more transparent and controllable.
Embed this episode
NOW PLAYING
What’s In My Human Feedback? Learning Interpretable Descriptions of Preference Data
No transcript for this episode yet
Similar Episodes
No similar episodes found.