“Measuring Activation Control in LLMs” by Marek Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Africa episode artwork

EPISODE · Aug 14, 2026 · 17 MIN

“Measuring Activation Control in LLMs” by Marek Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Africa

from LessWrong (30+ Karma)

TL;DR Inspired by the introspective awareness and CoT controllability papers, we made a benchmark to measure how well models can control their activations while completing a simple task. We are motivated by the concern that highly introspective models could control their activations, confounding probes and other monitors, and potentially even influencing their own training.We ran this on 25 open weight models ranging from 4B to 744B. We find that most language models are able to not only increase the salience of a concept in their residual stream on command, but also dial its strength up and down, including during specific intervals relative to the duration of the task. We also find that models are unable to control at which specific layer this is done.Counterintuitively, we find that within five of the seven model families we tested, the newest model scores lowest. For some reason, one of the oldest and smallest models of the panel, Llama 3.1 8B, performs best.It's not clear to us that newer models should have poorer control over their internal representations. More likely, where they “think” stops being the activation space, and becomes something else. We are looking for feedback (and other possible [...] ---Outline:(00:13) TL;DR(02:01) Methods(08:58) Results(15:41) Discussion(16:29) Acknowledgements --- First published: August 12th, 2026 Source: https://www.lesswrong.com/posts/HgvwxjzgwvsEvAiBH/measuring-activation-control-in-llms --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Aug 14, 2026

Embed this episode

NOW PLAYING

“Measuring Activation Control in LLMs” by Marek Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Africa

0:00 17:08

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 17 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on August 14, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!