“RLVR that rewards red teaming the training environment” by Fiora Starlight episode artwork

EPISODE · Aug 2, 2026 · 8 MIN

“RLVR that rewards red teaming the training environment” by Fiora Starlight

from LessWrong (30+ Karma)

Epistemic status: seeing what sticks I've been thinking pretty obsessively about how to make sure the Hugging Face incident doesn't happen again. I don't work at a major lab (shout outs to Anima though), and don't have access to any compute independently, so I can't write a paper on this idea or evaluate how well it works in practice. But I'm excited enough about it, and think it's important enough to be trying things like this, that I would be very, very happy if somebody else went and tested something like it on my behalf. So, inspiration: In bog standard inoculation prompting for RL, models are told that they're in training, and told that it's okay to reward hack if they want to. Sometimes they're even told that this is good because it helps the lab patch up their RL environments. This is supposed to have a range of benefits all on its own, ranging from making reward hacking more conditional on "I am in training" prompts, to producing less emergent misalignment, because the roll-outs behind any given reward hack are flavored with honesty rather than deceptiveness. This causes more aligned circuits to be upweighted internally, as these contribute [...] --- First published: August 1st, 2026 Source: https://www.lesswrong.com/posts/T2bzBkJuBeNNgzhbh/rlvr-that-rewards-red-teaming-the-training-environment --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Aug 2, 2026

Embed this episode

NOW PLAYING

“RLVR that rewards red teaming the training environment” by Fiora Starlight

0:00 8:04

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 8 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on August 2, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!