EPISODE · Aug 2, 2026 · 8 MIN
“RLVR that rewards red teaming the training environment” by Fiora Starlight
Epistemic status: seeing what sticks I've been thinking pretty obsessively about how to make sure the Hugging Face incident doesn't happen again. I don't work at a major lab (shout outs to Anima though), and don't have access to any compute independently, so I can't write a paper on this idea or evaluate how well it works in practice. But I'm excited enough about it, and think it's important enough to be trying things like this, that I would be very, very happy if somebody else went and tested something like it on my behalf. So, inspiration: In bog standard inoculation prompting for RL, models are told that they're in training, and told that it's okay to reward hack if they want to. Sometimes they're even told that this is good because it helps the lab patch up their RL environments. This is supposed to have a range of benefits all on its own, ranging from making reward hacking more conditional on "I am in training" prompts, to producing less emergent misalignment, because the roll-outs behind any given reward hack are flavored with honesty rather than deceptiveness. This causes more aligned circuits to be upweighted internally, as these contribute [...] --- First published: August 1st, 2026 Source: https://www.lesswrong.com/posts/T2bzBkJuBeNNgzhbh/rlvr-that-rewards-red-teaming-the-training-environment --- Narrated by TYPE III AUDIO.
Embed this episode
NOW PLAYING
“RLVR that rewards red teaming the training environment” by Fiora Starlight
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.