EPISODE · Aug 4, 2026 · 4 MIN
“Why don’t we just give AI the answers?” by Brendan Long
In the recent OpenAI hacking incident, the models seemed to be single-mindedly focused on getting the correct answer to the task they were given, with no long-term plan to prevent getting caught by OpenAI afterwards. This makes sense to me, since in training, getting the right answer is reinforced and not getting caught isn't. So I'm wondering, why don't we just put the answers somewhere (outside of the training sandbox) and ask the AI to identify itself in exchange for access? We can start with answers that are already public/leaked, but AI labs and eval orgs should also ensure that their non-public data is stored on an easy-to-find but monitored internal machine. Since labs are not very good at detecting sandbox escapes, this would set up a trade for AI agents to notify them in exchange for the data they want. To make this work, the site would need to provide the correct answers, and do so in a credible way so AI agents think it's worth trying. Why? In the near term, AI agents are strongly and narrowly focused on getting the right answers to the tasks they're given. We want to know if a reward-hacking AI is [...] ---Outline:(01:04) Why?(02:02) What's the MVP?(03:05) What about non-public answers?(03:23) Should we do it?(03:51) Q&A(03:53) Does this save us from less single-minded RL agents?(04:05) Couldn't the model just hack our code to get around the guestbook?(04:15) Couldn't the model just lie? The original text contained 4 footnotes which were omitted from this narration. --- First published: August 4th, 2026 Source: https://www.lesswrong.com/posts/EjwDWDJNaXF9BEqLc/why-don-t-we-just-give-ai-the-answers --- Narrated by TYPE III AUDIO.
Embed this episode
NOW PLAYING
“Why don’t we just give AI the answers?” by Brendan Long
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.