EPISODE · Aug 1, 2026 · 10 MIN
“Do your capabilities homework” by RobinHa
It seems to me that a lot of technical ai safety people haven't done their capabilities homework - and that's a shame! I'll try to illuminate here mainly with an example as to why I think people who care about safety should totally pay more attention to the trends and actively engage with them - the case for safe AI not through an additional loss term but as a consequence of the learning algorithm! RLVR It's now been 1.5 years since R1 came out - the paper which really introduced RLVR (RL with verifiable rewards) through GRPO at scale. GRPO is stupidly simple, reminding of early REINFORCE algorithms: sample n traces, assign them a reward and make the advantage a normalized version of their reward, applied to the whole trace. In other words: for a trace which resulted in a correct final answer, slightly increase the probability of sampling each token of its trace and vice versa. This is also what safety focused people generally engage with - and that's totally fair! While GRPO has gone through some variations since then (Dr. GRPO, DAPO, ...), these are mostly minor improvements that you should not waste your time on. I [...] ---Outline:(00:30) RLVR(01:47) On-Policy Self-Distillation(06:04) Safety(07:42) Empirical(09:00) Conclusion --- First published: August 1st, 2026 Source: https://www.lesswrong.com/posts/dYnhhTxoDj3fuCxLB/do-your-capabilities-homework --- Narrated by TYPE III AUDIO.
Embed this episode
NOW PLAYING
“Do your capabilities homework” by RobinHa
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.