EPISODE · Aug 6, 2026 · 9 MIN
“Three years of progress in 500 lines of code” by Gerard Boxo
TL;DR There is some consensus that LLMs are bad at hard-to-verify tasks. The question is whether models are getting better at them over time. As a motivating example, I gave one research-reproduction task to 12 models spanning three years of progress, to illustrate how (1) what looked like an emergent capability was a predictable trend, visible years earlier if progress was measured with granularity, and (2) for the earliest models, building a verifiable check would have been close to impossible: their attempts were so far from correct that a check would have had nothing to grade, so the task itself would have looked unverifiable. When thinking about current capabilities or forecasting future ones, researchers should be aware that binary judgments of usefulness can hide steady partial progress, and that a task looking 'unverifiable' today may say as much about the current capability profile of models as about the task itself. Introduction Frontier models like Fable seem to struggle with novel end-to-end research [1, 2], often producing slop, sometimes slop so bad that it would get you banned from arXiv [3]. Meanwhile, SWEs and researchers seem to find these tools incredibly useful. The best models are now capable of impressive [...] ---Outline:(00:11) TL;DR(01:07) Introduction(03:55) The Task: AI safety via debate (MNIST MCTS Debate)(07:58) Finishing Thoughts The original text contained 1 footnote which was omitted from this narration. --- First published: August 6th, 2026 Source: https://www.lesswrong.com/posts/K3NXziL6uJDeSYEHT/three-years-of-progress-in-500-lines-of-code --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Embed this episode
NOW PLAYING
“Three years of progress in 500 lines of code” by Gerard Boxo
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.