"How to game the METR plot" by shash42 episode artwork

EPISODE · Dec 21, 2025 · 12 MIN

"How to game the METR plot" by shash42

from LessWrong (Curated & Popular)

TL;DR: In 2025, we were in the 1-4 hour range, which has only 14 samples in METR's underlying data. The topic of each sample is public, making it easy to game METR horizon length measurements for a frontier lab, sometimes inadvertently. Finally, the “horizon length” under METR's assumptions might be adding little information beyond benchmark accuracy. None of this is to criticize METR—in research, its hard to be perfect on the first release. But I’m tired of what is being inferred from this plot, pls stop! 14 prompts ruled AI discourse in 2025 The METR horizon length plot was an excellent idea: it proposed measuring the length of tasks models can complete (in terms of estimated human hours needed) instead of accuracy. I'm glad it shifted the community toward caring about long-horizon tasks. They are a better measure of automation impacts, and economic outcomes (for example, labor laws are often based on number of hours of work). However, I think we are overindexing on it, far too much. Especially the AI Safety community, which based on it, makes huge updates in timelines, and research priorities. I suspect (from many anecdotes, including roon's) the METR plot has influenced significant investment [...] ---Outline:(01:24) 14. prompts ruled AI discourse in 2025(04:58) To improve METR horizon length, train on cybersecurity contests(07:12) HCAST Accuracy alone predicts log-linear trend in METR Horizon Lengths --- First published: December 20th, 2025 Source: https://www.lesswrong.com/posts/2RwDgMXo6nh42egoC/how-to-game-the-metr-plot --- Narrated by TYPE III AUDIO. ---Images from the article:

Episode metadata supplied by the publisher feed · Published Dec 21, 2025

Embed this episode

TL;DR: In 2025, we were in the 1-4 hour range, which has only 14 samples in METR's underlying data. The topic of each sample is public, making it easy to game METR horizon length measurements for a frontier lab, sometimes inadvertently. Finally, the “horizon length” under METR's assumptions might be adding little information beyond benchmark accuracy. None of this is to criticize METR—in research, its hard to be perfect on the first release. But I’m tired of what is being inferred from this ...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

"How to game the METR plot" by shash42

0:00 12:05

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (Curated & Popular)?

This episode is 12 minutes long.

When was this LessWrong (Curated & Popular) episode published?

This episode was published on December 21, 2025.

Can I download this LessWrong (Curated & Popular) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!