“Single Forward Pass Evals on Fable, Opus 5, and GPT-5.6-Sol” by Christine Corry episode artwork

EPISODE · Aug 3, 2026 · 15 MIN

“Single Forward Pass Evals on Fable, Opus 5, and GPT-5.6-Sol” by Christine Corry

from LessWrong (30+ Karma)

This is a research update for an on-going replication of single-forward-pass evals done as part of the Second Look Fellowship. In following posts, we will run more comprehensive replications of previous work and release open source tooling for single forward pass eval elicitation. Code can be found here. tl;dr We replicate experiments from Greenblatt 2025 and Greenblatt 2026 on one baseline model from the original post, Opus 4.5. Our evaluations agree with the trends and quantitative values described in the original posts.We run similar evaluations on Claude Fable 5, Opus 5, and GPT-5.6-Sol and find that the newer models show a substantial jump in performance on some evals. Fable 5 gets 87.6% accuracy on Gen-Arithmetic with 10 problem repeats whereas previous SOTA around 60%.GPT-5.6-Sol experiences significant uplift from filler tokens and problem repeats on all 4 datasets; filler tokens/repeats double performance from baseline on 3-hop. Figure 1: Baseline (no-CoT) vs. each model's peak repeat-or-filler condition on Gen-Arithmetic and 2-Hop reasoning. Error bars are 95% paired-bootstrap CIs; * marks a significant gain over baseline (paired t-test, Holm-Bonferroni corrected). Background If models can successfully do complex computations in a single forward pass, they may be able [...] ---Outline:(00:32) tl;dr(01:48) Background(02:31) Previous Work(03:13) Datasets(04:36) Evaluation Design(05:55) Eliciting no-CoT(07:00) Results(07:23) Gen-Arithmetic(08:11) Comp-Math(08:38) 2-Hop(09:11) 3-Hop(09:43) Per-model profiles(09:56) Trends over repeat and filler conditions(10:23) Conclusion(10:43) Appendix(10:46) Are we sure they aren't reasoning?(12:35) Temperature(12:59) Performance with CoT(13:39) Prompt structure The original text contained 5 footnotes which were omitted from this narration. --- First published: August 2nd, 2026 Source: https://www.lesswrong.com/posts/bxaWTNrdgJpkLXmgm/single-forward-pass-evals-on-fable-opus-5-and-gpt-5-6-sol --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Aug 3, 2026

Embed this episode

NOW PLAYING

“Single Forward Pass Evals on Fable, Opus 5, and GPT-5.6-Sol” by Christine Corry

0:00 15:25

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 15 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on August 3, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!