“We need 3rd party Training-Run Assessments” by Alex Meinke episode artwork

EPISODE · Jul 5, 2026 · 35 MIN

“We need 3rd party Training-Run Assessments” by Alex Meinke

from LessWrong (30+ Karma)

Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety. By a Training-Run Assessment, or TRA, I mean an in-depth analysis of the post-training pipeline and dynamics leading up to a frontier model release. A TRA can look at intermediate checkpoints, training rollouts, RL environments, reward signals, SFT datasets, and the process by which the developer responded to warning signs.[1] In this post I will argue that: Final-checkpoint evaluations will be insufficient to assess scheming risks.TRAs can be more effective at detecting scheming.Frontier developers should involve third parties to do TRAs or verify safety claims by the developers. The rest of the post lays out a taxonomy of TRAs and sketches a path toward a 3rd party ecosystem for them. We, at Apollo Research, are intending to conduct 3rd party Training-Run Assessments in the future. Detecting Scheming may require Training-Run Assessments By scheming I mean an AI covertly pursuing misaligned goals while deliberately concealing its intentions or capabilities from its developers. I restrict attention to “coherent” forms of scheming where the model pursues somewhat stable misaligned goals across context windows, rather than misalignment that surfaces only as isolated, context-dependent defections. [...] ---Outline:(01:23) Detecting Scheming may require Training-Run Assessments(03:55) Why 3rd parties should perform Training-Run Assessments(04:12) Developers may lack incentives to adequately assess scheming(04:49) Developers' safety assessments lack credibility(05:31) External evaluators can bundle expertise for assessing scheming(06:17) 3rd party TRAs can be developed gradually(08:50) Checkpoint evals(08:54) What?(10:30) How?(11:15) Data inspections(11:19) What?(12:03) Why?(13:30) How?(15:11) Process reviews(15:15) What?(15:46) Why?(17:15) How?(18:04) Additional considerations(19:36) Conclusion(21:23) Appendix(21:26) How plausible is scheming that can be detected during post-training but not after training?(21:58) If scheming arises, it will be detectable at some point during post-training(26:55) Scheming will be much harder to detect in the final checkpoint(30:46) Other use cases for TRAs(31:55) Secret loyalties(32:47) Developer-implanted sandbagging(33:35) Strong incapability arguments(34:12) Models that can't even be evaluated without posing significant risk The original text contained 3 footnotes which were omitted from this narration. --- First published: July 5th, 2026 Source: https://www.lesswrong.com/posts/3HvvjffA65mHLwaWm/we-need-3rd-party-training-run-assessments --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Jul 5, 2026

Embed this episode

NOW PLAYING

“We need 3rd party Training-Run Assessments” by Alex Meinke

0:00 35:04

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 35 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 5, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!