“SOTA alignment assessments don’t strongly update us against misalignment” by Alexa Pan episode artwork

EPISODE · Aug 1, 2026 · 46 MIN

“SOTA alignment assessments don’t strongly update us against misalignment” by Alexa Pan

from LessWrong (30+ Karma)

Anthropic concluded in the April Mythos Preview alignment risk update that the model "does not possess any unknown propensities that would increase alignment risk." The report argues that if Mythos Preview were coherently misaligned, it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”). While I agree with the report on the above bottom-line conclusions (substantially on priors), I think there are gaps in its argument which weaken the current assessment and might invalidate future assessments. In particular, the report often uses weak evidence to justify reliability. The report gives fairly weak experimental evidence for Mythos Preview having insufficient capabilities to evade monitoring. The model is plausibly often eval-aware and underelicited in the relevant capability evaluations. So, it might silently sandbag if coherently misaligned, or unintentionally underperform if otherwise misaligned. This limitation is important: one could argue that lack of covert capabilities for sophisticated sabotage (a subset of the capabilities I discuss here) is the single most load bearing argument in alignment risk reports.Authors of the report could have made calibrated guesses about Mythos Preview's covert capabilities, especially for covert sabotage, based on other factors despite the relatively weak [...] ---Outline:(03:14) How reliability fits into the overall safety argument(05:20) Reliability claims by AI companies(05:56) Reliability claims by external evaluators(06:32) Alignment assessments are less reliable than developers claim(07:15) 1: Measuring capabilities to covertly undermine alignment assessments(10:05) Issues with evaluation awareness(13:28) Issues with underestimating covert capabilities(16:34) Issues with sandbagging rule-out(19:10) 2: Stress-testing alignment assessments with auditing games(20:16) An auditing failure with Mythos(22:08) AuditBench results(24:04) 3: Conditioning on misalignment should make us think that certain covert capabilities are better than expected(26:04) Bottom line on the strength of current alignment assessments(28:52) Conclusion(29:29) Appendix:(29:32) Why I focus on motive / alignment assessments in alignment risk reports(30:44) Auditability vs. Trustedness(33:11) More reliability claims by developers and third party evaluators(33:27) Mythos Alignment Risk Update(34:41) Opus 4.6 Sabotage Risk Report(35:24) GPT 5.5 System card(36:21) Muse Spark system card(37:08) Mythos Alignment Risk Update, safety arguments against sandbagging(37:15) From the Mythos Alignment Update, §5.3.4, p. 24:(38:09) UK AISI evaluations for Opus 4.7(39:28) Past auditing games by Anthropic(41:45) Anti-auditing capability measurements(43:12) Conditioning on coherent misalignment updates us on certain covert capabilities The original text contained 46 footnotes which were omitted from this narration. --- First published: July 31st, 2026 Source: https://www.lesswrong.com/posts/oirrSj3itFLSyscW8/sota-alignment-assessments-don-t-strongly-update-us-against --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Episode metadata supplied by the publisher feed · Published Aug 1, 2026

Embed this episode

NOW PLAYING

“SOTA alignment assessments don’t strongly update us against misalignment” by Alexa Pan

0:00 46:20

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 46 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on August 1, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!