“Differential acceleration of alignment-relevant capabilities is a bad bet” by Zephaniah Roe episode artwork

EPISODE · Jul 21, 2026 · 15 MIN

“Differential acceleration of alignment-relevant capabilities is a bad bet” by Zephaniah Roe

from LessWrong (30+ Karma)

There is an idea floating around in the rough shape of "we need to accelerate capabilities that are differentially useful for safety research so AIs can help us make the future go better." The capabilities targeted are typically things bottlenecking alignment research, such as philosophical or conceptual reasoning. I feel nervous about this for two reasons. The first is that it's plausible that AI safety and AI R&D are bottlenecked by many of the same factors: AIs have poor epistemics, are bad at messy conceptual reasoning, and are unreliable at tasks without ground truth. Speeding up progress in any of these areas seems likely to speed up general AI R&D, giving everyone else less time to execute time-bottlenecked agendas (e.g., trying to do Plan A). The second reason I don't feel good about this is because I'm less confident it will help make handoff/deference/superalignment go well. To hand off conceptual alignment research to AIs we need to trust them to 1. be good at this research and 2. be generally trustworthy/aligned. We still don't know how to reliably prevent prosaic outer misalignment issues (e.g., sycophancy or going off-constitution), let alone worse issues that will make AIs less trustworthy in [...] ---Outline:(02:15) Examples of arguments for the acceleration of alignment-relevant capabilities(06:13) Why we should not do this kind of differential acceleration(06:18) Alignment bottlenecks are also capabilities bottlenecks(08:18) This could hurt time-bottlenecked strategies(09:48) I don't think that this would unblock superalignment/hand-off plans(11:53) But what about things like philosophy?(12:46) But isn't it pretty unlikely that you help labs make real capabilities progress?(14:03) My epistemic status The original text contained 4 footnotes which were omitted from this narration. --- First published: July 21st, 2026 Source: https://www.lesswrong.com/posts/FzGqnCkdKeTnZ9tjE/differential-acceleration-of-alignment-relevant-capabilities --- Narrated by TYPE III AUDIO.

Episode metadata supplied by the publisher feed · Published Jul 21, 2026

Embed this episode

NOW PLAYING

“Differential acceleration of alignment-relevant capabilities is a bad bet” by Zephaniah Roe

0:00 15:44

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LessWrong (30+ Karma)?

This episode is 15 minutes long.

When was this LessWrong (30+ Karma) episode published?

This episode was published on July 21, 2026.

Can I download this LessWrong (30+ Karma) episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!