Pretraining Safety w/ Ethan Roland episode artwork

EPISODE · Jul 9, 2026 · 1H 28M

Pretraining Safety w/ Ethan Roland

from Into AI Safety · host Jacob Haimes

What if the safest AI models weren't built by adding guardrails after training, but by shaping what gets learned in the first place? Ethan Roland, senior alignment researcher at AE Studio and first author on an ICML 2026 spotlight paper, joins Jacob to talk about gradient routing, a technique that routes dangerous capabilities into isolated parts of a model's architecture where they can be locked or removed entirely. They get into the absorption effect, KYC-style access control frameworks, and what it would actually take for frontier labs to adopt this kind of work before it's needed rather than after. Chapters(00:00) - Introduction (06:39) - Inside AE Studio (15:26) - China & the Alignment vs. Controllability Framing (18:23) - Data Filtering & Gradient Routing (Aside) (30:39) - Mixture of Experts Explained (Aside) (36:25) - Why Pre-Training Interventions Are Rare (42:43) - Ethan's Theory of Change (56:17) - Access Control Governance and KYC (Aside) (01:04:47) - The Researcher's Role in Policy Advocacy (01:11:38) - Speed Round (01:27:55) - Outro LinksBelow are the most important links for this episode. For more, visit the episode page on Kairos.fm.Ethan's websiteThe paper landing pagePreprint - Gradient Routing: Masking Gradients to Localize Computation in Neural NetworksICLR paper and webpage - Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMsWikipedia article - CBRN defenseNTI tutorials on bioweapons and nuclear testing

Episode metadata supplied by the publisher feed · Published Jul 9, 2026

Embed this episode

Ready to play

Pretraining Safety w/ Ethan Roland

0:00 1:28:57

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of Into AI Safety?

This episode is 1 hour and 28 minutes long.

When was this Into AI Safety episode published?

This episode was published on July 9, 2026.

Can I download this Into AI Safety episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!