TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale episode artwork

EPISODE · Aug 12, 2026

TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale

from AI Post Transformers

This episode explores TFGN, an architectural approach to continual pre-training of large language models that claims to solve catastrophic forgetting without four common crutches: replay buffers, task identifiers, small-scale toy benchmarks, and external penalty terms like Fisher-information regularization. The hosts trace the lineage of the forgetting problem back to 1989, explain why popular fixes like LoRA-based parameter-efficient fine-tuning don't actually address forgetting (they just shrink the blast radius), and why classic regularization methods like Elastic Weight Consolidation break down at billion-parameter scale. They also clarify why long-context windows and prompt-based knowledge aren't a substitute for genuinely updating model weights on massive, unbounded corpora like full codebases or legal archives. The conversation lays out TFGN's core mechanism as a dense, input-conditioned overlay operating inside each transformer block, contrasting it with sparse mixture-of-experts routing, and sets up backward transfer as the key metric for measuring whether old knowledge survives new training. Listeners interested in how production LLMs might eventually absorb new domains without expensive retraining or fragile adapter stacking will find the framing of this open problem sharply drawn. Sources: 1. TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale — Anurup Ganguli, 2026 http://arxiv.org/abs/2605.15053 2. Overcoming catastrophic forgetting in neural networks (EWC) — J. Kirkpatrick et al., 2017 https://scholar.google.com/scholar?q=Overcoming+catastrophic+forgetting+in+neural+networks+%28EWC%29 3. Loss of plasticity in deep continual learning — S. Dohare et al., 2024, Nature https://scholar.google.com/scholar?q=Loss+of+plasticity+in+deep+continual+learning 4. Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-Tuning — O. Y. L. Imanov, 2026, arXiv:2601.18699 https://scholar.google.com/scholar?q=Mechanistic+Analysis+of+Catastrophic+Forgetting+in+Large+Language+Models+During+Continual+Fine-Tuning 5. Examining Forgetting in Continual Pre-training of Aligned Large Language Models — C.-A. Li and H.-Y. Lee, 2024, arXiv:2401.03129 https://scholar.google.com/scholar?q=Examining+Forgetting+in+Continual+Pre-training+of+Aligned+Large+Language+Models 6. Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models — I. Abbes, G. Subbaraj, M. Riemer, et al., 2025, arXiv:2508.01908 https://scholar.google.com/scholar?q=Revisiting+Replay+and+Gradient+Alignment+for+Continual+Pre-Training+of+Large+Language+Models 7. Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thought — Y. Zhang, B. Tang, T. Ju, S. Duan, G. Liu, 2025, arXiv:2512.21711 https://scholar.google.com/scholar?q=Do+Latent+Tokens+Think%3F+A+Causal+and+Adversarial+Analysis+of+Chain-of-Continuous-Thought Interactive Visualization: TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale

Episode metadata supplied by the publisher feed · Published Aug 12, 2026

Embed this episode

NOW PLAYING

TFGN: Replay-Free, Task-Free Continual Pre-Training at Scale

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on August 12, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!