Long Context Pre-Training with Lighthouse Attention episode artwork

EPISODE · May 13, 2026

Long Context Pre-Training with Lighthouse Attention

from AI Post Transformers

This episode explores a May 7, 2026 arXiv paper on Lighthouse Attention and asks whether long-context language models can be pretrained cheaply with hierarchical sparse attention, then switched back to standard dense attention late in training without losing dense-model quality. It explains why long-context training is so expensive even with FlashAttention, contrasting dense quadratic attention with sparse and hierarchical schemes that try to narrow which tokens interact. The discussion walks through the paper’s core design: building multi-level pooled query/key/value pyramids, using a gradient-free top-K selector to choose relevant causal subsequences, running ordinary FlashAttention on that smaller set, and scattering the results back. Listeners would find it interesting because it frames the method as a practical systems bet with potentially major implications for 128K- to million-token pretraining, while also stressing that the evidence is still preliminary and far from proving it works at frontier scale. Sources: 1. Long Context Pre-Training with Lighthouse Attention https://arxiv.org/pdf/2605.06554 2. H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences — Zhenhai Zhu, Radu Soricut, 2021 https://scholar.google.com/scholar?q=H-Transformer-1D%3A+Fast+One-Dimensional+Hierarchical+Attention+for+Sequences 3. LongT5: Efficient Text-To-Text Transformer for Long Sequences — Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontanon, Jianmo Ni, Yun-Hsuan Sung, Yinfei Yang, 2021 https://scholar.google.com/scholar?q=LongT5%3A+Efficient+Text-To-Text+Transformer+for+Long+Sequences 4. HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention — Yufei Xu, Fanxu Meng, Fan Jiang, Yuxuan Wang, Ruijie Zhou, Jiexi Wu, Zhixin Pan, Zhaohui Wang, Xiaojuan Tang, Wenjie Pei, Tongxuan Liu, Di Yin, Xing Sun, Muhan Zhang, 2026 https://scholar.google.com/scholar?q=HISA%3A+Efficient+Hierarchical+Indexing+for+Fine-Grained+Sparse+Attention 5. Long Context Pre-Training with Lighthouse Attention — Bowen Peng, Subho Ghosh, Jeffrey Quesnelle, 2026 https://scholar.google.com/scholar?q=Long+Context+Pre-Training+with+Lighthouse+Attention 6. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Y. X. Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, Wangding Zeng, 2025 https://scholar.google.com/scholar?q=Native+Sparse+Attention%3A+Hardware-Aligned+and+Natively+Trainable+Sparse+Attention 7. MoBA: Mixture of Block Attention for Long-Context LLMs — Enzhe Lu, Zhejun Jiang, Jingyuan Liu, Yulun Du, Tao Jiang, Chao Hong, Shaowei Liu, Weiran He, Enming Yuan, Yuzhi Wang, Zhiqi Huang, Huan Yuan, Suting Xu, Xinran Xu, Guokun Lai, Yanru Chen, Huabin Zheng, Junjie Yan, Jianlin Su, Yuxin Wu, Neo Y. Zhang, Zhilin Yang, Xinyu Zhou, Mingxing Zhang, Jiezhong Qiu, 2025 https://scholar.google.com/scholar?q=MoBA%3A+Mixture+of+Block+Attention+for+Long-Context+LLMs 8. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Re, 2022 https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness 9. Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models — Xiang Hu, Zhanchao Zhou, Ruiqi Liang, Zehuan Li, Wei Wu, Jianguo Li, 2025 https://scholar.google.com/scholar?q=Every+Token+Counts%3A+Generalizing+16M+Ultra-Long+Context+in+Large+Language+Models 10. Hyperattention: Long-context attention in near-linear time — not verified from snippet, recent (not verified from snippet) https://scholar.google.com/scholar?q=Hyperattention%3A+Long-context+attention+in+near-linear+time 11. When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models — not verified from snippet, recent (not verified from snippet) https://scholar.google.com/scholar?q=When+Does+Content-Based+Routing+Work%3F+Representation+Requirements+for+Selective+Attention+in+Hybrid+Sequence+Models 12. Delta attention: Fast and accurate sparse attention inference by delta correction — not verified from snippet, recent (not verified from snippet) https://scholar.google.com/scholar?q=Delta+attention%3A+Fast+and+accurate+sparse+attention+inference+by+delta+correction 13. Spargeattention: Accurate and training-free sparse attention accelerating any model inference — not verified from snippet, recent (not verified from snippet) https://scholar.google.com/scholar?q=Spargeattention%3A+Accurate+and+training-free+sparse+attention+accelerating+any+model+inference 14. Post-training sparse attention with double sparsity — not verified from snippet, recent (not verified from snippet) https://scholar.google.com/scholar?q=Post-training+sparse+attention+with+double+sparsity 15. AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-07-memory-sparse-attention-for-100m-token-s-377cff.mp3 16. AI Post Transformers: RetrievalAttention for Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-17-retrievalattention-for-long-context-llm-ddf566.mp3 17. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3 18. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3 19. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3 20. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3 Interactive Visualization: Long Context Pre-Training with Lighthouse Attention

Episode metadata supplied by the publisher feed · Published May 13, 2026

Embed this episode

NOW PLAYING

Long Context Pre-Training with Lighthouse Attention

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on May 13, 2026.

Is there a transcript available for this episode?

Yes, a full transcript is available for this episode. You can read the complete transcript on the episode page.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!