Do Transformers Need Three Projections? episode artwork

EPISODE · Jun 11, 2026

Do Transformers Need Three Projections?

from AI Post Transformers

This episode explores whether transformers really need separate query, key, and value projections, treating the problem as weight tying inside attention rather than as a brand-new model design. It explains why KV-cache size and memory bandwidth are major bottlenecks for long-context, on-device decoding, then compares increasingly aggressive sharing schemes, especially the difference between tying keys and values versus tying queries and keys. The discussion emphasizes that the broader sweep happens at 300M parameters, while only the shared-K/V variant is carried to 1.2B scale and remains in contention against practical baselines like grouped-query and multi-query attention. Listeners get a concrete deployment tradeoff: shared K/V can reduce KV-cache memory by about 50 percent at roughly a 3.1 percent perplexity cost, making the episode especially interesting for anyone focused on efficient inference. Sources: 1. Do Transformers Need Three Projections? Systematic Study of QKV Variants — Ali Kayyam, Anusha Madan Gopal, M Anthony Lewis, 2026 http://arxiv.org/abs/2606.04032v2 2. Attention Is All You Need — Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, 2017 https://arxiv.org/abs/1706.03762 3. Fast Transformer Decoding: One Write-Head is All You Need — Noam Shazeer, 2019 https://arxiv.org/abs/1911.02150 4. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, Sumit Sanghai, 2023 https://arxiv.org/abs/2305.13245 5. Do Transformers Need Three Projections? Systematic Study of QKV Variants — Ali Kayyam, Anusha Madan Gopal, M. Anthony Lewis, 2026 https://arxiv.org/abs/2606.04032 6. Using the Output Embedding to Improve Language Models — Ofir Press, Lior Wolf, 2017 https://arxiv.org/abs/1608.05859 7. Linformer: Self-Attention with Linear Complexity — Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, Hao Ma, 2020 https://scholar.google.com/scholar?q=Linformer%3A+Self-Attention+with+Linear+Complexity 8. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention — Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, Francois Fleuret, 2020 https://scholar.google.com/scholar?q=Transformers+are+RNNs%3A+Fast+Autoregressive+Transformers+with+Linear+Attention 9. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023 https://scholar.google.com/scholar?q=Mamba%3A+Linear-Time+Sequence+Modeling+with+Selective+State+Spaces 10. AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations — Qian Tao et al., 2024 https://arxiv.org/abs/2410.13212 11. LongHeads: Multi-Head Attention is Secretly a Long Context Processor — Yi Lu et al., 2024 https://arxiv.org/abs/2402.10685 12. MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads — Weihao Liu et al., 2025 https://arxiv.org/abs/2502.13963 13. Squeezed Attention: Accelerating Long Context Length LLM Inference — Coleman Hooper et al., 2024 https://arxiv.org/abs/2411.09688 14. Beyond Uniform Query Distribution: Key-Driven Grouped Query Attention — Zohaib Khan et al., 2024 https://arxiv.org/abs/2408.08454 15. Weight Decay Induces Low-Rank Attention Layers — Seijin Kobayashi et al., 2024 https://arxiv.org/abs/2410.23819 16. Dissecting Query-Key Interaction in Vision Transformers — Xu Pan et al., 2024 https://arxiv.org/abs/2405.14880 17. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3 18. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3 19. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3 20. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3 21. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026 https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3 Interactive Visualization: Do Transformers Need Three Projections?

Episode metadata supplied by the publisher feed · Published Jun 11, 2026

Embed this episode

NOW PLAYING

Do Transformers Need Three Projections?

0:00 0:00

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

When was this AI Post Transformers episode published?

This episode was published on June 11, 2026.

Can I download this AI Post Transformers episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!