EPISODE · Apr 26, 2026 · 12 MIN
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
from Mastering Language Models: From Architecture to Optimization
Episode six of Topic 3 stops spreading work across machines and crawls inside a single GPU. FlashAttention's accusation is that attention was slow for the wrong reason: not too much math, too much traffic — the quadratic score matrix hauled back and forth between high-bandwidth memory and the tiny on-chip SRAM beside the compute units. Maya and Leo open at a laundromat where the machines were never the problem, walk through tiling and the online-softmax running tally that keeps blockwise attention mathematically exact, then stage the field's real fight: approximate the asymptote or engineer the exact computation's route. The stopwatch settles it — end-to-end wall-clock wins over approximate methods that cut FLOPs but not traffic — before the honest concession that the square law survives, and the closing diagnostic: after a faster kernel, which bottleneck inherited the crown? Sources: • FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness: https://arxiv.org/pdf/2205.14135
Embed this episode
NOW PLAYING
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.