FlashAttention
Exact attention, made fast by never writing the big matrix to memory.
Dao et al. · NeurIPS 2022 · Kernels. Read the paper ↗
A free, interactive, animated visual explainer of FlashAttention — every exhibit computed from the real formulas, with verbatim quotes from the source.
Questions
- What is FlashAttention?
- Exact attention, made fast by never writing the big matrix to memory.
- Who published FlashAttention, and where?
- Dao et al. — NeurIPS 2022 (arXiv:2205.14135).
- Where can I find a visual explainer of FlashAttention?
- Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.
Related explainers
- Attention Is All You Need
- Ring Attention with Blockwise Transformers for Near-Infinite Context
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling