FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Re-cutting the same tiles across warps — ~2× faster by fixing the work partition, not the math.
Dao · ICLR 2024 · Kernels. Read the paper ↗
A free, interactive, animated visual explainer of FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning — every exhibit computed from the real formulas, with verbatim quotes from the source.
Questions
- What is FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning?
- Re-cutting the same tiles across warps — ~2× faster by fixing the work partition, not the math.
- Who published FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning, and where?
- Dao — ICLR 2024 (arXiv:2307.08691).
- Where can I find a visual explainer of FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning?
- Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.
Related explainers
- FlashAttention
- GSPMD: General and Scalable Parallelization for ML Computation Graphs
- Ring Attention with Blockwise Transformers for Near-Infinite Context
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- Differential Transformer
- MiniMax-M1: Scaling Test-Time Compute with Lightning Attention