FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Re-cutting the same tiles across warps — ~2× faster by fixing the work partition, not the math.

Dao · ICLR 2024 · Kernels. Read the paper ↗

A free, interactive, animated visual explainer of FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning — every exhibit computed from the real formulas, with verbatim quotes from the source.

Questions

What is FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning?
Re-cutting the same tiles across warps — ~2× faster by fixing the work partition, not the math.
Who published FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning, and where?
Dao — ICLR 2024 (arXiv:2307.08691).
Where can I find a visual explainer of FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning?
Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.

Related explainers