FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Rebuilds attention for Hopper — async warps + FP8 — for 740 TFLOPs/s, 1.5-2.0x over FA-2.

Shah et al. · NeurIPS 2024 · Kernels. Read the paper ↗

A free, interactive, animated visual explainer of FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision — every exhibit computed from the real formulas, with verbatim quotes from the source.

Questions

What is FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision?
Rebuilds attention for Hopper — async warps + FP8 — for 740 TFLOPs/s, 1.5-2.0x over FA-2.
Who published FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision, and where?
Shah et al. — NeurIPS 2024 (arXiv:2407.08608).
Where can I find a visual explainer of FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision?
Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.

Related explainers