Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Selective SSMs and masked attention are one structured matrix, computed two ways.

Dao & Gu · ICML 2024 · Model Architectures. Read the paper ↗

A free, interactive, animated visual explainer of Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality — every exhibit computed from the real formulas, with verbatim quotes from the source.

Questions

What is Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality?
Selective SSMs and masked attention are one structured matrix, computed two ways.
Who published Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, and where?
Dao & Gu — ICML 2024 (arXiv:2405.21060).
Where can I find a visual explainer of Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality?
Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.

Related explainers