MiniMax-01
Near-linear attention at 456B — lightning attention, with a softmax layer every eighth block.
MiniMax · 2025 · Model Architectures. Read the paper ↗
A free, interactive, animated visual explainer of MiniMax-01 — every exhibit computed from the real formulas, with verbatim quotes from the source.
Questions
- What is MiniMax-01?
- Near-linear attention at 456B — lightning attention, with a softmax layer every eighth block.
- Who published MiniMax-01, and where?
- MiniMax — 2025 (arXiv:2501.08313).
- Where can I find a visual explainer of MiniMax-01?
- Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.
Related explainers
- DeepSeek-V3
- Qwen3
- OLMo 2
- Gemma 4
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Scalable Diffusion Models with Transformers
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints