Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
A fixed compute budget, spent unevenly — tokens route around blocks they don't need.
Raposo et al. · arXiv 2024 · Model Architectures. Read the paper ↗
A free, interactive, animated visual explainer of Mixture-of-Depths: Dynamically allocating compute in transformer-based language models — every exhibit computed from the real formulas, with verbatim quotes from the source.
Questions
- What is Mixture-of-Depths: Dynamically allocating compute in transformer-based language models?
- A fixed compute budget, spent unevenly — tokens route around blocks they don't need.
- Who published Mixture-of-Depths: Dynamically allocating compute in transformer-based language models, and where?
- Raposo et al. — arXiv 2024 (arXiv:2404.02258).
- Where can I find a visual explainer of Mixture-of-Depths: Dynamically allocating compute in transformer-based language models?
- Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.