DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
236B MoE, 21B active per token — MLA folds the whole KV cache into one latent vector
DeepSeek-AI · arXiv 2024 · Model Architectures. Read the paper ↗
A free, interactive, animated visual explainer of DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — every exhibit computed from the real formulas, with verbatim quotes from the source.
Questions
- What is DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model?
- 236B MoE, 21B active per token — MLA folds the whole KV cache into one latent vector
- Who published DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model, and where?
- DeepSeek-AI — arXiv 2024 (arXiv:2405.04434).
- Where can I find a visual explainer of DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model?
- Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.