DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

236B MoE, 21B active per token — MLA folds the whole KV cache into one latent vector

DeepSeek-AI · arXiv 2024 · Model Architectures. Read the paper ↗

A free, interactive, animated visual explainer of DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — every exhibit computed from the real formulas, with verbatim quotes from the source.

Questions

What is DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model?
236B MoE, 21B active per token — MLA folds the whole KV cache into one latent vector
Who published DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model, and where?
DeepSeek-AI — arXiv 2024 (arXiv:2405.04434).
Where can I find a visual explainer of DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model?
Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.

Related explainers