DeepSeek-V3
A 671B mixture-of-experts that activates only 37B — via latent-KV attention and loss-free routing.
DeepSeek-AI · 2024 · Model Architectures. Read the paper ↗
A free, interactive, animated visual explainer of DeepSeek-V3 — every exhibit computed from the real formulas, with verbatim quotes from the source.
Questions
- What is DeepSeek-V3?
- A 671B mixture-of-experts that activates only 37B — via latent-KV attention and loss-free routing.
- Who published DeepSeek-V3, and where?
- DeepSeek-AI — 2024 (arXiv:2412.19437).
- Where can I find a visual explainer of DeepSeek-V3?
- Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.
Related explainers
- Qwen3
- OLMo 2
- MiniMax-01
- Gemma 4
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Scalable Diffusion Models with Transformers
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints