The DeepSeek lineage
Follow a single frontier lab's innovations in order: the math-RL algorithm they invented, the attention and MoE efficiency tricks, and the reasoning model those choices enabled.
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
- DeepSeek-V3
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- DeepSeek-OCR: Contexts Optical Compression
- DeepSeek-R1