The Transformer, end to end
Build the Transformer from its core operation outward: attention and position, the encoder/decoder lineage, the efficiency tricks that let it scale, the state-space challengers, and the laws that govern all of it.
- Attention Is All You Need
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- BERT: Pre-training of Deep Bidirectional Transformers
- GPT-3: Language Models are Few-Shot Learners
- FlashAttention
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
- Differential Transformer
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Mixtral of Experts
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- RWKV: Reinventing RNNs for the Transformer Era
- Scaling Laws for Neural Language Models
- Training Compute-Optimal Large Language Models
- Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
- Titans: Learning to Memorize at Test Time
- Byte Latent Transformer: Patches Scale Better Than Tokens
- Efficiently Modeling Long Sequences with Structured State Spaces
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer