An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Cut an image into 16×16 patches, call each a word, feed a plain Transformer.
Dosovitskiy et al. · ICLR 2021 · Foundations. Read the paper ↗
A free, interactive, animated visual explainer of An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — every exhibit computed from the real formulas, with verbatim quotes from the source.
Questions
- What is An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale?
- Cut an image into 16×16 patches, call each a word, feed a plain Transformer.
- Who published An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, and where?
- Dosovitskiy et al. — ICLR 2021 (arXiv:2010.11929).
- Where can I find a visual explainer of An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale?
- Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.
Related explainers
- Attention Is All You Need
- GPT-3: Language Models are Few-Shot Learners
- Mixtral of Experts
- Training Compute-Optimal Large Language Models
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- BERT: Pre-training of Deep Bidirectional Transformers
- Scaling Laws for Neural Language Models
- Adam: A Method for Stochastic Optimization