BERT: Pre-training of Deep Bidirectional Transformers
Read the whole sentence at once — pre-train by filling in the blanks, then fine-tune anywhere.
Devlin et al. · NAACL 2019 · Foundations. Read the paper ↗
A free, interactive, animated visual explainer of BERT: Pre-training of Deep Bidirectional Transformers — every exhibit computed from the real formulas, with verbatim quotes from the source.
Questions
- What is BERT: Pre-training of Deep Bidirectional Transformers?
- Read the whole sentence at once — pre-train by filling in the blanks, then fine-tune anywhere.
- Who published BERT: Pre-training of Deep Bidirectional Transformers, and where?
- Devlin et al. — NAACL 2019 (arXiv:1810.04805).
- Where can I find a visual explainer of BERT: Pre-training of Deep Bidirectional Transformers?
- Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.
Related explainers
- Attention Is All You Need
- GPT-3: Language Models are Few-Shot Learners
- Mixtral of Experts
- Training Compute-Optimal Large Language Models
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Scaling Laws for Neural Language Models
- Adam: A Method for Stochastic Optimization
- Deep Residual Learning for Image Recognition