Training Compute-Optimal Large Language Models

Given a fixed compute budget, double the model and double the data — in equal proportion.

Hoffmann et al. · NeurIPS 2022 · Foundations. Read the paper ↗

A free, interactive, animated visual explainer of Training Compute-Optimal Large Language Models — every exhibit computed from the real formulas, with verbatim quotes from the source.

Questions

What is Training Compute-Optimal Large Language Models?
Given a fixed compute budget, double the model and double the data — in equal proportion.
Who published Training Compute-Optimal Large Language Models, and where?
Hoffmann et al. — NeurIPS 2022 (arXiv:2203.15556).
Where can I find a visual explainer of Training Compute-Optimal Large Language Models?
Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.

Related explainers