Efficiently Scaling Transformer Inference

Chop a 540B model across a TPU pod: 29ms/token, 76% MFU, 32x longer context

Pope et al. · MLSys 2023 · Serving. Read the paper ↗

A free, interactive, animated visual explainer of Efficiently Scaling Transformer Inference — every exhibit computed from the real formulas, with verbatim quotes from the source.

Questions

What is Efficiently Scaling Transformer Inference?
Chop a 540B model across a TPU pod: 29ms/token, 76% MFU, 32x longer context
Who published Efficiently Scaling Transformer Inference, and where?
Pope et al. — MLSys 2023 (arXiv:2211.05102).
Where can I find a visual explainer of Efficiently Scaling Transformer Inference?
Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.

Related explainers