DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Split a request's timeline into prefill and decode GPU pools — 4.48x more requests under SLO.

Zhong et al. · OSDI 2024 · Serving. Read the paper ↗

A free, interactive, animated visual explainer of DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — every exhibit computed from the real formulas, with verbatim quotes from the source.

Questions

What is DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving?
Split a request's timeline into prefill and decode GPU pools — 4.48x more requests under SLO.
Who published DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving, and where?
Zhong et al. — OSDI 2024 (arXiv:2401.09670).
Where can I find a visual explainer of DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving?
Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.

Related explainers