EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
Draft one layer down: autoregress on features, not tokens — 2.7–3.5× faster, losslessly.
Li et al. · ICML 2024 · Serving. Read the paper ↗
A free, interactive, animated visual explainer of EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty — every exhibit computed from the real formulas, with verbatim quotes from the source.
Questions
- What is EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty?
- Draft one layer down: autoregress on features, not tokens — 2.7–3.5× faster, losslessly.
- Who published EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty, and where?
- Li et al. — ICML 2024 (arXiv:2401.15077).
- Where can I find a visual explainer of EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty?
- Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.
Related explainers
- PagedAttention (vLLM)
- Efficiently Scaling Transformer Inference
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
- Fast Inference from Transformers via Speculative Decoding
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
- CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
- Efficient Streaming Language Models with Attention Sinks