Efficient Streaming Language Models with Attention Sinks
Pin 4 "attention-sink" tokens + a rolling window — stream 4M tokens, no fine-tuning.
Xiao et al. · ICLR 2024 · Serving. Read the paper ↗
A free, interactive, animated visual explainer of Efficient Streaming Language Models with Attention Sinks — every exhibit computed from the real formulas, with verbatim quotes from the source.
Questions
- What is Efficient Streaming Language Models with Attention Sinks?
- Pin 4 "attention-sink" tokens + a rolling window — stream 4M tokens, no fine-tuning.
- Who published Efficient Streaming Language Models with Attention Sinks, and where?
- Xiao et al. — ICLR 2024 (arXiv:2309.17453).
- Where can I find a visual explainer of Efficient Streaming Language Models with Attention Sinks?
- Right here — a free, interactive, animated walkthrough of the whole paper, with exhibits computed from the real formulas and verbatim quotes from the source.
Related explainers
- PagedAttention (vLLM)
- Efficiently Scaling Transformer Inference
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
- Fast Inference from Transformers via Speculative Decoding
- EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving