ECHO: Efficient KV Cache Offloading with Lossless Prefetching for Serving Native Sparse Attention LLMs

Guangda Liu, Wenhao Chen, Chengwei Li, and Zhenyu Ning, Shanghai Jiao Tong University; Jing Lin and Yiwu Yao, Huawei; Quan Chen, Shixuan Sun, and Jieru Zhao, Shanghai Jiao Tong University; Minyi Guo, Guizhou University and Shanghai Jiao Tong University

Native sparse attention has emerged as a promising approach for efficient long-context LLM inference without compromising accuracy. While it significantly reduces the attention computation and KV cache access costs, the KV cache size exhibits steeper linear growth with the context length. As a result, GPU HBM capacity becomes the bottleneck, limiting the concurrency of long-context requests and leading to poor hardware utilization and low generation throughput. We introduce ECHO, a serving system designed for native sparse-attention LLMs that employs KV cache offloading to overcome GPU HBM capacity limits. ECHO incorporates a graph-friendly cache manager that enables efficient dynamic KV cache eviction and recall entirely within GPU graphs, minimizing management overhead. Furthermore, by exploiting the numerical predictability of index scores and the sequential processing of queries, ECHO enables lossless intra-query prefetching for decoding and inter-query prefetching for prefill. By applying a fully pipelined fused GPU kernel, ECHO overlaps the recall overhead with indexer computation. Experiments show that ECHO delivers up to 2.1× higher generation throughput than state-of-the-art systems such as SGLang and vLLM under long-context workloads, while maintaining comparable latency under light load.

OSDI '26 Open Access Sponsored by
King Abdullah University of Science and Technology (KAUST)

Open Access Media

USENIX is committed to Open Access to the research presented at our events. Papers and proceedings are freely available to everyone once the event begins. Any video, audio, and/or slides that are posted after the event are also free and open to everyone. Support USENIX and our commitment to Open Access.

BibTeX
@inproceedings {318385,
author = {Guangda Liu and Wenhao Chen and Chengwei Li and Zhenyu Ning and Jing Lin and Yiwu Yao and Quan Chen and Shixuan Sun and Jieru Zhao and Minyi Guo},
title = {{ECHO}: Efficient {KV} Cache Offloading with Lossless Prefetching for Serving Native Sparse Attention {LLMs}},
booktitle = {20th USENIX Symposium on Operating Systems Design and Implementation (OSDI 26)},
year = {2026},
isbn = {978-1-939133-55-7},
address = {Seattle, WA},
pages = {17--37},
url = {https://www.usenix.org/conference/osdi26/presentation/liu-guangda},
publisher = {USENIX Association},
month = jul
}