Yapeng Jiang, Sun Yat-sen University, Peng Cheng Laboratory, and Zhuhai Key Laboratory of Trusted Large Language Models; Minghao Gan, Sun Yat-sen University; Zicong Hong, École Polytechnique Fédérale de Lausanne; Wuhui Chen, Sun Yat-sen University and Peng Cheng Laboratory; Junyuan Liang, Sun Yat-sen University; Yue Yu, Peng Cheng Laboratory; Meng Guo, Qilu University of Technology; Zibin Zheng, Sun Yat-sen University
Hybrid LLM inference systems exploit both GPU and CPU for computation and memory but are bottlenecked by the CPU’s lower computational capabilities. Recent approaches leverage activation sparsity by offline partitioning FFN neurons into "hot" (frequently activated) and "cold" (rarely activated) sets. This approach retains critical computations on the GPU, yet static partitioning struggles to adapt to runtime activation changes, leading to suboptimal throughput.
We present Kairox, an adaptive GPU-CPU hybrid inference system that addresses these limitations through online neuron balancing, a mechanism that dynamically redistributes neurons between the GPU and CPU based on activation patterns. To realize this, Kairox introduces a Live Pipeline designed to prefetch neurons by predicting next-layer activation patterns. Furthermore, leveraging activation locality, we develop a Temporal Activation Momentum cache policy to prioritize neurons with sustained utility while minimizing transient, wasteful transfers. Finally, an Adaptive Neuron Balancer modulates the balancing intensity according to runtime resource conditions, maintaining an optimal equilibrium between competing system bottlenecks. For standard completion on consumer-grade PCs, Kairox improves end-to-end throughput by up to 7.57×, 3.70×, 6.35×, and 3.76× over llama.cpp, PowerInfer, Neuralink, and Q-Infer, respectively. Across all evaluated settings, it achieves geomean speedups of 3.15× and 3.93× over llama.cpp on two representative PCs and around 2.1× over the three sparse baselines.
OSDI '26 Open Access Sponsored by
King Abdullah University of Science and Technology (KAUST)
Open Access Media
USENIX is committed to Open Access to the research presented at our events. Papers and proceedings are freely available to everyone once the event begins. Any video, audio, and/or slides that are posted after the event are also free and open to everyone. Support USENIX and our commitment to Open Access.

author = {Yapeng Jiang and Minghao Gan and Zicong Hong and Wuhui Chen and Junyuan Liang and Yue Yu and Meng Guo and Zibin Zheng},
title = {Kairox: Adaptive {GPU-CPU} Hybrid {LLM} Inference via Online Neuron Balancing},
booktitle = {20th USENIX Symposium on Operating Systems Design and Implementation (OSDI 26)},
year = {2026},
isbn = {978-1-939133-55-7},
address = {Seattle, WA},
pages = {1839--1856},
url = {https://www.usenix.org/conference/osdi26/presentation/jiang-yapeng},
publisher = {USENIX Association},
month = jul
}
