Youhui Bai
Associate Researcher, University of Science and Technology of China
Youhui Bai received his Ph.D. in Computer Architecture from the University of Science and Technology of China (USTC) in 2021. After graduation, he joined Huawei’s 2012 Laboratories as a Principal Engineer and returned to USTC in 2025, where he is currently an Associate Researcher in the School of Computer Science and Technology. His research focuses on distributed parallel computing for large language models, with an emphasis on addressing data movement, communication, memory, and sparse computing bottlenecks in LLM training and inference. He has published 12 papers in top-tier international conferences and journals, including SOSP, OSDI, CVPR, AAAI, HPCA, and TPDS, with 11 CCF-A papers and 7 papers as first or corresponding author. He has filed more than 10 national invention patents and one major technical secret. His work has received several honors, including the 2026 USENIX OSDI Jay Lepreau Best Paper Award (the first recipient from China), the 2024 World Artificial Intelligence Conference Youth Outstanding Paper Award (one of 10 papers selected worldwide and the only systems paper), and the 2023 ACM SIGCSE Outstanding Doctoral Dissertation Award (one of two recipients in China). He has led projects funded by the National Natural Science Foundation of China and industry partners including ByteDance and Tencent, and has served as a key technical contributor to major national science and technology projects and Hefei municipal innovation programs.
Topic
Sparse Attention-Driven KVCache Management for Domestic AI Supernodes
In long-context inference, persistent KVCache occupancy in High-Bandwidth Memory (HBM) limits service concurrency and decoding throughput. This talk presents a sparse attention-driven KVCache management system designed for domestic AI supernodes. The system uses Top-K indices to enable on-demand, token-level data access. Complete historical KV data is stored in a host memory pool within the supernode, while hot data is cached on the NPU and the data required for the current computation is assembled on demand. An index-based copy operator enables compatibility with graph execution on NPUs. The talk also shares integration experience with SGLang, including a redesign of the supernode-based KVCache memory pool, and discusses approaches to capacity scaling and performance optimization for long-context inference. Agenda Bottlenecks in Long-Context Inference: Analyze how KVCache capacity limits decoding concurrency, and the mismatch between page-level data movement and token-level sparse access. Sparse-Driven System Design: Introduce the host storage, device cache, and index mapping of SparseKVCacheManager, and explain how the UniDex index-copy operator enables compatibility with NPU graph execution. SGLang Integration and Performance Analysis: Share integration experience on domestic NPUs, analyze the relationship between HBM savings, larger decoding batches, and throughput improvements, and examine the overhead of on-demand data movement. Scaling for Domestic AI Supernodes: Present validation of remote DRAM sparse access based on HCCS and MemFabric, and discuss shared KV memory pool design and memory utilization in scenarios with separated prefill and decoding. Key Takeaways Understand the capacity and bandwidth bottlenecks of KVCache in long-context inference, and learn how sparse access can reduce device-resident data and increase service concurrency. Learn how Top-K index-driven, token-level KV management and data movement work, as well as engineering approaches for adapting dynamic sparse access to NPU graph execution. Understand the architectural considerations behind a shared KV memory pool for domestic AI supernodes, including how high-speed interconnects, remote DRAM access, and cross-node index addressing can work together to scale available memory capacity. Learn from SGLang integration and performance analysis how to evaluate the applicability of offloading solutions based on HBM utilization, data-transfer overhead, decoding throughput, and request latency.