{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T19:06:34Z","timestamp":1779131194816,"version":"3.51.4"},"reference-count":63,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62325205"],"award-info":[{"award-number":["62325205"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272215"],"award-info":[{"award-number":["62272215"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62502193"],"award-info":[{"award-number":["62502193"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100004608","name":"Natural Science Foundation of Jiangsu Province","doi-asserted-by":"crossref","award":["BK20243053"],"award-info":[{"award-number":["BK20243053"]}],"id":[{"id":"10.13039\/501100004608","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,5,18]]},"abstract":"<jats:p>\n                    Prefix-sharing among multiple prompts presents opportunities to combine the operations of the shared prefix, while attention computation in the decode stage, which becomes a critical bottleneck with increasing context lengths, is a memory-intensive process requiring heavy memory access on the key-value (KV) cache of the prefixes. Therefore, in this paper, we explore the potential of prefix-sharing in the attention computation of the decode stage. However, the tree structure of the prefix-sharing mechanism presents significant challenges for attention computation in efficiently processing shared KV cache access patterns while managing complex dependencies and balancing irregular workloads. To address the above challenges, we propose a dedicated attention kernel to\n                    <jats:underline>co<\/jats:underline>\n                    mbine the memory access of shared prefixes in the\n                    <jats:underline>dec<\/jats:underline>\n                    oding stage, namely CoDec. CoDec delivers two key innovations: a novel shared-prefix attention kernel that optimizes memory hierarchy and exploits both intra-block and inter-block parallelism, and a comprehensive workload balancing mechanism that efficiently estimates cost, divides tasks, and schedules execution. Experimental results show that CoDec achieves an average 1.9\u00d7 speedup and 120.9\u00d7 memory access reduction compared to the state-of-the-art FlashDecoding kernel regarding attention computation in the decode stage and 3.8\u00d7 end-to-end time per output token compared to vLLM.\n                  <\/jats:p>","DOI":"10.1145\/3802028","type":"journal-article","created":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T18:19:16Z","timestamp":1779128356000},"page":"1-27","source":"Crossref","is-referenced-by-count":0,"title":["CoDec: Prefix-Shared Decoding Kernel for LLMs"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9204-4075","authenticated-orcid":false,"given":"Zhibin","family":"Wang","sequence":"first","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-6516-5067","authenticated-orcid":false,"given":"Rui","family":"Ning","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3430-1189","authenticated-orcid":false,"given":"Chao","family":"Fang","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6933-7375","authenticated-orcid":false,"given":"Zhonghui","family":"Zhang","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2284-6842","authenticated-orcid":false,"given":"Xi","family":"Lin","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-6173-3983","authenticated-orcid":false,"given":"Shaobo","family":"Ma","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-5551-0765","authenticated-orcid":false,"given":"Mo","family":"Zhou","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5713-7225","authenticated-orcid":false,"given":"Xue","family":"Li","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7227-4786","authenticated-orcid":false,"given":"Zhongfeng","family":"Wang","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3154-3580","authenticated-orcid":false,"given":"Chengying","family":"Huan","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1565-9997","authenticated-orcid":false,"given":"Rong","family":"Gu","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6782-6689","authenticated-orcid":false,"given":"Kun","family":"Yang","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6934-1685","authenticated-orcid":false,"given":"Guihai","family":"Chen","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6581-8730","authenticated-orcid":false,"given":"Sheng","family":"Zhong","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2710-7628","authenticated-orcid":false,"given":"Chen","family":"Tian","sequence":"additional","affiliation":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,5,18]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2020. NVIDIA A100 Tensor Core GPU. https:\/\/www.nvidia.com\/en-us\/data-center\/a100\/. [Accessed 15-04-2025]."},{"key":"e_1_2_1_2_1","unstructured":"2024. meta-llama\/Llama-3.1-8B. https:\/\/huggingface.co\/meta-llama\/Llama-3.1-8B. [Accessed 15-04-2025]."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3725273"},{"key":"e_1_2_1_4_1","volume-title":"SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills. arXiv:2308.16369 [cs.LG] https:\/\/arxiv.org\/abs\/2308.16369","author":"Agrawal Amey","year":"2023","unstructured":"Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, and Ramachandran Ramjee. 2023. SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills. arXiv:2308.16369 [cs.LG] https:\/\/arxiv.org\/abs\/2308.16369"},{"key":"e_1_2_1_5_1","volume-title":"GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. arXiv:2305.13245 [cs.CL] https:\/\/arxiv.org\/abs\/2305.13245","author":"Ainslie Joshua","year":"2023","unstructured":"Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebr\u00f3n, and Sumit Sanghai. 2023. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. arXiv:2305.13245 [cs.CL] https:\/\/arxiv.org\/abs\/2305.13245"},{"key":"e_1_2_1_6_1","unstructured":"Tom B. Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell Sandhini Agarwal Ariel Herbert-Voss Gretchen Krueger Tom Henighan Rewon Child Aditya Ramesh Daniel M. Ziegler Jeffrey Wu Clemens Winter Christopher Hesse Mark Chen Eric Sigler Mateusz Litwin Scott Gray Benjamin Chess Jack Clark Christopher Berner Sam McCandlish Alec Radford Ilya Sutskever and Dario Amodei. 2020. Language Models are Few-Shot Learners. arXiv:2005.14165 [cs.CL] https: \/\/arxiv.org\/abs\/2005.14165"},{"key":"e_1_2_1_7_1","volume-title":"International Conference on Learning Representations (ICLR).","author":"Dao Tri","year":"2024","unstructured":"Tri Dao. 2024. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning. In International Conference on Learning Representations (ICLR)."},{"key":"e_1_2_1_8_1","unstructured":"Tri Dao Daniel Haziza Francisco Massa and Grigory Sizov. 2023. Flash-Decoding for long-context inference. https:\/\/crfm.stanford.edu\/2023\/10\/12\/flashdecoding.html. [Accessed 08-04-2025]."},{"key":"e_1_2_1_9_1","unstructured":"DeepSeek. 2024. DeepSeek-R1-Lite-Preview is now live: unleashing supercharged reasoning power. https:\/\/api-docs.deepseek.com\/news\/news1120."},{"key":"e_1_2_1_10_1","unstructured":"DeepSeek-AI Aixin Liu Bei Feng Bin Wang Bingxuan Wang Bo Liu Chenggang Zhao Chengqi Dengr Chong Ruan Damai Dai Daya Guo Dejian Yang Deli Chen Dongjie Ji Erhang Li Fangyun Lin Fuli Luo Guangbo Hao Guanting Chen Guowei Li H. Zhang Hanwei Xu Hao Yang Haowei Zhang Honghui Ding Huajian Xin Huazuo Gao Hui Li Hui Qu J. L. Cai Jian Liang Jianzhong Guo Jiaqi Ni Jiashi Li Jin Chen Jingyang Yuan Junjie Qiu Junxiao Song Kai Dong Kaige Gao Kang Guan Lean Wang Lecong Zhang Lei Xu Leyi Xia Liang Zhao Liyue Zhang Meng Li Miaojun Wang Mingchuan Zhang Minghua Zhang Minghui Tang Mingming Li Ning Tian Panpan Huang Peiyi Wang Peng Zhang Qihao Zhu Qinyu Chen Qiushi Du R. J. Chen R. L. Jin Ruiqi Ge Ruizhe Pan Runxin Xu Ruyi Chen S. S. Li Shanghao Lu Shangyan Zhou Shanhuang Chen Shaoqing Wu Shengfeng Ye Shirong Ma Shiyu Wang Shuang Zhou Shuiping Yu Shunfeng Zhou Size Zheng T. Wang Tian Pei Tian Yuan Tianyu Sun W. L. Xiao Wangding Zeng Wei An Wen Liu Wenfeng Liang Wenjun Gao Wentao Zhang X. Q. Li Xiangyue Jin Xianzu Wang Xiao Bi Xiaodong Liu Xiaohan Wang Xiaojin Shen Xiaokang Chen Xiaosha Chen Xiaotao Nie Xiaowen Sun Xiaoxiang Wang Xin Liu Xin Xie Xingkai Yu Xinnan Song Xinyi Zhou Xinyu Yang Xuan Lu Xuecheng Su Y. Wu Y. K. Li Y. X. Wei Y. X. Zhu Yanhong Xu Yanping Huang Yao Li Yao Zhao Yaofeng Sun Yaohui Li Yaohui Wang Yi Zheng Yichao Zhang Yiliang Xiong Yilong Zhao Ying He Ying Tang Yishi Piao Yixin Dong Yixuan Tan Yiyuan Liu Yongji Wang Yongqiang Guo Yuchen Zhu Yuduan Wang Yuheng Zou Yukun Zha Yunxian Ma Yuting Yan Yuxiang You Yuxuan Liu Z. Z. Ren Zehui Ren Zhangli Sha Zhe Fu Zhen Huang Zhen Zhang Zhenda Xie Zhewen Hao Zhihong Shao Zhiniu Wen Zhipeng Xu Zhongyu Zhang Zhuoshu Li Zihan Wang Zihui Gu Zilin Li and Ziwei Xie. 2024. DeepSeek-V2: A Strong Economical and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434 [cs.CL] https:\/\/arxiv.org\/abs\/2405.04434"},{"key":"e_1_2_1_11_1","unstructured":"DeepSeek-AI Aixin Liu Bei Feng Bing Xue Bingxuan Wang Bochao Wu Chengda Lu Chenggang Zhao Chengqi Deng Chenyu Zhang Chong Ruan Damai Dai Daya Guo Dejian Yang Deli Chen Dongjie Ji Erhang Li Fangyun Lin Fucong Dai Fuli Luo Guangbo Hao Guanting Chen Guowei Li H. Zhang Han Bao Hanwei Xu Haocheng Wang Haowei Zhang Honghui Ding Huajian Xin Huazuo Gao Hui Li Hui Qu J. L. Cai Jian Liang Jianzhong Guo Jiaqi Ni Jiashi Li Jiawei Wang Jin Chen Jingchang Chen Jingyang Yuan Junjie Qiu Junlong Li Junxiao Song Kai Dong Kai Hu Kaige Gao Kang Guan Kexin Huang Kuai Yu Lean Wang Lecong Zhang Lei Xu Leyi Xia Liang Zhao Litong Wang Liyue Zhang Meng Li Miaojun Wang Mingchuan Zhang Minghua Zhang Minghui Tang Mingming Li Ning Tian Panpan Huang Peiyi Wang Peng Zhang Qiancheng Wang Qihao Zhu Qinyu Chen Qiushi Du R. J. Chen R. L. Jin Ruiqi Ge Ruisong Zhang Ruizhe Pan Runji Wang Runxin Xu Ruoyu Zhang Ruyi Chen S. S. Li Shanghao Lu Shangyan Zhou Shanhuang Chen Shaoqing Wu Shengfeng Ye Shengfeng Ye Shirong Ma Shiyu Wang Shuang Zhou Shuiping Yu Shunfeng Zhou Shuting Pan T. Wang Tao Yun Tian Pei Tianyu Sun W. L. Xiao Wangding Zeng Wanjia Zhao Wei An Wen Liu Wenfeng Liang Wenjun Gao Wenqin Yu Wentao Zhang X. Q. Li Xiangyue Jin Xianzu Wang Xiao Bi Xiaodong Liu Xiaohan Wang Xiaojin Shen Xiaokang Chen Xiaokang Zhang Xiaosha Chen Xiaotao Nie Xiaowen Sun Xiaoxiang Wang Xin Cheng Xin Liu Xin Xie Xingchao Liu Xingkai Yu Xinnan Song Xinxia Shan Xinyi Zhou Xinyu Yang Xinyuan Li Xuecheng Su Xuheng Lin Y. K. Li Y. Q. Wang Y. X. Wei Y. X. Zhu Yang Zhang Yanhong Xu Yanhong Xu Yanping Huang Yao Li Yao Zhao Yaofeng Sun Yaohui Li Yaohui Wang Yi Yu Yi Zheng Yichao Zhang Yifan Shi Yiliang Xiong Ying He Ying Tang Yishi Piao Yisong Wang Yixuan Tan Yiyang Ma Yiyuan Liu Yongqiang Guo Yu Wu Yuan Ou Yuchen Zhu Yuduan Wang Yue Gong Yuheng Zou Yujia He Yukun Zha Yunfan Xiong Yunxian Ma Yuting Yan Yuxiang Luo Yuxiang You Yuxuan Liu Yuyang Zhou Z. F. Wu Z. Z. Ren Zehui Ren Zhangli Sha Zhe Fu Zhean Xu Zhen Huang Zhen Zhang Zhenda Xie Zhengyan Zhang Zhewen Hao Zhibin Gou Zhicheng Ma Zhigang Yan Zhihong Shao Zhipeng Xu Zhiyu Wu Zhongyu Zhang Zhuoshu Li Zihui Gu Zijia Zhu Zijun Liu Zilin Li Ziwei Xie Ziyang Song Ziyi Gao and Zizheng Pan. 2025. DeepSeek-V3 Technical Report. arXiv:2412.19437 [cs.CL] https:\/\/arxiv.org\/abs\/2412.19437"},{"key":"e_1_2_1_12_1","unstructured":"Dom Eccleston. 2023. ShareGPT. https:\/\/github.com\/domeccleston\/sharegpt."},{"key":"e_1_2_1_13_1","volume-title":"2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 146-1481","author":"Fang Chao","year":"2025","unstructured":"Chao Fang, Man Shi, Robin Geens, Arne Symons, Zhongfeng Wang, and Marian Verhelst. 2025. Anda: Unlocking efficient LLM inference with a variable-length grouped activation data format. In 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 146-1481."},{"key":"e_1_2_1_14_1","volume-title":"Break the sequential dependency of llm inference using lookahead decoding. arXiv preprint arXiv:2402.02057","author":"Fu Yichao","year":"2024","unstructured":"Yichao Fu, Peter Bailis, Ion Stoica, and Hao Zhang. 2024. Break the sequential dependency of llm inference using lookahead decoding. arXiv preprint arXiv:2402.02057 (2024)."},{"key":"e_1_2_1_15_1","unstructured":"Github. 2024. Accelerate your development speed with copilot. https:\/\/copilot.github.com."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1002\/j.1538-7305.1966.tb01709.x"},{"key":"e_1_2_1_17_1","unstructured":"Aaron Grattafiori Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle et al. 2024. The Llama 3 Herd of Models. arXiv:2407.21783 [cs.AI] https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_2_1_18_1","unstructured":"Zhicheng Guo Sijie Cheng Hao Wang Shihao Liang Yujia Qin Peng Li Zhiyuan Liu Maosong Sun and Yang Liu. 2024. StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models. arXiv:2403.07714 [cs.CL] https:\/\/arxiv.org\/abs\/2403.07714"},{"key":"e_1_2_1_19_1","volume-title":"Large Language Models are Zero-Shot Rankers for Recommender Systems. ArXiv abs\/2305.08845","author":"Hou Yupeng","year":"2023","unstructured":"Yupeng Hou, Junjie Zhang, Zihan Lin, Hongyu Lu, Ruobing Xie, Julian McAuley, and Wayne Xin Zhao. 2023. Large Language Models are Zero-Shot Rankers for Recommender Systems. ArXiv abs\/2305.08845 (2023). https:\/\/api.semanticscholar.org\/CorpusID:258686540"},{"key":"e_1_2_1_20_1","unstructured":"Cunchen Hu Heyang Huang Junhao Hu Jiang Xu Xusheng Chen Tao Xie Chenxi Wang Sa Wang Yungang Bao Ninghui Sun and Yizhou Shan. 2024. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool. arXiv:2406.17565 [cs.DC] https:\/\/arxiv.org\/abs\/2406.17565"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.307"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613165"},{"key":"e_1_2_1_23_1","unstructured":"Jiaqi Li Mengmeng Wang Zilong Zheng and Muhan Zhang. 2024. LooGLE: Can Long-Context Language Models Understand Long Contexts? arXiv:2311.04939 [cs.CL] https:\/\/arxiv.org\/abs\/2311.04939"},{"key":"e_1_2_1_24_1","volume-title":"Sequence Parallelism: Long Sequence Training from System Perspective. arXiv:2105.13120 [cs.LG] https:\/\/arxiv.org\/abs\/2105.13120","author":"Li Shenggui","year":"2022","unstructured":"Shenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li, and Yang You. 2022. Sequence Parallelism: Long Sequence Training from System Perspective. arXiv:2105.13120 [cs.LG] https:\/\/arxiv.org\/abs\/2105.13120"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3749168"},{"key":"e_1_2_1_26_1","unstructured":"Zachary C. Lipton John Berkowitz and Charles Elkan. 2015. A Critical Review of Recurrent Neural Networks for Sequence Learning. arXiv:1506.00019 [cs.LG] https:\/\/arxiv.org\/abs\/1506.00019"},{"key":"e_1_2_1_27_1","unstructured":"Junyu Luo Weizhi Zhang Ye Yuan Yusheng Zhao Junwei Yang Yiyang Gu Bohan Wu Binqi Chen Ziyue Qiao Qingqing Long Rongcheng Tu Xiao Luo Wei Ju Zhiping Xiao Yifan Wang Meng Xiao Chenwu Liu Jingyang Yuan Shichang Zhang Yiqiao Jin Fan Zhang Xian Wu Hanqing Zhao Dacheng Tao Philip S. Yu and Ming Zhang. 2025. Large Language Model Agent: A Survey on Methodology Applications and Challenges. arXiv:2503.21460 [cs.CL] https:\/\/arxiv.org\/abs\/2503.21460"},{"key":"e_1_2_1_28_1","volume-title":"APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration","author":"Ma Shaobo","year":"2025","unstructured":"Shaobo Ma, Chao Fang, Haikuo Shao, and Zhongfeng Wang. 2025. APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (2025)."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3620666.3651335"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3620665.3640383"},{"key":"e_1_2_1_31_1","unstructured":"OpenAI. 2024. ChatGPT. https:\/\/chat.openai.com. Accessed: 2024-08-09."},{"key":"e_1_2_1_32_1","unstructured":"OpenAI. 2024. Learning to reason with LLMs. https:\/\/openai.com\/index\/learning-to-reason-with-llms\/."},{"key":"e_1_2_1_33_1","unstructured":"Reiner Pope Sholto Douglas Aakanksha Chowdhery Jacob Devlin James Bradbury Anselm Levskaya Jonathan Heek Kefan Xiao Shivani Agrawal and Jeff Dean. 2022. Efficiently Scaling Transformer Inference. arXiv:2211.05102 [cs.LG] https:\/\/arxiv.org\/abs\/2211.05102"},{"key":"e_1_2_1_34_1","volume-title":"Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving. arXiv:2407.00079 [cs.DC] https:\/\/arxiv.org\/abs\/2407. 00079","author":"Qin Ruoyu","year":"2024","unstructured":"Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. 2024. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving. arXiv:2407.00079 [cs.DC] https:\/\/arxiv.org\/abs\/2407. 00079"},{"key":"e_1_2_1_35_1","unstructured":"QWen. 2024. QwQ: Reflect Deeply on the Boundaries of the Unknown. https:\/\/qwenlm.github.io\/blog\/qwq-32b-preview\/."},{"key":"e_1_2_1_36_1","doi-asserted-by":"crossref","unstructured":"Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm. In Extended abstracts of the 2021 CHI conference on human factors in computing systems. 1-7.","DOI":"10.1145\/3411763.3451760"},{"key":"e_1_2_1_37_1","volume-title":"Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, et al.","author":"Roziere Baptiste","year":"2023","unstructured":"Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, et al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950 (2023)."},{"key":"e_1_2_1_38_1","unstructured":"Noam Shazeer. 2019. Fast Transformer Decoding: One Write-Head is All You Need. arXiv:1911.02150 [cs.NE] https:\/\/arxiv.org\/abs\/1911.02150"},{"key":"e_1_2_1_39_1","volume-title":"Megatron- LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv:1909.08053 [cs.CL] https:\/\/arxiv.org\/abs\/1909.08053","author":"Shoeybi Mohammad","year":"2020","unstructured":"Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2020. Megatron- LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. arXiv:1909.08053 [cs.CL] https:\/\/arxiv.org\/abs\/1909.08053"},{"key":"e_1_2_1_40_1","unstructured":"Significant-Gravitas. 2023. AutoGPT: Build Deploy and Run AI Agents. https:\/\/github.com\/Significant-Gravitas\/AutoGPT."},{"key":"e_1_2_1_41_1","doi-asserted-by":"crossref","unstructured":"Chan Hee Song Jiaman Wu Clayton Washington Brian M. Sadler Wei-Lun Chao and Yu Su. 2023. LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models. arXiv:2212.04088 [cs.AI] https: \/\/arxiv.org\/abs\/2212.04088","DOI":"10.1109\/ICCV51070.2023.00280"},{"key":"e_1_2_1_42_1","volume-title":"Preble: Efficient Distributed Prompt Scheduling for LLM Serving. arXiv:2407.00023 [cs.DC] https:\/\/arxiv.org\/abs\/2407.00023","author":"Srivatsa Vikranth","year":"2024","unstructured":"Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, and Yiying Zhang. 2024. Preble: Efficient Distributed Prompt Scheduling for LLM Serving. arXiv:2407.00023 [cs.DC] https:\/\/arxiv.org\/abs\/2407.00023"},{"key":"e_1_2_1_43_1","volume-title":"Gemini: A Family of Highly Capable Multimodal Models. arXiv:2312.11805 [cs.CL] https:\/\/arxiv.org\/abs\/2312.11805","author":"Team Gemini","year":"2024","unstructured":"Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, et al. 2024. Gemini: A Family of Highly Capable Multimodal Models. arXiv:2312.11805 [cs.CL] https:\/\/arxiv.org\/abs\/2312.11805"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_2_1_45_1","unstructured":"vLLM Project. 2024. vLLM Automatic Prefix Caching. https:\/\/docs.vllm.ai\/en\/latest\/features\/automatic_prefix_caching. html. Accessed: 2024-08-09."},{"key":"e_1_2_1_46_1","volume-title":"Aakanksha Chowdhery, and Denny Zhou.","author":"Wang Xuezhi","year":"2023","unstructured":"Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171 [cs.CL] https:\/\/arxiv.org\/abs\/2203.11171"},{"key":"e_1_2_1_47_1","volume-title":"Echo: Efficient Co-Scheduling of Hybrid Online-Offline Tasks for Large Language Model Serving. arXiv:2504.03651 [cs.DC] https:\/\/arxiv.org\/abs\/2504.03651","author":"Wang Zhibin","year":"2025","unstructured":"Zhibin Wang, Shipeng Li, Xue Li, Yuhang Zhou, Zhonghui Zhang, Zibo Wang, Rong Gu, Chen Tian, Kun Yang, and Sheng Zhong. 2025. Echo: Efficient Co-Scheduling of Hybrid Online-Offline Tasks for Large Language Model Serving. arXiv:2504.03651 [cs.DC] https:\/\/arxiv.org\/abs\/2504.03651"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"},{"key":"e_1_2_1_49_1","unstructured":"Bingyang Wu Shengyu Liu Yinmin Zhong Peng Sun Xuanzhe Liu and Xin Jin. 2024. LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism. arXiv:2404.09526 [cs.DC] https:\/\/arxiv.org\/abs\/2404.09526"},{"key":"e_1_2_1_50_1","unstructured":"Yangzhen Wu Zhiqing Sun Shanda Li Sean Welleck and Yiming Yang. 2024. Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models. arXiv:2408.00724 [cs.AI] https:\/\/arxiv.org\/abs\/2408.00724"},{"key":"e_1_2_1_51_1","doi-asserted-by":"crossref","unstructured":"Junbin Xiao Xindi Shang Angela Yao and Tat-Seng Chua. 2021. NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions. arXiv:2105.08276 [cs.CV] https:\/\/arxiv.org\/abs\/2105.08276","DOI":"10.1109\/CVPR46437.2021.00965"},{"key":"e_1_2_1_52_1","unstructured":"Amy Yang Jingyi Yang Aya Ibrahim Xinfeng Xie Bangsheng Tang Grigory Sizov Jeremy Reizenstein Jongsoo Park and Jianyu Huang. 2024. Context Parallelism for Scalable Million-Token Inference. arXiv:2411.01783 [cs.DC] https:\/\/arxiv.org\/abs\/2411.01783"},{"key":"e_1_2_1_53_1","unstructured":"Jinwei Yao Kaiqi Chen Kexun Zhang Jiaxuan You Binhang Yuan Zeke Wang and Tao Lin. 2025. DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference. arXiv:2404.00242 [cs.CL] https:\/\/arxiv.org\/abs\/2404.00242"},{"key":"e_1_2_1_54_1","doi-asserted-by":"crossref","unstructured":"Jiayi Yao Hanchen Li Yuhan Liu Siddhant Ray Yihua Cheng Qizheng Zhang Kuntai Du Shan Lu and Junchen Jiang. 2025. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion. arXiv:2405.16444 [cs.LG] https:\/\/arxiv.org\/abs\/2405.16444","DOI":"10.1145\/3790254"},{"key":"e_1_2_1_55_1","unstructured":"Shunyu Yao Dian Yu Jeffrey Zhao Izhak Shafran Thomas L. Griffiths Yuan Cao and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv:2305.10601 [cs.CL] https:\/\/arxiv.org\/abs\/2305.10601"},{"key":"e_1_2_1_56_1","doi-asserted-by":"crossref","unstructured":"Lu Ye Ze Tao Yong Huang and Yang Li. 2024. ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition. arXiv:2402.15220 [cs.LG] https:\/\/arxiv.org\/abs\/2402.15220","DOI":"10.18653\/v1\/2024.acl-long.623"},{"key":"e_1_2_1_57_1","volume-title":"Cascade Inference: Memory Bandwidth Efficient Shared Prefix Batch Decoding. https:\/\/flashinfer.ai\/2024\/02\/02\/cascade-inference.html","author":"Ye Zihao","year":"2024","unstructured":"Zihao Ye, Ruihang Lai, Bo-Ru Lu, Chien-Yu Lin, Size Zheng, Lequn Chen, Tianqi Chen, and Luis Ceze. 2024. Cascade Inference: Memory Bandwidth Efficient Shared Prefix Batch Decoding. https:\/\/flashinfer.ai\/2024\/02\/02\/cascade-inference.html"},{"key":"e_1_2_1_58_1","unstructured":"Yilong Zhao Shuo Yang Kan Zhu Lianmin Zheng Baris Kasikci Yang Zhou Jiarong Xing and Ion Stoica. 2024. BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching. arXiv:2411.16102 [cs.LG] https:\/\/arxiv.org\/abs\/2411.16102"},{"key":"e_1_2_1_59_1","unstructured":"Zihuai Zhao Wenqi Fan Jiatong Li Yunqing Liu Xiaowei Mei Yiqi Wang Zhen Wen Fei Wang Xiangyu Zhao Jiliang Tang and Qing Li. 2024. Recommender Systems in the Era of Large Language Models (LLMs). arXiv:2307.02046 [cs.IR] https:\/\/arxiv.org\/abs\/2307.02046"},{"key":"e_1_2_1_60_1","volume-title":"Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng.","author":"Zheng Lianmin","year":"2024","unstructured":"Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. SGLang: Efficient Execution of Structured Language Model Programs. arXiv:2312.07104 [cs.AI] https:\/\/arxiv.org\/abs\/2312.07104"},{"key":"e_1_2_1_61_1","unstructured":"Zhen Zheng Xin Ji Taosong Fang Fanghao Zhou Chuanjie Liu and Gang Peng. 2024. BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching. arXiv:2412.03594 [cs.CL] https:\/\/arxiv.org\/abs\/2412.03594"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3676641.3716243"},{"key":"e_1_2_1_63_1","unstructured":"Zixuan Zhou Xuefei Ning Ke Hong Tianyu Fu Jiaming Xu Shiyao Li Yuming Lou Luning Wang Zhihang Yuan Xiuhong Li Shengen Yan Guohao Dai Xiao-Ping Zhang Yuhan Dong and Yu Wang. 2024. A Survey on Efficient Inference for Large Language Models. arXiv:2404.14294 [cs.CL] https:\/\/arxiv.org\/abs\/2404.14294"}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3802028","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T18:26:16Z","timestamp":1779128776000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3802028"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,18]]},"references-count":63,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,5,18]]}},"alternative-id":["10.1145\/3802028"],"URL":"https:\/\/doi.org\/10.1145\/3802028","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,18]]}}}