{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,28]],"date-time":"2026-05-28T06:01:45Z","timestamp":1779948105638,"version":"3.53.1"},"reference-count":38,"publisher":"Wiley","issue":"4","license":[{"start":{"date-parts":[[2026,1,21]],"date-time":"2026-01-21T00:00:00Z","timestamp":1768953600000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,1,21]],"date-time":"2026-01-21T00:00:00Z","timestamp":1768953600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62572462"],"award-info":[{"award-number":["62572462"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100021171","name":"Basic and Applied Basic Research Foundation of Guangdong Province","doi-asserted-by":"publisher","award":["2024A1515010251"],"award-info":[{"award-number":["2024A1515010251"]}],"id":[{"id":"10.13039\/501100021171","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100021171","name":"Basic and Applied Basic Research Foundation of Guangdong Province","doi-asserted-by":"publisher","award":["2023B1515130002"],"award-info":[{"award-number":["2023B1515130002"]}],"id":[{"id":"10.13039\/501100021171","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Softw Pract Exp"],"published-print":{"date-parts":[[2026,4]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Objective<\/jats:title>\n                    <jats:p>Large Language Models (LLMs) are increasingly deployed in modern AI infrastructure, creating a strong demand for high\u2010throughput and resource\u2010efficient serving systems. Disaggregated LLM serving, which decouples prompt prefill from auto\u2010regressive decode to accommodate their heterogeneous compute and memory characteristics, has emerged as a promising architecture. However, existing disaggregated serving systems suffer from three fundamental limitations: static resource allocation that fails to adapt to highly dynamic workloads, severe load imbalance between compute\u2010bound prefill and memory\u2010bound decode stages, and prefix\u2010cache\u2010aware routing that skews load distribution and creates performance hotspots. These issues collectively limit resource utilization, scalability, and the ability to meet service level objectives (SLOs) under real\u2010world workloads.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Methods<\/jats:title>\n                    <jats:p>To address these challenges, we propose BanaServe, a dynamic orchestration framework for disaggregated LLM serving that continuously rebalances both computational and memory resources across prefill and decode instances. BanaServe introduces three key mechanisms: (i) layer\u2010level weight migration to enable coarse\u2010grained redistribution of computation, (ii) attention\u2010level Key\u2013Value (KV) cache migration for fine\u2010grained memory load balancing, and (iii) a Global KV Cache Store with layer\u2010wise overlapped transmission to decouple routing decisions from cache placement. Together, these mechanisms eliminate cache\u2010induced hotspots and allow routers to perform purely load\u2010aware scheduling with minimal latency overhead. BanaServe is implemented on top of state\u2010of\u2010the\u2010art LLM serving frameworks, including vLLM and DistServe.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>We evaluate BanaServe under diverse and challenging workloads, including long\u2010context inference, bursty request arrivals, and mixed prompt\u2013generation patterns. Experimental results show that, compared to vLLM, BanaServe improves throughput by 1.2\u20133.9\u00d7 and reduces total processing time by 3.9%\u201378.4%. In comparison with DistServe, BanaServe achieves 1.1\u20132.8\u00d7 higher throughput while reducing latency by 1.4%\u201370.1%. These gains are consistent across workload variations, demonstrating BanaServe's robustness under highly dynamic serving conditions.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusion<\/jats:title>\n                    <jats:p>BanaServe demonstrates that dynamic, multi\u2010granularity resource rebalancing and cache\u2010decoupled routing are essential for efficient disaggregated LLM serving. By jointly addressing resource elasticity, stage imbalance, and cache\u2010induced load skew, BanaServe substantially improves throughput, latency, and resource utilization in real\u2010world deployments. This work provides a practical and scalable foundation for next\u2010generation LLM serving systems operating under dynamic and heterogeneous workloads.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1002\/spe.70054","type":"journal-article","created":{"date-parts":[[2026,1,21]],"date-time":"2026-01-21T08:58:51Z","timestamp":1768985931000},"page":"424-444","update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure"],"prefix":"10.1002","volume":"56","author":[{"given":"Yiyuan","family":"He","sequence":"first","affiliation":[{"name":"Southern University of Science and Technology  Shenzhen China"},{"name":"Shenzhen Institutes of Advanced Technology Chinese Academy of Sciences  Shenzhen China"},{"name":"AIOS Team Alibaba Group Inc.  HangZhou China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Minxian","family":"Xu","sequence":"additional","affiliation":[{"name":"Shenzhen Institutes of Advanced Technology Chinese Academy of Sciences  Shenzhen China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jingfeng","family":"Wu","sequence":"additional","affiliation":[{"name":"Shenzhen Institutes of Advanced Technology Chinese Academy of Sciences  Shenzhen China"},{"name":"University of Chinese Academy of Sciences  China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianmin","family":"Hu","sequence":"additional","affiliation":[{"name":"Southern University of Science and Technology  Shenzhen China"},{"name":"Shenzhen Institutes of Advanced Technology Chinese Academy of Sciences  Shenzhen China"},{"name":"AIOS Team Alibaba Group Inc.  HangZhou China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chong","family":"Ma","sequence":"additional","affiliation":[{"name":"AIOS Team Alibaba Group Inc.  HangZhou China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Min","family":"Shen","sequence":"additional","affiliation":[{"name":"AIOS Team Alibaba Group Inc.  HangZhou China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Le","family":"Chen","sequence":"additional","affiliation":[{"name":"AIOS Team Alibaba Group Inc.  HangZhou China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chengzhong","family":"Xu","sequence":"additional","affiliation":[{"name":"State Key Lab of IOTSC, Faculty of Science and Technology University of Macau  Macau SAR China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lin","family":"Qu","sequence":"additional","affiliation":[{"name":"AIOS Team Alibaba Group Inc.  HangZhou China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kejiang","family":"Ye","sequence":"additional","affiliation":[{"name":"Shenzhen Institutes of Advanced Technology Chinese Academy of Sciences  Shenzhen China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2026,1,21]]},"reference":[{"key":"e_1_2_11_2_1","article-title":"Gpt\u20104 Technical Report","author":"Achiam J.","year":"2023","journal-title":"arXiv Preprint, arXiv:2303.08774"},{"key":"e_1_2_11_3_1","article-title":"Llama: Open and Efficient Foundation Language Models","author":"Touvron H.","year":"2023","journal-title":"arXiv Preprint, arXiv:2302.13971"},{"key":"e_1_2_11_4_1","unstructured":"Anthropic Claude \u201cLarge Language Model \u201d2025 https:\/\/claude.ai."},{"key":"e_1_2_11_5_1","doi-asserted-by":"publisher","DOI":"10.1002\/spe.3432"},{"key":"e_1_2_11_6_1","doi-asserted-by":"publisher","DOI":"10.1002\/spe.70005"},{"key":"e_1_2_11_7_1","doi-asserted-by":"publisher","DOI":"10.1002\/spe.3422"},{"key":"e_1_2_11_8_1","first-page":"6000","article-title":"Attention is All You Need","author":"Vaswani A.","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_11_9_1","unstructured":"Z.Zhou X.Ning K.Hong et al. \u201cA Survey on Efficient Inference for Large Language Models \u201d arXiv preprint arXiv:2404.14294 (2024):1\u201336."},{"key":"e_1_2_11_10_1","first-page":"34661","article-title":"H2O: Heavy\u2010Hitter Oracle for Efficient Generative Inference of Large Language Models","volume":"36","author":"Zhang Z.","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_11_11_1","first-page":"1","volume-title":"The Twelfth International Conference on Learning Representations","author":"Ge S.","year":"2023"},{"key":"e_1_2_11_12_1","first-page":"193","volume-title":"18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Zhong Y.","year":"2024"},{"key":"e_1_2_11_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA59077.2024.00019"},{"key":"e_1_2_11_14_1","first-page":"62557","volume-title":"Sglang: Efficient Execution of Structured Language Model Programs","author":"Zheng L.","year":"2024"},{"key":"e_1_2_11_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613165"},{"key":"e_1_2_11_16_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"key":"e_1_2_11_17_1","unstructured":"R.Taori I.Gulrajani T.Zhang et al. \u201cStanford Alpaca: An Instruction\u2010Following LLaMA Model \u201d2023 https:\/\/github.com\/tatsu\u2010lab\/stanford_alpaca."},{"key":"e_1_2_11_18_1","unstructured":"NVIDIA Corporation \u201cNVIDIA Dynamo: Distributed Orchestration for Large\u2010Scale LLM Serving \u201d2025 https:\/\/developer.nvidia.com\/nvidia\u2010dynamo."},{"key":"e_1_2_11_19_1","first-page":"3119","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics","author":"Bai Y.","year":"2024"},{"key":"e_1_2_11_20_1","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1189"},{"key":"e_1_2_11_21_1","first-page":"68658","article-title":"Flashattention\u20103: Fast and Accurate Attention With Asynchrony and Low\u2010Precision","volume":"37","author":"Shah J.","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_11_22_1","unstructured":"S.Mukherjee A.Mitra G.Jawahar S.Agarwal H.Palangi andA.Awadallah \u201cOrca: Progressive Learning From Complex Explanation Traces of Gpt\u20104 \u201d arXiv Preprint arXiv:2306.02707 (2023):1\u201351."},{"key":"e_1_2_11_23_1","first-page":"117","volume-title":"18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Agrawal A.","year":"2024"},{"key":"e_1_2_11_24_1","first-page":"275","volume-title":"19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25)","author":"Zhang D.","year":"2025"},{"key":"e_1_2_11_25_1","unstructured":"C.Holmes M.Tanaka M.Wyatt et al. \u201cDeepspeed\u2010Fastgen: High\u2010Throughput Text Generation for LLMS via MII and Deepspeed\u2010Inference \u201d arXiv Preprint arXiv:2401.08671 (2024):1\u201314."},{"key":"e_1_2_11_26_1","unstructured":"NVIDIA Corporation \u201cTensorRT\u2010LLM: High\u2010Performance LLM Inference Library From NVIDIA \u201dhttps:\/\/developer.nvidia.com\/tensorrt\u2010llm."},{"key":"e_1_2_11_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3694715.3695948"},{"key":"e_1_2_11_28_1","first-page":"148","article-title":"Flashdecoding++: Faster Large Language Model Inference With Asynchronization, Flat Gemm Optimization, and Heuristics","volume":"6","author":"Hong K.","year":"2024","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_2_11_29_1","doi-asserted-by":"crossref","unstructured":"C.Hu H.Huang L.Xu et al. ShuffleInfer: Disaggregate LLM Inference for Mixed Downstream Workloads ACM Transactions on Architecture and Code Optimization.2025 Association for Computing Machinery 1544\u20133566 https:\/\/doi.org\/10.1145\/3732941.","DOI":"10.1145\/3732941"},{"key":"e_1_2_11_30_1","unstructured":"R.Qin Z.Li W.He et al. \u201cMooncake: Trading More Storage for Less Computation \u2014 A KVCache\u2010Centric Architecture for Serving LLM Chatbot \u201d in23rd USENIX Conference on File and Storage Technologies (FAST 25)(USENIX Association 2025) 155\u2013170 https:\/\/www.usenix.org\/conference\/fast25\/presentation\/qin."},{"key":"e_1_2_11_31_1","unstructured":"C.Hu H.Huang J.Hu et al. \u201cMemserve: Context Caching for Disaggregated Llm Serving With Elastic Memory Pool \u201d arXiv preprint arXiv:2406.17565 (2024): 1\u201314."},{"key":"e_1_2_11_32_1","unstructured":"S.Chen R.Jiang D.Yu et al. \u201cKVDirect: Distributed Disaggregated LLM Inference \u201d arXiv preprint arXiv:2501.14743 (2024)."},{"key":"e_1_2_11_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3695053.3730999"},{"key":"e_1_2_11_34_1","unstructured":"K.Hong L.Chen Z.Wang et al. \u201cSemi\u2010PD: Towards Efficient LLM Serving via Phase\u2010Wise Disaggregated Computation and Unified Storage \u201d arXiv preprint arXiv:2504.19867 (2025)."},{"key":"e_1_2_11_35_1","first-page":"173","volume-title":"18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Sun B.","year":"2024"},{"key":"e_1_2_11_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3620665.3640411"},{"key":"e_1_2_11_37_1","first-page":"218","volume-title":"Service\u2010Oriented Computing: 22nd International Conference, ICSOC 2024, Tunis, Tunisia, December 3\u20136, 2024, Proceedings, Part I","author":"He Y.","year":"2024"},{"key":"e_1_2_11_38_1","unstructured":"M.Azure \u201cAzure LLM Inference Dataset 2024 \u201d2024 https:\/\/github.com\/Azure\/AzurePublicDataset\/blob\/master\/AzureLLMInferenceDataset2024.md."},{"key":"e_1_2_11_39_1","unstructured":"M.Azure \u201cAzure Multimodal Model Inference Dataset 2025 \u201d2025 https:\/\/github.com\/Azure\/AzurePublicDataset\/blob\/master\/AzureLMMInferenceDataset2025.md."}],"container-title":["Software: Practice and Experience"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/spe.70054","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1002\/spe.70054","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/spe.70054","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,28]],"date-time":"2026-05-28T05:02:38Z","timestamp":1779944558000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/spe.70054"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,21]]},"references-count":38,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4]]}},"alternative-id":["10.1002\/spe.70054"],"URL":"https:\/\/doi.org\/10.1002\/spe.70054","archive":["Portico"],"relation":{},"ISSN":["0038-0644","1097-024X"],"issn-type":[{"value":"0038-0644","type":"print"},{"value":"1097-024X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,21]]},"assertion":[{"value":"2025-09-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-02","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}