{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T14:47:51Z","timestamp":1782571671377,"version":"3.54.5"},"reference-count":88,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T00:00:00Z","timestamp":1782518400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62572304"],"award-info":[{"award-number":["62572304"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    Applications based on Large Language Models (LLMs) contain a series of tasks to address real-world problems with boosted capability, which have dynamic demand volumes on diverse backends. Existing serving systems treat the resource demands of LLM applications as a blackbox, compromising end-to-end efficiency due to improper queuing order and backend warm up latency. We find that the resource demands of LLM applications can be modeled in a general and accurate manner with\n                    <jats:italic toggle=\"yes\">Probabilistic Demand Graph<\/jats:italic>\n                    (PDGraph). We then propose Hermes, which leverages PDGraph for efficient serving of LLM applications. Confronting probabilistic demand description, Hermes applies the Gittins policy to determine the scheduling order that can minimize the average application completion time. It also uses the PDGraph model to help prewarm cold backends at proper moments. Experiments with diverse LLM applications confirm that Hermes can effectively improve the application serving efficiency, reducing the average completion time by over 70% and the P95 completion time by over 80%.\n                  <\/jats:p>","DOI":"10.1145\/3803390","type":"journal-article","created":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T11:35:51Z","timestamp":1775561751000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Hermes: Efficient Serving of LLM Applications with Probabilistic Demand Modeling"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-1328-8119","authenticated-orcid":false,"given":"Yifei","family":"Liu","sequence":"first","affiliation":[{"name":"Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-8285-2499","authenticated-orcid":false,"given":"Zuo","family":"Gan","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3006-1943","authenticated-orcid":false,"given":"Zhenghao","family":"Gan","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5548-9458","authenticated-orcid":false,"given":"Weiye","family":"Wang","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9480-5632","authenticated-orcid":false,"given":"Chen","family":"Chen","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-9519-0546","authenticated-orcid":false,"given":"Yizhou","family":"Shan","sequence":"additional","affiliation":[{"name":"Huawei Technologies Co Ltd","place":["Shenzhen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2807-9780","authenticated-orcid":false,"given":"Xusheng","family":"Chen","sequence":"additional","affiliation":[{"name":"Huawei Technologies Co Ltd","place":["Shenzhen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2880-7100","authenticated-orcid":false,"given":"Zhenhua","family":"Han","sequence":"additional","affiliation":[{"name":"Unaffiliated","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4352-6507","authenticated-orcid":false,"given":"Yifei","family":"Zhu","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4060-9438","authenticated-orcid":false,"given":"Shixuan","family":"Sun","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0034-2302","authenticated-orcid":false,"given":"Minyi","family":"Guo","sequence":"additional","affiliation":[{"name":"Computer Science, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,27]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"2025. Development Tool to Measure Monitor and Analyze the Memory Behavior of Python Objects in a Running Python Application. Retrieved July 16 2025 from https:\/\/pympler.readthedocs.io\/en\/latest\/"},{"key":"e_1_3_2_3_2","unstructured":"2025. Multi-Agent-as-a-Service. Retrieved July 16 2025 from https:\/\/medium.com\/data-science\/multi-agent-as-a-service-a-senior-engineers-overview-fc759f5bbcfa"},{"key":"e_1_3_2_4_2","unstructured":"2025. OpenAI Assistants API. Retrieved July 16 2025 from https:\/\/platform.openai.com\/docs\/assistants\/overview"},{"key":"e_1_3_2_5_2","unstructured":"2025. OpenAI Function Calling. Retrieved July 16 2025 from https:\/\/platform.openai.com\/docs\/guides\/function-calling"},{"key":"e_1_3_2_6_2","unstructured":"2025. The Origin Code of ALFWorld Interaction. Retrieved July 16 2025 from https:\/\/github.com\/ysymyth\/ReAct\/blob\/6bdb3a1fd38b8188fc7ba4102969fe483df8fdc9\/alfworld.ipynb"},{"key":"e_1_3_2_7_2","unstructured":"2025. The Origin Code of Code Checking. Retrieved July 16 2025 from https:\/\/github.com\/GAIR-NLP\/factool\/tree\/3f3914bc090b644be044b7e0005113c135d8b20f\/factool\/code"},{"key":"e_1_3_2_8_2","unstructured":"2025. The Origin Code of Code Generation. Retrieved July 16 2025 from https:\/\/github.com\/microsoft\/autogen\/tree\/0560bdd645dfbc579a71f2f0fea98ea83dd3bb3f?tab=readme-ov-file#quickstart"},{"key":"e_1_3_2_9_2","unstructured":"2025. The Origin Code of Document Merging. Retrieved July 16 2025 from https:\/\/github.com\/spcl\/graph-of-thoughts\/tree\/a939a4577c07c80b8ecb194793b5a4169d99b31b\/examples\/doc_merge"},{"key":"e_1_3_2_10_2","unstructured":"2025. The Origin Code of Equation Verification. Retrieved July 16 2025 from https:\/\/github.com\/GAIR-NLP\/factool\/tree\/3f3914bc090b644be044b7e0005113c135d8b20f\/factool\/math"},{"key":"e_1_3_2_11_2","unstructured":"2025. The Origin Code of Fact Extraction and Verification. Retrieved July 16 2025 from https:\/\/github.com\/ysymyth\/ReAct\/blob\/6bdb3a1fd38b8188fc7ba4102969fe483df8fdc9\/FEVER.ipynb"},{"key":"e_1_3_2_12_2","unstructured":"2025. The Origin Code of Knowledge-Based-Query-Answering Verification. Retrieved July 16 2025 from https:\/\/github.com\/GAIR-NLP\/factool\/tree\/3f3914bc090b644be044b7e0005113c135d8b20f\/factool\/knowledge_qa"},{"key":"e_1_3_2_13_2","unstructured":"2025. The Origin Code of Plan-and-Execution. Retrieved July 16 2025 from https:\/\/github.com\/microsoft\/JARVIS\/tree\/c62e0faac76c4a2907cabe2cfe4bbe5f2e613400"},{"key":"e_1_3_2_14_2","unstructured":"2025. The Shift from Models to Compound AI Systems. Retrieved July 16 2025 from https:\/\/bair.berkeley.edu\/blog\/2024\/02\/18\/compound-ai-systems\/"},{"key":"e_1_3_2_15_2","unstructured":"2025. The Tutorials of How to Implement MapReduce Summarization. Retrieved July 16 2025 from https:\/\/python.langchain.com\/v0.2\/docs\/tutorials\/summarization\/#go-deeper-1"},{"key":"e_1_3_2_16_2","unstructured":"2025. vLLM: Easy Fast and Cheap LLM Serving for Everyone. Retrieved July 16 2025 from https:\/\/docs.vllm.ai\/en\/stable\/"},{"key":"e_1_3_2_17_2","unstructured":"2025. ZeroMQ - An open-source universal messaging library. Retrieved July 16 2025 from https:\/\/platform.openai.com\/docs\/overview"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","unstructured":"Samuli Aalto Urtzi Ayesta and Rhonda Righter. 2009. On the gittins index in the M\/G\/1 queue. Queueing Systems: Theory and Applications 63 1\u20134 (2009) 437\u2013458.","DOI":"10.1007\/s11134-009-9141-x"},{"key":"e_1_3_2_19_2","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Abhyankar Reyna","year":"2024","unstructured":"Reyna Abhyankar, Zijian He, Vikranth Srivatsa, Hao Zhang, and Yiying Zhang. 2024. InferCept: Efficient intercept support for augmented large language model inference. In Proceedings of the 41st International Conference on Machine Learning."},{"key":"e_1_3_2_20_2","volume-title":"Proceedings of the USENIX Conference on Operating Systems Design and Implementation (USENIX OSDI)","author":"Agrawal Amey","year":"2024","unstructured":"Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024. Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve. In Proceedings of the USENIX Conference on Operating Systems Design and Implementation (USENIX OSDI)."},{"key":"e_1_3_2_21_2","unstructured":"Amey Agrawal Haoran Qiu Junda Chen \u00cd\u00f1igo Goiri Chaojie Zhang Rayyan Shahid Ramachandran Ramjee Alexey Tumanov and Esha Choukse. 2024. No request left behind: Tackling heterogeneity in long-context LLM inference with medha. arXiv preprint arXiv:2409.17264 (2024)."},{"key":"e_1_3_2_22_2","first-page":"469","volume-title":"Proceedings of the 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17)","author":"Alipourfard Omid","year":"2017","unstructured":"Omid Alipourfard, Hongqiang Harry Liu, Jianshu Chen, Shivaram Venkataraman, Minlan Yu, and Ming Zhang. 2017. \\(\\lbrace\\) CherryPick \\(\\rbrace\\) : Adaptively unearthing the best cloud configurations for big data analytics. In Proceedings of the 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). 469\u2013482."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/INFCOM.2000.832234"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/INFCOM.1996.497885"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i16.29720"},{"key":"e_1_3_2_26_2","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell and others. 2020. Language models are few-shot learners. Advances in Neural Information Processing Systems 33 (2020) 1877\u20131901."},{"key":"e_1_3_2_27_2","unstructured":"Tianle Cai Yuhong Li Zhengyang Geng Hongwu Peng Jason D. Lee huai De-Chen and Tri Dao. 2024. Medusa: Simple LLM inference acceleration framework with multiple decoding heads. In International Conference on Machine Learning. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:267061277"},{"key":"e_1_3_2_28_2","unstructured":"Lingjiao Chen Matei Zaharia and James Zou. 2024. FrugalGPT: How to use large language models while reducing cost and improving performance. (2024). Retrieved from https:\/\/openreview.net\/forum?id=cSimKw5p6R"},{"key":"e_1_3_2_29_2","unstructured":"Shouyuan Chen Sherman Wong Liangjian Chen and Yuandong Tian. 2023. Extending context window of large language models via positional interpolation. arXiv preprint arXiv:2306.15595 (2023)."},{"key":"e_1_3_2_30_2","unstructured":"Ethan Chern Steffi Chern Shiqi Chen Weizhe Yuan Kehua Feng Chunting Zhou Junxian He Graham Neubig and Pengfei Liu. 2025. FacTool: Factuality detection in generative AI \u2013 A tool augmented framework for multi-task and multi-domain scenarios. In Second Conference on Language Modeling. Retrieved from https:\/\/openreview.net\/forum?id=hJkQL9VtWT"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","unstructured":"Tri Dao Daniel Y. Fu Stefano Ermon Atri Rudra and Christopher R\u00e9. 2022. Flashattention: Fast and memory-efficient exact attention with IO-awareness. In Proceedings of the 36th International Conference on Neural Information Processing Systems. 16344\u201316359.","DOI":"10.52202\/068431-1189"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/REAL.1993.393496"},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","unstructured":"Luciano Floridi and Massimo Chiriatti. 2020. GPT-3: Its nature scope limits and consequences. Minds and Machines 30 4 (2020) 681\u2013694.","DOI":"10.1007\/s11023-020-09548-1"},{"key":"e_1_3_2_34_2","unstructured":"Elias Frantar Saleh Ashkboos Torsten Hoefler and Dan Alistarh. 2022. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323 (2022)."},{"key":"e_1_3_2_35_2","unstructured":"Yichao Fu Junda Chen Siqi Zhu Zheyu Fu Zhongdongming Dai Aurick Qiao and Hao Zhang. 2024. Efficiently serving llm reasoning programs with certaindex. arXiv e-prints (2024) arXiv\u20132412."},{"key":"e_1_3_2_36_2","unstructured":"Yichao Fu Siqi Zhu Runlong Su Aurick Qiao Ion Stoica and Hao Zhang. 2024. Efficient LLM scheduling by learning to rank. In The Thirty-eighth Annual Conference on Neural Information Processing Systems. Retrieved from https:\/\/openreview.net\/forum?id=wlLjYl0Gi6"},{"key":"e_1_3_2_37_2","first-page":"111","volume-title":"Proceedings of the 2024 USENIX Annual Technical Conference (USENIX ATC 24)","author":"Gao Bin","year":"2024","unstructured":"Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo. 2024. \\(\\lbrace\\) Cost-efficient \\(\\rbrace\\) large language model serving for multi-turn conversations with \\(\\lbrace\\) cachedattention \\(\\rbrace\\) . In Proceedings of the 2024 USENIX Annual Technical Conference (USENIX ATC 24). 111\u2013126."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1002\/9780470980033"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3575693.3575721"},{"key":"e_1_3_2_40_2","unstructured":"Ke Hong Guohao Dai Jiaming Xu Qiuli Mao Xiuhong Li Jun Liu Kangdi Chen Yuhan Dong and Yu Wang. 2024. Flashdecoding++: Faster large language model inference with asynchronization flat gemm optimization and heuristics. In Proceedings of Machine Learning and Systems 6 (2024) 148\u2013161."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","unstructured":"Cunchen Hu Heyang Huang Liangliang Xu Xusheng Chen Chenxi Wang Jiang Xu Shuang Chen Hao Feng Sa Wang Yungang Bao Ninghui Sun and Yizhou Shan. 2025. ShuffleInfer: Disaggregate LLM inference for mixed downstream workloads. ACM Trans. Archit. Code Optim. 22 2 Article 77 (July 2025). DOI:10.1145\/3732941","DOI":"10.1145\/3732941"},{"key":"e_1_3_2_42_2","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Liang Wang Weizhu Chen and others. 2022. Lora: Low-rank adaptation of large language models. ICLR 1 2 (2022) 3."},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3575693.3575705"},{"key":"e_1_3_2_44_2","unstructured":"Redwan Ibne Seraj Khan Kunal Jain Haiying Shen Ankur Mallick Anjaly Parayil Anoop Kulkarni Steve Kofsky Pankhuri Choudhary Ren\u00e8e St Amant Rujia Wang et\u00a0al. 2024. Ensuring fair LLM serving amid diverse applications. arXiv preprint arXiv:2411.15997 (2024)."},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613175"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3571730"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.123"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","unstructured":"Yang Zhang Hanlei Jin Dan Meng Jun Wang and Jinghua Tan. 2026. A comprehensive survey on automatic text summarization with exploration of LLM-based methods. Neurocomputing 663 (2026) 131928. DOI:10.1016\/j.neucom.2025.131928","DOI":"10.1016\/j.neucom.2025.131928"},{"key":"e_1_3_2_49_2","doi-asserted-by":"crossref","unstructured":"Yunho Jin Chun-Feng Wu David Brooks and Gu-Yeon Wei. 2023. \\(S^3\\) : Increasing GPU utilization during generative inference for higher throughput. Advances in Neural Information Processing Systems 36 (2023) 18015\u201318027.","DOI":"10.52202\/075280-0791"},{"key":"e_1_3_2_50_2","unstructured":"Sehoon Kim Coleman Richard Charles Hooper Amir Gholami Zhen Dong Xiuyu Li Sheng Shen Michael W. Mahoney and Kurt Keutzer. 2024. SqueezeLLM: Dense-and-sparse quantization. In Proceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research). PMLR 23901\u201323923. Retrieved from https:\/\/proceedings.mlr.press\/v235\/kim24f.html"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613165"},{"key":"e_1_3_2_52_2","unstructured":"Suyi Li Hanfeng Lu Tianyuan Wu Minchen Yu Qizhen Weng Xusheng Chen Yizhou Shan Binhang Yuan and Wei Wang. 2024. Caraserve: Cpu-assisted and rank-aware lora serving for generative LLM inference. arXiv preprint arXiv:2401.11240 (2024)."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3604237.3626869"},{"key":"e_1_3_2_54_2","unstructured":"Bin Lin Chen Zhang Tao Peng Hanyu Zhao Wencong Xiao Minmin Sun Anmin Liu Zhipeng Zhang Lanbo Li Xiafei Qiu et\u00a0al. 2024. Infinite-LLM: efficient LLM service for long context with distattention and distributed kvcache. arXiv preprint arXiv:2401.02669 (2024)."},{"key":"e_1_3_2_55_2","volume-title":"Proceedings of the USENIX Conference on Operating Systems Design and Implementation (USENIX OSDI)","author":"Lin Chaofan","year":"2024","unstructured":"Chaofan Lin, Zhenhua Han, Chengruidong Zhang, Yuqing Yang, Fan Yang, Chen Chen, and Lili Qiu. 2024. Parrot: Efficient serving of LLM-based applications with semantic variable. In Proceedings of the USENIX Conference on Operating Systems Design and Implementation (USENIX OSDI)."},{"key":"e_1_3_2_56_2","unstructured":"Aixin Liu Bei Feng Bing Xue Bingxuan Wang Bochao Wu Chengda Lu Chenggang Zhao Chengqi Deng Chenyu Zhang Chong Ruan et\u00a0al. 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437 (2024)."},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","unstructured":"Christos Makridis. 2025. The impact of generative artificial intelligence on artists. Available at SSRN (2025).","DOI":"10.2139\/ssrn.5179390"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","unstructured":"Xupeng Miao Gabriele Oliaro Zhihao Zhang Xinhao Cheng Zeyu Wang Zhengxin Zhang Rae Ying Yee Wong Alan Zhu Lijie Yang Xiaoxiang Shi Chunan Shi Zhuoming Chen Daiyaan Arfeen Reyna Abhyankar and Zhihao Jia. 2024. SpecInfer: Accelerating large language model serving with tree-based speculative inference and verification. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201924). Association for Computing Machinery La Jolla CA USA 3 (2024) 932\u2013949. DOI:10.1145\/3620666.3651335","DOI":"10.1145\/3620666.3651335"},{"key":"e_1_3_2_59_2","volume-title":"Proceedings of the USENIX Conference on Operating Systems Design and Implementation (USENIX OSDI)","author":"Narayanan Deepak","year":"2020","unstructured":"Deepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee, and Matei Zaharia. 2020. Heterogeneity-aware cluster scheduling policies for deep learning workloads. In Proceedings of the USENIX Conference on Operating Systems Design and Implementation (USENIX OSDI)."},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/90.234856"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3698038.3698523"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3698038.3698523"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3190508.3190517"},{"key":"e_1_3_2_64_2","volume-title":"Proceedings of the 15th  \\(\\lbrace\\) USENIX \\(\\rbrace\\)  Symposium on Operating Systems Design and Implementation (OSDI\u201921)","author":"Qiao Aurick","year":"2021","unstructured":"Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R. Ganger, and Eric P. Xing. 2021. Pollux: Co-adaptive cluster scheduling for goodput-optimized deep learning. In Proceedings of the 15th \\(\\lbrace\\) USENIX \\(\\rbrace\\) Symposium on Operating Systems Design and Implementation (OSDI\u201921)."},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","unstructured":"Ruoyu Qin Zheming Li Weiran He Jialei Cui Heyi Tang Feng Ren Teng Ma Shangming Cai Yineng Zhang Mingxing Zhang Yongwei Wu Weimin Zheng and Xinran Xu. 2025. Mooncake: A Kvcache-centric disaggregated architecture for LLM serving. (2025). DOI:10.1145\/3773772","DOI":"10.1145\/3773772"},{"key":"e_1_3_2_66_2","unstructured":"Haoran Qiu Weichao Mao Archit Patke Shengkun Cui Saurabh Jha Chen Wang Hubertus Franke Zbigniew T. Kalbarczyk Tamer Basar and Ravishankar K. Iyer. 2024. Efficient interactive LLM serving with proxy model-based sequence length prediction. In International Conference on Architectural Support for Programming Languages and Operating Systems."},{"key":"e_1_3_2_67_2","unstructured":"Gemini Team Petko Georgiev Ving Ian Lei Ryan Burnell Libin Bai Anmol Gulati Garrett Tanzer Damien Vincent Zhufeng Pan Shibo Wang et\u00a0al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530 (2024)."},{"key":"e_1_3_2_68_2","unstructured":"Shuo Ren Can Xie Pu Jian Zhenjiang Ren Chunlin Leng and Jiajun Zhang. 2025. Towards scientific intelligence: A survey of LLM-based scientific agents. arXiv preprint arXiv:2503.24047 (2025)."},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1287\/opre.16.3.687"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.23919\/WiOpt52861.2021.9589051"},{"key":"e_1_3_2_71_2","doi-asserted-by":"crossref","unstructured":"Philip Sedgwick. 2012. Pearson\u2019s correlation coefficient. BMJ 345 (2012) e4483.","DOI":"10.1136\/bmj.e4483"},{"key":"e_1_3_2_72_2","unstructured":"Rana Shahout Eran Malach Chunwei Liu Weifan Jiang Minlan Yu and Michael Mitzenmacher. 2024. Don\u2019t stop me now: Embedding based scheduling for LLMs. arXiv preprint arXiv:2410.01035 (2024)."},{"key":"e_1_3_2_73_2","doi-asserted-by":"crossref","unstructured":"Yongliang Shen Kaitao Song Xu Tan Dongsheng Li Weiming Lu and Yueting Zhuang. 2023. HuggingGPT: Solving AI tasks with chatgpt and its friends in hugging face. Advances in Neural Information Processing Systems 36 (2023) 38154\u201338180.","DOI":"10.52202\/075280-1657"},{"key":"e_1_3_2_74_2","unstructured":"Zhuocheng Shen. 2024. LLM with tools: A survey. arXiv preprint arXiv:2409.18807 (2024)."},{"key":"e_1_3_2_75_2","unstructured":"Ying Sheng Shiyi Cao Dacheng Li Coleman Hooper Nicholas Lee Shuo Yang Christopher Chou Banghua Zhu Lianmin Zheng Kurt Keutzer Joseph E. Gonzalez and Ion Stoica. 2024. S-LoRA: Scalable serving of thousands of LoRA adapters. In Proceedings of Machine Learning and Systems P. Gibbons G. Pekhimenko and C. De Sa (Eds.). 6 (2024) 296\u2013311. Retrieved from https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2024\/file\/906419cd502575b617cc489a1a696a67-Paper-Conference.pdf"},{"key":"e_1_3_2_76_2","first-page":"965","volume-title":"Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Sheng Ying","year":"2024","unstructured":"Ying Sheng, Shiyi Cao, Dacheng Li, Banghua Zhu, Zhuohan Li, Danyang Zhuo, Joseph E. Gonzalez, and Ion Stoica. 2024. Fairness in serving large language models. In Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). 965\u2013988."},{"key":"e_1_3_2_77_2","doi-asserted-by":"crossref","unstructured":"Yuki Sonoda Ryo Kurokawa Yuta Nakamura Jun Kanzawa Mariko Kurokawa Yuji Ohizumi Wataru Gonoi and Osamu Abe. 2024. Diagnostic performances of GPT-4o claude 3 opus and gemini 1.5 pro in \u201cdiagnosis please\u201d cases. Japanese Journal of Radiology 42 11 (2024) 1231\u20131235.","DOI":"10.1007\/s11604-024-01619-y"},{"key":"e_1_3_2_78_2","volume-title":"Proceedings of the USENIX Conference on Operating Systems Design and Implementation (USENIX OSDI)","author":"Sun Biao","year":"2024","unstructured":"Biao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao, Xinyi Zhang, Yong Li, and Wei Lin. 2024. Llumnix: Dynamic scheduling for large language model serving. In Proceedings of the USENIX Conference on Operating Systems Design and Implementation (USENIX OSDI)."},{"key":"e_1_3_2_79_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timothee Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et\u00a0al. 2023. Llama: Open and efficient foundation language models."},{"key":"e_1_3_2_80_2","article-title":"Code execution is now by default inside docker container","author":"Vrousgou Olga","year":"2025","unstructured":"Olga Vrousgou. 2025. Code execution is now by default inside docker container. Retrieved July 16, 2025 from https:\/\/microsoft.github.io\/autogen\/0.2\/blog\/2024\/01\/23\/Code-execution-in-docker\/","journal-title":"https:\/\/microsoft.github.io\/autogen\/0.2\/blog\/2024\/01\/23\/Code-execution-in-docker\/"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.1145\/3694715.3695948"},{"key":"e_1_3_2_82_2","unstructured":"Bingyang Wu Yinmin Zhong Zili Zhang Shengyu Liu Fangyue Liu Yuanhang Sun Gang Huang Xuanzhe Liu and Xin Jin. 2023. Fast distributed inference serving for large language models. arXiv preprint arXiv:2305.05920 (2023)."},{"key":"e_1_3_2_83_2","unstructured":"Shunyu Yao Jeffrey Zhao Dian Yu Nan Du Izhak Shafran Karthik R Narasimhan and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations."},{"key":"e_1_3_2_84_2","volume-title":"Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (USENIX OSDI)","author":"Yu Gyeong-In","year":"2022","unstructured":"Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022. ORCA: A distributed serving system for transformer-based generative models. In Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (USENIX OSDI)."},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","unstructured":"Yilong Zhao Shuo Yang Kan Zhu Lianmin Zheng Baris Kasikci Yifan Qiao Yang Zhou Jiarong Xing and Ion Stoica. 2026. BlendServe: Optimizing offline inference with resource-aware batching. In Proceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems 2 (2026) 255\u2013273. DOI:10.1145\/3779212.3790133","DOI":"10.1145\/3779212.3790133"},{"key":"e_1_3_2_86_2","first-page":"91","volume-title":"Proceedings of the 2020 IFIP Networking Conference (Networking)","author":"Zheng Peng","year":"2020","unstructured":"Peng Zheng, Wendi Feng, Arvind Narayanan, and Zhi-Li Zhang. 2020. NFV performance profiling on multi-core servers. In Proceedings of the 2020 IFIP Networking Conference (Networking). IEEE, 91\u201399."},{"key":"e_1_3_2_87_2","volume-title":"Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (USENIX NSDI)","author":"Zheng Pengfei","year":"2023","unstructured":"Pengfei Zheng, Rui Pan, Tarannum Khan, Shivaram Venkataraman, and Aditya Akella. 2023. Shockwave: Fair and efficient cluster scheduling for dynamic adaptation in machine learning. In Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (USENIX NSDI)."},{"key":"e_1_3_2_88_2","doi-asserted-by":"crossref","unstructured":"Zangwei Zheng Xiaozhe Ren Fuzhao Xue Yang Luo Xin Jiang and Yang You. 2023. Response length perception and sequence scheduling: An LLM-empowered LLM inference pipeline. Advances in Neural Information Processing Systems 36 (2023) 65517\u201365530.","DOI":"10.52202\/075280-2859"},{"key":"e_1_3_2_89_2","volume-title":"Proceedings of the USENIX Conference on Operating Systems Design and Implementation (USENIX OSDI)","author":"Zhong Yinmin","year":"2024","unstructured":"Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024. DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. In Proceedings of the USENIX Conference on Operating Systems Design and Implementation (USENIX OSDI)."}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3803390","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T14:17:15Z","timestamp":1782569835000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3803390"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,27]]},"references-count":88,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3803390"],"URL":"https:\/\/doi.org\/10.1145\/3803390","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,27]]},"assertion":[{"value":"2025-07-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}