{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T09:08:21Z","timestamp":1784538501856,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":23,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T00:00:00Z","timestamp":1776038400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,4,13]]},"DOI":"10.1145\/3788550.3794877","type":"proceedings-article","created":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T08:46:11Z","timestamp":1784537171000},"page":"142-147","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["ASTL: An Adaptive Serving Stack for LLMs to Balance Utilization and Tail Latency"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9726-5408","authenticated-orcid":false,"given":"Farhoud","family":"Jafari Kaleibar","sequence":"first","affiliation":[{"name":"York University, Toronto, ON, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0383-920X","authenticated-orcid":false,"given":"Marin","family":"Litoiu","sequence":"additional","affiliation":[{"name":"York University, Toronto, ON, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,20]]},"reference":[{"key":"e_1_3_3_1_2_2","volume-title":"LLM Inference Performance Engineering: Best Practices","author":"Agarwal Megha","year":"2023","unstructured":"Megha Agarwal, Asfandyar Qureshi, Nikhil Sardana, Linden Li, Julian Quevedo, and Daya Khudia. 2023. LLM Inference Performance Engineering: Best Practices. https:\/\/www.databricks.com\/blog\/llm-inference-performance-engineering-best-practices Databricks Blog."},{"key":"e_1_3_3_1_3_2","first-page":"117","volume-title":"18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Agrawal Amey","year":"2024","unstructured":"Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024. Taming { Throughput-Latency} tradeoff in { LLM} inference with { Sarathi-Serve}. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). 117\u2013134."},{"key":"e_1_3_3_1_4_2","unstructured":"Amey Agrawal Haoran Qiu Junda Chen \u00cd\u00f1igo Goiri Chaojie Zhang Rayyan Shahid Ramachandran Ramjee Alexey Tumanov and Esha Choukse. 2024. Medha: Efficiently Serving Multi-Million Context Length LLM Inference Requests Without Approximations. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2409.17264 (2024)."},{"key":"e_1_3_3_1_5_2","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared\u00a0D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et\u00a0al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020) 1877\u20131901."},{"key":"e_1_3_3_1_6_2","unstructured":"Mark Chen and et.al. 2021. Evaluating Large Language Models Trained on Code. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2107.03374 (2021). https:\/\/arxiv.org\/abs\/2107.03374 Benchmark: HumanEval."},{"key":"e_1_3_3_1_7_2","unstructured":"Karl Cobbe and et.al. 2021. Training Verifiers to Solve Math Word Problems. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2110.14168 (2021). https:\/\/arxiv.org\/abs\/2110.14168 Dataset: GSM8K."},{"key":"e_1_3_3_1_8_2","doi-asserted-by":"crossref","unstructured":"Tri Dao Dan Fu Stefano Ermon Atri Rudra and Christopher R\u00e9. 2022. Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in neural information processing systems 35 (2022) 16344\u201316359.","DOI":"10.52202\/068431-1189"},{"key":"e_1_3_3_1_9_2","unstructured":"Yichao Fu Siqi Zhu Runlong Su Aurick Qiao Ion Stoica and Hao Zhang. 2024. Efficient LLM Scheduling by Learning to Rank. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2408.15792 (2024). https:\/\/arxiv.org\/abs\/2408.15792"},{"key":"e_1_3_3_1_10_2","unstructured":"Yinsicheng Jiang Yao Fu Yeqi Huang Ping Nie Zhan Lu Leyang Xue Congjie He Man-Kit Sit Jilong Xue Li Dong et\u00a0al. 2024. MoE-CAP: Benchmarking Cost Accuracy and Performance of Sparse Mixture-of-Experts Systems. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2412.07067 (2024)."},{"key":"e_1_3_3_1_11_2","doi-asserted-by":"crossref","first-page":"611","DOI":"10.1145\/3600006.3613165","volume-title":"Proceedings of the 29th symposium on operating systems principles","author":"Kwon Woosuk","year":"2023","unstructured":"Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody\u00a0Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles. 611\u2013626."},{"key":"e_1_3_3_1_12_2","first-page":"19274","volume-title":"International Conference on Machine Learning","author":"Leviathan Yaniv","year":"2023","unstructured":"Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning. PMLR, 19274\u201319286."},{"key":"e_1_3_3_1_13_2","first-page":"78","volume-title":"2023 7th Iranian Conference on Advances in Enterprise Architecture (ICAEA)","author":"Nia Amirhossein\u00a0Hossein","year":"2023","unstructured":"Amirhossein\u00a0Hossein Nia, Farhoud\u00a0Jafari Kaleibar, Fatemehzahra Feizi, Fatemeh Rahimi, and Houman Kashfi. 2023. Unlocking the power of data in telecom: building an effective MLOps infrastructure for model deployment. In 2023 7th Iranian Conference on Advances in Enterprise Architecture (ICAEA). IEEE, 78\u201384."},{"key":"e_1_3_3_1_14_2","unstructured":"Bowen Pang Kai Li and Feifan Wang. 2025. Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2503.05248 (2025)."},{"key":"e_1_3_3_1_15_2","doi-asserted-by":"crossref","first-page":"843","DOI":"10.18653\/v1\/2025.acl-srw.61","volume-title":"Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop)","author":"Piotrowski Grzegorz","year":"2025","unstructured":"Grzegorz Piotrowski, Mateusz Bystro\u0144ski, Miko\u0142aj Ho\u0142ysz, Jakub Binkowski, Grzegorz Chodak, and Tomasz\u00a0Jan Kajdanowicz. 2025. When Will the Tokens End? Graph-Based Forecasting for LLMs Output Length. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop). 843\u2013848."},{"key":"e_1_3_3_1_16_2","doi-asserted-by":"crossref","first-page":"2383","DOI":"10.18653\/v1\/D16-1264","volume-title":"Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Rajpurkar Pranav","year":"2016","unstructured":"Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2383\u20132392. https:\/\/arxiv.org\/abs\/1606.05250"},{"key":"e_1_3_3_1_17_2","doi-asserted-by":"crossref","unstructured":"Pol\u00a0G Recasens Ferran Agullo Yue Zhu Chen Wang Eun\u00a0Kyung Lee Olivier Tardieu Jordi Torres and Josep\u00a0Ll Berral. 2025. Mind the memory gap: Unveiling gpu bottlenecks in large-batch llm inference. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2503.08311 (2025).","DOI":"10.1109\/CLOUD67622.2025.00036"},{"key":"e_1_3_3_1_18_2","first-page":"1","volume-title":"2024 34th International Conference on Collaborative Advances in Software and COmputiNg (CASCON)","author":"Sarda Komal","year":"2024","unstructured":"Komal Sarda, Zakeya Namrud, Ian Watts, Larisa Shwartz, Seema Nagar, Prateeti Mohapatra, and Marin Litoiu. 2024. Augmenting Automatic Root-Cause Identification with Incident Alerts Using LLM. In 2024 34th International Conference on Collaborative Advances in Software and COmputiNg (CASCON). IEEE, 1\u201310."},{"key":"e_1_3_3_1_19_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et\u00a0al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2307.09288 (2023)."},{"key":"e_1_3_3_1_20_2","unstructured":"Meixuan Wang Yinyu Ye and Zijie Zhou. 2025. LLM Serving Optimization with Variable Prefill and Decode Lengths. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2508.06133 (2025)."},{"key":"e_1_3_3_1_21_2","doi-asserted-by":"crossref","unstructured":"Danny Weyns Radu Calinescu Raffaela Mirandola Kenji Tei Maribel Acosta Nelly Bencomo Amel Bennaceur Nicolas Boltz Tomas Bures Javier Camara et\u00a0al. 2023. Towards a research agenda for understanding and managing uncertainty in self-adaptive systems. ACM SIGSOFT Software Engineering Notes 48 4 (2023) 20\u201336.","DOI":"10.1145\/3617946.3617951"},{"key":"e_1_3_3_1_22_2","first-page":"640","volume-title":"Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles","author":"Wu Bingyang","year":"2024","unstructured":"Bingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun, Xuanzhe Liu, and Xin Jin. 2024. Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles. 640\u2013654."},{"key":"e_1_3_3_1_23_2","first-page":"273","volume-title":"2024 22nd International Symposium on Modeling and Optimization in Mobile, Ad Hoc, and Wireless Networks (WiOpt)","author":"Yang Yuqing","year":"2024","unstructured":"Yuqing Yang, Lei Jiao, and Yuedong Xu. 2024. A queueing theoretic perspective on low-latency llm inference with variable token length. In 2024 22nd International Symposium on Modeling and Optimization in Mobile, Ad Hoc, and Wireless Networks (WiOpt). IEEE, 273\u2013280."},{"key":"e_1_3_3_1_24_2","first-page":"193","volume-title":"18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Zhong Yinmin","year":"2024","unstructured":"Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024. { DistServe} : Disaggregating prefill and decoding for goodput-optimized large language model serving. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). 193\u2013210."}],"event":{"name":"SEAMS '26: 21st International Conference on Software Engineering for Adaptive and Self-Managing Systems","location":"Rio de Janeiro Brazil","acronym":"SEAMS '26","sponsor":["SIGSOFT ACM Special Interest Group on Software Engineering","IEEE CS"]},"container-title":["Proceedings of the 21st International Conference on Software Engineering for Adaptive and Self-Managing Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3788550.3794877","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T08:47:52Z","timestamp":1784537272000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3788550.3794877"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,13]]},"references-count":23,"alternative-id":["10.1145\/3788550.3794877","10.1145\/3788550"],"URL":"https:\/\/doi.org\/10.1145\/3788550.3794877","relation":{},"subject":[],"published":{"date-parts":[[2026,4,13]]},"assertion":[{"value":"2026-07-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}