{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T16:46:35Z","timestamp":1782405995512,"version":"3.54.5"},"reference-count":73,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    Serverless computing offers a compelling cloud model for online inference services. However, existing serverless platforms lack efficient support for GPUs, hindering their ability to deliver high-performance inference. In this article, we present\n                    <jats:monospace>Torpor<\/jats:monospace>\n                    , a serverless platform for GPU-efficient, low-latency inference. To enable efficient sharing of a node\u2019s GPUs among numerous inference functions,\n                    <jats:monospace>Torpor<\/jats:monospace>\n                    maintains models in main memory and dynamically swaps them onto GPUs upon request arrivals (i.e., late binding with model swapping).\n                    <jats:monospace>Torpor<\/jats:monospace>\n                    uses various techniques, including asynchronous API redirection, GPU runtime sharing, pipelined model execution, and efficient GPU memory management, to minimize latency overhead caused by model swapping. Additionally, we design an interference-aware request scheduling algorithm that utilizes high-speed GPU interconnects to meet latency service-level objectives (SLOs) for individual inference functions. We have implemented\n                    <jats:monospace>Torpor<\/jats:monospace>\n                    and evaluated its performance in a production environment. Utilizing late binding and model swapping,\n                    <jats:monospace>Torpor<\/jats:monospace>\n                    can concurrently serve hundreds of inference functions on a worker node with 4 GPUs, while achieving latency performance comparable to native execution, where each model is cached exclusively on a GPU. Pilot deployment in a leading commercial serverless cloud shows that\n                    <jats:monospace>Torpor<\/jats:monospace>\n                    reduces the GPU provisioning cost by 70% and 65% for users and the platform, respectively.\n                  <\/jats:p>","DOI":"10.1145\/3800690","type":"journal-article","created":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T11:06:09Z","timestamp":1776078369000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Enabling Low-Latency, GPU-Efficient Serverless Inference with Model Swapping"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6797-9028","authenticated-orcid":false,"given":"Minchen","family":"Yu","sequence":"first","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen","place":["Shenzhen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-1018-0957","authenticated-orcid":false,"given":"Ao","family":"Wang","sequence":"additional","affiliation":[{"name":"Alibaba Group","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5210-564X","authenticated-orcid":false,"given":"Bohui","family":"Wu","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen","place":["Shenzhen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-6785-5053","authenticated-orcid":false,"given":"Yuxuan","family":"Liu","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen","place":["Shenzhen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0167-8432","authenticated-orcid":false,"given":"Dong","family":"Chen","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology","place":["Hong Kong, Hong Kong"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-9677-7960","authenticated-orcid":false,"given":"Haoxuan","family":"Yu","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology","place":["Hong Kong, Hong Kong"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4585-4152","authenticated-orcid":false,"given":"Wei","family":"Wang","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology","place":["Hong Kong, Hong Kong"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5060-8411","authenticated-orcid":false,"given":"Ruichuan","family":"Chen","sequence":"additional","affiliation":[{"name":"Nokia Bell Labs","place":["Stuttgart, Germany"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-1381-6834","authenticated-orcid":false,"given":"Dapeng","family":"Nie","sequence":"additional","affiliation":[{"name":"Alibaba Group","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-8889-4132","authenticated-orcid":false,"given":"Haoran","family":"Yang","sequence":"additional","affiliation":[{"name":"Alibaba Group","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-1179-9888","authenticated-orcid":false,"given":"Yu","family":"Ding","sequence":"additional","affiliation":[{"name":"Alibaba Group","place":["Hangzhou, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"2025. AWS Lambda. Retrieved from https:\/\/aws.amazon.com\/lambda\/Accessed February 24 2026."},{"key":"e_1_3_2_3_2","unstructured":"2026. Alibaba Cloud Function Compute. Retrieved from https:\/\/www.alibabacloud.com\/product\/function-computeAccessed February 24 2026."},{"key":"e_1_3_2_4_2","unstructured":"2026. Aliyun cGPU. Retrieved from https:\/\/www.alibabacloud.com\/help\/en\/container-service-for-kubernetes\/latest\/cgpu-overviewAccessed February 24 2026."},{"key":"e_1_3_2_5_2","unstructured":"2026. Aliyun Function Compute Billing Scheme. Retrieved from https:\/\/www.alibabacloud.com\/help\/en\/function-compute\/latest\/billing-billingAccessed February 24 2026."},{"key":"e_1_3_2_6_2","unstructured":"2026. Aliyun Function Compute Instance Types and Modes. Retrieved from https:\/\/www.alibabacloud.com\/help\/en\/function-compute\/latest\/instance-types-and-instance-modesAccessed February 24 2026."},{"key":"e_1_3_2_7_2","unstructured":"2026. Amazon SageMaker. Retrieved from https:\/\/aws.amazon.com\/sagemaker\/Accessed February 24 2026."},{"key":"e_1_3_2_8_2","unstructured":"2026. AWS Lambda Provisioned Concurrency. Retrieved from https:\/\/docs.aws.amazon.com\/lambda\/latest\/dg\/provisioned-concurrency.htmlAccessed February 24 2026."},{"key":"e_1_3_2_9_2","unstructured":"2026. Azure Functions. Retrieved from https:\/\/azure.microsoft.com\/en-us\/services\/functions\/."},{"key":"e_1_3_2_10_2","unstructured":"2026. GVirtuS. Retrieved from https:\/\/github.com\/gvirtus\/GVirtuSAccessed February 24 2026."},{"key":"e_1_3_2_11_2","unstructured":"2026. Llama2. Retrieved from https:\/\/www.llama.com\/llama2Accessed February 24 2026."},{"key":"e_1_3_2_12_2","unstructured":"2026. Llama3. Retrieved from https:\/\/www.llama.com\/models\/llama-3Accessed February 24 2026."},{"key":"e_1_3_2_13_2","unstructured":"2026. Memory Management on Modern GPU Architectures. Retrieved from https:\/\/developer.download.nvidia.com\/video\/gputechconf\/gtc\/2019\/presentation\/s9727-memory-management-on-modern-gpu-architectures.pdfAccessed February 24 2026."},{"key":"e_1_3_2_14_2","unstructured":"2026. Nvidia Multi-Instance GPU. Retrieved from https:\/\/www.nvidia.com\/en-us\/technologies\/multi-instance-gpu\/."},{"key":"e_1_3_2_15_2","unstructured":"2026. Nvidia Multi-Process Service. Retrieved from https:\/\/docs.nvidia.com\/deploy\/mps\/Accessed February 24 2026."},{"key":"e_1_3_2_16_2","unstructured":"2026. Nvidia Virtual GPU. Retrieved from https:\/\/www.nvidia.com\/en-us\/data-center\/virtual-solutions\/."},{"key":"e_1_3_2_17_2","unstructured":"2026. Qwen. Retrieved from https:\/\/github.com\/QwenLM\/QwenAccessed February 24 2026."},{"key":"e_1_3_2_18_2","unstructured":"2026. ResNet in PyTorch. Retrieved from https:\/\/pytorch.org\/vision\/stable\/models\/resnet.htmlAccessed February 24 2026."},{"key":"e_1_3_2_19_2","unstructured":"2026. Stable Diffusion. Retrieved from https:\/\/huggingface.co\/stable-diffusion-v1-5\/stable-diffusion-v1-5Accessed February 24 2026."},{"key":"e_1_3_2_20_2","first-page":"419","volume-title":"Proceedings of the 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI 20)","author":"Agache Alexandru","year":"2020","unstructured":"Alexandru Agache, Marc Brooker, Alexandra Iordache, Anthony Liguori, Rolf Neugebauer, Phil Piwonka, and Diana-Maria Popa. 2020. Firecracker: Lightweight virtualization for serverless applications. In Proceedings of the 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI 20). Santa Clara, CA, 419\u2013434."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00073"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126933"},{"key":"e_1_3_2_23_2","volume-title":"Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)","author":"Bai Zhihao","year":"2020","unstructured":"Zhihao Bai, Zhen Zhang, Yibo Zhu, and Xin Jin. 2020. PipeSwitch: Fast pipelined context switching for deep learning applications. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)."},{"key":"e_1_3_2_24_2","first-page":"625","volume-title":"Proceedings of the 2022 USENIX Annual Technical Conference (USENIX ATC 22)","author":"Choi Sangjin","year":"2022","unstructured":"Sangjin Choi, Taeksoo Kim, Jinwoo Jeong, Rachata Ausavarungnirun, Myeongjae Jeon, Youngjin Kwon, and Jeongseob Ahn. 2022. Memory harvesting in multi-GPU systems with hierarchical unified virtual memory. In Proceedings of the 2022 USENIX Annual Technical Conference (USENIX ATC 22). Carlsbad, CA, 625\u2013638."},{"key":"e_1_3_2_25_2","first-page":"199","volume-title":"Proceedings of the 2022 USENIX Annual Technical Conference (USENIX ATC 22)","author":"Choi Seungbeom","year":"2022","unstructured":"Seungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park, Youngjin Kwon, and Jaehyuk Huh. 2022. Serving heterogeneous machine learning models on Multi-GPU Servers with Spatio-Temporal Sharing. In Proceedings of the 2022 USENIX Annual Technical Conference (USENIX ATC 22). Carlsbad, CA, 199\u2013216."},{"key":"e_1_3_2_26_2","first-page":"613","volume-title":"Proceedings of the 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17)","author":"Crankshaw Daniel","year":"2017","unstructured":"Daniel Crankshaw, Xin Wang, Guilio Zhou, Michael J. Franklin, Joseph E. Gonzalez, and Ion Stoica. 2017. Clipper: A low-latency online prediction serving system. In Proceedings of the 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). Boston, MA, 613\u2013627."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503222.3507732"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCS.2010.5547126"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS53621.2022.00077"},{"key":"e_1_3_2_31_2","first-page":"135","volume-title":"Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Fu Yao","year":"2024","unstructured":"Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai. 2024. ServerlessLLM: Low-Latency serverless inference for large language models. In Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). Santa Clara, CA, 135\u2013153."},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.5555\/1887695.1887738"},{"key":"e_1_3_2_33_2","first-page":"443","volume-title":"Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)","author":"Gujarati Arpan","year":"2020","unstructured":"Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace. 2020. Serving DNNs like Clockwork: Performance predictability from the bottom Up. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). 443\u2013462."},{"key":"e_1_3_2_34_2","first-page":"539","volume-title":"Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)","author":"Han Mingcong","year":"2022","unstructured":"Mingcong Han, Hanze Zhang, Rong Chen, and Haibo Chen. 2022. Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences. In Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). Carlsbad, CA, 539\u2013558."},{"key":"e_1_3_2_35_1","volume-title":"Proceedings of the 2025 USENIX Annual Technical Conference (USENIX ATC 2025)","author":"Hu Junhao","unstructured":"Junhao Hu, Jiang Xu, Zhixia Liu, Yulong He, Yuetao Chen, Hao Xu, Jiang Liu, Jie Meng, Baoquan Zhang, and Shining Wan. [n. d.]. DEEPSERVE: Serverless large language model serving at scale. In Proceedings of the 2025 USENIX Annual Technical Conference (USENIX ATC 2025)."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378530"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3552326.3567508"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446701"},{"key":"e_1_3_2_39_2","first-page":"463","volume-title":"Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)","author":"Jiang Yimin","year":"2020","unstructured":"Yimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi, Yong Cui, and Chuanxiong Guo. 2020. A unified architecture for accelerating distributed DNN training in heterogeneous GPU\/CPU clusters. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). 463\u2013479."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378529"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/365628.365655"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359654"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613165"},{"key":"e_1_3_2_44_2","first-page":"611","volume-title":"Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Lee Yunseong","year":"2018","unstructured":"Yunseong Lee, Alberto Scolari, Byung-Gon Chun, Marco Domenico Santambrogio, Markus Weimer, and Matteo Interlandi. 2018. PRETZEL: Opening the black box of machine learning prediction serving systems. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). Carlsbad, CA, 611\u2013626."},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3669940.3707251"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378505"},{"key":"e_1_3_2_47_2","volume-title":"Proceedings of the 23rd USENIX Conference on File and Storage Technologies (FAST 25)","author":"Qin Ruoyu","year":"2025","unstructured":"Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. 2025. Mooncake: Trading more storage for less computation \u2014 A KVCache-centric Architecture for Serving LLM Chatbot. In Proceedings of the 23rd USENIX Conference on File and Storage Technologies (FAST 25). Santa Clara, CA."},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/HiPC.2012.6507485"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783721"},{"key":"e_1_3_2_50_2","first-page":"397","volume-title":"Proceedings of 2021 USENIX Annual Technical Conference (USENIX ATC 21)","author":"Romero Francisco","year":"2021","unstructured":"Francisco Romero, Qian Li, Neeraja J. Yadwadkar, and Christos Kozyrakis. 2021. INFaaS: Automated model-less inference serving. In Proceedings of 2021 USENIX Annual Technical Conference (USENIX ATC 21). 397\u2013411."},{"key":"e_1_3_2_51_2","first-page":"205","volume-title":"Proceedings of the 2020 USENIX Annual Technical Conference (USENIX ATC 20)","author":"Shahrad Mohammad","year":"2020","unstructured":"Mohammad Shahrad, Rodrigo Fonseca, Inigo Goiri, Gohar Chaudhry, Paul Batum, Jason Cooke, Eduardo Laureano, Colby Tresness, Mark Russinovich, and Ricardo Bianchini. 2020. Serverless in the Wild: Characterizing and optimizing the serverless workload at a large cloud provider. In Proceedings of the 2020 USENIX Annual Technical Conference (USENIX ATC 20). 205\u2013218."},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359658"},{"key":"e_1_3_2_53_2","first-page":"1288:1\u20131288:23","volume-title":"Proceedings of the 40th International Conference on Machine Learning (ICML 2023)","author":"Sheng Ying","year":"2023","unstructured":"Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher R\u00e9, Ion Stoica, and Ce Zhang. 2023. FlexGen: High-throughput generative inference of large language models with a Single GPU. In Proceedings of the 40th International Conference on Machine Learning (ICML 2023). 1288:1\u20131288:23."},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3627703.3629578"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3542929.3563470"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446714"},{"key":"e_1_3_2_57_2","first-page":"443","volume-title":"Proceedings of 2021 USENIX Annual Technical Conference (USENIX ATC 21)","author":"Wang Ao","year":"2021","unstructured":"Ao Wang, Shuai Chang, Huangshi Tian, Hongqi Wang, Haoran Yang, Huiba Li, Rui Du, and Yue Cheng. 2021. FaaSNet: Scalable and fast provisioning of custom serverless container runtimes at alibaba cloud function compute. In Proceedings of 2021 USENIX Annual Technical Conference (USENIX ATC 21). 443\u2013457."},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3731569.3764813"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3767295.3769336"},{"key":"e_1_3_2_60_2","first-page":"59","volume-title":"Proceedings of the 2024 USENIX Annual Technical Conference (USENIX ATC 24)","author":"Wu Hao","year":"2024","unstructured":"Hao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim, Song Wu, Hao Fan, Ziyue Cheng, and Hai Jin. 2024. StreamBox: A lightweight GPU SandBox for serverless inference workflow. In Proceedings of the 2024 USENIX Annual Technical Conference (USENIX ATC 24). Santa Clara, CA, 59\u201373."},{"key":"e_1_3_2_61_2","first-page":"533","volume-title":"Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)","author":"Xiao Wencong","year":"2020","unstructured":"Wencong Xiao, Shiru Ren, Yong Li, Yang Zhang, Pengyang Hou, Zhi Li, Yihui Feng, Wei Lin, and Yangqing Jia. 2020. AntMan: Dynamic scaling on GPU clusters for deep learning. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). 533\u2013548."},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503222.3507709"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378466"},{"key":"e_1_3_2_64_2","first-page":"1489","volume-title":"Proceedings of the 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23)","author":"Yu Minchen","year":"2023","unstructured":"Minchen Yu, Tingjia Cao, Wei Wang, and Ruichuan Chen. 2023. Following the data, not the function: Rethinking function orchestration in serverless computing. In Proceedings of the 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). Boston, MA, 1489\u20131504."},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS51616.2021.00022"},{"key":"e_1_3_2_66_2","volume-title":"Proceedings of the 2025 USENIX Annual Technical Conference (USENIX ATC 25)","author":"Yu Minchen","year":"2025","unstructured":"Minchen Yu, Ao Wang, Dong Chen, Haoxuan Yu, Xiaonan Luo, Zhuohao Li, Wei Wang, Ruichuan Chen, Dapeng Nie, Haoran Yang, and Yu Ding. 2025. Torpor: GPU-enabled serverless computing for low-latency, resource-efficient inference. In Proceedings of the 2025 USENIX Annual Technical Conference (USENIX ATC 25)."},{"key":"e_1_3_2_67_2","volume-title":"Proceedings of the 9th Annual Conference on Machine Learning and Systems (MLSys 2026)","author":"Yu Minchen","year":"2026","unstructured":"Minchen Yu, Rui Yang, Chaobo Jia, Zhaoyuan Su, Sheng Yao, Tingfeng Lan, Yuchen Yang, Yue Cheng, Wei Wang, Ao Wang, and Ruichuan Chen. 2026. FaaScale: Unlocking fast LLM scaling for serverless inference. In Proceedings of the 9th Annual Conference on Machine Learning and Systems (MLSys 2026)."},{"key":"e_1_3_2_68_2","volume-title":"Proceedings of the 3rd Conference on Machine Learning and Systems (MLSys 2020)","author":"Yu Peifeng","year":"2020","unstructured":"Peifeng Yu and Mosharaf Chowdhury. 2020. Salus: Fine-grained GPU sharing primitives for deep learning applications. In Proceedings of the 3rd Conference on Machine Learning and Systems (MLSys 2020)."},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/3669940.3707285"},{"key":"e_1_3_2_70_2","first-page":"1049","volume-title":"Proceedings of the 2019 USENIX Annual Technical Conference (USENIX ATC 19)","author":"Zhang Chengliang","year":"2019","unstructured":"Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan. 2019. MArk: Exploiting cloud services for cost-effective, SLO-Aware machine learning inference serving. In Proceedings of the 2019 USENIX Annual Technical Conference (USENIX ATC 19). Renton, WA, 1049\u20131062."},{"key":"e_1_3_2_71_2","first-page":"275","volume-title":"Proceedings of the 19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25)","author":"Zhang Dingyan","year":"2025","unstructured":"Dingyan Zhang, Haotian Wang, Yang Liu, Xingda Wei, Yizhou Shan, Rong Chen, and Haibo Chen. 2025. BlitzScale: Fast and live large model autoscaling with O(1) host caching. In Proceedings of the 19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25). 275\u2013293."},{"key":"e_1_3_2_72_2","first-page":"653","volume-title":"Proceedings of the 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21)","author":"Zhang Hong","year":"2021","unstructured":"Hong Zhang, Yupeng Tang, Anurag Khandelwal, Jingrong Chen, and Ion Stoica. 2021. Caerus: NIMBLE task scheduling for serverless analytics. In Proceedings of the 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21). 653\u2013669."},{"key":"e_1_3_2_73_2","volume-title":"Proceedings of the 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23)","author":"Zhang Hong","year":"2023","unstructured":"Hong Zhang, Yupeng Tang, Anurag Khandelwal, and Ion Stoica. 2023. SHEPHERD: Serving DNNs in the wild. In Proceedings of the 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). Boston, MA."},{"key":"e_1_3_2_74_2","first-page":"193","volume-title":"Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Zhong Yinmin","year":"2024","unstructured":"Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024. DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. In Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). Santa Clara, CA, 193\u2013210."}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3800690","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:55:33Z","timestamp":1782402933000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3800690"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":73,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3800690"],"URL":"https:\/\/doi.org\/10.1145\/3800690","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2025-10-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-17","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}