{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T00:16:50Z","timestamp":1777421810153,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":81,"publisher":"ACM","license":[{"start":{"date-parts":[[2024,11,20]],"date-time":"2024-11-20T00:00:00Z","timestamp":1732060800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100006374","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62302300"],"award-info":[{"award-number":["62302300"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]},{"name":"National Key R&D Program of China","award":["2024QY1202"],"award-info":[{"award-number":["2024QY1202"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,11,20]]},"DOI":"10.1145\/3698038.3698510","type":"proceedings-article","created":{"date-parts":[[2024,11,14]],"date-time":"2024-11-14T06:32:43Z","timestamp":1731565963000},"page":"415-433","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["On-demand and Parallel Checkpoint\/Restore for GPU Applications"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-1080-2924","authenticated-orcid":false,"given":"Yanning","family":"Yang","sequence":"first","affiliation":[{"name":"Institute of Parallel and Distributed Systems, SEIEE, Shanghai Jiao Tong University and Engineering Research Center for Domain-specific Operating Systems, Ministry of Education"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7945-8430","authenticated-orcid":false,"given":"Dong","family":"Du","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, SEIEE, Shanghai Jiao Tong University and Engineering Research Center for Domain-specific Operating Systems, Ministry of Education"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2113-2224","authenticated-orcid":false,"given":"Haitao","family":"Song","sequence":"additional","affiliation":[{"name":"Shanghai Artificial Intelligence Research Institute, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6558-5298","authenticated-orcid":false,"given":"Yubin","family":"Xia","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, SEIEE, Shanghai Jiao Tong University and Engineering Research Center for Domain-specific Operating Systems, Ministry of Education"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,11,20]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2024. AMD GPU DMA buffer. https:\/\/github.com\/torvalds\/linux\/blob\/v6.8\/drivers\/gpu\/drm\/amd\/amdgpu\/amdgpu_dma_buf.h. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_2_1","unstructured":"2024. Apache OpenWhisk is a serverless open source cloud platform. http:\/\/openwhisk.apache.org\/. Referenced 2024."},{"key":"e_1_3_2_1_3_1","unstructured":"2024. Apache TVM. https:\/\/tvm.apache.org\/. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_4_1","volume-title":"AWS Lambda - Serverless Compute. https:\/\/aws.amazon.com\/lambda\/. Referenced","year":"2024","unstructured":"2024. AWS Lambda - Serverless Compute. https:\/\/aws.amazon.com\/lambda\/. Referenced Jan. 2024."},{"key":"e_1_3_2_1_5_1","unstructured":"2024. Best practices for GPU-accelerated instances. https:\/\/www.alibabacloud.com\/help\/en\/function-compute\/latest\/development-guide. Accessed: 2024-07-15."},{"key":"e_1_3_2_1_6_1","unstructured":"2024. Buffer Sharing and Synchronization (dma-buf) --- The Linux Kernel documentation. https:\/\/docs.kernel.org\/driver-api\/dma-buf.html. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_7_1","unstructured":"2024. Compute - Amazon EC2 Instance Types - AWS. https:\/\/aws.amazon.com\/ec2\/instance-types\/. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_8_1","unstructured":"2024. CR in namespace - CRIU. https:\/\/criu.org\/CR_in_namespace. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_9_1","unstructured":"2024. CRIU. https:\/\/criu.org\/Main_Page. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_10_1","unstructured":"2024. CUDA Toolkit - Free Tools and Training | NVIDIA Developer. https:\/\/developer.nvidia.com\/cuda-toolkit. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_11_1","unstructured":"2024. Google gVisor: Application Kernel for Containers. https:\/\/github.com\/google\/gvisor. Referenced 2024-07-16."},{"key":"e_1_3_2_1_12_1","unstructured":"2024. GPU memory --- ROCm Documentation. https:\/\/ROCm.docs.amd.com\/en\/latest\/conceptual\/gpu-memory.html. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_13_1","unstructured":"2024. GPUDirect | NVIDIA Developer. https:\/\/developer.nvidia.com\/gpudirect. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_14_1","unstructured":"2024. Heterogeneous Memory Management (HMM) --- The Linux Kernel documentation. https:\/\/www.kernel.org\/doc\/html\/v5.0\/vm\/hmm.html. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_15_1","unstructured":"2024. HIP Runtime API Reference: PeerToPeer Device Memory Access --- HIP 6.2.41134 Documentation. https:\/\/rocm.docs.amd.com\/projects\/HIP\/en\/latest\/doxygen\/html\/group___peer_to_peer.html. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_16_1","unstructured":"2024. hoytech\/vmtouch: Portable file system cache diagnostics and control. https:\/\/github.com\/hoytech\/vmtouch. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_17_1","unstructured":"2024. Introducing ChatGPT | OpenAI. https:\/\/openai.com\/index\/chatgpt\/. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_18_1","unstructured":"2024. Models and pre-trained weights --- Torchvision 0.18 documentation. https:\/\/PyTorch.org\/vision\/stable\/models.html. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_19_1","unstructured":"2024. NVIDIA Tesla V100 | NVIDIA. https:\/\/www.nvidia.com\/en-gb\/data-center\/tesla-v100\/. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_20_1","unstructured":"2024. NVIDIA\/cuda-checkpoint: CUDA checkpoint and restore utility. https:\/\/github.com\/NVIDIA\/cuda-checkpoint. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_21_1","unstructured":"2024. NVLink & NVSwitch: Fastest HPC Data Center Platform | NVIDIA. https:\/\/www.nvidia.com\/en-us\/data-center\/nvlink\/. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_22_1","unstructured":"2024. opencontainers\/runc. https:\/\/github.com\/opencontainers\/runc.git. Referenced 2024-07-15."},{"key":"e_1_3_2_1_23_1","unstructured":"2024. ROCm Software. https:\/\/www.amd.com\/en\/products\/software\/rocm.html. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_24_1","unstructured":"2024. ROCm\/rocBLAS: Next generation BLAS implementation for ROCm platform. https:\/\/github.com\/ROCm\/rocBLAS. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_25_1","unstructured":"2024. Safetensors: ML Safer for All. https:\/\/github.com\/huggingface\/safetensors. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_26_1","volume-title":"firecracker-microvm\/firecracker. https:\/\/github.com\/firecracker-microvm\/firecracker\/issues\/1184. Referenced","author":"Full","year":"2024","unstructured":"2024. [Snaps] Full snapshot + restore, firecracker-microvm\/firecracker. https:\/\/github.com\/firecracker-microvm\/firecracker\/issues\/1184. Referenced April 2024."},{"key":"e_1_3_2_1_27_1","unstructured":"2024. Sora | OpenAI. https:\/\/openai.com\/index\/sora\/. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_28_1","unstructured":"2024. Supporting ROCm with CRIU. https:\/\/github.com\/checkpoint-restore\/criu\/tree\/criu-dev\/plugins\/amdgpu. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_29_1","unstructured":"2024. Unified Memory for CUDA Beginners | NVIDIA Technical Blog. https:\/\/developer.nvidia.com\/blog\/unified-memory-cuda-beginners\/. Accessed: 2024-07-04."},{"key":"e_1_3_2_1_30_1","volume-title":"17th USENIX symposium on networked systems design and implementation (NSDI 20)","author":"Agache Alexandru","year":"2020","unstructured":"Alexandru Agache, Marc Brooker, Alexandra Iordache, Anthony Liguori, Rolf Neugebauer, Phil Piwonka, and Diana-Maria Popa. 2020. Firecracker: Lightweight virtualization for serverless applications. In 17th USENIX symposium on networked systems design and implementation (NSDI 20). 419--434."},{"key":"e_1_3_2_1_31_1","volume-title":"SAND: Towards High-Performance Serverless Computing. In 2018 USENIX Annual Technical Conference (USENIX ATC 18)","author":"Akkus Istemi Ekin","year":"2018","unstructured":"Istemi Ekin Akkus, Ruichuan Chen, Ivica Rimac, Manuel Stein, Klaus Satzke, Andre Beck, Paarijaat Aditya, and Volker Hilt. 2018. SAND: Towards High-Performance Serverless Computing. In 2018 USENIX Annual Technical Conference (USENIX ATC 18). 923--935."},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00073"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3492321.3524270"},{"key":"e_1_3_2_1_34_1","unstructured":"The KServe Authors. 2023. KServe. https:\/\/github.com\/kserve\/kserve. Accessed on 2024-06-22."},{"key":"e_1_3_2_1_35_1","volume-title":"PipeSwitch: Fast Pipelined Context Switching for Deep Learning Applications. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)","author":"Bai Zhihao","year":"2020","unstructured":"Zhihao Bai, Zhen Zhang, Yibo Zhu, and Xin Jin. 2020. PipeSwitch: Fast Pipelined Context Switching for Deep Learning Applications. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). USENIX Association, 499--514. https:\/\/www.usenix.org\/conference\/osdi20\/presentation\/bai"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3472883.3486992"},{"key":"e_1_3_2_1_37_1","volume-title":"Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal Sharing. In 2022 USENIX Annual Technical Conference (USENIX ATC 22)","author":"Choi Seungbeom","year":"2022","unstructured":"Seungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park, Youngjin Kwon, and Jaehyuk Huh. 2022. Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal Sharing. In 2022 USENIX Annual Technical Conference (USENIX ATC 22). USENIX Association, Carlsbad, CA, 199--216. https:\/\/www.usenix.org\/conference\/atc22\/presentation\/choi-seungbeom"},{"key":"e_1_3_2_1_38_1","volume-title":"Clipper: A Low-Latency Online Prediction Serving System. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17)","author":"Crankshaw Daniel","year":"2017","unstructured":"Daniel Crankshaw, Xin Wang, Guilio Zhou, Michael J. Franklin, Joseph E. Gonzalez, and Ion Stoica. 2017. Clipper: A Low-Latency Online Prediction Serving System. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). USENIX Association, Boston, MA, 613--627. https:\/\/www.usenix.org\/conference\/nsdi17\/technical-sessions\/presentation\/crankshaw"},{"key":"e_1_3_2_1_39_1","volume-title":"DVABatch: Diversity-aware Multi-Entry Multi-Exit Batching for Efficient Processing of DNN Services on GPUs. In 2022 USENIX Annual Technical Conference (USENIX ATC 22)","author":"Cui Weihao","year":"2022","unstructured":"Weihao Cui, Han Zhao, Quan Chen, Hao Wei, Zirui Li, Deze Zeng, Chao Li, and Minyi Guo. 2022. DVABatch: Diversity-aware Multi-Entry Multi-Exit Batching for Efficient Processing of DNN Services on GPUs. In 2022 USENIX Annual Technical Conference (USENIX ATC 22). USENIX Association, Carlsbad, CA, 183--198. https:\/\/www.usenix.org\/conference\/atc22\/presentation\/cui"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3419111.3421284"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503222.3507732"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378512"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS53621.2022.00077"},{"key":"e_1_3_2_1_44_1","volume-title":"ServerlessLLM: Low-Latency Serverless Inference for Large Language Models. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Fu Yao","year":"2024","unstructured":"Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai. 2024. ServerlessLLM: Low-Latency Serverless Inference for Large Language Models. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). USENIX Association, Santa Clara, CA, 135--153. https:\/\/www.usenix.org\/conference\/osdi24\/presentation\/fu"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446757"},{"key":"e_1_3_2_1_46_1","volume-title":"14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)","author":"Gujarati Arpan","year":"2020","unstructured":"Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace. 2020. Serving DNNs like Clockwork: Performance Predictability from the Bottom Up. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). USENIX Association, 443--462. https:\/\/www.usenix.org\/conference\/osdi20\/presentation\/gujarati"},{"key":"e_1_3_2_1_47_1","volume-title":"Microsecond-scale Preemption for Concurrent GPU-accelerated DNN Inferences. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)","author":"Han Mingcong","year":"2022","unstructured":"Mingcong Han, Hanze Zhang, Rong Chen, and Haibo Chen. 2022. Microsecond-scale Preemption for Concurrent GPU-accelerated DNN Inferences. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). USENIX Association, Carlsbad, CA, 539--558. https:\/\/www.usenix.org\/conference\/osdi22\/presentation\/han"},{"key":"e_1_3_2_1_48_1","volume-title":"2018 IEEE International Conference on Cloud Engineering (IC2E)","author":"Isahagian Vatche","year":"2017","unstructured":"Vatche Isahagian, Vinod Muthusamy, and Aleksander Slominski. 2017. Serving Deep Learning Models in a Serverless Platform. 2018 IEEE International Conference on Cloud Engineering (IC2E) (2017), 257--262. https:\/\/api.semanticscholar.org\/CorpusID:21724528"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3552326.3567508"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/2731186.2731192"},{"key":"e_1_3_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3342195.3387547"},{"key":"e_1_3_2_1_52_1","volume-title":"2022 USENIX Annual Technical Conference (USENIX ATC 22)","author":"Li Jie","year":"2022","unstructured":"Jie Li, Laiping Zhao, Yanan Yang, Kunlin Zhan, and Keqiu Li. 2022. Tetris: Memory-efficient Serverless Inference through Tensor Sharing. In 2022 USENIX Annual Technical Conference (USENIX ATC 22). USENIX Association, Carlsbad, CA. https:\/\/www.usenix.org\/conference\/atc22\/presentation\/li-jie"},{"key":"e_1_3_2_1_53_1","volume-title":"Help Rather Than Recycle: Alleviating Cold Startup in Serverless Computing Through Inter-Function Container Sharing. In 2022 USENIX Annual Technical Conference (USENIX ATC 22)","author":"Li Zijun","year":"2022","unstructured":"Zijun Li, Linsong Guo, Quan Chen, Jiagan Cheng, Chuhao Xu, Deze Zeng, Zhuo Song, Tao Ma, Yong Yang, Chao Li, and Minyi Guo. 2022. Help Rather Than Recycle: Alleviating Cold Startup in Serverless Computing Through Inter-Function Container Sharing. In 2022 USENIX Annual Technical Conference (USENIX ATC 22). USENIX Association, Carlsbad, CA, 69--84. https:\/\/www.usenix.org\/conference\/atc22\/presentation\/li-zijun-help"},{"key":"e_1_3_2_1_54_1","volume-title":"AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23)","author":"Li Zhuohan","year":"2023","unstructured":"Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23). USENIX Association, Boston, MA, 663--679. https:\/\/www.usenix.org\/conference\/osdi23\/presentation\/li-zhouhan"},{"key":"e_1_3_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3477113.3487273"},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3620678.3624785"},{"key":"e_1_3_2_1_57_1","volume-title":"Harmonizing Efficiency and Practicability: Optimizing Resource Utilization in Serverless Computing with Jiagu. In 2024 USENIX Annual Technical Conference (USENIX ATC 24)","author":"Liu Qingyuan","year":"2024","unstructured":"Qingyuan Liu, Yanning Yang, Dong Du, Yubin Xia, Ping Zhang, Jia Feng, James R. Larus, and Haibo Chen. 2024. Harmonizing Efficiency and Practicability: Optimizing Resource Utilization in Serverless Computing with Jiagu. In 2024 USENIX Annual Technical Conference (USENIX ATC 24). USENIX Association, Santa Clara, CA, 1--17. https:\/\/www.usenix.org\/conference\/atc24\/presentation\/liu-qingyuan"},{"key":"e_1_3_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503222.3507752"},{"key":"e_1_3_2_1_59_1","unstructured":"Microsoft. 2023. Azure ML. https:\/\/learn.microsoft.com\/en-us\/azure\/machine-learning. Accessed on 2024-06-22."},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2020.01.004"},{"key":"e_1_3_2_1_61_1","volume-title":"SOCK: Rapid Task Provisioning with Serverless-Optimized Containers. In 2018 USENIX Annual Technical Conference (USENIX ATC 18)","author":"Oakes Edward","year":"2018","unstructured":"Edward Oakes, Leon Yang, Dennis Zhou, Kevin Houck, Tyler Harter, Andrea Arpaci-Dusseau, and Remzi Arpaci-Dusseau. 2018. SOCK: Rapid Task Provisioning with Serverless-Optimized Containers. In 2018 USENIX Annual Technical Conference (USENIX ATC 18). 57--70."},{"key":"e_1_3_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3620678.3624664"},{"key":"e_1_3_2_1_63_1","volume-title":"INFaaS: Automated Model-less Inference Serving. In 2021 USENIX Annual Technical Conference (USENIX ATC 21)","author":"Romero Francisco","year":"2021","unstructured":"Francisco Romero, Qian Li, Neeraja J. Yadwadkar, and Christos Kozyrakis. 2021. INFaaS: Automated Model-less Inference Serving. In 2021 USENIX Annual Technical Conference (USENIX ATC 21). USENIX Association, 397--411. https:\/\/www.usenix.org\/conference\/atc21\/presentation\/romero"},{"key":"e_1_3_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503222.3507750"},{"key":"e_1_3_2_1_65_1","unstructured":"AWS SageMaker. 2023. Machine Learning Service - Amazon SageMaker. https:\/\/aws.amazon.com\/pm\/sagemaker\/."},{"key":"e_1_3_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3492321.3524272"},{"key":"e_1_3_2_1_67_1","volume-title":"2020 USENIX Annual Technical Conference (USENIX ATC 20)","author":"Shahrad Mohammad","year":"2020","unstructured":"Mohammad Shahrad, Rodrigo Fonseca, \u00cd\u00f1igo Goiri, Gohar Chaudhry, Paul Batum, Jason Cooke, Eduardo Laureano, Colby Tresness, Mark Russinovich, and Ricardo Bianchini. 2020. Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider. In 2020 USENIX Annual Technical Conference (USENIX ATC 20). 205--218."},{"key":"e_1_3_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359658"},{"key":"e_1_3_2_1_69_1","volume-title":"2020 USENIX Annual Technical Conference (USENIX ATC 20)","author":"Shillaker Simon","year":"2020","unstructured":"Simon Shillaker and Peter Pietzuch. 2020. Faasm: lightweight isolation for efficient stateful serverless computing. In 2020 USENIX Annual Technical Conference (USENIX ATC 20). 419--433."},{"key":"e_1_3_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/3423211.3425682"},{"key":"e_1_3_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446714"},{"key":"e_1_3_2_1_72_1","first-page":"1","volume-title":"Proceedings of the Fourteenth EuroSys Conference","author":"Amy Wang Kai-Ting","year":"2019","unstructured":"Kai-Ting Amy Wang, Rayson Ho, and Peng Wu. 2019. Replayable execution optimized for page sharing for a managed runtime environment. In Proceedings of the Fourteenth EuroSys Conference 2019.1-16."},{"key":"e_1_3_2_1_73_1","volume-title":"KR-CORE: A Microsecond-scale RDMA Control Plane for Elastic Computing. In 2022 USENIX Annual Technical Conference (USENIX ATC 22)","author":"Wei Xingda","year":"2022","unstructured":"Xingda Wei, Fangming Lu, Rong Chen, and Haibo Chen. 2022. KR-CORE: A Microsecond-scale RDMA Control Plane for Elastic Computing. In 2022 USENIX Annual Technical Conference (USENIX ATC 22). USENIX Association, Carlsbad, CA, 121--136. https:\/\/www.usenix.org\/conference\/atc22\/presentation\/wei"},{"key":"e_1_3_2_1_74_1","volume-title":"No Provisioned Concurrency: Fast RDMA-codesigned Remote Fork for Serverless Computing. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23)","author":"Wei Xingda","year":"2023","unstructured":"Xingda Wei, Fangming Lu, Tianxia Wang, Jinyu Gu, Yuhan Yang, Rong Chen, and Haibo Chen. 2023. No Provisioned Concurrency: Fast RDMA-codesigned Remote Fork for Serverless Computing. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23). USENIX Association, Boston, MA, 497--517. https:\/\/www.usenix.org\/conference\/osdi23\/presentation\/wei-rdma"},{"key":"e_1_3_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"key":"e_1_3_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503222.3507709"},{"key":"e_1_3_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/3617232.3624871"},{"key":"e_1_3_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS51616.2021.00022"},{"key":"e_1_3_2_1_79_1","unstructured":"Minchen Yu Ao Wang Dong Chen Haoxuan Yu Xiaonan Luo Zhuohao Li Wei Wang Ruichuan Chen Dapeng Nie and Haoran Yang. 2024. FaaSwap: SLO-Aware GPU-Efficient Serverless Inference via Model Swapping. arXiv:2306.03622 [cs.DC] https:\/\/arxiv.org\/abs\/2306.03622"},{"key":"e_1_3_2_1_80_1","volume-title":"Proceedings of the 2019 USENIX Conference on Usenix Annual Technical Conference (Renton, WA, USA) (USENIX ATC '19). USENIX Association, USA, 1049--1062","author":"Zhang Chengliang","year":"2019","unstructured":"Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan. 2019. MArk: exploiting cloud services for cost-effective, SLO-aware machine learning inference serving. In Proceedings of the 2019 USENIX Conference on Usenix Annual Technical Conference (Renton, WA, USA) (USENIX ATC '19). USENIX Association, USA, 1049--1062."},{"key":"e_1_3_2_1_81_1","unstructured":"Han Zhao Weihao Cui Quan Chen Shulai Zhang Zijun Li Jingwen Leng Chao Li Deze Zeng and Minyi Guo. 2024. Towards Fast Setup and High Throughput of GPU Serverless Computing. arXiv:2404.14691 [cs.DC] https:\/\/arxiv.org\/abs\/2404.14691"}],"event":{"name":"SoCC '24: ACM Symposium on Cloud Computing","location":"Redmond WA USA","acronym":"SoCC '24","sponsor":["SIGMOD ACM Special Interest Group on Management of Data","SIGOPS ACM Special Interest Group on Operating Systems"]},"container-title":["Proceedings of the ACM Symposium on Cloud Computing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3698038.3698510","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3698038.3698510","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T19:01:05Z","timestamp":1755889265000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3698038.3698510"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,20]]},"references-count":81,"alternative-id":["10.1145\/3698038.3698510","10.1145\/3698038"],"URL":"https:\/\/doi.org\/10.1145\/3698038.3698510","relation":{},"subject":[],"published":{"date-parts":[[2024,11,20]]},"assertion":[{"value":"2024-11-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}