{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T02:34:30Z","timestamp":1785465270646,"version":"3.56.0"},"reference-count":123,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62432010, 62272291, 62132014"],"award-info":[{"award-number":["62432010, 62272291, 62132014"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Comput. Syst."],"published-print":{"date-parts":[[2026,2,28]]},"abstract":"<jats:p>Many intelligent applications, such as autonomous driving and virtual reality, require running both latency-critical (real-time) and best-effort deep neural network (DNN) inference tasks to achieve both real-time and work-conserving on the GPU. However, commodity GPUs lack efficient preemptive scheduling support, and existing state-of-the-art approaches either have to monopolize GPU or let real-time tasks to wait for best-effort tasks to complete, resulting in low utilization, high latency, or both.<\/jats:p>\n                  <jats:p>\n                    This article presents\n                    <jats:sc>Reef<\/jats:sc>\n                    , the first GPU-accelerated DNN inference serving system that achieves low-latency and work-conserving for concurrent real-time and best-effort tasks.\n                    <jats:sc>Reef<\/jats:sc>\n                    accomplishes this by enabling microsecond-scale kernel preemption and controlled concurrent execution in GPU scheduling.\n                    <jats:sc>Reef<\/jats:sc>\n                    is novel in two ways. First, based on the observation that DNN inference kernels are mostly idempotent,\n                    <jats:sc>Reef<\/jats:sc>\n                    devises a reset-based preemption scheme that launches a real-time kernel on the GPU by proactively killing and restoring best-effort kernels at microsecond-scale. Second, since DNN inference kernels have varied parallelism and predictable latency,\n                    <jats:sc>Reef<\/jats:sc>\n                    proposes a dynamic kernel padding mechanism that dynamically pads the real-time kernel with appropriate best-effort kernels to fully utilize the GPU with negligible overhead. Evaluation using a new DNN inference serving benchmark (DISB) with diverse workloads and a real-world trace on both NVIDIA and AMD GPUs shows that\n                    <jats:sc>Reef<\/jats:sc>\n                    only incurs less than 5% overhead in end-to-end latency for real-time tasks but increases the overall throughput by up to 1.53\u00d7, compared to scheduling tasks sequentially. To demonstrate the practical benefits of our approach, we compare\n                    <jats:sc>Reef<\/jats:sc>\n                    with Triton, a widely-adopted production-level serving system. Our evaluation shows that\n                    <jats:sc>Reef<\/jats:sc>\n                    outperforms Triton by 1.12\u00d7 to 5.20\u00d7 in end-to-end latency for real-time tasks, while maintaining comparable throughput.\n                  <\/jats:p>","DOI":"10.1145\/3768622","type":"journal-article","created":{"date-parts":[[2025,9,18]],"date-time":"2025-09-18T11:36:32Z","timestamp":1758195392000},"page":"1-42","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Real-time, Work-conserving GPU Scheduling for Concurrent DNN Inference"],"prefix":"10.1145","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-1536-7485","authenticated-orcid":false,"given":"Mingcong","family":"Han","sequence":"first","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6115-8130","authenticated-orcid":false,"given":"Rong","family":"Chen","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-2568-177X","authenticated-orcid":false,"given":"Weihang","family":"Shen","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-0579-5707","authenticated-orcid":false,"given":"Hanze","family":"Zhang","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-7732-8773","authenticated-orcid":false,"given":"Jinrong","family":"Yang","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9720-0361","authenticated-orcid":false,"given":"Haibo","family":"Chen","sequence":"additional","affiliation":[{"name":"Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University","place":["Shanghai, China"]},{"name":"China and Key Laboratory of System Software, Chinese Academy of Sciences","place":["Shanghai, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,11,7]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","unstructured":"Jacob T. Adriaens Katherine Compton Nam Sung Kim and Michael J. Schulte. 2012. The case for GPGPU spatial multitasking. In Proceedings of the 2012 IEEE 18th International Symposium on High-Performance Computer Architecture (HPCA\u201912). IEEE Computer Society USA 1\u201312. DOI:10.1109\/HPCA.2012.6168946","DOI":"10.1109\/HPCA.2012.6168946"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","unstructured":"Miguel Alcon Hamid Tabani Leonidas Kosmidis Enrico Mezzetti Jaume Abella and Francisco J. Cazorla. 2020. Timing of autonomous driving software: problem analysis and prospects for future solutions. In 2020 IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS). 267\u2013280. DOI:10.1109\/RTAS48715.2020.000-1","DOI":"10.1109\/RTAS48715.2020.000-1"},{"key":"e_1_3_2_4_2","unstructured":"AMD ROCm. 2022. AMD ROCm Platform Documentation. Retrieved from https:\/\/rocmdocs.amd.com\/"},{"key":"e_1_3_2_5_2","unstructured":"Apollo Auto. 2022. Apollo: Architecture\/Hardware Connection. Retrieved from https:\/\/github.com\/ApolloAuto\/apollo"},{"key":"e_1_3_2_6_2","unstructured":"Apollo Auto. 2022. Apollo Perception Module. Retrieved from https:\/\/github.com\/ApolloAuto\/apollo\/tree\/master\/modules\/perception"},{"key":"e_1_3_2_7_2","unstructured":"Junjie Bai Fang Lu Ke Zhang et\u00a0al. 2019. ONNX: Open Neural Network Exchange. Retrieved from https:\/\/github.com\/onnx\/onnx"},{"key":"e_1_3_2_8_2","first-page":"499","volume-title":"14th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201920)","author":"Bai Zhihao","year":"2020","unstructured":"Zhihao Bai, Zhen Zhang, Yibo Zhu, and Xin Jin. 2020. PipeSwitch: Fast pipelined context switching for deep learning applications. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201920). 499\u2013514. Retrieved from https:\/\/www.usenix.org\/conference\/osdi20\/presentation\/bai"},{"key":"e_1_3_2_9_2","unstructured":"Baidu. 2022. Apollo. Retrieved from https:\/\/apollo.auto\/"},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","first-page":"54","DOI":"10.1109\/RTAS58335.2023.00012","article-title":"Hardware compute partitioning on NVIDIA GPUs*","author":"Bakita Joshua","year":"2023","unstructured":"Joshua Bakita and James H. Anderson. 2023. Hardware compute partitioning on NVIDIA GPUs*. 2023 IEEE 29th Real-Time and Embedded Technology and Applications Symposium (RTAS) (2023), 54\u201366. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:259235797","journal-title":"2023 IEEE 29th Real-Time and Embedded Technology and Applications Symposium (RTAS)"},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"287","DOI":"10.1109\/ECRTS.2012.15","volume-title":"24th Euromicro Conference on Real-Time Systems (ECRTS\u201912)","author":"Basaran C.","year":"2012","unstructured":"C. Basaran and K. Kang. 2012. Supporting preemptive task executions and memory copies in GPGPUs. In 24th Euromicro Conference on Real-Time Systems (ECRTS\u201912). 287\u2013296."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","unstructured":"Karsten Behrendt Libor Novak and Rami Botros. 2017. A deep learning approach to traffic lights: Detection tracking and classification. In 2017 IEEE International Conference on Robotics and Automation (ICRA). 1370\u20131377. DOI:10.1109\/ICRA.2017.7989163","DOI":"10.1109\/ICRA.2017.7989163"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","unstructured":"Nicola Capodieci Roberto Cavicchioli Marko Bertogna and Aingara Paramakuru. 2018. Deadline-Based Scheduling for GPU with Preemption Support. In 2018 IEEE Real-Time Systems Symposium (RTSS). 119\u2013130. DOI:10.1109\/RTSS.2018.00021","DOI":"10.1109\/RTSS.2018.00021"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","unstructured":"Shuai Che Michael Boyer Jiayuan Meng David Tarjan Jeremy W. Sheaffer Sang-Ha Lee and Kevin Skadron. 2009. Rodinia: A benchmark suite for heterogeneous computing. In 2009 IEEE International Symposium on Workload Characterization (IISWC). 44\u201354. DOI:10.1109\/IISWC.2009.5306797","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","unstructured":"Guoyang Chen Yue Zhao Xipeng Shen and Huiyang Zhou. 2017. EffiSha: A software framework for enabling effficient preemptive scheduling of GPU. In Proceedings of the 22nd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP\u201917). Association for Computing Machinery Austin Texas USA 3\u201316. DOI:10.1145\/3018743.3018748","DOI":"10.1145\/3018743.3018748"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","unstructured":"Quan Chen Hailong Yang Minyi Guo Ram Srivatsa Kannan Jason Mars and Lingjia Tang. 2017. Prophet: Precise QoS prediction on non-preemptive accelerators to improve utilization in warehouse-scale computers. In Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201917). Association for Computing Machinery Xi\u2019an China 17\u201332. DOI:10.1145\/3037697.3037700","DOI":"10.1145\/3037697.3037700"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","unstructured":"Quan Chen Hailong Yang Jason Mars and Lingjia Tang. 2016. Baymax: QoS awareness and increased utilization for non-preemptive accelerators in warehouse scale computers. In Proceedings of the Twenty-First International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201916). Association for Computing Machinery Atlanta Georgia USA 681\u2013696. DOI:10.1145\/2872362.2872368","DOI":"10.1145\/2872362.2872368"},{"key":"e_1_3_2_18_2","volume-title":"13th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201918)","author":"Chen T.","year":"2018","unstructured":"T. Chen, T. Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Q. Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, L. Ceze, Carlos Guestrin, and A. Krishnamurthy. 2018. TVM: An automated end-to-end optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201918)."},{"key":"e_1_3_2_19_2","unstructured":"Sharan Chetlur Cliff Woolley Philippe Vandermersch Jonathan M. Cohen John Tran Bryan Catanzaro and Evan Shelhamer. 2014. cuDNN: Efficient primitives for deep learning. arXiv:1410.0759. Retrieved from https:\/\/arxiv.org\/abs\/\/11410.0759"},{"key":"e_1_3_2_20_2","unstructured":"Green Car Congress. 2017. New ultrafast camera for self-driving vehicles and drones. Retrieved from https:\/\/www.greencarcongress.com\/2017\/02\/20170217-ntu.html"},{"key":"e_1_3_2_21_2","unstructured":"Brian F. Cooper. 2022. YCSB Core Workloads. Retrieved from https:\/\/github.com\/brianfrankcooper\/YCSB\/wiki\/Core-Workloads"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/1807128.1807152"},{"key":"e_1_3_2_23_2","unstructured":"NVIDIA Corporation. Triton Inference Server: An Optimized Cloud and Edge Inferencing Solution. Retrieved from https:\/\/github.com\/triton-inference-server. (n. d.)."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","unstructured":"Alexander Craik Yongtian He and Jos\u00e9 Luis Contreras-Vidal. 2019. Deep learning for electroencephalogram (EEG) classification tasks: A review. Journal of Neural Engineering 16 3 (June 2019) 031001. DOI:10.1088\/1741-2552\/ab0ab5","DOI":"10.1088\/1741-2552\/ab0ab5"},{"key":"e_1_3_2_25_2","volume-title":"14th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201917)","author":"Crankshaw D.","year":"2017","unstructured":"D. Crankshaw, Xin Wang, Giulio Zhou, M. Franklin, Joseph E. Gonzalez, and I. Stoica. 2017. Clipper: A low-latency online prediction serving system. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI\u201917)."},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","unstructured":"Weihao Cui Mengze Wei Quan Chen Xiaoxin Tang Jingwen Leng Li Li and Mingyi Guo. 2019. Ebird: Elastic batch for improving responsiveness and throughput of deep learning services. In 2019 IEEE 37th International Conference on Computer Design (ICCD). 497\u2013505. DOI:10.1109\/ICCD46524.2019.00075","DOI":"10.1109\/ICCD46524.2019.00075"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","unstructured":"Weihao Cui Han Zhao Quan Chen Ningxin Zheng Jingwen Leng Jieru Zhao Zhuo Song Tao Ma Yong Yang Chao Li and Minyi Guo. 2021. Enable simultaneous DNN services based on deterministic operator overlap and precise latency prediction. In Proceedings of the International Conference for High Performance Computing Networking Storage and Analysis (SC\u201921). Association for Computing Machinery St. Louis Missouri. DOI:10.1145\/3458817.3476143","DOI":"10.1145\/3458817.3476143"},{"key":"e_1_3_2_28_2","doi-asserted-by":"crossref","first-page":"475","DOI":"10.1145\/2254064.2254120","volume-title":"33rd ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI\u201912)","author":"Kruijf Marc A. de","year":"2012","unstructured":"Marc A. de Kruijf, Karthikeyan Sankaralingam, and Somesh Jha. 2012. Static analysis and compiler design for idempotent processing. In 33rd ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI\u201912). 475\u2013486. DOI:10.1145\/2254064.2254120"},{"key":"e_1_3_2_29_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/\/1810.04805"},{"key":"e_1_3_2_30_2","unstructured":"ROCm documentation. 2022. GCN Native ISA LLVM Code Generator: Kernel Dispatch. Retrieved from https:\/\/rocmdocs.amd.com\/en\/latest\/ROCm_Compiler_SDK\/ROCm-Native-ISA.html"},{"issue":"1","key":"e_1_3_2_31_2","doi-asserted-by":"crossref","first-page":"24","DOI":"10.1038\/s41591-018-0316-z","article-title":"A guide to deep learning in healthcare","volume":"25","author":"Esteva Andre","year":"2019","unstructured":"Andre Esteva, Alexandre Robicquet, Bharath Ramsundar, Volodymyr Kuleshov, Mark DePristo, Katherine Chou, Claire Cui, Greg Corrado, Sebastian Thrun, and Jeff Dean. 2019. A guide to deep learning in healthcare. Nature medicine 25, 1 (2019), 24\u201329.","journal-title":"Nature medicine"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","unstructured":"Jiarui Fang Yang Yu Chengduo Zhao and Jie Zhou. 2021. TurboTransformers: an efficient GPU serving system for transformer models. In Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP\u201921). Association for Computing Machinery Virtual Event Republic of Korea 389\u2013402. DOI:10.1145\/3437801.3441578","DOI":"10.1145\/3437801.3441578"},{"key":"e_1_3_2_33_2","first-page":"135","volume-title":"18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Fu Yao","year":"2024","unstructured":"Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai. 2024. [ServerlessLLM]:[Low-Latency] serverless inference for large language models. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). 135\u2013153."},{"key":"e_1_3_2_34_2","unstructured":"Google. 2024. Gemma-2B-IT. Retrieved from https:\/\/huggingface.co\/google\/gemma-2b-it"},{"key":"e_1_3_2_35_2","volume-title":"4th USENIX Workshop on Hot Topics in Parallelism (HotPar\u201912)","author":"Gregg Chris","year":"2012","unstructured":"Chris Gregg, Jonathan Dorn, K. Hazelwood, and K. Skadron. 2012. Fine-grained resource sharing for concurrent GPGPU kernels. In 4th USENIX Workshop on Hot Topics in Parallelism (HotPar\u201912)."},{"key":"e_1_3_2_36_2","volume-title":"14th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201920)","author":"Gujarati A.","year":"2020","unstructured":"A. Gujarati, Reza Karimi, Safya Alzayat, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace. 2020. Serving DNNs like Clockwork: Performance predictability from the bottom up. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201920)."},{"key":"e_1_3_2_37_2","unstructured":"Mingcong Han Weihang Shen Rong Chen Binyu Zang and Haibo Chen. 2025. Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained Multi-XPU Abstraction. (2025). arXiv:2508.09503. Retrieved from https:\/\/arxiv.org\/abs\/\/2508.09503"},{"key":"e_1_3_2_38_2","unstructured":"Mingcong Han Weihang Shen Guanwen Peng Rong Chen and Haibo Chen. 2024. Microsecond-scale Dynamic Validation of Idempotency for GPU Kernels. (2024). arXiv:2410.23661. Retrieved from https:\/\/arxiv.org\/abs\/\/2410.23661"},{"key":"e_1_3_2_39_2","first-page":"539","volume-title":"16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)","author":"Han Mingcong","year":"2022","unstructured":"Mingcong Han, Hanze Zhang, Rong Chen, and Haibo Chen. 2022. Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). USENIX Association, Carlsbad, CA, 539\u2013558. Retrieved from https:\/\/www.usenix.org\/conference\/osdi22\/presentation\/han"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","unstructured":"Johann Hauswald Yiping Kang Michael A. Laurenzano Quan Chen Cheng Li Trevor Mudge Ronald G. Dreslinski Jason Mars and Lingjia Tang. 2015. DjiNN and Tonic: DNN as a service and its implications for future warehouse scale computers. In Proceedings of the 42nd Annual International Symposium on Computer Architecture (ISCA\u201915) Association for Computing Machinery Portland Oregon 27\u201340. DOI:10.1145\/2749469.2749472","DOI":"10.1145\/2749469.2749472"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","unstructured":"Kaiming He X. Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (CVPR\u201916). 770\u2013778. DOI:10.1109\/CVPR.2016.90","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_42_2","unstructured":"Jeremy Hermann and Mike Del Balso. 2017. Meet Michelangelo: Uber\u2019s Machine Learning Platform. Retrieved from https:\/\/eng.uber.com\/michelangelo-machine-learning-platform\/"},{"issue":"6","key":"e_1_3_2_43_2","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1109\/MSP.2012.2205597","article-title":"Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups","volume":"29","author":"Hinton Geoffrey","year":"2012","unstructured":"Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N. Sainath, et\u00a0al. 2012. Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal Processing Magazine 29, 6 (2012), 82\u201397.","journal-title":"IEEE Signal Processing Magazine"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","unstructured":"Connor Holmes Daniel Mawhirter Yuxiong He Feng Yan and Bo Wu. 2019. GRNN: Low-latency and scalable RNN inference on GPUs. In Proceedings of the Fourteenth EuroSys Conference 2019 (EuroSys\u201919). Association for Computing Machinery Dresden Germany. DOI:10.1145\/3302424.3303949","DOI":"10.1145\/3302424.3303949"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","unstructured":"Qingda Hu Jiwu Shu Jie Fan and Youyou Lu. 2016. Run-time performance estimation and fairness-oriented scheduling policy for concurrent GPGPU applications. In 2016 45th International Conference on Parallel Processing (ICPP). 57\u201366. DOI:10.1109\/ICPP.2016.14","DOI":"10.1109\/ICPP.2016.14"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","unstructured":"Saksham Jain Iljoo Baek Shige Wang and Ragunathan Rajkumar. 2019. Fractional GPUs: Software-based compute and memory bandwidth reservation for GPUs. In 2019 IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS). 29\u201341. DOI:10.1109\/RTAS.2019.00011","DOI":"10.1109\/RTAS.2019.00011"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","unstructured":"Wonseok Jang Hansaem Jeong Kyungtae Kang Nikil Dutt and Jong-Chan Kim. 2020. R-TOD: Real-time object detector with minimized end-to-end delay for autonomous driving. In 2020 IEEE Real-Time Systems Symposium (RTSS). 191\u2013204. DOI:10.1109\/RTSS49844.2020.00027","DOI":"10.1109\/RTSS49844.2020.00027"},{"key":"e_1_3_2_48_2","doi-asserted-by":"crossref","first-page":"249","DOI":"10.1145\/3552326.3567508","volume-title":"Proceedings of the Eighteenth European Conference on Computer Systems (EuroSys\u201923)","author":"Jeong Jinwoo","year":"2023","unstructured":"Jinwoo Jeong, Seungsu Baek, and Jeongseob Ahn. 2023. Fast and efficient model serving using multi-GPUs with direct-host-access. In Proceedings of the Eighteenth European Conference on Computer Systems (EuroSys\u201923). Association for Computing Machinery, New York, NY, USA, 249\u2013265. DOI:10.1145\/3552326.3567508"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","unstructured":"Zhihao Jia Oded Padon James Thomas Todd Warszawski Matei Zaharia and Alex Aiken. 2019. TASO: optimizing deep learning computation with automatic generation of graph substitutions. In Proceedings of the 27th ACM Symposium on Operating Systems Principles (SOSP\u201919). Association for Computing Machinery Huntsville Ontario Canada 47\u201362. DOI:10.1145\/3341301.3359630","DOI":"10.1145\/3341301.3359630"},{"key":"e_1_3_2_50_2","first-page":"207","volume-title":"Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (ASPLOS 2023)","author":"Jung Jaehoon","year":"2023","unstructured":"Jaehoon Jung, Jinpyo Kim, and Jaejin Lee. 2023. DeepUM: Tensor migration and prefetching in unified memory. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (ASPLOS 2023). Association for Computing Machinery, New York, NY, USA, 207\u2013221. DOI:10.1145\/3575693.3575736"},{"key":"e_1_3_2_51_2","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1109\/RTSS.2011.13","article-title":"RGEM: A Responsive GPGPU Execution Model for Runtime Engines","author":"Kato Shinpei","year":"2011","unstructured":"Shinpei Kato, Karthik Lakshmanan, Aman Kumar, Mihir Kelkar, Yutaka Ishikawa, and Ragunathan Raj Rajkumar. 2011. RGEM: A Responsive GPGPU Execution Model for Runtime Engines. 2011 IEEE 32nd Real-Time Systems Symposium (2011), 57\u201366. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:8258940","journal-title":"2011 IEEE 32nd Real-Time Systems Symposium"},{"key":"e_1_3_2_52_2","volume-title":"USENIX Annual Technical Conference","author":"Kato Shinpei","year":"2011","unstructured":"Shinpei Kato, Karthik Lakshmanan, Ragunathan Raj Rajkumar, and Yutaka Ishikawa. 2011. TimeGraph: GPU scheduling for real-time multi-tasking environments. In USENIX Annual Technical Conference. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:18344830"},{"key":"e_1_3_2_53_2","volume-title":"USENIX Annual Technical Conference","author":"Kato Shinpei","year":"2012","unstructured":"Shinpei Kato, Michael McThrow, Carlos Maltzahn, and Scott A. Brandt. 2012. Gdev: First-class GPU resource management in the operating system. In USENIX Annual Technical Conference. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:3200101"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","unstructured":"Hyeonsu Lee Hyunjun Kim Cheolgi Kim Hwansoo Han and Euiseong Seo. 2021. Idempotence-based preemptive GPU kernel scheduling for embedded systems. IEEE Transactions on Computers 70 3 (2021) 332\u2013346. DOI:10.1109\/TC.2020.2988251","DOI":"10.1109\/TC.2020.2988251"},{"key":"e_1_3_2_55_2","unstructured":"TIMOTHY B. LEE. 2019. Tesla\u2019s autonomy event:Impressive progress with an unrealistic timeline. Retrieved from https:\/\/arstechnica.com\/cars\/2019\/04\/teslas-autonomy-event-impressive-progress-with-an-unrealistic-timeline\/"},{"key":"e_1_3_2_56_2","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1109\/HPCA47549.2020.00014","article-title":"Asymmetric resilience: Exploiting task-level idempotency for transient error recovery in accelerator-based systems","author":"Leng Jingwen","year":"2020","unstructured":"Jingwen Leng, Alper Buyuktosunoglu, Ramon Bertran Monfort, Pradip Bose, Quan Chen, Minyi Guo, and Vijay Janapa Reddi. 2020. Asymmetric resilience: Exploiting task-level idempotency for transient error recovery in accelerator-based systems. 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) (2020), 44\u201357.","journal-title":"2020 IEEE International Symposium on High Performance Computer Architecture (HPCA)"},{"key":"e_1_3_2_57_2","unstructured":"LG Electronics Inc.2022. Running Apollo 5.0 with SVL Simulator. Retrieved from https:\/\/www.svlsimulator.com\/docs\/system-under-test\/apollo5-0-instructions\/"},{"key":"e_1_3_2_58_2","first-page":"663","volume-title":"17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23)","author":"Li Zhuohan","year":"2023","unstructured":"Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E. Gonzalez, et\u00a0al. 2023. AlpaServe: Statistical multiplexing with model parallelism for deep learning serving. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23). USENIX Association, Boston, MA, 663\u2013679. Retrieved from https:\/\/www.usenix.org\/conference\/osdi23\/presentation\/li-zhouhan"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","unstructured":"Yun Liang Huynh Phung Huynh Kyle Rupnow Rick Siow Mong Goh and Deming Chen. 2015. Efficient GPU spatial-temporal multitasking. IEEE Trans. Parallel Distrib. Syst. 26 3 (March 2015) 748\u2013760. DOI:10.1109\/TPDS.2014.2313342","DOI":"10.1109\/TPDS.2014.2313342"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","unstructured":"Zhen Lin Lars Nyland and Huiyang Zhou. 2016. Enabling efficient preemption for SIMT architectures with lightweight context switching. In SC\u201916: Proceedings of the International Conference for High Performance Computing Networking Storage and Analysis. 898\u2013908. DOI:10.1109\/SC.2016.76","DOI":"10.1109\/SC.2016.76"},{"key":"e_1_3_2_61_2","unstructured":"LLVM. 2021. User Guide for AMDGPU Backend. Retrieved from https:\/\/llvm.org\/docs\/AMDGPUUsage.html"},{"key":"e_1_3_2_62_2","unstructured":"Justin Luitjens. CUDA Streams\u2014Best Practices and Common Pitfalls. Retrieved from http:\/\/on-demand.gputechconf.com\/gtc\/2014\/presentations\/S4158-cuda-streams-best-practices-common-pitfalls.pdf. (n. d.)."},{"key":"e_1_3_2_63_2","first-page":"881","volume-title":"14th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201920)","author":"Ma Lingxiao","year":"2020","unstructured":"Lingxiao Ma, Z. Xie, Zhi Yang, J. Xue, Youshan Miao, Wei Cui, W. Hu, Fan Yang, Lintao Zhang, and Lidong Zhou. 2020. Rammer: Enabling holistic deep learning compiler optimizations with rTasks. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201920). 881\u2013897."},{"key":"e_1_3_2_64_2","unstructured":"Meta AI. 2024. Llama-3.2-3B. Retrieved from https:\/\/huggingface.co\/meta-llama\/Llama-3.2-3B. (2024). Hugging Face Model."},{"key":"e_1_3_2_65_2","unstructured":"MLC team. 2023-2025. MLC-LLM. (2023-2025). Retrieved from https:\/\/github.com\/mlc-ai\/mlc-llm"},{"key":"e_1_3_2_66_2","first-page":"1","article-title":"Multi-sensor System for driver\u2019s hand-gesture recognition","volume":"1","author":"Molchanov Pavlo","year":"2015","unstructured":"Pavlo Molchanov, Shalini Gupta, Kihwan Kim, and Kari Pulli. 2015. Multi-sensor System for driver\u2019s hand-gesture recognition. 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition 1 (2015), 1\u20138.","journal-title":"11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition"},{"key":"e_1_3_2_67_2","first-page":"20","volume-title":"NeurIPS Workshop on Systems for Machine Learning","author":"Narayanan Deepak","year":"2018","unstructured":"Deepak Narayanan, Keshav Santhanam, Amar Phanishayee, and Matei Zaharia. 2018. Accelerating deep learning workloads through efficient multi-model execution. In NeurIPS Workshop on Systems for Machine Learning. 20."},{"key":"e_1_3_2_68_2","first-page":"595","volume-title":"Proceedings of the 29th Symposium on Operating Systems Principles (SOSP\u201923)","author":"Ng Kelvin K. W.","year":"2023","unstructured":"Kelvin K. W. Ng, Henri Maxime Demoulin, and Vincent Liu. 2023. Paella: Low-latency model serving with software-defined GPU scheduling. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP\u201923). Association for Computing Machinery, New York, NY, USA, 595\u2013610. DOI:10.1145\/3600006.3613163"},{"key":"e_1_3_2_69_2","unstructured":"NVIDIA. NVIDIA TensorRT. Retrieved from https:\/\/developer.nvidia.com\/tensorrt. (n. d.)."},{"key":"e_1_3_2_70_2","unstructured":"NVIDIA. NVIDIA\/open-gpu-kernel-modules: NVIDIA Linux open GPU kernel module source. https:\/\/github.com\/NVIDIA\/open-gpu-kernel-modules. (n. d.). Retrieved from https:\/\/github.com\/NVIDIA\/open-gpu-kernel-modules"},{"key":"e_1_3_2_71_2","unstructured":"NVIDIA. 2016. NVIDIA Tesla P100. Retrieved from http:\/\/www.nvidia.com\/object\/pascal-architecture-whitepaper.html"},{"key":"e_1_3_2_72_2","unstructured":"NVIDIA. 2021. CUDA Toolkit: Develop Optimize and Deploy GPU-Accelerated Apps. Retrieved from https:\/\/developer.nvidia.com\/cuda-toolkit"},{"key":"e_1_3_2_73_2","unstructured":"NVIDIA. 2024. Blackwell Instruction Set. Retrieved from https:\/\/docs.nvidia.com\/cuda\/cuda-binary-utilities\/index.html#blackwell-blackwell-instruction-set"},{"key":"e_1_3_2_74_2","volume-title":"Workshop on ML Systems at NIPS 2017","author":"Olston Christopher","year":"2017","unstructured":"Christopher Olston, Fangwei Li, Jeremiah Harmsen, Jordan Soyke, Kiril Gorovoy, Li Lao, Noah Fiedel, Sukriti Ramesh, and Vinu Rajashekhar. 2017. TensorFlow-serving: Flexible, high-performance ML serving. In Workshop on ML Systems at NIPS 2017."},{"key":"e_1_3_2_75_2","volume-title":"Eighteenth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201913)","author":"Pai S.","year":"2013","unstructured":"S. Pai, M. J. Thazhuthaveetil, and R. Govindarajan. 2013. Improving GPGPU concurrency with elastic kernels. In Eighteenth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201913)."},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","unstructured":"Jason Jong Kyu Park Yongjun Park and Scott Mahlke. 2015. Chimera: Collaborative preemption for multitasking on a shared GPU. In Proceedings of the Twentieth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201915). Association for Computing Machinery Istanbul Turkey 593\u2013606. DOI:10.1145\/2694344.2694346","DOI":"10.1145\/2694344.2694346"},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","unstructured":"Jason Jong Kyu Park Yongjun Park and Scott Mahlke. 2017. Dynamic resource management for efficient utilization of multitasking GPUs. In Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201917). Association for Computing Machinery Xi\u2019an China 527\u2013540. DOI:10.1145\/3037697.3037707","DOI":"10.1145\/3037697.3037707"},{"key":"e_1_3_2_78_2","volume-title":"33rd International Conference on Neural Information Processing Systems (NeurIPS\u201919)","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, S. Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, N. Gimelshein, L. Antiga, et\u00a0al. 2019. PyTorch: An imperative style, high-performance deep learning library. In 33rd International Conference on Neural Information Processing Systems (NeurIPS\u201919)."},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","unstructured":"Reid Pinkham Andrew Berkovich and Zhengya Zhang. 2021. Near-sensor distributed DNN processing for augmented and virtual reality. IEEE Journal on Emerging and Selected Topics in Circuits and Systems 11 4 (2021) 663\u2013676. DOI:10.1109\/JETCAS.2021.3121259","DOI":"10.1109\/JETCAS.2021.3121259"},{"key":"e_1_3_2_80_2","unstructured":"Qwen Team. 2024. Qwen2.5-3B-Instruct. Retrieved from https:\/\/huggingface.co\/Qwen\/Qwen2.5-3B-Instruct. (2024). Hugging Face Model."},{"key":"e_1_3_2_81_2","unstructured":"Alec Radford Jeff Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. OpenAI."},{"key":"e_1_3_2_82_2","doi-asserted-by":"publisher","unstructured":"Joseph Redmon Santosh Divvala Ross Girshick and Ali Farhadi. 2016. You only look once: Unified real-time object detection. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 779\u2013788. DOI:10.1109\/CVPR.2016.91","DOI":"10.1109\/CVPR.2016.91"},{"key":"e_1_3_2_83_2","unstructured":"Joseph Redmon and Ali Farhadi. 2018. YOLOv3: An Incremental Improvement. (2018). Retrieved from http:\/\/arxiv.org\/abs\/1804.02767cite arxiv:1804.02767Comment: Tech Report."},{"key":"e_1_3_2_84_2","unstructured":"Steve Rennich. CUDA C\/C++ Streams and Concurrency. Retrieved from https:\/\/developer.download.nvidia.cn\/CUDA\/training\/StreamsAndConcurrencyWebinar.pdf. (n. d.)."},{"key":"e_1_3_2_85_2","unstructured":"ROCm Core Technology. 2022. AMD GPU kernel driver with KFD. Retrieved from https:\/\/github.com\/RadeonOpenCompute\/ROCK-Kernel-Driver"},{"key":"e_1_3_2_86_2","unstructured":"ROCm Core Technology. 2022. AMD GPU kernel driver with KFD: unmap_queues_cpsch. Retrieved from https:\/\/github.com\/RadeonOpenCompute\/ROCK-Kernel-Driver\/blob\/master\/drivers\/gpu\/drm\/amd\/amdkfd\/kfd_device_queue_manager.c"},{"key":"e_1_3_2_87_2","unstructured":"ROCm Developer Tools. 2022. HIP: C++ Heterogeneous-Compute Interface for Portability. Retrieved from https:\/\/github.com\/ROCm-Developer-Tools\/HIP"},{"key":"e_1_3_2_88_2","first-page":"397","volume-title":"USENIX Annual Technical Conference (ATC\u201921)","author":"Romero Francisco","year":"2021","unstructured":"Francisco Romero, Qian Li, Neeraja J. Yadwadkar, and Christos Kozyrakis. 2021. INFaaS: Automated model-less inference serving. In USENIX Annual Technical Conference (ATC\u201921). 397\u2013411."},{"key":"e_1_3_2_89_2","first-page":"1","volume-title":"IEEE 23rd International Conference on Intelligent Transportation Systems Conference (ITSC\u201920)","author":"Rong Guodong","year":"2020","unstructured":"Guodong Rong, Byung Hyun Shin, Hadi Tabatabaee, Qiang Lu, Steve Lemke, M\u0101rti\u0146\u0161 Mo\u017eeiko, Eric Boise, Geehoon Uhm, Mark Gerow, Shalin Mehta, et\u00a0al. 2020. LGSVL Simulator: A High Fidelity Simulator for Autonomous Driving. In IEEE 23rd International Conference on Intelligent Transportation Systems Conference (ITSC\u201920). 1\u20136. DOI:10.1109\/ITSC45102.2020.9294422"},{"key":"e_1_3_2_90_2","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1145\/2043556.2043579","volume-title":"Proceedings of the Twenty-Third ACM Symposium on Operating Systems Principles (SOSP\u201911)","author":"Rossbach Christopher J.","year":"2011","unstructured":"Christopher J. Rossbach, Jon Currey, Mark Silberstein, Baishakhi Ray, and Emmett Witchel. 2011. PTask: Operating system abstractions to manage GPUs as compute devices. In Proceedings of the Twenty-Third ACM Symposium on Operating Systems Principles (SOSP\u201911). Association for Computing Machinery, New York, NY, USA, 233\u2013248. DOI:10.1145\/2043556.2043579"},{"key":"e_1_3_2_91_2","doi-asserted-by":"crossref","DOI":"10.1145\/3341301.3359658","article-title":"Nexus: A GPU cluster engine for accelerating DNN-based video analysis","author":"Shen Haichen","year":"2019","unstructured":"Haichen Shen, Lequn Chen, Yuchen Jin, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, and Ravi Sundaram. 2019. Nexus: A GPU cluster engine for accelerating DNN-based video analysis. 27th ACM Symposium on Operating Systems Principles (2019).","journal-title":"27th ACM Symposium on Operating Systems Principles"},{"key":"e_1_3_2_92_2","first-page":"671","volume-title":"19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25)","author":"Shen Weihang","year":"2025","unstructured":"Weihang Shen, Mingcong Han, Jialong Liu, Rong Chen, and Haibo Chen. 2025. XSched: Preemptive scheduling for diverse XPUs. In 19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25). USENIX Association, Boston, MA, 671\u2013692. Retrieved from https:\/\/www.usenix.org\/conference\/osdi25\/presentation\/shen-weihang"},{"key":"e_1_3_2_93_2","first-page":"701","volume-title":"17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23)","author":"Shi Yining","year":"2023","unstructured":"Yining Shi, Zhi Yang, Jilong Xue, Lingxiao Ma, Yuqing Xia, Ziming Miao, Yuxiao Guo, Fan Yang, and Lidong Zhou. 2023. Welder: Scheduling deep learning memory access via tile-graph. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23). 701\u2013718."},{"key":"e_1_3_2_94_2","first-page":"947","volume-title":"18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)","author":"Shubha Sudipta Saha","year":"2024","unstructured":"Sudipta Saha Shubha, Haiying Shen, and Anand Iyer. 2024. USHER: Holistic interference avoidance for resource optimized ML inference. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). USENIX Association, Santa Clara, CA, 947\u2013964. Retrieved from https:\/\/www.usenix.org\/conference\/osdi24\/presentation\/shubha"},{"key":"e_1_3_2_95_2","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556. Retrieved from https:\/\/arxiv.org\/abs\/\/1409.1556"},{"key":"e_1_3_2_96_2","doi-asserted-by":"crossref","first-page":"1075","DOI":"10.1145\/3627703.3629578","volume-title":"Proceedings of the Nineteenth European Conference on Computer Systems (EuroSys\u201924)","author":"Strati Foteini","year":"2024","unstructured":"Foteini Strati, Xianzhe Ma, and Ana Klimovic. 2024. Orion: Interference-aware, fine-grained GPU sharing for ML applications. In Proceedings of the Nineteenth European Conference on Computer Systems (EuroSys\u201924). Association for Computing Machinery, New York, NY, USA, 1075\u20131092. DOI:10.1145\/3627703.3629578"},{"key":"e_1_3_2_97_2","doi-asserted-by":"crossref","unstructured":"I. Tanasi\u0107 Isaac Gelado Javier Cabezas A. Ram\u00edrez N. Navarro and M. Valero. 2014. Enabling preemptive multiprogramming on GPUs. In Proceeding of the 41st Annual International Symposium on Computer Architecuture (ISCA\u201914). IEEE Press Minneapolis Minnesota USA 193\u2013204.","DOI":"10.1109\/ISCA.2014.6853208"},{"key":"e_1_3_2_98_2","unstructured":"TESLARATI. 2021. AMD confirms Tesla\u2019s new Model S and Model X will boast RDNA 2 GPUs. Retrieved from https:\/\/www.teslarati.com\/tesla-model-s-model-x-mcu3-specs-amd-gpu-confirmed-video\/"},{"key":"e_1_3_2_99_2","unstructured":"Apache TVM. 2021. Apache TVM: An End to End Machine Learning Compiler Framework for CPUs GPUs and accelerators. Retrieved from https:\/\/tvm.apache.org\/"},{"key":"e_1_3_2_100_2","doi-asserted-by":"publisher","unstructured":"Oreste Villa Mark Stephenson David Nellans and Stephen W. Keckler. 2019. NVBit: A dynamic binary instrumentation framework for NVIDIA GPUs. In Proceedings of the 52nd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO-52). Association for Computing Machinery Columbus OH USA 372\u2013383. DOI:10.1145\/3352460.3358307","DOI":"10.1145\/3352460.3358307"},{"key":"e_1_3_2_101_2","doi-asserted-by":"publisher","unstructured":"Guibin Wang YiSong Lin and Wei Yi. 2010. Kernel Fusion: An effective method for better power efficiency on multithreaded GPU. In Proceedings of the 2010 IEEE\/ACM Int\u2019l Conference on Green Computing and Communications & Int\u2019l Conference on Cyber Physical and Social Computing (GREENCOM-CPSCOM\u201910). IEEE Computer Society USA 344\u2013350. DOI:10.1109\/GreenCom-CPSCom.2010.102","DOI":"10.1109\/GreenCom-CPSCom.2010.102"},{"key":"e_1_3_2_102_2","volume-title":"USENIX Annual Technical Conference (ATC\u201925)","author":"Wang Jiali","year":"2025","unstructured":"Jiali Wang, Yankui Wang, Mingcong Han, and Rong Chen. 2025. Colocating ML inference and training with fast GPU memory handover. In USENIX Annual Technical Conference (ATC\u201925)."},{"key":"e_1_3_2_103_2","unstructured":"Yuxuan Wang R. J. Skerry-Ryan Daisy Stanton Yonghui Wu Ron J. Weiss Navdeep Jaitly Zongheng Yang Ying Xiao Zhifeng Chen Samy Bengio et\u00a0al. 2017. Tacotron: A fully end-to-end text-to-speech synthesis model. arXiv:1703.10135. Retrieved from https:\/\/arxiv.org\/abs\/\/1703.10135"},{"key":"e_1_3_2_104_2","doi-asserted-by":"publisher","unstructured":"Zhenning Wang Jun Yang Rami Melhem Bruce Childers Youtao Zhang and Minyi Guo. 2016. Simultaneous Multikernel GPU: Multi-tasking throughput processors via fine-grained sharing. In 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA). 358\u2013369. DOI:10.1109\/HPCA.2016.7446078","DOI":"10.1109\/HPCA.2016.7446078"},{"key":"e_1_3_2_105_2","doi-asserted-by":"publisher","unstructured":"Bo Wu Xu Liu Xiaobo Zhou and Changjun Jiang. 2017. FLEP: Enabling flexible and efficient preemption on GPUs. In Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201917). Association for Computing Machinery Xi\u2019an China 483\u2013496. DOI:10.1145\/3037697.3037742","DOI":"10.1145\/3037697.3037742"},{"key":"e_1_3_2_106_2","doi-asserted-by":"publisher","unstructured":"Yecheng Xiang and Hyoseung Kim. 2019. Pipelined data-parallel CPU\/GPU scheduling for multi-DNN real-time inference. In 2019 IEEE Real-Time Systems Symposium (RTSS). 392\u2013405. DOI:10.1109\/RTSS46320.2019.00042","DOI":"10.1109\/RTSS46320.2019.00042"},{"key":"e_1_3_2_107_2","doi-asserted-by":"publisher","unstructured":"Feng Yan Yuxiong He Olatunji Ruwase and Evgenia Smirni. 2018. Efficient deep neural network serving: Fast and furious. IEEE Transactions on Network and Service Management 15 1 (2018) 112\u2013126. DOI:10.1109\/TNSM.2018.2808352","DOI":"10.1109\/TNSM.2018.2808352"},{"key":"e_1_3_2_108_2","volume-title":"Proceedings of the 16th ACM SIGOPS Asia-Pacific Workshop on Systems (APSys\u201925)","author":"Yang Jinrong","year":"2025","unstructured":"Jinrong Yang, Zimeng Wang, Rong Chen, and Haibo Chen. 2025. A system-level abstraction and service for flourishing AI-powered applications. In Proceedings of the 16th ACM SIGOPS Asia-Pacific Workshop on Systems (APSys\u201925). Association for Computing Machinery, New York, NY, USA, 9. DOI:10.1145\/3725783.3764406"},{"key":"e_1_3_2_109_2","doi-asserted-by":"publisher","unstructured":"Ming Yang Shige Wang Joshua Bakita Thanh Vu F. Donelson Smith James H. Anderson and Jan-Michael Frahm. 2019. Re-thinking CNN frameworks for time-sensitive autonomous-driving applications: Addressing an industrial challenge. In 2019 IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS). 305\u2013317. DOI:10.1109\/RTAS.2019.00033","DOI":"10.1109\/RTAS.2019.00033"},{"key":"e_1_3_2_110_2","doi-asserted-by":"crossref","first-page":"768","DOI":"10.1145\/3503222.3507709","volume-title":"Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems","author":"Yang Yanan","year":"2022","unstructured":"Yanan Yang, Laiping Zhao, Yiming Li, Huanyu Zhang, Jie Li, Mingyang Zhao, Xingzhen Chen, and Keqiu Li. 2022. INFless: A native serverless system for low-latency, high-throughput inference. In Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. 768\u2013781."},{"key":"e_1_3_2_111_2","unstructured":"Wai Chee Yau. 2017. How Zendesk Serves TensorFlow Models in Production. Retrieved from https:\/\/zendesk.engineering\/how-zendesk-serves-tensorflow-models-in-production-751ee22f0f4b"},{"key":"e_1_3_2_112_2","doi-asserted-by":"publisher","unstructured":"Tsung Tai Yeh Matthew D. Sinclair Bradford M. Beckmann and Timothy G. Rogers. 2021. Deadline-aware offloading for high-throughput accelerators. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 479\u2013492. DOI:10.1109\/HPCA51647.2021.00048","DOI":"10.1109\/HPCA51647.2021.00048"},{"key":"e_1_3_2_113_2","first-page":"1","volume-title":"26th Annual International Conference on Mobile Computing and Networking (MobiCom\u201920)","author":"Yi Juheon","year":"2020","unstructured":"Juheon Yi and Youngki Lee. 2020. Heimdall: mobile GPU coordination platform for augmented reality applications. In 26th Annual International Conference on Mobile Computing and Networking (MobiCom\u201920). 1\u201314."},{"key":"e_1_3_2_114_2","doi-asserted-by":"publisher","unstructured":"Sebastian Zepf Javier Hernandez Alexander Schmitt Wolfgang Minker and Rosalind W. Picard. 2020. Driver emotion recognition for intelligent vehicles: A survey. ACM Comput. Surv. 53 3 (July 2020). DOI:10.1145\/3388790","DOI":"10.1145\/3388790"},{"key":"e_1_3_2_115_2","volume-title":"Symposium on Networked Systems Design and Implementation","author":"Zhang Hong","year":"2023","unstructured":"Hong Zhang, Yupeng Tang, Anurag Khandelwal, and Ion Stoica. 2023. SHEPHERD: Serving DNNs in the Wild. In Symposium on Networked Systems Design and Implementation. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:258559393"},{"key":"e_1_3_2_116_2","doi-asserted-by":"crossref","first-page":"1451","DOI":"10.1109\/TPDS.2021.3115630","article-title":"A survey of GPU multitasking methods supported by hardware architecture","volume":"33","author":"Zhao Chen","year":"2022","unstructured":"Chen Zhao, Wu Gao, Feiping Nie, and Huiyang Zhou. 2022. A survey of GPU multitasking methods supported by hardware architecture. IEEE Transactions on Parallel and Distributed Systems 33 (2022), 1451\u20131463. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:239197299","journal-title":"IEEE Transactions on Parallel and Distributed Systems"},{"key":"e_1_3_2_117_2","unstructured":"Hengyu Zhao Yubo Zhang Pingfan Meng Hui Shi Erran L. Li Tiancheng Lou and Jishen Zhao. 2019. Towards safety-aware computing system design in autonomous vehicles. arXiv:1905.08453. Retrieved from https:\/\/arxiv.org\/abs\/\/1905.08453"},{"key":"e_1_3_2_118_2","doi-asserted-by":"publisher","unstructured":"Wenyi Zhao Quan Chen and Minyi Guo. 2018. KSM: Online application-level performance slowdown prediction for spatial multitasking GPGPU. IEEE Comput. Archit. Lett. 17 2 (July 2018) 187\u2013191. DOI:10.1109\/LCA.2018.2851207","DOI":"10.1109\/LCA.2018.2851207"},{"key":"e_1_3_2_119_2","doi-asserted-by":"publisher","unstructured":"Wenyi Zhao Quan Chen Hao Lin Jianfeng Zhang Jingwen Leng Chao Li Wenli Zheng Li Li and Minyi Guo. 2019. Themis: Predicting and reining in application-level slowdown on spatial multitasking GPUs. In 2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 653\u2013663. DOI:10.1109\/IPDPS.2019.00074","DOI":"10.1109\/IPDPS.2019.00074"},{"key":"e_1_3_2_120_2","doi-asserted-by":"publisher","unstructured":"Xia Zhao Magnus Jahre and Lieven Eeckhout. 2020. HSM: A hybrid slowdown model for multitasking GPUs. In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201920). Association for Computing Machinery Lausanne Switzerland 1371\u20131385. DOI:10.1145\/3373376.3378457","DOI":"10.1145\/3373376.3378457"},{"key":"e_1_3_2_121_2","volume-title":"14th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201920)","author":"Zheng Lianmin","year":"2020","unstructured":"Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, et\u00a0al. 2020. Ansor: Generating high-performance tensor programs for deep learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI\u201920)."},{"key":"e_1_3_2_122_2","doi-asserted-by":"crossref","first-page":"190","DOI":"10.1109\/RTAS.2018.00028","article-title":"S3DNN: Supervised streaming and scheduling for GPU-accelerated real-time DNN workloads","author":"Zhou Husheng","year":"2018","unstructured":"Husheng Zhou, Soroush Bateni, and Cong Liu. 2018. S3DNN: Supervised streaming and scheduling for GPU-accelerated real-time DNN workloads. 2018 IEEE Real-Time and Embedded Technology and Applications Symposium (2018), 190\u2013201.","journal-title":"2018 IEEE Real-Time and Embedded Technology and Applications Symposium"},{"key":"e_1_3_2_123_2","doi-asserted-by":"publisher","unstructured":"Husheng Zhou Guangmo Tong and Cong Liu. 2015. GPES: A preemptive execution system for GPGPU computing. In 21st IEEE Real-Time and Embedded Technology and Applications Symposium. 87\u201397. DOI:10.1109\/RTAS.2015.7108420","DOI":"10.1109\/RTAS.2015.7108420"},{"key":"e_1_3_2_124_2","first-page":"233","volume-title":"16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)","author":"Zhu Hongyu","year":"2022","unstructured":"Hongyu Zhu, Ruofan Wu, Yijia Diao, Shanbin Ke, Haoyu Li, Chen Zhang, Jilong Xue, Lingxiao Ma, Yuqing Xia, Wei Cui, et\u00a0al. 2022. [ROLLER]: Fast and efficient tensor compilation for deep learning. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). 233\u2013248."}],"container-title":["ACM Transactions on Computer Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3768622","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,7]],"date-time":"2025-11-07T14:01:07Z","timestamp":1762524067000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3768622"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,7]]},"references-count":123,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,2,28]]}},"alternative-id":["10.1145\/3768622"],"URL":"https:\/\/doi.org\/10.1145\/3768622","relation":{},"ISSN":["0734-2071","1557-7333"],"issn-type":[{"value":"0734-2071","type":"print"},{"value":"1557-7333","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,7]]},"assertion":[{"value":"2024-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-05","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}