{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T16:46:40Z","timestamp":1782406000111,"version":"3.54.5"},"reference-count":43,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>GPU sharing is commonly employed in GPU clusters to improve utilization, with spatial sharing being one of the most widely adopted techniques. However, spatial sharing can lead to resource interference, making task execution times difficult to predict. Predictable execution times for each task are crucial in GPU cluster management and task scheduling. In this article, we propose a performance predictor for multi-DNN training tasks in GPU spatial sharing environments. We first conduct experiments on spatial sharing for multiple DNN workloads on a single GPU, demonstrating that concurrent execution of multiple tasks improves overall performance and GPU resource utilization compared to serial execution. By analyzing warp stall reasons collected during task execution, we investigate the interference for computation and memory resources under MPS on GPUs. Finally, we design a performance predictor that predicts the execution time of a target DNN training task when it runs concurrently with other tasks under GPU spatial sharing via MPS. The predictor is capable of predicting the execution time of each task for previously unseen combinations of DNN training tasks. Extensive evaluations on modern GPUs show that compared to other baseline methods, our approach exhibits higher prediction accuracy, as well as improved stability and robustness. Experiments on multiple GPU architectures, as well as at higher concurrency levels, further demonstrate that our method possesses strong generalization and scalability. We also conducted a performance analysis under diverse workload pattern and a case study to validate the practical applicability of our predictor in real scheduling environments.<\/jats:p>","DOI":"10.1145\/3807454","type":"journal-article","created":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T09:55:28Z","timestamp":1777110928000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Performance Prediction of Concurrent DNN Training Tasks in GPU Spatial Sharing Environments"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-5696-4944","authenticated-orcid":false,"given":"Sichao","family":"Chen","sequence":"first","affiliation":[{"name":"Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7502-7094","authenticated-orcid":false,"given":"Desheng","family":"Wang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen","place":["Shenzhen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4783-876X","authenticated-orcid":false,"given":"Weizhe","family":"Zhang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0043-4370","authenticated-orcid":false,"given":"Meng","family":"Hao","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8709-5625","authenticated-orcid":false,"given":"Yu-Chu","family":"Tian","sequence":"additional","affiliation":[{"name":"Queensland University of Technology","place":["Brisbane, Australia"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6494-775X","authenticated-orcid":false,"given":"Hui","family":"He","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2019. Deploy machine learning models in production environments. Retrieved April 30 2026 from https:\/\/learn.microsoft.com\/en-us\/azure\/cloud-adoption-framework\/innovate\/best-practices\/ml-deployment-inference"},{"key":"e_1_3_1_3_2","first-page":"265","volume-title":"Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et\u00a0al. 2016. TensorFlow: A system for Large-Scale machine learning. In Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation. 265\u2013283."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00039"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3431731"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3689031.3696074"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2018.00027"},{"key":"e_1_3_1_8_2","first-page":"503","volume-title":"Proceedings of the 2021 USENIX Annual Technical Conference","author":"Geoffrey X. Yu","year":"2021","unstructured":"X. Yu Geoffrey, Yubo Gao, Pavel Golikov, and Gennady Pekhimenko. 2021. Habitat: A runtime-based computational performance predictor for deep neural network training. In Proceedings of the 2021 USENIX Annual Technical Conference. 503\u2013521."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/BDCloud.2018.00077"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2951218"},{"key":"e_1_3_1_11_2","first-page":"539","volume-title":"Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation","author":"Han Mingcong","year":"2022","unstructured":"Mingcong Han, Hanze Zhang, Rong Chen, and Haibo Chen. 2022. Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences. In Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation. 539\u2013558."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2024.3430063"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2018.00059"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.5555\/3358807.3358888"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3542929.3563510"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2025.3577796"},{"key":"e_1_3_1_19_2","first-page":"82","article-title":"Plink: Discovering and exploiting locality for accelerated distributed training on the public cloud","volume":"2","author":"Luo Liang","year":"2020","unstructured":"Liang Luo, Peter West, Jacob Nelson, Arvind Krishnamurthy, and Luis Ceze. 2020. Plink: Discovering and exploiting locality for accelerated distributed training on the public cloud. Proceedings of Machine Learning and Systems 2 (2020), 82\u201397.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3522712"},{"key":"e_1_3_1_21_2","article-title":"MXNet","year":"2017","unstructured":"MXNet. 2017. MXNet. Retrieved April 30, 2026 from https:\/\/mxnet.apache.org\/","journal-title":"Retrieved April 30, 2026 from"},{"key":"e_1_3_1_22_2","unstructured":"NVIDIA. 2020. CUDA Multi-Process Service. Retrieved April 30 2026 from https:\/\/docs.nvidia.com\/deploy\/pdf\/CUDA_Multi_Process_Service_Overview.pdf"},{"key":"e_1_3_1_23_2","unstructured":"NVIDIA. 2020. Nvidia multi-instance GPU (MIG). Retrieved April 30 2026 from https:\/\/www.nvidia.com\/en-us\/technologies\/multi-instance-gpu\/"},{"key":"e_1_3_1_24_2","article-title":"NVIDIA System Management Interface (nvidia-smi)","year":"2020","unstructured":"NVIDIA. 2020. NVIDIA System Management Interface (nvidia-smi). Retrieved April 30, 2026 from https:\/\/docs.nvidia.com\/deploy\/nvidia-smi\/index.html","journal-title":"Retrieved April 30, 2026 from"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS.2017.3"},{"key":"e_1_3_1_26_2","year":"2019","unstructured":"PyTorch. 2019. Retrieved April 30, 2026 from https:\/\/pytorch.org\/","journal-title":"Retrieved April 30, 2026 from"},{"key":"e_1_3_1_27_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Qi Hang","year":"2017","unstructured":"Hang Qi, Evan R. Sparks, and Ameet Talwalkar. 2017. Paleo: A performance model for deep neural networks. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359658"},{"key":"e_1_3_1_29_2","unstructured":"Ryan Smith. 2017. Nvidia volta unveiled: Gv100 GPU and tesla v100 accelerator announced."},{"key":"e_1_3_1_30_2","year":"2016","unstructured":"TensorFlow. 2016. Retrieved April 30, 2026 from https:\/\/www.tensorflow.org\/","journal-title":"Retrieved April 30, 2026 from"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3731599.3767396"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid49817.2020.00-15"},{"key":"e_1_3_1_33_2","first-page":"945","volume-title":"Proceedings of the 19th USENIX Symposium on Networked Systems Design and Implementation","author":"Weng Qizhen","year":"2022","unstructured":"Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, and Yu Ding. 2022. MLaaS in the wild: Workload analysis and scheduling in Large-Scale heterogeneous GPU clusters. In Proceedings of the 19th USENIX Symposium on Networked Systems Design and Implementation. 945\u2013960."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2021.3097287"},{"key":"e_1_3_1_35_2","first-page":"69","volume-title":"Proceedings of the 20th USENIX Symposium on Networked Systems Design and Implementation","author":"Wu Bingyang","year":"2023","unstructured":"Bingyang Wu, Zili Zhang, Zhihao Bai, Xuanzhe Liu, and Xin Jin. 2023. Transparent GPU sharing in container clouds for deep learning workloads. In Proceedings of the 20th USENIX Symposium on Networked Systems Design and Implementation. 69\u201385."},{"key":"e_1_3_1_36_2","first-page":"795","article-title":"Sustainable ai: Environmental implications, challenges and opportunities","volume":"4","author":"Wu Carole-Jean","year":"2022","unstructured":"Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et\u00a0al. 2022. Sustainable ai: Environmental implications, challenges and opportunities. Proceedings of Machine Learning and Systems 4 (2022), 795\u2013813.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_1_37_2","first-page":"595","volume-title":"Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation","author":"Xiao Wencong","year":"2018","unstructured":"Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, et\u00a0al. 2018. Gandiva: Introspective cluster scheduling for deep learning. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation. 595\u2013610."},{"key":"e_1_3_1_38_2","first-page":"533","volume-title":"Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation","author":"Xiao Wencong","year":"2020","unstructured":"Wencong Xiao, Shiru Ren, Yong Li, Yang Zhang, Pengyang Hou, Zhi Li, Yihui Feng, Wei Lin, and Yangqing Jia. 2020. AntMan: Dynamic scaling on GPU clusters for deep learning. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation. 533\u2013548."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.5555\/3357034.3357051"},{"key":"e_1_3_1_40_2","first-page":"98","article-title":"Fine-grained GPU sharing primitives for deep learning applications","volume":"2","author":"Yu Peifeng","year":"2020","unstructured":"Peifeng Yu and Mosharaf Chowdhury. 2020. Fine-grained GPU sharing primitives for deep learning applications. Proceedings of Machine Learning and Systems 2 (2020), 98\u2013111.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3673038.3673089"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3712285.3759857"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3689031.3696070"},{"key":"e_1_3_1_44_2","unstructured":"Maohua Zhu Liu Liu Chao Wang and Yuan Xie. 2016. Cnnlab: A novel parallel framework for neural networks using gpu and fpga-a practical study with trade-off analysis. arXiv:1606.06234. Retrieved from https:\/\/arxiv.org\/abs\/1606.06234"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3807454","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:55:49Z","timestamp":1782402949000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3807454"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":43,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3807454"],"URL":"https:\/\/doi.org\/10.1145\/3807454","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2025-09-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}