{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T15:42:04Z","timestamp":1783784524664,"version":"3.55.0"},"reference-count":35,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2024,11,20]],"date-time":"2024-11-20T00:00:00Z","timestamp":1732060800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-sa\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation Award","doi-asserted-by":"crossref","award":["CNS-2008072"],"award-info":[{"award-number":["CNS-2008072"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2024,12,31]]},"abstract":"<jats:p>Deep Neural Networks (DNNs) have been applied as an effective machine learning algorithm to tackle problems in different domains. However, the endeavor to train sophisticated DNN models can stretch from days into weeks, presenting substantial obstacles in the realm of research focused on large-scale DNN architectures. Distributed Deep Learning (DDL) contributes to accelerating DNN training by distributing training workloads across multiple computation accelerators, for example, graphics processing units (GPUs). Despite the considerable amount of research directed toward enhancing DDL training, the influence of data loading on GPU utilization and overall training efficacy remains relatively overlooked. It is non-trivial to optimize data-loading in DDL applications that need intensive central processing unit (CPU) and input\/output (I\/O) resources to process enormous training data. When multiple DDL applications are deployed on a system (e.g., Cloud and High-Performance Computing (HPC) system), the lack of a practical and efficient technique for data-loader allocation incurs GPU idleness and degrades the training throughput. Therefore, our work first focuses on investigating the impact of data-loading on the global training throughput. We then propose a throughput prediction model to predict the maximum throughput for an individual DDL training application. By leveraging the predicted results, A-Dloader is designed to dynamically allocate CPU and I\/O resources to concurrently running DDL applications and use the data-loader allocation as a knob to reduce GPU idle intervals and thus improve the overall training throughput. We implement and evaluate A-Dloader in a DDL framework for a series of DDL applications arriving and completing across the runtime. Our experimental results show that A-Dloader can achieve a 28.9% throughput improvement and a 10% makespan improvement compared with allocating resources evenly across applications.<\/jats:p>","DOI":"10.1145\/3680546","type":"journal-article","created":{"date-parts":[[2024,8,22]],"date-time":"2024-08-22T11:46:47Z","timestamp":1724327207000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["A Data-Loader Tunable Knob to Shorten GPU Idleness for Distributed Deep Learning"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2858-0505","authenticated-orcid":false,"given":"Danlin","family":"Jia","sequence":"first","affiliation":[{"name":"Samsung Semiconductor Inc USA, San Jose, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9844-992X","authenticated-orcid":false,"given":"Geng","family":"Yuan","sequence":"additional","affiliation":[{"name":"University of Georgia School of Computing, Athens, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-8841-9699","authenticated-orcid":false,"given":"Yiming","family":"Xie","sequence":"additional","affiliation":[{"name":"Northeastern University College of Engineering, Boston, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6210-8883","authenticated-orcid":false,"given":"Xue","family":"Lin","sequence":"additional","affiliation":[{"name":"Northeastern University College of Engineering, Boston, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0934-1287","authenticated-orcid":false,"given":"Ningfang","family":"Mi","sequence":"additional","affiliation":[{"name":"Northeastern University College of Engineering, Boston, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,11,20]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Discovery Cluster. (n.d.) Retrieved from https:\/\/rc-docs.northeastern.edu\/en\/latest\/using-discovery\/workingwithgpu.html"},{"key":"e_1_3_2_3_2","unstructured":"NCCL. (n.d.) NVIDIA Collective Communications Library. Retrieved from https:\/\/developer.nvidia.com\/nccl"},{"key":"e_1_3_2_4_2","unstructured":"THOP. (n.d.) PyTorch-OpCounter. Retrieved from https:\/\/pypi.org\/project\/thop\/"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3320060"},{"key":"e_1_3_2_6_2","article-title":"Faster neural network training with data echoing","author":"Choi Dami","year":"2019","unstructured":"Dami Choi, Alexandre Passos, Christopher J. Shallue, and George E. Dahl. 2019. Faster neural network training with data echoing. arXiv preprint arXiv:1907.05550 (2019).","journal-title":"arXiv preprint arXiv:1907.05550"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476181"},{"key":"e_1_3_2_8_2","volume-title":"International Conference on Machine Learning","author":"Evci Utku","year":"2020","unstructured":"Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen. 2020. Rigging the lottery: Making all tickets winners. In International Conference on Machine Learning. PMLR."},{"key":"e_1_3_2_9_2","volume-title":"2021 USENIX Annual Technical Conference (USENIX ATC 21)","author":"Geoffrey X. Yu","year":"2021","unstructured":"X. Yu Geoffrey, Yubo Gao, Pavel Golikov, and Gennady Pekhimenko. 2021. Habitat: A \\(\\lbrace\\) Runtime-Based \\(\\rbrace\\) computational performance predictor for deep neural network training. In 2021 USENIX Annual Technical Conference (USENIX ATC 21)."},{"key":"e_1_3_2_10_2","first-page":"485","volume-title":"16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19)","author":"Gu Juncheng","year":"2019","unstructured":"Juncheng Gu, Mosharaf Chowdhury, Kang G. Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo. 2019. Tiresias: A \\(\\lbrace\\) GPU \\(\\rbrace\\) cluster manager for distributed deep learning. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). 485\u2013500."},{"key":"e_1_3_2_11_2","article-title":"(n.d.) TicTac: Accelerating distributed deep learning with communication scheduling","author":"Hashemi Sayed Hadi","unstructured":"Sayed Hadi Hashemi, Sangeetha Abdu Jyothi, and Roy H. Campbell. (n.d.) TicTac: Accelerating distributed deep learning with communication scheduling. SysML 2019.","journal-title":"SysML 2019."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_13_2","article-title":"MobileNets: Efficient convolutional neural networks for mobile vision applications","author":"Howard Andrew G.","year":"2017","unstructured":"Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017).","journal-title":"arXiv preprint arXiv:1704.04861"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00070"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CLOUD55607.2022.00068"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1989.1.4.541"},{"key":"e_1_3_2_17_2","article-title":"Pytorch distributed: Experiences on accelerating data parallel training","author":"Li Shen","year":"2020","unstructured":"Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, et\u00a0al. 2020. Pytorch distributed: Experiences on accelerating data parallel training. arXiv preprint arXiv:2006.15704 (2020).","journal-title":"arXiv preprint arXiv:2006.15704"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3107554"},{"key":"e_1_3_2_19_2","first-page":"1","article-title":"Training CNNs in presence of JPEG compression: Multimedia forensics vs. computer vision","author":"Mandelli Sara","year":"2020","unstructured":"Sara Mandelli, Nicol\u00f2 Bonettini, Paolo Bestagini, and Stefano Tubaro. 2020. Training CNNs in presence of JPEG compression: Multimedia forensics vs. computer vision. 2020 IEEE International Workshop on Information Forensics and Security (WIFS) (2020), 1\u20136.","journal-title":"2020 IEEE International Workshop on Information Forensics and Security (WIFS)"},{"key":"e_1_3_2_20_2","article-title":"Analyzing and mitigating data stalls in DNN training","author":"Mohan Jayashree","year":"2020","unstructured":"Jayashree Mohan, Amar Phanishayee, Ashish Raniwala, and Vijay Chidambaram. 2020. Analyzing and mitigating data stalls in DNN training. arXiv preprint arXiv:2007.06775 (2020).","journal-title":"arXiv preprint arXiv:2007.06775"},{"key":"e_1_3_2_21_2","first-page":"1","article-title":"Auto-Driving policies in highway based on distributional deep reinforcement learning","author":"Molaie Mahdi","year":"2021","unstructured":"Mahdi Molaie, Abdollah Amirkhani, and Hossein Kashiani. 2021. Auto-Driving policies in highway based on distributional deep reinforcement learning. 2021 5th International Conference on Pattern Recognition and Image Analysis (IPRIA) (2021), 1\u20136.","journal-title":"2021 5th International Conference on Pattern Recognition and Image Analysis (IPRIA)"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2916550"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3190508.3190517"},{"key":"e_1_3_2_24_2","unstructured":"Hang Qi Evan R. Sparks and Ameet Talwalkar. 2016. Paleo: A performance model for deep neural networks. (2016)."},{"key":"e_1_3_2_25_2","volume-title":"15th  \\(\\lbrace\\) USENIX \\(\\rbrace\\)  Symposium on Operating Systems Design and Implementation ( \\(\\lbrace\\) OSDI \\(\\rbrace\\)  21)","author":"Qiao Aurick","year":"2021","unstructured":"Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R. Ganger, and Eric P. Xing. 2021. Pollux: Co-adaptive cluster scheduling for goodput-optimized deep learning. In 15th \\(\\lbrace\\) USENIX \\(\\rbrace\\) Symposium on Operating Systems Design and Implementation ( \\(\\lbrace\\) OSDI \\(\\rbrace\\) 21)."},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_2_27_2","article-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014).","journal-title":"arXiv preprint arXiv:1409.1556"},{"key":"e_1_3_2_28_2","first-page":"6105","volume-title":"International Conference on Machine Learning","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc Le. 2019. EfficientNet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning. PMLR, 6105\u20136114."},{"key":"e_1_3_2_29_2","first-page":"595","volume-title":"13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Xiao Wencong","year":"2018","unstructured":"Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, et\u00a0al. 2018. Gandiva: Introspective cluster scheduling for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). 595\u2013610."},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/HiPC.2019.00037"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3470496.3533044"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC.2019.8916403"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2018.8573476"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/MASCOTS.2018.00023"},{"key":"e_1_3_2_35_2","article-title":"Importance of data loading pipeline in training deep neural networks","author":"Zolnouri Mahdi","year":"2020","unstructured":"Mahdi Zolnouri, Xinlin Li, and Vahid Partovi Nia. 2020. Importance of data loading pipeline in training deep neural networks. arXiv preprint arXiv:2005.02130 (2020).","journal-title":"arXiv preprint arXiv:2005.02130"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00907"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3680546","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3680546","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:11Z","timestamp":1750295891000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3680546"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,20]]},"references-count":35,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,12,31]]}},"alternative-id":["10.1145\/3680546"],"URL":"https:\/\/doi.org\/10.1145\/3680546","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11,20]]},"assertion":[{"value":"2023-04-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-07-02","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-11-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}