{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T05:15:16Z","timestamp":1781673316187,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":84,"publisher":"ACM","license":[{"start":{"date-parts":[[2024,4,12]],"date-time":"2024-04-12T00:00:00Z","timestamp":1712880000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,4,12]]},"DOI":"10.1145\/3597503.3639232","type":"proceedings-article","created":{"date-parts":[[2024,4,12]],"date-time":"2024-04-12T16:43:26Z","timestamp":1712940206000},"page":"1-13","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":24,"title":["An Empirical Study on Low GPU Utilization of Deep Learning Jobs"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1899-8561","authenticated-orcid":false,"given":"Yanjie","family":"Gao","sequence":"first","affiliation":[{"name":"Microsoft Research, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-7357-9221","authenticated-orcid":false,"given":"Yichen","family":"He","sequence":"additional","affiliation":[{"name":"Microsoft Research, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-8466-9336","authenticated-orcid":false,"given":"Xinze","family":"Li","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2967-5963","authenticated-orcid":false,"given":"Bo","family":"Zhao","sequence":"additional","affiliation":[{"name":"Microsoft Research, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9148-5861","authenticated-orcid":false,"given":"Haoxiang","family":"Lin","sequence":"additional","affiliation":[{"name":"Microsoft Research, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-1661-0013","authenticated-orcid":false,"given":"Yoyo","family":"Liang","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-2760-9024","authenticated-orcid":false,"given":"Jing","family":"Zhong","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3063-9425","authenticated-orcid":false,"given":"Hongyu","family":"Zhang","sequence":"additional","affiliation":[{"name":"Chongqing University, Chongqing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-2172-588X","authenticated-orcid":false,"given":"Jingzhou","family":"Wang","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-1304-9265","authenticated-orcid":false,"given":"Yonghua","family":"Zeng","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-1298-5321","authenticated-orcid":false,"given":"Keli","family":"Gui","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-9812-0187","authenticated-orcid":false,"given":"Jie","family":"Tong","sequence":"additional","affiliation":[{"name":"Microsoft, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-6455-3898","authenticated-orcid":false,"given":"Mao","family":"Yang","sequence":"additional","affiliation":[{"name":"Microsoft Research, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,4,12]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"TensorFlow: A System for Large-Scale Machine Learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16)","author":"Abadi Martin","year":"2016","unstructured":"Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. 2016. TensorFlow: A System for Large-Scale Machine Learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). USENIX Association, Savannah, GA, 265--283."},{"key":"e_1_3_2_1_2_1","unstructured":"Amazon. 2023. Amazon SageMaker. https:\/\/aws.amazon.com\/sagemaker."},{"key":"e_1_3_2_1_3_1","unstructured":"Weights & Biases. 2023. Current Best Practices for Training LLMs from Scratch. https:\/\/wandb.ai\/site\/llm-whitepaper."},{"key":"e_1_3_2_1_4_1","volume-title":"Workshop on ML Systems, NIPS.","author":"Boag Scott","year":"2017","unstructured":"Scott Boag, Parijat Dube, Benjamin Herta, Waldemar Hummer, Vatche Ishakian, K JAYARAM, Michael Kalantar, Vinod Muthusamy, Priya NAG-PURKAR, and Florian Rosenberg. 2017. Scalable multi-framework multi-tenant lifecycle management of deep learning training jobs. In Workshop on ML Systems, NIPS."},{"key":"e_1_3_2_1_5_1","volume-title":"Proceedings of the Summer Simulation Multi-Conference","author":"Briongos Samira","unstructured":"Samira Briongos, Pedro Malag\u00f3n, Jos\u00e9 L. Risco, and Jos\u00e9 M. Moya. 2017. Building Accurate Models to Determine the Current CPU Utilization of a Host within a Virtual Machine Allocated on It. In Proceedings of the Summer Simulation Multi-Conference (Bellevue, Washington) (SummerSim '17). Society for Computer Simulation International, San Diego, CA, USA, Article 33, 12 pages."},{"key":"e_1_3_2_1_6_1","first-page":"5","article-title":"Borg, Omega, and","volume":"59","author":"Burns Brendan","year":"2016","unstructured":"Brendan Burns, Brian Grant, David Oppenheimer, Eric Brewer, and John Wilkes. 2016. Borg, Omega, and Kubernetes. Commun. ACM 59, 5 (apr 2016), 50--57.","journal-title":"Kubernetes. Commun. ACM"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3540250.3549123"},{"key":"e_1_3_2_1_8_1","volume-title":"TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation. USENIX Association","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation. USENIX Association, Carlsbad, CA, 578--594."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.3115\/1072064.1072067"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1177\/001316446002000104"},{"key":"e_1_3_2_1_11_1","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","volume":"1","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1. Association for Computational Linguistics, Minneapolis, Minnesota, 4171--4186."},{"key":"e_1_3_2_1_12_1","volume-title":"19th USENIX Symposium on Networked Systems Design and Implementation. USENIX Association","author":"Eisenman Assaf","year":"2022","unstructured":"Assaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere, Raghuraman Krishnamoorthi, Krishnakumar Nair, Misha Smelyanskiy, and Murali Annavaram. 2022. Check-N-Run: a Checkpointing System for Training Deep Learning Recommendation Models. In 19th USENIX Symposium on Networked Systems Design and Implementation. USENIX Association, Renton, WA, 929--943."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Yanjie Gao Xianyu Gu Hongyu Zhang Haoxiang Lin and Mao Yang. 2023. Runtime Performance Prediction for Deep Learning Models with Graph Neural Network. In 2023 IEEE\/ACM 45th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). 368--380.","DOI":"10.1109\/ICSE-SEIP58684.2023.00039"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510077"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3368089.3417050"},{"key":"e_1_3_2_1_16_1","volume-title":"An Empirical Study on Quality Issues of Deep Learning Platform. In 2023 IEEE\/ACM 45th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). 455--466","author":"Gao Yanjie","year":"2023","unstructured":"Yanjie Gao, Xiaoxiang Shi, Haoxiang Lin, Hongyu Zhang, Hao Wu, Rui Li, and Mao Yang. 2023. An Empirical Study on Quality Issues of Deep Learning Platform. In 2023 IEEE\/ACM 45th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). 455--466."},{"key":"e_1_3_2_1_17_1","volume-title":"Resource-Guided Configuration Space Reduction for Deep Learning Models. In 2021 IEEE\/ACM 43rd International Conference on Software Engineering (ICSE). 175--187","author":"Gao Yanjie","year":"2021","unstructured":"Yanjie Gao, Yonghao Zhu, Hongyu Zhang, Haoxiang Lin, and Mao Yang. 2021. Resource-Guided Configuration Space Reduction for Deep Learning Models. In 2021 IEEE\/ACM 43rd International Conference on Software Engineering (ICSE). 175--187."},{"key":"e_1_3_2_1_18_1","volume-title":"Deep Learning","author":"Goodfellow Ian","unstructured":"Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning. MIT Press. http:\/\/www.deeplearningbook.org."},{"key":"e_1_3_2_1_19_1","unstructured":"Google. 2022. Best Practices for Performance and Cost Optimization for Machine Learning. http:\/\/web.archive.org\/web\/20220521055530\/https:\/\/cloud.google.com\/architecture\/best-practices-for-ml-performance-cost."},{"key":"e_1_3_2_1_20_1","unstructured":"Google. 2023. Google Vertex AI. https:\/\/cloud.google.com\/vertex-ai."},{"key":"e_1_3_2_1_21_1","volume-title":"DeepProf: Performance Analysis for Deep Learning Applications via Mining GPU Execution Patterns. CoRR abs\/1707.03750","author":"Gu Jiazhen","year":"2017","unstructured":"Jiazhen Gu, Huan Liu, Yangfan Zhou, and Xin Wang. 2017. DeepProf: Performance Analysis for Deep Learning Applications via Mining GPU Execution Patterns. CoRR abs\/1707.03750 (2017). arXiv:1707.03750"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2851613.2851872"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2019.06.100"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2961111.2962602"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476223"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3338906.3338955"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.23919\/ICITST.2017.8356346"},{"key":"e_1_3_2_1_28_1","volume-title":"Proceedings of the 2019 USENIX Conference on Usenix Annual Technical Conference (Renton, WA, USA) (USENIX ATC '19). USENIX Association, USA, 947--960","author":"Jeon Myeongjae","year":"2019","unstructured":"Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, unjie Qian, Wencong Xiao, and Fan Yang. 2019. Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads. In Proceedings of the 2019 USENIX Conference on Usenix Annual Technical Conference (Renton, WA, USA) (USENIX ATC '19). USENIX Association, USA, 947--960."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-59410-7_40"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2254064.2254075"},{"key":"e_1_3_2_1_31_1","unstructured":"Jupyter. 2023. Project Jupyter. https:\/\/jupyter.org."},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476224"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1109\/TC.2008.90","article-title":"Adaptive Fault Management of Parallel Applications for High-Performance Computing","volume":"57","author":"Lan Zhiling","year":"2008","unstructured":"Zhiling Lan and Yawei Li. 2008. Adaptive Fault Management of Parallel Applications for High-Performance Computing. IEEE Trans. Comput. 57, 12 (dec 2008), 1647--1660.","journal-title":"IEEE Trans. Comput."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/HiPC56025.2022.00044"},{"key":"e_1_3_2_1_35_1","unstructured":"Haoyuan Li. 2018. Alluxio: A Virtual Distributed File System. Ph.D. Dissertation. EECS Department University of California Berkeley."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3190508.3190552"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2568225.2568229"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_2_1_39_1","volume-title":"The Eleventh International Conference on Learning Representations.","author":"Lu Yucheng","year":"2023","unstructured":"Yucheng Lu, Conglong Li, Minjia Zhang, Christopher De Sa, and Yuxiong He. 2023. Maximizing Communication Efficiency for Large-scale Training via 0\/1 Adam. In The Eleventh International Conference on Learning Representations."},{"key":"e_1_3_2_1_40_1","volume-title":"Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI '20)","author":"Ma Lingxiao","year":"2020","unstructured":"Lingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue, Youshan Miao, Wei Cui, Wenxiang Hu, Fan Yang, Lintao Zhang, and Lidong Zhou. 2020. Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI '20). USENIX Association, 881--897."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3487043"},{"key":"e_1_3_2_1_42_1","volume-title":"An Empirical Model of Large-Batch Training. CoRR abs\/1812.06162","author":"McCandlish Sam","year":"2018","unstructured":"Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team. 2018. An Empirical Model of Large-Batch Training. CoRR abs\/1812.06162 (2018)."},{"key":"e_1_3_2_1_43_1","volume-title":"GPU Occupancy Prediction of Deep Learning Models Using Graph Neural Network. In 2023 IEEE International Conference on Cluster Computing (CLUSTER). 318--329","author":"Mei Hengquan","year":"2023","unstructured":"Hengquan Mei, Huaizhi Qu, Jingwei Sun, Yanjie Gao, Haoxiang Lin, and Guangzhong Sun. 2023. GPU Occupancy Prediction of Deep Learning Models Using Graph Neural Network. In 2023 IEEE International Conference on Cluster Computing (CLUSTER). 318--329."},{"key":"e_1_3_2_1_44_1","volume-title":"Docker: Lightweight Linux Containers for Consistent Development and Deployment. Linux J.","author":"Merkel Dirk","year":"2014","unstructured":"Dirk Merkel. 2014. Docker: Lightweight Linux Containers for Consistent Development and Deployment. Linux J. 2014, 239, Article 2 (mar 2014)."},{"key":"e_1_3_2_1_45_1","unstructured":"Microsoft. 2018. NNI (Neural Network Intelligence): an open source AutoML toolkit for AutoML lifecycle. https:\/\/github.com\/microsoft\/nni."},{"key":"e_1_3_2_1_46_1","unstructured":"Microsoft. 2023. AzureML Large Scale Deep Learning Best Practices. https:\/\/github.com\/Azure\/azureml-examples\/tree\/main\/best-practices\/largescale-deep-learning."},{"key":"e_1_3_2_1_47_1","unstructured":"Microsoft. 2023. Microsoft Azure Machine Learning. https:\/\/azure.microsoft.com\/en-us\/services\/machine-learning-service."},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"crossref","unstructured":"Ben Mildenhall Pratul P. Srinivasan Matthew Tancik Jonathan T. Barron Ravi Ramamoorthi and Ren Ng. 2020. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. In ECCV.","DOI":"10.1007\/978-3-030-58452-8_24"},{"key":"e_1_3_2_1_49_1","volume-title":"Fine-Grained DNN Checkpointing. In 19th USENIX Conference on File and Storage Technologies (FAST 21)","author":"Mohan Jayashree","year":"2021","unstructured":"Jayashree Mohan, Amar Phanishayee, and Vijay Chidambaram. 2021. CheckFreq: Frequent, Fine-Grained DNN Checkpointing. In 19th USENIX Conference on File and Storage Technologies (FAST 21). USENIX Association, 203--216."},{"key":"e_1_3_2_1_50_1","volume-title":"Using LSTM and SARIMA Models to Forecast Cluster CPU Usage. CoRR abs\/2007.08092","author":"Nashold Langston","year":"2020","unstructured":"Langston Nashold and Rayan Krishnan. 2020. Using LSTM and SARIMA Models to Forecast Cluster CPU Usage. CoRR abs\/2007.08092 (2020). arXiv:2007.08092"},{"key":"e_1_3_2_1_51_1","volume-title":"DeepFreeze: Towards Scalable Asynchronous Checkpointing of Deep Learning Models. In 2020 20th IEEE\/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGRID). 172--181","author":"Nicolae Bogdan","year":"2020","unstructured":"Bogdan Nicolae, Jiali Li, Justin M. Wozniak, George Bosilca, Matthieu Dorier, and Franck Cappello. 2020. DeepFreeze: Towards Scalable Asynchronous Checkpointing of Deep Learning Models. In 2020 20th IEEE\/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGRID). 172--181."},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2015.100"},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-4009"},{"key":"e_1_3_2_1_55_1","volume-title":"Proceedings of the 3rd International Conference on Distributed Computing Systems. 22--30","author":"Ousterhout J.K.","year":"1982","unstructured":"J.K. Ousterhout. 1982. Scheduling techniques for concurrent systems. In Proceedings of the 3rd International Conference on Distributed Computing Systems. 22--30."},{"key":"e_1_3_2_1_56_1","volume-title":"High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32","volume":"32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32, Vol. 32. Curran Associates, Inc., 8024--8035."},{"key":"e_1_3_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078195"},{"key":"e_1_3_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476311.3476390"},{"key":"e_1_3_2_1_59_1","unstructured":"PyTorch. 2022. Data Loading Utility. https:\/\/pytorch.org\/docs\/1.12\/data.html."},{"key":"e_1_3_2_1_60_1","volume-title":"Proceedings of ICLR.","author":"Qi Hang","year":"2017","unstructured":"Hang Qi, Evan R. Sparks, and Ameet Talwalkar. 2017. Paleo: A Performance Model for Deep Neural Networks. In Proceedings of ICLR."},{"key":"e_1_3_2_1_61_1","volume-title":"Prometheus: A Next-Generation Monitoring System (Talk). In SREcon15 Europe","author":"Rabenstein Bj\u00f6rn","year":"2015","unstructured":"Bj\u00f6rn Rabenstein and Julius Volz. 2015. Prometheus: A Next-Generation Monitoring System (Talk). In SREcon15 Europe. USENIX Association, Dublin."},{"key":"e_1_3_2_1_62_1","unstructured":"Alec Radford Jeff Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language Models are Unsupervised MultitaskLearners. (2019)."},{"key":"e_1_3_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458336.3465287"},{"key":"e_1_3_2_1_64_1","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis","author":"Rajbhandari Samyam","year":"2020","unstructured":"Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020. ZeRO: Memory Optimizations toward Training Trillion Parameter Models. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (Atlanta, Georgia) (SC '20). IEEE Press, Article 20, 16 pages."},{"key":"e_1_3_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3406703"},{"key":"e_1_3_2_1_66_1","volume-title":"2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 1--13","author":"Rhu Minsoo","unstructured":"Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, and Stephen W. Keckler. 2016. vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design. In 2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 1--13."},{"key":"e_1_3_2_1_67_1","volume-title":"Horovod: fast and easy distributed deep learning in TensorFlow. CoRR abs\/1802.05799","author":"Sergeev Alexander","year":"2018","unstructured":"Alexander Sergeev and Mike Del Balso. 2018. Horovod: fast and easy distributed deep learning in TensorFlow. CoRR abs\/1802.05799 (2018). arXiv:1802.05799"},{"key":"e_1_3_2_1_68_1","volume-title":"Measuring the Effects of Data Parallelism on Neural Network Training. Journal of Machine Learning Research (JMLR)","author":"Shallue Chris","year":"2018","unstructured":"Chris Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-dickstein, Roy Frostig, and George Dahl. 2018. Measuring the Effects of Data Parallelism on Neural Network Training. Journal of Machine Learning Research (JMLR) (2018)."},{"key":"e_1_3_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3468264.3468591"},{"key":"e_1_3_2_1_70_1","unstructured":"StackOverflow. 2011. Why is CUDA pinned memory so fast? https:\/\/stackoverflow.com\/questions\/5736968\/why-is-cuda-pinned-memory-so-fast."},{"key":"e_1_3_2_1_71_1","unstructured":"TensorFlow. 2023. Get Started with TensorFlow Transform. https:\/\/www.tensorflow.org\/tfx\/transform\/get_started."},{"key":"e_1_3_2_1_72_1","volume-title":"Manso","author":"Thompson Neil C.","year":"2020","unstructured":"Neil C. Thompson, Kristjan H. Greenewald, Keeheon Lee, and Gabriel F. Manso. 2020. The Computational Limits of Deep Learning. CoRR abs\/2007.05558 (2020)."},{"key":"e_1_3_2_1_73_1","unstructured":"Kenton Varda et al. 2013. Cap'n Proto serialization\/RPC system - core tools and C++ library. https:\/\/github.com\/capnproto\/capnproto."},{"key":"e_1_3_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1145\/3437984.3458831"},{"key":"e_1_3_2_1_75_1","volume-title":"HuggingFace's Transformers: State-of-the-art Natural Language Processing. CoRR abs\/1910.03771","author":"Wolf Thomas","year":"2019","unstructured":"Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R\u00e9mi Louf, Morgan Funtowicz, and Jamie Brew. 2019. HuggingFace's Transformers: State-of-the-art Natural Language Processing. CoRR abs\/1910.03771 (2019). arXiv:1910.03771"},{"key":"e_1_3_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3064966"},{"key":"e_1_3_2_1_77_1","volume-title":"Proceedings of the 13th USENIX Conference on Operating Systems Design and Implementation. USENIX Association, USA, 595--610","author":"Xiao Wencong","year":"2018","unstructured":"Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, Fan Yang, and Lidong Zhou. 2018. Gandiva: Introspective Cluster Scheduling for Deep Learning. In Proceedings of the 13th USENIX Conference on Operating Systems Design and Implementation. USENIX Association, USA, 595--610."},{"key":"e_1_3_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00087"},{"key":"e_1_3_2_1_79_1","volume-title":"Towards GPU Utilization Prediction for Cloud Deep Learning. In 12th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 20)","author":"Yeung Gingfung","year":"2020","unstructured":"Gingfung Yeung, Damian Borowiec, Adrian Friday, Richard Harper, and Peter Garraghan. 2020. Towards GPU Utilization Prediction for Cloud Deep Learning. In 12th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 20). USENIX Association."},{"key":"e_1_3_2_1_80_1","volume-title":"Scaling SGD Batch Size to 32K for ImageNet Training. CoRR abs\/1708.03888","author":"You Yang","year":"2017","unstructured":"Yang You, Igor Gitman, and Boris Ginsburg. 2017. Scaling SGD Batch Size to 32K for ImageNet Training. CoRR abs\/1708.03888 (2017). arXiv:1708.03888"},{"key":"e_1_3_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.1145\/2934664"},{"key":"e_1_3_2_1_82_1","volume-title":"An Empirical Study on Program Failures of Deep Learning Jobs. In 2020 IEEE\/ACM 42nd International Conference on Software Engineering (ICSE). 1159--1170","author":"Zhang Ru","year":"2020","unstructured":"Ru Zhang, Wencong Xiao, Hongyu Zhang, Yu Liu, Haoxiang Lin, and Mao Yang. 2020. An Empirical Study on Program Failures of Deep Learning Jobs. In 2020 IEEE\/ACM 42nd International Conference on Software Engineering (ICSE). 1159--1170."},{"key":"e_1_3_2_1_83_1","volume-title":"An Empirical Study of Common Challenges in Developing Deep Learning Applications. In 2019 IEEE 30th International Symposium on Software Reliability Engineering (ISSRE). 104--115","author":"Zhang Tianyi","year":"2019","unstructured":"Tianyi Zhang, Cuiyun Gao, Lei Ma, Michael Lyu, and Miryung Kim. 2019. An Empirical Study of Common Challenges in Developing Deep Learning Applications. In 2019 IEEE 30th International Symposium on Software Reliability Engineering (ISSRE). 104--115."},{"key":"e_1_3_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.1145\/3213846.3213866"},{"key":"e_1_3_2_1_85_1","doi-asserted-by":"publisher","DOI":"10.1145\/3552326.3567499"}],"event":{"name":"ICSE '24: IEEE\/ACM 46th International Conference on Software Engineering","location":"Lisbon Portugal","acronym":"ICSE '24","sponsor":["SIGSOFT ACM Special Interest Group on Software Engineering","IEEE CS","Faculty of Engineering of University of Porto"]},"container-title":["Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3597503.3639232","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3597503.3639232","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:50:12Z","timestamp":1750287012000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3597503.3639232"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,12]]},"references-count":84,"alternative-id":["10.1145\/3597503.3639232","10.1145\/3597503"],"URL":"https:\/\/doi.org\/10.1145\/3597503.3639232","relation":{},"subject":[],"published":{"date-parts":[[2024,4,12]]},"assertion":[{"value":"2024-04-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}