{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,28]],"date-time":"2026-07-28T03:07:07Z","timestamp":1785208027175,"version":"3.55.0"},"reference-count":80,"publisher":"Association for Computing Machinery (ACM)","issue":"5","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,1]]},"abstract":"<jats:p>\n            Training Deep Neural Networks (DNNs) is resource-intensive and time-consuming. While prior research has explored many different ways of reducing DNN training time, the impact of\n            <jats:italic>input data pipeline<\/jats:italic>\n            , i.e., fetching raw data items from storage and performing data pre-processing in memory, has been relatively unexplored. This paper makes the following contributions: (1) We present the first comprehensive analysis of how the input data pipeline affects the training time of widely-used computer vision and audio Deep Neural Networks (DNNs), that typically involve complex data pre-processing. We analyze nine different models across three tasks and four datasets while varying factors such as the amount of memory, number of CPU threads, storage device, GPU generation etc on servers that are a part of a large production cluster at Microsoft. We find that in many cases, DNN training time is dominated by\n            <jats:italic>data stall time<\/jats:italic>\n            : time spent waiting for data to be fetched and pre-processed. (2) We build a tool, DS-Analyzer to precisely measure data stalls using a differential technique, and perform predictive what-if analysis on data stalls. (3) Finally, based on the insights from our analysis, we design and implement three simple but effective techniques in a data-loading library, CoorDL, to mitigate data stalls. Our experiments on a range of DNN tasks, models, datasets, and hardware configs show that when PyTorch uses CoorDL instead of the state-of-the-art DALI data loading library, DNN training time is reduced significantly (by as much as 5X on a single server).\n          <\/jats:p>","DOI":"10.14778\/3446095.3446100","type":"journal-article","created":{"date-parts":[[2021,3,23]],"date-time":"2021-03-23T16:36:58Z","timestamp":1616517418000},"page":"771-784","source":"Crossref","is-referenced-by-count":92,"title":["Analyzing and mitigating data stalls in DNN training"],"prefix":"10.14778","volume":"14","author":[{"given":"Jayashree","family":"Mohan","sequence":"first","affiliation":[{"name":"University of Texas at Austin"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Amar","family":"Phanishayee","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ashish","family":"Raniwala","sequence":"additional","affiliation":[{"name":"Microsoft"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Vijay","family":"Chidambaram","sequence":"additional","affiliation":[{"name":"University of Texas at Austin &amp; VMWare Research"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,3,23]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2020. AWS Instance Types. https:\/\/aws.amazon.com\/ec2\/instance-types\/#p3.  2020. AWS Instance Types. https:\/\/aws.amazon.com\/ec2\/instance-types\/#p3."},{"key":"e_1_2_1_2_1","unstructured":"2020. AWS Instance Types. https:\/\/aws.amazon.eom\/ec2\/instance-types\/#p2.  2020. AWS Instance Types. https:\/\/aws.amazon.eom\/ec2\/instance-types\/#p2."},{"key":"e_1_2_1_3_1","unstructured":"2020. Blobfuse. https:\/\/github.com\/Azure\/azure-storage-fuse.  2020. Blobfuse. https:\/\/github.com\/Azure\/azure-storage-fuse."},{"key":"e_1_2_1_4_1","unstructured":"2020. Cloud TPU Tools. https:\/\/cloud.google.com\/tpu\/docs\/cloud-tpu-tools.  2020. Cloud TPU Tools. https:\/\/cloud.google.com\/tpu\/docs\/cloud-tpu-tools."},{"key":"e_1_2_1_5_1","unstructured":"2020. DALI: Supported Operations. https:\/\/docs.nvidia.com\/deeplearning\/dali\/user-guide\/docs\/supported_ops.html#nvidia.dali.ops.FileReader.  2020. DALI: Supported Operations. https:\/\/docs.nvidia.com\/deeplearning\/dali\/user-guide\/docs\/supported_ops.html#nvidia.dali.ops.FileReader."},{"key":"e_1_2_1_6_1","volume-title":"EBS"},{"key":"e_1_2_1_7_1","unstructured":"2020. Fast AI Data Preprocessing with NVIDIA DALI. https:\/\/devblogs.nvidia.com\/fast-ai-data-preprocessing-with-nvidia-dali\/.  2020. Fast AI Data Preprocessing with NVIDIA DALI. https:\/\/devblogs.nvidia.com\/fast-ai-data-preprocessing-with-nvidia-dali\/."},{"key":"e_1_2_1_8_1","unstructured":"2020. ImageNet-22k. http:\/\/www.image-net.org\/releases.  2020. ImageNet-22k. http:\/\/www.image-net.org\/releases."},{"key":"e_1_2_1_9_1","unstructured":"2020. Microsoft Philly Traces. https:\/\/github.com\/msr-fiddle\/philly-traces.  2020. Microsoft Philly Traces. https:\/\/github.com\/msr-fiddle\/philly-traces."},{"key":"e_1_2_1_10_1","unstructured":"2020. NVIDIA DGX-2: Enterprise AI Research System. https:\/\/www.nvidia.com\/en-us\/data-center\/dgx-2\/.  2020. NVIDIA DGX-2: Enterprise AI Research System. https:\/\/www.nvidia.com\/en-us\/data-center\/dgx-2\/."},{"key":"e_1_2_1_11_1","unstructured":"2020. NVIDIA Object Detection. https:\/\/github.com\/NVIDIA\/DeepLearningExamples\/tree\/master\/PyTorch\/Detection\/SSD.  2020. NVIDIA Object Detection. https:\/\/github.com\/NVIDIA\/DeepLearningExamples\/tree\/master\/PyTorch\/Detection\/SSD."},{"key":"e_1_2_1_12_1","unstructured":"2020. NVIDIA Profiler. https:\/\/docs.nvidia.com\/cuda\/profiler-users-guide\/index.html.  2020. NVIDIA Profiler. https:\/\/docs.nvidia.com\/cuda\/profiler-users-guide\/index.html."},{"key":"e_1_2_1_13_1","unstructured":"2020. Profiling MXNet models. https:\/\/mxnet.apache.org\/api\/python\/docs\/tutorials\/performance\/backend\/profiler.html.  2020. Profiling MXNet models. https:\/\/mxnet.apache.org\/api\/python\/docs\/tutorials\/performance\/backend\/profiler.html."},{"key":"e_1_2_1_14_1","unstructured":"2020. TorchAudio classifier. https:\/\/pytorch.org\/tutorials\/beginner\/audio_classifier_tutorial.html?highlight=audio.  2020. TorchAudio classifier. https:\/\/pytorch.org\/tutorials\/beginner\/audio_classifier_tutorial.html?highlight=audio."},{"key":"e_1_2_1_15_1","unstructured":"2020. TorchVision models. https:\/\/pytorch.org\/docs\/stable\/torchvision\/models.html.  2020. TorchVision models. https:\/\/pytorch.org\/docs\/stable\/torchvision\/models.html."},{"key":"e_1_2_1_16_1","unstructured":"2020. Training a Champion: Building Deep Neural Nets for Big Data Analytics. https:\/\/www.kdnuggets.com\/training-a-champion-building-deep-neural-nets-for-big-data-analytics.html\/.  2020. Training a Champion: Building Deep Neural Nets for Big Data Analytics. https:\/\/www.kdnuggets.com\/training-a-champion-building-deep-neural-nets-for-big-data-analytics.html\/."},{"key":"e_1_2_1_17_1","unstructured":"2020. WMT16. http:\/\/www.statmt.org\/wmt16\/.  2020. WMT16. http:\/\/www.statmt.org\/wmt16\/."},{"key":"e_1_2_1_18_1","volume-title":"Youtube-8m: A large-scale video classification benchmark. arXiv preprint arXiv:1609.08675","author":"Abu-El-Haija Sami","year":"2016"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/2503308.2188395"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3342195.3387555"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.5555\/3291168.3291211"},{"key":"e_1_2_1_22_1","volume-title":"Training Deep Nets with Sublinear Memory Cost. arXiv preprint arXiv:1604.06174","author":"Chen Tianqi","year":"2016"},{"key":"e_1_2_1_23_1","volume-title":"Ubershuffle: Communication-efficient data shuffling for sgd via coding theory. Advances in Neural Information Processing Systems (NIPS)","author":"Chung Jichan","year":"2017"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952190"},{"key":"e_1_2_1_25_1","volume-title":"Fma: A dataset for music analysis. arXiv preprint arXiv:1612.01840","author":"Defferrard Micha\u00ebl","year":"2016"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/n19-1423"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098043"},{"key":"e_1_2_1_28_1","unstructured":"Mel Gorman. 2020. Understanding the Linux Virtual Memory Manager. https:\/\/www.kernel.org\/doc\/gorman\/html\/understand\/understand013.html.  Mel Gorman. 2020. Understanding the Linux Virtual Memory Manager. https:\/\/www.kernel.org\/doc\/gorman\/html\/understand\/understand013.html."},{"key":"e_1_2_1_29_1","volume-title":"large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677","author":"Goyal Priya","year":"2017"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.5555\/3323234.3323274"},{"key":"e_1_2_1_32_1","volume-title":"Proceedings of Machine Learning and Systems 2019","author":"Hashemi Sayed Hadi","year":"2019"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.5555\/3294771.3294936"},{"key":"e_1_2_1_35_1","volume-title":"SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and &lt","author":"Iandola Forrest N","year":"2016"},{"key":"e_1_2_1_36_1","volume-title":"Game Development Tools.","author":"Iyer Kumar"},{"key":"e_1_2_1_37_1","unstructured":"Max Jaderberg Valentin Dalibard Simon Osindero Wojciech M Czarnecki Jeff Donahue Ali Razavi Oriol Vinyals Tim Green Iain Dunning Karen Simonyan etal 2017. Population based training of neural networks. arXiv preprint arXiv:1711.09846 (2017).  Max Jaderberg Valentin Dalibard Simon Osindero Wojciech M Czarnecki Jeff Donahue Ali Razavi Oriol Vinyals Tim Green Iain Dunning Karen Simonyan et al. 2017. Population based training of neural networks. arXiv preprint arXiv:1711.09846 (2017)."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00070"},{"key":"e_1_2_1_39_1","volume-title":"Proceedings of Machine Learning and Systems 2019","author":"Jayarajan Anand","year":"2019"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.5555\/3358807.3358888"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359630"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.5555\/3357034.3357049"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.5555\/2999134.2999257"},{"key":"e_1_2_1_44_1","volume-title":"Quiver: An Informed Storage Cache for Deep Learning. In 18th USENIX Conference on File and Storage Technologies, FAST 2020","author":"Kumar Abhishek Vijaya","year":"2020"},{"key":"e_1_2_1_45_1","unstructured":"Alina Kuznetsova Hassan Rom Neil Alldrin Jasper Uijlings Ivan Krasin Jordi Pont-Tuset Shahab Kamali Stefan Popov Matteo Malloci Tom Duerig etal 2018. The open images dataset v4: Unified image classification object detection and visual relationship detection at scale. arXiv preprint arXiv:1811.00982 (2018).  Alina Kuznetsova Hassan Rom Neil Alldrin Jasper Uijlings Ivan Krasin Jordi Pont-Tuset Shahab Kamali Stefan Popov Matteo Malloci Tom Duerig et al. 2018. The open images dataset v4: Unified image classification object detection and visual relationship detection at scale. arXiv preprint arXiv:1811.00982 (2018)."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.5555\/3122009.3242042"},{"key":"e_1_2_1_47_1","volume-title":"Tune: A research platform for distributed model selection and training. arXiv preprint arXiv:1807.05118","author":"Liaw Richard","year":"2018"},{"key":"e_1_2_1_48_1","volume-title":"Use Local SGD. In 8th International Conference on Learning Representations, ICLR 2020","author":"Lin Tao","year":"2020"},{"key":"e_1_2_1_49_1","volume-title":"Deep gradient compression: Reducing the communication bandwidth for distributed training. arXiv preprint arXiv:1712.01887","author":"Lin Yujun","year":"2017"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"e_1_2_1_51_1","volume-title":"Themis: Fair and Efficient GPU Cluster Scheduling. In 17th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2020","author":"Mahajan Kshiteej","year":"2020"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2019.01.037"},{"key":"e_1_2_1_53_1","unstructured":"Apache Mesos. 2020. RecordIO Data Format. https:\/\/mesos.apache.org\/documentation\/latest\/recordio\/.  Apache Mesos. 2020. RecordIO Data Format. https:\/\/mesos.apache.org\/documentation\/latest\/recordio\/."},{"key":"e_1_2_1_54_1","unstructured":"MLPerf. 2020. MLPerf Training Results v0.7. https:\/\/github.com\/mlperf\/training_results_v0.7.  MLPerf. 2020. MLPerf Training Results v0.7. https:\/\/github.com\/mlperf\/training_results_v0.7."},{"key":"e_1_2_1_55_1","volume-title":"Analyzing and Mitigating Data Stalls in DNN Training. CoRR abs\/2007.06775","author":"Mohan Jayashree","year":"2020"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407816"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359646"},{"key":"e_1_2_1_58_1","volume-title":"NIPS Workshop on Systems for Machine Learning (December","author":"Narayanan Deepak","year":"2018"},{"key":"e_1_2_1_59_1","unstructured":"Andrew NG. 2020. Data and DNNs. https:\/\/www.wired.com\/brandlab\/2015\/05\/andrew-ng-deep-learning-mandate-humans-not-just-machines\/.  Andrew NG. 2020. Data and DNNs. https:\/\/www.wired.com\/brandlab\/2015\/05\/andrew-ng-deep-learning-mandate-humans-not-just-machines\/."},{"key":"e_1_2_1_60_1","volume-title":"The effectiveness of data augmentation in image classification using deep learning. arXiv preprint arXiv:1712.04621","author":"Perez Luis","year":"2017"},{"key":"e_1_2_1_61_1","article-title":"Tunability: Importance of Hyperparameters of Machine Learning Algorithms","volume":"20","author":"Probst Philipp","year":"2019","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_2_1_62_1","unstructured":"PyTorch. 2020. PyTorch Training Examples. https:\/\/github.com\/pytorch\/examples\/tree\/master\/imagenet.  PyTorch. 2020. PyTorch Training Examples. https:\/\/github.com\/pytorch\/examples\/tree\/master\/imagenet."},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.5555\/3195638.3195660"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_2_1_66_1","unstructured":"Supheakmungkol Sarin Knot Pipatsrisawat Khiem Pham Anurag Batra and Lu\u00eds Valente. 2019. Crowdsource by Google: A Platform for Collecting Inclusive and Representative Machine Learning Data.  Supheakmungkol Sarin Knot Pipatsrisawat Khiem Pham Anurag Batra and Lu\u00eds Valente. 2019. Crowdsource by Google: A Platform for Collecting Inclusive and Representative Machine Learning Data."},{"key":"e_1_2_1_67_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014"},{"key":"e_1_2_1_68_1","volume-title":"Don't decay the learning rate, increase the batch size. arXiv preprint arXiv:1711.00489","author":"Smith Samuel L","year":"2017"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/358699.358703"},{"key":"e_1_2_1_70_1","unstructured":"Mustafa Suleyman. 2020. Using AI to give doctors a 48-hour head start on life-threatening illness. https:\/\/deepmind.com\/blog\/article\/predicting-patient-deterioration.  Mustafa Suleyman. 2020. Using AI to give doctors a 48-hour head start on life-threatening illness. https:\/\/deepmind.com\/blog\/article\/predicting-patient-deterioration."},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/258623.258680"},{"key":"e_1_2_1_72_1","volume-title":"Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions. arXiv preprint arXiv:1802.04730","author":"Vasilache Nicolas","year":"2018"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.515"},{"key":"e_1_2_1_74_1","unstructured":"Yonghui Wu Mike Schuster Zhifeng Chen Quoc V Le Mohammad Norouzi Wolfgang Macherey Maxim Krikun Yuan Cao Qin Gao Klaus Macherey etal 2016. Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation. arXiv preprint arXiv:1609.08144 (2016).  Yonghui Wu Mike Schuster Zhifeng Chen Quoc V Le Mohammad Norouzi Wolfgang Macherey Maxim Krikun Yuan Cao Qin Gao Klaus Macherey et al. 2016. Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation. arXiv preprint arXiv:1609.08144 (2016)."},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.5555\/3291168.3291212"},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.5555\/3154690.3154708"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00716"},{"key":"e_1_2_1_78_1","volume-title":"Daydream: Accurately Estimating the Efficacy of Optimizations for DNN Training. In 2020 USENIX Annual Technical Conference, USENIX ATC 2020","author":"Zhu Hongyu","year":"2020"},{"key":"e_1_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.1109\/MASCOTS.2018.00023"},{"key":"e_1_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.11"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3446095.3446100","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:21:29Z","timestamp":1672226489000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3446095.3446100"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,1]]},"references-count":80,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2021,1]]}},"alternative-id":["10.14778\/3446095.3446100"],"URL":"https:\/\/doi.org\/10.14778\/3446095.3446100","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,1]]}}}