{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T20:26:37Z","timestamp":1782937597577,"version":"3.54.5"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2019,6,30]],"date-time":"2019-06-30T00:00:00Z","timestamp":1561852800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"U.S. Department of Energy, Office of Science, Advanced Scientific Computing Research","award":["DE-AC02-06CH11357"],"award-info":[{"award-number":["DE-AC02-06CH11357"]}]},{"name":"NSF XPS","award":["CCF-1337131"],"award-info":[{"award-number":["CCF-1337131"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Parallel Comput."],"published-print":{"date-parts":[[2019,6,30]]},"abstract":"<jats:p>Scalable deep neural network training has been gaining prominence because of the increasing importance of deep learning in a multitude of scientific and commercial domains. Consequently, a number of researchers have investigated techniques to optimize deep learning systems. Much of the prior work has focused on runtime and algorithmic enhancements to optimize the computation and communication. Despite these enhancements, however, deep learning systems still suffer from scalability limitations, particularly with respect to data I\/O. This situation is especially true for training models where the computation can be effectively parallelized, leaving I\/O as the major bottleneck. In fact, our analysis shows that I\/O can take up to 90% of the total training time. Thus, in this article, we first analyze LMDB, the most widely used I\/O subsystem of deep learning frameworks, to understand the causes of this I\/O inefficiency. Based on our analysis, we propose LMDBIO\u2014an optimized I\/O plugin for scalable deep learning. LMDBIO includes six novel optimizations that together address the various shortcomings in existing I\/O for deep learning. Our experimental results show that LMDBIO significantly outperforms LMDB in all cases and improves overall application performance by up to 65-fold on a 9,216-core system.<\/jats:p>","DOI":"10.1145\/3331526","type":"journal-article","created":{"date-parts":[[2019,7,1]],"date-time":"2019-07-01T19:19:22Z","timestamp":1562008762000},"page":"1-34","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":40,"title":["Scalable Deep Learning via I\/O Analysis and Optimization"],"prefix":"10.1145","volume":"6","author":[{"given":"Sarunya","family":"Pumma","sequence":"first","affiliation":[{"name":"Department of Computer Science, Virginia Tech, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Min","family":"Si","sequence":"additional","affiliation":[{"name":"Mathematics and Computer Science Division, Argonne National Laboratory, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wu-Chun","family":"Feng","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Virginia Tech, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pavan","family":"Balaji","sequence":"additional","affiliation":[{"name":"Mathematics and Computer Science Division, Argonne National Laboratory, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,7]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"{n.d.}. NVIDIA Collective Communications Library (NCCL): Multi-GPU and Multi-Node Collective Communication Primitives. Retrieved from https:\/\/developer.nvidia.com\/nccl.  {n.d.}. NVIDIA Collective Communications Library (NCCL): Multi-GPU and Multi-Node Collective Communication Primitives. Retrieved from https:\/\/developer.nvidia.com\/nccl."},{"key":"e_1_2_1_2_1","unstructured":"2015. Caffe-MPI for Deep Learning. Retrieved from https:\/\/github.com\/Caffe-MPI\/Caffe-MPI.github.io.  2015. Caffe-MPI for Deep Learning. Retrieved from https:\/\/github.com\/Caffe-MPI\/Caffe-MPI.github.io."},{"key":"e_1_2_1_3_1","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dan Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. Retrieved from http:\/\/tensorflow.org\/ Software available from tensorflow.org.  Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dan Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. Retrieved from http:\/\/tensorflow.org\/ Software available from tensorflow.org."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3018743.3018769"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2017.2752706"},{"key":"e_1_2_1_6_1","unstructured":"Nicolas Castet. 2018. Distributed deep learning with Horovod and PowerAI DDL. Retrieved from https:\/\/developer.ibm.com\/linuxonpower\/2018\/08\/24\/distributed-deep-learning-horovod-powerai-ddl\/.  Nicolas Castet. 2018. Distributed deep learning with Horovod and PowerAI DDL. Retrieved from https:\/\/developer.ibm.com\/linuxonpower\/2018\/08\/24\/distributed-deep-learning-horovod-powerai-ddl\/."},{"key":"e_1_2_1_7_1","volume-title":"MXNET: A flexible and efficient machine learning library for heterogeneous distributed systems. arXiv preprint arXiv:1512.01274","author":"Chen Tianqi","year":"2015"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/PDSW-DISCS.2018.00011"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/38.56302"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_11_1","unstructured":"Facebook. 2017. Gloo. Retrieved from https:\/\/github.com\/facebookincubator\/gloo\/blob\/master\/docs\/readme.md.  Facebook. 2017. Gloo. Retrieved from https:\/\/github.com\/facebookincubator\/gloo\/blob\/master\/docs\/readme.md."},{"key":"e_1_2_1_12_1","unstructured":"Michael Feldman. 2017. Intel Spills Details on Knights Mill Processor. Retrieved from https:\/\/www.top500.org\/news\/intel-spills-details-on-knights-mill-processor\/.  Michael Feldman. 2017. Intel Spills Details on Knights Mill Processor. Retrieved from https:\/\/www.top500.org\/news\/intel-spills-details-on-knights-mill-processor\/."},{"key":"e_1_2_1_13_1","unstructured":"Andrew Gibiansky. {n.d.}. Bringing HPC Techniques to Deep Learning. Retrieved from http:\/\/andrew.gibiansky.com.  Andrew Gibiansky. {n.d.}. Bringing HPC Techniques to Deep Learning. Retrieved from http:\/\/andrew.gibiansky.com."},{"key":"e_1_2_1_14_1","unstructured":"Google. 2018. Cloud Tensor Processing Units (TPUs). Retrieved from https:\/\/cloud.google.com\/tpu\/docs\/tpus.  Google. 2018. Cloud Tensor Processing Units (TPUs). Retrieved from https:\/\/cloud.google.com\/tpu\/docs\/tpus."},{"key":"e_1_2_1_15_1","volume-title":"large minibatch SGD: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677","author":"Goyal Priya","year":"2017"},{"key":"e_1_2_1_16_1","unstructured":"The HDF Group. 2012. Enabling a Strict Consistency Semantics Model in Parallel HDF5. Retrieved from https:\/\/support.hdfgroup.org\/HDF5\/doc\/Advanced\/PHDF5FileConsistencySemantics\/PHDF5FileConsistencySemantics.pdf.  The HDF Group. 2012. Enabling a Strict Consistency Semantics Model in Parallel HDF5. Retrieved from https:\/\/support.hdfgroup.org\/HDF5\/doc\/Advanced\/PHDF5FileConsistencySemantics\/PHDF5FileConsistencySemantics.pdf."},{"key":"e_1_2_1_17_1","volume-title":"Deep Learning with Keras","author":"Gulli Antonio"},{"key":"e_1_2_1_18_1","unstructured":"Mark Harris. 2017. NVIDIA DGX-1: The Fastest Deep Learning System. Retrieved from https:\/\/devblogs.nvidia.com\/dgx-1-fastest-deep-learning-system\/.  Mark Harris. 2017. NVIDIA DGX-1: The Fastest Deep Learning System. Retrieved from https:\/\/devblogs.nvidia.com\/dgx-1-fastest-deep-learning-system\/."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3208040.3208045"},{"key":"e_1_2_1_21_1","unstructured":"Jeremy Hsu. 2016. Fujitsu Memory Tech Speeds Up Deep-Learning AI. Retrieved from https:\/\/spectrum.ieee.org\/tech-talk\/computing\/software\/fujitsu-memory-tech-speeds-up-deep-learning-ai.  Jeremy Hsu. 2016. Fujitsu Memory Tech Speeds Up Deep-Learning AI. Retrieved from https:\/\/spectrum.ieee.org\/tech-talk\/computing\/software\/fujitsu-memory-tech-speeds-up-deep-learning-ai."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.284"},{"key":"e_1_2_1_23_1","volume-title":"LIRS: Enabling efficient machine learning on NVM-based storage via a lightweight implementation of random shuffling. arXiv preprint arXiv:1810.04509","author":"Ke Zhi-Lin","year":"2018"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICISCT.2016.7777390"},{"key":"e_1_2_1_25_1","volume-title":"One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv:1404.5997","author":"Krizhevsky Alex","year":"2014"},{"key":"e_1_2_1_27_1","volume-title":"Efficient training of convolutional neural nets on large distributed systems. arXiv preprint arXiv:1711.00705","author":"Kumar Sameer","year":"2017"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2018.00054"},{"key":"e_1_2_1_29_1","volume-title":"Why M heads are better than one: Training a diverse ensemble of deep networks. arXiv","author":"Lee Stefan","year":"2015"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1048935.1050189"},{"key":"e_1_2_1_31_1","volume-title":"Taylor","author":"Ma He","year":"2016"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2018.00068"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CCGRID.2018.00072"},{"key":"e_1_2_1_34_1","unstructured":"Microsoft. {n.d.}. Cognitive Toolkit: Multiple GPUs and Machines. Retrieved from https:\/\/docs.microsoft.com\/en-us\/cognitive-toolkit\/multiple-gpus-and-machines.  Microsoft. {n.d.}. Cognitive Toolkit: Multiple GPUs and Machines. Retrieved from https:\/\/docs.microsoft.com\/en-us\/cognitive-toolkit\/multiple-gpus-and-machines."},{"key":"e_1_2_1_35_1","unstructured":"Timothy Prickett Morgan. 2017. Machine Learning Gets an InfiniBand Boost with Caffe2. Retrieved from https:\/\/www.nextplatform.com\/2017\/04\/19\/machine-learning-gets-infiniband-boost-caffe2\/.  Timothy Prickett Morgan. 2017. Machine Learning Gets an InfiniBand Boost with Caffe2. Retrieved from https:\/\/www.nextplatform.com\/2017\/04\/19\/machine-learning-gets-infiniband-boost-caffe2\/."},{"key":"e_1_2_1_36_1","unstructured":"NVIDIA. 2018. NVIDIA Deep Learning Platform: Giant Leaps in Performance and Efficiency for AI Services From the Data Center to the Network\u2019s Edge. Retrieved from https:\/\/images.nvidia.com\/content\/pdf\/inference-technical-overview.pdf.  NVIDIA. 2018. NVIDIA Deep Learning Platform: Giant Leaps in Performance and Efficiency for AI Services From the Data Center to the Network\u2019s Edge. Retrieved from https:\/\/images.nvidia.com\/content\/pdf\/inference-technical-overview.pdf."},{"key":"e_1_2_1_37_1","volume-title":"A Guide to NumPy","author":"Oliphant Travis E."},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the Conference and Workshop on Neural Information Processing Systems (NIPS-W\u201917)","author":"Paszke Adam","year":"2017"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPADS.2017.00097"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCC-SmartCity-DSS.2017.29"},{"key":"e_1_2_1_41_1","volume-title":"Summer School on Machine Learning","author":"Rasmussen Carl Edward"},{"key":"e_1_2_1_42_1","unstructured":"Baidu Research. {n.d.}. baidu-allreduce. Retrieved from https:\/\/github.com\/baidu-research\/baidu-allreduce.  Baidu Research. {n.d.}. baidu-allreduce. Retrieved from https:\/\/github.com\/baidu-research\/baidu-allreduce."},{"key":"e_1_2_1_43_1","unstructured":"Microsoft Research. 2017. The Microsoft Cognitive Toolkit. Retrieved from https:\/\/docs.microsoft.com\/en-us\/cognitive-toolkit\/.  Microsoft Research. 2017. The Microsoft Cognitive Toolkit. Retrieved from https:\/\/docs.microsoft.com\/en-us\/cognitive-toolkit\/."},{"key":"e_1_2_1_44_1","unstructured":"Karl Rupp. 2018. Microprocessor Trend Data. Retrieved from https:\/\/github.com\/karlrupp\/microprocessor-trend-data.  Karl Rupp. 2018. Microprocessor Trend Data. Retrieved from https:\/\/github.com\/karlrupp\/microprocessor-trend-data."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_2_1_46_1","volume-title":"Horovod: Fast and easy distributed deep learning in TensorFlow. arXiv preprint arXiv:1802.05799","author":"Sergeev Alexander","year":"2018"},{"key":"e_1_2_1_47_1","unstructured":"Facebook Open Source. {n.d.}. Caffe2 A New Lightweight Modular and Scalable Deep Learning Framework. Retrieved from https:\/\/caffe2.ai.  Facebook Open Source. {n.d.}. Caffe2 A New Lightweight Modular and Scalable Deep Learning Framework. Retrieved from https:\/\/caffe2.ai."},{"key":"e_1_2_1_48_1","unstructured":"TensorFlow. {n.d.}. How To Compile and Use MPI-Enabled TensorFlow. Retrieved from https:\/\/github.com\/tensorflow\/tensorflow\/tree\/master\/tensorflow\/contrib\/mpi.  TensorFlow. {n.d.}. How To Compile and Use MPI-Enabled TensorFlow. Retrieved from https:\/\/github.com\/tensorflow\/tensorflow\/tree\/master\/tensorflow\/contrib\/mpi."},{"key":"e_1_2_1_49_1","volume-title":"Proceedings of the IEEE\/ACM Conference on Supercomputing (SC\u201998)","author":"Thakur R."},{"key":"e_1_2_1_50_1","volume-title":"Proceedings of the 7th Symposium on the Frontiers of Massively Parallel Computation","author":"Thakur R."},{"key":"e_1_2_1_52_1","volume-title":"MVAPICH: MPI over InfiniBand, 10GigE\/iWARP and RoCE.","author":"The Ohio State University","year":"2014"},{"key":"e_1_2_1_53_1","volume-title":"Theano: A Python framework for fast computation of mathematical expressions. arXiv e-prints abs\/1605.02688 (May","author":"Team Theano Development","year":"2016"},{"key":"e_1_2_1_54_1","volume-title":"Proceedings of the Workshop on Machine Learning Systems (LearningSys) in the 29th Annual Conference on Neural Information Processing Systems (NIPS\u201915)","volume":"5","author":"Tokui Seiya","year":"2015"},{"key":"e_1_2_1_55_1","volume-title":"Distributed TensorFlow with MPI. CoRR abs\/1603.02339","author":"Vishnu Abhinav","year":"2016"},{"key":"e_1_2_1_56_1","volume-title":"The SGI\u00ae AltixTM 3000 global shared-memory architecture. Silicon Graphics","author":"Woodacre Michael","year":"2005"},{"key":"e_1_2_1_57_1","unstructured":"Joe Yaworski. 2017. Intel Omni-Path Architecture Enables Deep Learning Training on HPC. Retrieved from https:\/\/itpeernetwork.intel.com\/intel-omni-path-deep-learning-training\/.  Joe Yaworski. 2017. Intel Omni-Path Architecture Enables Deep Learning Training on HPC. Retrieved from https:\/\/itpeernetwork.intel.com\/intel-omni-path-deep-learning-training\/."},{"key":"e_1_2_1_58_1","volume-title":"Image classification at supercomputer scale. arXiv preprint arXiv:1811.06992","author":"Ying Chris","year":"2018"},{"key":"e_1_2_1_59_1","volume-title":"Scaling SGD batch size to 32k for ImageNet training. arXiv preprint arXiv:1708.03888","author":"You Yang","year":"2017"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3225058.3225069"},{"key":"e_1_2_1_61_1","volume-title":"Wide residual Networks. arXiv preprint arXiv:1605.07146","author":"Zagoruyko Sergey","year":"2016"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2014.2319813"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/MASCOTS.2018.00023"}],"container-title":["ACM Transactions on Parallel Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3331526","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3331526","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:13:38Z","timestamp":1750202018000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3331526"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,6,30]]},"references-count":61,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2019,6,30]]}},"alternative-id":["10.1145\/3331526"],"URL":"https:\/\/doi.org\/10.1145\/3331526","relation":{},"ISSN":["2329-4949","2329-4957"],"issn-type":[{"value":"2329-4949","type":"print"},{"value":"2329-4957","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,6,30]]},"assertion":[{"value":"2018-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-07-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}