{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,8,13]],"date-time":"2024-08-13T18:50:49Z","timestamp":1723575049006},"reference-count":77,"publisher":"Association for Computing Machinery (ACM)","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,2]]},"abstract":"<jats:p>Despite the fact that GPUs and accelerators are more efficient in deep learning (DL), commercial clouds like Facebook and Amazon now heavily use CPUs in DL computation because there are large numbers of CPUs which would otherwise sit idle during off-peak periods. Following the trend, CPU vendors have not only released high-performance many-core CPUs but also developed efficient math kernel libraries. However, current DL platforms cannot scale well to a large number of CPU cores, making many-core CPUs inefficient in DL computation. We analyze the memory access patterns of various layers and identify the root cause of the low scalability, i.e., the per-layer barriers that are implicitly imposed by current platforms which assign one single instance (i.e., one batch of input data) to a CPU. The barriers cause severe memory bandwidth contention and CPU starvation in the access-intensive layers (like activation and BN).<\/jats:p>\n          <jats:p>This paper presents a novel approach called ParaX, which boosts the performance of DL on many-core CPUs by effectively alleviating bandwidth contention and CPU starvation. Our key idea is to assign one instance to each CPU core instead of to the entire CPU, so as to remove the per-layer barriers on the executions of the many cores. ParaX designs an ultralight scheduling policy which sufficiently overlaps the access-intensive layers with the compute-intensive ones to avoid contention, and proposes a NUMA-aware gradient server mechanism for training which leverages shared memory to substantially reduce the overhead of per-iteration parameter synchronization. We have implemented ParaX on MXNet. Extensive evaluation on a two-NUMA Intel 8280 CPU shows that ParaX significantly improves the training\/inference throughput for all tested models (for image recognition and natural language processing) by 1.73X ~ 2.93X.<\/jats:p>","DOI":"10.14778\/3447689.3447692","type":"journal-article","created":{"date-parts":[[2021,4,12]],"date-time":"2021-04-12T16:20:06Z","timestamp":1618244406000},"page":"864-877","source":"Crossref","is-referenced-by-count":6,"title":["ParaX"],"prefix":"10.14778","volume":"14","author":[{"given":"Lujia","family":"Yin","sequence":"first","affiliation":[{"name":"NUDT, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yiming","family":"Zhang","sequence":"additional","affiliation":[{"name":"NUDT, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhaoning","family":"Zhang","sequence":"additional","affiliation":[{"name":"NUDT, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuxing","family":"Peng","sequence":"additional","affiliation":[{"name":"NUDT, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peng","family":"Zhao","sequence":"additional","affiliation":[{"name":"Intel Research, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,4,12]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"https:\/\/www.cs.toronto.edu\/~kriz\/cifar.html. Online","year":"2021"},{"key":"e_1_2_1_2_1","volume-title":"http:\/\/yann.lecun.com\/exdb\/mnist\/. Online","year":"2021"},{"key":"e_1_2_1_3_1","volume-title":"https:\/\/www.nvidia.com\/en-us\/data-center\/tesla-p100\/. Online","year":"2021"},{"key":"e_1_2_1_4_1","volume-title":"https:\/\/software.intel.com\/en-us\/articles\/intel-avx-512-instructions. Online","year":"2021"},{"key":"e_1_2_1_5_1","volume-title":"https:\/\/software.intel.com\/en-us\/articles\/performance-boosting-in-seldon. Online","year":"2021"},{"key":"e_1_2_1_6_1","volume-title":"https:\/\/www.nvidia.com\/en-us\/design-visualization\/quadro\/rtx-8000\/. Online","year":"2021"},{"key":"e_1_2_1_7_1","volume-title":"https:\/\/blog.udacity.com\/2020\/08\/machine-learning-for-big-data.html. Online","year":"2021"},{"key":"e_1_2_1_8_1","volume-title":"https:\/\/www.kaggle.com\/idevji1\/sherlock-holmes-stories. Online","year":"2021"},{"key":"e_1_2_1_9_1","volume-title":"https:\/\/software.intel.com\/en-us\/vtune. Online","year":"2021"},{"key":"e_1_2_1_10_1","volume-title":"https:\/\/github.com\/nicexlab\/parax-source. Online","year":"2021"},{"key":"e_1_2_1_11_1","volume-title":"https:\/\/pytorch.org\/. Online","year":"2021"},{"key":"e_1_2_1_12_1","volume-title":"docs.aws.amazon.com\/dlami\/latest\/devguide\/deep-learning-containers-eks-tutorials-cpu-training.html. Online","year":"2021"},{"key":"e_1_2_1_13_1","volume-title":"https:\/\/github.com\/oneapi-src\/oneDNN. Online","year":"2021"},{"key":"e_1_2_1_14_1","volume-title":"https:\/\/developer.nvidia.com\/about-cuda. Online","year":"2021"},{"key":"e_1_2_1_15_1","volume-title":"https:\/\/docs.nvidia.com\/deeplearning\/sdk\/cudnn-developer-guide\/index.html. Online","year":"2021"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/3026877.3026899"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.5555\/355074"},{"key":"e_1_2_1_18_1","volume-title":"Revisiting Distributed Synchronous SGD. CoRR abs\/1604.00981","author":"Chen Jianmin","year":"2016"},{"key":"e_1_2_1_19_1","volume-title":"Revisiting distributed synchronous SGD. arXiv preprint arXiv:1604.00981","author":"Chen Jianmin","year":"2016"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3093337.3037700"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2980024.2872368"},{"key":"e_1_2_1_22_1","volume-title":"Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems. arXiv preprint arXiv:1512.01274","author":"Chen Tianqi","year":"2015"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/3291168.3291211"},{"key":"e_1_2_1_24_1","volume-title":"cudnn: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759","author":"Chetlur Sharan","year":"2014"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463676.2465338"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969442.2969588"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2959100.2959190"},{"key":"e_1_2_1_28_1","volume-title":"Distributed deep learning using synchronous stochastic gradient descent. arXiv preprint arXiv:1602.06709","author":"Das Dipankar","year":"2016"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/2999134.2999271"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/W14-3309"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2019.2910506"},{"key":"e_1_2_1_33_1","volume-title":"High-Performance Deep Learning via a Single Building Block. arXiv preprint arXiv:1906.06440","author":"Georganas Evangelos","year":"2019"},{"key":"e_1_2_1_34_1","volume-title":"large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677","author":"Goyal Priya","year":"2017"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00084"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00047"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2018.00059"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3038912.3052569"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/963770.963772"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2008.209"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_2_1_44_1","volume-title":"Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861","author":"Howard Andrew G","year":"2017"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00083"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.5555\/3045118.3045167"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00286"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654889"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00070"},{"key":"e_1_2_1_50_1","volume-title":"On large-batch training for deep learning: Generalization gap and sharp minima. arXiv preprint arXiv:1609.04836","author":"Keskar Nitish Shirish","year":"2016"},{"key":"e_1_2_1_51_1","first-page":"105","article-title":"Operating system for a non-uniform memory access multiprocessor system","volume":"6","author":"Kimmel Jeffrey S","year":"2000","journal-title":"US Patent"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2009.263"},{"key":"e_1_2_1_53_1","volume-title":"One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv:1404.5997","author":"Krizhevsky Alex","year":"2014"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.5555\/2999134.2999257"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3054775"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358284"},{"key":"e_1_2_1_57_1","volume-title":"Ternary weight networks. arXiv preprint arXiv:1605.04711","author":"Li Fengfu","year":"2016"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.5555\/2685048.2685095"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/2623330.2623612"},{"key":"e_1_2_1_60_1","volume-title":"enhanced Smith-Waterman protein database search on CUDA-enabled GPUs based on SIMT and virtualized SIMD abstractions. BMC research notes 3, 1","author":"Liu Yongchao","year":"2010"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.5555\/311445"},{"key":"e_1_2_1_62_1","doi-asserted-by":"crossref","unstructured":"Tom\u00e1\u0161 Mikolov Martin Karafi\u00e1t Luk\u00e1\u0161 Burget Jan \u010cernock\u1ef3 and Sanjeev Khudanpur. 2010. Recurrent neural network based language model. In Eleventh annual conference of the international speech communication association.  Tom\u00e1\u0161 Mikolov Martin Karafi\u00e1t Luk\u00e1\u0161 Burget Jan \u010cernock\u1ef3 and Sanjeev Khudanpur. 2010. Recurrent neural network based language model. In Eleventh annual conference of the international speech communication association.","DOI":"10.21437\/Interspeech.2010-343"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/2508148.2485927"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCAS.2018.8351352"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00647"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_32"},{"key":"e_1_2_1_67_1","volume-title":"Glow: Graph lowering compiler techniques for neural networks. arXiv preprint arXiv:1805.00907","author":"Rotem Nadav","year":"2018"},{"key":"e_1_2_1_68_1","volume-title":"Horovod: fast and easy distributed deep learning in TensorFlow. CoRR abs\/1802.05799","author":"Sergeev Alexander","year":"2018"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_2_1_70_1","volume-title":"Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions. arXiv preprint arXiv:1802.04730","author":"Vasilache Nicolas","year":"2018"},{"key":"e_1_2_1_71_1","volume-title":"Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation. CoRR abs\/1609.08144","author":"Wu Yonghui","year":"2016"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/216585.216588"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/3225058.3225069"},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1145\/1755913.1755940"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732977.2733001"},{"key":"e_1_2_1_76_1","volume-title":"2019 USENIX Conference on Operational Machine Learning, OpML 2019","author":"Zhang Minjia","year":"2019"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305890.3306107"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3447689.3447692","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:17:07Z","timestamp":1672226227000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3447689.3447692"}},"subtitle":["boosting deep learning for big data analytics on many-core CPUs"],"short-title":[],"issued":{"date-parts":[[2021,2]]},"references-count":77,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2021,2]]}},"alternative-id":["10.14778\/3447689.3447692"],"URL":"https:\/\/doi.org\/10.14778\/3447689.3447692","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,2]]}}}