{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,24]],"date-time":"2025-08-24T01:33:20Z","timestamp":1755999200661},"reference-count":37,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"8","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Inf. &amp; Syst."],"published-print":{"date-parts":[[2021,8,1]]},"DOI":"10.1587\/transinf.2020edp7201","type":"journal-article","created":{"date-parts":[[2021,7,31]],"date-time":"2021-07-31T22:16:10Z","timestamp":1627769770000},"page":"1332-1339","source":"Crossref","is-referenced-by-count":7,"title":["Hybrid Electrical\/Optical Switch Architectures for Training Distributed Deep Learning in Large-Scale"],"prefix":"10.1587","volume":"E104.D","author":[{"given":"Thao-Nguyen","family":"TRUONG","sequence":"first","affiliation":[{"name":"National Institute of Advanced Industrial Science and Technology (AIST)"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ryousei","family":"TAKANO","sequence":"additional","affiliation":[{"name":"National Institute of Advanced Industrial Science and Technology (AIST)"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","unstructured":"[1] T. Ben-Nun and T. Hoefler, \u201cDemystifying parallel and distributed deep learning: An in-depth concurrency analysis,\u201d CoRR, vol.abs\/1802.09941, 2018."},{"key":"2","doi-asserted-by":"crossref","unstructured":"[2] Y. You, Z. Zhang, C.-J. Hsieh, J. Demmel, and K. Keutzer, \u201cImageNet Training in Minutes,\u201d Proceedings of the 47th International Conference on Parallel Processing (ICPP 2018), New York, NY, USA, pp.1:1-1:10, ACM, 2018. 10.1145\/3225058.3225069","DOI":"10.1145\/3225058.3225069"},{"key":"3","doi-asserted-by":"crossref","unstructured":"[3] Y. You, A. Bulu\u00e7, and J. Demmel, \u201cScaling deep learning on GPU and knights landing clusters,\u201d Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pp.1-9, ACM, 2017. 10.1145\/3126908.3126912","DOI":"10.1145\/3126908.3126912"},{"key":"4","unstructured":"[4] K. Simonyan and A. Zisserman, \u201cVery deep convolutional networks for large-scale image recognition,\u201d arXiv preprint arXiv:1409.1556, 2014."},{"key":"5","unstructured":"[5] Y. Huang, Y. Cheng, D. Chen, H. Lee, J. Ngiam, Q.V. Le, and Z. Chen, \u201cGPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism,\u201d CoRR, vol.abs\/1811.06965, 2018."},{"key":"6","unstructured":"[6] M. Shoeybi et al., \u201cMegatron-lm: Training multi-billion parameter language models using model parallelism,\u201d ArXiv, vol.abs\/ 1909.08053, 2019."},{"key":"7","unstructured":"[7] \u201cTuring-NLG: A 17-billion-parameter language model by Microsoft.\u201d https:\/\/www.microsoft.com\/en-us\/research\/blog\/turing-nlg-a-17-billion-parameter-language-model-by-microsoft\/ [01 April 2020]."},{"key":"8","doi-asserted-by":"crossref","unstructured":"[8] P. Patarasuk and X. Yuan, \u201cBandwidth efficient all-reduce operation on tree topologies,\u201d IEEE International Parallel and Distributed Processing Symposium, IPDPS 2007, pp.1-8, IEEE, 2007. 10.1109\/ipdps.2007.370405","DOI":"10.1109\/IPDPS.2007.370405"},{"key":"9","doi-asserted-by":"crossref","unstructured":"[9] M. Bayatpour, J.M. Hashmi, S. Chakraborty, H. Subramoni, P. Kousha, and D.K. Panda, \u201cSalar: Scalable and adaptive designs for large message reduction collectives,\u201d IEEE International Conference on Cluster Computing, CLUSTER 2018, Belfast, UK, pp.12-23, 2018. 10.1109\/cluster.2018.00014","DOI":"10.1109\/CLUSTER.2018.00014"},{"key":"10","doi-asserted-by":"crossref","unstructured":"[10] T.T. Nguyen, M. Wahib, and R. Takano, \u201cHierarchical distributed-memory multi-leader mpi-allreduce for deep learning workloads,\u201d Sixth International Symposium on Computing and Networking, CANDAR Workshops 2018, Takayama, Japan, pp.216-222, IEEE, 2018. 10.1109\/candarw.2018.00048","DOI":"10.1109\/CANDARW.2018.00048"},{"key":"11","unstructured":"[11] F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu, \u201c1-bit stochastic gradient descent and application to data-parallel distributed training of speech DNNs,\u201d Fifteenth Annual Conference of the International Speech Communication Association (INTERSPEECH 2014), pp.1058-1062, Sept. 2014."},{"key":"12","unstructured":"[12] W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li, \u201cTernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning,\u201d Advances in Neural Information Processing Systems 30, ed. I. Guyon, U.V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, pp.1509-1519, Curran Associates, Inc., 2017."},{"key":"13","doi-asserted-by":"crossref","unstructured":"[13] N. Strom, \u201cScalable distributed DNN training using commodity GPU cloud computing,\u201d Sixteenth Annual Conference of the International Speech Communication Association (Interspeech 2015), 2015.","DOI":"10.21437\/Interspeech.2015-354"},{"key":"14","doi-asserted-by":"crossref","unstructured":"[14] F. Sattler, S. Wiedemann, K. M\u00fcller, and W. Samek, \u201cSparse binary compression: Towards distributed deep learning with minimal communication,\u201d CoRR, vol.abs\/1805.08768, 2018.","DOI":"10.1109\/IJCNN.2019.8852172"},{"key":"15","unstructured":"[15] M. Yamazaki, A. Kasagi, A. Tabuchi, T. Honda, M. Miwa, N. Fukumoto, T. Tabaru, A. Ike, and K. Nakashima, \u201cYet Another Accelerated SGD: ResNet-50 Training on ImageNet in 74.7 seconds,\u201d CoRR, vol.abs\/1903.12650, 2019."},{"key":"16","doi-asserted-by":"crossref","unstructured":"[16] N. Dryden, N. Maruyama, T. Moon, T. Benson, A. Yoo, M. Snir, and B. Van Essen, \u201cAluminum: An asynchronous, gpu-aware communication library optimized for large-scale training of deep neural networks on hpc systems,\u201d 2018 IEEE\/ACM Machine Learning in HPC Environments (MLHPC), pp.1-13, Nov. 2018. 10.1109\/mlhpc.2018.8638639","DOI":"10.1109\/MLHPC.2018.8638639"},{"key":"17","doi-asserted-by":"crossref","unstructured":"[17] T.T. Nguyen and R. Takano, \u201cOn the feasibility of hybrid electrical\/optical switch architecture for large-scale training of distributed deep learning,\u201d 2019 IEEE\/ACM Workshop on Photonics-Optics Technology Oriented Networking, Information and Computing Systems (PHOTONICS), pp.7-14, 2019. 10.1109\/photonics49561.2019.00007","DOI":"10.1109\/PHOTONICS49561.2019.00007"},{"key":"18","unstructured":"[18] \u201cNVIDIA Collective Communications Library (NCCL), Multi-GPU and multi-node collective communication primitives.\u201d https:\/\/developer.nvidia.com\/nccl, accessed April. 1. 2019."},{"key":"19","unstructured":"[19] \u201cBaidu Allreduce.\u201d https:\/\/github.com\/baidu-research\/baidu-allreduce, accessed April. 1. 2019."},{"key":"20","doi-asserted-by":"publisher","unstructured":"[20] P. Patarasuk and X. Yuan, \u201cBandwidth optimal all-reduce algorithms for clusters of workstations,\u201d Journal of Parallel and Distributed Computing, vol.69, no.2, pp.117-124, 2009. 10.1016\/j.jpdc.2008.09.002","DOI":"10.1016\/j.jpdc.2008.09.002"},{"key":"21","unstructured":"[21] Top 500 Supercomputer Sites. http:\/\/www.top500.org\/"},{"key":"22","unstructured":"[22] NVIDIA, \u201cWhite Paper NVIDIA DGX-1 System Architecture, The Fastest Platform for Deep Learning.\u201d https:\/\/www.azken.com\/images\/dgx1_images\/dgx1-system-architecture-whitepaper1.pdf, accessed April. 1. 2019."},{"key":"23","unstructured":"[23] A. Ishii, D.F. withEric Anderson, B. Dally, G.D. Dennison, M. Hummel, and J. Schafer, \u201cNvswitch and dgx-2 nvlink-switching chip and scale-up compute server,\u201d 2018 IEEE Hot Chips 30 Symposium (HCS), pp.1-30, Aug. 2018."},{"key":"24","unstructured":"[24] Mellanox, \u201cCS7500 InfiniBand EDR 100Gb\/s Switch System.\u201d http:\/\/www.mellanox.com\/related-docs\/prod_ib_switch_systems\/pb_cs7500.pdf, accessed April. 1. 2019."},{"key":"25","unstructured":"[25] Mellanox, \u201cSB7800 InfiniBand EDR 100Gb\/s Switch System.\u201d http:\/\/www.mellanox.com\/related-docs\/prod_ib_switch_systems\/pb_sb7800.pdf, accessed April. 1. 2019."},{"key":"26","unstructured":"[26] PLX Technology, \u201cProduct Brief: PEX 8716, PCI Express Gen 3 Switch, 16 Lanes, 4 Ports..\u201d https:\/\/docs.broadcom.com\/docs\/12351846, accessed April. 1. 2019."},{"key":"27","unstructured":"[27] A. Li, S.L. Song, J. Chen, J. Li, X. Liu, N.R. Tallent, and K.J. Barker, \u201cEvaluating Modern GPU Interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect,\u201d CoRR, vol.abs\/1903.04611, 2019."},{"key":"28","doi-asserted-by":"publisher","unstructured":"[28] A.A. Awan, K. Hamidouche, J.M. Hashmi, and D.K. Panda, \u201cS-caffe: co-designing MPI runtimes and caffe for scalable deep learning on modern GPU clusters,\u201d ACM Sigplan Notices, vol.52, no.8, pp.193-205, 2017. 10.1145\/3155284.3018769","DOI":"10.1145\/3155284.3018769"},{"key":"29","unstructured":"[29] A.R. Mamidala, G. Kollias, C. Ward, and F. Artico, \u201cMXNET-MPI: Embedding MPI parallelism in Parameter Server Task Model for scaling Deep Learning,\u201d arXiv preprint arXiv:1801.03855, 2018."},{"key":"30","doi-asserted-by":"crossref","unstructured":"[30] T.T. Nguyen, H. Matsutani, and M. Koibuchi, \u201cLow-reliable low-latency networks optimized for hpc parallel applications,\u201d 2018 IEEE 17th International Symposium on Network Computing and Applications (NCA), pp.1-10, Nov. 2018. 10.1109\/nca.2018.8548063","DOI":"10.1109\/NCA.2018.8548063"},{"key":"31","unstructured":"[31] K.J. Barker, A. Benner, R. Hoare, A. Hoisie, A.K. Jones, D.K. Kerbyson, D. Li, R. Melhem, R. Rajamony, E. Schenfeld, S. Shao, C. Stunkel, and P. Walker, \u201cOn the feasibility of optical circuit switching for high performance computing systems,\u201d SC &apos;05: Proceedings of the 2005 ACM\/IEEE Conference on Supercomputing, p.16, Nov. 2005. 10.1109\/sc.2005.48"},{"key":"32","doi-asserted-by":"publisher","unstructured":"[32] C. Kachris, K. Kanonakis, and I. Tomkos, \u201cOptical interconnection networks in data centers: recent trends and future challenges,\u201d IEEE Communications Magazine, vol.51, no.9, pp.39-45, Sept. 2013. 10.1109\/mcom.2013.6588648","DOI":"10.1109\/MCOM.2013.6588648"},{"key":"33","doi-asserted-by":"publisher","unstructured":"[33] J. Chen, Y. Gong, M. Fiorani, and S. Aleksic, \u201cOptical interconnects at the top of the rack for energy-efficient data centers,\u201d IEEE Communications Magazine, vol.53, no.8, pp.140-148, Aug. 2015. 10.1109\/mcom.2015.7180521","DOI":"10.1109\/MCOM.2015.7180521"},{"key":"34","doi-asserted-by":"publisher","unstructured":"[34] K.I. Kitayama, Y.C. Huang, Y. Yoshida, R. Takahashi, T. Segawa, S. Ibrahim, T. Nakahara, Y. Suzaki, M. Hayashitani, Y. Hasegawa, Y. Mizukoshi, and A. Hiramatsu, \u201cTorus-topology data center network based on optical packet\/agile circuit switching with intelligent flow management,\u201d J. Lightwave Technol., vol.33, no.5, pp.1063-1071, March 2015. 10.1109\/jlt.2015.2394384","DOI":"10.1109\/JLT.2015.2394384"},{"key":"35","doi-asserted-by":"crossref","unstructured":"[35] K. Sato, \u201cRealization and application of large-scale fast optical circuit switch for data center networking,\u201d 2017 European Conference on Optical Communication (ECOC), pp.1-3, 2017. 10.1109\/ecoc.2017.8346155","DOI":"10.1109\/ECOC.2017.8346155"},{"key":"36","unstructured":"[36] SimGrid: Versatile Simulation of Distributed Systems. http:\/\/simgrid.gforge.inria.fr\/, accessed April. 1. 2019."},{"key":"37","doi-asserted-by":"crossref","unstructured":"[37] K. He, X. Zhang, S. Ren, and J. Sun, \u201cDeep residual learning for image recognition,\u201d Proceedings of the IEEE conference on computer vision and pattern recognition, pp.770-778, IEEE, 2016. 10.1109\/cvpr.2016.90","DOI":"10.1109\/CVPR.2016.90"}],"container-title":["IEICE Transactions on Information and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E104.D\/8\/E104.D_2020EDP7201\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,6]],"date-time":"2023-01-06T01:38:57Z","timestamp":1672969137000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E104.D\/8\/E104.D_2020EDP7201\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,1]]},"references-count":37,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2021]]}},"URL":"https:\/\/doi.org\/10.1587\/transinf.2020edp7201","relation":{},"ISSN":["0916-8532","1745-1361"],"issn-type":[{"value":"0916-8532","type":"print"},{"value":"1745-1361","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,8,1]]},"article-number":"2020EDP7201"}}