{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:27:44Z","timestamp":1750220864989,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":18,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,12,20]],"date-time":"2019-12-20T00:00:00Z","timestamp":1576800000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,12,20]]},"DOI":"10.1145\/3377170.3377245","type":"proceedings-article","created":{"date-parts":[[2020,3,20]],"date-time":"2020-03-20T11:16:22Z","timestamp":1584702982000},"page":"120-125","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Distributed Deep Neural Network Training with Important Gradient Filtering, Delayed Update and Static Filtering"],"prefix":"10.1145","author":[{"given":"Kairu","family":"Li","sequence":"first","affiliation":[{"name":"Institute of Microelectronics, Tsinghua University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongyu","family":"Wu","sequence":"additional","affiliation":[{"name":"Computer Science and Technology Tsinghua University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jia","family":"Tian","sequence":"additional","affiliation":[{"name":"Cortex Labs"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wentao","family":"Tian","sequence":"additional","affiliation":[{"name":"Cortex Labs"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zuochang","family":"Ye","sequence":"additional","affiliation":[{"name":"Institute of Microelectrionics, Tsinghua University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,3,20]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Sparse communication for distributed gradient descent. arXiv preprint arXiv:1704.05021","author":"Aji A. F.","year":"2017","unstructured":"Aji , A. F. and Heafield , K . ( 2017 ). Sparse communication for distributed gradient descent. arXiv preprint arXiv:1704.05021 . Aji, A. F. and Heafield, K. (2017). Sparse communication for distributed gradient descent. arXiv preprint arXiv:1704.05021."},{"key":"e_1_3_2_1_2_1","volume-title":"End to end speech recognition in en- glish and mandarin","author":"Amodei D.","year":"2016","unstructured":"Amodei , D. , Anubhai , R. , Battenberg , E. , Case , C. , Casper , J. , Catanzaro , B. , Chen , J. , Chrzanowski , M. , Coates , A. , Diamos , G. , ( 2016 ). End to end speech recognition in en- glish and mandarin . Amodei, D., Anubhai, R., Battenberg, E., Case, C., Casper, J., Catanzaro, B., Chen, J., Chrzanowski, M., Coates, A., Diamos, G., et al. (2016). End to end speech recognition in en- glish and mandarin."},{"key":"e_1_3_2_1_3_1","first-page":"571","volume-title":"OSDI","volume":"14","author":"Chilimbi","year":"2014","unstructured":"[ Chilimbi et al. , 2014 ] Chilimbi, T. M., Suzue, Y., Apacible, J., and Kalyanaraman, K. (2014). Project adam: Building an efficient and scalable deep learning training system . In OSDI , volume 14 , pages 571 -- 582 . [Chilimbi et al., 2014] Chilimbi, T. M., Suzue, Y., Apacible, J., and Kalyanaraman, K. (2014). Project adam: Building an efficient and scalable deep learning training system. In OSDI, volume 14, pages 571--582."},{"key":"e_1_3_2_1_4_1","first-page":"1337","volume-title":"International Conference on Machine Learning","author":"Coates A.","year":"2013","unstructured":"Coates , A. , Huval , B. , Wang , T. , Wu , D. , Catanzaro , B. , and Andrew , N . ( 2013 ). Deep learning with cots hpc systems . In International Conference on Machine Learning , pages 1337 -- 1345 . Coates, A., Huval, B., Wang, T., Wu, D., Catanzaro, B., and Andrew, N. (2013). Deep learning with cots hpc systems. In International Conference on Machine Learning, pages 1337--1345."},{"key":"e_1_3_2_1_5_1","first-page":"1223","volume-title":"Advances in neural information processing systems","author":"Dean J.","year":"2012","unstructured":"Dean , J. , Corrado , G. , Monga , R. , Chen , K. , Devin , M. , Mao , M. , Senior , A. , Tucker , P. , Yang , K. , Le , Q. V. , ( 2012 b). Large scale distributed deep networks . In Advances in neural information processing systems , pages 1223 -- 1231 . Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Senior, A., Tucker, P., Yang, K., Le, Q. V., et al. (2012b). Large scale distributed deep networks. In Advances in neural information processing systems, pages 1223--1231."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_1_7_1","volume-title":"Accurate, large mini-batch sgd: training imagenet in 1 hour. arXiv preprint arXiv:1706.02677","author":"Goyal P.","year":"2017","unstructured":"Goyal , P. , Doll\u00e1r , P. , Girshick , R. , Noordhuis , P. , Wesolowski , L. , Kyrola , A. , Tulloch , A. , Jia , Y. , and He , K . ( 2017 ). Accurate, large mini-batch sgd: training imagenet in 1 hour. arXiv preprint arXiv:1706.02677 . Goyal, P., Doll\u00e1r, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. (2017). Accurate, large mini-batch sgd: training imagenet in 1 hour. arXiv preprint arXiv:1706.02677."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_9_1","volume-title":"Quantization and training of neural networks for efficient integer-arithmetic- only inference. arXiv preprint arXiv:1712.05877","author":"Jacob B.","year":"2017","unstructured":"Jacob , B. , Kligys , S. , Chen , B. , Zhu , M. , Tang , M. , Howard , A. , Adam , H. , and Kalenichenko , D . ( 2017 ). Quantization and training of neural networks for efficient integer-arithmetic- only inference. arXiv preprint arXiv:1712.05877 . Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., and Kalenichenko, D. (2017). Quantization and training of neural networks for efficient integer-arithmetic- only inference. arXiv preprint arXiv:1712.05877."},{"key":"e_1_3_2_1_10_1","volume-title":"Highly scalable deep learning training system with mixed-precision: Training imagenet in four minutes. arXiv preprint arXiv:1807.11205","author":"Jia X.","year":"2018","unstructured":"Jia , X. , Song , S. , He , W. , Wang , Y. , Rong , H. , Zhou , F. , Xie , L. , Guo , Z. , Yang , Y. , Yu , L. , ( 2018 ). Highly scalable deep learning training system with mixed-precision: Training imagenet in four minutes. arXiv preprint arXiv:1807.11205 . Jia, X., Song, S., He, W., Wang, Y., Rong, H., Zhou, F., Xie, L., Guo, Z., Yang, Y., Yu, L., et al. (2018). Highly scalable deep learning training system with mixed-precision: Training imagenet in four minutes. arXiv preprint arXiv:1807.11205."},{"key":"e_1_3_2_1_11_1","volume-title":"How to scale distributed deep learning? arXiv preprint arXiv:1611.04581","author":"Jin P. H.","year":"2016","unstructured":"Jin , P. H. , Yuan , Q. , Iandola , F. , and Keutzer , K . ( 2016 ). How to scale distributed deep learning? arXiv preprint arXiv:1611.04581 . Jin, P. H., Yuan, Q., Iandola, F., and Keutzer, K. (2016). How to scale distributed deep learning? arXiv preprint arXiv:1611.04581."},{"key":"e_1_3_2_1_12_1","unstructured":"[Krizhevsky et al. 2012] Krizhevsky A. Sutskever I. and Hinton G. E. (2012). Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems pages 1097--1105.  [Krizhevsky et al. 2012] Krizhevsky A. Sutskever I. and Hinton G. E. (2012). Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems pages 1097--1105."},{"key":"e_1_3_2_1_13_1","unstructured":"[Lin et al. 2017] Lin Y. Han S. Mao H. Wang Y. and Dally W. J. (2017). Deep gradient compression: Reducing the communication bandwidth for distributed training. arXiv preprint arXiv:1712.01887.  [Lin et al. 2017] Lin Y. Han S. Mao H. Wang Y. and Dally W. J. (2017). Deep gradient compression: Reducing the communication bandwidth for distributed training. arXiv preprint arXiv:1712.01887."},{"key":"e_1_3_2_1_14_1","volume-title":"Modeling and evaluation of synchronous stochastic gradient descent in distributed deep learning on multiple gpus. arXiv preprint arXiv:1805.03812","author":"Shi S.","year":"2018","unstructured":"Shi , S. , Wang , Q. , Chu , X. , and Li , B . ( 2018 ). Modeling and evaluation of synchronous stochastic gradient descent in distributed deep learning on multiple gpus. arXiv preprint arXiv:1805.03812 . Shi, S., Wang, Q., Chu, X., and Li, B. (2018). Modeling and evaluation of synchronous stochastic gradient descent in distributed deep learning on multiple gpus. arXiv preprint arXiv:1805.03812."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2015-354"},{"key":"e_1_3_2_1_16_1","first-page":"1509","volume-title":"Advances in neural information processing systems","author":"Wen W.","year":"2017","unstructured":"Wen , W. , Xu , C. , Yan , F. , Wu , C. , Wang , Y. , Chen , Y. , and Li , H . ( 2017 ). Terngrad: Ternary gradients to reduce communication in distributed deep learning . In Advances in neural information processing systems , pages 1509 -- 1519 . Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H. (2017). Terngrad: Ternary gradients to reduce communication in distributed deep learning. In Advances in neural information processing systems, pages 1509--1519."},{"key":"e_1_3_2_1_17_1","volume-title":"Asynchronous stochastic gradient descent with delay compensation for distributed deep learning. CoRR, abs\/1609.08326","author":"Zheng S.","year":"2016","unstructured":"Zheng , S. , Meng , Q. , Wang , T. , Chen , W. , Yu , N. , Ma , Z. , and Liu , T . ( 2016 ). Asynchronous stochastic gradient descent with delay compensation for distributed deep learning. CoRR, abs\/1609.08326 . Zheng, S., Meng, Q., Wang, T., Chen, W., Yu, N., Ma, Z., and Liu, T. (2016). Asynchronous stochastic gradient descent with delay compensation for distributed deep learning. CoRR, abs\/1609.08326."},{"key":"e_1_3_2_1_18_1","volume-title":"Dorefa-net: Training low bitwidth convolutional neural net- works with low bitwidth gradients. arXiv preprint arXiv:1606.06160.","author":"Zhou","year":"2016","unstructured":"Zhou et al. , 2016 ] Zhou, S. , Wu, Y., Ni, Z., Zhou, X., Wen, H., and Zou, Y. ( 2016). Dorefa-net: Training low bitwidth convolutional neural net- works with low bitwidth gradients. arXiv preprint arXiv:1606.06160. Zhou et al., 2016] Zhou, S., Wu, Y., Ni, Z., Zhou, X., Wen, H., and Zou, Y. (2016). Dorefa-net: Training low bitwidth convolutional neural net- works with low bitwidth gradients. arXiv preprint arXiv:1606.06160."}],"event":{"name":"ICIT 2019: IoT and Smart City","sponsor":["Shanghai Jiao Tong University Shanghai Jiao Tong University","The Hong Kong Polytechnic The Hong Kong Polytechnic University","University of Malaya University of Malaya"],"location":"Shanghai China","acronym":"ICIT 2019"},"container-title":["Proceedings of the 2019 7th International Conference on Information Technology: IoT and Smart City"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3377170.3377245","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3377170.3377245","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:23:41Z","timestamp":1750202621000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3377170.3377245"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,12,20]]},"references-count":18,"alternative-id":["10.1145\/3377170.3377245","10.1145\/3377170"],"URL":"https:\/\/doi.org\/10.1145\/3377170.3377245","relation":{},"subject":[],"published":{"date-parts":[[2019,12,20]]},"assertion":[{"value":"2020-03-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}