{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:33:05Z","timestamp":1750221185041,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":48,"publisher":"ACM","license":[{"start":{"date-parts":[[2018,8,13]],"date-time":"2018-08-13T00:00:00Z","timestamp":1534118400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2018,8,13]]},"DOI":"10.1145\/3225058.3225077","type":"proceedings-article","created":{"date-parts":[[2018,8,8]],"date-time":"2018-08-08T19:13:06Z","timestamp":1533755586000},"page":"1-10","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["GLP4NN"],"prefix":"10.1145","author":[{"given":"Hao","family":"Fu","sequence":"first","affiliation":[{"name":"Tianjin University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shanjiang","family":"Tang","sequence":"additional","affiliation":[{"name":"Tianjin University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bingsheng","family":"He","sequence":"additional","affiliation":[{"name":"National University of Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ce","family":"Yu","sequence":"additional","affiliation":[{"name":"Tianjin University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jizhou","family":"Sun","sequence":"additional","affiliation":[{"name":"Tianjin University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,8,13]]},"reference":[{"key":"e_1_3_2_1_1_1","first-page":"265","article-title":"TensorFlow: A System for Large-Scale Machine Learning","volume":"16","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi , Paul Barham , Jianmin Chen , Zhifeng Chen , Andy Davis , Jeffrey Dean , Matthieu Devin , Sanjay Ghemawat , Geoffrey Irving , Michael Isard , 2016 . TensorFlow: A System for Large-Scale Machine Learning .. In OSDI , Vol. 16. 265 -- 283 . Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. 2016. TensorFlow: A System for Large-Scale Machine Learning.. In OSDI, Vol. 16. 265--283.","journal-title":"OSDI"},{"key":"e_1_3_2_1_2_1","volume-title":"NIPS 2011, BigLearning Workshop","volume":"3","author":"Bergstra James","year":"2011","unstructured":"James Bergstra , Fr\u00e9d\u00e9ric Bastien , Olivier Breuleux , Pascal Lamblin , Razvan Pascanu , Olivier Delalleau , Guillaume Desjardins , David Warde-Farley , Ian Goodfellow , Arnaud Bergeron , 2011 . Theano: Deep learning on gpus with python . In NIPS 2011, BigLearning Workshop , Granada, Spain , Vol. 3 . Citeseer. James Bergstra, Fr\u00e9d\u00e9ric Bastien, Olivier Breuleux, Pascal Lamblin, Razvan Pascanu, Olivier Delalleau, Guillaume Desjardins, David Warde-Farley, Ian Goodfellow, Arnaud Bergeron, et al. 2011. Theano: Deep learning on gpus with python. In NIPS 2011, BigLearning Workshop, Granada, Spain, Vol. 3. Citeseer."},{"key":"e_1_3_2_1_3_1","volume-title":"Revisiting Distributed Synchronous SGD. In International Conference on Learning Representations Workshop Track. https:\/\/arxiv.org\/abs\/1604","author":"Chen Jianmin","year":"2016","unstructured":"Jianmin Chen , Rajat Monga , Samy Bengio , and Rafal Jozefowicz . 2016 . Revisiting Distributed Synchronous SGD. In International Conference on Learning Representations Workshop Track. https:\/\/arxiv.org\/abs\/1604 .00981 Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz. 2016. Revisiting Distributed Synchronous SGD. In International Conference on Learning Representations Workshop Track. https:\/\/arxiv.org\/abs\/1604.00981"},{"key":"e_1_3_2_1_4_1","volume-title":"MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. arXiv preprint arXiv:1512.01274","author":"Chen Tianqi","year":"2015","unstructured":"Tianqi Chen , Mu Li , Yutian Yi , Min Lin , Naiyan Wang , Minjie Wang , Tianjun Xiao , Bing Xu , Chiyuan Zhang , and Zheng Zhang . 2015. MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. arXiv preprint arXiv:1512.01274 ( 2015 ). Tianqi Chen, Mu Li, Yutian Yi, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang. 2015. MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. arXiv preprint arXiv:1512.01274 (2015)."},{"key":"e_1_3_2_1_5_1","volume-title":"cuDNN: Efficient Primitives for Deep Learning. arXiv preprint arXiv:1410.0759","author":"Chetlur Sharan","year":"2014","unstructured":"Sharan Chetlur , Cliff Woolley , Philippe Vandermersch , Jonathan Cohen , John Tran , Bryan Catanzaro , and Evan Shelhamer . 2014. cuDNN: Efficient Primitives for Deep Learning. arXiv preprint arXiv:1410.0759 ( 2014 ). Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014. cuDNN: Efficient Primitives for Deep Learning. arXiv preprint arXiv:1410.0759 (2014)."},{"key":"e_1_3_2_1_6_1","first-page":"22","article-title":"Solving the Straggler Problem with Bounded Staleness","volume":"13","author":"Cipar James","year":"2013","unstructured":"James Cipar , Qirong Ho , Jin Kyu Kim , Seunghak Lee , Gregory R. Ganger , Garth Gibson , Kimberly Keeton , and Eric P. Xing . 2013 . Solving the Straggler Problem with Bounded Staleness . In HotOS , Vol. 13. 22 -- 22 . James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Gregory R. Ganger, Garth Gibson, Kimberly Keeton, and Eric P. Xing. 2013. Solving the Straggler Problem with Bounded Staleness. In HotOS, Vol. 13. 22--22.","journal-title":"HotOS"},{"key":"e_1_3_2_1_7_1","volume-title":"NIPS Workshop.","author":"Collobert Ronan","year":"2011","unstructured":"Ronan Collobert , Koray Kavukcuoglu , and Cl\u00e9ment Farabet . 2011 . Torch7: A matlab-like environment for machine learning. In BigLearn , NIPS Workshop. Ronan Collobert, Koray Kavukcuoglu, and Cl\u00e9ment Farabet. 2011. Torch7: A matlab-like environment for machine learning. In BigLearn, NIPS Workshop."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/E17-1104"},{"volume-title":"2014 USENIX Annual Technical Conference (USENIX ATC 14)","author":"Cui Henggang","key":"e_1_3_2_1_9_1","unstructured":"Henggang Cui , James Cipar , Qirong Ho , Jin Kyu Kim , Seunghak Lee , Abhimanu Kumar , Jinliang Wei , Wei Dai , Gregory R. Ganger , Phillip B. Gibbons , Garth A. Gibson , and Eric P. Xing . 2014. Exploiting Bounded Staleness to Speed Up Big Data Analytics . In 2014 USENIX Annual Technical Conference (USENIX ATC 14) . USENIX Association, 37--48. Henggang Cui, James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Abhimanu Kumar, Jinliang Wei, Wei Dai, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, and Eric P. Xing. 2014. Exploiting Bounded Staleness to Speed Up Big Data Analytics. In 2014 USENIX Annual Technical Conference (USENIX ATC 14). USENIX Association, 37--48."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2901318.2901323"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_3_2_1_12_1","volume-title":"Xing","author":"Dai Wei","year":"2015","unstructured":"Wei Dai , Abhimanu Kumar , Jinliang Wei , Qirong Ho , Garth A. Gibson , and Eric P . Xing . 2015 . High-Performance Distributed ML at Scale through Parameter Server Consistency Models. In AAAI. 79--87. Wei Dai, Abhimanu Kumar, Jinliang Wei, Qirong Ho, Garth A. Gibson, and Eric P. Xing. 2015. High-Performance Distributed ML at Scale through Parameter Server Consistency Models. In AAAI. 79--87."},{"key":"e_1_3_2_1_13_1","volume-title":"Ng","author":"Dean Jeffrey","year":"2012","unstructured":"Jeffrey Dean , Greg Corrado , Rajat Monga , Kai Chen , Matthieu Devin , Mark Mao , Mar\u0107aurelio Ranzato , Andrew Senior , Paul Tucker , Ke Yang , Quoc V. Le , and Andrew Y . Ng . 2012 . Large scale distributed deep networks. In Advances in Neural Information Processing Systems. Curran Associates, Inc ., 1223--1231. Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Mar\u0107aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, Quoc V. Le, and Andrew Y. Ng. 2012. Large scale distributed deep networks. In Advances in Neural Information Processing Systems. Curran Associates, Inc., 1223--1231."},{"volume-title":"Retrieved","year":"2017","key":"e_1_3_2_1_14_1","unstructured":"Facebook. 2017 . Caffe2 . Retrieved March 31, 2018 from https:\/\/caffe2.ai\/ Facebook. 2017. Caffe2. Retrieved March 31, 2018 from https:\/\/caffe2.ai\/"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654902"},{"key":"e_1_3_2_1_16_1","volume-title":"Phillip B. Gibbons, Garth A. Gibson, Greg Ganger, and Eric P. Xing.","author":"Ho Qirong","year":"2013","unstructured":"Qirong Ho , James Cipar , Henggang Cui , Seunghak Lee , Jin Kyu Kim , Phillip B. Gibbons, Garth A. Gibson, Greg Ganger, and Eric P. Xing. 2013 . More effective distributed ml via a stale synchronous parallel parameter server. In Advances in Neural Information Processing Systems. Curran Associates, Inc ., 1223--1231. Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B. Gibbons, Garth A. Gibson, Greg Ganger, and Eric P. Xing. 2013. More effective distributed ml via a stale synchronous parallel parameter server. In Advances in Neural Information Processing Systems. Curran Associates, Inc., 1223--1231."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654889"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2901318.2901331"},{"key":"e_1_3_2_1_19_1","volume-title":"One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv:1404.5997","author":"Krizhevsky Alex","year":"2014","unstructured":"Alex Krizhevsky . 2014. One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv:1404.5997 ( 2014 ). Alex Krizhevsky. 2014. One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv:1404.5997 (2014)."},{"key":"e_1_3_2_1_20_1","unstructured":"Alex Krizhevsky and Geoffrey Hinton. 2009. Learning multiple layers of features from tiny images. (2009).  Alex Krizhevsky and Geoffrey Hinton. 2009. Learning multiple layers of features from tiny images. (2009)."},{"key":"e_1_3_2_1_21_1","volume-title":"Proceedings of the 25th International Conference on Neural Information Processing Systems -","volume":"1","author":"Krizhevsky Alex","unstructured":"Alex Krizhevsky , Ilya Sutskever , and Geoffrey E. Hinton . 2012. ImageNet Classification with Deep Convolutional Neural Networks . In Proceedings of the 25th International Conference on Neural Information Processing Systems - Volume 1 (NIPS'12). Curran Associates Inc., USA, 1097--1105. http:\/\/dl.acm.org\/citation.cfm?id=2999134.2999257 Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. ImageNet Classification with Deep Convolutional Neural Networks. In Proceedings of the 25th International Conference on Neural Information Processing Systems - Volume 1 (NIPS'12). Curran Associates Inc., USA, 1097--1105. http:\/\/dl.acm.org\/citation.cfm?id=2999134.2999257"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.435"},{"key":"e_1_3_2_1_23_1","volume-title":"SC16: International Conference for. 633--644","author":"Li Chao","year":"2016","unstructured":"Chao Li , Yi Yang , Min Feng , Srimat Chakradhar , and Huiyang Zhou . 2016 . Optimizing memory efficiency for deep convolutional neural networks on GPUs. In High Performance Computing, Networking, Storage and Analysis , SC16: International Conference for. 633--644 . Chao Li, Yi Yang, Min Feng, Srimat Chakradhar, and Huiyang Zhou. 2016. Optimizing memory efficiency for deep convolutional neural networks on GPUs. In High Performance Computing, Networking, Storage and Analysis, SC16: International Conference for. 633--644."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.5555\/2685048.2685095"},{"key":"e_1_3_2_1_25_1","volume-title":"Fast training of convolutional networks through ffts. arXiv preprint arXiv:1312.5851","author":"Mathieu Michael","year":"2013","unstructured":"Michael Mathieu , Mikael Henaff , and Yann LeCun . 2013. Fast training of convolutional networks through ffts. arXiv preprint arXiv:1312.5851 ( 2013 ). Michael Mathieu, Mikael Henaff, and Yann LeCun. 2013. Fast training of convolutional networks through ffts. arXiv preprint arXiv:1312.5851 (2013)."},{"key":"e_1_3_2_1_26_1","unstructured":"Chen Meng Minmin Sun Jun Yang Minghui Qiu and Yang Gu. 2017. Training Deeper Models by GPU Memory Optimization on TensorFlow. (2017).  Chen Meng Minmin Sun Jun Yang Minghui Qiu and Yang Gu. 2017. Training Deeper Models by GPU Memory Optimization on TensorFlow. (2017)."},{"key":"e_1_3_2_1_27_1","volume-title":"Retrieved","author":"NVIDIA.","year":"2008","unstructured":"NVIDIA. 2008 . NVIDIA Visual Profiler . Retrieved March 31, 2018 from https:\/\/developer.nvidia.com\/nvidia-visual-profiler NVIDIA. 2008. NVIDIA Visual Profiler. Retrieved March 31, 2018 from https:\/\/developer.nvidia.com\/nvidia-visual-profiler"},{"key":"e_1_3_2_1_28_1","volume-title":"Retrieved","author":"NVIDIA.","year":"2015","unstructured":"NVIDIA. 2015 . DIGITS . Retrieved March 31, 2018 from https:\/\/developer.nvidia.com\/digits NVIDIA. 2015. DIGITS. Retrieved March 31, 2018 from https:\/\/developer.nvidia.com\/digits"},{"key":"e_1_3_2_1_29_1","volume-title":"Retrieved","author":"NVIDIA.","year":"2016","unstructured":"NVIDIA. 2016 . NVIDIA Collective Communications Library . Retrieved March 31, 2018 from https:\/\/developer.nvidia.com\/nccl NVIDIA. 2016. NVIDIA Collective Communications Library. Retrieved March 31, 2018 from https:\/\/developer.nvidia.com\/nccl"},{"key":"e_1_3_2_1_30_1","volume-title":"Retrieved","author":"NVIDIA.","year":"2017","unstructured":"NVIDIA. 2017 . cuBLAS Library . Retrieved March 31, 2018 from https:\/\/developer.nvidia.com\/cublas NVIDIA. 2017. cuBLAS Library. Retrieved March 31, 2018 from https:\/\/developer.nvidia.com\/cublas"},{"key":"e_1_3_2_1_31_1","volume-title":"Retrieved","author":"NVIDIA.","year":"2017","unstructured":"NVIDIA. 2017 . NVIDIA CUDA Profiling Tools Interface . Retrieved March 31, 2018 from https:\/\/developer.nvidia.com\/cuda-profiling-tools-interface NVIDIA. 2017. NVIDIA CUDA Profiling Tools Interface. Retrieved March 31, 2018 from https:\/\/developer.nvidia.com\/cuda-profiling-tools-interface"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2915224"},{"volume-title":"Microarchitecture (MICRO), 2016 49th Annual IEEE\/ACM International Symposium on. 1--13","author":"Rhu Minsoo","key":"e_1_3_2_1_33_1","unstructured":"Minsoo Rhu , Natalia Gimelshein , Jason Clemons , Arslan Zulfiqar , and Stephen W. Keckler . 2016. vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design . In Microarchitecture (MICRO), 2016 49th Annual IEEE\/ACM International Symposium on. 1--13 . Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, and Stephen W. Keckler. 2016. vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design. In Microarchitecture (MICRO), 2016 49th Annual IEEE\/ACM International Symposium on. 1--13."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2145816.2145819"},{"key":"e_1_3_2_1_36_1","volume-title":"Going Deeper With Convolutions. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Szegedy Christian","year":"2015","unstructured":"Christian Szegedy , Wei Liu , Yangqing Jia , Pierre Sermanet , Scott Reed , Dragomir Anguelov , Dumitru Erhan , Vincent Vanhoucke , and Andrew Rabinovich . 2015 . Going Deeper With Convolutions. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015. Going Deeper With Convolutions. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2851141.2851158"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.5555\/3014904.3015005"},{"volume-title":"Retrieved","year":"2007","key":"e_1_3_2_1_39_1","unstructured":"Vampir.eu. 2007 . Vampir . Retrieved March 31, 2018 from https:\/\/www.vampir.eu\/ Vampir.eu. 2007. Vampir. Retrieved March 31, 2018 from https:\/\/www.vampir.eu\/"},{"key":"e_1_3_2_1_40_1","volume-title":"Fast convolutional nets with fbfft: A GPU performance evaluation. arXiv preprint arXiv:1412.7580","author":"Vasilache Nicolas","year":"2014","unstructured":"Nicolas Vasilache , Jeff Johnson , Michael Mathieu , Soumith Chintala , Serkan Piantino , and Yann LeCun . 2014. Fast convolutional nets with fbfft: A GPU performance evaluation. arXiv preprint arXiv:1412.7580 ( 2014 ). Nicolas Vasilache, Jeff Johnson, Michael Mathieu, Soumith Chintala, Serkan Piantino, and Yann LeCun. 2014. Fast convolutional nets with fbfft: A GPU performance evaluation. arXiv preprint arXiv:1412.7580 (2014)."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2733373.2806232"},{"key":"e_1_3_2_1_42_1","volume-title":"Xing","author":"Wei Jinliang","year":"2013","unstructured":"Jinliang Wei , Wei Dai , Abhimanu Kumar , Xun Zheng , Qirong Ho , and Eric P . Xing . 2013 . Consistent bounded-asynchronous parameter servers for distributed ML. arXiv preprint arXiv:1312.7869 (2013). Jinliang Wei, Wei Dai, Abhimanu Kumar, Xun Zheng, Qirong Ho, and Eric P. Xing. 2013. Consistent bounded-asynchronous parameter servers for distributed ML. arXiv preprint arXiv:1312.7869 (2013)."},{"key":"e_1_3_2_1_43_1","volume-title":"Efficient Gradient Boosted Decision Tree Training on GPUs. In Parallel and Distributed Processing Symposium (IPDPS)","author":"Wen Zeyi","year":"2018","unstructured":"Zeyi Wen , Bingsheng He , Kotagiri Ramamohanarao , Shengliang Lu , and Jiashuai Shi . 2018 . Efficient Gradient Boosted Decision Tree Training on GPUs. In Parallel and Distributed Processing Symposium (IPDPS) , 2018 IEEE International. Vancouver, British Columbia, Canada. Zeyi Wen, Bingsheng He, Kotagiri Ramamohanarao, Shengliang Lu, and Jiashuai Shi. 2018. Efficient Gradient Boosted Decision Tree Training on GPUs. In Parallel and Distributed Processing Symposium (IPDPS), 2018 IEEE International. Vancouver, British Columbia, Canada."},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654931"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/TBDATA.2015.2472014"},{"volume-title":"Advances in Neural Information Processing Systems. Curran associates","author":"Xu Li","key":"e_1_3_2_1_46_1","unstructured":"Li Xu , Jimmy SJ. Ren , Ce Liu , and Jiaya Jia . 2014. Deep Convolutional Neural Network for Image Deconvolution . In Advances in Neural Information Processing Systems. Curran associates , Inc ., 1790--1798. Li Xu, Jimmy SJ. Ren, Ce Liu, and Jiaya Jia. 2014. Deep Convolutional Neural Network for Image Deconvolution. In Advances in Neural Information Processing Systems. Curran associates, Inc., 1790--1798."},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2013.257"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2017.06.002"}],"event":{"name":"ICPP 2018: 47th International Conference on Parallel Processing","sponsor":["University of Oregon University of Oregon"],"location":"Eugene OR USA","acronym":"ICPP 2018"},"container-title":["Proceedings of the 47th International Conference on Parallel Processing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3225058.3225077","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3225058.3225077","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T01:39:06Z","timestamp":1750210746000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3225058.3225077"}},"subtitle":["A Convergence-invariant and Network-agnostic Light-weight Parallelization Framework for Deep Neural Networks on Modern GPUs"],"short-title":[],"issued":{"date-parts":[[2018,8,13]]},"references-count":48,"alternative-id":["10.1145\/3225058.3225077","10.1145\/3225058"],"URL":"https:\/\/doi.org\/10.1145\/3225058.3225077","relation":{},"subject":[],"published":{"date-parts":[[2018,8,13]]},"assertion":[{"value":"2018-08-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}