{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:30:37Z","timestamp":1750221037282,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":41,"publisher":"ACM","license":[{"start":{"date-parts":[[2018,10,1]],"date-time":"2018-10-01T00:00:00Z","timestamp":1538352000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["1652328, 1718158"],"award-info":[{"award-number":["1652328, 1718158"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2018,10]]},"DOI":"10.1145\/3240302.3240422","type":"proceedings-article","created":{"date-parts":[[2019,1,4]],"date-time":"2019-01-04T13:33:56Z","timestamp":1546608836000},"page":"279-290","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Leveraging MLC STT-RAM for energy-efficient CNN training"],"prefix":"10.1145","author":[{"given":"Hengyu","family":"Zhao","sequence":"first","affiliation":[{"name":"University of California"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jishen","family":"Zhao","sequence":"additional","affiliation":[{"name":"University of California"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,10]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2016. NVIDIA cuDNN: GPU accelerated deep learning. (2016).  2016. NVIDIA cuDNN: GPU accelerated deep learning. (2016)."},{"key":"e_1_3_2_1_2_1","unstructured":"http:\/\/docs.nvidia.com\/cuda\/profiler-users-guide\/. NVIDIA Profiler user's guide. (http:\/\/docs.nvidia.com\/cuda\/profiler-users-guide\/).  http:\/\/docs.nvidia.com\/cuda\/profiler-users-guide\/. NVIDIA Profiler user's guide. (http:\/\/docs.nvidia.com\/cuda\/profiler-users-guide\/)."},{"key":"e_1_3_2_1_3_1","unstructured":"https:\/\/github.com\/amd\/OpenCL-caffe. AMD: Caffe for OpenCL. (https:\/\/github.com\/amd\/OpenCL-caffe).  https:\/\/github.com\/amd\/OpenCL-caffe. AMD: Caffe for OpenCL. (https:\/\/github.com\/amd\/OpenCL-caffe)."},{"key":"e_1_3_2_1_4_1","unstructured":"https:\/\/www.micron.com\/\/media\/documents\/products\/data-sheet\/dram\/gddr5\/8gb_gddr5x_sgram_brief.pdf. Micron GDDR5X SGRAM specification. (https:\/\/www.micron.com\/\/media\/documents\/products\/data-sheet\/dram\/gddr5\/8gb_gddr5x_sgram_brief.pdf).  https:\/\/www.micron.com\/\/media\/documents\/products\/data-sheet\/dram\/gddr5\/8gb_gddr5x_sgram_brief.pdf. Micron GDDR5X SGRAM specification. (https:\/\/www.micron.com\/\/media\/documents\/products\/data-sheet\/dram\/gddr5\/8gb_gddr5x_sgram_brief.pdf)."},{"key":"e_1_3_2_1_5_1","unstructured":"https:\/\/www.nvidia.com\/en-us\/geforce\/products\/. NVIDIA GeForce GTX 1080 Ti. (https:\/\/www.nvidia.com\/en-us\/geforce\/products\/).  https:\/\/www.nvidia.com\/en-us\/geforce\/products\/. NVIDIA GeForce GTX 1080 Ti. (https:\/\/www.nvidia.com\/en-us\/geforce\/products\/)."},{"key":"e_1_3_2_1_6_1","unstructured":"https:\/\/www.nvidia.com\/en-us\/geforce\/products\/10series\/titan-xp\/. NVIDIA TITAN Xp. (https:\/\/www.nvidia.com\/en-us\/geforce\/products\/10series\/titan-xp\/).  https:\/\/www.nvidia.com\/en-us\/geforce\/products\/10series\/titan-xp\/. NVIDIA TITAN Xp. (https:\/\/www.nvidia.com\/en-us\/geforce\/products\/10series\/titan-xp\/)."},{"key":"e_1_3_2_1_7_1","unstructured":"http:\/\/www.image-net.org\/challenges\/LSVRC\/. ImageNet Large Scale Visual Recognition Challenge(ILSVRC). (http:\/\/www.image-net.org\/challenges\/LSVRC\/).  http:\/\/www.image-net.org\/challenges\/LSVRC\/. ImageNet Large Scale Visual Recognition Challenge(ILSVRC). (http:\/\/www.image-net.org\/challenges\/LSVRC\/)."},{"volume-title":"12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16)","author":"Abadi Martin","key":"e_1_3_2_1_8_1","unstructured":"Martin Abadi , Paul Barham , and Jianmin Chen et al. 2016. TensorFlow: A system for large-scale machine learning . In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16) . 265--283. Martin Abadi, Paul Barham, and Jianmin Chen et al. 2016. TensorFlow: A system for large-scale machine learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). 265--283."},{"key":"e_1_3_2_1_9_1","volume-title":"VLSI Technology (VLSIT), 2013 Symposium on. T134--T135","author":"Aoki M","year":"2013","unstructured":"M Aoki , H Noshiro , K Tsunoda , Y Iba , A Hatada , M Nakabayashi , A Takahashi , C Yoshida , Y Yamazaki , T Takenaga , and others. 2013 . Novel highly scalable multi-level cell for STT-MRAM with stacked perpendicular MTJs . In VLSI Technology (VLSIT), 2013 Symposium on. T134--T135 . M Aoki, H Noshiro, K Tsunoda, Y Iba, A Hatada, M Nakabayashi, A Takahashi, C Yoshida, Y Yamazaki, T Takenaga, and others. 2013. Novel highly scalable multi-level cell for STT-MRAM with stacked perpendicular MTJs. In VLSI Technology (VLSIT), 2013 Symposium on. T134--T135."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1102351.1102363"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2016.2625245"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.58"},{"volume-title":"Circuits and Systems (MWSCAS), 2010 53rd IEEE International Midwest Symposium on. 1109--1112","author":"Chen Yiran","key":"e_1_3_2_1_13_1","unstructured":"Yiran Chen , Xiaobin Wang , and Wenzhong et al. Zhu. 2010. Access scheme of multi-level cell spin-transfer torque random access memory and its optimization . In Circuits and Systems (MWSCAS), 2010 53rd IEEE International Midwest Symposium on. 1109--1112 . Yiran Chen, Xiaobin Wang, and Wenzhong et al. Zhu. 2010. Access scheme of multi-level cell spin-transfer torque random access memory and its optimization. In Circuits and Systems (MWSCAS), 2010 53rd IEEE International Midwest Symposium on. 1109--1112."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.40"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.13"},{"volume-title":"Computer-Aided Design (ICCAD), 2014 IEEE\/ACM International Conference on. IEEE, 301--308","author":"Chi Ping","key":"e_1_3_2_1_16_1","unstructured":"Ping Chi , Cong Xu , and Tao et al. Zhang. 2014. Using multi-level cell STT-RAM for fast and energy-efficient local checkpointing . In Computer-Aided Design (ICCAD), 2014 IEEE\/ACM International Conference on. IEEE, 301--308 . Ping Chi, Cong Xu, and Tao et al. Zhang. 2014. Using multi-level cell STT-RAM for fast and energy-efficient local checkpointing. In Computer-Aided Design (ICCAD), 2014 IEEE\/ACM International Conference on. IEEE, 301--308."},{"key":"e_1_3_2_1_17_1","volume-title":"Proceedings of the 33rd International Conference on International Conference on Machine Learning -","volume":"48","author":"Diamos Gregory","year":"2016","unstructured":"Gregory Diamos , Shubho Sengupta , Bryan Catanzaro , Mike Chrzanowski , Adam Coates , Erich Elsen , Jesse Engel , Awni Hannun , and Sanjeev Satheesh . 2016 . Persistent RNNs: Stashing Recurrent Weights On-chip . In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48 . Gregory Diamos, Shubho Sengupta, Bryan Catanzaro, Mike Chrzanowski, Adam Coates, Erich Elsen, Jesse Engel, Awni Hannun, and Sanjeev Satheesh. 2016. Persistent RNNs: Stashing Recurrent Weights On-chip. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48."},{"volume-title":"Emerging Memory Technologies","author":"Dong Xiangyu","key":"e_1_3_2_1_18_1","unstructured":"Xiangyu Dong , Cong Xu , Norm Jouppi , and Yuan Xie . 2014. NVSim: A circuit-level performance, energy, and area model for emerging non-volatile memory . In Emerging Memory Technologies . Springer , 15--50. Xiangyu Dong, Cong Xu, Norm Jouppi, and Yuan Xie. 2014. NVSim: A circuit-level performance, energy, and area model for emerging non-volatile memory. In Emerging Memory Technologies. Springer, 15--50."},{"key":"e_1_3_2_1_19_1","unstructured":"Xiaohua Lou et al. 2008. Demonstration of multilevel cell spin transfer switching in MgO magnetic tunnel junctions. Applied Physics Letters (2008).  Xiaohua Lou et al. 2008. Demonstration of multilevel cell spin transfer switching in MgO magnetic tunnel junctions. Applied Physics Letters (2008)."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2429384.2429498"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.30"},{"key":"e_1_3_2_1_22_1","volume-title":"International Conference on Learning Representations (ICLR)","author":"Han Song","year":"2016","unstructured":"Song Han , Huizi Mao , and William J Dally . 2016 . Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding . International Conference on Learning Representations (ICLR) (2016). Song Han, Huizi Mao, and William J Dally. 2016. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. International Conference on Learning Representations (ICLR) (2016)."},{"key":"e_1_3_2_1_23_1","unstructured":"Stephen Jos\u00e9 Hanson and Lorien Y Pratt. 1989. Comparing biases for minimal network construction with back-propagation. In Advances in neural information processing systems. 177--185.   Stephen Jos\u00e9 Hanson and Lorien Y Pratt. 1989. Comparing biases for minimal network construction with back-propagation. In Advances in neural information processing systems. 177--185."},{"key":"e_1_3_2_1_24_1","unstructured":"Babak Hassibi and David G Stork. 1993. Second order derivatives for network pruning: Optimal brain surgeon. In Advances in neural information processing systems. 164--171.   Babak Hassibi and David G Stork. 1993. Second order derivatives for network pruning: Optimal brain surgeon. In Advances in neural information processing systems. 164--171."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"crossref","unstructured":"T. Ishigaki et al. 2010. A multi-level-cell spin-transfer torque memory with series-stacked magnetotunnel junctions. In VLSIT.  T. Ishigaki et al. 2010. A multi-level-cell spin-transfer torque memory with series-stacked magnetotunnel junctions. In VLSIT.","DOI":"10.1109\/VLSIT.2010.5556126"},{"key":"e_1_3_2_1_27_1","unstructured":"Jeff Janzen. The Micron system-power calculator. (????). http:\/\/www.micron.com\/products\/dram\/syscalc.html.  Jeff Janzen. The Micron system-power calculator. (????). http:\/\/www.micron.com\/products\/dram\/syscalc.html."},{"key":"e_1_3_2_1_28_1","volume-title":"Caffe: Convolutional Architecture for Fast Feature Embedding. arXiv preprint arXiv:1408.5093","author":"Jia Yangqing","year":"2014","unstructured":"Yangqing Jia , Evan Shelhamer , and Jeff Donahue et al. 2014 . Caffe: Convolutional Architecture for Fast Feature Embedding. arXiv preprint arXiv:1408.5093 (2014). Yangqing Jia, Evan Shelhamer, and Jeff Donahue et al. 2014. Caffe: Convolutional Architecture for Fast Feature Embedding. arXiv preprint arXiv:1408.5093 (2014)."},{"key":"e_1_3_2_1_29_1","volume-title":"Raquel Urtasun, and Andreas Moshovos.","author":"Judd Patrick","year":"2015","unstructured":"Patrick Judd , Jorge Albericio , Tayler Hetherington , Tor Aamodt , Natalie Enright Jerger , Raquel Urtasun, and Andreas Moshovos. 2015 . Reduced-precision strategies for bounded memory in deep neural nets. arXiv preprint arXiv:1511.05236 (2015). Patrick Judd, Jorge Albericio, Tayler Hetherington, Tor Aamodt, Natalie Enright Jerger, Raquel Urtasun, and Andreas Moshovos. 2015. Reduced-precision strategies for bounded memory in deep neural nets. arXiv preprint arXiv:1511.05236 (2015)."},{"key":"e_1_3_2_1_30_1","unstructured":"Heehoon Kim Hyoungwook Nam Wookeun Jung and Jaejin Lee. 2017. Performance Analysis of CNN Frameworks for GPUs. In ISPASS.  Heehoon Kim Hyoungwook Nam Wookeun Jung and Jaejin Lee. 2017. Performance Analysis of CNN Frameworks for GPUs. In ISPASS."},{"key":"e_1_3_2_1_31_1","unstructured":"A. Krizhevsky. 2014. One Weird Trick For Parallelizing Convolutional Neural Networks. In arxiv.org.  A. Krizhevsky. 2014. One Weird Trick For Parallelizing Convolutional Neural Networks. In arxiv.org."},{"key":"e_1_3_2_1_32_1","unstructured":"Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems. 1097--1105.   Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems. 1097--1105."},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669172"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080254"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.32"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"crossref","unstructured":"Minsoo Rhu Natalia Gimelshein and Jason Clemons et al. 2016. vDNN: Virtualized Deep Neural Networks for Scalable Memory-Efficient Neural Network Design. In MICRO.   Minsoo Rhu Natalia Gimelshein and Jason Clemons et al. 2016. vDNN: Virtualized Deep Neural Networks for Scalable Memory-Efficient Neural Network Design. In MICRO.","DOI":"10.1109\/MICRO.2016.7783721"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.12"},{"key":"e_1_3_2_1_38_1","volume-title":"Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1--9.","author":"Szegedy Christian","key":"e_1_3_2_1_39_1","unstructured":"Christian Szegedy , Wei Liu , and Yangqing Jia et al. 2015. Going deeper with convolutions . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1--9. Christian Szegedy, Wei Liu, and Yangqing Jia et al. 2015. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1--9."},{"key":"e_1_3_2_1_40_1","unstructured":"The Next Platform. 2016. Baidu Eyes Deep Learning Strategy in Wake of New GPU Options. In www.nextplatform.com.  The Next Platform. 2016. Baidu Eyes Deep Learning Strategy in Wake of New GPU Options. In www.nextplatform.com."},{"key":"e_1_3_2_1_41_1","volume-title":"Theano: A Python framework for fast computation of mathematical expressions. arXiv e-prints abs\/1605.02688 (May","author":"Team Theano Development","year":"2016","unstructured":"Theano Development Team . 2016 . Theano: A Python framework for fast computation of mathematical expressions. arXiv e-prints abs\/1605.02688 (May 2016). Theano Development Team. 2016. Theano: A Python framework for fast computation of mathematical expressions. arXiv e-prints abs\/1605.02688 (May 2016)."}],"event":{"name":"MEMSYS '18: The International Symposium on Memory Systems","acronym":"MEMSYS '18","location":"Alexandria Virginia USA"},"container-title":["Proceedings of the International Symposium on Memory Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3240302.3240422","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3240302.3240422","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3240302.3240422","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:43:42Z","timestamp":1750207422000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3240302.3240422"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,10]]},"references-count":41,"alternative-id":["10.1145\/3240302.3240422","10.1145\/3240302"],"URL":"https:\/\/doi.org\/10.1145\/3240302.3240422","relation":{},"subject":[],"published":{"date-parts":[[2018,10]]},"assertion":[{"value":"2018-10-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}