{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,22]],"date-time":"2026-01-22T00:27:20Z","timestamp":1769041640316,"version":"3.49.0"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2024,9,14]],"date-time":"2024-09-14T00:00:00Z","timestamp":1726272000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key R&D","award":["2021YFB0300300"],"award-info":[{"award-number":["2021YFB0300300"]}]},{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["62172430, U22A2027,62172155"],"award-info":[{"award-number":["62172430, U22A2027,62172155"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"STIP of Hunan Province","award":["2022RC3065"],"award-info":[{"award-number":["2022RC3065"]}]},{"name":"Foundation of PDL","award":["2023-JKWPDL-02"],"award-info":[{"award-number":["2023-JKWPDL-02"]}]},{"name":"Key Laboratory of Advanced Microprocessor Chips and Systems"},{"name":"Hunan Postgraduate Research Innovation Project","award":["CX20220077"],"award-info":[{"award-number":["CX20220077"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2024,9,30]]},"abstract":"<jats:p>As the Convolutional Neural Network (CNN) goes deeper and more complex, the network becomes memory-intensive and computation-intensive. To address this issue, the lightweight neural network reduces parameters and Multiplication-and-Accumulation (MAC) operations by using the Depthwise Separable Convolution (DSC) to improve speed and efficiency. Nonetheless, the energy efficiency of classical Von Neumann architectures for CNNs is limited due to the memory wall challenge. Spin-based architectures have the potential to address this challenge thanks to the integration of memory and computing with ultra-high energy efficiency. However, deploying the DSC on spin-based architectures with the traditional dataflow leads to huge activation movements and low hardware utilization. Moreover, the inter-layer data dependency of neural networks increases latency. These factors become the bottleneck of improving energy efficiency and performance.<\/jats:p><jats:p>Inspired by these challenges, we propose a novel dataflow on Spin-based Architectures for Lightweight neural networks (SAL). The novel dataflow replaces convolution unrolling by selecting activations in the crossbar according to the convolution window and also realizes the inter-layer data reuse. Moreover, the novel dataflow also reduces the latency due to the data dependency between layers, realizing higher performance. To the best of our knowledge, this is the first design to use hybrid dataflow for the PIM architecture. We also optimize the structure of the spin-based crossbar and the pipeline based on the dataflow to achieve better data reuse and computational parallelism. For deploying the MobileNet V1, the novel dataflow improves the hardware utilization by 23\u00d7\u223c 105\u00d7 and reduces the data traffic by 1.09\u00d7\u223c 18.6\u00d7. Compared with the NEBULA, a spin-based non-Von Neumann architecture, the SAL reduces the energy consumption by 4\u00d7 and improves the performance by 7.3\u00d7, which are 0.32<jats:italic>mJ<\/jats:italic>and 10.43<jats:italic>GOPs<\/jats:italic><jats:sup>-1<\/jats:sup>, respectively. Moreover, the SAL improves power efficiency over 29 times more than the NEBULA. Compared with the Eyeriss, the SAL improves the energy efficiency by four orders of magnitude.<\/jats:p>","DOI":"10.1145\/3673654","type":"journal-article","created":{"date-parts":[[2024,6,14]],"date-time":"2024-06-14T11:29:10Z","timestamp":1718364550000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["SAL: Optimizing the Dataflow of Spin-based Architectures for Lightweight Neural Networks"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5600-3740","authenticated-orcid":false,"given":"Yunping","family":"Zhao","sequence":"first","affiliation":[{"name":"Institute for Quantum Information &amp; State Key Laboratory of High Performance Computing, National University of Defense Technology, changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1710-4060","authenticated-orcid":false,"given":"Sheng","family":"Ma","sequence":"additional","affiliation":[{"name":"School of Computer, National University of Defense Technology, changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4912-0364","authenticated-orcid":false,"given":"Hengzhu","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer, National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9743-2034","authenticated-orcid":false,"given":"Dongsheng","family":"Li","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Parallel and Distributed Computing, National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,9,14]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2022.3222966"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2023.3298705"},{"key":"e_1_3_1_4_2","first-page":"C264\u2013C265","volume-title":"Proceedings of the 2013 Symposium on VLSI Circuits","author":"Chen Vanessa H-C","year":"2013","unstructured":"Vanessa H-C Chen and Lawrence Pileggi. 2013. An 8.5 mW 5GS\/s 6b flash ADC with dynamic offset calibration in 32nm CMOS SOI. In Proceedings of the 2013 Symposium on VLSI Circuits. IEEE, C264\u2013C265."},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","first-page":"609","DOI":"10.1109\/MICRO.2014.58","volume-title":"Proceedings of the 2014 47th Annual IEEE\/ACM International Symposium on Microarchitecture","author":"Chen Yunji","year":"2014","unstructured":"Yunji Chen, Tao Luo, Shaoli Liu, Shijin Zhang, Liqiang He, Jia Wang, Ling Li, Tianshi Chen, Zhiwei Xu, Ninghui Sun, and Olivier Temam. 2014. DaDianNao: A machine-learning supercomputer. In Proceedings of the 2014 47th Annual IEEE\/ACM International Symposium on Microarchitecture. 609\u2013622."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2016.2616357"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2910232"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001140"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.195"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2020.2981068"},{"key":"e_1_3_1_11_2","article-title":"Designing efficient accelerator of depthwise separable convolutional neural network on FPGA","author":"Ding Wei","year":"2020","unstructured":"Wei Ding, Zeyu Huang, Zunkai Huang, Li Tian, Hui Wang, and Songlin Feng. 2020. Designing efficient accelerator of depthwise separable convolutional neural network on FPGA. Journal of Systems Architecture 97, 1 (2020), 163\u2013173.","journal-title":"Journal of Systems Architecture"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","first-page":"51","DOI":"10.1109\/SISPAD.2011.6035047","volume-title":"Proceedings of the 2011 International Conference on Simulation of Semiconductor Processes and Devices","author":"Fong Xuanyao","year":"2011","unstructured":"Xuanyao Fong, Sumeet K. Gupta, Niladri N. Mojumder, Sri Harsha Choday, Charles Augustine, and Kaushik Roy. 2011. KNACK: A hybrid spin-charge mixed-mode simulator for evaluating different genres of spin-transfer torque MRAM bit-cells. In Proceedings of the 2011 International Conference on Simulation of Semiconductor Processes and Devices. IEEE, 51\u201354."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41928-019-0360-9"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2014.2353793"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_3_1_16_2","unstructured":"Andrew G. Howard Menglong Zhu Bo Chen Dmitry Kalenichenko Weijun Wang Tobias Weyand Marco Andreetto and Hartwig Adam. 2017. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv:1704.04861. Retrieved from https:\/\/arxiv.org\/abs\/1704.04861"},{"key":"e_1_3_1_17_2","first-page":"19:1\u201319:6","volume-title":"Proceedings of the 53rd Annual Design Automation Conference, DAC 2016, Austin, TX, USA, June 5-9, 2016","author":"Hu Miao","year":"2016","unstructured":"Miao Hu, John Paul Strachan, Zhiyong Li, Emmanuelle M. Grafals, Noraica Davila, Catherine Graves, Sity Lam, Ning Ge, Jianhua Joshua Yang, and R. Stanley Williams. 2016. Dot-product engine for neuromorphic computing: Programming 1T1M crossbar to accelerate matrix-vector multiplication. In Proceedings of the 53rd Annual Design Automation Conference, DAC 2016, Austin, TX, USA, June 5-9, 2016. ACM, 19:1\u201319:6."},{"issue":"1","key":"e_1_3_1_18_2","first-page":"6:1\u20136:17","article-title":"Rescuing ReRAM-based neural computing systems from device variation","volume":"28","author":"Huang Chenglong","year":"2023","unstructured":"Chenglong Huang, Nuo Xu, Junwei Zeng, Wenqing Wang, Yihong Hu, Liang Fang, Desheng Ma, and Yanting Chen. 2023. Rescuing ReRAM-based neural computing systems from device variation. ACM Transactions on Design Automation of Electronic Systems 28, 1 (2023), 6:1\u20136:17.","journal-title":"ACM Transactions on Design Automation of Electronic Systems"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10070992"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3299874.3319450"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSI.2021.3078541"},{"issue":"1","key":"e_1_3_1_22_2","first-page":"70","article-title":"Optimizing depthwise separable convolution operations on GPUs","volume":"33","author":"Lu Gangzhao","year":"2021","unstructured":"Gangzhao Lu, Weizhe Zhang, and Zheng Wang. 2021. Optimizing depthwise separable convolution operations on GPUs. IEEE Transactions on Parallel and Distributed Systems 33, 1 (2021), 70\u201387.","journal-title":"IEEE Transactions on Parallel and Distributed Systems"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_8"},{"issue":"4","key":"e_1_3_1_24_2","doi-asserted-by":"crossref","first-page":"1333","DOI":"10.1109\/TCSI.2019.2958568","article-title":"Optimizing weight mapping and data flow for convolutional neural networks on processing-in-memory architectures","volume":"67","author":"Peng Xiaochen","year":"2019","unstructured":"Xiaochen Peng, Rui Liu, and Shimeng Yu. 2019. Optimizing weight mapping and data flow for convolutional neural networks on processing-in-memory architectures. IEEE Transactions on Circuits and Systems I: Regular Papers 67, 4 (2019), 1333\u20131343.","journal-title":"IEEE Transactions on Circuits and Systems I: Regular Papers"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSI.2011.2107214"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","first-page":"4557","DOI":"10.1109\/IJCNN.2017.7966434","volume-title":"Proceedings of the 2017 International Joint Conference on Neural Networks (IJCNN)","author":"Sengupta Abhronil","year":"2017","unstructured":"Abhronil Sengupta, Aayush Ankit, and Kaushik Roy. 2017. Performance analysis and benchmarking of all-spin spiking neural networks (special session paper). In Proceedings of the 2017 International Joint Conference on Neural Networks (IJCNN). IEEE, 4557\u20134563."},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","first-page":"544","DOI":"10.1109\/BioCAS.2016.7833852","volume-title":"Proceedings of the 2016 IEEE Biomedical Circuits and Systems Conference (BioCAS)","author":"Sengupta Abhronil","year":"2016","unstructured":"Abhronil Sengupta, Bing Han, and Kaushik Roy. 2016. Toward a spintronic deep learning spiking neural processor. In Proceedings of the 2016 IEEE Biomedical Circuits and Systems Conference (BioCAS). IEEE, 544\u2013547."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TBCAS.2016.2525823"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001139"},{"key":"e_1_3_1_31_2","first-page":"1","volume-title":"Proceedings of the 2012 International Joint Conference on Neural Networks (IJCNN)","author":"Sharad Mrigank","year":"2012","unstructured":"Mrigank Sharad, Charles Augustine, Georgios Panagopoulos, and Kaushik Roy. 2012. Spin based neuron-synapse module for ultra low power programmable computational networks. In Proceedings of the 2012 International Joint Conference on Neural Networks (IJCNN). IEEE, 1\u20137."},{"issue":"2","key":"e_1_3_1_32_2","doi-asserted-by":"crossref","first-page":"1033","DOI":"10.1021\/acs.nanolett.9b04200","article-title":"Magnetic domain wall-based synaptic and activation function generator for neuromorphic accelerators","volume":"20","author":"Siddiqui Saima A.","year":"2019","unstructured":"Saima A. Siddiqui, Sumit Dutta, Astera Tang, Luqiao Liu, Caroline A. Ross, and Marc A. Baldo. 2019. Magnetic domain wall-based synaptic and activation function generator for neuromorphic accelerators. Nano Letters 20, 2 (2019), 1033\u20131040.","journal-title":"Nano Letters"},{"key":"e_1_3_1_33_2","first-page":"363","volume-title":"Proceedings of the 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA)","author":"Singh Sonali","year":"2020","unstructured":"Sonali Singh, Anup Sarma, Nicholas Jao, Ashutosh Pattnaik, Sen Lu, and Kezhou Yang. 2020. NEBULA: A neuromorphic spin-based ultra-low power architecture for SNNs and ANNs. In Proceedings of the 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). IEEE, 363\u2013376."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3604802"},{"issue":"11","key":"e_1_3_1_35_2","first-page":"2860","article-title":"Heterogeneous systolic array architecture for compact CNNs hardware accelerators","volume":"33","author":"Xu Rui","year":"2021","unstructured":"Rui Xu, Sheng Ma, and Yaohua Wang. 2021. Heterogeneous systolic array architecture for compact CNNs hardware accelerators. IEEE Transactions on Parallel and Distributed Systems 33, 11 (2021), 2860\u20132871.","journal-title":"IEEE Transactions on Parallel and Distributed Systems"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460776"},{"key":"e_1_3_1_37_2","first-page":"11.4.1\u201311.4.4","volume-title":"Proceedings of the 2017 IEEE International Electron Devices Meeting (IEDM)","author":"Yan Bonan","year":"2017","unstructured":"Bonan Yan, Chenchen Liu, Xiaoxiao Liu, Yiran Chen, and Hai Li. 2017. Understanding the tradeoffs of device, circuit and application in ReRAM-based neuromorphic computing systems. In Proceedings of the 2017 IEEE International Electron Devices Meeting (IEDM). 11.4.1\u201311.4.4."},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","first-page":"1916","DOI":"10.1109\/ACSSC.2017.8335698","volume-title":"Proceedings of the 2017 51st Asilomar Conference on Signals, Systems, and Computers","author":"Yang Tien-Ju","year":"2017","unstructured":"Tien-Ju Yang, Yu-Hsin Chen, and Joel Emer. 2017. A method to estimate the energy consumption of deep neural networks. In Proceedings of the 2017 51st Asilomar Conference on Signals, Systems, and Computers. IEEE, 1916\u20131920."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-020-1942-4"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TBCAS.2016.2533798"},{"key":"e_1_3_1_41_2","volume-title":"Proceedings of the 2016 IEEE\/ACM International Symposium on Nanoscale Architectures (NANOARCH)","author":"Zhang Deming","year":"2016","unstructured":"Deming Zhang, Lang Zeng, Youguang Zhang, Weisheng Zhao, and Jacques Olivier Klein. 2016. Stochastic spintronic device based synapses and spiking neurons for neuromorphic computation. In Proceedings of the 2016 IEEE\/ACM International Symposium on Nanoscale Architectures (NANOARCH). IEEE."},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1002\/pssr.201900029"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00716"},{"issue":"7","key":"e_1_3_1_44_2","first-page":"128","article-title":"Spintronic devices for neuromorphic computing","volume":"63","author":"Zhang YaJun","year":"2020","unstructured":"YaJun Zhang, Qi Zheng, XiaoRui Zhu, Zhe Yuan, and Ke Xia. 2020. Spintronic devices for neuromorphic computing. Science China (Physics, Mechanics and Astronomy) 63, 7 (2020), 128\u2013130.","journal-title":"Science China (Physics, Mechanics and Astronomy)"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2023.3297968"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3643134"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3632957"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-021-3472-9"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2022.3184464"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3673654","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3673654","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:06:07Z","timestamp":1750291567000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3673654"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,14]]},"references-count":48,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,9,30]]}},"alternative-id":["10.1145\/3673654"],"URL":"https:\/\/doi.org\/10.1145\/3673654","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,9,14]]},"assertion":[{"value":"2024-01-24","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-06-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-09-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}