{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:16:20Z","timestamp":1750220180932,"version":"3.41.0"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2022,10,26]],"date-time":"2022-10-26T00:00:00Z","timestamp":1666742400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Science Foundation","award":["1948201, 2000851"],"award-info":[{"award-number":["1948201, 2000851"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["J. Emerg. Technol. Comput. Syst."],"published-print":{"date-parts":[[2022,10,31]]},"abstract":"<jats:p>Spiking neural networks (SNNs) are brain-inspired event-driven models of computation with promising ultra-low energy dissipation. Rich network dynamics emergent in recurrent spiking neural networks (R-SNNs) can form temporally based memory, offering great potential in processing complex spatiotemporal data. However, recurrence in network connectivity produces tightly coupled data dependency in both space and time, rendering hardware acceleration of R-SNNs challenging. We present the first work to exploit spatiotemporal parallelisms to accelerate the R-SNN-based inference on systolic arrays using an architecture called SaARSP. We decouple the processing of feedforward synaptic connections from that of recurrent connections to allow for the exploitation of parallelisms across multiple time points. We propose a novel time window size optimization (TWSO) technique, to further explore the temporal granularity of the proposed decoupling in terms of optimal time window size and reconfiguration of the systolic array considering layer-dependent connectivity to boost performance. Stationary dataflow and time window size are jointly optimized to trade off between weight data reuse and movements of partial sums, the two bottlenecks in latency and energy dissipation of the accelerator. The proposed systolic-array architecture offers a unifying solution to an acceleration of both feedforward and recurrent SNNs, and delivers 4,000X EDP improvement on average for different R-SNN benchmarks over a conventional baseline.<\/jats:p>","DOI":"10.1145\/3510854","type":"journal-article","created":{"date-parts":[[2022,6,27]],"date-time":"2022-06-27T12:53:46Z","timestamp":1656334426000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["SaARSP: An Architecture for Systolic-Array Acceleration of Recurrent Spiking Neural Networks"],"prefix":"10.1145","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7370-889X","authenticated-orcid":false,"given":"Jeong-Jun","family":"Lee","sequence":"first","affiliation":[{"name":"University of California, Santa Barbara, California, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1004-4499","authenticated-orcid":false,"given":"Wenrui","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of California, Santa Barbara, California, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2093-1788","authenticated-orcid":false,"given":"Yuan","family":"Xie","sequence":"additional","affiliation":[{"name":"University of California, Santa Barbara, California, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3548-4589","authenticated-orcid":false,"given":"Peng","family":"Li","sequence":"additional","affiliation":[{"name":"University of California, Santa Barbara, California, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,26]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.05.044"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2015.2474396"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.3389\/fnins.2018.00023"},{"key":"e_1_3_1_5_2","first-page":"787","volume-title":"Advances in Neural Information Processing Systems","author":"Bellec Guillaume","year":"2018","unstructured":"Guillaume Bellec, Darjan Salaj, Anand Subramoney, Robert Legenstein, and Wolfgang Maass. 2018. Long short-term memory and learning-to-learn in networks of spiking neurons. In Advances in Neural Information Processing Systems. 787\u2013797."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1523\/JNEUROSCI.18-24-10464.1998"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3304103"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-014-0788-3"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1146\/annurev.neuro.31.060407.125639"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2013.6707077"},{"key":"e_1_3_1_11_2","first-page":"33","volume-title":"Proceedings of the 2012 Design, Automation & Test in Europe Conference & Exhibition (DATE)","author":"Chen Ke","year":"2012","unstructured":"Ke Chen, Sheng Li, Naveen Muralimanohar, Jung Ho Ahn, Jay B. Brockman, and Norman P. Jouppi. 2012. Cacti-3DD: Architecture-Level Modeling for 3D Die-Stacked dram main memory. In Proceedings of the 2012 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 33\u201338."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/2541940.2541967"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2017.54"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2016.2616357"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33269-2_15"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CICC.2019.8780116"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/icassp40776.2020.9053856"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2018.112130359"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1146\/annurev-psych-113011-143750"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750389"},{"key":"e_1_3_1_21_2","article-title":"Timit acoustic phonetic continuous speech corpus","author":"Garofolo John S.","year":"1993","unstructured":"John S. Garofolo. 1993. Timit acoustic phonetic continuous speech corpus. In Linguistic Data Consortium, (1993).","journal-title":"Linguistic Data Consortium,"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511815706"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2014.106"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3299874.3317966"},{"key":"e_1_3_1_25_2","article-title":"Toward the optimal design and FPGA implementation of spiking neural networks","author":"Guo Wenzhe","year":"2021","unstructured":"Wenzhe Guo, Hasan Erdem Yantir, Mohammed E. Fouda, Ahmed M. Eltawil, and Khaled Nabil Salama. 2021. Toward the optimal design and FPGA implementation of spiking neural networks. IEEE Transactions on Neural Networks and Learning Systems (2021).","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001163"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISVLSI49217.2020.00088"},{"key":"e_1_3_1_28_2","volume-title":"Proceedings of the Conference on Neural Information Processing Systems","author":"Jin Yingyezhe","year":"2018","unstructured":"Yingyezhe Jin, Wenrui Zhang, and Peng Li. 2018. Hybrid macro\/micro level backpropagation for training deep spiking neural networks. In Proceedings of the Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2017.12.005"},{"key":"e_1_3_1_30_2","first-page":"1","article-title":"Introducing Qualcomm zeroth processors: Brain-inspired computing","author":"Kumar Samir","year":"2013","unstructured":"Samir Kumar. 2013. Introducing Qualcomm zeroth processors: Brain-inspired computing. Qualcomm ONQ Blog (2013), 1\u201311.","journal-title":"Qualcomm ONQ Blog"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358252"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.3389\/fnins.2016.00508"},{"key":"e_1_3_1_34_2","unstructured":"Mark Liberman Robert Amsler kKen Church Ed Fox Carole Hafner Judy Klavans Mitch Marcus Bob Mercer Jan Pedersen Paul Roossin Don Walker Susan Warwick and Antonio Zampolli. 1991. TI 46-word IDC93s9. https:\/\/catalog.ldc.upenn.edu\/docs\/ldc93s9\/ti46.readme.html."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1982.1171644"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0893-6080(97)00011-7"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1162\/089976602760407955"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TBCAS.2017.2759700"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00038"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2013.2294916"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.3389\/fnins.2015.00437"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2016.2572164"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD.2013.6657019"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.3389\/fnins.2015.00141"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-18338-7_22"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11265-020-01546-x"},{"key":"e_1_3_1_47_2","article-title":"Scale-sim: Systolic CNN accelerator simulator","author":"Samajdar Ananda","year":"2018","unstructured":"Ananda Samajdar, Yuhao Zhu, Paul Whatmough, Matthew Mattina, and Tushar Krishna. 2018. Scale-sim: Systolic CNN accelerator simulator. arxiv preprint arxiv:1811.02883 (2018).","journal-title":"arxiv preprint arxiv:1811.02883"},{"key":"e_1_3_1_48_2","first-page":"1412","volume-title":"Advances in Neural Information Processing Systems","author":"Shrestha Sumit Bam","year":"2018","unstructured":"Sumit Bam Shrestha and Garrick Orchard. 2018. Slayer: Spike layer error reassignment in time. In Advances in Neural Information Processing Systems. 1412\u20131421."},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3243176.3243184"},{"key":"e_1_3_1_50_2","article-title":"A power-efficient binary-weight spiking neural network architecture for real-time object classification","author":"Tan Pai-Yu","year":"2020","unstructured":"Pai-Yu Tan, Po-Yao Chuang, Yen-Ting Lin, Cheng-Wen Wu, and Juin-Ming Lu. 2020. A power-efficient binary-weight spiking neural network architecture for real-time object classification. arxiv preprint arxiv:2003.06310 (2020).","journal-title":"arxiv preprint arxiv:2003.06310"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2018.12.002"},{"key":"e_1_3_1_52_2","article-title":"Rethinking full connectivity in recurrent neural networks","author":"Keirsbilck Matthijs van","year":"2019","unstructured":"Matthijs van Keirsbilck, Alexander Keller, and Xiaodong Yang. 2019. Rethinking full connectivity in recurrent neural networks. arxiv preprint arxiv:1905.12340 (2019).","journal-title":"arxiv preprint arxiv:1905.12340"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11390-020-9686-z"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.3389\/fnins.2018.00331"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41563-019-0291-x"},{"key":"e_1_3_1_56_2","article-title":"Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms","author":"Xiao Han","year":"2017","unstructured":"Han Xiao, Kashif Rasul, and Roland Vollgraf. 2017. Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms. arxiv preprint arxiv:1708.07747 (2017).","journal-title":"arxiv preprint arxiv:1708.07747"},{"key":"e_1_3_1_57_2","first-page":"7802","volume-title":"Advances in Neural Information Processing Systems","author":"Zhang Wenrui","year":"2019","unstructured":"Wenrui Zhang and Peng Li. 2019. Spike-train level backpropagation for training deep recurrent spiking neural networks. In Advances in Neural Information Processing Systems. 7802\u20137813."},{"key":"e_1_3_1_58_2","article-title":"Temporal spike sequence learning via backpropagation for deep spiking neural networks","author":"Zhang Wenrui","year":"2020","unstructured":"Wenrui Zhang and Peng Li. 2020. Temporal spike sequence learning via backpropagation for deep spiking neural networks. arxiv preprint arxiv:2002.10085 (2020).","journal-title":"arxiv preprint arxiv:2002.10085"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco_a_01393"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2015.2388544"}],"container-title":["ACM Journal on Emerging Technologies in Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3510854","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3510854","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:02:12Z","timestamp":1750186932000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3510854"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,26]]},"references-count":59,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2022,10,31]]}},"alternative-id":["10.1145\/3510854"],"URL":"https:\/\/doi.org\/10.1145\/3510854","relation":{},"ISSN":["1550-4832","1550-4840"],"issn-type":[{"type":"print","value":"1550-4832"},{"type":"electronic","value":"1550-4840"}],"subject":[],"published":{"date-parts":[[2022,10,26]]},"assertion":[{"value":"2021-03-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-01-09","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-10-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}