{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,10]],"date-time":"2025-11-10T21:21:22Z","timestamp":1762809682899,"version":"3.41.0"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2025,1,10]],"date-time":"2025-01-10T00:00:00Z","timestamp":1736467200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U23A20301, 62102433, 62272477"],"award-info":[{"award-number":["U23A20301, 62102433, 62272477"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"NUDT Foundation","award":["No.ZK2023-16, 2022-KJWPDL-08"],"award-info":[{"award-number":["No.ZK2023-16, 2022-KJWPDL-08"]}]},{"name":"NSF of Hunan Province","award":["2022JJ10066"],"award-info":[{"award-number":["2022JJ10066"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>\n            Digital SRAM-based CIM architectures must balance three critical factors: quantized neural network bitwidth, accuracy loss, and computational efficiency, each crucial to optimizing performance and efficiency. In Domain Specific Accelerators (DSAs), flexible and specific hardware design, when incorporated with tailored Power-of-2 (P-2) quantization schemes, addresses this issue. However, in CIMs, the absence of flexible and specific hardware to support dynamic switching between general and tailored quantization schemes hinders the adoption of efficient quantization methods. In this article, we propose the\n            <jats:bold>I<\/jats:bold>\n            n-situ\n            <jats:bold>S<\/jats:bold>\n            hift\n            <jats:bold>O<\/jats:bold>\n            peration based\n            <jats:bold>Acc<\/jats:bold>\n            elerator (\n            <jats:bold>ISOAcc<\/jats:bold>\n            ) for efficient SRAM-based multiplication. The key idea is to introduce transmission gates near the SRAM array to enable the selection of bits from either the same or the neighbor line when data flows from one row to another. This functionally equals a shift operation. By configuring the transmission gates array in a cascade manner, ISOAcc can support 0 to 15-bit shift with a negligible overhead. The ISOAcc can directly leverage P-2 quantization schemes in hardware, thereby greatly reducing multiplication cycles. We have chosen five well-known neural networks to evaluate ISOAcc. The evaluations show that ISOAcc achieves an average performance improvement of 3.24\u00d7 and an energy reduction of 75%, compared with the state-of-the-art (SOTA) SRAM-based CIM design, Bit-Parallel.\n          <\/jats:p>","DOI":"10.1145\/3707205","type":"journal-article","created":{"date-parts":[[2024,12,5]],"date-time":"2024-12-05T09:46:21Z","timestamp":1733391981000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["ISOAcc: In-situ Shift Operation-based Accelerator For Efficient in-SRAM Multiplication"],"prefix":"10.1145","volume":"30","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5620-2355","authenticated-orcid":false,"given":"Gaoyang","family":"Zhao","sequence":"first","affiliation":[{"name":"National University of Defense Technology College of Computer Science and Technology, Changsha, China and Key Labtoratory of Advanced Microprocessor Chips and Systems, National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6233-6800","authenticated-orcid":false,"given":"Junzhong","family":"Shen","sequence":"additional","affiliation":[{"name":"National University of Defense Technology College of Computer Science and Technology, Changsha, China and Key Labtoratory of Advanced Microprocessor Chips and Systems, National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5561-4685","authenticated-orcid":false,"given":"Rongzhen","family":"Lin","sequence":"additional","affiliation":[{"name":"National University of Defense Technology College of Computer Science and Technology, Changsha, China and Key Labtoratory of Advanced Microprocessor Chips and Systems, National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-9792-6426","authenticated-orcid":false,"given":"Hua","family":"Li","sequence":"additional","affiliation":[{"name":"National University of Defense Technology College of Computer Science and Technology, Changsha, China and Key Laboratory of Advanced Microprocessor Chips and Systems, National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9556-5535","authenticated-orcid":false,"given":"Yaohua","family":"Wang","sequence":"additional","affiliation":[{"name":"National University of Defense Technology College of Computer Science and Technology, Changsha, China and Key Laboratory of Advanced Microprocessor Chips and Systems, National University of Defense Technology, Changsha, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,1,10]]},"reference":[{"doi-asserted-by":"publisher","key":"e_1_3_3_2_2","DOI":"10.1109\/HPCA.2017.21"},{"doi-asserted-by":"publisher","key":"e_1_3_3_3_2","DOI":"10.1109\/HPCA56546.2023.10071074"},{"unstructured":"Aojun Zhou Anbang Yao Yiwen Guo Lin Xu and Yurong Chen. 2017. Incremental network quantization: Towards lossless CNNs with low-precision weights. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=HyQJ-mclg","key":"e_1_3_3_4_2"},{"doi-asserted-by":"publisher","key":"e_1_3_3_5_2","DOI":"10.1145\/3079856.3080231"},{"doi-asserted-by":"publisher","key":"e_1_3_3_6_2","DOI":"10.1145\/3296957.3173177"},{"key":"e_1_3_3_7_2","series-title":"NIPS \u201920","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS \u201920). Curran Associates Inc., Red Hook, NY, Article 159, 25 pages."},{"doi-asserted-by":"publisher","key":"e_1_3_3_8_2","DOI":"10.1109\/ISSCC49657.2024.10454507"},{"doi-asserted-by":"publisher","key":"e_1_3_3_9_2","DOI":"10.1109\/HPCA51647.2021.00027"},{"doi-asserted-by":"publisher","key":"e_1_3_3_10_2","DOI":"10.1109\/MICRO.2014.58"},{"doi-asserted-by":"publisher","key":"e_1_3_3_11_2","DOI":"10.1109\/MM.2021.3061394"},{"unstructured":"Taiwan Semiconductor Manufacturing Company. 2024. 28nm Technology. Retrieved from https:\/\/www.tsmc.com\/english\/dedicatedFoundry\/technology\/logic\/l_28nm. 2024.04.15.","key":"e_1_3_3_12_2"},{"unstructured":"Semiconductor Manufacturing International Corporation. 2024. 28nm. Retrieved from https:\/\/www.smics.com\/jp\/site\/technology_advanced_Te. 2024.04.15.","key":"e_1_3_3_13_2"},{"key":"e_1_3_3_14_2","volume-title":"Synthesis Lectures on Computer Architecture","author":"Daichi Fujiki","year":"2021","unstructured":"Fujiki Daichi, Wang Xiaowei, Subramaniyan Arun, and Das Reetuparna. 2021. In-\/near-memory computing. In Synthesis Lectures on Computer Architecture. Morgan & Claypools. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:238679778"},{"doi-asserted-by":"publisher","key":"e_1_3_3_15_2","DOI":"10.1109\/TCAD.2012.2185930"},{"doi-asserted-by":"publisher","key":"e_1_3_3_16_2","DOI":"10.1109\/TCAD.2023.3330819"},{"doi-asserted-by":"publisher","key":"e_1_3_3_17_2","DOI":"10.1109\/ISCA.2018.00040"},{"doi-asserted-by":"publisher","key":"e_1_3_3_18_2","DOI":"10.1109\/ISSCC49657.2024.10454278"},{"doi-asserted-by":"publisher","key":"e_1_3_3_19_2","DOI":"10.1109\/ISSCC42615.2023.10067260"},{"doi-asserted-by":"publisher","key":"e_1_3_3_20_2","DOI":"10.1109\/ISSCC42613.2021.9365989"},{"key":"e_1_3_3_21_2","volume-title":"The Art of Analog Layout","author":"Hastings Alan","year":"2012","unstructured":"Alan Hastings and JF Hastings. 2012. The Art of Analog Layout. Pearson."},{"doi-asserted-by":"publisher","key":"e_1_3_3_22_2","DOI":"10.1109\/CVPR.2016.90"},{"doi-asserted-by":"publisher","key":"e_1_3_3_23_2","DOI":"10.1109\/ISSCC42615.2023.10067305"},{"doi-asserted-by":"publisher","key":"e_1_3_3_24_2","DOI":"10.1145\/3489517.3530446"},{"doi-asserted-by":"publisher","key":"e_1_3_3_25_2","DOI":"10.1109\/ARITH48897.2020.00029"},{"doi-asserted-by":"publisher","key":"e_1_3_3_26_2","DOI":"10.1109\/MSP.2012.2205597"},{"doi-asserted-by":"publisher","key":"e_1_3_3_27_2","DOI":"10.1162\/neco.1997.9.8.1735"},{"doi-asserted-by":"publisher","key":"e_1_3_3_28_2","DOI":"10.1109\/ISSCC42615.2023.10067335"},{"unstructured":"Itay Hubara Matthieu Courbariaux Daniel Soudry Ran El-Yaniv and Yoshua Bengio. 2016. Binarized neural networks. In Proceedings of the 30th International Conference on Neural Information Processing Systems (Barcelona Spain) (NIPS\u201916). Curran Associates Inc. Red Hook NY USA 4114\u20134122.","key":"e_1_3_3_29_2"},{"doi-asserted-by":"publisher","key":"e_1_3_3_30_2","DOI":"10.1109\/JBHI.2023.3345897"},{"doi-asserted-by":"publisher","key":"e_1_3_3_31_2","DOI":"10.1109\/IPDPSW55747.2022.00088"},{"doi-asserted-by":"publisher","key":"e_1_3_3_32_2","DOI":"10.1145\/3579371.3589350"},{"doi-asserted-by":"publisher","key":"e_1_3_3_33_2","DOI":"10.1109\/JSSC.2020.3039206"},{"doi-asserted-by":"publisher","key":"e_1_3_3_34_2","DOI":"10.1145\/3065386"},{"unstructured":"RIOS Laboratory. 2024. OpenRPDK28: Open PDK for 28nm Process Technology. Retrieved from https:\/\/github.com\/RIOSLaboratory\/OpenRPDK28. Accessed: 2024-12-02.","key":"e_1_3_3_35_2"},{"doi-asserted-by":"publisher","key":"e_1_3_3_36_2","DOI":"10.1109\/5.726791"},{"doi-asserted-by":"publisher","key":"e_1_3_3_37_2","DOI":"10.1109\/DAC18072.2020.9218567"},{"key":"e_1_3_3_38_2","volume-title":"AAAI Conference on Artificial Intelligence","author":"Leng Cong","year":"2017","unstructured":"Cong Leng, Hao Li, Shenghuo Zhu, and Rong Jin. 2017. Extremely low bit neural network: Squeeze the last bit out with ADMM. In AAAI Conference on Artificial Intelligence. AAAI. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:8840788"},{"doi-asserted-by":"publisher","unstructured":"Bin Liu Fengfu Li Xiaoxing Wang Bo Zhang and Junchi Yan. 2023. Ternary weight networks. In ICASSP 2023-2023 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP) 1\u20135. DOI:10.1109\/ICASSP49357.2023.10094626","key":"e_1_3_3_39_2","DOI":"10.1109\/ICASSP49357.2023.10094626"},{"doi-asserted-by":"publisher","key":"e_1_3_3_40_2","DOI":"10.1109\/ISSCC42615.2023.10067555"},{"doi-asserted-by":"publisher","key":"e_1_3_3_41_2","DOI":"10.1109\/EMC2-NIPS53020.2019.00017"},{"doi-asserted-by":"publisher","key":"e_1_3_3_42_2","DOI":"10.1109\/TC.2023.3285095"},{"doi-asserted-by":"publisher","key":"e_1_3_3_43_2","DOI":"10.1109\/Confluence52989.2022.9734142"},{"doi-asserted-by":"publisher","key":"e_1_3_3_44_2","DOI":"10.1109\/HCS49909.2020.9220641"},{"doi-asserted-by":"publisher","key":"e_1_3_3_45_2","DOI":"10.1109\/CVPR.2016.91"},{"doi-asserted-by":"publisher","key":"e_1_3_3_46_2","DOI":"10.1109\/CVPR.2018.00474"},{"doi-asserted-by":"publisher","key":"e_1_3_3_47_2","DOI":"10.1145\/196244.196430"},{"doi-asserted-by":"publisher","key":"e_1_3_3_48_2","DOI":"10.1109\/JSSC.2019.2952773"},{"doi-asserted-by":"publisher","key":"e_1_3_3_49_2","DOI":"10.1109\/TCAD.2022.3172600"},{"doi-asserted-by":"publisher","key":"e_1_3_3_50_2","DOI":"10.1109\/ISSCC42614.2022.9731762"},{"doi-asserted-by":"publisher","key":"e_1_3_3_51_2","DOI":"10.1109\/ISSCC42614.2022.9731645"},{"doi-asserted-by":"publisher","key":"e_1_3_3_52_2","DOI":"10.1109\/ISSCC42615.2023.10067526"},{"doi-asserted-by":"publisher","key":"e_1_3_3_53_2","DOI":"10.1109\/ISSCC42614.2022.9731545"},{"key":"e_1_3_3_54_2","doi-asserted-by":"crossref","first-page":"574","DOI":"10.1109\/ISSCC49657.2024.10454489","article-title":"34.5 A 818-4094TOPS\/W capacitor-reconfigured CIM macro for unified acceleration of CNNs and transformers","volume":"67","author":"Yoshioka Kentaro","year":"2024","unstructured":"Kentaro Yoshioka. 2024. 34.5 A 818-4094TOPS\/W capacitor-reconfigured CIM macro for unified acceleration of CNNs and transformers. 2024 IEEE International Solid-State Circuits Conference (ISSCC) 67 (2024), 574\u2013576. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:268375824","journal-title":"2024 IEEE International Solid-State Circuits Conference (ISSCC)"},{"doi-asserted-by":"publisher","key":"e_1_3_3_55_2","DOI":"10.1109\/ISSCC49657.2024.10454313"},{"doi-asserted-by":"publisher","key":"e_1_3_3_56_2","DOI":"10.1007\/978-3-030-01237-3_23"},{"doi-asserted-by":"publisher","key":"e_1_3_3_57_2","DOI":"10.1145\/3643134"},{"doi-asserted-by":"publisher","key":"e_1_3_3_58_2","DOI":"10.1145\/3632957"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3707205","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3707205","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:14Z","timestamp":1750295894000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3707205"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,10]]},"references-count":57,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3707205"],"URL":"https:\/\/doi.org\/10.1145\/3707205","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"type":"print","value":"1084-4309"},{"type":"electronic","value":"1557-7309"}],"subject":[],"published":{"date-parts":[[2025,1,10]]},"assertion":[{"value":"2024-06-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-11-26","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}