{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,8]],"date-time":"2025-11-08T13:59:31Z","timestamp":1762610371350,"version":"build-2065373602"},"reference-count":37,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2026,1,31]]},"abstract":"<jats:p>SRAM is widely used in computing-in-memory (CIM) neural network accelerators because of its relatively mature technology and good compatibility with complementary metal oxide semiconductor logic process. Digital SRAM-CIM is favored by researchers because of its stability and accuracy. However, the current digital SRAM-CIM macro only supports the weight-stationary dataflows, which means the repeated movement of graph data. Some special deep neural network layers, such as depth-wise, make the utilization of computing resources inside CIM low.<\/jats:p>\n                  <jats:p>To overcome these problems, we propose C-CIM, which can switch between input-stationary and weight-stationary dataflows and support matrix multiplication as well as convolution operations with multiple mainstream convolution kernel sizes (1\u00d71, 3\u00d73, 5\u00d75 and 7\u00d77). The C-CIM achieves an average performance of 27.31TOPS\/W@8b at a frequency of 1GHz. Experimental results show that our proposed SRAM-CIM successfully outperforms baseline in terms of performance optimization, achieving up to 7.6\u00d7 performance speedup and up to 86.84% reduction in activation relocation.<\/jats:p>","DOI":"10.1145\/3769859","type":"journal-article","created":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T11:30:11Z","timestamp":1760095811000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["C-CIM: A Multi-Mode\n                    <u>C<\/u>\n                    onvolution-Capable SRAM-\n                    <u>CIM<\/u>"],"prefix":"10.1145","volume":"31","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-5632-2600","authenticated-orcid":false,"given":"Renyu","family":"Yang","sequence":"first","affiliation":[{"name":"Key Laboratory of Advanced Microprocessor Chips and Systems, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-0875-196X","authenticated-orcid":false,"given":"Xin","family":"Ju","sequence":"additional","affiliation":[{"name":"Key Laboratory of Advanced Microprocessor Chips and Systems, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5875-3297","authenticated-orcid":false,"given":"Mei","family":"Wen","sequence":"additional","affiliation":[{"name":"Key Laboratory of Advanced Microprocessor Chips and Systems, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-9896-8862","authenticated-orcid":false,"given":"Jinjin","family":"Deng","sequence":"additional","affiliation":[{"name":"Key Laboratory of Advanced Microprocessor Chips and Systems, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-1772-4883","authenticated-orcid":false,"given":"Yi","family":"Wen","sequence":"additional","affiliation":[{"name":"Key Laboratory of Advanced Microprocessor Chips and Systems, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6233-6800","authenticated-orcid":false,"given":"Junzhong","family":"Shen","sequence":"additional","affiliation":[{"name":"Key Laboratory of Advanced Microprocessor Chips and Systems, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7592-5484","authenticated-orcid":false,"given":"Bin","family":"Liang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Advanced Microprocessor Chips and Systems, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8345-3608","authenticated-orcid":false,"given":"Tianyu","family":"Wang","sequence":"additional","affiliation":[{"name":"Shenzhen University","place":["Shenzhen, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9526-6634","authenticated-orcid":false,"given":"Zhaoyan","family":"Shen","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Shandong University","place":["Qingdao, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2173-2847","authenticated-orcid":false,"given":"Zili","family":"Shao","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Chinese University of Hong Kong","place":["New Territories, Hong Kong"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,11,8]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Kumar Chellapilla Sidd Puri and Patrice Y. Simard. 2006. High performance convolutional neural networks for document processing. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:14936779"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC42613.2021.9365766"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","unstructured":"Hyungmin Cho. 2021. RiSA: A Reinforced systolic array for depthwise convolutions and embedded tensor reshaping. ACM Trans. Embed. Comput. Syst. 20 5s (September 2021). DOI:10.1145\/3476984","DOI":"10.1145\/3476984"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","unstructured":"Jack Choquette. 2023. NVIDIA hopper H100 GPU: Scaling performance. IEEE Micro 43 3 (2023) 9\u201317. DOI:10.1109\/MM.2023.3256796","DOI":"10.1109\/MM.2023.3256796"},{"key":"e_1_3_1_6_2","volume-title":"ISSCC","author":"Y.D. C.","year":"2021","unstructured":"C. Y.D., et\u00a0al. 2021. 16.4 An 89TOPS\/W and 16.3TOPS\/mm2 All-Digital SRAM-based full-precision compute-in memory macro in 22nm for machine-learning edge applications. In ISSCC."},{"key":"e_1_3_1_7_2","volume-title":"ISSCC","author":"Hidehiro F.","year":"2022","unstructured":"F. Hidehiro, et\u00a0al. 2022. A 5-nm 254-TOPS\/W 221-TOPS\/mm2 fully-digital computing-in-memory macro supporting wide-range dynamic-voltage-frequency scaling and simultaneous MAC and write operations. In ISSCC."},{"key":"e_1_3_1_8_2","article-title":"Adapt-Flow: A flexible DNN accelerator architecture for heterogeneous dataflow implementation","author":"Yang J.","year":"2022","unstructured":"J. Yang, et\u00a0al. 2022. Adapt-Flow: A flexible DNN accelerator architecture for heterogeneous dataflow implementation. VLSI (2022).","journal-title":"VLSI"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","unstructured":"Hyunjoon Kim Taegeun Yoo Tony Tae-Hyoung Kim and Bongjin Kim. 2021. Colonnade: A reconfigurable SRAM-based digital bit-serial compute-In-memory macro for processing neural networks. IEEE Journal of Solid-State Circuits 56 7 (2021) 2221\u20132233. DOI:10.1109\/JSSC.2021.3061508","DOI":"10.1109\/JSSC.2021.3061508"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","unstructured":"Y. Lecun L. Bottou Y. Bengio and P. Haffner. 1998. Gradient-based learning applied to document recognition. Proceedings of the IEEE 86 11 (1998) 2278\u20132324. DOI:10.1109\/5.726791","DOI":"10.1109\/5.726791"},{"key":"e_1_3_1_11_2","volume-title":"DAC","author":"YC. L.","year":"2023","unstructured":"L. YC., et\u00a0al. 2023. Morphable CIM: Improving operation intensity and depthwise capability for SRAM-CIM architecture. In DAC."},{"key":"e_1_3_1_12_2","volume-title":"ICCD","author":"Xu. R.","year":"2020","unstructured":"R. Xu., et\u00a0al. 2020. CMSA: Configurable multi-directional systolic array for convolutional neural networks. In ICCD."},{"key":"e_1_3_1_13_2","volume-title":"ISSCC","author":"Bo. W.","year":"2023","unstructured":"W. Bo., et\u00a0al. 2023. A 28nm horizontal-weight-shift and vertical-feature-shift-based separate-WL 6T-SRAM computation-in-memory unit-macro for edge depthwise neural-networks. In ISSCC."},{"key":"e_1_3_1_14_2","volume-title":"ICCAD","author":"Qilin Z.","year":"2020","unstructured":"Z. Qilin, et\u00a0al. 2020. MobiLattice: A depth-wise DCNN accelerator with hybrid digital\/analog nonvolatile processing-in-memory block. In ICCAD."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC42613.2021.9366000"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.3390\/electronics11060945"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2024.3375251"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_3_1_19_2","unstructured":"Andrew G. Howard Menglong Zhu Bo Chen Dmitry Kalenichenko Weijun Wang Tobias Weyand Marco Andreetto and Hartwig Adam. 2017. MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv:1704.04861. Retrieved from https:\/\/arxiv.org\/abs\/\/1704.04861. https:\/\/api.semanticscholar.org\/CorpusID:12670695"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/NVMSA63038.2024.10693675"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10070992"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2021.3061521"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3583781.3590306"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2023.3311948"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD63220.2024.00078"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18072.2020.9218724"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAD51958.2021.9643497"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC19947.2020.9062995"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2019.00027"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC56929.2023.10247928"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41928-023-01053-4"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC42614.2022.9731545"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC42614.2022.9731545"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD56317.2022.00068"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2024.3522879"},{"key":"e_1_3_1_37_2","first-page":"1","volume-title":"2020 IEEE\/ACM International Conference On Computer Aided Design (ICCAD)","author":"Zheng Qilin","year":"2020","unstructured":"Qilin Zheng, Xingchen Li, Zongwei Wang, Guangyu Sun, Yimao Cai, Ru Huang, Yiran Chen, and Hai Li. 2020. MobiLattice: A depth-wise DCNN accelerator with hybrid digital\/analog nonvolatile processing-in-memory block. In 2020 IEEE\/ACM International Conference On Computer Aided Design (ICCAD). 1\u20139."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA53966.2022.00042"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3769859","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,8]],"date-time":"2025-11-08T13:56:34Z","timestamp":1762610194000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3769859"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,8]]},"references-count":37,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1,31]]}},"alternative-id":["10.1145\/3769859"],"URL":"https:\/\/doi.org\/10.1145\/3769859","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"type":"print","value":"1084-4309"},{"type":"electronic","value":"1557-7309"}],"subject":[],"published":{"date-parts":[[2025,11,8]]},"assertion":[{"value":"2024-10-06","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}