{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T16:55:22Z","timestamp":1781110522776,"version":"3.54.1"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"5s","license":[{"start":{"date-parts":[[2023,9,9]],"date-time":"2023-09-09T00:00:00Z","timestamp":1694217600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"EC H2020 WiPLASH","award":["863337"],"award-info":[{"award-number":["863337"]}]},{"name":"EC H2020 FVLLMONTI","award":["101016776"],"award-info":[{"award-number":["101016776"]}]},{"name":"ACCESS \u2013 AI Chip Center for Emerging Smart Systems"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2023,10,31]]},"abstract":"<jats:p>Compute memories are memory arrays augmented with dedicated logic to support arithmetic. They support the efficient execution of data-centric computing patterns, such as those characterizing Artificial Intelligence (AI) algorithms. These architectures can provide computing capabilities as part of the memory array structures (In-Memory Computing, IMC) or at their immediate periphery (Near-Memory Computing, NMC). By bringing the processing elements inside (or very close to) storage, compute memories minimize the cost of data access. Moreover, highly parallel (and, hence, high-performance) computations are enabled by exploiting the regular structure of memory arrays. However, the regular layout of memory elements also constrains the data range of inputs and outputs, since the bitwidths of operands and results stored at each address cannot be freely varied. Addressing this challenge, we herein propose a HW\/SW co-design methodology combining careful per-layer quantization and inter-layer scaling with lightweight hardware support for overflow-free computation of dot-vector operations. We demonstrate their use to implement the convolutional and fully connected layers of AI models. We embody our strategy in two implementations, based on IMC and NMC, respectively. Experimental results highlight that an area overhead of only 10.5% (for IMC) and 12.9% (for NMC) is required when interfacing with a 2KB subarray. Furthermore, inferences on benchmark CNNs show negligible accuracy degradation due to quantization for equivalent floating-point implementations.<\/jats:p>\n          <jats:p\/>","DOI":"10.1145\/3609387","type":"journal-article","created":{"date-parts":[[2023,9,9]],"date-time":"2023-09-09T13:33:18Z","timestamp":1694266398000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Overflow-free Compute Memories for Edge AI Acceleration"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9662-498X","authenticated-orcid":false,"given":"Flavio","family":"Ponzina","sequence":"first","affiliation":[{"name":"\u00c9cole Polytechnique F\u00e9d\u00e9rale de Lausanne (EPFL), Embedded Systems Laboratory, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8251-6390","authenticated-orcid":false,"given":"Marco","family":"Rios","sequence":"additional","affiliation":[{"name":"\u00c9cole Polytechnique F\u00e9d\u00e9rale de Lausanne (EPFL), Embedded Systems Laboratory, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8984-9793","authenticated-orcid":false,"given":"Alexandre","family":"Levisse","sequence":"additional","affiliation":[{"name":"\u00c9cole Polytechnique F\u00e9d\u00e9rale de Lausanne (EPFL), Embedded Systems Laboratory, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8940-3775","authenticated-orcid":false,"given":"Giovanni","family":"Ansaloni","sequence":"additional","affiliation":[{"name":"\u00c9cole Polytechnique F\u00e9d\u00e9rale de Lausanne (EPFL), Embedded Systems Laboratory, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9536-4947","authenticated-orcid":false,"given":"David","family":"Atienza","sequence":"additional","affiliation":[{"name":"\u00c9cole Polytechnique F\u00e9d\u00e9rale de Lausanne (EPFL), Embedded Systems Laboratory, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,9,9]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1109\/IC2E52221.2021.00016","volume-title":"2021 IEEE International Conference on Cloud Engineering (IC2E\u201921)","author":"Baller Stephan Patrick","year":"2021","unstructured":"Stephan Patrick Baller, Anshul Jindal, Mohak Chadha, and Michael Gerndt. 2021. DeepEdgeBench: Benchmarking deep neural networks on edge devices. In 2021 IEEE International Conference on Cloud Engineering (IC2E\u201921). IEEE, 20\u201330."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1080\/23746149.2016.1259585"},{"issue":"4","key":"e_1_3_2_4_2","first-page":"44","article-title":"An efficient CNN accelerator for low-cost edge systems","volume":"21","author":"Choi Kyubaik","year":"2022","unstructured":"Kyubaik Choi and Gerald E. Sobelman. 2022. An efficient CNN accelerator for low-cost edge systems. ACM Trans. Embed. Comput. Syst. 21, 4, Article 44 (aug2022), 20 pages.","journal-title":"ACM Trans. Embed. Comput. Syst."},{"issue":"8","key":"e_1_3_2_5_2","doi-asserted-by":"crossref","first-page":"675","DOI":"10.1038\/s42256-021-00356-5","article-title":"Automatic heterogeneous quantization of deep neural networks for low-latency inference on the edge for particle detectors","volume":"3","author":"Jr Claudionor N. Coelho","year":"2021","unstructured":"Claudionor N. Coelho Jr, Aki Kuusela, Shan Li, Hao Zhuang, Jennifer Ngadiuba, Thea Klaeboe Aarrestad, Vladimir Loncar, Maurizio Pierini, Adrian Alan Pol, and Sioni Summers. 2021. Automatic heterogeneous quantization of deep neural networks for low-latency inference on the edge for particle detectors. Nature Machine Intelligence 3, 8 (2021), 675\u2013686.","journal-title":"Nature Machine Intelligence"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1109\/HPCA.2015.7056040","volume-title":"2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA\u201915)","author":"Farmahini-Farahani Amin","year":"2015","unstructured":"Amin Farmahini-Farahani, Jung Ho Ahn, Katherine Morrow, and Nam Sung Kim. 2015. NDA: Near-DRAM acceleration architecture leveraging commodity DRAM devices and standard memory modules. In 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA\u201915). 283\u2013295."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2022.3174101"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC.2014.6757323"},{"key":"e_1_3_2_10_2","article-title":"Mobilenets: Efficient convolutional neural networks for mobile vision applications","author":"Howard Andrew G.","year":"2017","unstructured":"Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017).","journal-title":"arXiv preprint arXiv:1704.04861"},{"key":"e_1_3_2_11_2","article-title":"Binarized neural networks","volume":"29","author":"Hubara Itay","year":"2016","unstructured":"Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2016. Binarized neural networks. Advances in Neural Information Processing Systems 29 (2016).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00286"},{"key":"e_1_3_2_13_2","first-page":"1","volume-title":"2021 IEEE Globecom Workshops (GC Wkshps\u201921)","author":"Kamruzzaman M. M.","year":"2021","unstructured":"M. M. Kamruzzaman. 2021. New opportunities, challenges, and applications of edge-AI for connected healthcare in smart cities. In 2021 IEEE Globecom Workshops (GC Wkshps\u201921). IEEE, 1\u20136."},{"key":"e_1_3_2_14_2","series-title":"Proceedings of the 38th International Conference on Machine Learning","first-page":"5506","volume":"139","author":"Kim Sehoon","year":"2021","unstructured":"Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. 2021. I-BERT: Integer-only BERT quantization. In Proceedings of the 38th International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol. 139), Marina Meila and Tong Zhang (Eds.). PMLR, 5506\u20135518."},{"key":"e_1_3_2_15_2","article-title":"ALPINE: Analog in-memory acceleration with tight processor integration for deep learning","author":"Klein Joshua","year":"2022","unstructured":"Joshua Klein, Irem Boybat, Yasir Qureshi, Martino Dazzi, Alexandre Levisse, Giovanni Ansaloni, Marina Zapater, Abu Sebastian, and David Atienza. 2022. ALPINE: Analog in-memory acceleration with tight processor integration for deep learning. IEEE Trans. Comput. (2022).","journal-title":"IEEE Trans. Comput."},{"key":"e_1_3_2_16_2","unstructured":"Alex Krizhevsky Geoffrey Hinton et\u00a0al. 2009. Learning multiple layers of features from tiny images. (2009)."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3065386"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","first-page":"350","DOI":"10.1109\/ISSCC42613.2021.9365862","volume-title":"2021 IEEE International Solid-State Circuits Conference (ISSCC\u201921)","volume":"64","author":"Kwon Young-Cheon","year":"2021","unstructured":"Young-Cheon Kwon, Suk Han Lee, Jaehoon Lee, Sang-Hyuk Kwon, Je Min Ryu, Jong-Pil Son, O. Seongil, Hak-Soo Yu, Haesuk Lee, Soo Young Kim, et\u00a0al. 2021. 25.4 a 20nm 6gb function-in-memory dram, based on hbm2 with a 1.2 tflops programmable computing unit using bank-level parallelism, for machine learning applications. In 2021 IEEE International Solid-State Circuits Conference (ISSCC\u201921), Vol. 64. IEEE, 350\u2013352."},{"key":"e_1_3_2_19_2","first-page":"1","volume-title":"2020 57th ACM\/IEEE Design Automation Conference (DAC\u201920)","author":"Lee Kyeongho","year":"2020","unstructured":"Kyeongho Lee, Jinho Jeong, Sungsoo Cheon, Woong Choi, and Jongsun Park. 2020. Bit parallel 6T SRAM in-memory computing with reconfigurable bit-precision. In 2020 57th ACM\/IEEE Design Automation Conference (DAC\u201920). IEEE, 1\u20136."},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","DOI":"10.1109\/JIOT.2022.3176400","article-title":"A survey on the convergence of edge computing and AI for UAVs: Opportunities and challenges","author":"McEnroe Patrick","year":"2022","unstructured":"Patrick McEnroe, Shen Wang, and Madhusanka Liyanage. 2022. A survey on the convergence of edge computing and AI for UAVs: Opportunities and challenges. IEEE Internet of Things Journal (2022).","journal-title":"IEEE Internet of Things Journal"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.3390\/mi13071143"},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","first-page":"164","DOI":"10.1109\/ISVLSI51109.2021.00039","volume-title":"2021 IEEE Computer Society Annual Symposium on VLSI (ISVLSI\u201921)","author":"Ponzina Flavio","year":"2021","unstructured":"Flavio Ponzina, Marco Rios, Giovanni Ansaloni, Alexandre Levisse, and David Atienza. 2021. A flexible in-memory computing architecture for heterogeneously quantized CNNs. In 2021 IEEE Computer Society Annual Symposium on VLSI (ISVLSI\u201921). IEEE, 164\u2013169."},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","first-page":"88","DOI":"10.1109\/MICRO50266.2020.00020","volume-title":"2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201920)","author":"Ramanathan Akshay Krishna","year":"2020","unstructured":"Akshay Krishna Ramanathan, Gurpreet S. Kalsi, Srivatsa Srinivasa, Tarun Makesh Chandran, Kamlesh R. Pillai, Om J. Omer, Vijaykrishnan Narayanan, and Sreenivas Subramoney. 2020. Look-Up table based energy efficient processing in cache support for neural network acceleration. In 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201920). 88\u2013101."},{"key":"e_1_3_2_24_2","first-page":"1","volume-title":"Proceedings of the 55th Annual Design Automation Conference","author":"Reagen Brandon","year":"2018","unstructured":"Brandon Reagen, Udit Gupta, Lillian Pentecost, Paul Whatmough, Sae Kyu Lee, Niamh Mulholland, David Brooks, and Gu-Yeon Wei. 2018. Ares: A framework for quantifying the resilience of deep neural networks. In Proceedings of the 55th Annual Design Automation Conference. 1\u20136."},{"key":"e_1_3_2_25_2","first-page":"1","volume-title":"2021 IEEE High Performance Extreme Computing Conference (HPEC\u201921)","author":"Reuther Albert","year":"2021","unstructured":"Albert Reuther, Peter Michaleas, Michael Jones, Vijay Gadepally, Siddharth Samsi, and Jeremy Kepner. 2021. AI accelerator survey and trends. In 2021 IEEE High Performance Extreme Computing Conference (HPEC\u201921). IEEE, 1\u20139."},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3526241.3530351"},{"issue":"01","key":"e_1_3_2_27_2","first-page":"1","article-title":"Bit-line computing for CNN accelerators co-design in edge AI inference","author":"Rios M.","year":"2023","unstructured":"M. Rios, F. Ponzina, A. Levisse, G. Ansaloni, and D. Atienza. 2023. Bit-line computing for CNN accelerators co-design in edge AI inference. IEEE Transactions on Emerging Topics in Computing 01 (2023), 1\u201314.","journal-title":"IEEE Transactions on Emerging Topics in Computing"},{"key":"e_1_3_2_28_2","first-page":"34","volume-title":"2019 IFIP\/IEEE 27th International Conference on Very Large Scale Integration (VLSI-SoC\u201919)","author":"Rios Marco","year":"2019","unstructured":"Marco Rios, William Simon, Alexandre Levisse, Marina Zapater, and David Atienza. 2019. An associativity-agnostic in-cache computing architecture optimized for multiplication. In 2019 IFIP\/IEEE 27th International Conference on Very Large Scale Integration (VLSI-SoC\u201919). IEEE, 34\u201339."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41565-020-0655-z"},{"key":"e_1_3_2_30_2","doi-asserted-by":"crossref","first-page":"273","DOI":"10.1145\/3123939.3124544","volume-title":"2017 50th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201917)","author":"Seshadri Vivek","year":"2017","unstructured":"Vivek Seshadri, Donghyuk Lee, Thomas Mullins, Hasan Hassan, Amirali Boroumand, Jeremie Kim, Michael A. Kozuch, Onur Mutlu, Phillip B. Gibbons, and Todd C. Mowry. 2017. Ambit: In-memory accelerator for bulk bitwise operations using commodity DRAM technology. In 2017 50th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201917). 273\u2013287."},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","first-page":"295","DOI":"10.1145\/2989081.2989087","volume-title":"Proceedings of the Second International Symposium on Memory Systems","author":"Siegl Patrick","year":"2016","unstructured":"Patrick Siegl, Rainer Buchty, and Mladen Berekovic. 2016. Data-centric computing frontiers: A survey on processing-in-memory. In Proceedings of the Second International Symposium on Memory Systems. 295\u2013308."},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2020.2972528"},{"key":"e_1_3_2_33_2","article-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014).","journal-title":"arXiv preprint arXiv:1409.1556"},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","first-page":"547","DOI":"10.1109\/ISQED51717.2021.9424263","volume-title":"2021 22nd International Symposium on Quality Electronic Design (ISQED\u201921)","author":"Srinivasa Srivatsa","year":"2021","unstructured":"Srivatsa Srinivasa, Akshay Krishna Ramanathan, Jainaveen Sundaram, Dileep Kurian, Srinivasan Gopal, Nilesh Jain, Anuradha Srinivasan, Ravi Iyer, Vijaykrishnan Narayanan, and Tanay Karnik. 2021. Trends and opportunities for SRAM based in-memory and near-memory computation. In 2021 22nd International Symposium on Quality Electronic Design (ISQED\u201921). 547\u2013552."},{"key":"e_1_3_2_35_2","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1109\/PEHC54839.2021.00010","volume-title":"2021 IEEE\/ACM Programming Environments for Heterogeneous Computing (PEHC\u201921)","author":"Sukumar Sreenivas R.","year":"2021","unstructured":"Sreenivas R. Sukumar, Jacob A. Balma, Cong Xu, and Sergey Serebryakov. 2021. Survival of the fittest amidst the cambrian explosion of processor architectures for artificial intelligence. In 2021 IEEE\/ACM Programming Environments for Heterogeneous Computing (PEHC\u201921). IEEE, 34\u201343."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3135690"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSSC.2019.2922889"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.521"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41563-019-0291-x"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.634"},{"issue":"3","key":"e_1_3_2_42_2","article-title":"Auto-tuning fixed-point precision with TVM on RISC-V packed SIMD extension","volume":"28","author":"Yang Chun-Chieh","year":"2023","unstructured":"Chun-Chieh Yang, Yi-Ru Chen, Hui-Hsin Liao, Yuan-Ming Chang, and Jenq-Kuen Lee. 2023. Auto-tuning fixed-point precision with TVM on RISC-V packed SIMD extension. ACM Trans. Des. Autom. Electron. Syst. 28, 3 (2023).","journal-title":"ACM Trans. Des. Autom. Electron. Syst."},{"key":"e_1_3_2_43_2","first-page":"1","volume-title":"2019 IEEE\/ACM International Symposium on Low Power Electronics and Design (ISLPED\u201919)","author":"Yoo Taegeun","year":"2019","unstructured":"Taegeun Yoo, Hyunjoon Kim, Qian Chen, Tony Tae-Hyoung Kim, and Bongjin Kim. 2019. A logic compatible 4T dual embedded DRAM array for in-memory computation of deep neural networks. In 2019 IEEE\/ACM International Symposium on Low Power Electronics and Design (ISLPED\u201919). 1\u20136."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2017.2776302"},{"key":"e_1_3_2_45_2","doi-asserted-by":"crossref","first-page":"217","DOI":"10.1109\/A-SSCC47793.2019.9056933","volume-title":"2019 IEEE Asian Solid-State Circuits Conference (A-SSCC\u201919)","author":"Zhang Zhixiao","year":"2019","unstructured":"Zhixiao Zhang, Jia-Jing Chen, Xin Si, Yung-Ning Tu, Jian-Wei Su, Wei-Hsing Huang, Jing-Hong Wang, Wei-Chen Wei, Yen-Cheng Chiu, Je-Min Hong, et\u00a0al. 2019. A 55nm 1-to-8 bit configurable 6T SRAM based computing-in-memory unit-macro for CNN-based AI edge processors. In 2019 IEEE Asian Solid-State Circuits Conference (A-SSCC\u201919). IEEE, 217\u2013218."}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3609387","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3609387","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:23Z","timestamp":1750178783000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3609387"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,9]]},"references-count":44,"journal-issue":{"issue":"5s","published-print":{"date-parts":[[2023,10,31]]}},"alternative-id":["10.1145\/3609387"],"URL":"https:\/\/doi.org\/10.1145\/3609387","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,9]]},"assertion":[{"value":"2023-03-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-13","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}