{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T06:27:30Z","timestamp":1780468050178,"version":"3.54.1"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"5s","funder":[{"name":"Swiss NSF Edge-Companions","award":["10002812"],"award-info":[{"award-number":["10002812"]}]},{"name":"EC H2020 FVLLMONTI","award":["101016776"],"award-info":[{"award-number":["101016776"]}]},{"name":"Swiss State Secretariat for Education, Research, and Innovation"},{"name":"ACCESS\u2014AI Chip Center for Emerging Smart Systems"},{"name":"InnoHK initiative of the Innovation and Technology Commission of the Hong Kong Special Administrative Region Government"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>\n            By interfacing computing logic directly to the DRAM banks, bank-level Compute-near-Memory (CnM) architectures promise to mitigate the bottleneck at the memory interconnect. While this computation paradigm heavily reduces the energy requirements for data movement across the system, current solutions fail to co-optimize hardware and software to further increase efficiency. Instead, in this manuscript, we present\n            <jats:bold>SideDRAM<\/jats:bold>\n            , a co-designed bank-level CnM architecture to enable massively parallel and energy-efficient computations near DRAM. In contrast with past solutions, we support flexible data typing and heterogeneous quantization, relying on the robustness of workloads to employ small bitwidths, and enable a row-wide access to the banks to exploit parallelism and spatial locality. As a result, SideDRAM integrates (1) software-defined SIMD (SoftSIMD) datapaths, supporting low-energy computing with flexible precision, (2) an interface to the banks based on very wide registers (VWRs), enabling asymmetric data access to both utilize the full DRAM bank bandwidth and leverage data locality at the datapath, and (3) a low-overhead distributed control plane, allowing the efficient handling of variable data typing. We benchmark SideDRAM as a near-DRAM solution by analyzing the area, performance, and energy consumption of an HBM2 CnM channel executing heterogeneously quantized machine learning models. The results show that, compared to the state-of-the-art FIMDRAM design, energy improvements of up to 67% are achieved when a DeiT-S inference is executed with a batch size of 16 under the same area constraints, resulting in energy-delay-area product (EDAP) savings that reach 83%. When comparing to a massively parallel mixed-signal CnM solution, SideDRAM consistently obtains similar performance and better energy efficiency results (geomean of 15\u00d7 improvement across workloads) at a lower area overhead.\n          <\/jats:p>","DOI":"10.1145\/3762641","type":"journal-article","created":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T11:51:43Z","timestamp":1755863503000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["SideDRAM: Integrating SoftSIMD Datapaths near DRAM Banks for Energy-Efficient Variable Precision Computation"],"prefix":"10.1145","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1349-5351","authenticated-orcid":false,"given":"Rafael","family":"Medina Morillas","sequence":"first","affiliation":[{"name":"Embedded Systems Laboratory (ESL), EPFL","place":["Lausanne, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-8221-8608","authenticated-orcid":false,"given":"Pengbo","family":"Yu","sequence":"additional","affiliation":[{"name":"Embedded Systems Laboratory (ESL), EPFL","place":["Lausanne, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8984-9793","authenticated-orcid":false,"given":"Alexandre","family":"Levisse","sequence":"additional","affiliation":[{"name":"Embedded Systems Laboratory (ESL), EPFL","place":["Lausanne, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1087-3433","authenticated-orcid":false,"given":"Dwaipayan","family":"Biswas","sequence":"additional","affiliation":[{"name":"IMEC","place":["Leuven, Belgium"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6971-1965","authenticated-orcid":false,"given":"Marina","family":"Zapater","sequence":"additional","affiliation":[{"name":"REDS Institute, HEIG-VD","place":["Yverdon-les-Bains, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8940-3775","authenticated-orcid":false,"given":"Giovanni","family":"Ansaloni","sequence":"additional","affiliation":[{"name":"Embedded Systems Laboratory (ESL), EPFL","place":["Lausanne, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3599-8515","authenticated-orcid":false,"given":"Francky","family":"Catthoor","sequence":"additional","affiliation":[{"name":"Microlab, NTUA","place":["Zografou, Greece"]},{"name":"IMEC","place":["Zografou, Greece"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9536-4947","authenticated-orcid":false,"given":"David","family":"Atienza","sequence":"additional","affiliation":[{"name":"Embedded Systems Laboratory (ESL), EPFL","place":["Lausanne, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,26]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3566097.3567867"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TEC.1961.5219227"},{"key":"e_1_3_3_4_2","unstructured":"Mart van Baalen Andrey Kuzmin Suparna S. Nair Yuwei Ren Eric Mahurin Chirag Patel Sundar Subramanian Sanghyuk Lee Markus Nagel Joseph Soriaga et\u00a0al. 2023. FP8 versus INT8 for efficient deep learning inference. arXiv:2303.17951. Retrieved from https:\/\/arxiv.org\/abs\/2303.17951"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2024.3367822"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3665314.3670806"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-90-481-9528-2"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","unstructured":"Seunghwan Cho Haerang Choi Eunhyeok Park Hyunsung Shin and Sungjoo Yoo. 2020. McDRAM v2: In-dynamic random access memory systolic array accelerator to address the large model problem in deep neural networks on the edge. IEEE Access 8 (2020) 135223\u2013135243. DOI:10.1109\/ACCESS.2020.3011265","DOI":"10.1109\/ACCESS.2020.3011265"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2022.3167391"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3489517.3530980"},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/HOTCHIPS.2019.8875680"},{"key":"e_1_3_3_12_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 4171\u20134186."},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1201\/9781003162810-13"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2024.3373763"},{"key":"e_1_3_3_15_2","unstructured":"Rishabh Goyal Joaquin Vanschoren Victor Van Acht and Stephan Nijssen. 2021. Fixed-point quantization of convolutional neural networks for quantized inference on embedded platforms. arXiv:2102.02147. Retrieved from https:\/\/arxiv.org\/abs\/2102.02147"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446749"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO50266.2020.00040"},{"key":"e_1_3_3_19_2","unstructured":"Andrew G. Howard Menglong Zhu Bo Chen Dmitry Kalenichenko Weijun Wang Tobias Weyand Marco Andreetto and Hartwig Adam. 2017. MobileNets: Efficient convolutional neural networks for mobile vision applications. Retrieved from https:\/\/arxiv.org\/abs\/1704.04861"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-01746-9"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2005.92"},{"key":"e_1_3_3_22_2","unstructured":"JEDEC. 2021. High bandwidth memory (HBM) DRAM. JESD235D. https:\/\/www.jedec.org\/standards-documents\/docs\/jesd235a"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2021.3097700"},{"key":"e_1_3_3_25_2","unstructured":"Asif Ali Khan Jo\u00e3o Paulo C. De Lima Hamid Farzaneh and Jeronimo Castrillon. 2024. The Landscape of Compute-near-memory and Compute-in-memory: A Research and Commercial Overview. arXiv:2401.14428. Retrieved from https:\/\/arxiv.org\/abs\/2401.14428"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/HCS59251.2023.10254711"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","unstructured":"Yoongu Kim Weikun Yang and Onur Mutlu. 2016. Ramulator: A fast and extensible DRAM simulator. IEEE Computer Architecture Letters 15 1 (2016) 45\u201349. DOI:10.1109\/LCA.2015.2414456","DOI":"10.1109\/LCA.2015.2414456"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/DATE.2007.364485"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/HCS55958.2022.9895629"},{"key":"e_1_3_3_30_2","volume-title":"Proceedings of the ISCA","author":"Lee Sukhan","year":"2021","unstructured":"Sukhan Lee, Shin-haeng Kang, Jaehoon Lee, Hyeonsu Kim, Eojin Lee, Seungwoo Seo, Hosang Yoon, Seungwon Lee, Kyounghwan Lim, Hyunsung Shin, Jinhyun Kim, Seongil O, Anand Iyer, David Wang, Kyomin Sohn, and Nam Sung Kim. 2021. Hardware architecture and software stack for PIM based on commercial DRAM technology. In Proceedings of the ISCA."},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2018.00062"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123939.3123977"},{"key":"e_1_3_3_33_2","unstructured":"Ji Lin Jiaming Tang Haotian Tang Shang Yang Wei-Ming Chen Wei-Chen Wang Guangxuan Xiao Xingyu Dang Chuang Gan and Song Han. 2024. AWQ: Activation-aware weight quantization for on-device llm compression and acceleration. In Proceedings of Machine Learning and Systems 6 (2024) 87\u2013100. Retrieved from https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2024\/file\/42a452cbafa9dd64e9ba4aa95cc1ef21-Paper-Conference.pdf"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3579371.3589101"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2024.3442989"},{"key":"e_1_3_3_36_2","unstructured":"Markus Nagel Marios Fournarakis Rana Ali Amjad Yelysei Bondarenko Mart Van Baalen and Tijmen Blankevoort. 2021. A white paper on neural network quantization. arXiv:2106.08295. Retrieved from https:\/\/arxiv.org\/abs\/2106.08295"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA57654.2024.00024"},{"key":"e_1_3_3_38_2","volume-title":"Binary Arithmetic for Finite-Word-Length Linear Controllers: MEMS Applications","author":"Oudjida Abdelkrim Kamel","year":"2014","unstructured":"Abdelkrim Kamel Oudjida. 2014. Binary Arithmetic for Finite-Word-Length Linear Controllers: MEMS Applications. Ph.D. Dissertation. L\u2019Universit\u00e9 de Franche-Compt\u00e9."},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","unstructured":"Sang-Soo Park KyungSoo Kim Jinin So Jin Jung Jonggeon Lee Kyoungwan Woo Nayeon Kim Younghyun Lee Hyungyo Kim Yongsuk Kwon Jinhyun Kim Jieun Lee YeonGon Cho Yongmin Tai Jeonghyeon Cho Hoyoung Song Jung Ho Ahn and Nam Sung Kim. 2024. An LPDDR-based CXL-PNM platform for TCO-efficient inference of transformer-based large language models. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 970\u2013982. DOI:10.1109\/HPCA57654.2024.00078","DOI":"10.1109\/HPCA57654.2024.00078"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW60793.2023.00138"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/DATE.2007.364435"},{"key":"e_1_3_3_42_2","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv:1409.1556. Retrieved from https:\/\/arxiv.org\/abs\/1409.1556"},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/VLSI-SoC62099.2024.10767834"},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2022.3172774"},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-15074-6_23"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394885.3431522"},{"key":"e_1_3_3_47_2","unstructured":"Yu-Shan Tai An-Yeu and Wu. 2024. MPTQ-ViT: Mixed-Precision Post-Training Quantization for Vision Transformer. Retrieved from https:\/\/arxiv.org\/abs\/2401.14895"},{"key":"e_1_3_3_48_2","first-page":"10347","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Touvron Hugo","year":"2021","unstructured":"Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv\u00e9 J\u00e9gou. 2021. Training data-efficient image transformers and distillation through attention. In Proceedings of the International Conference on Machine Learning. PMLR, 10347\u201310357."},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2024.3357597"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00006"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/216585.216588"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA61900.2025.00054"},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2024.3375793"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3762641","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T13:44:36Z","timestamp":1759239876000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3762641"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,26]]},"references-count":52,"journal-issue":{"issue":"5s","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3762641"],"URL":"https:\/\/doi.org\/10.1145\/3762641","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,26]]},"assertion":[{"value":"2025-08-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-07","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-26","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}