{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T13:47:36Z","timestamp":1782481656317,"version":"3.54.5"},"reference-count":19,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T00:00:00Z","timestamp":1782432000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>The growing demand for high-performance, energy-efficient execution of data-parallel workloads has driven the resurgence of vector processors, yet their expanding instruction sets exacerbate the area-performance tradeoff of vector processing units (VPUs). Processing-in-memory (PIM) technique offers a promising path to mitigate this tradeoff by offloading vector operations near data. However, integrating the computing-capable SRAM (C-SRAM) with conventional VPUs introduces significant architectural challenges, including inefficient coordination between heterogeneous devices, the lack of a unified hardware\/software interface, and underutilized parallelism within the C-SRAM arrays due to unoptimized data handling.<\/jats:p>\n                  <jats:p>To address these challenges, this article proposes a heterogeneous VPU (HVPU) that seamlessly integrates a standard VPU inside a vector processor with a C-SRAM for more efficient vector processing. HVPU introduces a standardized interface between the processor frontend and the C-SRAM, enabling instruction dispatch, dynamic hazard resolution, and concurrent execution. Furthermore, it employs multiple independent PIM blocks coupled with dual controlling pipelines inside the C-SRAM for further performance improvement. This design facilitates pipelined data loading and computing, effectively hiding memory latency and fully unlocking the parallel potential of C-SRAM. The system is supported by a user-friendly and generic programming model featuring a two-layer extended ISA system and a vector batch pipelining mechanism.<\/jats:p>\n                  <jats:p>Experimental results show that our HVPU-enhanced processor achieves significant speedups of 5.11\u00d7 to 39.0\u00d7 over the Xuantie-910 baseline on vector benchmarks, respectively, while reducing energy consumption by 75% on average. Meanwhile, it outperforms state-of-the-art PIM accelerators by 1.33\u00d7 to 1.97\u00d7 on the same benchmarks, with minimal area overhead. This demonstrates that our architectural co-design effectively alleviates the area-performance tradeoff in vector processors and offers a scalable heterogeneous architecture template for efficient vector processing.<\/jats:p>","DOI":"10.1145\/3817062","type":"journal-article","created":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T11:31:03Z","timestamp":1780054263000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Enabling Efficient Vector Processing: A Heterogeneous Vector Architecture with in-SRAM Computing"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-5154-0086","authenticated-orcid":false,"given":"Ruoxi","family":"Wang","sequence":"first","affiliation":[{"name":"College of Computer Science and Technology, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2384-5359","authenticated-orcid":false,"given":"Zhang","family":"Dunbo","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7217-1712","authenticated-orcid":false,"given":"Shangshang","family":"Yao","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-9456-8695","authenticated-orcid":false,"given":"Lang","family":"Qingjie","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-5299-7890","authenticated-orcid":false,"given":"Junyi","family":"Zhu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9043-2998","authenticated-orcid":false,"given":"Li","family":"Shen","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, National University of Defense Technology","place":["Changsha, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,26]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.21"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750385"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10071074"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00016"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/1950413.1950420"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00040"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613424.3614268"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322257"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","unstructured":"Nastaran Hajinazar Geraldo F. Oliveira Sven Gregorio Jo\u00e3o Dinis Ferreira Nika Mansouri Ghiasi Minesh Patel Mohammed Alser Saugata Ghose Juan G\u00f3mez-Luna and Onur Mutlu. 2021. SIMDRAM: A framework for bit-serial SIMD processing using DRAM. In Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS\u201921). Association for Computing Machinery New York NY USA 329\u2013345. DOI:10.1145\/3445814.3446749","DOI":"10.1145\/3445814.3446749"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3743136"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00071"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3690824"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3762641"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA57654.2024.00024"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC42613.2021.9366056"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123939.3124544"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11390-025-4555-4"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3631528"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.7544\/issn1000-1239.202330151"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3817062","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T12:57:44Z","timestamp":1782478664000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3817062"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,26]]},"references-count":19,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3817062"],"URL":"https:\/\/doi.org\/10.1145\/3817062","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,26]]},"assertion":[{"value":"2026-02-09","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-16","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-26","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}