{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T16:46:36Z","timestamp":1782405996496,"version":"3.54.5"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    Transformers have revolutionized AI in natural language processing and computer vision, but their enormous computation and memory demands pose significant challenges for hardware acceleration. In practice, end-to-end throughput is often limited by paged data movement and interconnect bandwidth, not just raw MAC count. This work proposes a unified\n                    <jats:bold>system\u2013accelerator co-design<\/jats:bold>\n                    approach to efficiently accelerate transformer inference by jointly optimizing a novel hardware matrix accelerator and its system integration with paged, streaming dataflows and explicit overlap of compute and transfer. On the hardware side, we introduce\n                    <jats:bold>MatrixFlow<\/jats:bold>\n                    , a loosely\u2010coupled 16\u00d7 16 systolic\u2010array accelerator featuring a\n                    <jats:bold>block\u2010based<\/jats:bold>\n                    matrix multiplication method that is page-aligned (4 KB tiles), uses only a small (\u224820 KB) on-chip buffer, and runs a pipelined schedule of DMA, compute, and DMA-out to fully utilize interconnect bandwidth, emphasizing standard DMA-driven streaming rather than large on-chip reuse. On the system side, we develop\n                    <jats:bold>Gem5\u2010AcceSys<\/jats:bold>\n                    , an extension of the gem5 full system simulator allowing exploration of\n                    <jats:bold>standard interconnects<\/jats:bold>\n                    (PCIe) and\n                    <jats:bold>configurable memory hierarchies<\/jats:bold>\n                    including Direct-Memory (DM), Direct-Cache (DC), and Device-Memory (DevMem) modes with SMMU\/TLB effects. Through co\u2010design, MatrixFlow\u2019s novel dataflow and the Gem5\u2010AcceSys platform are tuned in tandem to alleviate data\u2010movement bottlenecks without requiring specialized CPU\u2010instruction\u2010set modifications. We validate our approach with gem5 simulations on representative transformer models (BERT and ViT) across multiple data types and system setups.\n                    <jats:bold>Results<\/jats:bold>\n                    demonstrate up to\n                    <jats:bold>22\u00d7 speed\u2010up<\/jats:bold>\n                    in end\u2010to\u2010end inference over a CPU\u2010only baseline and performance gains of\n                    <jats:bold>5\u00d7\u20138\u00d7<\/jats:bold>\n                    over state\u2010of\u2010the\u2010art loosely\u2010 and tightly\u2010coupled accelerators. Furthermore, we show that a standard PCIe-based host memory design can achieve \u223c80% of the performance of on\u2010device HBM memory. Overall,\n                    <jats:bold>paged streaming and pipeline overlap<\/jats:bold>\n                    , not large local SRAMs, emerge as the most effective knobs for efficient transformer inference under realistic system constraints.\n                  <\/jats:p>","DOI":"10.1145\/3803551","type":"journal-article","created":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T11:09:11Z","timestamp":1775041751000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Mitigating the Bandwidth Wall via Data-Streaming System\u2013Accelerator Co-Design"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7410-502X","authenticated-orcid":false,"given":"Qunyou","family":"Liu","sequence":"first","affiliation":[{"name":"Embedded Systems Laboratory, \u00c9cole polytechnique f\u00e9d\u00e9rale de Lausanne Facult\u00e9 des sciences et techniques de l'Ing\u00e9nieur","place":["Lausanne, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6971-1965","authenticated-orcid":false,"given":"Marina","family":"Zapater","sequence":"additional","affiliation":[{"name":"Embedded Systems Laboratory, \u00c9cole Polytechnique F\u00e9d\u00e9rale de Lausanne Facult\u00e9 des Sciences et Techniques de l'Ing\u00e9nieur","place":["Lausanne, Switzerland"]},{"name":"School of Engineering and Management Vaud (HEIG-VD), University of Applied Sciences Western Switzerland (HES-SO)","place":["Lausanne, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9536-4947","authenticated-orcid":false,"given":"David","family":"Atienza","sequence":"additional","affiliation":[{"name":"Embedded Systems Laboratory, \u00c9cole Polytechnique F\u00e9d\u00e9rale de Lausanne Facult\u00e9 des Sciences et Techniques de l'Ing\u00e9nieur","place":["Lausanne, Switzerland"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_1_2_2","volume-title":"AMD64 Architecture Programmer\u2019s Manual, Volume 3: General-Purpose and System Instructions","author":"Inc. Advanced Micro Devices,","year":"2018","unstructured":"Advanced Micro Devices, Inc.2018. AMD64 Architecture Programmer\u2019s Manual, Volume 3: General-Purpose and System Instructions. AMD, Santa Clara, CA. Retrieved from https:\/\/developer.amd.com\/resources\/developer-guides-manuals\/Section 4 documents SSE and AVX instruction support."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2018.8573496"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3566097.3567867"},{"key":"e_1_3_1_5_2","article-title":"Architecture and Implementation of the ARM Cortex-A8 Microprocessor","author":"Ltd. ARM","year":"2005","unstructured":"ARM Ltd.2005. Architecture and Implementation of the ARM Cortex-A8 Microprocessor. Design & Reuse Articles. Retrieved May 10, 2025 from https:\/\/www.design-reuse.com\/articles\/11580\/architecture-and-implementation-of-the-arm-cortex-a8-microprocessor.html","journal-title":"Design & Reuse Articles"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3472456.3472461"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/2024716.2024718"},{"key":"e_1_3_1_8_2","article-title":"Amazon\u2019s Custom ML Accelerators: AWS Trainium and Inferentia","author":"Bondalapati Vijay","year":"2023","unstructured":"Vijay Bondalapati, Dilip Sankar, and Mert Catak. 2023. Amazon\u2019s Custom ML Accelerators: AWS Trainium and Inferentia. Cloud Optimo Blog. Retrieved May 10, 2025 from https:\/\/www.cloudoptimo.com\/blog\/amazons-custom-ml-accelerators-aws-trainium-and-inferentia\/","journal-title":"Cloud Optimo Blog"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/2541940.2541967"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC.2016.7418007"},{"key":"e_1_3_1_11_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Chen Zhe","year":"2023","unstructured":"Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao. 2023. Vision transformer adapter for dense predictions. In Proceedings of the International Conference on Learning Representations (ICLR). OpenReview.net, Kigali, Rwanda, 1\u201322. ICLR Spotlight."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3579654.3579751"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_3_1_14_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et\u00a0al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations (ICLR). OpenReview.net, Virtual, 1\u201321. arxiv:2010.11929 [cs.CV]. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18074.2021.9586216"},{"key":"e_1_3_1_16_2","article-title":"Streaming SIMD Extensions (SSE) in the Pentium III Processor","author":"Corporation Intel","year":"1999","unstructured":"Intel Corporation. 1999. Streaming SIMD Extensions (SSE) in the Pentium III Processor. Intel Technology Brief. Retrieved May 10, 2025 from https:\/\/www.intel.com\/content\/www\/us\/en\/support\/articles\/000005779\/processors.html","journal-title":"Intel Technology Brief"},{"key":"e_1_3_1_17_2","article-title":"Intel\u00ae Advanced Vector Extensions (Intel AVX) Introduces 256-bit Vector Processing","author":"Corporation Intel","year":"2012","unstructured":"Intel Corporation. 2012. Intel\u00ae Advanced Vector Extensions (Intel AVX) Introduces 256-bit Vector Processing. Intel Developer Documentation. Retrieved May 10, 2025 from https:\/\/www.intel.com\/content\/www\/us\/en\/support\/articles\/000005779\/processors.html","journal-title":"Intel Developer Documentation"},{"key":"e_1_3_1_18_2","unstructured":"Andrei Ivanov Nikoli Dryden Tal Ben-Nun Shigang Li and Torsten Hoefler. 2020. Data movement is all you need: A case study on optimizing transformers. arxiv:2007.00072 [cs.LG]. Retrieved from https:\/\/arxiv.org\/abs\/2007.00072"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","unstructured":"Norman P. Jouppi Cliff Young Nishant Patil David A. Patterson Gururaj S. Agrawal Raminder Bajwa Sarah Bates Suresh Bhatia Nan Boden Alexander Borchers et\u00a0al. 2017. In-datacenter performance analysis of a tensor processing unit. In Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA). ACM Toronto ON Canada 1\u201312. DOI:10.1145\/3079856.3080246","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","unstructured":"Rachid Karami Chakshu Moar Sheng-Chun Kao and Hyoukjun Kwon. 2024. NonGEMM bench: Understanding the performance horizon of the latest ML workloads with NonGEMM workloads. arXiv:2404.11788. Retrieved from https:\/\/arxiv.org\/abs\/2404.11788","DOI":"10.1109\/ISPASS64960.2025.00011"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00047"},{"key":"e_1_3_1_23_2","first-page":"256","volume-title":"Proceedings of the SIAM Conference on Sparse Matrix Processing","author":"Kung H. T.","year":"1979","unstructured":"H. T. Kung. 1979. Systolic arrays (for VLSI). In Proceedings of the SIAM Conference on Sparse Matrix Processing. SIAM, Knoxville, TN, USA, 256\u2013282."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","unstructured":"Qunyou Liu Pengbo Yu Marina Zapater and David Atienza. 2026. SigmaQuant: Hardware-aware heterogeneous quantization method for edge DNN inference. IEEE Transactions on Circuits and Systems for Artificial Intelligence. 1\u201314. DOI:10.1109\/TCASAI.2026.3666506","DOI":"10.1109\/TCASAI.2026.3666506"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3659207"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","unstructured":"Qunyou Liu Marina Zapater and David Atienza. 2025. Gem5-AcceSys: Enabling system-level exploration of standard interconnects for novel accelerators. In 2025 62nd ACM\/IEEE Design Automation Conference (DAC). IEEE\/ACM. DOI:10.1109\/DAC63849.2025.11133394","DOI":"10.1109\/DAC63849.2025.11133394"},{"key":"e_1_3_1_27_2","unstructured":"Qunyou Liu Marina Zapater and David Atienza. 2025. MatrixFlow: System-accelerator co-design for high-performance transformer applications. arxiv:2503.05290 [cs.AR]. Retrieved from https:\/\/arxiv.org\/abs\/2503.05290"},{"key":"e_1_3_1_28_2","unstructured":"Simone Machetti Pasquale Davide Schiavone Lara Orlandic Darong Huang Deniz Kasap Giovanni Ansaloni and David Atienza. 2025. e-GPU: An open-source and configurable RISC-V graphic processing unit for TinyAI applications. arxiv:2505.08421 [cs.AR]. Retrieved from https:\/\/arxiv.org\/abs\/2505.08421"},{"key":"e_1_3_1_29_2","article-title":"Tenstorrent\u2019s Blackhole chips boast 768 RISC-V cores and almost as many FLOPS","author":"Mann Tobias","year":"2024","unstructured":"Tobias Mann. 2024. Tenstorrent\u2019s Blackhole chips boast 768 RISC-V cores and almost as many FLOPS. The Register (Online). Retrieved May 10, 2025 from https:\/\/www.theregister.com\/2024\/08\/27\/tenstorrent_ai_blackhole\/","journal-title":"The Register (Online)"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3422575.3422781"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3461662"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01196"},{"key":"e_1_3_1_33_2","article-title":"INT8 Transformers for Inference Acceleration","author":"Rock Andy","year":"2022","unstructured":"Andy Rock, Omar Khalil, Ofer Shai, and Paul Grouchy. 2022. INT8 Transformers for Inference Acceleration. Poster, 2nd Workshop on Efficient Natural Language and Speech Processing (ENLSP) at NeurIPS 2022. Retrieved from https:\/\/neurips2022-enlsp.github.io\/papers\/paper_52.pdf","journal-title":"Poster, 2nd Workshop on Efficient Natural Language and Speech Processing (ENLSP) at NeurIPS 2022"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO50266.2020.00047"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783751"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10766-022-00727-4"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/EPEPS48591.2020.9231458"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTAS.2010.21"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3424669"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2024.3375793"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISVLSI61997.2024.00046"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3803551","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:55:54Z","timestamp":1782402954000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3803551"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":41,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3803551"],"URL":"https:\/\/doi.org\/10.1145\/3803551","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2025-06-02","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-06","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}