{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T16:11:34Z","timestamp":1781194294258,"version":"3.54.1"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2026,3,31]]},"abstract":"<jats:p>Deploying Vision Transformers (ViTs) on edge devices poses significant challenges due to their high computational demands and memory access overheads, which severely hinder real-time inference efficiency. This article proposes a modular and adaptive ViT acceleration architecture targeting the AMD Versal ACAP platform. By leveraging heterogeneous resource collaboration and fine-grained dataflow optimizations, the proposed design addresses performance bottlenecks effectively. We introduce a resource-efficient attention computation module that localizes self-attention operations within AI Engine (AIE) core clusters, thereby reducing inter-module communication and minimizing MAC resource usage. In parallel, a resource-aware multi-stage pipeline scheduling strategy dynamically partitions and parallelizes the computation-intensive feed-forward network (FFN), improving computation reuse and module-level coordination. The architecture integrates parameter tiling and a PLIO-based broadcasting mechanism to construct a decoupled compute-communication dataflow engine, alleviating memory bottlenecks. Experimental results on the Xilinx VCK5000 ACAP platform demonstrate that the proposed design achieves 33.2 TOPS throughput at INT8 precision\u2014outperforming the state-of-the-art EQ-ViT accelerator by 27%\u2014while maintaining a competitive efficiency of 510.6 GOPS\/W. Scalability evaluations on ViT-Base and DeiT-Tiny confirm the design\u2019s adaptability in edge scenarios, offering a resource-efficient and reconfigurable hardware paradigm for high-density Transformer inference.<\/jats:p>","DOI":"10.1145\/3779444","type":"journal-article","created":{"date-parts":[[2025,12,12]],"date-time":"2025-12-12T03:00:07Z","timestamp":1765508407000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["REATA: An Efficient Vision Transformer Accelerator Featuring a Resource-Optimized Attention Design on\u00a0Versal\u00a0ACAP"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8601-802X","authenticated-orcid":false,"given":"Wenbo","family":"Zhang","sequence":"first","affiliation":[{"name":"College of Computer Science, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5628-6943","authenticated-orcid":false,"given":"Yan","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Computer Science, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-8262-2185","authenticated-orcid":false,"given":"Yiqi","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Computer Science, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-2197-0928","authenticated-orcid":false,"given":"Lingjie","family":"Wu","sequence":"additional","affiliation":[{"name":"College of Computer Science, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-6281-8515","authenticated-orcid":false,"given":"Xingtong","family":"Hu","sequence":"additional","affiliation":[{"name":"College of Computer Science, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,3,6]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/HOTCHIPS.2019.8875639"},{"key":"e_1_3_1_3_2","unstructured":"AMD. 2023. Versal Adaptive SoC AI Engine Architecture Manual (AM009). Retrieved from https:\/\/docs.amd.com\/r\/en-US\/am009-versal-ai-engineReleased by AMD"},{"key":"e_1_3_1_4_2","unstructured":"AMD. 2024. AI Engine Kernel and Graph Programming Guide (UG1079). Retrieved from https:\/\/docs.amd.com\/r\/2024.1-English\/ug1079-ai-engine-kernel-codingDescribes the intricacies of AI Engine kernel and graph programming"},{"key":"e_1_3_1_5_2","unstructured":"AMD. 2024. AI Engine Tools and Flows User Guide (UG1076). Retrieved from https:\/\/docs.amd.com\/r\/2024.1-English\/ug1076-ai-engine-environment"},{"key":"e_1_3_1_6_2","unstructured":"AMD (Xilinx). 2025. Power Design Manager (PDM) User Guide. AMDInc. Retrieved from https:\/\/www.amd.com\/en\/products\/software\/adaptive-socs-and-fpgas\/power-design-manager.html"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2025.107728"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3656177"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV61041.2025.00929"},{"key":"e_1_3_1_12_2","unstructured":"Tri Dao. 2023. Flashattention-2: Faster attention with better parallelism and work partitioning. arXiv:2307.08691. Retrieved from https:\/\/arxiv.org\/abs\/2307.08691"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1189"},{"key":"e_1_3_1_14_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10071047"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2024.3443692"},{"key":"e_1_3_1_17_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3289602.3293906"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2024.3517751"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01565"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL57034.2022.00027"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3676536.3676766"},{"key":"e_1_3_1_24_2","unstructured":"Xiang Liu Yijun Song Xia Li Yifei Sun Huiying Lan Zemin Liu Linshan Jiang and Jialin Li. 2024. ED-ViT: Splitting vision transformer for distributed inference on edge devices. arXiv:2410.11650. Retrieved from https:\/\/arxiv.org\/abs\/2410.11650"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV61041.2025.00064"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/HiPC58850.2023.00039"},{"key":"e_1_3_1_28_2","unstructured":"NVIDIA. 2023. FasterTransformer. Retrieved from https:\/\/github.com\/NVIDIA\/FasterTransformer [Online]."},{"key":"e_1_3_1_29_2","unstructured":"NVIDIA. 2023. NVIDIA AWS A10G GPU Data Sheet. Retrieved June 19 2025 from https:\/\/aws.amazon.com\/ec2\/instance-types\/p4\/"},{"key":"e_1_3_1_30_2","unstructured":"Hannes Vanholder. 2016. Efficient Inference with TensorRT. Retrieved June 19 2025 from https:\/\/developer.nvidia.com\/blog\/efficient-inference-tensorrt\/"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2022.3197489"},{"key":"e_1_3_1_32_2","unstructured":"Tianzhu Ye Li Dong Yuqing Xia Yutao Sun Yi Zhu Gao Huang and Furu Wei. 2024. Differential transformer. arXiv:2410.05258. Retrieved from https:\/\/arxiv.org\/abs\/2410.05258"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10071027"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00564"},{"key":"e_1_3_1_35_2","unstructured":"Dong Zhang Rui Yan Pingcheng Dong and Kwang-Ting Cheng. 2025. Memory efficient transformer adapter for dense predictions. arXiv:2502.01962. Retrieved from https:\/\/arxiv.org\/abs\/2502.01962"},{"key":"e_1_3_1_36_2","unstructured":"Wenbo Zhang Yiqi Liu and Zhenshan Bao. 2024. CAT: Customized transformer accelerator framework on versal ACAP. arXiv:2409.09689. Retrieved from https:\/\/arxiv.org\/abs\/2409.09689"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3678010"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW54120.2021.00314"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00911"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01388"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3543622.3573210"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3686163"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626202.3637569"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3779444","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T11:15:28Z","timestamp":1773227728000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3779444"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,6]]},"references-count":42,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,3,31]]}},"alternative-id":["10.1145\/3779444"],"URL":"https:\/\/doi.org\/10.1145\/3779444","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,6]]},"assertion":[{"value":"2025-06-20","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-24","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-06","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}