{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T17:03:20Z","timestamp":1774631000415,"version":"3.50.1"},"reference-count":26,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,11,9]],"date-time":"2023-11-09T00:00:00Z","timestamp":1699488000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62090021"],"award-info":[{"award-number":["62090021"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Sino-German Mobility Programme (M-0187) by Sino-German Center for Research Promotion"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2023,11,30]]},"abstract":"<jats:p>\n            Transposed convolution has been prevailing in convolutional neural networks (CNNs), playing an important role in multiple scenarios such as image segmentation and back-propagation process of training CNNs. This mainly benefits from the ability to up-sample the input feature maps by interpolating new information from the input feature pixels. However, the\n            <jats:italic>backward-stencil computation<\/jats:italic>\n            constrains its performance and hindered its wide application in diverse platforms. Moreover, in contrast to the efforts on accelerating the convolution, there is a rare investigation on the acceleration of transposed convolution that is identically compute-intensive as the former.\n          <\/jats:p>\n          <jats:p>\n            For acceleration of transposed convolution, we propose an\n            <jats:italic>intermediate-centric<\/jats:italic>\n            dataflow scheme, in which we decouple the generation of the intermediate patch from its further process, aim at efficiently performing the\n            <jats:italic>backward-stencil computation<\/jats:italic>\n            . The\n            <jats:italic>intermediate-centric<\/jats:italic>\n            dataflow breaks the transposed convolution into several phases\/stages, achieving feeding the input feature maps and performing the\n            <jats:italic>backward-stencil computation<\/jats:italic>\n            in a pipelining manner. It also provides four-degree computation parallelism and efficient data reuse of input feature maps\/weights. Furthermore, we also theoretically analyze the irregular data dependence leveraging the polyhedral model, which constrains the parallel computing of transposed convolution. Additionally, we devise an optimization problem to explore the design space and automatically generate the optimal design configurations for different transposed convolutional layers and hardware platforms. By selecting the representative transposed convolutional layers from DCGAN, FSRCNN, and FCN, we generate the corresponding accelerator arrays of\n            <jats:italic>intermediate-centric<\/jats:italic>\n            dataflow on the Xilinx Alveo U200 platform and reach the performance of 3.92 TOPS, 2.72 TOPS, and 4.76 TOPS, respectively.\n          <\/jats:p>","DOI":"10.1145\/3561053","type":"journal-article","created":{"date-parts":[[2022,9,1]],"date-time":"2022-09-01T11:48:35Z","timestamp":1662032915000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["An\n            <i>Intermediate-Centric<\/i>\n            Dataflow for Transposed Convolution Acceleration on FPGA"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6764-7437","authenticated-orcid":false,"given":"Zhengzheng","family":"Ma","sequence":"first","affiliation":[{"name":"School of Computer Science, Center for Energy-efficient Computing and Applications, Peking University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0684-7299","authenticated-orcid":false,"given":"Tuo","family":"Dai","sequence":"additional","affiliation":[{"name":"School of Computer Science, Center for Energy-efficient Computing and Applications, Peking University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0996-2260","authenticated-orcid":false,"given":"Xuechao","family":"Wei","sequence":"additional","affiliation":[{"name":"School of Computer Science, Center for Energy-efficient Computing and Applications, Peking University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4932-3655","authenticated-orcid":false,"given":"Guojie","family":"Luo","sequence":"additional","affiliation":[{"name":"School of Computer Science, Center for Energy-efficient Computing and Applications, Peking University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,11,9]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"Proceedings of the Design Automation Conference","author":"Bachrach Jonathan","year":"2012","unstructured":"Jonathan Bachrach, Huy Vo, Brian Richards, Yunsup Lee, Andrew Waterman, Rimas Avi\u017eienis, John Wawrzynek, and Krste Asanovi\u0107. 2012. Chisel: Constructing hardware in a scala embedded language. In Proceedings of the Design Automation Conference."},{"key":"e_1_3_2_3_2","doi-asserted-by":"crossref","unstructured":"Marco Bevilacqua Aline Roumy Christine Guillemot and Marie Line Alberi-Morel. 2012. Low-complexity single-image super-resolution based on nonnegative neighbor embedding. (2012).","DOI":"10.5244\/C.26.135"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/2503210.2503289"},{"key":"e_1_3_2_5_2","volume-title":"Proceedings of the International Conference on Field-Programmable Technology","author":"Chan Long Chung","year":"2019","unstructured":"Long Chung Chan, Gurshaant Malik, and Nachiket Kapre. 2019. Partitioning FPGA-optimized systolic arrays for fun and profit. In Proceedings of the International Conference on Field-Programmable Technology."},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","unstructured":"Jung-Woo Chang Keon-Woo Kang and Suk-Ju Kang. 2020. An energy-efficient FPGA-based deconvolutional neural networks accelerator for single image super-resolution. IEEE Transactions on Circuits and Systems for Video Technology 30 1 (2020) 281\u2013295.","DOI":"10.1109\/TCSVT.2018.2888898"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_25"},{"key":"e_1_3_2_8_2","volume-title":"Proceedings of the IEEE 25th Annual International Symposium on Field-Programmable Custom Computing Machines","author":"Guan Yijin","year":"2017","unstructured":"Yijin Guan, Liang Hao, Ning Xu, Wenqing Wang, Xi Chen, Guangyu Sun, Wei Zhang, and Jason Cong. 2017. FP-DNN: An automated framework for mapping deep neural networks onto FPGAs with RTL-HLS hybrid templates. In Proceedings of the IEEE 25th Annual International Symposium on Field-Programmable Custom Computing Machines."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","unstructured":"Norman P. Jouppi Cliff Young Nishant Patil David Patterson Gaurav Agrawal Raminder Bajwa Sarah Bates Suresh Bhatia Nan Boden Al Borchers et\u00a0al. 2017. In-datacenter performance analysis of a tensor processing unit. In Proceedings of the 44th Annual International Symposium on Computer Architecture .","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3242900"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2020.3000519"},{"key":"e_1_3_2_14_2","volume-title":"Proceedings of the International Symposium on Field-Programmable Gate Arrays","author":"Qiu Jiantao","year":"2016","unstructured":"Jiantao Qiu, Jie Wang, Song Yao, Kaiyuan Guo, Boxun Li, Erjin Zhou, Jincheng Yu, Tianqi Tang, Ningyi Xu, Sen Song, et\u00a0al. 2016. Going deeper with embedded FPGA platform for convolutional neural network. In Proceedings of the International Symposium on Field-Programmable Gate Arrays."},{"key":"e_1_3_2_15_2","unstructured":"Alec Radford Luke Metz and Soumith Chintala. 2015. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv:1511.06434 Retrieved from https:\/\/arxiv.org\/abs\/1511.06434."},{"key":"e_1_3_2_16_2","first-page":"1","volume-title":"Proceedings of the 2019 IEEE International Symposium on Circuits and Systems","author":"Wang Deguang","year":"2019","unstructured":"Deguang Wang, Junzhong Shen, Mei Wen, and Chunyuan Zhang. 2019. Towards a uniform architecture for the efficient implementation of 2D and 3D deconvolutional neural networks on FPGAs. In Proceedings of the 2019 IEEE International Symposium on Circuits and Systems. 1\u20135."},{"key":"e_1_3_2_17_2","volume-title":"Proceedings of the International Symposium on Field-Programmable Gate Arrays","author":"Wang Jie","year":"2021","unstructured":"Jie Wang, Licheng Guo, and Jason Cong. 2021. AutoSA: A polyhedral compiler for high-performance systolic arrays on FPGA. In Proceedings of the International Symposium on Field-Programmable Gate Arrays."},{"key":"e_1_3_2_18_2","volume-title":"Proceedings of the Design Automation Conference","author":"Wei Xuechao","year":"2017","unstructured":"Xuechao Wei, Cody Hao Yu, Peng Zhang, Youxiang Chen, Yuxin Wang, Han Hu, Yun Liang, and Jason Cong. 2017. Automated systolic array architecture synthesis for high throughput CNN inference on FPGAs. In Proceedings of the Design Automation Conference."},{"key":"e_1_3_2_19_2","doi-asserted-by":"crossref","first-page":"100008","DOI":"10.1016\/j.hcc.2021.100008","article-title":"A survey of federated learning for edge computing: Research problems and solutions","author":"Xia Qi","year":"2021","unstructured":"Qi Xia, Winson Ye, Zeyi Tao, Jindi Wu, and Qun Li. 2021. A survey of federated learning for edge computing: Research problems and solutions. High-Confidence Computing 1, 1 (2021), 100008.","journal-title":"High-Confidence Computing"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240765.3240810"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2018.2857258"},{"key":"e_1_3_2_22_2","volume-title":"Proceedings of the International Symposium on Field-Programmable Custom Computing Machines","author":"Yazdanbakhsh Amir","year":"2018","unstructured":"Amir Yazdanbakhsh, Michael Brzozowski, Behnam Khaleghi, Soroush Ghodrati, Kambiz Samadi, Nam Sung Kim, and Hadi Esmaeilzadeh. 2018. FlexiGAN: An end-to-end solution for FPGA acceleration of generative adversarial networks. In Proceedings of the International Symposium on Field-Programmable Custom Computing Machines."},{"key":"e_1_3_2_23_2","volume-title":"Proceedings of the International Symposium on Computer Architecture","author":"Yazdanbakhsh Amir","year":"2018","unstructured":"Amir Yazdanbakhsh, Kambiz Samadi, Nam Sung Kim, and Hadi Esmaeilzadeh. 2018. GANAX: A unified MIMD-SIMD acceleration for generative adversarial networks. In Proceedings of the International Symposium on Computer Architecture."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2020.2995741"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689060"},{"key":"e_1_3_2_26_2","volume-title":"Proceedings of the International Symposium on Circuits and Systems","author":"Zhang Jiaxi","year":"2019","unstructured":"Jiaxi Zhang, Wentai Zhang, Guojie Luo, Xuechao Wei, Yun Liang, and Jason Cong. 2019. Frequency improvement of systolic array-based CNNs on FPGAs. In Proceedings of the International Symposium on Circuits and Systems."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240765.3240801"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3561053","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3561053","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:15Z","timestamp":1750182555000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3561053"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,9]]},"references-count":26,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,11,30]]}},"alternative-id":["10.1145\/3561053"],"URL":"https:\/\/doi.org\/10.1145\/3561053","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,9]]},"assertion":[{"value":"2021-12-31","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-08-10","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}