{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,24]],"date-time":"2025-10-24T16:45:11Z","timestamp":1761324311996,"version":"build-2065373602"},"reference-count":31,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2020,8,25]],"date-time":"2020-08-25T00:00:00Z","timestamp":1598313600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Regione Calabria - Italy","award":["POR Calabria FSE\/FESR 2014-2020 \u2013 International Mobility of PhD students and research grants\/type A Researchers - Actions 10.5.6 and 10.5.12"],"award-info":[{"award-number":["POR Calabria FSE\/FESR 2014-2020 \u2013 International Mobility of PhD students and research grants\/type A Researchers - Actions 10.5.6 and 10.5.12"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Today, convolutional and deconvolutional neural network models are exceptionally popular thanks to the impressive accuracies they have been proven in several computer-vision applications. To speed up the overall tasks of these neural networks, purpose-designed accelerators are highly desirable. Unfortunately, the high computational complexity and the huge memory demand make the design of efficient hardware architectures, as well as their deployment in resource- and power-constrained embedded systems, still quite challenging. This paper presents a novel purpose-designed hardware accelerator to perform 2D deconvolutions. The proposed structure applies a hardware-oriented computational approach that overcomes the issues of traditional deconvolution methods, and it is suitable for being implemented within any virtually system-on-chip based on field-programmable gate array devices. In fact, the novel accelerator is simply scalable to comply with resources available within both high- and low-end devices by adequately scaling the adopted parallelism. As an example, when exploited to accelerate the Deep Convolutional Generative Adversarial Network model, the novel accelerator, running as a standalone unit implemented within the Xilinx Zynq XC7Z020 System-on-Chip (SoC) device, performs up to 72 GOPs. Moreover, it dissipates less than 500mW@200MHz and occupies 5.6%, 4.1%, 17%, and 96%, respectively, of the look-up tables, flip-flops, random access memory, and digital signal processors available on-chip. When accommodated within the same device, the whole embedded system equipped with the novel accelerator performs up to 54 GOPs and dissipates less than 1.8W@150MHz. Thanks to the increased parallelism exploitable, more than 900 GOPs can be executed when the high-end Virtex-7 XC7VX690T device is used as the implementation platform. Moreover, in comparison with state-of-the-art competitors implemented within the Zynq XC7Z045 device, the system proposed here reaches a computational capability up to 20% higher, and saves more than 60% and 80% of power consumption and logic resources requirement, respectively, using 5.7\u00d7 fewer on-chip memory resources.<\/jats:p>","DOI":"10.3390\/jimaging6090085","type":"journal-article","created":{"date-parts":[[2020,8,25]],"date-time":"2020-08-25T09:30:07Z","timestamp":1598347807000},"page":"85","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Efficient Deconvolution Architecture for Heterogeneous Systems-on-Chip"],"prefix":"10.3390","volume":"6","author":[{"given":"Stefania","family":"Perri","sequence":"first","affiliation":[{"name":"Department of Mechanical, Energy and Management Engineering, University of Calabria, 87036 Rende, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cristian","family":"Sestito","sequence":"additional","affiliation":[{"name":"Department of Informatics, Modeling, Electronics and System Engineering, University of Calabria, 87036 Rende, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2197-4563","authenticated-orcid":false,"given":"Fanny","family":"Spagnolo","sequence":"additional","affiliation":[{"name":"Department of Informatics, Modeling, Electronics and System Engineering, University of Calabria, 87036 Rende, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pasquale","family":"Corsonello","sequence":"additional","affiliation":[{"name":"Department of Informatics, Modeling, Electronics and System Engineering, University of Calabria, 87036 Rende, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,8,25]]},"reference":[{"key":"ref_1","unstructured":"Goodfellow, I.J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014, January 8\u201313). Generative adversarial nets. Proceedings of the 27th International Conference on Neural Information Processing Systems\u2014Volume 2, Montreal, QC, Canada."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1016\/j.asoc.2018.05.018","article-title":"A review on deep learning techniques for image and video semantic segmentation","volume":"70","author":"Oprea","year":"2018","journal-title":"Appl. Soft Comput."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"295","DOI":"10.1109\/TPAMI.2015.2439281","article-title":"Image super-resolution using deep convolutional networks","volume":"38","author":"Dong","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_4","unstructured":"Dumoulin, V., and Visin, F. (2019, July 14). A Guide to Convolution Arithmetic for Deep Learning. Available online: https:\/\/arxiv.org\/abs\/1603.07285."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2295","DOI":"10.1109\/JPROC.2017.2761740","article-title":"Efficient processing of deep neural networks: A tutorial and survey","volume":"105","author":"Sze","year":"2017","journal-title":"Proc. IEEE"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1016\/j.vlsi.2019.07.005","article-title":"Computer vision algorithms and hardware implementations: A survey","volume":"69","author":"Feng","year":"2019","journal-title":"Integration"},{"key":"ref_7","first-page":"1","article-title":"A survey of FPGA-based accelerators for convolutional neural networks","volume":"32","author":"Mittal","year":"2018","journal-title":"Neural Comput. Appl."},{"key":"ref_8","unstructured":"Yazdanbakhsh, A., Brzozowski, M., Khaleghi, B., Ghodrati, S., Samadi, K., Kim, N.S., and Esmaeilzadeh, H. (May, January 29). FlexiGAN: An End-to-End Solution for FPGA Acceleration of Generative Adversarial Networks. Proceedings of the 26th IEEE International Symposium on Field-Programmable Custom Computing Machines, Boulder, CO, USA."},{"key":"ref_9","unstructured":"Zhang, X., Das, S., Neopane, O., and Kreutz-Delgado, K. (2019, July 14). A Design Methodology for Efficient Implementation of Deconvolutional Neural Networks on an FPGA. Available online: https:\/\/arxiv.org\/abs\/1705.02583."},{"key":"ref_10","first-page":"1","article-title":"Optimizing CNN-based Segmentation with Deeply Customized Convolutional and Deconvolutional Architectures on FPGA","volume":"11","author":"Liu","year":"2018","journal-title":"ACM Trans. Rec. Technol. Syst."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Liu, S., Zeng, C., Fan, H., Ng, H.-C., Meng, J., and Luk, W. (2018, January 10\u201314). Memory-Efficient Architecture for Accelerating Generative Networks on FPGAs. Proceedings of the IEEE International Conference on Field Programmable Technology, Naha, Okinawa, Japan.","DOI":"10.1109\/FPT.2018.00016"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Liu, S., and Luk, W. (2019, January 9\u201313). Towards an Efficient Accelerator for DNN-Based Remote Sensing Image Segmentation on FPGAs. Proceedings of the 29th International Conference on Field Programmable Logic and Applications, Barcelona, Spain.","DOI":"10.1109\/FPL.2019.00037"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Chang, J.-W., and Kang, S.-J. (2018, January 22\u201325). Optimizing FPGA-based convolutional neural networks accelerator for image super-resolution. Proceedings of the 23rd Asia and South Pacific Design Automation Conference, Jeju, South Korea.","DOI":"10.1109\/ASPDAC.2018.8297347"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"281","DOI":"10.1109\/TCSVT.2018.2888898","article-title":"An Energy-Efficient FPGA-Based Deconvolutional Neural Networks Accelerator for Single Image Super-Resolution","volume":"30","author":"Chang","year":"2020","journal-title":"IEEE Trans. Circ. Sys. Video Technol."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Wang, D., Shen, J., Wen, M., and Zhang, C. (2019). Efficient Implementation of 2D and 3D Sparse Deconvolutional Neural Networks with a Uniform Architecture on FPGAs. Electronics, 8.","DOI":"10.3390\/electronics8070803"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Chang, J.-W., Ahn, S., Kang, K.-W., and Kang, S.-J. (2020, January 13\u201316). Towards Design Methodology of Efficient Fast Algorithms for Accelerating Generative Adversarial Networks on FPGAs. Proceedings of the 25th Asia and South Pacific Design Automation Conference, Beijing, China.","DOI":"10.1109\/ASP-DAC47756.2020.9045214"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1545","DOI":"10.1109\/TVLSI.2020.2995741","article-title":"Uni-OPU: An FPGA-Based Uniform Accelerator for Convolutional and Transposed Convolutional Networks","volume":"28","author":"Yu","year":"2020","journal-title":"IEEE Trans. VLSI Syst."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Di, X., Yang, H.-G., Jia, Y., Huang, Z., and Mao, N. (2020). Exploring Efficient Acceleration Architecture for Winograd-Transformed Transposed Convolution of GANs on FPGAs. Electronics, 9.","DOI":"10.3390\/electronics9020286"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Spagnolo, F., Perri, S., Frustaci, F., and Corsonello, P. (2020). Energy-Efficient Architecture for CNNs Inference on Heterogeneous FPGA. J. Low Power Electron. Appl., 10.","DOI":"10.3390\/jlpea10010001"},{"key":"ref_20","unstructured":"(2020, July 14). AMBA 4 AXI4, AXI4-Lite, and AXI4-Stream Protocol Assertions User Guide. Available online: http:\/\/infocenter.arm.com\/help\/index.jsp?topic=\/com.arm.doc.ihi0022d\/index.html."},{"key":"ref_21","unstructured":"Radford, A., Metz, L., and Chintala, S. (2016, January 2\u20134). Unsupervised representation learning with deep convolutional generative adversarial networks. Proceedings of the International Conference on Learning Representations, Caribe Hilton, San Juan, Puerto Rico."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Gaide, B., Gaitonde, D., Ravishankar, C., and Bauer, T. (2019, January 24\u201326). Xilinx Adaptive Compute Acceleration Platform: Versal\u2122Architecture. Proceedings of the ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201919), Seaside, CA, USA.","DOI":"10.1145\/3289602.3293906"},{"key":"ref_23","unstructured":"(2020, July 14). Zynq-7000 SoC Technical Reference Manual (UG585 v. 1.12.2), July 2018. Available online: https:\/\/www.xilinx.com\/support\/documentation\/user_guides\/ug585-Zynq-7000-TRM.pdf."},{"key":"ref_24","unstructured":"(2020, July 14). 7 Series FPGAs Data Sheet: Overview (DS180 v. 2.6), February 2018. Available online: https:\/\/www.xilinx.com\/support\/documentation\/data_sheets\/ds180_7Series_Overview.pdf."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully Convolutional Networks for Semantic Segmentation. Proceedings of the 28th IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Noh, H., Hong, S., and Han, B. (2015, January 7\u201313). Learning deconvolution network for semantic segmentation. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.178"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zeiler, M.D., Krishnan, D., Taylor, G.W., and Fergus, R. (2010, January 13\u201318). Deconvolutional networks. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539957"},{"key":"ref_29","unstructured":"(2020, July 14). AXI DMA v7.1, LogiCORE IP Product Guide (PG021). Available online: https:\/\/www.xilinx.com\/support\/documentation\/ip_documentation\/axi_dma\/v7_1\/pg021_axi_dma.pdf."},{"key":"ref_30","unstructured":"(2020, July 14). AXI Video Direct Memory Access v6.2, LogiCORE IP Product Guide (PG020). Available online: https:\/\/www.xilinx.com\/support\/documentation\/ip_documentation\/axi_vdma\/v6_2\/pg020_axi_vdma.pdf."},{"key":"ref_31","unstructured":"(2020, July 14). AXI4-Stream Infrastructure IP Suite v3.0 LogiCORE IP Product Guide (PG085). Available online: https:\/\/www.xilinx.com\/support\/documentation\/ip_documentation\/axis_infrastructure_ip_suite\/v1_1\/pg085-axi4stream-infrastructure.pdf."}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/6\/9\/85\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:06:17Z","timestamp":1760177177000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/6\/9\/85"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,8,25]]},"references-count":31,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2020,9]]}},"alternative-id":["jimaging6090085"],"URL":"https:\/\/doi.org\/10.3390\/jimaging6090085","relation":{},"ISSN":["2313-433X"],"issn-type":[{"type":"electronic","value":"2313-433X"}],"subject":[],"published":{"date-parts":[[2020,8,25]]}}}