{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:18:03Z","timestamp":1760239083480,"version":"build-2065373602"},"reference-count":26,"publisher":"MDPI AG","issue":"19","license":[{"start":{"date-parts":[[2020,9,28]],"date-time":"2020-09-28T00:00:00Z","timestamp":1601251200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100019081","name":"Hunan Provincial Science and Technology Plan Project","doi-asserted-by":"publisher","award":["2018XK2102"],"award-info":[{"award-number":["2018XK2102"]}],"id":[{"id":"10.13039\/501100019081","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Due to the high throughput and high computing capability of convolutional neural networks (CNNs), researchers are paying increasing attention to the design of CNNs hardware accelerator architecture. Accordingly, in this paper, we propose a block parallel computing algorithm based on the matrix transformation computing algorithm (MTCA) to realize the convolution expansion and resolve the block problem of the intermediate matrix. It enables high parallel implementation on hardware. Moreover, we also provide a specific calculation method for the optimal partition of matrix multiplication to optimize performance. In our evaluation, our proposed method saves more than 60% of hardware storage space compared with the im2col(image to column) approach. More specifically, in the case of large-scale convolutions, it saves nearly 82% of storage space. Under the accelerator architecture framework designed in this paper, we realize the performance of 26.7GFLOPS-33.4GFLOPS (depending on convolution type) on FPGA(Field Programmable Gate Array) by reducing bandwidth and improving data reusability. It is 1.2\u00d7\u20134.0\u00d7 faster than memory-efficient convolution (MEC) and im2col, respectively, and represents an effective solution for a large-scale convolution accelerator.<\/jats:p>","DOI":"10.3390\/s20195558","type":"journal-article","created":{"date-parts":[[2020,9,28]],"date-time":"2020-09-28T10:39:58Z","timestamp":1601289598000},"page":"5558","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["An Accelerator Design Using a MTCA Decomposition Algorithm for CNNs"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5600-3740","authenticated-orcid":false,"given":"Yunping","family":"Zhao","sequence":"first","affiliation":[{"name":"College of Computer, National University of Defense Technology, Changsha 410073, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianzhuang","family":"Lu","sequence":"additional","affiliation":[{"name":"College of Computer, National University of Defense Technology, Changsha 410073, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaowen","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Computer, National University of Defense Technology, Changsha 410073, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,9,28]]},"reference":[{"key":"ref_1","first-page":"1097","article-title":"ImageNet classification with deep convolutional neural networks","volume":"60","author":"Krizhevsky","year":"2012","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_2","first-page":"96","article-title":"Target recognition in SAR images via sparse representation in the frequency domain","volume":"12","author":"Dong","year":"2019","journal-title":"Pattern Recognit."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"154","DOI":"10.1007\/s11263-013-0620-5","article-title":"Selective search for object recognition","volume":"2","author":"Uijlings","year":"2013","journal-title":"Int. J. Comput. Vis."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014). Rich feature hierarchies for accurate object detection and semantic segmentation. Proc. IEEE CVPR, 580\u2013587.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Noh, H., Hong, S., and Han, B. (2015). Learning deconvolution net-work for semantic segmentation. Proc. IEEE ICCV, 1520\u20131528.","DOI":"10.1109\/ICCV.2015.178"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"393","DOI":"10.1145\/3007787.3001179","article-title":"Cambricon: An instruction set architecture for neural networks","volume":"44","author":"Liu","year":"2016","journal-title":"ACM Sigarch Comput. Archit. News"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Lavin, A., and Gray, S. (2016). Fast algorithms for convolutional neural net-works. Proc. IEEE CVPR, 4013\u20134021.","DOI":"10.1109\/CVPR.2016.435"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Chen, Y.H., Krishna, T., Emer, J.S., and Sze, V. (2016, January 18\u201322). Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks. Proceedings of the 2016 ACM\/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA), Seoul, Korea.","DOI":"10.1109\/ISCA.2016.40"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"968","DOI":"10.1109\/JSSC.2017.2778281","article-title":"A high energy efficient reconfigurable hybrid neural network processor for deep learning applications","volume":"53","author":"Yin","year":"2018","journal-title":"IEEE J. Solid-State Circuits"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Desoli, G. (2017, January 5\u20139). A 2.9TOPS\/W deep convolutional neural network SoC in FD-SOI 28 nm for intelligent embedded systems. Proceedings of the IEEE Int. Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA.","DOI":"10.1109\/ISSCC.2017.7870349"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Shin, D., Lee, J., and Yoo, H.J. (2017, January 5\u20139). DNPU: An 8.1TOPS\/W reconfigurable CNN-RNN processor for general-purpose deep neural networks. Proceedings of the IEEE International Solid-State Circuits Conference (ISSCC), San Francisco, CA, USA.","DOI":"10.1109\/ISSCC.2017.7870350"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1941","DOI":"10.1109\/TCSI.2017.2767204","article-title":"Efficient hardware architectures for deep convolutional neural network","volume":"65","author":"Wang","year":"2018","journal-title":"IEEE Trans. Circuits Syst. I"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1354","DOI":"10.1109\/TVLSI.2018.2815603","article-title":"Optimizing the convolution operation to accelerate deep neural networks on FPGA","volume":"26","author":"Ma","year":"2018","journal-title":"IEEE Trans. Very Large Scale Integr. (VLSI) Syst."},{"key":"ref_14","first-page":"1349","article-title":"An architecture to accelerate convolution in deep neural networks","volume":"65","author":"Ardakani","year":"2018","journal-title":"IEEE Trans. Very Large Scale Integr. (VLSI) Syst."},{"key":"ref_15","first-page":"217","article-title":"Optimization method of convolution calculation based on matrix transformation","volume":"45","author":"Fang","year":"2019","journal-title":"Comput. Eng."},{"key":"ref_16","unstructured":"Kung, H.T., and Leiserson, C.E. (1978). Systolic Arrays. Handbook of Signal Processing Systems, Springer."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"269","DOI":"10.1145\/2654822.2541967","article-title":"DianNao: A small-footprint high-throuhput accelerator for ubiquitous machine-learning","volume":"49","author":"Chen","year":"2014","journal-title":"ACM SIGARCH Comput. Archit. News"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Chen, Y., Lou, T., and Liu, S. (2014). DaDianNao: A machine-learning supercomputer. ACM Int. Symp. Microarchit., 609\u2013622.","DOI":"10.1109\/MICRO.2014.58"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1109\/JSSC.2016.2616357","article-title":"Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural net-works","volume":"52","author":"Chen","year":"2017","journal-title":"IEEE J. Solid-State Circuits"},{"key":"ref_20","first-page":"10","article-title":"MALMM: A Multi-array Architecture for Large-scale Matrix Multiplication on FPGA","volume":"15","author":"You","year":"2018","journal-title":"IEICE Electron. Express"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"515","DOI":"10.1016\/j.ces.2017.10.006","article-title":"Parallel computing method of two-dimensional matrix convolution","volume":"52","author":"Zhang","year":"2018","journal-title":"Eng. Sci."},{"key":"ref_22","unstructured":"Jing, S., Haoqi, R., Zhifeng, Z., Jun, W., and Zhenyu, J. (2020, January 16\u201319). A High-Performance Systolic Array Accelerator Dedicated for CNN. Proceedings of the 2019 IEEE 19th International Conference on Communication Technology (ICCT), Xi\u2019an, China."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1953","DOI":"10.1109\/TVLSI.2020.3002779","article-title":"An Efficient Hardware Accelerator for Structured Sparse Convolutional Neural Networks on FPGAs","volume":"28","author":"Chaoyang","year":"2020","journal-title":"IEEE Trans. Very Large Scale Integr. (VLSI) Syst."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Maurizio, C., Beatrice, B., Alberto, M., and Muhammad, S. (2020). An Updated Survey of Efficient Hardware Architectures for Accelerating Deep Convolutional Neural Networks. Future Internet, 12.","DOI":"10.3390\/fi12070113"},{"key":"ref_25","unstructured":"Cho, M., and Brand, D. (2017, January 6\u201311). MEC: Memory-efficient convolution for deep neural network. Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_26","first-page":"2251","article-title":"Matrix multiplication and vectorization for multi-core vector processors","volume":"41","author":"Liu","year":"2018","journal-title":"J. Comput. Sci."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/19\/5558\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:14:32Z","timestamp":1760177672000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/19\/5558"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,28]]},"references-count":26,"journal-issue":{"issue":"19","published-online":{"date-parts":[[2020,10]]}},"alternative-id":["s20195558"],"URL":"https:\/\/doi.org\/10.3390\/s20195558","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2020,9,28]]}}}