{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T15:26:15Z","timestamp":1781537175673,"version":"3.54.5"},"reference-count":32,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2022,4,18]],"date-time":"2022-04-18T00:00:00Z","timestamp":1650240000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Institute of Information &amp; Communications Technology Planning &amp; Evaluation(IITP) grant funded by the Korea government (MSIT)","award":["2020-0-01297"],"award-info":[{"award-number":["2020-0-01297"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>We propose a method for minimizing global buffer access within a deep learning accelerator for convolution operations by maximizing the data reuse through a local register file, thereby substituting the local register file access for the power-hungry global buffer access. To fully exploit the merits of data reuse, this study proposes a rearrangement of the computational sequence in a deep learning accelerator. Once input data are read from the global buffer, repeatedly reading the same data is performed only through the local register file, saving significant power consumption. Furthermore, different from prior works that equip local register files in each computation unit, the proposed method enables sharing a local register file along the column of the 2D computation array, saving resources and controlling overhead. The proposed accelerator is implemented on an off-the-shelf field-programmable gate array to verify the functionality and resource utilization. Then, the performance improvement of the proposed method is demonstrated relative to popular deep learning accelerators. Our evaluation indicates that the proposed deep learning accelerator reduces the number of global-buffer accesses to nearly 86.8%, consequently saving up to 72.3% of the power consumption for the input data memory access with a minor increase in resource usage compared to a conventional deep learning accelerator.<\/jats:p>","DOI":"10.3390\/s22083095","type":"journal-article","created":{"date-parts":[[2022,4,19]],"date-time":"2022-04-19T02:39:31Z","timestamp":1650335971000},"page":"3095","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Minimizing Global Buffer Access in a Deep Learning Accelerator Using a Local Register File with a Rearranged Computational Sequence"],"prefix":"10.3390","volume":"22","author":[{"given":"Minjae","family":"Lee","sequence":"first","affiliation":[{"name":"Department of Electronic Engineering, Hanyang University, Seoul 04763, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhongfeng","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Electronic Engineering, Hanyang University, Seoul 04763, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6915-6693","authenticated-orcid":false,"given":"Seungwon","family":"Choi","sequence":"additional","affiliation":[{"name":"Department of Electronic Engineering, Hanyang University, Seoul 04763, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jungwook","family":"Choi","sequence":"additional","affiliation":[{"name":"Department of Electronic Engineering, Hanyang University, Seoul 04763, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,4,18]]},"reference":[{"key":"ref_1","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20136). ImageNet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Graves, A., Mohamed, A.-R., and Hinton, G. (2013, January 26\u201331). Speech recognition with deep recurrent neural networks. Proceedings of the 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"ref_4","unstructured":"Ba, J., Mnih, V., and Kavukcuoglu, K. (2015). Multiple object recognition with visual attention. arXiv."},{"key":"ref_5","unstructured":"Simonyan, K., and Zisserman, A. (2014, January 8\u201313). Two-stream convolutional networks for action recognition in videos. Proceedings of the Advances in Neural Information Processing Systems 27 (NIPS 2014), Montreal, QC, Canada."},{"key":"ref_6","unstructured":"Vanhoucke, V., Senior, A., and Mao, M.Z. (2011, January 12\u201317). Improving the speed of neural networks on CPUs. Proceedings of the Deep Learning and Unsupervised Feature Learning Workshop, NIPS 2011, Granada, Spain."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Jagannathan, S., Mody, M., and Mathew, M. (2016, January 7\u201311). Optimizing convolutional neural network on DSP. Proceedings of the 2016 IEEE International Conference on Consumer Electronics (ICCE), Las Vegas, NV, USA.","DOI":"10.1109\/ICCE.2016.7430652"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"121987","DOI":"10.1109\/ACCESS.2020.3006773","article-title":"An FPGA-based hardware accelerator for real-time block-matching and 3D filtering","volume":"8","author":"Wang","year":"2020","journal-title":"IEEE Access"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Xu, R., Han, F., and Ta, Q. (2018, January 12). Deep Learning at Scale on NVIDIA V100 Accelerators. Proceedings of the Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS), Dallas, TX, USA.","DOI":"10.1109\/PMBS.2018.8641600"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"101635","DOI":"10.1016\/j.sysarc.2019.101635","article-title":"A survey of techniques for optimizing deep learning on GPUs","volume":"99","author":"Mittal","year":"2019","journal-title":"In J. Syst. Archit."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2962131","article-title":"Understanding GPU power: A survey of profiling, modeling, and simulation methods","volume":"49","author":"Bridges","year":"2016","journal-title":"ACM Comput. Surv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Yang, T.-J., Chen, Y.-H., Emer, J., and Sze, V. (November, January 29). A method to estimate the energy consumption of deep neural networks. Proceedings of the 2017 51st Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, USA.","DOI":"10.1109\/ACSSC.2017.8335698"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"398","DOI":"10.1109\/JETCAS.2019.2908937","article-title":"MAX2: An ReRAM-based neural network accelerator that maximizes data reuse and area utilization","volume":"9","author":"Mao","year":"2019","journal-title":"IEEE J. Emerg. Sel. Topics Circuits Syst."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Xu, H., Shiomi, J., and Onodera, H. (2020, January 7\u20139). On-chip memory optimized CNN accelerator with efficient partial-sum accumulation. Proceedings of the 2020 on Great Lakes Symposium on VLSI, Beijing, China.","DOI":"10.1145\/3386263.3406925"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Ma, Y., Cao, Y., Vrudhula, S., and Seo, J.-S. (2017, January 22\u201324). Optimizing loop operation and dataflow in FPGA acceleration of deep convolutional neural networks. Proceedings of the 2017 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays, Monterey, CA, USA.","DOI":"10.1145\/3020078.3021736"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1354","DOI":"10.1109\/TVLSI.2018.2815603","article-title":"Optimizing the convolution operation to accelerate deep neural networks on FPGA","volume":"26","author":"Ma","year":"2018","journal-title":"IEEE Trans. Very Large Scale Integr. (VLSI) Syst."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1109\/JSSC.2016.2616357","article-title":"Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks","volume":"52","author":"Chen","year":"2017","journal-title":"IEEE J. Solid-State Circuits"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Kwon, H., Samajdar, A., and Krishna, T. (2017, January 19\u201320). Rethinking NoCs for spatial neural network accelerators. Proceedings of the 2017 Eleventh IEEE\/ACM International Symposium on Networks-on-Chip (NOCS), Seoul, Korea.","DOI":"10.1145\/3130218.3130230"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Kwon, H., Samajdar, A., and Krishna, T. (2018, January 24\u201328). Maeri: Enabling flexible dataflow mapping over dnn accelerators via reconfigurable interconnects. Proceedings of the 23rd International Conference on Architectural Support for Programming Languages and Operating Systems, Williamsburg, VA, USA.","DOI":"10.1145\/3173162.3173176"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1109\/MM.2017.54","article-title":"Using dataflow to optimize energy efficiency of deep neural network accelerators","volume":"37","author":"Chen","year":"2017","journal-title":"IEEE Micro"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"292","DOI":"10.1109\/JETCAS.2019.2910232","article-title":"Eyeriss v2: A flexible accelerator for emerging deep neural networks on mobile devices","volume":"9","author":"Chen","year":"2019","journal-title":"IEEE J. Emerg. Sel. Topics Circuits Syst."},{"key":"ref_22","unstructured":"NVDLA (2022, February 01). Unit Description. Available online: http:\/\/nvdla.org\/hw\/v1\/ias\/unit_description.html."},{"key":"ref_23","first-page":"1217","article-title":"A resource-limited hardware accelerator for convolutional neural networks in embedded vision applications","volume":"64","author":"Moini","year":"2017","journal-title":"IEEE Trans. Circuits Syst. II Express Briefs"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"187754","DOI":"10.1109\/ACCESS.2020.3031055","article-title":"CNN acceleration with hardware-efficient dataflow for super-resolution","volume":"8","author":"Lee","year":"2020","journal-title":"IEEE Access"},{"key":"ref_25","unstructured":"(2022, February 01). Xilinx Zynq UltraScale+ MPSoC ZCU102 Evaluation Kit. Available online: https:\/\/www.xilinx.com\/products\/boards-and-kits\/ek-u1-zcu102-g.html."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"El-Sawy, A., El-Bakry, H., and Loey, M. (2016, January 24\u201326). CNN for handwritten arabic digits recognition based on LeNet-5. Proceedings of the International Conference on advanced INTELLIGENT Systems and Informatics, Cairo, Egypt.","DOI":"10.1007\/978-3-319-48308-5_54"},{"key":"ref_27","unstructured":"(2022, February 01). Modified National Institute of Standards and Technology (MNIST) Database. Available online: http:\/\/yann.lecun.com\/exdb\/mnist\/."},{"key":"ref_28","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_30","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. (2018, January 18\u201323). MobileNetV2: Inverted residuals and linear bottlenecks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_32","unstructured":"Tan, M., and Le, Q. (2019). Efficientnet: Rethinking model scaling for convolutional neural networks. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/8\/3095\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:56:12Z","timestamp":1760136972000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/8\/3095"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,18]]},"references-count":32,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2022,4]]}},"alternative-id":["s22083095"],"URL":"https:\/\/doi.org\/10.3390\/s22083095","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,4,18]]}}}