{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,2]],"date-time":"2026-08-02T17:41:45Z","timestamp":1785692505009,"version":"3.56.0"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"19","license":[{"start":{"date-parts":[[2022,5,29]],"date-time":"2022-05-29T00:00:00Z","timestamp":1653782400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,5,29]],"date-time":"2022-05-29T00:00:00Z","timestamp":1653782400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001778","name":"Deakin University","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001778","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2022,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Progress is being made to deploy convolutional neural networks (CNNs) into the Internet of Things (IoT) edge devices for handling image analysis tasks locally. These tasks require low-latency and low-power computation on low-resource IoT edge devices. However, CNN-based algorithms, e.g. YOLOv2, typically contain millions of parameters. With the increase in the CNN\u2019s depth, filters are increased by a power of two. A large number of filters and operations could lead to frequent off-chip memory access that affects the operation speed and power consumption of the device. Therefore, it is a challenge to map a deep CNN into a low-resource edge IoT platform. To address this challenge, we present a resource-constrained Field-Programmable Gate Array implementation of YOLOv2 with optimized data transfer and computing efficiency. Firstly, a scalable cross-layer dataflow strategy is proposed which allows on-chip data transfer between different types of layers, and offers flexible off-chip data transfer when the intermediate results are unaffordable on-chip. Next, a filter-level data-reuse dataflow strategy together with a filter-level parallel multiply-accumulate operation computing processing elements array is developed. Finally, multi-level sliding buffers are developed to optimize the convolutional computing loop and reuse the input feature maps and weights. Experiment results show that our implementation has achieved 4.8\u00a0W of low-power consumption for executing YOLOv2, an 8-bit deep CNN containing 50.6\u00a0MB weights, using low-resource of 8.3 Mbits on-chip memory. The throughput and power efficiency are 100.33 GOP\/s and 20.90 GOP\/s\/W, respectively.<\/jats:p>","DOI":"10.1007\/s00521-022-07351-w","type":"journal-article","created":{"date-parts":[[2022,5,28]],"date-time":"2022-05-28T23:03:25Z","timestamp":1653779005000},"page":"16989-17006","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":23,"title":["Resource-constrained FPGA implementation of YOLOv2"],"prefix":"10.1007","volume":"34","author":[{"given":"Zhichao","family":"Zhang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"M. A. Parvez","family":"Mahmud","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6292-1214","authenticated-orcid":false,"given":"Abbas Z.","family":"Kouzani","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,5,29]]},"reference":[{"issue":"4","key":"7351_CR1","doi-asserted-by":"publisher","first-page":"2167","DOI":"10.1109\/COMST.2020.3007787","volume":"22","author":"Y Shi","year":"2020","unstructured":"Shi Y, Yang K, Jiang T, Zhang J, Letaief KB (2020) Communication-efficient edge AI: algorithms and systems. IEEE Commun Surv Tutor 22(4):2167\u20132191","journal-title":"IEEE Commun Surv Tutor"},{"key":"7351_CR2","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2020.3041781","author":"C Xu","year":"2020","unstructured":"Xu C, Jiang S, Luo G, Sun G, An N, Huang G, Liu X (2020) The case for FPGA-based edge computing. IEEE Trans Mob Comput. https:\/\/doi.org\/10.1109\/TMC.2020.3041781","journal-title":"IEEE Trans Mob Comput"},{"issue":"3","key":"7351_CR3","doi-asserted-by":"publisher","first-page":"243","DOI":"10.1145\/3007787.3001163","volume":"44","author":"S Han","year":"2016","unstructured":"Han S, Liu X, Mao H, Pu J, Pedram A, Horowitz MA, Dally WJ (2016) EIE: Efficient inference engine on compressed deep neural network. ACM SIGARCH Comput Archit News 44(3):243\u2013254","journal-title":"ACM SIGARCH Comput Archit News"},{"key":"7351_CR4","doi-asserted-by":"crossref","unstructured":"Liu Z, Zheng T, Xu G, Yang Z, Liu H, Cai D (2020) Training-time-friendly network for real-time object detection. In: proceedings of the AAAI conference on artificial intelligence, vol 07. pp 11685\u201311692","DOI":"10.1609\/aaai.v34i07.6838"},{"key":"7351_CR5","unstructured":"Zou Z, Shi Z, Guo Y, Ye J (2019) Object detection in 20 years: a survey. arXiv preprint arXiv:190505055"},{"issue":"5","key":"7351_CR6","doi-asserted-by":"publisher","first-page":"1327","DOI":"10.1007\/s00521-019-04550-w","volume":"32","author":"Z Zhang","year":"2020","unstructured":"Zhang Z, Kouzani AZ (2020) Implementation of DNNs on IoT devices. Neural Comput Appl 32(5):1327\u20131356","journal-title":"Neural Comput Appl"},{"key":"7351_CR7","doi-asserted-by":"crossref","unstructured":"Arshad MA, Shahriar S, Sagahyroon A (2020) On the Use of FPGAs to Implement CNNs: a Brief Review. In: 2020 International conference on computing, electronics & communications engineering (iCCECE), IEEE, pp 230\u2013236","DOI":"10.1109\/iCCECE49321.2020.9231243"},{"key":"7351_CR8","unstructured":"Murshed M, Murphy C, Hou D, Khan N, Ananthanarayanan G, Hussain F (2019) Machine learning at the network edge: A survey. arXiv preprint arXiv:190800080"},{"key":"7351_CR9","doi-asserted-by":"crossref","unstructured":"Garg D, Sharma K, Singla A (2018) Designing a green data processing device using different input\/output standards on FPGA. In: 2018 fifth international conference on parallel, distributed and grid computing (PDGC), IEEE, pp 75\u201379","DOI":"10.1109\/PDGC.2018.8745716"},{"issue":"11","key":"7351_CR10","doi-asserted-by":"publisher","first-page":"3212","DOI":"10.1109\/TNNLS.2018.2876865","volume":"30","author":"Z-Q Zhao","year":"2019","unstructured":"Zhao Z-Q, Zheng P, Xu S-t, Wu X (2019) Object detection with deep learning: a review. IEEE Trans Neural Netw Learn Syst 30(11):3212\u20133232","journal-title":"IEEE Trans Neural Netw Learn Syst"},{"issue":"2","key":"7351_CR11","doi-asserted-by":"publisher","first-page":"85","DOI":"10.1007\/s13748-019-00203-0","volume":"9","author":"A Dhillon","year":"2020","unstructured":"Dhillon A, Verma GK (2020) Convolutional neural network: a review of models, methodologies and applications to object detection. Prog Artif Intell 9(2):85\u2013112","journal-title":"Prog Artif Intell"},{"key":"7351_CR12","doi-asserted-by":"crossref","unstructured":"Redmon J, Divvala S, Girshick R, Farhadi A (2016) You only look once: unified, real-time object detection. In: proceedings of the IEEE conference on computer vision and pattern recognition, pp 779\u2013788","DOI":"10.1109\/CVPR.2016.91"},{"issue":"10","key":"7351_CR13","doi-asserted-by":"publisher","first-page":"9318","DOI":"10.1109\/JIOT.2020.2990215","volume":"7","author":"J Sanchez","year":"2020","unstructured":"Sanchez J, Sawant A, Neff C, Tabkhi H (2020) AWARE-CNN: automated workflow for application-aware real-time edge acceleration of CNNs. IEEE Internet Things J 7(10):9318\u20139329","journal-title":"IEEE Internet Things J"},{"key":"7351_CR14","doi-asserted-by":"crossref","unstructured":"Ahmad A, Pasha MA, Raza GJ (2020) Accelerating Tiny YOLOv3 using FPGA-Based Hardware\/Software Co-Design. In: 2020 IEEE international symposium on circuits and systems (ISCAS), IEEE, pp 1\u20135","DOI":"10.1109\/ISCAS45731.2020.9180843"},{"key":"7351_CR15","doi-asserted-by":"crossref","unstructured":"Yu Z, Bouganis C-S (2020) A parameterisable FPGA-tailored architecture for YOLOv3-tiny. In: international symposium on applied reconfigurable computing, Springer, pp 330-344","DOI":"10.1007\/978-3-030-44534-8_25"},{"issue":"6","key":"7351_CR16","doi-asserted-by":"publisher","first-page":"2450","DOI":"10.1109\/TCSVT.2020.3020569","volume":"31","author":"DT Nguyen","year":"2020","unstructured":"Nguyen DT, Kim H, Lee H-J (2020) Layer-specific optimization for mixed data flow with mixed precision in FPGA design for CNN-based object detectors. IEEE Trans Circuits Syst Video Technol 31(6):2450\u20132464","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"7351_CR17","doi-asserted-by":"crossref","unstructured":"Bozorgzadeh B, Covey DP, Heidenreich BA, Garris PA, Mohseni P (2014) Real-time processing of fast-scan cyclic voltammetry (FSCV) data using a field-programmable gate array (FPGA). In: 2014 36th annual international conference of the IEEE engineering in medicine and biology society, IEEE, pp 2036\u20132039","DOI":"10.1109\/EMBC.2014.6944016"},{"key":"7351_CR18","doi-asserted-by":"crossref","unstructured":"Xu J, Nie Y, Wang P, L\u00f3pez AM (2019) Training a binary weight object detector by knowledge transfer for autonomous driving. In: 2019 international conference on robotics and automation (ICRA), IEEE, pp 2379\u20132384","DOI":"10.1109\/ICRA.2019.8793743"},{"key":"7351_CR19","doi-asserted-by":"crossref","unstructured":"Dinelli G, Meoni G, Rapuano E, Fanucci L (2020) Advantages and limitations of fully on-chip CNN FPGA-based hardware accelerator. In: 2020 IEEE international symposium on circuits and systems (ISCAS), IEEE, pp 1\u20135","DOI":"10.1109\/ISCAS45731.2020.9180867"},{"issue":"2020","key":"7351_CR20","doi-asserted-by":"publisher","first-page":"116569","DOI":"10.1109\/ACCESS.2020.3004198","volume":"8","author":"Z Wang","year":"2020","unstructured":"Wang Z, Xu K, Wu S, Liu L, Liu L, Wang D (2020) Sparse-YOLO: hardware\/software co-design of an FPGA accelerator for YOLOv2. IEEE Access 8(2020):116569\u2013116585","journal-title":"IEEE Access"},{"issue":"2020","key":"7351_CR21","doi-asserted-by":"publisher","first-page":"105455","DOI":"10.1109\/ACCESS.2020.3000009","volume":"8","author":"S Li","year":"2020","unstructured":"Li S, Luo Y, Sun K, Yadav N, Choi KK (2020) A novel FPGA accelerator design for real-time and ultra-low power deep convolutional neural networks compared with titan X GPU. IEEE Access 8(2020):105455\u2013105471","journal-title":"IEEE Access"},{"key":"7351_CR22","unstructured":"Gschwend D (2020) Zynqnet: an fpga-accelerated embedded convolutional neural network. arXiv preprint arXiv:200506892"},{"issue":"3","key":"7351_CR23","doi-asserted-by":"publisher","first-page":"481","DOI":"10.1007\/s11554-020-00977-w","volume":"18","author":"K Xu","year":"2021","unstructured":"Xu K, Wang X, Liu X, Cao C, Li H, Peng H, Wang D (2021) A dedicated hardware accelerator for real-time acceleration of YOLOv2. J Real-Time Image Proc 18(3):481\u2013492","journal-title":"J Real-Time Image Proc"},{"issue":"8","key":"7351_CR24","doi-asserted-by":"publisher","first-page":"1861","DOI":"10.1109\/TVLSI.2019.2905242","volume":"27","author":"DT Nguyen","year":"2019","unstructured":"Nguyen DT, Nguyen TN, Kim H, Lee H-J (2019) A high-throughput and power-efficient FPGA implementation of YOLO CNN for object detection. IEEE Trans Very Large Scale Integr (VLSI) Syst 27(8):1861\u20131873","journal-title":"IEEE Trans Very Large Scale Integr (VLSI) Syst"},{"key":"7351_CR25","doi-asserted-by":"crossref","unstructured":"Ding C, Wang S, Liu N, Xu K, Wang Y, Liang Y (2019) REQ-YOLO: a resource-aware, efficient quantization framework for object detection on FPGAs. In: proceedings of the 2019 ACM\/SIGDA international symposium on field-programmable gate arrays, pp 33\u201342","DOI":"10.1145\/3289602.3293904"},{"key":"7351_CR26","doi-asserted-by":"crossref","unstructured":"Jacob B, Kligys S, Chen B, Zhu M, Tang M, Howard A, Adam H, Kalenichenko D (2018) Quantization and training of neural networks for efficient integer-arithmetic-only inference. In: proceedings of the IEEE conference on computer vision and pattern recognition, pp 2704\u20132713","DOI":"10.1109\/CVPR.2018.00286"},{"issue":"2020","key":"7351_CR27","doi-asserted-by":"publisher","first-page":"70","DOI":"10.1016\/j.neunet.2019.12.027","volume":"125","author":"Y Yang","year":"2020","unstructured":"Yang Y, Deng L, Wu S, Yan T, Xie Y, Li G (2020) Training high-performance and large-scale deep neural networks with full 8-bit integers. Neural Netw 125(2020):70\u201382","journal-title":"Neural Netw"},{"key":"7351_CR28","doi-asserted-by":"crossref","unstructured":"Abdiyeva K, Tibeyev T, Lukac M (2020) Capacity limits of fully binary CNN. In: 2020 IEEE 50th international symposium on multiple-valued logic (ISMVL), IEEE, pp 206\u2013211","DOI":"10.1109\/ISMVL49045.2020.000-4"},{"key":"7351_CR29","doi-asserted-by":"crossref","unstructured":"Guan Y, Liang H, Xu N, Wang W, Shi S, Chen X, Sun G, Zhang W, Cong J (2017) FP-DNN: an automated framework for mapping deep neural networks onto FPGAs with RTL-HLS hybrid templates. In: 2017 IEEE 25th annual international symposium on field-programmable custom computing machines (FCCM), IEEE, pp 152\u2013159","DOI":"10.1109\/FCCM.2017.25"},{"key":"7351_CR30","doi-asserted-by":"crossref","unstructured":"Redmon J, Farhadi A (2017) YOLO9000: better, faster, stronger. In: proceedings of the IEEE conference on computer vision and pattern recognition, pp 7263\u20137271","DOI":"10.1109\/CVPR.2017.690"},{"issue":"2","key":"7351_CR31","doi-asserted-by":"publisher","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","volume":"88","author":"M Everingham","year":"2010","unstructured":"Everingham M, Van Gool L, Williams CK, Winn J, Zisserman A (2010) The pascal visual object classes (voc) challenge. Int J Comput Vision 88(2):303\u2013338","journal-title":"Int J Comput Vision"},{"key":"7351_CR32","unstructured":"Redmon J (2018) yolov2-voc.cfg. https:\/\/github.com\/pjreddie\/darknet\/blob\/master\/cfg\/yolov2-voc.cfg"},{"key":"7351_CR33","unstructured":"Joseph R (2016) YOLO: real-time object detection. https:\/\/pjreddie.com\/darknet\/yolov2\/"},{"key":"7351_CR34","doi-asserted-by":"crossref","unstructured":"Stanisz J, Lis K, Gorgon M (2021) Implementation of the pointpillars network for 3D object detection in reprogrammable heterogeneous devices using FINN. J Signal Process Syst, 1\u201316","DOI":"10.36227\/techrxiv.12593555.v1"},{"issue":"3","key":"7351_CR35","doi-asserted-by":"publisher","first-page":"282","DOI":"10.3390\/electronics10030282","volume":"10","author":"N Zhang","year":"2021","unstructured":"Zhang N, Wei X, Chen H, Liu W (2021) FPGA implementation for CNN-based optical remote sensing object detection. Electronics 10(3):282","journal-title":"Electronics"},{"key":"7351_CR36","doi-asserted-by":"crossref","unstructured":"Wang J, Gu S (2021) FPGA implementation of object detection accelerator based on Vitis-AI. In: 2021 11th international conference on information science and technology (ICIST), IEEE, pp 571\u2013577","DOI":"10.1109\/ICIST52614.2021.9440554"},{"issue":"2021","key":"7351_CR37","first-page":"1","volume":"2","author":"J Kusyk","year":"2021","unstructured":"Kusyk J, Saeed SM, Uyar MU (2021) Survey on quantum circuit compilation for noisy intermediate-scale quantum computers: artificial intelligence to heuristics. IEEE Trans Quant Eng 2(2021):1\u201316","journal-title":"IEEE Trans Quant Eng"},{"key":"7351_CR38","unstructured":"Adaptable & real-time AI inference acceleration. (2022). https:\/\/github.com\/Xilinx\/Vitis-AI"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-022-07351-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-022-07351-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-022-07351-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,9,22]],"date-time":"2022-09-22T13:51:08Z","timestamp":1663854668000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-022-07351-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,5,29]]},"references-count":38,"journal-issue":{"issue":"19","published-print":{"date-parts":[[2022,10]]}},"alternative-id":["7351"],"URL":"https:\/\/doi.org\/10.1007\/s00521-022-07351-w","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,5,29]]},"assertion":[{"value":"13 September 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 April 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 May 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}