{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,19]],"date-time":"2026-07-19T03:23:28Z","timestamp":1784431408592,"version":"3.55.0"},"reference-count":25,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,1,4]],"date-time":"2024-01-04T00:00:00Z","timestamp":1704326400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,1,4]],"date-time":"2024-01-04T00:00:00Z","timestamp":1704326400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Shaanxi Provincial Department of Education 2022 General Special Research Program Projects","award":["22JK0471"],"award-info":[{"award-number":["22JK0471"]}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Foundation Project","doi-asserted-by":"crossref","award":["2022MD723841"],"award-info":[{"award-number":["2022MD723841"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["EURASIP J. Adv. Signal Process."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The purpose of human object detection is to obtain the number of people and their position in images, which is one of the core problems in the field of machine vision. However, the high missing detection rate from small- and medium-sized human bodies due to the large variety of human scale in human object detection tasks still influences the performance of human object detection. To solve the above problem, this paper proposed an improved ASPP_BiFPN_YOLOv4 (ABYOLOv4) method to detect human object detection. In detail, Atrous Spatial Pyramid Pooling (ASPP) module was used to replace the original Spatial Pyramid Pooling module to increase the receptive field level of the network and improve the perception ability of multi-scale targets. Then, the original Path Aggregation Network (PANet) multi-scale fusion module was replaced by the self-built bi-layer bidirectional feature pyramid network (Bi-FPN). Meanwhile, a new feature was imported into the proposed model to reuse the mid- and low-level features, which could enhance the ability of the network to express the characteristics of small- and medium-sized targets. Finally, the standard convolution in Bi-FPN was replaced by depth-separable convolution to make the network achieve the balance of accuracy and the number of parameters. To identify the performance of the proposed ABYOLOv4 model, the human object detection experiment is carried out by using the public data set of VOC2007 and VOC2012, the improved YOLOv4 algorithm is 0.5% higher than the original AP algorithm, and the weight file size of the model is reduced by 45.3\u00a0M. The experimental results demonstrated that the proposed ABYOLOv4 network has higher accuracy and lower computational cost for human target detection.<\/jats:p>","DOI":"10.1186\/s13634-023-01105-z","type":"journal-article","created":{"date-parts":[[2024,1,4]],"date-time":"2024-01-04T15:05:48Z","timestamp":1704380748000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["ABYOLOv4: improved YOLOv4 human object detection based on enhanced multi-scale feature fusion"],"prefix":"10.1186","volume":"2024","author":[{"given":"Rui","family":"Li","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xin","family":"Zeng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9144-4824","authenticated-orcid":false,"given":"Shiqiang","family":"Yang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qi","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"An","family":"Yan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dexin","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,1,4]]},"reference":[{"key":"1105_CR1","doi-asserted-by":"publisher","unstructured":"B. Hariharan, P. Arbelaez, R. Girshick, J. Malik, Simultaneous detection and segmentation. In Computer Vision\u2013ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, Proceedings, Part VII 13, 297\u2013312 (2014). https:\/\/doi.org\/10.5220\/0009142905550561","DOI":"10.5220\/0009142905550561"},{"key":"1105_CR2","doi-asserted-by":"publisher","unstructured":"B. Hariharan, P. Arbel\u00e1ez, R. Girshick, J. Malik, Hypercolumns for object segmentation and fine-grained localization, In Proceedings of the IEEE conference on computer vision and pattern recognition. 447\u2013456 (2015). https:\/\/doi.org\/10.1109\/cvpr.2015.7298642","DOI":"10.1109\/cvpr.2015.7298642"},{"key":"1105_CR3","doi-asserted-by":"publisher","unstructured":"J. Dai, K. He, and J. Sun, Instance-aware semantic segmentation via multi-task network cascades. In Proceedings of the IEEE conference on computer vision and pattern recognition, 3150\u20133158 (2016). https:\/\/doi.org\/10.1109\/cvpr.2016.343","DOI":"10.1109\/cvpr.2016.343"},{"key":"1105_CR4","doi-asserted-by":"publisher","unstructured":"K. He, G. Gkioxari, P. Dollar, and R. Girshick, Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, 2961\u20132969 (2017). https:\/\/doi.org\/10.48550\/arXiv.1703.06870","DOI":"10.48550\/arXiv.1703.06870"},{"key":"1105_CR5","doi-asserted-by":"publisher","unstructured":"A. Karpathy, L. Fei-Fei, Deep visual-semantic alignments for generating image descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, 3128\u20133137 (2015). https:\/\/doi.org\/10.1109\/cvpr.2015.7298932","DOI":"10.1109\/cvpr.2015.7298932"},{"key":"1105_CR6","doi-asserted-by":"publisher","unstructured":"Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov et al., Show, attend and tell: Neural image caption generation with visual attention. In International conference on machine learning, 2048\u20132057 (2015). https:\/\/doi.org\/10.1109\/cvpr.2015.7298935","DOI":"10.1109\/cvpr.2015.7298935"},{"issue":"6","key":"1105_CR7","doi-asserted-by":"publisher","first-page":"1367","DOI":"10.1109\/tpami.2017.2708709","volume":"40","author":"Q Wu","year":"2018","unstructured":"Q. Wu, C. Shen, P. Wang, A. Dick, A. van den Hengel, Image captioning and visual question answering based on attributes and external knowledge. IEEE Trans. Pattern Anal. 40(6), 1367\u20131381 (2018). https:\/\/doi.org\/10.1109\/tpami.2017.2708709","journal-title":"IEEE Trans. Pattern Anal."},{"issue":"10","key":"1105_CR8","doi-asserted-by":"publisher","first-page":"2896","DOI":"10.1109\/cvpr.2016.95","volume":"28","author":"K Kang","year":"2017","unstructured":"K. Kang, H. Li, J. Yan, X. Zeng, B. Yang, T. Xiao et al., T-cnn: Tubelets with convolutional neural networks for object detection from videos. IEEE Trans. Circ. Syst. Vid. 28(10), 2896\u20132907 (2017). https:\/\/doi.org\/10.1109\/cvpr.2016.95","journal-title":"IEEE Trans. Circ. Syst. Vid."},{"issue":"3","key":"1105_CR9","doi-asserted-by":"publisher","first-page":"309","DOI":"10.1145\/1015706.1015720","volume":"3","author":"C Rother","year":"2004","unstructured":"C. Rother, V. Kolmogorov, A. Blake, \"GrabCut\u201d: interactive foreground extraction using iterated graph cuts. CM Trans. Graph. 3(3), 309\u2013314 (2004). https:\/\/doi.org\/10.1145\/1015706.1015720","journal-title":"CM Trans. Graph."},{"issue":"3","key":"1105_CR10","doi-asserted-by":"publisher","first-page":"216","DOI":"10.1016\/j.patrec.2009.10.005","volume":"31","author":"SD Han","year":"2010","unstructured":"S.D. Han, W.B. Tao, X.L. Wu, X.C. Tai, T.J. Wang, Fast image segmentation based on multilevel banded closed-form method. Pattern Recogn. Lett. 31(3), 216\u2013225 (2010). https:\/\/doi.org\/10.1016\/j.patrec.2009.10.005","journal-title":"Pattern Recogn. Lett."},{"key":"1105_CR11","doi-asserted-by":"publisher","unstructured":"N. Dalal, B. Triggs, Histograms of oriented gradients for human detection. In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR'05), 1, 886\u2013893 (2005). https:\/\/doi.org\/10.1109\/cvpr.2005.177","DOI":"10.1109\/cvpr.2005.177"},{"key":"1105_CR12","doi-asserted-by":"publisher","unstructured":"R. Ronfard, C. Schmid, B. Triggs, Learning to parse pictures of people. Lect. Notes Comput. Sci. 700\u2013714 (2002). https:\/\/doi.org\/10.1007\/3-540-47979-1_47","DOI":"10.1007\/3-540-47979-1_47"},{"key":"1105_CR13","doi-asserted-by":"publisher","unstructured":"R. Girshick, J. Donahue, T. Darrell, J. Malik, Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition. 580\u2013587 (2014). https:\/\/doi.org\/10.1109\/cvpr.2014.81","DOI":"10.1109\/cvpr.2014.81"},{"key":"1105_CR14","doi-asserted-by":"publisher","unstructured":"R. Girshick, Fast r-cnn. In Proceedings of the IEEE international conference on computer vision. 1440\u20131448 (2015). https:\/\/doi.org\/10.1109\/iccv.2015.169","DOI":"10.1109\/iccv.2015.169"},{"issue":"6","key":"1105_CR15","doi-asserted-by":"publisher","first-page":"1137","DOI":"10.1109\/tpami.2016.2577031","volume":"39","author":"SQ Ren","year":"2017","unstructured":"S.Q. Ren, K.M. He, R. Girshick, J. Sun, Faster r-cnn: towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. 39(6), 1137\u20131149 (2017). https:\/\/doi.org\/10.1109\/tpami.2016.2577031","journal-title":"IEEE Trans. Pattern Anal."},{"key":"1105_CR16","doi-asserted-by":"publisher","unstructured":"J. Redmon, S. Divvala, R. Girshick, A. Farhadi, You only look once: unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition. 779\u2013788 (2016). https:\/\/doi.org\/10.1109\/cvpr.2016.91","DOI":"10.1109\/cvpr.2016.91"},{"key":"1105_CR17","doi-asserted-by":"publisher","unstructured":"T.Y. Lin, P. Goyal, R. Girshick, K. He, P. Doll\u00e1r, Focal loss for dense object detection. In Proceedings of the IEEE international conference on computer vision. 2980\u20132988 (2017). https:\/\/doi.org\/10.1109\/iccv.2017.324","DOI":"10.1109\/iccv.2017.324"},{"key":"1105_CR18","doi-asserted-by":"publisher","unstructured":"W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.Y. Fu, et al. SSD: single shot multibox detector. In Computer Vision\u2013ECCV 2016: 14th European Conference. Part I 14 (21\u201337) (2016). https:\/\/doi.org\/10.1007\/978-3-319-46448-02","DOI":"10.1007\/978-3-319-46448-02"},{"key":"1105_CR19","unstructured":"A. Bochkovskiy, C.Y. Wang, H.Y. M. Liao, Yolov4: optimal speed and accuracy of object detection, (2020). arXiv:2004.10934"},{"key":"1105_CR20","doi-asserted-by":"publisher","first-page":"105742","DOI":"10.1016\/j.compag.2020.105742","volume":"178","author":"D Wu","year":"2020","unstructured":"D. Wu, S. Lv, M. Jiang, H. Song, Using channel pruning-based YOLO v4 deep learning algorithm for the real-time and accurate detection of apple flowers in natural environments. Pattern Recogn. Lett. 178, 105742 (2020). https:\/\/doi.org\/10.1016\/j.compag.2020.105742","journal-title":"Pattern Recogn. Lett."},{"issue":"9","key":"1105_CR21","doi-asserted-by":"publisher","first-page":"1904","DOI":"10.1109\/tpami.2015.2389824","volume":"37","author":"K He","year":"2015","unstructured":"K. He, X. Zhang, S. Ren, J. Sun, Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE Trans. Pattern. Anal. 37(9), 1904\u20131916 (2015). https:\/\/doi.org\/10.1109\/tpami.2015.2389824","journal-title":"IEEE Trans. Pattern. Anal."},{"issue":"4","key":"1105_CR22","doi-asserted-by":"publisher","first-page":"834","DOI":"10.1109\/tpami.2017.2699184","volume":"40","author":"LC Chen","year":"2017","unstructured":"L.C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, A.L. Yuille, Deeplab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans. Pattern Anal. 40(4), 834\u2013848 (2017). https:\/\/doi.org\/10.1109\/tpami.2017.2699184","journal-title":"IEEE Trans. Pattern Anal."},{"key":"1105_CR23","doi-asserted-by":"publisher","unstructured":"M. Tan, R. Pang, Q.V. Le, Efficientdet: scalable and efficient object detection. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 10781\u201310790 (2020). https:\/\/doi.org\/10.1109\/cvpr42600.2020.01079","DOI":"10.1109\/cvpr42600.2020.01079"},{"key":"1105_CR24","doi-asserted-by":"publisher","unstructured":"S. Liu, L. Qi, H. Qin, J. Shi, J. Jia, Path aggregation network for instance segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition. 8759\u20138768 (2018). https:\/\/doi.org\/10.1109\/cvpr.2018.00913","DOI":"10.1109\/cvpr.2018.00913"},{"key":"1105_CR25","unstructured":"A.G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, Weyand, T. et al., Mobilenets: Efficient convolutional neural networks for mobile vision applications. (2017). arXiv:1704.04861"}],"container-title":["EURASIP Journal on Advances in Signal Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13634-023-01105-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13634-023-01105-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13634-023-01105-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,1,4]],"date-time":"2024-01-04T15:08:28Z","timestamp":1704380908000},"score":1,"resource":{"primary":{"URL":"https:\/\/asp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13634-023-01105-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,4]]},"references-count":25,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["1105"],"URL":"https:\/\/doi.org\/10.1186\/s13634-023-01105-z","relation":{},"ISSN":["1687-6180"],"issn-type":[{"value":"1687-6180","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,4]]},"assertion":[{"value":"29 July 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 December 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 January 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"6"}}