{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T07:07:25Z","timestamp":1777705645986,"version":"3.51.4"},"reference-count":8,"publisher":"SAGE Publications","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IFS"],"published-print":{"date-parts":[[2022,2,2]]},"abstract":"<jats:p>In recent years, the research on object detection has been intensified. A large number of object detection results are applied to our daily life, which greatly facilitates our work and life. In this paper, we propose a more effective object detection neural network model ENHANCE_YOLOV4. We studied the effects of several attention mechanisms on YOLOV4, and finally concluded that spatial attention mechanism had the best effect on YOLOV4. Therefore, based on previous studies, this paper introduces Dilated Convolution and one-by-one convolution into the spatial attention mechanism to expand the receptive field and combine channel information. Compared with CBAM and BAM, which are composed of spatial attention and channel attention, this improved spatial attention module reduces model parameters and improves detection capabilities. We built a new network model by embedding improved spatial attention module in the appropriate place in YOLOV4. And this paper proves that the detection accuracy of this network structure on the VOC data set is increased by 0.8%, and the detection accuracy on the coco data set is increased by 7%when the calculation performance is increased a little.<\/jats:p>","DOI":"10.3233\/jifs-211648","type":"journal-article","created":{"date-parts":[[2021,10,29]],"date-time":"2021-10-29T11:39:13Z","timestamp":1635507553000},"page":"2359-2368","source":"Crossref","is-referenced-by-count":15,"title":["An object detection network based on YOLOv4 and improved spatial attention mechanism"],"prefix":"10.1177","volume":"42","author":[{"given":"Zhixiong","family":"Chen","sequence":"first","affiliation":[{"name":"School of Software, Xin Jiang University, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shengwei","family":"Tian","sequence":"additional","affiliation":[{"name":"School of Software, Xin Jiang University, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Long","family":"Yu","sequence":"additional","affiliation":[{"name":"Network Center, Xin Jiang University, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liqiang","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Software, Xin Jiang University, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinyu","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Software, Xin Jiang University, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","reference":[{"key":"10.3233\/JIFS-211648_ref1","unstructured":"Krizhevsky A. , Sutskever I. and Hinton G. , ImageNet classificationwith deep convolutional neural networks[J]. Advances in Neural Information Processing Systems 25(2) (2012)."},{"issue":"9","key":"10.3233\/JIFS-211648_ref4","first-page":"1916","article-title":"Spatial pyramid pooling in deepconvolutional networks for visual recognition[J]","volume":"37","author":"He","year":"1904","journal-title":"IEEETransactions on Pattern Analysis and Machine Intelligence"},{"key":"10.3233\/JIFS-211648_ref9","doi-asserted-by":"crossref","unstructured":"Lin T.Y. , Goyal P. and Girshick R. , Focal loss for dense object detection[C]\/\/ IEEE Transactions onPattern Analysis & Machine Intelligence. IEEE (2017), 2999\u20133007.","DOI":"10.1109\/ICCV.2017.324"},{"key":"10.3233\/JIFS-211648_ref23","doi-asserted-by":"crossref","unstructured":"Woo S. , Park J. and Lee J.Y. , et al. CBAM: Convolutional Block Attention Module[C]\/\/ European Conference on Computer Vision. Springer, Cham (2018).","DOI":"10.1007\/978-3-030-01234-2_1"},{"issue":"11","key":"10.3233\/JIFS-211648_ref27","doi-asserted-by":"crossref","first-page":"1875","DOI":"10.1109\/TMM.2015.2477044","article-title":"Describing Multimedia ContentUsing Attention-Based Encoder-Decoder Networks[J]","volume":"17","author":"Cho","year":"1875","journal-title":"IEEETransactions on Multimedia"},{"key":"10.3233\/JIFS-211648_ref29","unstructured":"Mnih V. , Heess N.M.O. , Graves A. , et al. Recurrent Models of Visual Attention. MIT Press, 2014."},{"key":"10.3233\/JIFS-211648_ref36","doi-asserted-by":"crossref","unstructured":"Zeiler M.D. , Krishnan D. , Taylor G.W. , et al. Deconvolutional networks[C]\/\/ Computer Vision & Pattern Recognition. IEEE (2010).","DOI":"10.1109\/CVPR.2010.5539957"},{"key":"10.3233\/JIFS-211648_ref38","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"}],"container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"original-title":[],"link":[{"URL":"https:\/\/content.iospress.com\/download?id=10.3233\/JIFS-211648","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:44:34Z","timestamp":1777455874000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/full\/10.3233\/JIFS-211648"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,2]]},"references-count":8,"journal-issue":{"issue":"3"},"URL":"https:\/\/doi.org\/10.3233\/jifs-211648","relation":{},"ISSN":["1064-1246","1875-8967"],"issn-type":[{"value":"1064-1246","type":"print"},{"value":"1875-8967","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,2]]}}}