{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T01:41:54Z","timestamp":1781487714174,"version":"3.54.1"},"reference-count":45,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2020,4,11]],"date-time":"2020-04-11T00:00:00Z","timestamp":1586563200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,4,11]],"date-time":"2020-04-11T00:00:00Z","timestamp":1586563200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Yuyou Talent Support Plan of North China University of Technology","award":["107051360019XN132\/017"],"award-info":[{"award-number":["107051360019XN132\/017"]}]},{"name":"The Fundamental Research Funds for Beijing Universities","award":["110052971803\/037"],"award-info":[{"award-number":["110052971803\/037"]}]},{"name":"Special Research Foundation of North China University of Technology","award":["PXM2017_014212_000014"],"award-info":[{"award-number":["PXM2017_014212_000014"]}]},{"DOI":"10.13039\/501100004826","name":"Natural Science Foundation of Beijing Municipality","doi-asserted-by":"publisher","award":["4162022"],"award-info":[{"award-number":["4162022"]}],"id":[{"id":"10.13039\/501100004826","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Hum. Cent. Comput. Inf. Sci."],"published-print":{"date-parts":[[2020,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Object detection methods aim to identify all target objects in the target image and determine the categories and position information in order to achieve machine vision understanding. Numerous approaches have been proposed to solve this problem, mainly inspired by methods of computer vision and deep learning. However, existing approaches always perform poorly for the detection of small, dense objects, and even fail to detect objects with random geometric transformations. In this study, we compare and analyse mainstream object detection algorithms and propose a multi-scaled deformable convolutional object detection network to deal with the challenges faced by current methods. Our analysis demonstrates a strong performance on par, or even better, than state of the art methods. We use deep convolutional networks to obtain multi-scaled features, and add deformable convolutional structures to overcome geometric transformations. We then fuse the multi-scaled features by up sampling, in order to implement the final object recognition and region regress. Experiments prove that our suggested framework improves the accuracy of detecting small target objects with geometric deformation, showing significant improvements in the trade-off between accuracy and speed.<\/jats:p>","DOI":"10.1186\/s13673-020-00219-9","type":"journal-article","created":{"date-parts":[[2020,4,11]],"date-time":"2020-04-11T13:02:33Z","timestamp":1586610153000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":78,"title":["An improved object detection algorithm based on multi-scaled and deformable convolutional neural networks"],"prefix":"10.1186","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9779-9466","authenticated-orcid":false,"given":"Danyang","family":"Cao","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhixin","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lei","family":"Gao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2020,4,11]]},"reference":[{"key":"219_CR1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-020-08627-w","author":"L Shine","year":"2020","unstructured":"Shine L, Jiji CV (2020) Automated detection of helmet on motorcyclists from traffic surveillance videos: a comparative analysis using hand-crafted features and CNN. Multimed Tools Appl. https:\/\/doi.org\/10.1007\/s11042-020-08627-w","journal-title":"Multimed Tools Appl"},{"key":"219_CR2","doi-asserted-by":"publisher","DOI":"10.1007\/s12652-019-01344-9","author":"J Liu","year":"2019","unstructured":"Liu J, Yang Y, Lv S, Wang J, Chen H et al (2019) Attention-based BiGRU-CNN for Chinese question classification. J Ambient Intell Humaniz Comput. https:\/\/doi.org\/10.1007\/s12652-019-01344-9","journal-title":"J Ambient Intell Humaniz Comput"},{"issue":"24","key":"219_CR3","doi-asserted-by":"publisher","first-page":"35329","DOI":"10.1007\/s11042-019-08116-9","volume":"78","author":"D Cao","year":"2019","unstructured":"Cao D, Zhu M, Gao L et al (2019) An image caption method based on object detection. Multimed Tools Appl 78(24):35329\u201335350","journal-title":"Multimed Tools Appl"},{"key":"219_CR4","unstructured":"Xudong L, Mao Y, Tao L (2017) The survey of object detection based on convolutional neural networks. Appl Res Comput 34(10): 2881\u20132886\u2009+\u20092891"},{"issue":"5","key":"219_CR5","first-page":"1176","volume":"14","author":"M Aamir","year":"2018","unstructured":"Aamir M, Pu Y, Rahman Z, Abro WA, Naeem H, Ullah F, Badr AM (2018) A hybrid proposed framework for object detection and classification. J Inf Process Syst 14(5):1176\u20131194","journal-title":"J Inf Process Syst"},{"key":"219_CR6","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, et al (2016) Deep residual learning for image recognition. In: Paper presented at the IEEE conference on computer vision and pattern recognition, Las Vegas, Nevada, 26\u201330 June 2016, pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"key":"219_CR7","unstructured":"Krizhevsky A, Sutskever I, Hinton G E (2012) ImageNet classification with deep convolutional neural networks. In: Paper presented at the twenty-sixth annual conference on neural information processing systems, Lake Tahoe, Nevada, 3\u20136 December 2012, pp 1097\u20131105"},{"key":"219_CR8","doi-asserted-by":"crossref","unstructured":"Szegedy C, Liu W, Jia Y, Sermanet, P, Reed S (2015) Going deeper with convolutions. In: Paper presented at the IEEE conference on computer vision and pattern recognition, Boston, Massachusetts, 7\u201312 June 2015, pp 1\u20139","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"219_CR9","unstructured":"Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition. In: Paper presented at the international conference on learning representations, San Diego, California, 7\u20139 May 2015, pp 1\u201314"},{"key":"219_CR10","unstructured":"Andrew G, Menglong Zhu, Bo Chen, Dmitry Kalenichenko (2017) MobileNets: efficient convolutional neural networks for mobile vision. In: Paper presented at the IEEE conference on computer vision and pattern recognition, Honolulu, Hawaii, 21\u201326 July 2017"},{"issue":"3","key":"219_CR11","doi-asserted-by":"publisher","first-page":"178","DOI":"10.1049\/iet-cdt.2018.5026","volume":"13","author":"FF dos Santos","year":"2019","unstructured":"dos Santos FF, Carro L, Rech P (2019) Kernel and layer vulnerability factor to evaluate object detection reliability in GPUs. IET Comput Digital Tech 13(3):178\u2013186","journal-title":"IET Comput Digital Tech"},{"key":"219_CR12","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1186\/s13673-019-0191-8","volume":"9","author":"MJJ Ghrabat","year":"2019","unstructured":"Ghrabat MJJ, Ma G, Maolood IY et al (2019) An effective image retrieval based on optimized genetic algorithm utilized a novel SVM-based convolutional neural network classifier. Human-centric Comput Inf Sci 9:31","journal-title":"Human-centric Comput Inf Sci"},{"key":"219_CR13","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1186\/s13673-019-0203-8","volume":"9","author":"F Zhang","year":"2019","unstructured":"Zhang F, Wu T, Pan J et al (2019) Human motion recognition based on SVM in VR art media interaction environment. Human-centric Comput Inf Sci 9:40","journal-title":"Human-centric Comput Inf Sci"},{"key":"219_CR14","doi-asserted-by":"crossref","unstructured":"Girshick R, Donahue J, Darrell T, Malik J (2014) Rich feature hierarchies for accurate object detection and semantic segmentation. In: Paper presented at the IEEE conference on computer vision and pattern recognition, Columbus, Ohio, 23\u201328 June 2014","DOI":"10.1109\/CVPR.2014.81"},{"key":"219_CR15","doi-asserted-by":"crossref","unstructured":"Girshick R (2015) Fast R-CNN. In: Paper presented at IEEE international conference on computer vision, Santiago, Chile, 7\u201313 December 2015, pp 1440\u20131448","DOI":"10.1109\/ICCV.2015.169"},{"issue":"6","key":"219_CR16","doi-asserted-by":"publisher","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","volume":"39","author":"S Ren","year":"2015","unstructured":"Ren S, He K, Girshick R et al (2015) Faster R-CNN: towards real-time object detection with region proposal networks. IEEE Trans Pattern Anal Mach Intell 39(6):1137\u20131149","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"6","key":"219_CR17","doi-asserted-by":"publisher","first-page":"1300","DOI":"10.21629\/JSEE.2018.06.17","volume":"29","author":"C Jinbo","year":"2018","unstructured":"Jinbo C, Zhiheng W, Hengyu L (2018) Real-time object segmentation based on convolutional neural network with saliency optimization for picking. J Syst Eng Electron 29(6):1300\u20131307","journal-title":"J Syst Eng Electron"},{"key":"219_CR18","doi-asserted-by":"crossref","unstructured":"Redmon J, Divvala S, Girshick R, et al (2016) You only look once: unified, real-time object detection. In: Paper presented at the IEEE conference on computer vision and pattern recognition, Las Vegas, Nevada, 26\u201330 June 2016, pp 779\u2013788","DOI":"10.1109\/CVPR.2016.91"},{"key":"219_CR19","doi-asserted-by":"crossref","unstructured":"Liu W, Anguelov D, Erhan D, Szegedy C, Reed S, Fu C-Y, Berg AC (2016) SSD: single shot multibox detector. In: Paper presented at the 14th European conference on computer vision, Amsterdam, The Netherlands, 11\u201314 October 2016","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"219_CR20","doi-asserted-by":"crossref","unstructured":"Lin TY, Doll\u00e1r P, Girshick R, He K, Hariharan B, Belongie S (2017) Feature pyramid networks for object detection. In: Paper presented at the IEEE conference on computer vision and pattern recognition, Honolulu, Hawaii, 21\u201326 July 2017, pp 2117\u20132125","DOI":"10.1109\/CVPR.2017.106"},{"key":"219_CR21","doi-asserted-by":"crossref","unstructured":"Bodla N, Singh B, Chellappa R, et al (2017) Soft-NMS\u2014improving object detection with one line of code. In: Paper presented at IEEE international conference on computer vision, Venice, Italy, 22\u2013 October 2017","DOI":"10.1109\/ICCV.2017.593"},{"key":"219_CR22","doi-asserted-by":"publisher","first-page":"33","DOI":"10.1186\/s13673-018-0156-3","volume":"8","author":"A Sun","year":"2018","unstructured":"Sun A, Li Y, Huang Y et al (2018) Facial expression recognition using optimized active regions. Human-centric Comput Inf Sci 8:33","journal-title":"Human-centric Comput Inf Sci"},{"key":"219_CR23","doi-asserted-by":"publisher","first-page":"e1897","DOI":"10.1002\/cav.1897","volume":"30","author":"Y Hou","year":"2019","unstructured":"Hou Y, Luo H, Zhao W, Zhang X, Wang J, Peng J et al (2019) Multilayer feature descriptors fusion CNN models for fine-grained visual recognition. Comput Anim Virtual Worlds 30:e1897","journal-title":"Comput Anim Virtual Worlds"},{"issue":"193","key":"219_CR24","doi-asserted-by":"publisher","first-page":"102907","DOI":"10.1016\/j.cviu.2020.102907","volume":"4","author":"Longyin Wen","year":"2020","unstructured":"Wen Longyin, Dawei Du, Cai Zhaowei et al (2020) UA-DETRAC: a new benchmark and protocol for multi-object detection and tracking. Comput Vis Image Underst 4(193):102907","journal-title":"Comput Vis Image Underst"},{"key":"219_CR25","unstructured":"Redmon J, Farhadi A (2018) YOLOv3: an incremental improvement. arXiv preprint, arXiv:1804.02767v1 [cs.CV], Unpublished"},{"key":"219_CR26","unstructured":"Redmon J (2013\u20132016) Darknet: open source neural networks in c. http:\/\/pjreddie.com\/darknet\/. Accessed 30 July 2018"},{"key":"219_CR27","doi-asserted-by":"crossref","unstructured":"Redmon J, Farhadi A (2017) Yolo9000: better, faster, stronger. In: Paper presented at the IEEE conference on computer vision and pattern recognition, Honolulu, Hawaii, 21\u201326 July 2017, pp 6517\u20136525","DOI":"10.1109\/CVPR.2017.690"},{"key":"219_CR28","doi-asserted-by":"crossref","unstructured":"Girshick R, Donahue J, Darrell T, Malik J (2014) Rich feature hierarchies for accurate object detection and semantic segmentation. In: Paper presented at the IEEE conference on computer vision and pattern recognition, Columbus, Ohio, 23\u201328 June 2014, pp 580\u2013587","DOI":"10.1109\/CVPR.2014.81"},{"key":"219_CR29","doi-asserted-by":"crossref","unstructured":"Brink H, Vadapalli HB (2017) Deformable part models with CNN features for facial landmark detection under occlusion. In: Paper presented at the South African Institute of Computer Scientists and Information Technologists, ACM, Thaba\\\u201dNchu, South Africa, 26\u201328 September 2017, pp 1\u20139","DOI":"10.1145\/3129416.3129451"},{"key":"219_CR30","doi-asserted-by":"crossref","unstructured":"Jeon Y, Kim J (2017) Active convolution: learning the shape of convolution for image classification. In: Paper presented at the IEEE conference on computer vision and pattern recognition, Honolulu, Hawaii, 21\u201326 July 2017, pp 1846\u20131854","DOI":"10.1109\/CVPR.2017.200"},{"key":"219_CR31","unstructured":"Jifeng D, Haozhi Q, Yuwen X, Yi L, Guodong Z, Han H and Yichen W (2017) Deformable convolutional networks. In: Paper presented at IEEE international conference on computer vision, Venice, Italy, 22\u201329 October 2017, pp 764\u2013773"},{"key":"219_CR32","doi-asserted-by":"crossref","unstructured":"Mordan T, Thome N, Cord M, Henaff G (2017) Deformable part-based fully convolutional network for object detection. In: Paper presented at British machine vision conference (BMVC), London, United Kingdom, 4\u20137 Sep 2017","DOI":"10.5244\/C.31.88"},{"issue":"1","key":"219_CR33","first-page":"176","volume":"14","author":"H Zeng","year":"2018","unstructured":"Zeng H, Liu Y, Li S, Che J, Wang X (2018) Convolutional neural network based multi-feature fusion for non-rigid 3D model retrieval. J Inf Process Syst 14(1):176\u2013190","journal-title":"J Inf Process Syst"},{"issue":"9","key":"219_CR34","doi-asserted-by":"publisher","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","volume":"37","author":"K He","year":"2015","unstructured":"He K, Zhang X, Ren S et al (2015) Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE Trans Pattern Anal Mach Intell 37(9):1904\u20131916","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"4","key":"219_CR35","doi-asserted-by":"publisher","first-page":"191","DOI":"10.1049\/trit.2018.1026","volume":"3","author":"S Sun","year":"2018","unstructured":"Sun S, Yin Y, Wang X, Xu D, Wu W, Gu Q (2018) Fast object detection based on binary deep convolution neural networks. CAAI Trans Intell Technol 3(4):191\u2013197","journal-title":"CAAI Trans Intell Technol"},{"key":"219_CR36","doi-asserted-by":"publisher","first-page":"29","DOI":"10.1186\/s13673-018-0152-7","volume":"8","author":"W Song","year":"2018","unstructured":"Song W, Zou S, Tian Y, Fong S, Cho K (2018) Classifying 3D objects in LiDAR point clouds with a back-propagation neural network. Human-centric Comput Inf Sci 8:29","journal-title":"Human-centric Comput Inf Sci"},{"issue":"25","key":"219_CR37","doi-asserted-by":"publisher","first-page":"1433","DOI":"10.1049\/el.2018.6712","volume":"54","author":"K Zhao","year":"2018","unstructured":"Zhao K, Zhu X, Jiang H et al (2018) Dynamic loss for one-stage object detectors in computer vision. Electron Lett 54(25):1433\u20131434","journal-title":"Electron Lett"},{"key":"219_CR38","unstructured":"Krasin I, Duerig T, Alldrin N, Ferrari V, Abu-El-Haija S, Kuznetsova A, Rom H, Uijlings J, Popov S, Veit A, Belongie S, Gomes V, Gupta A, Sun C, Chechik G, Cai D, Feng Z, Narayanan D, Murphy K (2017) Openimages: a public dataset for large-scale multi-label and multi-class image classification. Dataset available from https:\/\/github.com\/openimages. Accessed 30 July 2018"},{"issue":"2","key":"219_CR39","doi-asserted-by":"publisher","first-page":"154","DOI":"10.1007\/s11263-013-0620-5","volume":"104","author":"JRR Uijlings","year":"2013","unstructured":"Uijlings JRR et al (2013) Selective search for object recognition. Int J Comput Vis 104(2):154\u2013171","journal-title":"Int J Comput Vis"},{"key":"219_CR40","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, et al (2009) ImageNet: a large-scale hierarchical image database. In: Paper presented at IEEE Conference on computer vision and pattern recognition, Miami, Florida, 20\u201325 June 2009, pp 248\u2013255","DOI":"10.1109\/CVPR.2009.5206848"},{"issue":"2","key":"219_CR41","doi-asserted-by":"publisher","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","volume":"88","author":"M Everingham","year":"2010","unstructured":"Everingham M, Van Gool L, Williams CK, Winn J, Zisserman A (2010) The pascal visual object classes (voc) challenge. Int J Comput Vis 88(2):303\u2013338","journal-title":"Int J Comput Vis"},{"key":"219_CR42","doi-asserted-by":"crossref","unstructured":"Long J, Shelhamer E, Darrell T (2015) Fully convolutional networks for semantic segmentation. Paper presented at the IEEE conference on computer vision and pattern recognition, Boston, Massachusetts, 7\u201312 June 2015, pp 3431\u20133440","DOI":"10.1109\/CVPR.2015.7298965"},{"issue":"8","key":"219_CR43","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1186\/s13673-018-0148-3","volume":"2018","author":"E Gultepe","year":"2018","unstructured":"Gultepe E, Makrehchi M (2018) Improving clustering performance using independent component analysis and unsupervised feature learning. Human-centric Computi Inf Sci 2018(8):25","journal-title":"Human-centric Computi Inf Sci"},{"key":"219_CR44","doi-asserted-by":"crossref","unstructured":"Wang K, Zhang D, Li Y, et al (2017) Cost-effective active learning for deep image classification. IEEE Trans Circuits Systems Video Technol (99):1\u20131","DOI":"10.1109\/TCSVT.2016.2589879"},{"key":"219_CR45","doi-asserted-by":"crossref","unstructured":"Huang J, Guadarrama S, Murphy K, et al (2017) Speed\/accuracy trade-offs for modern convolutional object detectors. In: Paper presented at the IEEE conference on computer vision and pattern recognition, Honolulu, Hawaii, 21\u201326 July 2017, pp 3296\u20133297","DOI":"10.1109\/CVPR.2017.351"}],"container-title":["Human-centric Computing and Information Sciences"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13673-020-00219-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13673-020-00219-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13673-020-00219-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,7,30]],"date-time":"2021-07-30T12:07:28Z","timestamp":1627646848000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1186\/s13673-020-00219-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,4,11]]},"references-count":45,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,12]]}},"alternative-id":["219"],"URL":"https:\/\/doi.org\/10.1186\/s13673-020-00219-9","relation":{},"ISSN":["2192-1962"],"issn-type":[{"value":"2192-1962","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,4,11]]},"assertion":[{"value":"26 September 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 March 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 April 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare that they have no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"14"}}