{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T11:36:29Z","timestamp":1783424189966,"version":"3.54.6"},"reference-count":24,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2021,11,29]],"date-time":"2021-11-29T00:00:00Z","timestamp":1638144000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Deep learning is a relatively new branch of machine learning in which computers are taught to recognize patterns in massive volumes of data. It primarily describes learning at various levels of representation, which aids in understanding data that includes text, voice, and visuals. Convolutional neural networks have been used to solve challenges in computer vision, including object identification, image classification, semantic segmentation and a lot more. Object detection in videos involves confirming the presence of the object in the image or video and then locating it accurately for recognition. In the video, modelling techniques suffer from high computation and memory costs, which may decrease performance measures such as accuracy and efficiency to identify the object accurately in real-time. The current object detection technique based on a deep convolution neural network requires executing multilevel convolution and pooling operations on the entire image to extract deep semantic properties from it. For large objects, detection models can provide superior results; however, those models fail to detect the varying size of the objects that have low resolution and are greatly influenced by noise because the features after the repeated convolution operations of existing models do not fully represent the essential characteristics of the objects in real-time. With the help of a multi-scale anchor box, the proposed approach reported in this paper enhances the detection accuracy by extracting features at multiple convolution levels of the object. The major contribution of this paper is to design a model to understand better the parameters and the hyper-parameters which affect the detection and the recognition of objects of varying sizes and shapes, and to achieve real-time object detection and recognition speeds by improving accuracy. The proposed model has achieved 84.49 mAP on the test set of the Pascal VOC-2007 dataset at 11 FPS, which is comparatively better than other real-time object detection models.<\/jats:p>","DOI":"10.3390\/fi13120307","type":"journal-article","created":{"date-parts":[[2021,11,30]],"date-time":"2021-11-30T04:48:37Z","timestamp":1638247717000},"page":"307","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["An Efficient Deep Convolutional Neural Network Approach for Object Detection and Recognition Using a Multi-Scale Anchor Box in Real-Time"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3752-7220","authenticated-orcid":false,"given":"Vijayakumar","family":"Varadarajan","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, The University of New South Wales, Sydney 1466, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6376-8346","authenticated-orcid":false,"given":"Dweepna","family":"Garg","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Devang Patel Institute of Advance Technology and Research (DEPSTAR), Faculty of Technology and Engineering (FTE), Charotar University of Science and Technology (CHARUSAT), Anand 388421, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2653-3780","authenticated-orcid":false,"given":"Ketan","family":"Kotecha","sequence":"additional","affiliation":[{"name":"Symbiosis Centre for Applied Artificial Intelligence, Symbiosis International (Deemed University), Pune 412115, India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2021,11,29]]},"reference":[{"key":"ref_1","first-page":"1","article-title":"A survey on deep learning: Algorithms, techniques, and applications","volume":"51","author":"Pouyanfar","year":"2018","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1016\/j.neucom.2015.09.116","article-title":"Deep learning for visual understanding: A review","volume":"187","author":"Guo","year":"2016","journal-title":"Neurocomputing"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1038\/nature16961","article-title":"Mastering the game of Go with deep neural networks and tree search","volume":"529","author":"Silver","year":"2016","journal-title":"Nature"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Palmer, R., West, G., and Tan, T. (2012, January 3\u20135). Scale Proportionate Histograms of Oriented Gradients for Object Detection in Co-Registered Visual and Range Data. Proceedings of the 2012 IEEE International Conference on Digital Image Computing Techniques and Applications (DICTA), Fremantle, Australia.","DOI":"10.1109\/DICTA.2012.6411699"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive Image Features from Scale-Invariant Keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_6","first-page":"407","article-title":"SURF: Speeded Up Robust Features","volume":"Volume 3951","author":"Leonardis","year":"2006","journal-title":"Computer Vision\u2013ECCV 2006, Proceedings of the 9th European Conference on Computer Vision, Graz, Austria, 7\u201313 May 2006"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","article-title":"The Pascal Visual Object Classes (VOC) Challenge","volume":"88","author":"Everingham","year":"2010","journal-title":"Int. J. Comput. Vis."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"290","DOI":"10.1016\/j.neucom.2021.06.072","article-title":"Probabilistic faster R-CNN with stochastic region proposing: Towards object detection and recognition in remote sensing imagery","volume":"459","author":"Yi","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 29). Mask r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"331","DOI":"10.2174\/2213275912666190429152319","article-title":"A Robust Real Time Object Detection and Recognition Algorithm for Multiple Objects","volume":"14","author":"Modwel","year":"2021","journal-title":"Recent Adv. Comput. Sci. Commun."},{"key":"ref_14","unstructured":"(2021, July 25). IoU Accuracy. Available online: https:\/\/www.pyimagesearch.com\/2016\/11\/07\/intersection-over-union-ioufor-object-detection\/."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, Faster, Stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_16","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_17","first-page":"171","article-title":"You Only Look Once (YOLOv3): Object Detection and Recognition for Indoor Environment","volume":"7","author":"Salam","year":"2021","journal-title":"Multicult. Educ."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"6083","DOI":"10.1109\/JSTARS.2021.3087555","article-title":"Multi-Scale Ship Detection From SAR and Optical Imagery Via A More Accurate YOLOv3","volume":"14","author":"Hong","year":"2021","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_19","unstructured":"Bochkovskiy, A., Wang, C.Y., and Liao, H.Y.M. (2020). YOLOv4 Optimal Speed and Accuracy of Object Detection. In Proceedings of the Computer Vision and Pattern Recognition. arXiv."},{"key":"ref_20","unstructured":"PyTorch Hub (2021, November 26). Ultralytics\/yolov5: v4.0\u2014nn.SiLU() Activations, Weights & Biases Logging, PyTorch Hub Integration. Available online: https:\/\/codechina.csdn.net\/ForrestGump92\/yolov5\/-\/tree\/silu."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Han, G., Zhang, X., and Li, C. (2017, January 17\u201320). Single shot object detection with top-down refinement. Proceedings of the 2017 IEEE International Conference on Image Processing (ICIP), Beijing, China.","DOI":"10.1109\/ICIP.2017.8296905"},{"key":"ref_22","first-page":"85","article-title":"CNN based detectors on planetary environments: A performance evaluation","volume":"14","author":"Rubio","year":"2020","journal-title":"Front. Neurorobot."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Duan, K., Bai, S., Xie, L., Qi, H., Huang, Q., and Tian, Q. (2019, January 27\u201328). CenterNet: Keypoint Triplets for Object Detection. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00667"},{"key":"ref_24","unstructured":"(2021, August 12). Canny Edge Detector Algorithm. Available online: https:\/\/towardsdatascience.com\/canny-edge-detectionstep-by-step-in-python-computer-vision-b49c3a2d8123."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/13\/12\/307\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:37:33Z","timestamp":1760168253000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/13\/12\/307"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,11,29]]},"references-count":24,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2021,12]]}},"alternative-id":["fi13120307"],"URL":"https:\/\/doi.org\/10.3390\/fi13120307","relation":{},"ISSN":["1999-5903"],"issn-type":[{"value":"1999-5903","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,11,29]]}}}