{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T05:40:32Z","timestamp":1774935632734,"version":"3.50.1"},"reference-count":40,"publisher":"MDPI AG","issue":"18","license":[{"start":{"date-parts":[[2020,9,7]],"date-time":"2020-09-07T00:00:00Z","timestamp":1599436800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61762025"],"award-info":[{"award-number":["61762025"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100017691","name":"Guangxi Key Research and Development Program","doi-asserted-by":"publisher","award":["AB17195053, AB18126053, AB18126063, AD18281002"],"award-info":[{"award-number":["AB17195053, AB18126053, AB18126063, AD18281002"]}],"id":[{"id":"10.13039\/501100017691","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Natural Science Foundation of Guangxi of China","award":["2017GXNSFAA198226, 2019GXNSFDA185007, 2019GXNSFDA185006"],"award-info":[{"award-number":["2017GXNSFAA198226, 2019GXNSFDA185007, 2019GXNSFDA185006"]}]},{"name":"Guilin Science and Technology Development Program","award":["20180107-4"],"award-info":[{"award-number":["20180107-4"]}]},{"name":"Innovation Project of GUET Graduate Education","award":["2019YCXS051, 2020YCXS052"],"award-info":[{"award-number":["2019YCXS051, 2020YCXS052"]}]},{"name":"Guangxi Colleges and Universities Key Laboratory of Intelligent Processing of Computer Image and Graphics","award":["GIIP201603"],"award-info":[{"award-number":["GIIP201603"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In recent years, increasing image data comes from various sensors, and object detection plays a vital role in image understanding. For object detection in complex scenes, more detailed information in the image should be obtained to improve the accuracy of detection task. In this paper, we propose an object detection algorithm by jointing semantic segmentation (SSOD) for images. First, we construct a feature extraction network that integrates the hourglass structure network with the attention mechanism layer to extract and fuse multi-scale features to generate high-level features with rich semantic information. Second, the semantic segmentation task is used as an auxiliary task to allow the algorithm to perform multi-task learning. Finally, multi-scale features are used to predict the location and category of the object. The experimental results show that our algorithm substantially enhances object detection performance and consistently outperforms other three comparison algorithms, and the detection speed can reach real-time, which can be used for real-time detection.<\/jats:p>","DOI":"10.3390\/s20185080","type":"journal-article","created":{"date-parts":[[2020,9,7]],"date-time":"2020-09-07T09:18:16Z","timestamp":1599470296000},"page":"5080","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":27,"title":["Convolutional Neural Networks-Based Object Detection Algorithm by Jointing Semantic Segmentation for Images"],"prefix":"10.3390","volume":"20","author":[{"given":"Baohua","family":"Qiang","sequence":"first","affiliation":[{"name":"Guangxi Colleges and Universities Key Laboratory of Intelligent Processing of Computer Image and Graphics, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3083-014X","authenticated-orcid":false,"given":"Ruidong","family":"Chen","sequence":"additional","affiliation":[{"name":"Guangxi Colleges and Universities Key Laboratory of Intelligent Processing of Computer Image and Graphics, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1874-3641","authenticated-orcid":false,"given":"Mingliang","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Computer Science, Chongqing University, 174 Shazheng Street, Shapingba District, Chongqing 400044, China"},{"name":"State Key Laboratory of Internet of Things for Smart City, Faculty of Science and Technology, University of Macau, Macau, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5667-2039","authenticated-orcid":false,"given":"Yuanchao","family":"Pang","sequence":"additional","affiliation":[{"name":"Guangxi Colleges and Universities Key Laboratory of Intelligent Processing of Computer Image and Graphics, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9132-1077","authenticated-orcid":false,"given":"Yijie","family":"Zhai","sequence":"additional","affiliation":[{"name":"Guangxi Colleges and Universities Key Laboratory of Intelligent Processing of Computer Image and Graphics, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5397-3576","authenticated-orcid":false,"given":"Minghao","family":"Yang","sequence":"additional","affiliation":[{"name":"Guangxi Colleges and Universities Key Laboratory of Intelligent Processing of Computer Image and Graphics, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,9,7]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Cvar, N., Trilar, J., Kos, A., Volk, M., and Stojmenova Duh, E. (2020). The Use of IoT Technology in Smart Cities and Smart Villages: Similarities, Differences, and Future Prospects. Sensors, 20.","DOI":"10.3390\/s20143897"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","article-title":"Object Detection with Discriminatively Trained Part-Based Models","volume":"32","author":"Felzenszwalb","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Wu, X., Duan, J., Zhong, M., Li, P., and Wang, J. (2020). VNF Chain Placement for Large Scale IoT of Intelligent Transportation. Sensors, 20.","DOI":"10.3390\/s20143819"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Qiang, B., Zhang, S., Zhan, Y., Xie, W., and Zhao, T. (2019). Improved Convolutional Pose Machines for Human Pose Estimation Using Image Sensor Data. Sensors, 19.","DOI":"10.3390\/s19030718"},{"key":"ref_5","first-page":"434","article-title":"An Evaluation of Pixel-Based Methods for the Detection of Floating Objects on the Sea Surface","volume":"33","author":"Borghgraef","year":"2010","journal-title":"EURASIP J. Adv. Signal Process."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"543","DOI":"10.1016\/j.jvcir.2011.03.009","article-title":"Robust real-time ship detection and tracking for visual surveillance of cage aquaculture","volume":"22","author":"Hu","year":"2011","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Shi, G., Suo, J., Liu, C., Wan, K., and Lv, X. (2017, January 3\u20135). Moving target detection algorithm in image sequences based on edge detection and frame difference. Proceedings of the 2017 IEEE 3rd Information Technology and Mechatronics Engineering Conference, Chongqing, China.","DOI":"10.1109\/ITOEC.2017.8122449"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Kang, Y., Huang, W., and Zheng, S. (2017, January 20\u201322). An improved frame difference method for moving target detection. Proceedings of the 2017 Chinese Automation Congress, Jinan, China.","DOI":"10.1109\/CAC.2017.8243011"},{"key":"ref_9","unstructured":"Viola, P., and Jones, M. (2001, January 8\u201314). Rapid object detection using a boosted cascade of simple features. Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Kauai, HI, USA."},{"key":"ref_10","unstructured":"Dalal, N., and Triggs, B. (2005, January 20\u201325). Histograms of oriented gradients for human detection. Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Diego, CA, USA."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Zeiler, M.D., Taylor, G.W., and Fergus, R. (2011, January 6\u201313). Adaptive deconvolutional networks for mid and high level feature learning. In Proceeding of the 2011 IEEE International Conference on Computer Vision (ICCV), Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126474"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 11\u201318). Fast R-CNN. Proceedings of the 2015 Ieee International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"386","DOI":"10.1109\/TPAMI.2018.2844175","article-title":"Mask R-CNN","volume":"42","author":"He","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2014, January 6\u201312). Spatial pyramid pooling in deep convolutional networks for visual recognition. In Proceeding of the 13th European Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10578-9_23"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Kong, T., Yao, A., Chen, Y., and Sun, F. (July, January 26). HyperNet: Towards accurate region proposal generation and joint object detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.98"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Cai, Z., and Vasconcelos, N. (2018, January 18\u201323). Cascade R-CNN: Delving into high quality object detection. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00644"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., and Berg, A.C. (2016, January 8\u201316). SSD: Single shot multibox detector. Proceedings of the 2016 European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"318","DOI":"10.1109\/TPAMI.2018.2858826","article-title":"Focal Loss for Dense Object Detection","volume":"42","author":"Lin","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhang, S., Wen, L., Bian, X., Lei, Z., and Li, S.Z. (2018, January 18\u201323). Single-shot refinement neural network for object detection. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00442"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Liu, S., and Huang, D. (2018, January 8\u201314). Receptive field block net for accurate and fast object detection. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01252-6_24"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Shen, Z., Liu, Z., Li, J., Jiang, Y.-G., Chen, Y., and Xue, X. (2017, January 22\u201329). DSOD: Learning deeply supervised object detectors from scratch. Proceedings of the 2017 IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.212"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., van der Maaten, L., and Weinberger, K.Q. (2017, January 21\u201326). Densely connected convolutional networks. Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"642","DOI":"10.1007\/s11263-019-01204-1","article-title":"CornerNet: Detecting Objects as Paired Keypoints","volume":"128","author":"Law","year":"2020","journal-title":"Int. J. Comput. Vis."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Duan, K., Bai, S., Xie, L., Qi, H., Huang, Q., and Tian, Q. (November, January 27). Centernet: Keypoint triplets for object detection. Proceedings of the 2019 IEEE\/Cvf International Conference on Computer Vision, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00667"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Tian, Z., Shen, C., Chen, H., and He, T. (November, January 27). FCOS: Fully convolutional one-stage object detection. Proceedings of the 2019 IEEE\/Cvf International Conference on Computer Vision, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00972"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Shinohara, T., Xiu, H., and Matsuoka, M. (2020). FWNet: Semantic Segmentation for Full-Waveform LiDAR Data Using Deep Learning. Sensors, 20.","DOI":"10.3390\/s20123568"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Saez, A., Bergasa, L.M., Lopez-Guillen, E., Romera, E., Tradacete, M., Gomez-Huelamo, C., and del Egido, J. (2019). Real-Time Semantic Segmentation for Fisheye Urban Driving Images Based on ERFNet. Sensors, 19.","DOI":"10.3390\/s19030503"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Eom, H., Lee, D., Han, S., Hariyani, Y.S., Lim, Y., Sohn, I., Park, K., and Park, C. (2020). End-To-End Deep Learning Architecture for Continuous Blood Pressure Estimation Using Attention Mechanism. Sensors, 20.","DOI":"10.3390\/s20082338"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Meng, F., Wang, X., Wang, D., Shao, F., and Fu, L. (2020). Spatial-Semantic and Temporal Attention Mechanism-Based Online Multi-Object Tracking. Sensors, 20.","DOI":"10.3390\/s20061653"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Jiang, B., Luo, R., Mao, J., Xiao, T., and Jiang, Y. (2018, January 8\u201314). Acquisition of localization confidence for accurate object detection. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_48"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"926","DOI":"10.1109\/LSP.2018.2822810","article-title":"Additive Margin Softmax for Face Verification","volume":"25","author":"Wang","year":"2018","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_35","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A Method for Stochastic Optimization. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","article-title":"The Pascal Visual Object Classes (VOC) Challenge","volume":"88","author":"Everingham","year":"2010","journal-title":"Int. J. Comput. Vis."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"640","DOI":"10.1109\/TPAMI.2016.2572683","article-title":"Fully Convolutional Networks for Semantic Segmentation","volume":"39","author":"Shelhamer","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Ponn, T., Kr\u00f6ger, T., and Diermeyer, F. (2020). Identification and Explanation of Challenging Conditions for Camera-Based Object Detection of Automated Vehicles. Sensors, 20.","DOI":"10.3390\/s20133699"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Dollar, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_40","unstructured":"Wang, R.J., Li, X., and Ling, C.X. (2018, January 2\u20138). Pelee: A real-time object detection system on mobile devices. Proceedings of the 32nd Conference on Neural Information Processing Systems (NIPS), Montreal, QC, Canada."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/18\/5080\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:07:38Z","timestamp":1760177258000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/18\/5080"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,7]]},"references-count":40,"journal-issue":{"issue":"18","published-online":{"date-parts":[[2020,9]]}},"alternative-id":["s20185080"],"URL":"https:\/\/doi.org\/10.3390\/s20185080","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,9,7]]}}}