{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T15:00:00Z","timestamp":1782313200255,"version":"3.54.5"},"reference-count":44,"publisher":"MDPI AG","issue":"14","license":[{"start":{"date-parts":[[2022,7,15]],"date-time":"2022-07-15T00:00:00Z","timestamp":1657843200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Iqra University, Pakistan","award":["015MEO-227"],"award-info":[{"award-number":["015MEO-227"]}]},{"name":"Universiti Teknologi PETRONAS (UTP), Malaysia","award":["015MEO-227"],"award-info":[{"award-number":["015MEO-227"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Semantic segmentation for accurate visual perception is a critical task in computer vision. In principle, the automatic classification of dynamic visual scenes using predefined object classes remains unresolved. The challenging problems of learning deep convolution neural networks, specifically ResNet-based DeepLabV3+ (the most recent version), are threefold. The problems arise due to (1) biased centric exploitations of filter masks, (2) lower representational power of residual networks due to identity shortcuts, and (3) a loss of spatial relationship by using per-pixel primitives. To solve these problems, we present a proficient approach based on DeepLabV3+, along with an added evaluation metric, namely, Unified DeepLabV3+ and S3core, respectively. The presented unified version reduced the effect of biased exploitations via additional dilated convolution layers with customized dilation rates. We further tackled the problem of representational power by introducing non-linear group normalization shortcuts to solve the focused problem of semi-dark images. Meanwhile, to keep track of the spatial relationships in terms of the global and local contexts, geometrically bunched pixel cues were used. We accumulated all the proposed variants of DeepLabV3+ to propose Unified DeepLabV3+ for accurate visual decisions. Finally, the proposed S3core evaluation metric was based on the weighted combination of three different accuracy measures, i.e., the pixel accuracy, IoU (intersection over union), and Mean BFScore, as robust identification criteria. Extensive experimental analysis performed over a CamVid dataset confirmed the applicability of the proposed solution for autonomous vehicles and robotics for outdoor settings. The experimental analysis showed that the proposed Unified DeepLabV3+ outperformed DeepLabV3+ by a margin of 3% in terms of the class-wise pixel accuracy, along with a higher S3core, depicting the effectiveness of the proposed approach.<\/jats:p>","DOI":"10.3390\/s22145312","type":"journal-article","created":{"date-parts":[[2022,7,18]],"date-time":"2022-07-18T01:53:22Z","timestamp":1658109202000},"page":"5312","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["Unified DeepLabV3+ for Semi-Dark Image Semantic Segmentation"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3921-0104","authenticated-orcid":false,"given":"Mehak Maqbool","family":"Memon","sequence":"first","affiliation":[{"name":"High Performance Cloud Computing Center (HPC3), Department of Computer and Information Sciences, Universiti Teknologi PETRONAS, Seri Iskandar 32610, Malaysia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6617-8149","authenticated-orcid":false,"given":"Manzoor Ahmed","family":"Hashmani","sequence":"additional","affiliation":[{"name":"High Performance Cloud Computing Center (HPC3), Department of Computer and Information Sciences, Universiti Teknologi PETRONAS, Seri Iskandar 32610, Malaysia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9815-2704","authenticated-orcid":false,"given":"Aisha Zahid","family":"Junejo","sequence":"additional","affiliation":[{"name":"High Performance Cloud Computing Center (HPC3), Department of Computer and Information Sciences, Universiti Teknologi PETRONAS, Seri Iskandar 32610, Malaysia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Syed Sajjad","family":"Rizvi","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Shaheed Zulfiqar Ali Bhutto Institute of Science and Technology, Karachi 75600, Pakistan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kamran","family":"Raza","sequence":"additional","affiliation":[{"name":"Faculty of Engineering Science and Technology, Iqra University, Karachi 75600, Pakistan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,7,15]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Memon, M.M., Hashmani, M.A., Junejo, A.Z., Rizvi, S.S., and Arain, A. (2021). A Novel Luminance-Based Algorithm for Classification of Semi-Dark Images. Appl. Sci., 11.","DOI":"10.3390\/app11188694"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Chen, C., Chen, Q., Xu, J., and Koltun, V. (2018, January 18\u201323). Learning to see in the dark. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00347"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Ouyang, S., and Li, Y. (2021). Combining deep semantic segmentation network and graph convolutional neural network for semantic segmentation of remote sensing imagery. Remote Sens., 13.","DOI":"10.3390\/rs13010119"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Yu, J., Zeng, P., Yu, Y., Yu, H., Huang, L., and Zhou, D. (2022). A Combined Convolutional Neural Network for Urban Land-Use Classification with GIS Data. Remote Sens., 14.","DOI":"10.3390\/rs14051128"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Senthilnathan, R. (2022). Deep Learning in Vision-Based Automated Inspection: Current State and Future Prospects. Machine Learning in Industry, Springer.","DOI":"10.1007\/978-3-030-75847-9_8"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Chen, L.-C., Yang, Y., Wang, J., Xu, W., and Yuille, A.L. (2016, January 27\u201330). Attention to scale: Scale-aware semantic image segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.396"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1016\/j.neucom.2019.02.003","article-title":"Survey on semantic segmentation using deep learning techniques","volume":"338","author":"Lateef","year":"2019","journal-title":"Neurocomputing"},{"key":"ref_9","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_13","unstructured":"Chen, L.-C., Papandreou, G., Schroff, F., and Adam, H. (2017). Rethinking atrous convolution for semantic image segmentation. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. (2018, January 8\u201314). Encoder-decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"ref_15","unstructured":"Zhang, C., Rameau, F., Lee, S., Kim, J., Benz, P., Argaw, D.M., Bazin, J.-C., and Kweon, I.S. (2019, January 9\u201312). Revisiting residual networks with nonlinear shortcuts. Proceedings of the BMVC, Cardiff, UK."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"McAllister, R., Gal, Y., Kendall, A., Van Der Wilk, M., Shah, A., Cipolla, R., and Weller, A. (2017, January 19\u201325). Concrete Problems for Autonomous Vehicle Safety: Advantages of Bayesian Deep Learning. Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence AI and Autonomy Track, Melbourne, Australia.","DOI":"10.24963\/ijcai.2017\/661"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1792","DOI":"10.1109\/LRA.2019.2896518","article-title":"Normalization in training U-Net for 2-D biomedical semantic segmentation","volume":"4","author":"Zhou","year":"2019","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zhao, W., Fu, Y., Wei, X., and Wang, H. (2018). An improved image semantic segmentation method based on superpixels and conditional random fields. Appl. Sci., 8.","DOI":"10.3390\/app8050837"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zheng, S., Jayasumana, S., Romera-Paredes, B., Vineet, V., Su, Z., Du, D., Huang, C., and Torr, P.H. (2015, January 7\u201313). Conditional random fields as recurrent neural networks. Proceedings of the IEEE International Conference on Computer Vision, Washington, DC, USA.","DOI":"10.1109\/ICCV.2015.179"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Plath, N., Toussaint, M., and Nakajima, S. (2009, January 14\u201318). Multi-class image segmentation using conditional random fields and global classification. Proceedings of the 26th Annual International Conference on Machine Learning, Montreal, QC, Canada.","DOI":"10.1145\/1553374.1553479"},{"key":"ref_21","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012). Imagenet Classification with Deep Convolutional Neural Networks, Association for Computing Machinery."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"98","DOI":"10.1007\/s11263-014-0733-5","article-title":"The pascal visual object classes challenge: A retrospective","volume":"111","author":"Everingham","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. (2016, January 27\u201330). The cityscapes dataset for semantic urban scene understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.350"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"88","DOI":"10.1016\/j.patrec.2008.04.005","article-title":"Semantic object classes in video: A high-definition ground truth database","volume":"30","author":"Brostow","year":"2009","journal-title":"Pattern Recognit. Lett."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Song, S., Lichtenberg, S.P., and Xiao, J. (2015, January 7\u201312). Sun rgb-d: A rgb-d scene understanding benchmark suite. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298655"},{"key":"ref_27","unstructured":"Cogswell, M., Lin, X., Purushwalkam, S., and Batra, D. (2014). Combining the best of graphical models and convnets for semantic segmentation. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1915","DOI":"10.1109\/TPAMI.2012.231","article-title":"Learning hierarchical features for scene labeling","volume":"35","author":"Farabet","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Liu, C., Yuen, J., Torralba, A., Sivic, J., and Freeman, W.T. (2008, January 12\u201318). Sift flow: Dense correspondence across different scenes. Proceedings of the European Conference on Computer Vision, Marseille, France.","DOI":"10.1007\/978-3-540-88690-7_3"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Tighe, J., and Lazebnik, S. (2010, January 5\u201311). Superparsing: Scalable nonparametric image parsing with superpixels. Proceedings of the European Conference on Computer Vision, Heraklion, Greece.","DOI":"10.1007\/978-3-642-15555-0_26"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Gould, S., Fulton, R., and Koller, D. (October, January 27). Decomposing a scene into geometric and semantically consistent regions. Proceedings of the 2009 IEEE 12th International Conference on Computer Vision, Kyoto, Japan.","DOI":"10.1109\/ICCV.2009.5459211"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Papandreou, G., Chen, L.-C., Murphy, K.P., and Yuille, A.L. (2015, January 7\u201313). Weakly-and semi-supervised learning of a deep convolutional network for semantic image segmentation. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.203"},{"key":"ref_33","unstructured":"Saito, S., Kerola, T., and Tsutsui, S. (2022, May 29). Superpixel Clustering with Deep Features for Unsupervised Road Segmentation. Available online: https:\/\/www.arxiv-vanity.com\/papers\/1711.05998\/."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"He, Y., Chiu, W.-C., Keuper, M., and Fritz, M. (2017, January 21\u201326). Std2p: Rgbd semantic segmentation using spatio-temporal data-driven pooling. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.757"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"196","DOI":"10.1016\/j.neucom.2019.01.016","article-title":"Superpixel based continuous conditional random field neural network for semantic segmentation","volume":"340","author":"Zhou","year":"2019","journal-title":"Neurocomputing"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Kae, A., Sohn, K., Lee, H., and Learned-Miller, E. (2013, January 23\u201328). Augmenting CRFs with Boltzmann machine shape priors for image labeling. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.263"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Smith, B.M., Zhang, L., Brandt, J., Lin, Z., and Yang, J. (2013, January 23\u201328). Exemplar-based face parsing. Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.447"},{"key":"ref_38","unstructured":"Fisher Yu, V.K. (2015). Multi-Scale Context Aggregation by Dilated Convolutions. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Wu, Y., and He, K. (2018, January 8\u201314). Group normalization. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01261-8_1"},{"key":"ref_40","unstructured":"Dumoulin, V., and Visin, F. (2016). A guide to convolution arithmetic for deep learning. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Brostow, G.J., Shotton, J., Fauqueur, J., and Cipolla, R. (2008, January 12\u201318). Segmentation and recognition using structure from motion point clouds. Proceedings of the European Conference on Computer Vision, Marseille, France.","DOI":"10.1007\/978-3-540-88682-2_5"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Csurka, G., Larlus, D., Perronnin, F., and Meylan, F.J.I.P. (2013). What is a good evaluation measure for semantic segmentation?. Proceedings of the British Machine Vision Conference, BMVA Press.","DOI":"10.5244\/C.27.32"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Fernandez-Moral, E., Martins, R., Wolf, D., and Rives, P. (2018, January 26\u201330). A new metric for evaluating semantic segmentation: Leveraging global and contour accuracy. Proceedings of the 2018 IEEE Intelligent Vehicles Symposium (iv), Suzhou, China.","DOI":"10.1109\/IVS.2018.8500497"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Saito, M., and Matsumoto, M. (2008). SIMD-oriented fast Mersenne Twister: A 128-bit pseudorandom number generator. Monte Carlo and Quasi-Monte Carlo Methods 2006, Springer.","DOI":"10.1007\/978-3-540-74496-2_36"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/14\/5312\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:51:39Z","timestamp":1760140299000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/14\/5312"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,15]]},"references-count":44,"journal-issue":{"issue":"14","published-online":{"date-parts":[[2022,7]]}},"alternative-id":["s22145312"],"URL":"https:\/\/doi.org\/10.3390\/s22145312","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,7,15]]}}}