{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T15:38:37Z","timestamp":1780501117046,"version":"3.54.1"},"reference-count":92,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2024,4,30]],"date-time":"2024-04-30T00:00:00Z","timestamp":1714435200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Image understanding plays a pivotal role in various computer vision tasks, such as extraction of essential features from images, object detection, and segmentation. At a higher level of granularity, both semantic and instance segmentation are necessary for fully grasping a scene. In recent times, the concept of panoptic segmentation has emerged as a field of study that unifies semantic and instance segmentation. This article sheds light on the pivotal role of panoptic segmentation as a visualization tool for understanding scene components, including object detection, categorization, and precise localization of scene elements. Advancements in achieving panoptic segmentation and suggested improvements to the predicted outputs through a top-down approach are discussed. Furthermore, datasets relevant to both scene recognition and panoptic segmentation are explored to facilitate a comparative analysis. Finally, the article outlines certain promising directions in image recognition and analysis by underlining the ongoing evolution in image understanding methodologies.<\/jats:p>","DOI":"10.3390\/a17050189","type":"journal-article","created":{"date-parts":[[2024,4,30]],"date-time":"2024-04-30T04:01:52Z","timestamp":1714449712000},"page":"189","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Insights into Image Understanding: Segmentation Methods for Object Recognition and Scene Classification"],"prefix":"10.3390","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1107-5879","authenticated-orcid":false,"given":"Sarfaraz Ahmed","family":"Mohammed","sequence":"first","affiliation":[{"name":"Department of Computer Science, College of Engineering and Applied Science, University of Cincinnati, Cincinnati, OH 45221-0030, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anca L.","family":"Ralescu","sequence":"additional","affiliation":[{"name":"Department of Computer Science, College of Engineering and Applied Science, University of Cincinnati, Cincinnati, OH 45221-0030, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,4,30]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.inffus.2020.12.009","article-title":"A selection framework of sensor combination feature subset for human motion phase segmentation","volume":"70","author":"Wang","year":"2021","journal-title":"Inf. Fusion"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.inffus.2016.10.003","article-title":"MRI segmentation fusion for brain tumor detection","volume":"36","author":"Cabria","year":"2017","journal-title":"Inf. Fusion"},{"key":"ref_3","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015). Medical Image Computing and Computer-Assisted Intervention\u2014MICCAI 2015, Springer. Lecture Notes in Computer Science."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"122017","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"SegNet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","unstructured":"Yao, J., Fidler, S., and Urtasun, R. (2012, January 16\u201321). Describing the scene as a whole: Joint object detection, scene classification and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Xiong, Y., Liao, R., Zhao, H., Hu, R., Bai, M., Yumer, E., and Urtasun, R. (2019, January 16\u201317). Upsnet: A unified panoptic segmentation network. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00902"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Liu, S., Jia, J., Fidler, S., and Urtasun, R. (2017, January 22\u201329). SGN: Sequential grouping networks for instance segmentation. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.378"},{"key":"ref_9","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C.L. (2014). Computer Vision e ECCV 2014, Springer. Lecture Notes in Computer Science."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Kirillov, A., He, K., Girshick, R., Rother, C., and Doll\u00e1r, P. (2019, January 16\u201317). Panoptic segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00963"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"107205","DOI":"10.1016\/j.patcog.2020.107205","article-title":"Scene recognition: A comprehensive survey","volume":"102","author":"Xie","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"131","DOI":"10.1007\/s11263-015-0807-z","article-title":"Guest editorial: Scene understanding","volume":"112","author":"Hoiem","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1023\/A:1011139631724","article-title":"Modeling the shape of the scene: A holistic representation of the spatial envelope","volume":"42","author":"Oliva","year":"2001","journal-title":"Int. J. Comput. Vis."},{"key":"ref_14","unstructured":"Zhou, B., Lapedriza, A., Xiao, J., Torralba, A., and Oliva, A. (2014, January 8\u201313). Learning deep features for scene recognition using places database. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (July, January 26). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2016, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_18","unstructured":"Schwing, A.G., and Urtasun, R. (2015). Fully connected deep structured networks. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Li, Y., Qi, H., Dai, J., Ji, X., and Wei, Y. (2016). Fully convolutional instance-aware semantic segmentation. arXiv.","DOI":"10.1109\/CVPR.2017.472"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Neuhold, G., Ollmann, T., Bulo, S.R., and Kontschieder, P. (2017, January 22\u201329). The Mapillary Vistas Dataset for Semantic Understanding of Street Scenes. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.534"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"Imagenet large scale visual recognition challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_23","unstructured":"Abu-El-Haija, S., Kothari, N., Lee, J., Natsev, P., Toderici, G., Varadarajan, B., and Vijayanarasimhan, S. (2016). Youtube-8m: A large-scale video classification benchmark. arXiv."},{"key":"ref_24","unstructured":"Brostow, G.J., Shotton, J., Fauqueur, J., and Cipolla, R. (2008). Computer Vision\u2013ECCV 2008: 10th European Conference on Computer Vision, Marseille, France, 12\u201318 October 2008, Proceedings, Part I 10, Springer."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The KITTI dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_26","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. (July, January 26). The Cityscapes dataset for semantic urban scene understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Leibe, B., Cornelis, N., Cornelis, K., and Van Gool, L. (2007, January 18\u201323). Dynamic 3D scene analysis from a moving vehicle. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Minneapolis, MN, USA.","DOI":"10.1109\/CVPR.2007.383146"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Tighe, J., and Lazebnik, S. (2013, January 23\u201328). Finding things: Image parsing with regions and per-exemplar detectors. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.386"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Tighe, J., Niethammer, M., and Lazebnik, S. (2014, January 23\u201328). Scene parsing with object instances and occlusion ordering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.479"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1370","DOI":"10.1109\/TPAMI.2013.193","article-title":"Relating things and stuff via object property interactions","volume":"36","author":"Sun","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_31","unstructured":"Wu, Y. (2024, March 25). Available online: https:\/\/github.com\/facebookresearch\/detectron2."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"530","DOI":"10.1109\/TPAMI.2004.1273918","article-title":"Learning to detect natural image boundaries using local brightness, color, and texture cues","volume":"26","author":"Martin","year":"2004","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_33","unstructured":"Van Rijsbergen, C. (1979). Information Retrieval, Butterworths."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Kirillov, A., Girshick, R., He, K., and Doll\u00e1r, P. (2019, January 16\u201317). Panoptic feature pyramid networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00656"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Li, Y., Chen, X., Zhu, Z., Xie, L., Huang, G., Du, D., and Wang, X. (2019, January 16\u201317). Attention-guided unified network for panoptic segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00719"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Porzi, L., Bulo, S.R., Colovic, A., and Kontschieder, P. (2019, January 16\u201317). Seamless scene segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00847"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Mohan, R., and Valada, A. (2021). EfficientPS: Efficient Panoptic Segmentation. arXiv.","DOI":"10.1109\/CVPR52688.2022.02035"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask R-CNN. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"103283","DOI":"10.1016\/j.dsp.2021.103283","article-title":"A survey on deep learning-based panoptic segmentation","volume":"120","author":"Li","year":"2022","journal-title":"Digit. Signal Process."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"de Geus, D., Meletis, P., and Dubbelman, G. (2019). Panoptic Segmentation with a Joint Semantic and Instance Segmentation Network. arXiv.","DOI":"10.1109\/LRA.2020.2969919"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Liu, H., Peng, C., Yu, C., Wang, J., Liu, X., Yu, G., and Jiang, W. (2019, January 16\u201317). An End-To-End Network for Panoptic Segmentation. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00633"},{"key":"ref_42","unstructured":"(2024, March 25). On-Device Panoptic Segmentation for Camera Using Transformers. October 2021. Available online: https:\/\/machinelearning.apple.com\/research\/panoptic-segmentation."},{"key":"ref_43","unstructured":"(2024, March 25). Fast Class-Agnostic Salient Object Segmentation. June 2023. Available online: https:\/\/machinelearning.apple.com\/research\/salient-object-segmentation."},{"key":"ref_44","unstructured":"Prakhar Bansal (2024, March 25). Panoptic Segmentation Explained. Available online: https:\/\/medium.com\/@prakhar.bansal\/panoptic-segmentation-explained-5fa7313591a3."},{"key":"ref_45","unstructured":"(2024, March 25). Using Panoptic Segmentation to Train Autonomous Vehicles. Mindy News Blog. December 2021. Available online: https:\/\/mindy-support.com\/news-post\/using-panoptic-segmentation-to-train-autonomous-vehicles\/."},{"key":"ref_46","unstructured":"(2024, March 25). COCO 2020 Panoptic Segmentation. Available online: https:\/\/cocodataset.org\/#panoptic-2020."},{"key":"ref_47","unstructured":"(2024, March 25). Cityscapes Dataset. Available online: https:\/\/www.cityscapes-dataset.com\/dataset-overview\/."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Yu, F., Chen, H., Wang, X., Xian, W., Chen, Y., Liu, F., Madhavan, V., and Darrell, T. (2020, January 14\u201319). BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA. Available online: https:\/\/arxiv.org\/abs\/1805.04687.pdf.","DOI":"10.1109\/CVPR42600.2020.00271"},{"key":"ref_49","unstructured":"(2024, March 25). Mapillary Vistas Dataset. Available online: https:\/\/www.mapillary.com\/dataset\/vistas."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"474","DOI":"10.1016\/j.patcog.2017.09.025","article-title":"Scene recognition with objectness","volume":"74","author":"Cheng","year":"2018","journal-title":"Pattern Recognit."},{"key":"ref_51","unstructured":"(2024, March 25). Semantic KITTI Dataset. Available online: http:\/\/www.semantic-kitti.org\/dataset.html."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Dai, A., Chang, A.X., Savva, M., Halber, M., Funkhouser, T., and Nie\u00dfner, M. (2017, January 21\u201326). ScanNet: Richly annotated 3D Reconstructions of Indoor Scenes. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.261"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O. (2020, January 13\u201319). nuScenes: A multimodal dataset for autonomous driving. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2020, Seattle, WA, USA. Available online: https:\/\/arXiv:1903.11027.pdf.","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Arnab, A., and Torr, P.H. (2017, January 21\u201326). Pixelwise instance segmentation with a dynamically instantiated network. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.100"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Bai, M., and Urtasun, R. (2017, January 21\u201326). Deep watershed transform for instance segmentation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.305"},{"key":"ref_56","unstructured":"Chang, C.-Y., Chang, S.-E., Hsiao, P.-Y., and Fu, L.-C. (December, January 30). Epsnet: Efficient panoptic segmentation network with cross-layer attention fusion. Proceedings of the Asian Conference on Computer Vision, Asian Conference on Computer Vision, Kyoto, Japan."},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"823","DOI":"10.1109\/TIP.2013.2295756","article-title":"mCENTRIST: A multi-channel feature generation mechanism for scene categorization","volume":"23","author":"Xiao","year":"2014","journal-title":"IEEE Trans. Image Process."},{"key":"ref_58","doi-asserted-by":"crossref","first-page":"373","DOI":"10.1016\/j.patcog.2011.06.012","article-title":"Building global image features for scene recognition","volume":"45","author":"Meng","year":"2012","journal-title":"Pattern Recognit."},{"key":"ref_59","doi-asserted-by":"crossref","first-page":"1489","DOI":"10.1109\/TPAMI.2010.224","article-title":"CENTRIST: A Visual Descriptor for Scene Categorization","volume":"33","author":"Wu","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_60","unstructured":"Lazebnik, S., Schmid, C., and Ponce, J. (2006, January 17\u201322). Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, New York, NY, USA."},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"3372","DOI":"10.1109\/TIP.2016.2567076","article-title":"A discriminative representation of convolutional features for indoor scene recognition","volume":"25","author":"Khan","year":"2016","journal-title":"IEEE Trans. Image Process."},{"key":"ref_62","unstructured":"Wu, J., and Rehg, J.M. (2009, January 20\u201325). Beyond the Euclidean distance: Creating effective visual codebooks using the histogram intersection kernel. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA."},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Gong, Y., Wang, L., and Guo, R. (2014, January 6\u201312). Multi-scale order less pooling of deep convolutional activation features. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10584-0_26"},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"Gao, S., Tsang, I.W.H., Chia, L.T., and Zhao, P. (2010, January 13\u201318). Local features are not lonely\u2014Laplacian sparse coding for image classification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539943"},{"key":"ref_65","doi-asserted-by":"crossref","first-page":"1182","DOI":"10.1109\/TMM.2019.2942478","article-title":"Hierarchical coding of convolutional features for scene recognition","volume":"22","author":"Xie","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"Zhang, C., Liu, J., Tian, Q., Xu, C., Lu, H., and Ma, S. (2011, January 20\u201325). Image classification by non-negative sparse coding, low-rank and sparse decomposition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995484"},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Jiang, Y., Yuan, J., and Yu, G. (2012, January 7\u201313). Randomized spatial partition for scene recognition. Proceedings of the European Conference on Computer Vision, Florence, Italy.","DOI":"10.1007\/978-3-642-33709-3_52"},{"key":"ref_68","doi-asserted-by":"crossref","first-page":"4829","DOI":"10.1109\/TIP.2016.2599292","article-title":"A spatial layout and scale invariant feature representation for indoor scene classification","volume":"25","author":"Hayat","year":"2016","journal-title":"IEEE Trans. Image Process."},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Yang, M., Li, B., Fan, H., and Jiang, Y. (2015, January 27\u201330). Randomized spatial pooling in deep convolutional networks for scene recognition. Proceedings of the IEEE Conference on Image Processing, Quebec City, QC, Canada.","DOI":"10.1109\/ICIP.2015.7350829"},{"key":"ref_70","unstructured":"Li, L.J., Su, H., Fei-Fei, L., and Xing, E. (2010, January 6\u201311). Object bank: A high-level image representation for scene classification and semantic feature sparsification. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"Zuo, Z., Wang, G., Shuai, B., Zhao, L., Yang, Q., and Jiang, X. (2014, January 6\u201312). Learning discriminative and shareable features for scene classification. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10590-1_36"},{"key":"ref_72","doi-asserted-by":"crossref","first-page":"45230","DOI":"10.1109\/ACCESS.2019.2908448","article-title":"Scene Categorization Model using Deep Visually Sensitive features","volume":"7","author":"Shi","year":"2019","journal-title":"IEEE Access"},{"key":"ref_73","doi-asserted-by":"crossref","unstructured":"Lin, D., Lu, C., Liao, R., and Jia, J. (2014, January 23\u201328). Learning important spatial pooling regions for scene regions for scene classification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.476"},{"key":"ref_74","doi-asserted-by":"crossref","unstructured":"Wu, R., Wang, B., Wang, W., and Yu, Y. (2015, January 7\u201313). Harvesting discriminative meta objects with deep CNN features for scene classification. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.152"},{"key":"ref_75","doi-asserted-by":"crossref","unstructured":"Yang, S., and Ramanan, D. (2015, January 7\u201313). Multi-scale recognition with DAG-CNNs. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.144"},{"key":"ref_76","doi-asserted-by":"crossref","unstructured":"Liu, Y., Chen, Q., Chen, W., and Wassell, I. (2018, January 2\u20137). Dictionary learning inspired deep network for scene recognition. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12312"},{"key":"ref_77","doi-asserted-by":"crossref","first-page":"82066","DOI":"10.1109\/ACCESS.2020.2989863","article-title":"FOSNet: An End-to End Trainable Deep Neural Network for Scene Recognition","volume":"8","author":"Seong","year":"2020","journal-title":"IEEE Access"},{"key":"ref_78","doi-asserted-by":"crossref","first-page":"1263","DOI":"10.1109\/TCSVT.2015.2511543","article-title":"Hybrid CNN and dictionary-based models for scene recognition and domain adaptation","volume":"27","author":"Xie","year":"2017","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_79","doi-asserted-by":"crossref","unstructured":"Xiao, J., Hays, J., Ehinger, K.A., Oliva, A., and Torralba, A. (2010, January 13\u201318). SUN database: Large-scale scene recognition from abbey to zoo. Proceedings of the IEEE International Conference on Computer Vision, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539970"},{"key":"ref_80","doi-asserted-by":"crossref","unstructured":"Quattoni, A., and Torralba, A. (2009, January 20\u201325). Recognizing Indoor Scenes. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206537"},{"key":"ref_81","doi-asserted-by":"crossref","unstructured":"Li, L., and Fei-Fei, L. (2007, January 14\u201321). What, where and who? Classifying events by scene and object recognition. Proceedings of the IEEE International Conference on Computer Vision, Rio de Janeiro, Brazil.","DOI":"10.1109\/ICCV.2007.4408872"},{"key":"ref_82","doi-asserted-by":"crossref","first-page":"1452","DOI":"10.1109\/TPAMI.2017.2723009","article-title":"Places: A 10 million image database for scene recognition","volume":"40","author":"Zhou","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_83","doi-asserted-by":"crossref","first-page":"2028","DOI":"10.1109\/TIP.2017.2666739","article-title":"Weakly supervised PatchNets: Describing and aggregating local patches for scene recognition","volume":"26","author":"Wang","year":"2017","journal-title":"IEEE Trans. Image Process."},{"key":"ref_84","unstructured":"Vogel, J., and Schiele, B. (2004). Pattern Recognition, Proceedings of the 26th DAGM Symposium, 30 August\u20131 September 2004, Springer."},{"key":"ref_85","doi-asserted-by":"crossref","first-page":"042616","DOI":"10.1117\/1.JRS.11.042616","article-title":"Deep feature extraction and combination for synthetic aperture radar target classification","volume":"11","author":"Amrani","year":"2017","journal-title":"J. Appl. Remote Sens."},{"key":"ref_86","doi-asserted-by":"crossref","unstructured":"Song, Y., Zhang, Z., Liu, L., Rahimpour, A., and Qi, H. (2017). Computer Vision\u2014ACCV 2016. ACCV 2016, Proceedings of the Asian Conference on Computer Vision, Taipei, Taiwan, 20\u201324 November 2016, Springer.","DOI":"10.1007\/978-3-319-54181-5_2"},{"key":"ref_87","doi-asserted-by":"crossref","first-page":"493","DOI":"10.1109\/TPAMI.2013.113","article-title":"Feature coding in image classification: A comprehensive study","volume":"36","author":"Huang","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_88","doi-asserted-by":"crossref","first-page":"1143","DOI":"10.1109\/LSP.2016.2641020","article-title":"Discovering class-specific spatial layouts for scene recognition","volume":"24","author":"Weng","year":"2017","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_89","first-page":"346","article-title":"Spatial pyramid pooling in deep convolutional networks for visual recognition","volume":"33","author":"He","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_90","doi-asserted-by":"crossref","unstructured":"Pandey, M., and Lazebnik, S. (2011, January 6\u201313). Scene recognition and weakly supervised object localization with deformable part-based models. Proceedings of the IEEE International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126383"},{"key":"ref_91","unstructured":"Elharrouss, O., Al-Maadeed, S., Subramanian, N., Ottakath, N., Almaadeed, N., and Himeur, Y. (2021). Panoptic Segmentation: A Review. arXiv."},{"key":"ref_92","doi-asserted-by":"crossref","unstructured":"Minaee, S., Boykov, Y., Porikli, F., Plaza, A., Kehtarnavaz, N., and Terzopoulos, D. (2020). Image Segmentation Using Deep Learning: A Survey. arXiv.","DOI":"10.1109\/TPAMI.2021.3059968"}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/17\/5\/189\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:36:47Z","timestamp":1760107007000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/17\/5\/189"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,30]]},"references-count":92,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2024,5]]}},"alternative-id":["a17050189"],"URL":"https:\/\/doi.org\/10.3390\/a17050189","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,30]]}}}