{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T07:58:06Z","timestamp":1784879886937,"version":"3.55.0"},"reference-count":62,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2024,5,14]],"date-time":"2024-05-14T00:00:00Z","timestamp":1715644800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>Indoor scene classification plays a pivotal role in enabling social robots to seamlessly adapt to their environments, facilitating effective navigation and interaction within diverse indoor scenes. By accurately characterizing indoor scenes, robots can autonomously tailor their behaviors, making informed decisions to accomplish specific tasks. Traditional methods relying on manually crafted features encounter difficulties when characterizing complex indoor scenes. On the other hand, deep learning models address the shortcomings of traditional methods by autonomously learning hierarchical features from raw images. Despite the success of deep learning models, existing models still struggle to effectively characterize complex indoor scenes. This is because there is high degree of intra-class variability and inter-class similarity within indoor environments. To address this problem, we propose a dual-stream framework that harnesses both global contextual information and local features for enhanced recognition. The global stream captures high-level features and relationships across the scene. The local stream employs a fully convolutional network to extract fine-grained local information. The proposed dual-stream architecture effectively distinguishes scenes that share similar global contexts but contain different localized objects. We evaluate the performance of the proposed framework on a publicly available benchmark indoor scene dataset. From the experimental results, we demonstrate the effectiveness of the proposed framework.<\/jats:p>","DOI":"10.3390\/computers13050121","type":"journal-article","created":{"date-parts":[[2024,5,14]],"date-time":"2024-05-14T06:28:12Z","timestamp":1715668092000},"page":"121","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":21,"title":["Indoor Scene Classification through Dual-Stream Deep Learning: A Framework for Improved Scene Understanding in Robotics"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7406-8441","authenticated-orcid":false,"given":"Sultan Daud","family":"Khan","sequence":"first","affiliation":[{"name":"Department of Computer Science, National University of Technology, Islamabad 44000, Pakistan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0217-0751","authenticated-orcid":false,"given":"Kamal M.","family":"Othman","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, College of Engineering, Umm Al-Qura University, Makkah 24382, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,5,14]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"7265","DOI":"10.1109\/TCYB.2021.3052499","article-title":"Indoor place category recognition for a cleaning robot by fusing a probabilistic approach and deep learning","volume":"52","author":"Choe","year":"2021","journal-title":"IEEE Trans. Cybern."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1016\/j.ejor.2021.01.019","article-title":"Planning and control of autonomous mobile robots for intralogistics: Literature review and research agenda","volume":"294","author":"Fragapane","year":"2021","journal-title":"Eur. J. Oper. Res."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Ozkil, A.G., Fan, Z., Dawids, S., Aanes, H., Kristensen, J.K., and Christensen, K.H. (2009, January 5\u20137). Service robots for hospitals: A case study of transportation tasks in a hospital. Proceedings of the 2009 IEEE International Conference on Automation and Logistics, Shenyang, China.","DOI":"10.1109\/ICAL.2009.5262912"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Kyrarini, M., Lygerakis, F., Rajavenkatanarayanan, A., Sevastopoulos, C., Nambiappan, H.R., Chaitanya, K.K., Babu, A.R., Mathew, J., and Makedon, F. (2021). A survey of robots in healthcare. Technologies, 9.","DOI":"10.3390\/technologies9010008"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"382","DOI":"10.1016\/j.chb.2017.02.064","article-title":"Shopping with a robotic companion","volume":"77","author":"Bertacchini","year":"2017","journal-title":"Comput. Hum. Behav."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"235","DOI":"10.1080\/01691864.2019.1698460","article-title":"Restock and straightening system for retail automation using compliant and mobile manipulation","volume":"34","author":"Okada","year":"2020","journal-title":"Adv. Robot."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"58","DOI":"10.1016\/j.cogr.2021.06.001","article-title":"Substantial capabilities of robotics in enhancing industry 4.0 implementation","volume":"1","author":"Javaid","year":"2021","journal-title":"Cogn. Robot."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"21901","DOI":"10.1109\/ACCESS.2017.2760201","article-title":"Research on automatic parking systems based on parking scene recognition","volume":"5","author":"Ma","year":"2017","journal-title":"IEEE Access"},{"key":"ref_9","first-page":"1","article-title":"An improved deep network-based scene classification method for self-driving cars","volume":"71","author":"Ni","year":"2022","journal-title":"IEEE Trans. Instrum. Meas."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"102655","DOI":"10.1016\/j.jvcir.2019.102655","article-title":"Scene categorization towards urban tunnel traffic by image quality assessment","volume":"65","author":"Zhou","year":"2019","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"28","DOI":"10.23919\/JSEE.2023.000031","article-title":"Autonomous landing scene recognition based on transfer learning for drones","volume":"34","author":"Du","year":"2023","journal-title":"J. Syst. Eng. Electron."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"O\u2019Mahony, N., Campbell, S., Krpalkova, L., Riordan, D., Walsh, J., Murphy, A., and Ryan, C. (2018, January 21\u201322). Deep learning for visual navigation of unmanned ground vehicles: A review. Proceedings of the 2018 29th Irish Signals and Systems Conference (ISSC), Belfast, UK.","DOI":"10.1109\/ISSC.2018.8585381"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Ekici, M., Se\u00e7kin, A.\u00c7., \u00d6zek, A., and Karpuz, C. (2022). Warehouse drone: Indoor positioning and product counter with virtual fiducial markers. Drones, 7.","DOI":"10.3390\/drones7010003"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"103068","DOI":"10.1016\/j.autcon.2019.103068","article-title":"An integrated UGV-UAV system for construction site data collection","volume":"112","author":"Asadi","year":"2020","journal-title":"Autom. Constr."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Wijayathunga, L., Rassau, A., and Chai, D. (2023). Challenges and solutions for autonomous ground robot scene understanding and navigation in unstructured outdoor environments: A review. Appl. Sci., 13.","DOI":"10.20944\/preprints202304.0373.v1"},{"key":"ref_16","unstructured":"Tagarakis, A.C., Kalaitzidis, D., Filippou, E., Benos, L., and Bochtis, D. (2022). Information and Communication Technologies for Agriculture\u2014Theme III: Decision, Springer."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"424","DOI":"10.1016\/j.patcog.2012.07.017","article-title":"Scene classification using a multi-resolution bag-of-features model","volume":"46","author":"Zhou","year":"2013","journal-title":"Pattern Recognit."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Khan, N.Y., McCane, B., and Wyvill, G. (2011, January 6\u20138). SIFT and SURF performance evaluation against various image deformations on benchmark dataset. Proceedings of the 2011 International Conference on Digital Image Computing: Techniques and Applications, Noosa, QLD, Australia.","DOI":"10.1109\/DICTA.2011.90"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Ayers, B., and Boutell, M. (2007, January 17\u201322). Home interior classification using SIFT keypoint histograms. Proceedings of the 2007 IEEE Conference on Computer Vision and Pattern Recognition, Minneapolis, MN, USA.","DOI":"10.1109\/CVPR.2007.383485"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"346","DOI":"10.1016\/j.cviu.2007.09.014","article-title":"Speeded-up robust features (SURF)","volume":"110","author":"Bay","year":"2008","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1223","DOI":"10.1007\/s11042-020-09759-9","article-title":"Scale-space multi-view bag of words for scene categorization","volume":"80","author":"Giveki","year":"2021","journal-title":"Multimed. Tools Appl."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"381","DOI":"10.1109\/TCSVT.2010.2041828","article-title":"Contextual bag-of-words for visual categorization","volume":"21","author":"Li","year":"2010","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Ergul, E., and Arica, N. (2010, January 23\u201326). Scene classification using spatial pyramid of latent topics. Proceedings of the 2010 20th International Conference on Pattern Recognition, Istanbul, Turkey.","DOI":"10.1109\/ICPR.2010.879"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"118","DOI":"10.1016\/j.patcog.2018.04.025","article-title":"Improved spatial pyramid matching for scene recognition","volume":"82","author":"Xie","year":"2018","journal-title":"Pattern Recognit."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"017201","DOI":"10.1117\/1.OE.51.1.017201","article-title":"Scene classification based on spatial pyramid representation by superpixel lattices and contextual visual features","volume":"51","author":"Gu","year":"2012","journal-title":"Opt. Eng."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"28405","DOI":"10.1007\/s11042-022-12481-3","article-title":"Indoor localization system using deep learning based scene recognition","volume":"81","author":"Labinghisa","year":"2022","journal-title":"Multimed. Tools Appl."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"116382","DOI":"10.1016\/j.eswa.2021.116382","article-title":"DeepScene: Scene classification via convolutional neural network with spatial pyramid pooling","volume":"193","author":"Yee","year":"2022","journal-title":"Expert Syst. Appl."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Wozniak, P., Afrisal, H., Esparza, R.G., and Kwolek, B. (2018, January 17\u201319). Scene recognition for indoor localization of mobile robots using deep CNN. Proceedings of the Computer Vision and Graphics: International Conference, ICCVG 2018, Warsaw, Poland. Proceedings.","DOI":"10.1007\/978-3-030-00692-1_13"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"2725","DOI":"10.1007\/s00371-022-02488-0","article-title":"NIR\/RGB image fusion for scene classification using deep neural networks","volume":"39","author":"Soroush","year":"2022","journal-title":"Vis. Comput."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Heikel, E., and Espinosa-Leal, L. (2022). Indoor scene recognition via object detection and TF-IDF. J. Imaging, 8.","DOI":"10.20944\/preprints202207.0070.v1"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Biswas, M., Buckchash, H., and Prasad, D.K. (2023). pNNCLR: Stochastic Pseudo Neighborhoods for Contrastive Learning based Unsupervised Representation Learning Problems. arXiv.","DOI":"10.1016\/j.neucom.2024.127810"},{"key":"ref_32","unstructured":"Swadzba, A., and Wachsmuth, S. (2010). Asian Conference on Computer Vision, Springer."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"646","DOI":"10.1016\/j.robot.2012.10.006","article-title":"A detailed analysis of a new 3D spatial feature vector for indoor scene classification","volume":"62","author":"Swadzba","year":"2014","journal-title":"Robot. Auton. Syst."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Li, X., and Guo, Y. (2014, January 6\u201312). Multi-level adaptive active learning for scene classification. Proceedings of the Computer Vision\u2013ECCV 2014: 13th European Conference, Zurich, Switzerland. Proceedings, Part VII 13.","DOI":"10.1007\/978-3-319-10584-0_16"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"483","DOI":"10.1016\/j.patcog.2012.08.006","article-title":"Pairwise constraints based multiview features fusion for scene classification","volume":"46","author":"Yu","year":"2013","journal-title":"Pattern Recognit."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"204","DOI":"10.1007\/s11263-014-0779-4","article-title":"Indoor scene understanding with geometric and semantic contexts","volume":"112","author":"Choi","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"683","DOI":"10.1109\/LSP.2011.2170165","article-title":"Efficient learning of sample-specific discriminative features for scene classification","volume":"18","author":"Han","year":"2011","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Zuo, Z., Wang, G., Shuai, B., Zhao, L., Yang, Q., and Jiang, X. (2014, January 6\u201312). Learning discriminative and shareable features for scene classification. Proceedings of the Computer Vision\u2013ECCV 2014: 13th European Conference, Zurich, Switzerland. Proceedings, Part I 13.","DOI":"10.1007\/978-3-319-10590-1_36"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Espinace, P., Kollar, T., Soto, A., and Roy, N. (2010, January 3\u20137). Indoor scene recognition through object detection. Proceedings of the 2010 IEEE International Conference on Robotics and Automation, Anchorage, AK, USA.","DOI":"10.1109\/ROBOT.2010.5509682"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Margolin, R., Zelnik-Manor, L., and Tal, A. (2014, January 6\u201312). Otc: A novel local descriptor for scene classification. Proceedings of the Computer Vision\u2013ECCV 2014: 13th European Conference, Zurich, Switzerland. Proceedings, Part VII 13.","DOI":"10.1007\/978-3-319-10584-0_25"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1016\/j.eswa.2016.10.038","article-title":"Growing random forest on deep convolutional neural networks for scene categorization","volume":"71","author":"Bai","year":"2017","journal-title":"Expert Syst. Appl."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Khan, S.H., Hayat, M., and Porikli, F. (2017, January 22\u201329). Scene categorization with spectral features. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.601"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Pereira, R., Gon\u00e7alves, N., Garrote, L., Barros, T., Lopes, A., and Nunes, U.J. (2020, January 15\u201317). Deep-learning based global and semantic feature fusion for indoor scene classification. Proceedings of the 2020 IEEE International Conference on Autonomous Robot Systems and Competitions (ICARSC), Ponta Delgada, Portugal.","DOI":"10.1109\/ICARSC49921.2020.9096068"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Pereira, R., Garrote, L., Barros, T., Lopes, A., and Nunes, U.J. (October, January 27). A deep learning-based indoor scene classification approach enhanced with inter-object distance semantic features. Proceedings of the 2021 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Prague, Czech Republic.","DOI":"10.1109\/IROS51168.2021.9636242"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"82066","DOI":"10.1109\/ACCESS.2020.2989863","article-title":"FOSNet: An end-to-end trainable deep neural network for scene recognition","volume":"8","author":"Seong","year":"2020","journal-title":"IEEE Access"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"4829","DOI":"10.1109\/TIP.2016.2599292","article-title":"A spatial layout and scale invariant feature representation for indoor scene classification","volume":"25","author":"Hayat","year":"2016","journal-title":"IEEE Trans. Image Process."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Guo, W., Wu, R., Chen, Y., and Zhu, X. (2018). Deep learning scene recognition method based on localization enhancement. Sensors, 18.","DOI":"10.3390\/s18103376"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"440","DOI":"10.1016\/j.procs.2020.03.253","article-title":"Indoor home scene recognition using capsule neural networks","volume":"167","author":"Basu","year":"2020","journal-title":"Procedia Comput. Sci."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Sun, N., Zhu, X., Liu, J., and Han, G. (2017, January 29\u201331). Indoor scene recognition based on deep learning and sparse representation. Proceedings of the 2017 13th International Conference on Natural Computation, Fuzzy Systems and Knowledge Discovery (ICNC-FSKD), Guilin, China.","DOI":"10.1109\/FSKD.2017.8393385"},{"key":"ref_50","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K.Q. (2017, January 21\u201326). Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_52","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"1217","DOI":"10.2991\/ijcis.d.210326.001","article-title":"Multi-scale person localization with multi-stage deep sequential framework","volume":"14","author":"Khan","year":"2021","journal-title":"Int. J. Comput. Intell. Syst."},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"864","DOI":"10.1109\/LGRS.2018.2888887","article-title":"Scale adaptive proposal network for object detection in remote sensing images","volume":"16","author":"Zhang","year":"2019","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_56","unstructured":"Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2024, March 23). Automatic Differentiation in Pytorch. Available online: https:\/\/openreview.net\/pdf\/25b8eee6c373d48b84e5e9c6e10e7cbbbce4ac73.pdf?ref=blog.premai.io."},{"key":"ref_57","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_58","first-page":"1","article-title":"SRIN: A new dataset for social robot indoor navigation","volume":"4","author":"Othman","year":"2020","journal-title":"Glob. J. Eng. Sci."},{"key":"ref_59","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20136). Imagenet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA."},{"key":"ref_60","unstructured":"Tan, M., and Le, Q. (2019, January 9\u201315). Efficientnet: Rethinking model scaling for convolutional neural networks. Proceedings of the International Conference on Machine Learning, PMLR, Long Beach, CA, USA."},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhou, X., Lin, M., and Sun, J. (2018, January 18\u201322). Shufflenet: An extremely efficient convolutional neural network for mobile devices. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00716"},{"key":"ref_62","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv."}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/13\/5\/121\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:41:57Z","timestamp":1760107317000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/13\/5\/121"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,14]]},"references-count":62,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2024,5]]}},"alternative-id":["computers13050121"],"URL":"https:\/\/doi.org\/10.3390\/computers13050121","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5,14]]}}}