{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T07:58:03Z","timestamp":1784879883223,"version":"3.55.0"},"reference-count":43,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2022,7,26]],"date-time":"2022-07-26T00:00:00Z","timestamp":1658793600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Indoor scene recognition and semantic information can be helpful for social robots. Recently, in the field of indoor scene recognition, researchers have incorporated object-level information and shown improved performances. This paper demonstrates that scene recognition can be performed solely using object-level information in line with these advances. A state-of-the-art object detection model was trained to detect objects typically found in indoor environments and then used to detect objects in scene data. These predicted objects were then used as features to predict room categories. This paper successfully combines approaches conventionally used in computer vision and natural language processing (YOLO and TF-IDF, respectively). These approaches could be further helpful in the field of embodied research and dynamic scene classification, which we elaborate on.<\/jats:p>","DOI":"10.3390\/jimaging8080209","type":"journal-article","created":{"date-parts":[[2022,7,27]],"date-time":"2022-07-27T04:59:16Z","timestamp":1658897956000},"page":"209","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":24,"title":["Indoor Scene Recognition via Object Detection and TF-IDF"],"prefix":"10.3390","volume":"8","author":[{"given":"Edvard","family":"Heikel","sequence":"first","affiliation":[{"name":"Department of Business Management and Analytics, Arcada University of Applied Sciences, 00550 Helsinki, Finland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6861-8024","authenticated-orcid":false,"given":"Leonardo","family":"Espinosa-Leal","sequence":"additional","affiliation":[{"name":"Department of Business Management and Analytics, Arcada University of Applied Sciences, 00550 Helsinki, Finland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,7,26]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Narasimhan, M., Wijmans, E., Chen, X., Darrell, T., Batra, D., Parikh, D., and Singh, A. (2020). Seeing the Un-Scene: Learning Amodal Semantic Maps for Room Navigation. arXiv.","DOI":"10.1007\/978-3-030-58523-5_30"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Othman, K., and Rad, A. (2019). An indoor room classification system for social robots via integration of CNN and ECOC. Appl. Sci., 9.","DOI":"10.3390\/app9030470"},{"key":"ref_3","unstructured":"Kwon, O., and Oh, S. (2020, January 13\u201316). Learning to use topological memory for visual navigation. Proceedings of the 20th International Conference on Control, Automation and Systems, Busan, Korea."},{"key":"ref_4","unstructured":"Zhu, Y., Mottaghi, R., Kolve, E., Lim, J., Gupta, A., Fei-Fei, L., and Farhadi, A. (June, January 29). Target-driven visual navigation in indoor scenes using deep reinforcement learning. Proceedings of the IEEE International Conference on Robotics and Automation, Singapore."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1227","DOI":"10.1007\/s00371-016-1348-3","article-title":"Indoor scene modeling from a single image using normal inference and edge features","volume":"33","author":"Liu","year":"2017","journal-title":"Vis. Comput."},{"key":"ref_6","unstructured":"Chaplot, D., Gandhi, D., Gupta, A., and Salakhutdinov, R. (2020). Object Goal Navigation using Goal-Oriented Semantic Exploration. arXiv."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2691","DOI":"10.1007\/s00371-021-02147-w","article-title":"Semantic scene synthesis: Application to assistive systems","volume":"38","author":"Zatout","year":"2021","journal-title":"Vis. Comput."},{"key":"ref_8","unstructured":"Yang, W., Wang, X., Farhadi, A., Gupta, G., and Mottaghi, R. (2018). Visual semantic navigation using scene priors. arXiv."},{"key":"ref_9","first-page":"975","article-title":"Text mining: Use of TF-IDF to example the relevance of words to documents","volume":"181","author":"Qaiser","year":"2018","journal-title":"Int. J. Comput. Appl."},{"key":"ref_10","unstructured":"Ramos, J. (2003, January 21\u201324). Using TF-IDF to determine word relevance in document queries. Proceedings of the First Instructional Conference on Machine Learning, Piscataway, NJ, USA."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Dadgar, S., Araghi, M., and Farahani, M. (2016, January 17\u201318). A novel text mining approach based on TF-IDF and support vector machine for news classification. Proceedings of the IEEE International Conference on Engineering and Technology, Coimbatore, India.","DOI":"10.1109\/ICETECH.2016.7569223"},{"key":"ref_12","unstructured":"Teder, M., Mayor-Torres, J., and Teufel, C. (2009). Deriving visual semantics from spatial context: An adaptation of LSA and Word2Vec to generate object and scene embeddings from images. arXiv."},{"key":"ref_13","unstructured":"Chen, B., Sahdev, R., Wu, D., Zhao, X., Papagelis, M., and Tsotsos, J. (2019). Scene Classification in Indoor Environments for Robots using Context Based Word Embeddings. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Quattoni, A., and Torralba, A. (2009, January 20\u201325). Recognizing indoor scenes. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPRW.2009.5206537"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Matei, A., Glavan, A., and Talavera, E. (2020, January 11\u201313). Deep learning for scene recognition from visual data: A survey. Proceedings of the International Conference on Hybrid Artificial Intelligence Systems, Gij\u00f3n, Spain.","DOI":"10.1007\/978-3-030-61705-9_64"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Yang, J., Jiang, Y.G., Hauptmann, A., and Ngo, C.W. (2007, January 24\u201329). Evaluating bag-of-visual-words representations in scene classification. Proceedings of the International Workshop on Multimedia Information Retrieval, Bavaria, Germany.","DOI":"10.1145\/1290082.1290111"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"2055","DOI":"10.1109\/TIP.2017.2675339","article-title":"Knowledge guided disambiguation for large-scale scene classification with multi-resolution","volume":"26","author":"Wang","year":"2017","journal-title":"CNNs IEEE Trans. Image"},{"key":"ref_18","unstructured":"Liao, Y., Kodagoda, S., Wang, Y., Shi, L., and Liu, Y. (2016, January 16\u201321). Understand scene categories by objects: A semantic regularized scene classifier using Convolutional Neural Networks. Proceedings of the IEEE International Conference on Robotics and Automation, Stockholm, Sweden."},{"key":"ref_19","unstructured":"Yao, J., Fidler, S., and Urtasun, R. (2012, January 16\u201321). Describing the scene as a whole: Joint object detection, scene classification and semantic segmentation. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Providence, RI, USA."},{"key":"ref_20","unstructured":"Li, L.J., Su, H., Li, F.F., and P Xing, E. (2010). Object bank: A high- level image representation for scene classification &amp, semantic feature sparsification. In Advances in Neural Information Processing Systems; Carnegie Mellon University."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1007\/s00371-008-0294-0","article-title":"Toward a higher-level visual representation for object-based image retrieval","volume":"25","author":"Zheng","year":"2009","journal-title":"Vis. Comput."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified real-time object detection. Proceedings of the 28th IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, Nevada, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"683","DOI":"10.1002\/wcs.1254","article-title":"Latent semantic analysis","volume":"4","author":"Evangelopoulos","year":"2013","journal-title":"Wiley Interdiscip. Rev. Cogn. Sci."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_25","unstructured":"Simonyan, J. (2015). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zhou, L., Cen, J., Wang, X., Sun, Z., Lam, T.L., and Xu, Y. (October, January 27). Borm: Bayesian object relation model for indoor scene recognition. Proceedings of the 2021 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Czech Republic.","DOI":"10.1109\/IROS51168.2021.9636024"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Song, S., Lichtenberg, S.P., and Xiao, J. (2015, January 15). Sun rgb-d: A rgb-d scene understanding benchmark suite. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298655"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1452","DOI":"10.1109\/TPAMI.2017.2723009","article-title":"Places: A 10 million image database for scene recognition","volume":"40","author":"Zhou","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Miao, B., Zhou, L., Mian, A.S., Lam, T.L., and Xu, Y. (October, January 27). Object-to-scene: Learning to transfer object knowledge to indoor scene recognition. Proceedings of the 2021 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Prague, Czech Republic.","DOI":"10.1109\/IROS51168.2021.9636700"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Labinghisa, B.A., and Lee, D.M. (2022). Indoor localization system using deep learning based scene recognition. Multimed. Tools Appl.","DOI":"10.1109\/ICAIIC51459.2021.9415278"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"1956","DOI":"10.1007\/s11263-020-01316-z","article-title":"The open images dataset V4: Unified image classification, object detection, and visual relationship detection at scale","volume":"128","author":"Kuznetsova","year":"2020","journal-title":"Int. J. Comput. Vis."},{"key":"ref_32","unstructured":"Jocher, G., and Yolov5 (2021, July 01). Code Repository. Available online: https:\/\/github.com\/ultralytics\/yolov5."},{"key":"ref_33","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Espinosa Leal, L., Chapman, A., and Westerlund, M. (2019, January 8\u201311). Reinforcement learning for extended reality: Designing self-play scenarios. Proceedings of the 52nd Hawaii International Conference on System Sciences, Grand Wailea, HI, USA.","DOI":"10.24251\/HICSS.2019.020"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"8427","DOI":"10.3233\/JIFS-189161","article-title":"Autonomous industrial management via reinforcement learning","volume":"39","author":"Chapman","year":"2020","journal-title":"J. Intell. Fuzzy Syst."},{"key":"ref_36","unstructured":"Bochkovskiy, A., Wang, C.Y., and Liao, H.Y.M. (2020). Yolov4: Optimal speed and accuracy of object detection. arXiv."},{"key":"ref_37","unstructured":"Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017). Automatic Differentiation in Pytorch, NIPS-Workshop."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, faster, stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Chollet, F. (2017, January 21\u201326). Xception: Deep learning with depthwise separable convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.195"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., and Torralba, A. (2017, January 21\u201326). Scene parsing through ade20k dataset. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.544"},{"key":"ref_41","first-page":"2825","article-title":"Scikit-learn: Machine learning in python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref_42","unstructured":"Baeza-Yates, R., and Ribeiro-Neto, B. (1999). Modern Information Retrieval, ACM Press."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Heikel, E., and Espinosa-Leal, L. (2021, July 01). Trained Models and Datasets for Indoor Scene Recognition via Object Detection and TF-IDF. 2022. Available online: https:\/\/doi.org\/10.5281\/zenodo.6792296.","DOI":"10.20944\/preprints202207.0070.v1"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/8\/8\/209\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:56:31Z","timestamp":1760140591000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/8\/8\/209"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,26]]},"references-count":43,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2022,8]]}},"alternative-id":["jimaging8080209"],"URL":"https:\/\/doi.org\/10.3390\/jimaging8080209","relation":{"has-preprint":[{"id-type":"doi","id":"10.20944\/preprints202207.0070.v1","asserted-by":"object"}]},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,7,26]]}}}