{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T07:57:52Z","timestamp":1784879872296,"version":"3.55.0"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2021,5,11]],"date-time":"2021-05-11T00:00:00Z","timestamp":1620691200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2021,5,31]]},"abstract":"<jats:p>In this work, we address the task of scene recognition from image data. A scene is a spatially correlated arrangement of various visual semantic contents also known as concepts, e.g., \u201cchair,\u201d\u00a0 \u201ccar,\u201d\u00a0 \u201csky,\u201d\u00a0 etc. Representation learning using visual semantic content can be regarded as one of the most trivial ideas as it mimics the human behavior of perceiving visual information. Semantic multinomial (SMN) representation is one such representation that captures semantic information using posterior probabilities of concepts. The core part of obtaining SMN representation is the building of concept models. Therefore, it is necessary to have ground-truth (true) concept labels for every concept present in an image. Moreover, manual labeling of concepts is practically not feasible due to the large number of images in the dataset. To address this issue, we propose an approach for generating pseudo-concepts in the absence of true concept labels. We utilize the pre-trained deep CNN-based architectures where activation maps (filter responses) from convolutional layers are considered as initial cues to the pseudo-concepts. The non-significant activation maps are removed using the proposed filter-specific threshold-based approach that leads to the removal of non-prominent concepts from data. Further, we propose a grouping mechanism to group the same pseudo-concepts using subspace modeling of filter responses to achieve a non-redundant representation. Experimental studies show that generated SMN representation using pseudo-concepts achieves comparable results for scene recognition tasks on standard datasets like MIT-67 and SUN-397 even in the absence of true concept labels.<\/jats:p>","DOI":"10.1145\/3436494","type":"journal-article","created":{"date-parts":[[2021,5,12]],"date-time":"2021-05-12T00:56:03Z","timestamp":1620780963000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["Visual Semantic-Based Representation Learning Using Deep CNNs for Scene Recognition"],"prefix":"10.1145","volume":"17","author":[{"given":"Shikha","family":"Gupta","sequence":"first","affiliation":[{"name":"Indian Institute of Technology Mandi, Mandi, H.P."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Krishan","family":"Sharma","sequence":"additional","affiliation":[{"name":"Indian Institute of Technology Mandi, Mandi, H.P."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dileep Aroor","family":"Dinesh","sequence":"additional","affiliation":[{"name":"Indian Institute of Technology Mandi, Mandi, H.P."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Veena","family":"Thenkanidiyoor","sequence":"additional","affiliation":[{"name":"National Institute of Technology Goa"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,5,11]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the International Conference on Image Processing (ICIP\u201903)","volume":"3","author":"Barla A.","year":"2003"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2004.03.009"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1961189.1961199"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.25.76"},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the British Machine Vision Conference (BMVC\u201914)","author":"Chatfield K."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2017.09.025"},{"key":"e_1_2_1_7_1","volume-title":"Chung and Fan Chung Graham","author":"Fan R.","year":"1997"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2005.177"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2013.2293512"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915)","author":"Dixit M.","year":"2015"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/3044805.3044879"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/1390681.1442794"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2015.2469281"},{"key":"e_1_2_1_15_1","volume-title":"Net2Vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks. arXiv preprint arXiv:1801.03454 (March","author":"Fong Ruth","year":"2018"},{"key":"e_1_2_1_16_1","volume-title":"Deep spatial pyramid: The devil is once again in the details. arXiv preprint arXiv:1504.05277","author":"Gao Bin-Bin","year":"2015"},{"key":"e_1_2_1_17_1","volume-title":"van Loan","author":"Golub Gene H.","year":"2013"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10584-0_26"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/NCC.2017.8077077"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5220\/0006596101410148"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1080\/13506280444000544"},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916)","author":"Herranz Luis"},{"key":"e_1_2_1_24_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201910)","author":"J\u00e9gou Herv\u00e9"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3231738"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2016.2567076"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.5555\/2999134.2999257"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2013.2271476"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-013-0660-x"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/2999792.2999899"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0945-y"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206536"},{"key":"e_1_2_1_33_1","article-title":"Visualizing data using t-SNE","author":"van der Maaten Laurens","year":"2008","journal-title":"Journal of Machine Learning Research 9"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1011139631724"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.5555\/2354409.2355028"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.5555\/1888089.1888101"},{"key":"e_1_2_1_37_1","volume-title":"National Conference on Computer Vision, Pattern Recognition, Image Processing, and Graphics. Springer, 400--409","author":"Pradhan Deepak Kumar","year":"2017"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206537"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2007.900138"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2011.175"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-013-0636-x"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2989863"},{"key":"e_1_2_1_43_1","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201918)","author":"Sharma Krishan"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2812600"},{"key":"e_1_2_1_45_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint:1409.1556 (September","author":"Simonyan Karen","year":"2014"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2925002"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2018.2848543"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2016.11.023"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2016.2614862"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-27814-6_27"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11222-007-9033-z"},{"key":"e_1_2_1_53_1","volume-title":"Proceedings of IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201902)","author":"Wan V."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.152"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2010.5539970"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2015.2511543"},{"key":"e_1_2_1_57_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201909)","author":"Yang Jianchao","year":"2009"},{"key":"e_1_2_1_58_1","volume-title":"Fisher kernel for deep neural activations. arXiv preprint arXiv:1412.1628","author":"Yoo Donggeun","year":"2014"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2015.7301274"},{"key":"e_1_2_1_60_1","volume-title":"Proceedings of the Deep Learning Workshop in International Conference on Machine Learning (ICML\u201915)","author":"Yosinski Jason","year":"2015"},{"key":"e_1_2_1_61_1","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201914)","author":"Matthew"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2817070"},{"key":"e_1_2_1_63_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915)","author":"Zhao Fang"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2723009"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.5555\/2968826.2968881"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3436494","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3436494","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T17:45:16Z","timestamp":1750268716000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3436494"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,11]]},"references-count":65,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2021,5,31]]}},"alternative-id":["10.1145\/3436494"],"URL":"https:\/\/doi.org\/10.1145\/3436494","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,5,11]]},"assertion":[{"value":"2019-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-05-11","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}