{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,26]],"date-time":"2026-02-26T16:10:41Z","timestamp":1772122241278,"version":"3.50.1"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2022,1,27]],"date-time":"2022-01-27T00:00:00Z","timestamp":1643241600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100010418","name":"Institute for Information & communications Technology Promotion","doi-asserted-by":"crossref","award":["2016-0-00406"],"award-info":[{"award-number":["2016-0-00406"]}],"id":[{"id":"10.13039\/501100010418","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"crossref","award":["NRF-2019R1A2C1006608"],"award-info":[{"award-number":["NRF-2019R1A2C1006608"]}],"id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"crossref"}]},{"name":"BK21 FOUR program of the Ministry of Education","award":["NRF5199991014091"],"award-info":[{"award-number":["NRF5199991014091"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2022,1,31]]},"abstract":"<jats:p>\n            In recent years, dynamic scene understanding has gained attention from researchers because of its widespread applications. The main important factor in successfully understanding the dynamic scenes lies in jointly representing the appearance and motion features to obtain an informative description. Numerous methods have been introduced to solve dynamic scene recognition problem, nevertheless, a few concerns still need to be investigated. In this article, we introduce a novel multi-modal network for dynamic scene understanding from video data, which captures both spatial appearance and temporal dynamics effectively. Furthermore, two-level joint tuning layers are proposed to integrate the global and local spatial features as well as spatial and temporal stream deep features. In order to extract the temporal information, we present a novel dynamic descriptor, namely,\n            <jats:bold>Volume Symmetric Gradient Local Graph Structure<\/jats:bold>\n            (\n            <jats:bold>VSGLGS<\/jats:bold>\n            ), which generates temporal feature maps similar to optical flow maps. However, this approach overcomes the issues of optical flow maps. Additionally,\n            <jats:bold>Volume Local Directional Transition Pattern<\/jats:bold>\n            (\n            <jats:bold>VLDTP<\/jats:bold>\n            ) based handcrafted spatiotemporal feature descriptor is also introduced, which extracts the directional information through exploiting edge responses. Lastly, a stacked\n            <jats:bold>Bidirectional Long Short-Term Memory<\/jats:bold>\n            (\n            <jats:bold>Bi-LSTM<\/jats:bold>\n            ) network along with a temporal mixed pooling scheme is designed to achieve the dynamic information without noise interference. The extensive experimental investigation proves that the proposed multi-modal network outperforms most of the state-of-the-art approaches for dynamic scene understanding.\n          <\/jats:p>","DOI":"10.1145\/3462218","type":"journal-article","created":{"date-parts":[[2022,1,27]],"date-time":"2022-01-27T19:44:21Z","timestamp":1643312661000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["A Novel Multi-Modal Network-Based Dynamic Scene Understanding"],"prefix":"10.1145","volume":"18","author":[{"given":"Md Azher","family":"Uddin","sequence":"first","affiliation":[{"name":"Department of Artificial Intelligence, Ajou University, Suwon-si, Gyeonggi-do, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joolekha Bibi","family":"Joolee","sequence":"additional","affiliation":[{"name":"Department of Computer Scienceand Engineering, Kyung Hee University, Yongin-si, Gyeonggi-do, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Young-Koo","family":"Lee","sequence":"additional","affiliation":[{"name":"Department of Computer Scienceand Engineering, Kyung Hee University, Yongin-si, Gyeonggi-do, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kyung-Ah","family":"Sohn","sequence":"additional","affiliation":[{"name":"Department of Software and Computer Engineering, and Department of Artificial Intelligence, Ajou University, Suwon-si, Gyeonggi-do, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,1,27]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2878865"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2014.04.006"},{"key":"e_1_3_1_4_2","unstructured":"S. Abu-El-Haija N. Kothari J. Lee P. Natsev G. Toderici B. Varadarajan and S. Vijayanarasimhan. 2016. Youtube-8M: A large-scale video classification benchmark. https:\/\/arxiv.org\/pdf\/1609.08675.pdf."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2016.2593719"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.5555\/2354409.2354989"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.04.110"},{"key":"e_1_3_1_8_2","article-title":"A convolution bidirectional long short-term memory neural network for driver emotion recognition","author":"Du G.","year":"2020","unstructured":"G. Du, Z. Wang, B. Gao, S. Mumtaz, K. M. Abualnaja, and C. Du. 2020. A convolution bidirectional long short-term memory neural network for driver emotion recognition. IEEE Transactions on Intelligent Transportation Systems (2020). DOI:https:\/\/doi.org\/10.1109\/TITS.2020.3007357","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.5244\/C.27.56"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.343"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2526008"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.786"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2016.08.035"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2016.06.055"},{"key":"e_1_3_1_15_2","first-page":"320","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201918)","author":"Hadji I.","year":"2010","unstructured":"I. Hadji and R. P. Wildes. 2010. A new large scale dynamic texture dataset with application to ConvNet understanding. In Proceedings of the European Conference on Computer Vision (ECCV\u201918). Springer, 320\u2013335."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2017.08.046"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2018.2823360"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018497"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/AVSS.2010.17"},{"key":"e_1_3_1_21_2","article-title":"Video-based depression level analysis by encoding deep spatiotemporal features","author":"Jazaery M. A.","year":"2018","unstructured":"M. A. Jazaery and G. Guo. 2018. Video-based depression level analysis by encoding deep spatiotemporal features. IEEE Transactions on Affective Computing (2018). DOI:https:\/\/doi.org\/10.1109\/TAFFC.2018.2870884","journal-title":"IEEE Transactions on Affective Computing"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3231738"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2010.09.137"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.68"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.dsp.2020.102742"},{"key":"e_1_3_1_26_2","volume-title":"Proceedings of the Digital Image Computing: Techniques and Applications (DICTA\u201920)","author":"Peng X.","year":"2020","unstructured":"X. Peng and A. Bouzerdoum. 2020. Part-based feature aggregation method for dynamic scene recognition. In Proceedings of the Digital Image Computing: Techniques and Applications (DICTA\u201920). IEEE."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2015.07.071"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2010.5539864"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.5555\/2968826.2968890"},{"key":"e_1_3_1_30_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201915)","author":"Simonyan K.","year":"2015","unstructured":"K. Simonyan and A. Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In Proceedings of the International Conference on Learning Representations (ICLR\u201915)."},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.5555\/3298023.3298188"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2016.11.023"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10489-019-01589-z"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.336"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIE.2018.2881943"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-69900-4_75"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2013.110"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1162\/089976602317318938"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.03.041"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2007.1110"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2017.2650899"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3462218","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3462218","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:48:53Z","timestamp":1750193333000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3462218"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,27]]},"references-count":41,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,1,31]]}},"alternative-id":["10.1145\/3462218"],"URL":"https:\/\/doi.org\/10.1145\/3462218","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,27]]},"assertion":[{"value":"2020-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-01-27","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}