{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T01:20:44Z","timestamp":1782868844544,"version":"3.54.5"},"reference-count":68,"publisher":"Association for Computing Machinery (ACM)","issue":"2s","license":[{"start":{"date-parts":[[2020,4,30]],"date-time":"2020-04-30T00:00:00Z","timestamp":1588204800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Provincial Natural Science Foundation of Zhejiang","award":["Y19F020071,Y18F020084"],"award-info":[{"award-number":["Y19F020071,Y18F020084"]}]},{"name":"Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University","award":["93K172016K08"],"award-info":[{"award-number":["93K172016K08"]}]},{"name":"Natural Science Foundation of the Jiangsu Higher Education Institutions of China","award":["19KJA230001"],"award-info":[{"award-number":["19KJA230001"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61972059, 61702055, 61773272"],"award-info":[{"award-number":["61972059, 61702055, 61773272"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100004608","name":"Natural Science Foundation of Jiangsu Province","doi-asserted-by":"crossref","award":["BK20191474"],"award-info":[{"award-number":["BK20191474"]}],"id":[{"id":"10.13039\/501100004608","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2020,4,30]]},"abstract":"<jats:p>Event recognition in surveillance video has gained extensive attention from the computer vision community. This process still faces enormous challenges due to the tiny inter-class variations that are caused by various facets, such as severe occlusion, cluttered backgrounds, and so forth. To address these issues, we propose a spatio-temporal deep residual network with hierarchical attentions (STDRN-HA) for video event recognition. In the first attention layer, the ResNet fully connected feature guides the Faster R-CNN feature to generate object-based attention (O-attention) for target objects. In the second attention layer, the O-attention further guides the ResNet convolutional feature to yield the holistic attention (H-attention) in order to perceive more details of the occluded objects and the global background. In the third attention layer, the attention maps use the deep features to obtain the attention-enhanced features. Then, the attention-enhanced features are input into a deep residual recurrent network, which is used to mine more event clues from videos. Furthermore, an optimized loss function named softmax-RC is designed, which embeds the residual block regularization and center loss to solve the vanishing gradient in a deep network and enlarge the distance between inter-classes. We also build a temporal branch to exploit the long- and short-term motion information. The final results are obtained by fusing the outputs of the spatial and temporal streams. Experiments on the four realistic video datasets, CCV, VIRAT 1.0, VIRAT 2.0, and HMDB51, demonstrate that the proposed method has good performance and achieves state-of-the-art results.<\/jats:p>","DOI":"10.1145\/3378026","type":"journal-article","created":{"date-parts":[[2020,6,22]],"date-time":"2020-06-22T02:49:20Z","timestamp":1592794160000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Spatio-Temporal Deep Residual Network with Hierarchical Attentions for Video Event Recognition"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1188-0483","authenticated-orcid":false,"given":"Yonggang","family":"Li","sequence":"first","affiliation":[{"name":"Jiaxing University, China and Soochow University, Jiaxing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chunping","family":"Liu","sequence":"additional","affiliation":[{"name":"Soochow University, Suzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yi","family":"Ji","sequence":"additional","affiliation":[{"name":"Soochow University, Suzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shengrong","family":"Gong","sequence":"additional","affiliation":[{"name":"Changshu Institute of Science and Technology, China and Soochow University, China and Beijing Jiaotong University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haibao","family":"Xu","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,6,21]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"The IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 1314--1321","author":"Mohamed"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2846411"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298878"},{"key":"e_1_2_1_6_1","unstructured":"Christoph Feichtenhofer Axel Pinz and Richard Wildes. 2016. Spatiotemporal residual networks for video action recognition. In Advances in Neural Information Processing Systems (NIPS). 3468--3476.  Christoph Feichtenhofer Axel Pinz and Richard Wildes. 2016. Spatiotemporal residual networks for video action recognition. In Advances in Neural Information Processing Systems (NIPS). 3468--3476."},{"key":"e_1_2_1_7_1","volume-title":"The IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 4768--4777","author":"Feichtenhofer Christoph"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.213"},{"key":"e_1_2_1_9_1","first-page":"1112","article-title":"Hierarchical LSTMs with adaptive attention for visual captioning","volume":"42","author":"Gao Lianli","year":"2019","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.227"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.337"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU.2013.6707742"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2884478"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00685"},{"key":"e_1_2_1_15_1","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. In Computer Vision and Pattern Recognition (CVPR).  Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. In Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8682583"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2016.7552981"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2904996"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1282280.1282352"},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of ACM International Conference on Multimedia Retrieval (ICMR).","author":"Jiang Yu Gang"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2771462"},{"key":"e_1_2_1_22_1","volume-title":"The IEEE International Conference on Computer Vision (ICCV). 2556--2563","author":"Kuehne Hilde","year":"2011"},{"key":"e_1_2_1_23_1","volume-title":"Recognition of human actions using CNN-GWO: A novel modeling of CNN for enhancement of classification performance. Multimedia Tools 8 Applications 77, 18","author":"Kumaran N.","year":"2018"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.115"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.394"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2670782"},{"key":"e_1_2_1_27_1","first-page":"2852","article-title":"Deep residual dual unidirectional DLSTM for video event recognition with spatial-temporal consistency","volume":"41","author":"Li Yonggang","year":"2018","journal-title":"Chinese Journal of Computers"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2017.10.011"},{"key":"e_1_2_1_29_1","volume-title":"The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Liu Jun"},{"key":"e_1_2_1_30_1","doi-asserted-by":"crossref","unstructured":"Xiang Long Chuang Gan Gerard de Melo Xiao Liu Yandong Li Fu Li and Shilei Wen. 2018. Multimodal keyless attention fusion for video classification. In AAAI.  Xiang Long Chuang Gan Gerard de Melo Xiao Liu Yandong Li Fu Li and Shilei Wen. 2018. Multimodal keyless attention fusion for video classification. In AAAI.","DOI":"10.1609\/aaai.v32i1.12319"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00817"},{"key":"e_1_2_1_32_1","volume-title":"AAAI Conference on Artificial Intelligence. 7218--7225","author":"Lu Pan","year":"2018"},{"key":"e_1_2_1_33_1","volume-title":"The IEEE International Conference on Computer Vision (ICCV). 5773--5782","author":"Mahmud Tahmida"},{"key":"e_1_2_1_34_1","volume-title":"Advances in Neural Information Processing Systems 27","author":"Mnih Volodymyr"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2844101"},{"key":"e_1_2_1_36_1","volume-title":"Saurajit Mukherjee, J. K. Aggarwal, Hyungtae Lee, Larry Davis, et\u00a0al.","author":"Oh Sangmin","year":"2011"},{"key":"e_1_2_1_37_1","volume-title":"The IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3138--3147","author":"Paul Sujoy"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.94"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2018.2808685"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.508"},{"key":"e_1_2_1_41_1","unstructured":"Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems 28. 91--99.  Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems 28. 91--99."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/78.650093"},{"key":"e_1_2_1_43_1","volume-title":"Action recognition using visual attention. ICLR","author":"Sharma Shikhar","year":"2016"},{"key":"e_1_2_1_44_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In Advances in Neural Information Processing Systems. 568--576.  Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In Advances in Neural Information Processing Systems. 568--576."},{"key":"e_1_2_1_45_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014"},{"key":"e_1_2_1_46_1","volume-title":"AAAI Conference on Artificial Intelligence. 4929--4936","author":"Tan Zhixing","year":"2018"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.441"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.387"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00751"},{"key":"e_1_2_1_51_1","doi-asserted-by":"crossref","unstructured":"Limin Wang Yu Qiao and Xiaoou Tang. 2015. Action recognition with trajectory-pooled deep-convolutional descriptors. In Computer Vision and Pattern Recognition (CVPR). 4305--4314.  Limin Wang Yu Qiao and Xiaoou Tang. 2015. Action recognition with trajectory-pooled deep-convolutional descriptors. In Computer Vision and Pattern Recognition (CVPR). 4305--4314.","DOI":"10.1109\/CVPR.2015.7299059"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"e_1_2_1_53_1","doi-asserted-by":"crossref","unstructured":"Xiaoyang Wang and Qiang Ji. 2014. A hierarchical context model for event recognition in surveillance video. In Computer Vision and Pattern Recognition (CVPR). 2561--2568.  Xiaoyang Wang and Qiang Ji. 2014. A hierarchical context model for event recognition in surveillance video. In Computer Vision and Pattern Recognition (CVPR). 2561--2568.","DOI":"10.1109\/CVPR.2014.328"},{"key":"e_1_2_1_54_1","doi-asserted-by":"crossref","unstructured":"Xiaoyang Wang and Qiang Ji. 2015. Video event recognition with deep hierarchical context model. In Computer Vision and Pattern Recognition (CVPR). 4418--4427.  Xiaoyang Wang and Qiang Ji. 2015. Video event recognition with deep hierarchical context model. In Computer Vision and Pattern Recognition (CVPR). 4418--4427.","DOI":"10.1109\/CVPR.2015.7299071"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2616308"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46478-7_31"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2733373.2806222"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2016.2589838"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2879749"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46478-7_28"},{"key":"e_1_2_1_61_1","volume-title":"The IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 1798--1807","author":"Xu Zhongwen"},{"key":"e_1_2_1_62_1","volume-title":"The IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS). 1--6.","author":"Yikang Li","year":"2018"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2907060"},{"key":"e_1_2_1_64_1","first-page":"2902268","article-title":"Task-aware attention model for clothing attribute prediction","volume":"2019","author":"Zhang Sanyi","year":"2019","journal-title":"DOI:https:\/\/doi.org\/10.1109\/TCSVT."},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.5555\/3172077.3172393"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2014.2369044"},{"key":"e_1_2_1_67_1","volume-title":"Roy-Chowdhury","author":"Zhu Yingying","year":"2013"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46604-0_47"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3378026","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3378026","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:41:00Z","timestamp":1750200060000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3378026"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,4,30]]},"references-count":68,"journal-issue":{"issue":"2s","published-print":{"date-parts":[[2020,4,30]]}},"alternative-id":["10.1145\/3378026"],"URL":"https:\/\/doi.org\/10.1145\/3378026","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,4,30]]},"assertion":[{"value":"2019-05-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-01-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-06-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}