{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T21:26:56Z","timestamp":1781990816737,"version":"3.54.5"},"reference-count":85,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,7,12]],"date-time":"2023-07-12T00:00:00Z","timestamp":1689120000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Technology and Innovation Major Project of the Ministry of Science and Technology of China","award":["2020AAA0108400 and 2020AAA0108402"],"award-info":[{"award-number":["2020AAA0108400 and 2020AAA0108402"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61976069, U21B2038, 62236008, 62022083, 61836002, and 61931008"],"award-info":[{"award-number":["61976069, U21B2038, 62236008, 62022083, 61836002, and 61931008"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100005090","name":"Beijing Nova Program","doi-asserted-by":"crossref","award":["Z201100006820023"],"award-info":[{"award-number":["Z201100006820023"]}],"id":[{"id":"10.13039\/501100005090","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,11,30]]},"abstract":"<jats:p>Recently, with the vigorous development of deep learning and multimedia technology, intelligent urban computing has received more and more extensive attention from academia and industry. Unfortunately, most of the related technologies are black-box paradigms that lack interpretability. Among them, video event recognition is a basic technology. Event contains multiple concepts and their rich interactions, which can assist us to construct explainable event recognition methods. However, the crucial concepts needed to recognize events have various temporal existing patterns, and the relationship between events and the temporal characteristics of concepts has not been fully exploited. This brings great challenges for concept-based event categorization. To address the above issues, we introduce the temporal concept receptive field, which is the length of the temporal window size required to capture key concepts for concept-based event recognition methods. Accordingly, we introduce the temporal dynamic convolution\u00a0(TDC) to model the temporal concept receptive field dynamically according to different events. Its core idea is to combine the results of multiple convolution layers with the learned coefficients from two complementary perspectives. These convolution layers contain a variety of kernel sizes, which can provide temporal concept receptive fields of different lengths. Similarly, we also propose the cross-domain temporal dynamic convolution\u00a0(CrTDC) with the help of the rich relationship between different concepts. Different coefficients can help us to capture suitable temporal concept receptive field sizes and highlight crucial concepts to obtain accurate and complete concept representations for event analysis. Based on the TDC and CrTDC, we introduce the temporal dynamic concept modeling network\u00a0(TDCMN) for explainable video event recognition. We evaluate TDCMN on large-scale and challenging datasets FCVID, ActivityNet, and CCV. Experimental results show that TDCMN significantly improves the event recognition performance of concept-based methods, and the explainability of our method inspires us to construct more explainable models from the perspective of the temporal concept receptive field.<\/jats:p>","DOI":"10.1145\/3568312","type":"journal-article","created":{"date-parts":[[2022,10,25]],"date-time":"2022-10-25T13:27:02Z","timestamp":1666704422000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Temporal Dynamic Concept Modeling Network for Explainable Video Event Recognition"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0042-7074","authenticated-orcid":false,"given":"Weigang","family":"Zhang","sequence":"first","affiliation":[{"name":"Harbin Institute of Technology, Weihai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9196-9818","authenticated-orcid":false,"given":"Zhaobo","family":"Qi","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Weihai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5931-0527","authenticated-orcid":false,"given":"Shuhui","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Computing Technology, Chinese Academy of Sciences and Peng Cheng Laboratory, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5117-8867","authenticated-orcid":false,"given":"Chi","family":"Su","sequence":"additional","affiliation":[{"name":"SmartMore, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4038-753X","authenticated-orcid":false,"given":"Li","family":"Su","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7542-296X","authenticated-orcid":false,"given":"Qingming","family":"Huang","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences, Institute of Computing Technology, Chinese Academy of Sciences and Peng Cheng Laboratory, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,7,12]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3306240"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3199668"},{"key":"e_1_3_1_4_2","first-page":"2235","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Bhattacharya Subhabrata","year":"2014","unstructured":"Subhabrata Bhattacharya, Mahdi M. Kalayeh, Rahul Sukthankar, and Mubarak Shah. 2014. Recognition of complex events: Exploiting temporal dynamics between underlying concepts. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2235\u20132242."},{"key":"e_1_3_1_5_2","first-page":"6214","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems","author":"Burkov Egor","year":"2018","unstructured":"Egor Burkov and Victor S. Lempitsky. 2018. Deep neural networks with box convolutions. In Proceedings of the Conference on Advances in Neural Information Processing Systems. 6214\u20136224."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","first-page":"581","DOI":"10.1145\/2733373.2806218","volume-title":"Proceedings of the 23rd ACM International Conference on Multimedia","author":"Chang Xiaojun","year":"2015","unstructured":"Xiaojun Chang, Yao-Liang Yu, Yi Yang, and Alexander G. Hauptmann. 2015. Searching persuasively: Joint event detection and evidence recounting with limited supervision. In Proceedings of the 23rd ACM International Conference on Multimedia. ACM, 581\u2013590."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.208"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2608901"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","unstructured":"Jiawei Chen Yin Cui Guangnan Ye Dong Liu and Shih-Fu Chang. 2014. Event-driven semantic concept discovery by exploiting weakly tagged internet images. In ACM International Conference on Multimedia Retrieval . 1.","DOI":"10.1145\/2578726.2578729"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","unstructured":"Weidong Chen Dexiang Hong Yuankai Qi Zhenjun Han Shuhui Wang Laiyun Qing Qingming Huang and Guorong Li. 2022. Multi-attention network for compressed video referring object segmentation. In Proceedings of the 30th ACM International Conference on Multimedia . 4416\u20134425.","DOI":"10.1145\/3503161.3547761"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3514250"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475534"},{"key":"e_1_3_1_15_2","first-page":"11793","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Chen Xiaodong","year":"2021","unstructured":"Xiaodong Chen, Xinchen Liu, Wu Liu, Xiaoping Zhang, Yongdong Zhang, and Tao Mei. 2021. Explainable person re-identification with attribute-guided metric distillation. In Proceedings of the IEEE International Conference on Computer Vision. 11793\u201311802."},{"key":"e_1_3_1_16_2","doi-asserted-by":"crossref","unstructured":"Yinpeng Chen Xiyang Dai Mengchen Liu Dongdong Chen Lu Yuan and Zicheng Liu. 2020. Dynamic convolution: Attention over convolution kernels. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 11027\u201311036.","DOI":"10.1109\/CVPR42600.2020.01104"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.86"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00630"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.213"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_21_2","unstructured":"Dang Ha The Hien. 2017. A guide to receptive field arithmetic for convolutional neural networks. https:\/\/syncedreview.com\/2017\/05\/11\/a-guide-to-receptive-field-arithmetic-for-convolutional-neural-networks\/."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3409332"},{"key":"e_1_3_1_23_2","first-page":"430","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Izadinia Hamid","year":"2012","unstructured":"Hamid Izadinia and Mubarak Shah. 2012. Recognizing complex events using large margin joint low-level event model. In Proceedings of the European Conference on Computer Vision. Springer, 430\u2013444."},{"key":"e_1_3_1_24_2","unstructured":"Xu Jia Bert De Brabandere Tinne Tuytelaars and Luc Van Gool. 2016. Dynamic filter networks. In Advances in Neural Information Processing Systems . 667\u2013675."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2012.2188038"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2670560"},{"key":"e_1_3_1_27_2","first-page":"29","volume-title":"Proceedings of the 1st ACM International Conference on Multimedia Retrieval","author":"Jiang Yu-Gang","year":"2011","unstructured":"Yu-Gang Jiang, Guangnan Ye, Shih-Fu Chang, Daniel Ellis, and Alexander C. Loui. 2011. Consumer video understanding: A benchmark database and an evaluation of human and machine performance. In Proceedings of the 1st ACM International Conference on Multimedia Retrieval. ACM, 29."},{"key":"e_1_3_1_28_2","first-page":"386","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201918)","author":"Kang Sunghun","year":"2018","unstructured":"Sunghun Kang, Junyeong Kim, Hyunsoo Choi, Sungjin Kim, and Chang D. Yoo. 2018. Pivot correlational neural network for multimodal video categorization. In Proceedings of the European Conference on Computer Vision (ECCV\u201918). 386\u2013401."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.223"},{"key":"e_1_3_1_30_2","article-title":"The kinetics human action video dataset","author":"Kay Will","year":"2017","unstructured":"Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev et\u00a0al. 2017. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950 (2017).","journal-title":"arXiv preprint arXiv:1705.06950"},{"key":"e_1_3_1_31_2","first-page":"953","article-title":"Lp-norm multiple kernel learning","volume":"12","author":"Kloft Marius","year":"2011","unstructured":"Marius Kloft, Ulf Brefeld, S\u00f6ren Sonnenburg, and Alexander Zien. 2011. Lp-norm multiple kernel learning. J. Mach. Learn. Res. 12, Mar. (2011), 953\u2013997.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_1_32_2","first-page":"6232","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Korbar Bruno","year":"2019","unstructured":"Bruno Korbar, Du Tran, and Lorenzo Torresani. 2019. SCSampler: Sampling salient clips from video for efficient action recognition. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 6232\u20136242."},{"key":"e_1_3_1_33_2","first-page":"675","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Lai Kuan-Ting","year":"2014","unstructured":"Kuan-Ting Lai, Dong Liu, Ming-Syan Chen, and Shih-Fu Chang. 2014. Recognizing complex events in videos by learning key static-dynamic evidences. In Proceedings of the European Conference on Computer Vision. Springer, 675\u2013688."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2670782"},{"key":"e_1_3_1_35_2","first-page":"510","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Xiang","year":"2019","unstructured":"Xiang Li, Wenhai Wang, Xiaolin Hu, and Jian Yang. 2019. Selective kernel networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 510\u2013519."},{"key":"e_1_3_1_36_2","first-page":"6053","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Li Yanghao","year":"2019","unstructured":"Yanghao Li, Yuntao Chen, Naiyan Wang, and Zhao-Xiang Zhang. 2019. Scale-aware trident networks for object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 6053\u20136062."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3378026"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3132068"},{"issue":"9","key":"e_1_3_1_39_2","first-page":"2070","article-title":"Deep collaborative embedding for social image understanding","volume":"41","author":"Li Zechao","year":"2018","unstructured":"Zechao Li, Jinhui Tang, and Tao Mei. 2018. Deep collaborative embedding for social image understanding. IEEE Trans. Pattern Anal. Mach. Intell. 41, 9 (2018), 2070\u20132083.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"3","key":"e_1_3_1_40_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2822907","article-title":"Multimedia news summarization in search","volume":"7","author":"Li Zechao","year":"2016","unstructured":"Zechao Li, Jinhui Tang, Xueming Wang, Jing Liu, and Hanqing Lu. 2016. Multimedia news summarization in search. ACM Trans. Intell. Syst. Technol. 7, 3 (2016), 1\u201320.","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"e_1_3_1_41_2","series-title":"Proceedings of the International Conference on Machine Learning","first-page":"6172","volume":"119","author":"Lioutas Vasileios","year":"2020","unstructured":"Vasileios Lioutas and Yuhong Guo. 2020. Time-aware large kernel convolutions. In Proceedings of the International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol. 119). 6172\u20136183."},{"key":"e_1_3_1_42_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Liu Kun","year":"2018","unstructured":"Kun Liu, Wu Liu, Chuang Gan, Mingkui Tan, and Huadong Ma. 2018. T-C3D: Temporal convolutional 3D network for real-time action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11280-018-0642-6"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2020.2984569"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3524497"},{"key":"e_1_3_1_46_2","doi-asserted-by":"crossref","unstructured":"Ze Liu Jia Ning Yue Cao Yixuan Wei Zheng Zhang Stephen Lin and Han Hu. 2022. Video Swin Transformer. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 3192\u20133201.","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"e_1_3_1_47_2","first-page":"86","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Meng Yue","year":"2020","unstructured":"Yue Meng, Chung-Ching Lin, Rameswar Panda, Prasanna Sattigeri, Leonid Karlinsky, Aude Oliva, Kate Saenko, and Rogerio Feris. 2020. AR-Net: Adaptive frame resolution for efficient action recognition. In Proceedings of the European Conference on Computer Vision. Springer, 86\u2013104."},{"key":"e_1_3_1_48_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Meng Yue","year":"2021","unstructured":"Yue Meng, Rameswar Panda, Chung-Ching Lin, Prasanna Sattigeri, Leonid Karlinsky, Kate Saenko, Aude Oliva, and Rog\u00e9rio Feris. 2021. AdaFuse: Adaptive temporal fusion network for efficient action recognition. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_49_2","volume-title":"Proceedings of the British Machine Vision Conference","author":"Nagel Markus","year":"2015","unstructured":"Markus Nagel, Thomas Mensink, and Cees G. M. Snoek. 2015. Event Fisher vectors: Robust encoding visual diversity of visual streams. In Proceedings of the British Machine Vision Conference. 178.1\u2013178.12."},{"key":"e_1_3_1_50_2","first-page":"3832","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence","author":"Pan Yingwei","year":"2016","unstructured":"Yingwei Pan, Yehao Li, Ting Yao, Tao Mei, Houqiang Li, and Yong Rui. 2016. Learning deep intrinsic video representation by exploring temporal coherence and graph structure. In Proceedings of the International Joint Conference on Artificial Intelligence. Citeseer, 3832\u20133838."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.94"},{"key":"e_1_3_1_52_2","first-page":"441","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Peng Zhimao","year":"2019","unstructured":"Zhimao Peng, Zechao Li, Junge Zhang, Yan Li, Guo-Jun Qi, and Jinhui Tang. 2019. Few-shot image recognition with knowledge transfer. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 441\u2013449."},{"key":"e_1_3_1_53_2","unstructured":"Zhaobo Qi Shuhui Wang Chi Su Li Su Qingming Huang and Qi Tian. 2020. Towards more explainability: Concept knowledge mining network for event recognition. In The 28th ACM International Conference on Multimedia . 3857\u20133865."},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3059923"},{"key":"e_1_3_1_55_2","doi-asserted-by":"crossref","unstructured":"Zhaobo Qi Shuhui Wang Chi Su Li Su Weigang Zhang and Qingming Huang. 2020. Modeling temporal concept receptive field dynamically for untrimmed video analysis. In The 28th ACM International Conference on Multimedia . 3798\u20133806.","DOI":"10.1145\/3394171.3413618"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.590"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_1_58_2","first-page":"568","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In Proceedings of the International Conference on Advances in Neural Information Processing Systems. 568\u2013576."},{"key":"e_1_3_1_59_2","first-page":"4561","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Singh Bharat","year":"2015","unstructured":"Bharat Singh, Xintong Han, Zhe Wu, Vlad I. Morariu, and Larry S. Davis. 2015. Selecting relevant web trained concepts for automated event retrieval. In Proceedings of the IEEE International Conference on Computer Vision. 4561\u20134569."},{"key":"e_1_3_1_60_2","doi-asserted-by":"crossref","unstructured":"John R. Smith Milind R. Naphade and Apostol Natsev. 2003. Multimedia semantic indexing using model vectors. In Proceedings of the IEEE International Conference on Multimedia and Expo . 445\u2013448.","DOI":"10.1109\/ICME.2003.1221649"},{"issue":"7","key":"e_1_3_1_61_2","doi-asserted-by":"crossref","first-page":"1769","DOI":"10.1109\/TMM.2019.2959426","article-title":"Spatio-temporal VLAD encoding of visual events using temporal ordering of the mid-level deep semantics","volume":"22","author":"Soltanian Mohammad","year":"2019","unstructured":"Mohammad Soltanian, Sajjad Amini, and Shahrokh Ghaemmaghami. 2019. Spatio-temporal VLAD encoding of visual events using temporal ordering of the mid-level deep semantics. IEEE Trans. Multim. 22, 7 (2019), 1769\u20131784.","journal-title":"IEEE Trans. Multim."},{"issue":"1","key":"e_1_3_1_62_2","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1109\/TMM.2018.2844101","article-title":"Hierarchical concept score postprocessing and concept-wise normalization in CNN-based video event recognition","volume":"21","author":"Soltanian Mohammad","year":"2018","unstructured":"Mohammad Soltanian and Shahrokh Ghaemmaghami. 2018. Hierarchical concept score postprocessing and concept-wise normalization in CNN-based video event recognition. IEEE Trans. Multim. 21, 1 (2018), 157\u2013172.","journal-title":"IEEE Trans. Multim."},{"key":"e_1_3_1_63_2","first-page":"2222","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems","author":"Srivastava Nitish","year":"2012","unstructured":"Nitish Srivastava and Ruslan R. Salakhutdinov. 2012. Multimodal learning with deep Boltzmann machines. In Proceedings of the International Conference on Advances in Neural Information Processing Systems. 2222\u20132230."},{"key":"e_1_3_1_64_2","doi-asserted-by":"crossref","unstructured":"Christian Szegedy Sergey Ioffe Vincent Vanhoucke and Alexander A. Alemi. 2017. Inception-v4 Inception-ResNet and the Impact of residual connections on learning. In Proceedings of the ThirtyFirst AAAI Conference on Artificial Intelligence . 4278\u20134284.","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"e_1_3_1_65_2","doi-asserted-by":"crossref","unstructured":"Christian Szegedy Vincent Vanhoucke Sergey Ioffe Jonathon Shlens and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In IEEE Conference on Computer Vision and Pattern Recognition . 2818\u20132826.","DOI":"10.1109\/CVPR.2016.308"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"e_1_3_1_68_2","doi-asserted-by":"crossref","unstructured":"Yulin Wang Zhaoxi Chen Haojun Jiang Shiji Song Yizeng Han and Gao Huang. 2021. Adaptive focus for efficient video recognition. In ICCV . 16229\u201316238.","DOI":"10.1109\/ICCV48922.2021.01594"},{"key":"e_1_3_1_69_2","first-page":"6222","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Wu Wenhao","year":"2019","unstructured":"Wenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen, and Shilei Wen. 2019. Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition. In Proceedings of the IEEE International Conference on Computer Vision. 6222\u20136231."},{"key":"e_1_3_1_70_2","first-page":"3112","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wu Zuxuan","year":"2016","unstructured":"Zuxuan Wu, Yanwei Fu, Yu-Gang Jiang, and Leonid Sigal. 2016. Harnessing object and scene semantics for large-scale video understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3112\u20133121."},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654931"},{"key":"e_1_3_1_72_2","first-page":"7778","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems","author":"Wu Zuxuan","year":"2019","unstructured":"Zuxuan Wu, Caiming Xiong, Yu-Gang Jiang, and Larry S. Davis. 2019. LiteEval: A coarse-to-fine framework for resource efficient video recognition. In Proceedings of the Conference on Advances in Neural Information Processing Systems. 7778\u20137787."},{"key":"e_1_3_1_73_2","first-page":"1278","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wu Zuxuan","year":"2019","unstructured":"Zuxuan Wu, Caiming Xiong, Chih-Yao Ma, Richard Socher, and Larry S. Davis. 2019. AdaFrame: Adaptive frame selection for fast video recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1278\u20131287."},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2879749"},{"issue":"6","key":"e_1_3_1_75_2","first-page":"1425","article-title":"Discovering latent discriminative patterns for multi-mode event representation","volume":"21","author":"Xie Wenlong","year":"2018","unstructured":"Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Tingting Han, Sicheng Zhao, and Tat-Seng Chua. 2018. Discovering latent discriminative patterns for multi-mode event representation. IEEE Trans. Multim. 21, 6 (2018), 1425\u20131436.","journal-title":"IEEE Trans. Multim."},{"key":"e_1_3_1_76_2","first-page":"97","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Xu Zhongwen","year":"2014","unstructured":"Zhongwen Xu, Ivor W. Tsang, Yi Yang, Zhigang Ma, and Alexander G. Hauptmann. 2014. Event detection using multi-level relevance labels and multiple features. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 97\u2013104."},{"key":"e_1_3_1_77_2","volume-title":"Proceedings of the 29th AAAI Conference on Artificial Intelligence","author":"Yan Yan","year":"2015","unstructured":"Yan Yan, Yi Yang, Haoquan Shen, Deyu Meng, Gaowen Liu, Alex Hauptmann, and Nicu Sebe. 2015. Complex event detection via event oriented dictionary learning. In Proceedings of the 29th AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_1_78_2","first-page":"1305","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems","author":"Yang Brandon","year":"2019","unstructured":"Brandon Yang, Gabriel Bender, Quoc V. Le, and Jiquan Ngiam. 2019. CondConv: Conditionally parameterized convolutions for efficient inference. In Proceedings of the International Conference on Advances in Neural Information Processing Systems. 1305\u20131316."},{"key":"e_1_3_1_79_2","doi-asserted-by":"publisher","DOI":"10.1145\/2962719"},{"key":"e_1_3_1_80_2","first-page":"471","volume-title":"Proceedings of the 23rd ACM International Conference on Multimedia","author":"Ye Guangnan","year":"2015","unstructured":"Guangnan Ye, Yitong Li, Hongliang Xu, Dong Liu, and Shih-Fu Chang. 2015. EventNet: A large scale structured concept library for complex event detection in video. In Proceedings of the 23rd ACM International Conference on Multimedia. ACM, 471\u2013480."},{"key":"e_1_3_1_81_2","unstructured":"Fisher Yu and Vladlen Koltun. 2016. Multi-scale context aggregation by dilated convolutions. In International Conference on Learning Representations ."},{"key":"e_1_3_1_82_2","article-title":"Exploiting mid-level semantics for large-scale complex video classification","author":"Zhang Ji","year":"2019","unstructured":"Ji Zhang, Kuizhi Mei, Yu Zheng, and Jianping Fan. 2019. Exploiting mid-level semantics for large-scale complex video classification. IEEE Trans. Multim. (2019).","journal-title":"IEEE Trans. Multim."},{"key":"e_1_3_1_83_2","unstructured":"Linguang Zhang Maciej Halber and Szymon Rusinkiewicz. 2019. Accelerating large-kernel convolution using summed-area tables. arXiv preprint arXiv:1906.11367 (2019)."},{"issue":"1","key":"e_1_3_1_84_2","first-page":"6:1\u20136:22","article-title":"Visual content recognition by exploiting semantic feature map with attention and multi-task learning","volume":"15","author":"Zhao Rui-Wei","year":"2019","unstructured":"Rui-Wei Zhao, Qi Zhang, Zuxuan Wu, Jianguo Li, and Yu-Gang Jiang. 2019. Visual content recognition by exploiting semantic feature map with attention and multi-task learning. ACM Trans. Multim. Comput., Commun. Applic. 15, 1s (2019), 6:1\u20136:22.","journal-title":"ACM Trans. Multim. Comput., Commun. Applic."},{"key":"e_1_3_1_85_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2017.2762344"},{"key":"e_1_3_1_86_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2723009"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3568312","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3568312","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:51:33Z","timestamp":1750182693000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3568312"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,12]]},"references-count":85,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,11,30]]}},"alternative-id":["10.1145\/3568312"],"URL":"https:\/\/doi.org\/10.1145\/3568312","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,12]]},"assertion":[{"value":"2022-02-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-09-25","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}