{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T15:43:40Z","timestamp":1784821420827,"version":"3.55.0"},"reference-count":77,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,2,25]],"date-time":"2023-02-25T00:00:00Z","timestamp":1677283200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"PRIN project PREVUE","award":["2017N2RK7K"],"award-info":[{"award-number":["2017N2RK7K"]}]},{"name":"EU H2020 project AI4Media","award":["951911"],"award-info":[{"award-number":["951911"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,8,31]]},"abstract":"<jats:p>Video processing and analysis have become an urgent task, as a huge amount of videos (e.g., YouTube, Hulu) are uploaded online every day. The extraction of representative key frames from videos is important in video processing and analysis since it greatly reduces computing resources and time. Although great progress has been made recently, large-scale video classification remains an open problem, as the existing methods have not well balanced the performance and efficiency simultaneously. To tackle this problem, this work presents an unsupervised method to retrieve the key frames, which combines the convolutional neural network and temporal segment density peaks clustering. The proposed temporal segment density peaks clustering is a generic and powerful framework, and it has two advantages compared with previous works. One is that it can calculate the number of key frames automatically. The other is that it can preserve the temporal information of the video. Thus, it improves the efficiency of video classification. Furthermore, a long short-term memory network is added on the top of the convolutional neural network to further elevate the performance of classification. Moreover, a weight fusion strategy of different input networks is presented to boost performance. By optimizing both video classification and key frame extraction simultaneously, we achieve better classification performance and higher efficiency. We evaluate our method on two popular datasets (i.e., HMDB51 and UCF101), and the experimental results consistently demonstrate that our strategy achieves competitive performance and efficiency compared with the state-of-the-art approaches.<\/jats:p>","DOI":"10.1145\/3571735","type":"journal-article","created":{"date-parts":[[2022,12,12]],"date-time":"2022-12-12T14:47:06Z","timestamp":1670856426000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":48,"title":["Deep Unsupervised Key Frame Extraction for Efficient Video Classification"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2077-1246","authenticated-orcid":false,"given":"Hao","family":"Tang","sequence":"first","affiliation":[{"name":"ETH Zurich, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0653-8373","authenticated-orcid":false,"given":"Lei","family":"Ding","sequence":"additional","affiliation":[{"name":"University of Trento, Trento, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9347-5395","authenticated-orcid":false,"given":"Songsong","family":"Wu","sequence":"additional","affiliation":[{"name":"Guangdong University of Petrochemical Technology, Maoming, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9790-1504","authenticated-orcid":false,"given":"Bin","family":"Ren","sequence":"additional","affiliation":[{"name":"University of Trento, Trento, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6597-7248","authenticated-orcid":false,"given":"Nicu","family":"Sebe","sequence":"additional","affiliation":[{"name":"University of Trento, Trento, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0663-5659","authenticated-orcid":false,"given":"Paolo","family":"Rota","sequence":"additional","affiliation":[{"name":"University of Trento, Trento, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,2,25]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"Proceedings of CVPR","author":"Bilen Hakan","year":"2016","unstructured":"Hakan Bilen, Basura Fernando, Efstratios Gavves, Andrea Vedaldi, and Stephen Gould. 2016. Dynamic image networks for action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_3_2","volume-title":"Proceedings of CVPR","author":"Cai Zhuowei","year":"2014","unstructured":"Zhuowei Cai, Limin Wang, Xiaojiang Peng, and Yu Qiao. 2014. Multi-view super vector for action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_4_2","volume-title":"Proceedings of CVPR","author":"Carreira Joao","year":"2017","unstructured":"Joao Carreira and Andrew Zisserman. 2017. Quo Vadis, action recognition? A new model and the kinetics dataset. In Proceedings of CVPR."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2005.856896"},{"key":"e_1_3_2_6_2","volume-title":"Proceedings of CVPR","author":"Choutas Vasileios","year":"2018","unstructured":"Vasileios Choutas, Philippe Weinzaepfel, J\u00e9r\u00f4me Revaud, and Cordelia Schmid. 2018. Potion: Pose motion representation for action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2011.2166951"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2010.08.004"},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of ECCV","author":"Souza C\u00e9sar Roberto de","year":"2016","unstructured":"C\u00e9sar Roberto de Souza, Adrien Gaidon, Eleonora Vig, and Antonio Manuel L\u00f3pez. 2016. Sympathy for the details: Dense trajectories and hybrid classification architectures for action recognition. In Proceedings of ECCV."},{"key":"e_1_3_2_10_2","volume-title":"Proceedings of CVPR","author":"Donahue Jeffrey","year":"2015","unstructured":"Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2015. Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of CVPR."},{"key":"e_1_3_2_11_2","volume-title":"Proceedings of CVPR","author":"Duta Ionut Cosmin","year":"2017","unstructured":"Ionut Cosmin Duta, Bogdan Ionescu, Kiyoharu Aizawa, and Nicu Sebe. 2017. Spatio-temporal vector of locally max pooled features for action recognition in videos. In Proceedings of CVPR."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jvcir.2012.06.013"},{"key":"e_1_3_2_13_2","volume-title":"Proceedings of NIPS","author":"Feichtenhofer Christoph","year":"2016","unstructured":"Christoph Feichtenhofer, Axel Pinz, and Richard Wildes. 2016. Spatiotemporal residual networks for video action recognition. In Proceedings of NIPS."},{"key":"e_1_3_2_14_2","volume-title":"Proceedings of CVPR","author":"Feichtenhofer Christoph","year":"2017","unstructured":"Christoph Feichtenhofer, Axel Pinz, and Richard P. Wildes. 2017. Spatiotemporal multiplier networks for video action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_15_2","volume-title":"Proceedings of CVPR","author":"Feichtenhofer Christoph","year":"2018","unstructured":"Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes, and Andrew Zisserman. 2018. What have we learned from deep representations for action recognition? In Proceedings of CVPR."},{"key":"e_1_3_2_16_2","volume-title":"Proceedings of CVPR","author":"Feichtenhofer Christoph","year":"2016","unstructured":"Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman. 2016. Convolutional two-stream network fusion for video action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_17_2","volume-title":"Proceedings of CVPR","author":"Fernando Basura","year":"2015","unstructured":"Basura Fernando, Efstratios Gavves, Jose M. Oramas, Amir Ghodrati, and Tinne Tuytelaars. 2015. Modeling video evolution for action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/tpami.2013.65"},{"key":"e_1_3_2_19_2","volume-title":"Proceedings of CVPR","author":"Gao Ruohan","year":"2018","unstructured":"Ruohan Gao, Bo Xiong, and Kristen Grauman. 2018. Im2Flow: Motion hallucination from static images for action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_20_2","volume-title":"Proceedings of ICASSP","author":"Gharbi Hana","year":"2017","unstructured":"Hana Gharbi, Sahbi Bahroun, Mohamed Massaoudi, and Ezzeddine Zagrouba. 2017. Key frames extraction using graph modularity clustering for efficient video summarization. In Proceedings of ICASSP."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2012.2214871"},{"key":"e_1_3_2_22_2","volume-title":"Proceedings of CVPR","author":"Hara Kensho","year":"2018","unstructured":"Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. 2018. Can spatiotemporal 3D CNNs retrace the history of 2D CNNS and ImageNet? In Proceedings of CVPR."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.59"},{"key":"e_1_3_2_25_2","volume-title":"Proceedings of CVPR","author":"Kar Amlan","year":"2017","unstructured":"Amlan Kar, Nishant Rai, Karan Sikka, and Gaurav Sharma. 2017. AdaScan: Adaptive scan pooling in deep convolutional neural networks for human action recognition in videos. In Proceedings of CVPR."},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.223"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jvcir.2013.08.003"},{"key":"e_1_3_2_28_2","volume-title":"Proceedings of ICCV","author":"Kuehne H.","year":"2011","unstructured":"H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre. 2011. HMDB: A large video database for human motion recognition. In Proceedings of ICCV."},{"key":"e_1_3_2_29_2","volume-title":"Proceedings of ICPR","author":"Kulhare Sourabh","year":"2016","unstructured":"Sourabh Kulhare, Shagan Sah, Suhas Pillai, and Raymond Ptucha. 2016. Key frame extraction for salient activity recognition. In Proceedings of ICPR."},{"key":"e_1_3_2_30_2","volume-title":"Proceedings of CVPR","author":"Lan Zhengzhong","year":"2015","unstructured":"Zhengzhong Lan, Ming Lin, Xuanchong Li, Alex G. Hauptmann, and Bhiksha Raj. 2015. Beyond Gaussian pyramid: Multi-skip feature stacking for action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2008.4587756"},{"issue":"2","key":"e_1_3_2_32_2","doi-asserted-by":"crossref","first-page":"125","DOI":"10.1016\/j.trit.2016.10.001","article-title":"Sequential bag-of-words model for human action classification","volume":"1","author":"Liu Hong","year":"2016","unstructured":"Hong Liu, Hao Tang, Wei Xiao, ZiYi Guo, Lu Tian, and Yuan Gao. 2016. Sequential bag-of-words model for human action classification. CAAI Transactions on Intelligence Technology 1, 2 (2016), 125\u2013136.","journal-title":"CAAI Transactions on Intelligence Technology"},{"key":"e_1_3_2_33_2","volume-title":"Proceedings of ICIP","author":"Liu Hong","year":"2015","unstructured":"Hong Liu, Lu Tian, Mengyuan Liu, and Hao Tang. 2015. SDM-BSM: A fusing depth scheme for human action recognition. In Proceedings of ICIP."},{"key":"e_1_3_2_34_2","volume-title":"Proceedings of AAAI","author":"Long Xiang","year":"2018","unstructured":"Xiang Long, Chuang Gan, Gerard de Melo, Xiao Liu, Yandong Li, Fu Li, and Shilei Wen. 2018. Multimodal keyless attention fusion for video classification. In Proceedings of AAAI."},{"key":"e_1_3_2_35_2","volume-title":"Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability","year":"1967","unstructured":"James MacQueen. 1967. Some methods for classification and analysis of multivariate observations. In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, Lucien M. Le Cam and Jerzy Neyman (Eds.). University of California Press, 281\u2013297."},{"key":"e_1_3_2_36_2","volume-title":"Proceedings of ICME","author":"Mei Shaohui","year":"2014","unstructured":"Shaohui Mei, Genliang Guan, Zhiyong Wang, Mingyi He, Xian-Sheng Hua, and David Dagan Feng. 2014. L2, 0 constrained sparse dictionary selection for video summarization. In Proceedings of ICME."},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2014.08.002"},{"key":"e_1_3_2_38_2","volume-title":"Proceedings of CVPR","author":"Meng Jingjing","year":"2016","unstructured":"Jingjing Meng, Hongxing Wang, Junsong Yuan, and Yap-Peng Tan. 2016. From keyframes to key objects: Video summarization by representative object proposal selection. In Proceedings of CVPR."},{"key":"e_1_3_2_39_2","volume-title":"Proceedings of CVPR","author":"Ni Bingbing","year":"2015","unstructured":"Bingbing Ni, Pierre Moulin, Xiaokang Yang, and Shuicheng Yan. 2015. Motion part regularization: Improving action recognition via trajectory selection. In Proceedings of CVPR."},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2009.2013517"},{"key":"e_1_3_2_41_2","volume-title":"Proceedings of ICPR","author":"Panda Rameswar","year":"2014","unstructured":"Rameswar Panda, Sanjay K. Kuanar, and Ananda S. Chowdhury. 2014. Scalable video summarization using skeleton graph and random walk. In Proceedings of ICPR."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2016.03.013"},{"key":"e_1_3_2_43_2","first-page":"581","volume-title":"Proceedings of ECCV","author":"Peng Xiaojiang","year":"2014","unstructured":"Xiaojiang Peng, Changqing Zou, Yu Qiao, and Qiang Peng. 2014. Action recognition with stacked Fisher vectors. In Proceedings of ECCV. 581\u2013595."},{"key":"e_1_3_2_44_2","volume-title":"Proceedings of CVPR","author":"Souza Cesar Roberto de","year":"2017","unstructured":"Cesar Roberto de Souza, Adrien Gaidon, Yohann Cabon, and Antonio Manuel Lopez. 2017. Procedural generation of videos to train deep action recognition networks. In Proceedings of CVPR."},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.1242072"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_2_47_2","volume-title":"Proceedings of CVPR","author":"Shou Zheng","year":"2019","unstructured":"Zheng Shou, Xudong Lin, Yannis Kalantidis, Laura Sevilla-Lara, Marcus Rohrbach, Shih-Fu Chang, and Zhicheng Yan. 2019. DMC-Net: Generating discriminative motion cues for fast compressed video action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_48_2","volume-title":"Proceedings of ECCV","author":"Sigurdsson Gunnar A.","year":"2016","unstructured":"Gunnar A. Sigurdsson, Xinlei Chen, and Abhinav Gupta. 2016. Learning visual storylines with skipping recurrent neural networks. In Proceedings of ECCV."},{"key":"e_1_3_2_49_2","volume-title":"Proceedings of NIPS","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In Proceedings of NIPS."},{"key":"e_1_3_2_50_2","article-title":"UCF101: A dataset of 101 human actions classes from videos in the wild","author":"Soomro Khurram","year":"2012","unstructured":"Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012).","journal-title":"arXiv preprint arXiv:1212.0402"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00931"},{"key":"e_1_3_2_52_2","volume-title":"Proceedings of ICCV","author":"Sun Lin","year":"2015","unstructured":"Lin Sun, Kui Jia, Dit-Yan Yeung, and Bertram E. Shi. 2015. Human action recognition using factorized spatio-temporal convolutional networks. In Proceedings of ICCV."},{"key":"e_1_3_2_53_2","volume-title":"Proceedings of CVPR","author":"Sun Shuyang","year":"2018","unstructured":"Shuyang Sun, Zhanghui Kuang, Lu Sheng, Wanli Ouyang, and Wei Zhang. 2018. Optical flow guided feature: A fast and robust motion representation for video action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_54_2","volume-title":"Proceedings of ACM MM","author":"Tang Hao","year":"2015","unstructured":"Hao Tang, Hong Liu, and Wei Xiao. 2015. Gender classification using pyramid segmentation for unconstrained back-facing video sequences. In Proceedings of ACM MM."},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2018.11.038"},{"key":"e_1_3_2_56_2","volume-title":"Proceedings of ICCV","author":"Tran Du","year":"2015","unstructured":"Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015. Learning spatiotemporal features with 3D convolutional networks. In Proceedings of ICCV."},{"key":"e_1_3_2_57_2","article-title":"Convnet architecture search for spatiotemporal feature learning","author":"Tran Du","year":"2017","unstructured":"Du Tran, Jamie Ray, Zheng Shou, Shih-Fu Chang, and Manohar Paluri. 2017. Convnet architecture search for spatiotemporal feature learning. arXiv preprint arXiv:1708.05038 (2017).","journal-title":"arXiv preprint arXiv:1708.05038"},{"key":"e_1_3_2_58_2","volume-title":"Proceedings of CVPR","author":"Tran Du","year":"2018","unstructured":"Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. 2018. A closer look at spatiotemporal convolutions for action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2712608"},{"issue":"7","key":"e_1_3_2_60_2","doi-asserted-by":"crossref","first-page":"770","DOI":"10.1016\/j.patrec.2012.12.009","article-title":"Spatio-temporal feature-based keyframe detection from video shots using spectral clustering","volume":"34","author":"V\u00e1zquez-Martn Ricardo","year":"2013","unstructured":"Ricardo V\u00e1zquez-Martn and Antonio Bandera. 2013. Spatio-temporal feature-based keyframe detection from video shots using spectral clustering. Pattern Recognition Letters 34, 7 (2013), 770\u2013779.","journal-title":"Pattern Recognition Letters"},{"key":"e_1_3_2_61_2","volume-title":"Proceedings of NIPS","author":"Vondrick Carl","year":"2016","unstructured":"Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. 2016. Generating videos with scene dynamics. In Proceedings of NIPS."},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2016.10.014"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0846-5"},{"key":"e_1_3_2_64_2","volume-title":"Proceedings of ICCV","author":"Wang Heng","year":"2013","unstructured":"Heng Wang and Cordelia Schmid. 2013. Action recognition with improved trajectories. In Proceedings of ICCV."},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2013.2295753"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299059"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0859-0"},{"key":"e_1_3_2_68_2","volume-title":"Proceedings of ECCV","author":"Wang Limin","year":"2016","unstructured":"Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2016. Temporal segment networks: Towards good practices for deep action recognition. In Proceedings of ECCV."},{"key":"e_1_3_2_69_2","volume-title":"Proceedings of CVPR","author":"Wang Xiaolong","year":"2016","unstructured":"Xiaolong Wang, Ali Farhadi, and Abhinav Gupta. 2016. Actions transformations. In Proceedings of CVPR."},{"key":"e_1_3_2_70_2","volume-title":"Proceedings of CVPR","author":"Wang Yunbo","year":"2017","unstructured":"Yunbo Wang, Mingsheng Long, Jianmin Wang, and Philip S. Yu. 2017. Spatiotemporal pyramid network for video action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_71_2","volume-title":"Proceedings of CVPR","author":"Wang Yali","year":"2018","unstructured":"Yali Wang, Lei Zhou, and Yu Qiao. 2018. Temporal hallucinating for action recognition with few still images. In Proceedings of CVPR."},{"key":"e_1_3_2_72_2","volume-title":"Proceedings of CVPR","author":"Wu Chao-Yuan","year":"2018","unstructured":"Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu, R. Manmatha, Alexander J. Smola, and Philipp Kr\u00e4henb\u00fchl. 2018. Compressed video action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_73_2","volume-title":"Proceedings of CVPR","author":"Yang Jianwei","year":"2016","unstructured":"Jianwei Yang, Devi Parikh, and Dhruv Batra. 2016. Joint unsupervised learning of deep representations and image clusters. In Proceedings of CVPR."},{"key":"e_1_3_2_74_2","volume-title":"Proceedings of CVPR","author":"Ng Joe Yue-Hei","year":"2015","unstructured":"Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. 2015. Beyond short snippets: Deep networks for video classification. In Proceedings of CVPR."},{"key":"e_1_3_2_75_2","volume-title":"Proceedings of CVPR","author":"Zhou Yizhou","year":"2018","unstructured":"Yizhou Zhou, Xiaoyan Sun, Zheng-Jun Zha, and Wenjun Zeng. 2018. MiCT: Mixed 3D\/2D convolutional tube for human action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_76_2","volume-title":"Proceedings of CVPR","author":"Zhu Wangjiang","year":"2016","unstructured":"Wangjiang Zhu, Jie Hu, Gang Sun, Xudong Cao, and Yu Qiao. 2016. A key volume mining deep framework for action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_77_2","volume-title":"Proceedings of CVPR","author":"Zhu Yi","year":"2018","unstructured":"Yi Zhu, Yang Long, Yu Guan, Shawn Newsam, and Ling Shao. 2018. Towards universal representation for unseen action recognition. In Proceedings of CVPR."},{"key":"e_1_3_2_78_2","volume-title":"Proceedings of ICIP","author":"Zhuang Yueting","year":"1998","unstructured":"Yueting Zhuang, Yong Rui, Thomas S. Huang, and Sharad Mehrotra. 1998. Adaptive key frame extraction using unsupervised clustering. In Proceedings of ICIP."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3571735","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3571735","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:33Z","timestamp":1750182573000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3571735"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,25]]},"references-count":77,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,8,31]]}},"alternative-id":["10.1145\/3571735"],"URL":"https:\/\/doi.org\/10.1145\/3571735","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,25]]},"assertion":[{"value":"2022-04-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-11-06","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}