{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,14]],"date-time":"2025-10-14T00:49:37Z","timestamp":1760402977257,"version":"build-2065373602"},"reference-count":37,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2021,4,12]],"date-time":"2021-04-12T00:00:00Z","timestamp":1618185600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Most existing video action recognition methods mainly rely on high-level semantic information from convolutional neural networks (CNNs) but ignore the discrepancies of different information streams. However, it does not normally consider both long-distance aggregations and short-range motions. Thus, to solve these problems, we propose hierarchical excitation aggregation and disentanglement networks (Hi-EADNs), which include multiple frame excitation aggregation (MFEA) and a feature squeeze-and-excitation hierarchical disentanglement (SEHD) module. MFEA specifically uses long-short range motion modelling and calculates the feature-level temporal difference. The SEHD module utilizes these differences to optimize the weights of each spatiotemporal feature and excite motion-sensitive channels. Moreover, without introducing additional parameters, this feature information is processed with a series of squeezes and excitations, and multiple temporal aggregations with neighbourhoods can enhance the interaction of different motion frames. Extensive experimental results confirm our proposed Hi-EADN method effectiveness on the UCF101 and HMDB51 benchmark datasets, where the top-5 accuracy is 93.5% and 76.96%.<\/jats:p>","DOI":"10.3390\/sym13040662","type":"journal-article","created":{"date-parts":[[2021,4,12]],"date-time":"2021-04-12T21:47:33Z","timestamp":1618264053000},"page":"662","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Hi-EADN: Hierarchical Excitation Aggregation and Disentanglement Frameworks for Action Recognition Based on Videos"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3673-356X","authenticated-orcid":false,"given":"Zeyuan","family":"Hu","sequence":"first","affiliation":[{"name":"Department of Information Communication Engineering, Tongmyong University, Busan 48520, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eung-Joo","family":"Lee","sequence":"additional","affiliation":[{"name":"Department of Information Communication Engineering, Tongmyong University, Busan 48520, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,4,12]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Yang, C., Xu, Y., Shi, J., Dai, B., and Zhou, B. (2020, January 13\u201319). Temporal pyramid network for action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00067"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"71","DOI":"10.1109\/MSP.2019.2909074","article-title":"Radar-on-chip\/in-package in autonomous driving vehicles and intelligent transport systems: Opportunities and challenges","volume":"36","author":"Saponara","year":"2019","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"59405","DOI":"10.1109\/ACCESS.2018.2874022","article-title":"Human action recognition algorithm based on adaptive initialization of deep learning model parameters and support vector machine","volume":"6","author":"An","year":"2018","journal-title":"IEEE Access"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.patcog.2018.07.028","article-title":"Asymmetric 3d convolutional neural networks for action recognition","volume":"85","author":"Yang","year":"2019","journal-title":"Pattern Recognit."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"3938","DOI":"10.1109\/TNNLS.2017.2740318","article-title":"Deep manifold learning combined with convolutional neural networks for action recognition","volume":"29","author":"Chen","year":"2017","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"4293","DOI":"10.1007\/s00521-019-04615-w","article-title":"Spatiotemporal neural networks for action recognition based on joint loss","volume":"32","author":"Jing","year":"2020","journal-title":"Neural Comput. Appl."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2990","DOI":"10.1109\/TMM.2020.2965434","article-title":"Spatio-temporal attention networks for action recognition and detection","volume":"22","author":"Li","year":"2020","journal-title":"IEEE Trans. Multimed."},{"key":"ref_8","unstructured":"Ji, S., Xu, W., Yang, M., and Yu, K. (2010, January 21\u201324). 3D Convolutional Neural Networks for Human Action Recognition. Proceedings of the 27th International Conference on Machine Learning (ICML-10), Haifa, Israel."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"615","DOI":"10.1167\/jov.20.11.615","article-title":"Weak integration of form and motion in two-stream CNNs for action recognition","volume":"20","author":"Peng","year":"2020","journal-title":"J. Vis."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"6954174","DOI":"10.1155\/2020\/6954174","article-title":"Human Action Recognition Algorithm Based on Improved ResNet and Skeletal Keypoints in Single Image","volume":"2020","author":"Lin","year":"2020","journal-title":"Math. Probl. Eng."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"688","DOI":"10.1049\/iet-ipr.2019.0985","article-title":"An Efficient Inception V2 based Deep Convolutional Neural Network for Real-Time Hand Action Recognition","volume":"14","author":"Bose","year":"2019","journal-title":"IET Image Process."},{"key":"ref_12","first-page":"4412","article-title":"Binary Hashing CNN Features for Action Recognition","volume":"12","author":"Li","year":"2018","journal-title":"TIIS"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"11360","DOI":"10.1166\/asl.2017.10283","article-title":"Deep CNN object features for improved action recognition in low quality videos","volume":"23","author":"Rahman","year":"2017","journal-title":"Adv. Sci. Lett."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"340","DOI":"10.1007\/s11263-018-1111-5","article-title":"Second-order Temporal Pooling for Action Recognition","volume":"127","author":"Cherian","year":"2019","journal-title":"Int. J. Comput. Vis."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K.Q. (2017, January 21\u201326). Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1317","DOI":"10.1016\/j.procs.2018.05.048","article-title":"Human Detection and Tracking using HOG for Action Recognition","volume":"132","author":"Seemanthini","year":"2018","journal-title":"Procedia Comput. Sci."},{"key":"ref_17","first-page":"886","article-title":"An Action Recognition Model Based on the Bayesian Networks","volume":"513","author":"Chen","year":"2014","journal-title":"Appl. Mech. Mater."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s10489-019-01572-8","article-title":"Multi-scale affined-HOF and dimension selection for view-unconstrained action recognition","volume":"50","author":"Tran","year":"2020","journal-title":"Appl. Intell."},{"key":"ref_19","unstructured":"Wang, L., Koniusz, P., and Huynh, D.Q. (2019). Hallucinating Bag-of-Words and Fisher Vector IDT terms for CNN-based Action Recognition. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wang, L., and Zhi-Pan, W.U. (2019). A Comparative Review of Recent Kinect-based Action Recognition Algorithms. arXiv.","DOI":"10.1109\/TIP.2019.2925285"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Jagadeesh, B., and Patil, C.M. (2016, January 20\u201321). Video based action detection and recognition human using optical flow and SVM classifier. Proceedings of the 2016 IEEE International Conference on Recent Trends in Electronics, Information Communication Technology (RTEICT), Bangalore, India.","DOI":"10.1109\/RTEICT.2016.7808136"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., and Fei-Fei, L. (2014, January 23\u201328). Large-scale video classification with convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA. Available online: https:\/\/dl.acm.org\/doi\/10.1109\/CVPR.2014.223.","DOI":"10.1109\/CVPR.2014.223"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Patil, G.G., and Banyal, R.K. (2019, January 29\u201331). Techniques of Deep Learning for Image Recognition. Proceedings of the 2019 IEEE 5th International Conference for Convergence in Technology (I2CT), Pune, India.","DOI":"10.1109\/I2CT45611.2019.9033628"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"449","DOI":"10.1016\/j.patrec.2020.01.024","article-title":"BshapeNet: Object Detection and Instance Segmentation with Bounding Shape Masks","volume":"131","author":"Kang","year":"2020","journal-title":"Pattern Recognit. Lett."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"187","DOI":"10.36548\/jiip.2020.4.003","article-title":"Comparative Study: Statistical Approach and Deep Learning Method for Automatic Segmentation Methods for Lung CT Image Segmentation","volume":"2","author":"Sungheetha","year":"2020","journal-title":"J. Innov. Image Process."},{"key":"ref_26","unstructured":"Simonyan, K., and Zisserman, A. (2014). Two-Stream Convolutional Networks for Action Recognition in Videos. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"677","DOI":"10.1109\/TPAMI.2016.2599174","article-title":"Long-term Recurrent Convolutional Networks for Visual Recognition and, Description","volume":"39","author":"Donahue","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., and Van Gool, L. (2016, January 8\u201316). Temporal segment networks: Towards good practices for deep action recognition. Proceedings of the 14th European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"ref_29","unstructured":"Feichtenhofer, C., Pinz, A., and Wildes, R.P. (2021, March 10). Spatiotemporal Residual Networks for Video Action Recognition. Available online: https:\/\/papers.nips.cc\/paper\/2016\/file\/3e7e0224018ab3cf51abb96464d518cd-Paper.pdf."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Yue-Hei Ng, J., Hausknecht, M., Vijayanarasimhan, S., Vinyals, O., Monga, R., and Toderici, G. (2015, January 7\u201312). Beyond short snippets: Deep networks for video classification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299101"},{"key":"ref_31","unstructured":"Li, C., Zhong, Q., Xie, D., and Pu, S. (2017, January 10\u201314). Skeleton-based Action Recognition with Convolutional Neural Networks. Proceedings of the 2017 IEEE International Conference on Multimedia & Expo Workshops (ICMEW), Hong Kong, China."},{"key":"ref_32","unstructured":"Liao, X., He, L., Yang, Z., and Zhang, C. (2018). Video-based Person Re-identification via 3D Convolutional Networks and Non-local Attention. Asian Conference on Computer Vision, Springer."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Kalfaoglu, M.E., Kalkan, S., and Alatan, A. (2020). Late Temporal Modeling in 3D CNN Architectures with BERT for Action Recognition. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-68238-5_48"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Anvarov, F., Kim, D.H., and Song, B.C. (2020). Action Recognition Using Deep 3D CNNs with Sequential Feature Aggregation and Attention. Electronics, 9.","DOI":"10.3390\/electronics9010147"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Jalal, M.A., Aftab, W., Moore, R.K., and Mihaylova, L. (2019, January 2\u20135). Dual stream spatio-temporal motion fusion with self-attention for action recognition. Proceedings of the 22nd International Conference on Information Fusion, Ottawa, ON, Canada.","DOI":"10.23919\/FUSION43075.2019.9011320"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1187","DOI":"10.1109\/LSP.2019.2923918","article-title":"Three-Stream Network with Bidirectional Self-Attention for Action Recognition in Extreme Low-Resolution Videos","volume":"26","author":"Purwanto","year":"2019","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"226","DOI":"10.1016\/j.patrec.2018.07.034","article-title":"Joint Spatial-Temporal Attention for Action Recognition","volume":"112","author":"Yu","year":"2018","journal-title":"Pattern Recognit. Lett."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/13\/4\/662\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T14:12:25Z","timestamp":1760364745000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/13\/4\/662"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,12]]},"references-count":37,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2021,4]]}},"alternative-id":["sym13040662"],"URL":"https:\/\/doi.org\/10.3390\/sym13040662","relation":{},"ISSN":["2073-8994"],"issn-type":[{"type":"electronic","value":"2073-8994"}],"subject":[],"published":{"date-parts":[[2021,4,12]]}}}