{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T14:50:41Z","timestamp":1781016641768,"version":"3.54.1"},"reference-count":50,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2023,2,3]],"date-time":"2023-02-03T00:00:00Z","timestamp":1675382400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the National Key Research and Development Program of China","award":["255"],"award-info":[{"award-number":["255"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In recent years, deep learning techniques have excelled in video action recognition. However, currently commonly used video action recognition models minimize the importance of different video frames and spatial regions within some specific frames when performing action recognition, which makes it difficult for the models to adequately extract spatiotemporal features from the video data. In this paper, an action recognition method based on improved residual convolutional neural networks (CNNs) for video frames and spatial attention modules is proposed to address this problem. The network can guide what and where to emphasize or suppress with essentially little computational cost using the video frame attention module and the spatial attention module. It also employs a two-level attention module to emphasize feature information along the temporal and spatial dimensions, respectively, highlighting the more important frames in the overall video sequence and the more important spatial regions in some specific frames. Specifically, we create the video frame and spatial attention map by successively adding the video frame attention module and the spatial attention module to aggregate the spatial and temporal dimensions of the intermediate feature maps of the CNNs to obtain different feature descriptors, thus directing the network to focus more on important video frames and more contributing spatial regions. The experimental results further show that the network performs well on the UCF-101 and HMDB-51 datasets.<\/jats:p>","DOI":"10.3390\/s23031707","type":"journal-article","created":{"date-parts":[[2023,2,6]],"date-time":"2023-02-06T02:06:43Z","timestamp":1675649203000},"page":"1707","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":32,"title":["Two-Level Attention Module Based on Spurious-3D Residual Networks for Human Action Recognition"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1030-5660","authenticated-orcid":false,"given":"Bo","family":"Chen","sequence":"first","affiliation":[{"name":"Science and Technology on Microsystem Laboratory, Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences, Shanghai 201800, China"},{"name":"School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fangzhou","family":"Meng","sequence":"additional","affiliation":[{"name":"Science and Technology on Microsystem Laboratory, Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences, Shanghai 201800, China"},{"name":"School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongying","family":"Tang","sequence":"additional","affiliation":[{"name":"Science and Technology on Microsystem Laboratory, Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences, Shanghai 201800, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guanjun","family":"Tong","sequence":"additional","affiliation":[{"name":"Science and Technology on Microsystem Laboratory, Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences, Shanghai 201800, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,2,3]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"4354","DOI":"10.1109\/TIP.2016.2590322","article-title":"Pedestrian Behavior Modeling from Stationary Crowds With Applications to Intelligent Surveillance","volume":"25","author":"Yi","year":"2016","journal-title":"IEEE Trans. Image Process."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zhuang, C., Zhou, H., and Sakane, S. (2016, January 3\u20137). Learning by showing: An end-to-end imitation leaning approach for robot action recognition and generation. Proceedings of the 2016 IEEE International Conference on Robotics and Biomimetics (ROBIO), Qingdao, China.","DOI":"10.1109\/ROBIO.2016.7866317"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"2071","DOI":"10.1007\/s11831-021-09649-9","article-title":"A Survey on Deep Learning Approaches to Medical Images and a Systematic Look up into Real-Time Object Detection","volume":"29","author":"Kaur","year":"2022","journal-title":"Arch. Comput. Methods Eng."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"166334","DOI":"10.1007\/s11704-021-0236-9","article-title":"ResLNet: Deep residual LSTM network with longer input for action recogntion","volume":"16","author":"Wang","year":"2022","journal-title":"Front. Comput. Sci."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Vrskova, R., Hudec, R., Kamencay, P., and Sykora, P. (2022). Human Activity Classification Using the 3DCNN Architecture. Appl. Sci. Basel, 12.","DOI":"10.3390\/app12020931"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"689","DOI":"10.1109\/TMM.2021.3058050","article-title":"Human Action Recognition by Discriminative Feature Pooling and Video Segment Attention Model","volume":"24","author":"Moniruzzaman","year":"2022","journal-title":"IEEE Trans. Multimed."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"3097","DOI":"10.1049\/ipr2.12541","article-title":"Video-based action recognition using spurious-3D residual attention networks","volume":"16","author":"Chen","year":"2022","journal-title":"Iet Image Process."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_10","unstructured":"Du, T., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 11\u201318). Learning Spatiotemporal Features with 3D Convolutional Networks. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile."},{"key":"ref_11","unstructured":"Lan, Z.Z., Lin, M., Li, X.C., Hauptmann, A.G., and Raj, B. (2015, January 7\u201312). Beyond Gaussian Pyramid: Multi-skip Feature Stacking for Action Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Peng, X.J., Zou, C.Q., Qiao, Y., and Peng, Q. (2014, January 6\u201312). Action Recognition with Stacked Fisher Vectors. Proceedings of the 13th European Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_38"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, H., and Schmid, C. (2013, January 1\u20138). Action Recognition with Improved Trajectories. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Sydney, Australia.","DOI":"10.1109\/ICCV.2013.441"},{"key":"ref_14","first-page":"84","article-title":"Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Commun. ACM"},{"key":"ref_15","unstructured":"Simonyan, K., and Zisserman, A. (2014, January 8\u201313). Two-Stream Convolutional Networks for Action Recognition in Videos. Proceedings of the 28th Conference on Neural Information Processing Systems (NIPS), Montreal, CA, USA."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., and Van Gool, L. (2016, January 8\u201316). Temporal Segment Networks: Towards Good Practices for Deep Action Recognition. Proceedings of the 14th European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Carreira, J., and Zisserman, A. (2017, January 21\u201326). Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. Proceedings of the 30th IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.502"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Sun, M., Yuan, Y.C., Zhou, F., and Ding, E.R. (2018, January 8\u201314). Multi-Attention Multi-Class Constraint for Fine-grained Image Recognition. Proceedings of the 15th European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01270-0_49"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zheng, H.L., Fu, J.L., Mei, T., and Luo, J.B. (2017, January 22\u201329). Learning Multi-Attention Convolutional Neural Network for Fine-Grained Image Recognition. Proceedings of the 16th IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.557"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wang, F., Jiang, M.Q., Qian, C., Yang, S., Li, C., Zhang, H.G., Wang, X., and Tang, X. (2017, January 21\u201326). Residual Attention Network for Image Classification. Proceedings of the 30th IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.683"},{"key":"ref_21","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017, January 4\u20139). Attention Is All You Need. Proceedings of the 31st Annual Conference on Neural Information Processing Systems (NIPS), Long Beach, CA, USA."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Li, H., Chen, J., Hu, R., Yu, M., Chen, H., and Xu, Z. (2019, January 8\u201311). Action Recognition Using Visual Attention with Reinforcement Learning. Proceedings of the 25th International Conference on MultiMedia Modeling (MMM), Thessaloniki, Greece.","DOI":"10.1007\/978-3-030-05716-9_30"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Ma, C.Y., Kadav, A., Melvin, I., Kira, Z., AlRegib, G., and Graf, H.P. (2018, January 18\u201323). Attend and Interact: Higher-Order Object Interactions for Video Understanding. Proceedings of the 31st IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00710"},{"key":"ref_24","unstructured":"Girdhar, R., and Ramanan, D. (2017, January 4\u20139). Attentional Pooling for Action Recognition. Proceedings of the 31st Annual Conference on Neural Information Processing Systems (NIPS), Long Beach, CA, USA."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_26","unstructured":"Mnih, V., Heess, N., Graves, A., and Kavukcuoglu, K. (2014, January 8\u201313). Recurrent Models of Visual Attention. Proceedings of the 28th Conference on Neural Information Processing Systems (NIPS), Montreal, QC, Canada."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"21789","DOI":"10.1007\/s11042-021-10752-z","article-title":"Spatial-temporal channel-wise attention network for action recognition","volume":"80","author":"Chen","year":"2021","journal-title":"Multimed. Tools Appl."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"16785","DOI":"10.1109\/ACCESS.2020.2968024","article-title":"Learning Attention-Enhanced Spatiotemporal Representation for Action Recognition","volume":"8","author":"Shi","year":"2020","journal-title":"IEEE Access"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Long, X., Gan, C., de Melo, G., Wu, J.J., Liu, X., and Wen, S. (2018, January 18\u201323). Attention Clusters: Purely Attention Based Local Feature Integration for Video Classification. Proceedings of the 31st IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00817"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhang, J.C., and Peng, Y.X. (2019, January 8\u201311). Hierarchical Vision-Language Alignment for Video Captioning. Proceedings of the 25th International Conference on MultiMedia Modeling (MMM), Thessaloniki, Greece.","DOI":"10.1007\/978-3-030-05710-7_4"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Zhang, J.C., Peng, Y.X., and Soc, I.C. (2019, January 16\u201320). Object-aware Aggregation with Bidirectional Temporal Graph for Video Captioning. Proceedings of the 32nd IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00852"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"773","DOI":"10.1109\/TCSVT.2018.2808685","article-title":"Two-Stream Collaborative Learning With Spatial-Temporal Attention for Video Classification","volume":"29","author":"Peng","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.-Y., and Kweon, I.S. (2018, January 8\u201314). CBAM: Convolutional Block Attention Module. Proceedings of the 15th European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_34","unstructured":"Wang, L., Xiong, Y., Wang, Z., and Qiao, Y. (2015). Towards Good Practices for Very Deep Two-Stream ConvNets. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2015, January 11\u201318). Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.123"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Qiu, Z., Yao, T., and Mei, T. (2017, January 22\u201329). Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks. Proceedings of the 16th IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.590"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Zhou, Y., Sun, X., Zha, Z.-J., and Zeng, W. (2018, January 18\u201323). MiCT: Mixed 3D\/2D Convolutional Tube for Human Action Recognition. Proceedings of the 31st IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00054"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"5783","DOI":"10.1109\/TIP.2020.2984904","article-title":"STA-CNN: Convolutional Spatial-Temporal Attention Learning for Action Recognition","volume":"29","author":"Yang","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Yang, G., Yang, Y., Lu, Z., Yang, J., Liu, D., Zhou, C., and Fan, Z. (2022). STA-TSN: Spatial-Temporal Attention Temporal Segment Network for action recognition in video. PLoS ONE, 17.","DOI":"10.1371\/journal.pone.0265115"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"1059","DOI":"10.1049\/iet-ipr.2019.0963","article-title":"Dual attention convolutional network for action recognition","volume":"14","author":"Li","year":"2020","journal-title":"Iet Image Process."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., and Paluri, M. (2018, January 18\u201323). A Closer Look at Spatiotemporal Convolutions for Action Recognition. Proceedings of the 31st IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00675"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"104122","DOI":"10.1016\/j.imavis.2021.104122","article-title":"2D progressive fusion module for action recognition","volume":"109","author":"Shen","year":"2021","journal-title":"Image Vis. Comput."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Zhang, Y. (2022). MEST: An Action Recognition Network with Motion Encoder and Spatio-Temporal Module. Sensors, 22.","DOI":"10.3390\/s22176595"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"9875","DOI":"10.1007\/s11042-022-11937-w","article-title":"Deep learning network model based on fusion of spatiotemporal features for action recognition","volume":"81","author":"Yang","year":"2022","journal-title":"Multimed. Tools Appl."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"2799","DOI":"10.1109\/TIP.2018.2890749","article-title":"Action-Stage Emphasized Spatiotemporal VLAD for Video Action Recognition","volume":"28","author":"Tu","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Wang, L., Tong, Z., Ji, B., Wu, G., and Ieee Comp, S.O.C. (2021, January 19\u201325). TDN: Temporal Difference Networks for Efficient Action Recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.00193"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"2119","DOI":"10.1587\/transinf.2022EDP7058","article-title":"Model-Agnostic Multi-Domain Learning with Domain-Specific Adapters for Action Recognition","volume":"105","author":"Omi","year":"2022","journal-title":"IEICE Trans. Inf. Syst."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"103406","DOI":"10.1016\/j.cviu.2022.103406","article-title":"TCLR: Temporal contrastive learning for video representation","volume":"219","author":"Dave","year":"2022","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"1155","DOI":"10.1109\/ACCESS.2017.2778011","article-title":"Action Recognition in Video Sequences using Deep Bi-Directional LSTM With CNN Features","volume":"6","author":"Ullah","year":"2018","journal-title":"IEEE Access"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"813","DOI":"10.1109\/TETCI.2020.3014367","article-title":"HAR-Depth: A Novel Framework for Human Action Recognition Using Sequential Learning and Depth Estimated History Images","volume":"5","author":"Sahoo","year":"2021","journal-title":"IEEE Trans. Emerg. Top. Comput. Intell."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/3\/1707\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:23:45Z","timestamp":1760120625000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/3\/1707"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,3]]},"references-count":50,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,2]]}},"alternative-id":["s23031707"],"URL":"https:\/\/doi.org\/10.3390\/s23031707","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,3]]}}}