{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T17:14:21Z","timestamp":1765041261464,"version":"build-2065373602"},"reference-count":46,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2022,3,4]],"date-time":"2022-03-04T00:00:00Z","timestamp":1646352000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62171328"],"award-info":[{"award-number":["62171328"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Hubei Technology Innovation Project","award":["2019AAA045"],"award-info":[{"award-number":["2019AAA045"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Temporal modeling is the key for action recognition in videos, but traditional 2D CNNs do not capture temporal relationships well. 3D CNNs can achieve good performance, but are computationally intensive and not well practiced on existing devices. Based on these problems, we design a generic and effective module called spatio-temporal motion network (SMNet). SMNet maintains the complexity of 2D and reduces the computational effort of the algorithm while achieving performance comparable to 3D CNNs. SMNet contains a spatio-temporal excitation module (SE) and a motion excitation module (ME). The SE module uses group convolution to fuse temporal information to reduce the number of parameters in the network, and uses spatial attention to extract spatial information. The ME module uses the difference between adjacent frames to extract feature-level motion patterns between adjacent frames, which can effectively encode motion features and help identify actions efficiently. We use ResNet-50 as the backbone network and insert SMNet into the residual blocks to form a simple and effective action network. The experiment results on three datasets, namely Something-Something V1, Something-Something V2, and Kinetics-400, show that it out performs state-of-the-arts motion recognition networks.<\/jats:p>","DOI":"10.3390\/e24030368","type":"journal-article","created":{"date-parts":[[2022,3,6]],"date-time":"2022-03-06T20:35:50Z","timestamp":1646598950000},"page":"368","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["A Spatio-Temporal Motion Network for Action Recognition Based on Spatial Attention"],"prefix":"10.3390","volume":"24","author":[{"given":"Qi","family":"Yang","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, Wuhan Institute of Technology, Wuhan 430205, China"},{"name":"Hubei Key Laboratory of Intelligent Robot, Wuhan Institute of Technology, Wuhan 430205, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3900-6456","authenticated-orcid":false,"given":"Tongwei","family":"Lu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Wuhan Institute of Technology, Wuhan 430205, China"},{"name":"Hubei Key Laboratory of Intelligent Robot, Wuhan Institute of Technology, Wuhan 430205, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huabing","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Wuhan Institute of Technology, Wuhan 430205, China"},{"name":"Hubei Key Laboratory of Intelligent Robot, Wuhan Institute of Technology, Wuhan 430205, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,3,4]]},"reference":[{"key":"ref_1","first-page":"568","article-title":"Two-stream convolutional networks for action recognition in videos","volume":"1","author":"Simonyan","year":"2014","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_2","first-page":"20","article-title":"Temporal segment networks: Towards good practices for deep action recognition","volume":"9912","author":"Wang","year":"2016","journal-title":"Comput. Vis."},{"key":"ref_3","unstructured":"Wang, L., Xiong, Y., Wang, Z., and Qiao, Y. (2015). Towards Good Practices for Very Deep Two-Stream ConvNets. arXiv."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1109\/TPAMI.2012.59","article-title":"3D Convolutional neural networks for human action recognition","volume":"35","author":"Ji","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning spatiotemporal features with 3D convolutional networks. Proceedings of the 2015 International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_6","first-page":"214","article-title":"A duality based approach for realtime TV-L1 optical flow","volume":"4713","author":"Zach","year":"2007","journal-title":"Jt. Pattern Recognit. Symp."},{"key":"ref_7","unstructured":"Zhu, Y., Lan, Z., Newsam, S., and Hauptmann, A. (2018). Hidden Two-Stream Convolutional Networks for Action Recognition. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"652","DOI":"10.1109\/TPAMI.2019.2938758","article-title":"Res2Net: A New Multi-scale Backbone Architecture","volume":"43","author":"Gao","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_9","first-page":"1647","article-title":"FlowNet 2.0: Evolution of optical flow estimation with deep networks","volume":"2017","author":"Ilg","year":"2017","journal-title":"IEEE Conf. Comput. Vis. Pattern Recognit."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018). CBAM: Convolutional block attention module. arXiv.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016). Deep residual learning for image recognition. arXiv.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Lin, J., Gan, C., Wang, K., and Han, S. (November, January 27). TSM: Temporal Shift Module for Efficient Video Understanding. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00718"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Pinz, A., and Zisserman, A. (2016). Convolutional Two-Stream Network Fusion for Video Action Recognition. arXiv.","DOI":"10.1109\/CVPR.2016.213"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Lee, M., Lee, S., Son, S., Park, G., and Kwak, N. (2018). Motion feature network: Fixed motion filter for action recognition. arXiv.","DOI":"10.1007\/978-3-030-01249-6_24"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Liu, Z., Luo, D., Wang, Y., Wang, L., Tai, Y., Wang, C., Li, J., Huang, F., and Lu, T. (2020). TEINet: Towards an efficient architecture for video recognition. arXiv.","DOI":"10.1609\/aaai.v34i07.6836"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Li, Y., Ji, B., Shi, X., Zhang, J., Kang, B., and Wang, L. (2020). TEA: Temporal Excitation and Aggregation for Action Recognition. arXiv.","DOI":"10.1109\/CVPR42600.2020.00099"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Jiang, B., Wang, M., Gan, W., Wu, W., and Yan, J. (2019). STM: Spatiotemporal and motion encoding for action recognition. arXiv.","DOI":"10.1109\/ICCV.2019.00209"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Carreira, J., and Zisserman, A. (2017). Quo Vadis, action recognition? A new model and the kinetics dataset. arXiv.","DOI":"10.1109\/CVPR.2017.502"},{"key":"ref_19","unstructured":"Diba, A., Fayyaz, M., Sharma, V., Karami, A.H., and Yousefzadeh, R. (2017). Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K.Q. (2017). Densely connected convolutional networks. arXiv.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Qiu, Z., Yao, T., and Mei, T. (2017). Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks. arXiv.","DOI":"10.1109\/ICCV.2017.590"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., and He, K. (2019). Slowfast networks for video recognition. arXiv.","DOI":"10.1109\/ICCV.2019.00630"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C. (2020). X3D: Expanding Architectures for Efficient Video Recognition. arXiv.","DOI":"10.1109\/CVPR42600.2020.00028"},{"key":"ref_24","unstructured":"Jaderberg, M., Simonyan, K., Zisserman, A., and Kavukcuoglu, K. (2015). Spatial transformer networks. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"2011","DOI":"10.1109\/TPAMI.2019.2913372","article-title":"Squeeze-and-Excitation Networks","volume":"42","author":"Hu","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wang, F., Jiang, M., Qian, C., Yang, S., Li, C., Zhang, H., Wang, X., and Tang, X. (2017). Residual attention network for image classification. arXiv.","DOI":"10.1109\/CVPR.2017.683"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1145\/3065386","article-title":"ImageNet classification with deep convolutional neural networks","volume":"60","author":"Krizhevsky","year":"2017","journal-title":"Commun. ACM"},{"key":"ref_28","unstructured":"Zagoruyko, S., and Komodakis, N. (2017). Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Goyal, R., Kahou, S.E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., and Mueller-Freitag, M. (2017). The \u2019Something Something\u2019 Video Database for Learning and Evaluating Visual Common Sense. arXiv.","DOI":"10.1109\/ICCV.2017.622"},{"key":"ref_30","unstructured":"Kay, W., Carreira, J., Simonyan, K., Zhang, B., and Zisserman, A. (2017). The Kinetics Human Action Video Dataset. arXiv."},{"key":"ref_31","unstructured":"Jia, D., Wei, D., Socher, R., Li, L.J., Kai, L., and Li, F.F. (2009, January 20\u201325). ImageNet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., Gupta, A., and He, K. (2018). Non-local Neural Networks. arXiv.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhou, B., Andonian, A., Oliva, A., and Torralba, A. (2018). Temporal Relational Reasoning in Videos. arXiv.","DOI":"10.1007\/978-3-030-01246-5_49"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zolfaghari, M., Singh, K., and Brox, T. (2018). ECO: Efficient Convolutional Network for Online Video Understanding, Springer.","DOI":"10.1007\/978-3-030-01216-8_43"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Liu, Z., Wang, L., Wu, W., Qian, C., and Lu, T. (2020). TAM: Temporal Adaptive Module for Video Recognition. arXiv.","DOI":"10.1109\/ICCV48922.2021.01345"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Xie, S., Sun, C., Huang, J., Tu, Z., and Murphy, K. (2018). Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. arXiv.","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Wang, H., Tran, D., Torresani, L., and Feiszli, M. (2020). Video modeling with correlation networks. arXiv.","DOI":"10.1109\/CVPR42600.2020.00043"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Li, X., Wang, Y., Zhou, Z., and Qiao, Y. (2020). SmallBigNet: Integrating Core and Contextual Views for Video Classification. arXiv.","DOI":"10.1109\/CVPR42600.2020.00117"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Tran, D., Wang, H., Torresani, L., Ray, J., Lecun, Y., and Paluri, M. (2018). A Closer Look at Spatiotemporal Convolutions for Action Recognition. arXiv.","DOI":"10.1109\/CVPR.2018.00675"},{"key":"ref_40","unstructured":"Fan, Q., Chen, C.F., Kuehne, H., Pistoia, M., and Cox, D. (2019). More Is Less: Learning Efficient Video Representations by Big-Little Network and Depthwise Temporal Aggregation. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Wang, L., Li, W., and Van Gool, L. (2018). Appearance-and-Relation Networks for Video Classification. arXiv.","DOI":"10.1109\/CVPR.2018.00155"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Li, X., Zhang, Y., Liu, C., Shuai, B., and Tighe, J. (2021). VidTr: Video Transformer Without Convolutions. arXiv.","DOI":"10.1109\/ICCV48922.2021.01332"},{"key":"ref_43","unstructured":"Bertasius, G., Wang, H., and Torresani, L. (2021). Is Space-Time Attention All You Need for Video Understanding?. arXiv."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lui, M., and Schmid, C. (2021). ViViT: A Video Vision Transformer. arXiv.","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Fan, H., Xiong, B., Mangalam, K., Li, Y., and Feichtenhofer, C. (2021). Multiscale Vision Transformers. arXiv.","DOI":"10.1109\/ICCV48922.2021.00675"},{"key":"ref_46","unstructured":"Patrick, M., Campbell, D., Asano, Y.M., Metze, I., and Henriques, J.F. (2021). Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers. arXiv."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/3\/368\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:31:52Z","timestamp":1760135512000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/3\/368"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,4]]},"references-count":46,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2022,3]]}},"alternative-id":["e24030368"],"URL":"https:\/\/doi.org\/10.3390\/e24030368","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2022,3,4]]}}}