{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T02:58:00Z","timestamp":1760151480426,"version":"build-2065373602"},"reference-count":42,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2022,3,22]],"date-time":"2022-03-22T00:00:00Z","timestamp":1647907200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62072051"],"award-info":[{"award-number":["62072051"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>In the field of video motion recognition, with the increase in network depth, there is an asymmetry in the amount of parameters and their accuracy. We propose the structure of micro-attention branches and a method of integrating attention branches in multi-branches 3D convolution networks. Our proposed attention branches can improve this asymmetry problem. Attention branches can be flexibly added to 3D convolution network in the form of plug-in, without changing the overall structure of the original network. Through this structure, the newly constructed network can fuse the attention features extracted by attention branches in real time in the process of feature extraction. By adding attention branches, the model can focus on the action subject more accurately, so as to improve the accuracy of the model. Moreover, in the case that there are multiple sub-branches in the construction module of the existing network, 3D micro-attention branches can well adapt to this scenario. In the kinetics dataset, we use our proposed micro-attention branch structure to construct a deep network, which is compared with the original network. The experimental results show that the recognition accuracy of the network with micro-attention branches is improved by 3.6% compared with the original network, while the amount of parameters to be trained is only increased by 0.6%.<\/jats:p>","DOI":"10.3390\/sym14040639","type":"journal-article","created":{"date-parts":[[2022,3,22]],"date-time":"2022-03-22T23:30:23Z","timestamp":1647991823000},"page":"639","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Micro-Attention Branch, a Flexible Plug-In That Enhances Existing 3D ConvNets"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5641-7441","authenticated-orcid":false,"given":"Yongsheng","family":"Li","sequence":"first","affiliation":[{"name":"State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tengfei","family":"Tu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6360-1116","authenticated-orcid":false,"given":"Huawei","family":"Wang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Yin","sequence":"additional","affiliation":[{"name":"National Computer Network Emergency Response Technical Team\/Coordination Center of China (CNCERT\/CC), Beijing 100029, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,3,22]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1109\/TPAMI.2012.59","article-title":"3D Convolutional Neural Networks for Human Action Recognition","volume":"35","author":"Ji","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Baccouche, M., Mamalet, F., Wolf, C., Garcia, C., and Baskurt, A. (2011). Sequential deep learning for human action recognition. International Workshop on Human Behavior Understanding, Springer.","DOI":"10.1007\/978-3-642-25446-8_4"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Carreira, J., and Zisserman, A. (2017, January 21\u201326). Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.502"},{"key":"ref_4","unstructured":"Diba, A., Fayyaz, M., Sharma, V., Karami, A.H., Mahdi Arzani, M., Yousefzadeh, R., and Van Gool, L. (2017). Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., and He, K. (November, January 27). SlowFast Networks for Video Recognition. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00630"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Ng, J.Y.H., and Davis, L.S. (2018, January 12\u201315). Temporal Difference Networks for Video Action Recognition. Proceedings of the 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe, NV, USA.","DOI":"10.1109\/WACV.2018.00176"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Sun, L., Jia, K., Yeung, D.Y., and Shi, B.E. (2015, January 7\u201313). Human Action Recognition Using Factorized Spatio-Temporal Convolutional Networks. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.522"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning Spatiotemporal Features with 3D Convolutional Networks. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Tran, D., Wang, H., Feiszli, M., and Torresani, L. (November, January 27). Video Classification With Channel-Separated Convolutional Networks. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00565"},{"key":"ref_10","unstructured":"Ferrari, V., Hebert, M., Sminchisescu, C., and Weiss, Y. (2018). Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification. Computer Vision\u2014ECCV 2018, Springer International Publishing."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Yao, L., Torabi, A., Cho, K., Ballas, N., Pal, C., Larochelle, H., and Courville, A. (2015, January 7\u201313). Describing Videos by Exploiting Temporal Structure. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.512"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Yue-Hei Ng, J., Hausknecht, M., Vijayanarasimhan, S., Vinyals, O., Monga, R., and Toderici, G. (2015, January 7\u201312). Beyond short snippets: Deep networks for video classification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299101"},{"key":"ref_13","first-page":"2208","article-title":"Trajectory Convolution for Action Recognition","volume":"Volume 31","author":"Bengio","year":"2018","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Ferrari, V., Hebert, M., Sminchisescu, C., and Weiss, Y. (2018). ECO: Efficient Convolutional Network for Online Video Understanding. Computer Vision\u2014ECCV 2018, Springer International Publishing.","DOI":"10.1007\/978-3-030-01252-6"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Yang, C., Xu, Y., Shi, J., Dai, B., and Zhou, B. (2020, January 14\u201319). Temporal Pyramid Network for Action Recognition. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00067"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C. (2020, January 14\u201319). X3D: Expanding Architectures for Efficient Video Recognition. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00028"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Willems, G., Tuytelaars, T., and Van Gool, L. (2008). An efficient dense and scale-invariant spatio-temporal interest point detector. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-540-88688-4_48"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Wang, H., and Schmid, C. (2013, January 1\u20138). Action recognition with improved trajectories. Proceedings of the IEEE International Conference on Computer Vision, Sydney, Australia.","DOI":"10.1109\/ICCV.2013.441"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Laptev, I., and Lindeberg, T. (2004, January 23\u201326). Velocity adaptation of space-time interest points. Proceedings of the 17th International Conference on Pattern Recognition (ICPR 2004), Washington, DC, USA.","DOI":"10.1109\/ICPR.2004.1334003"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Raptis, M., Kokkinos, I., and Soatto, S. (2012, January 16\u201321). Discovering discriminative action parts from mid-level video representations. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6247807"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Jain, A., Gupta, A., Rodriguez, M., and Davis, L.S. (2013, January 23\u201328). Representing videos using mid-level discriminative patches. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Washington, DC, USA.","DOI":"10.1109\/CVPR.2013.332"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Wang, L., Qiao, Y., and Tang, X. (2013, January 14\u201319). Motionlets: Mid-level 3d parts for human motion recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR.2013.345"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Qiu, Z., Yao, T., and Mei, T. (2017, January 22\u201329). Learning spatio-temporal representation with pseudo-3d residual networks. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.590"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., Gupta, A., and He, K. (2018, January 18\u201322). Non-local neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1360","DOI":"10.1109\/TIP.2005.852470","article-title":"Toward automatic phenotyping of developing embryos from videos","volume":"14","author":"Ning","year":"2005","journal-title":"IEEE Trans. Image Process."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"677","DOI":"10.1109\/TPAMI.2016.2599174","article-title":"Long-Term Recurrent Convolutional Networks for Visual Recognition and Description","volume":"39","author":"Donahue","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Wu, Z., Wang, X., Jiang, Y.G., Ye, H., and Xue, X. (2015, January 26\u201330). Modeling spatial-temporal clues in a hybrid deep learning framework for video classification. Proceedings of the 23rd ACM International Conference on Multimedia, Brisbane, Australia.","DOI":"10.1145\/2733373.2806222"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long Short-Term Memory","volume":"9","author":"Sepp","year":"1997","journal-title":"Neural Comput."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Taylor, G.W., Fergus, R., LeCun, Y., and Bregler, C. (2010). Convolutional learning of spatio-temporal features. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-642-15567-3_11"},{"key":"ref_30","first-page":"568","article-title":"Two-Stream Convolutional Networks for Action Recognition in Videos","volume":"Volume 1","author":"Simonyan","year":"2014","journal-title":"Proceedings of the 27th International Conference on Neural Information Processing Systems (NIPS\u201914)"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Wu, Z., Jiang, Y.G., Wang, X., Ye, H., and Xue, X. (2016, January 15\u201319). Multi-stream multi-class fusion of deep networks for video classification. Proceedings of the 24th ACM International Conference on Multimedia, Amsterdam, The Netherlands.","DOI":"10.1145\/2964284.2964328"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Ch\u00e9ron, G., Laptev, I., and Schmid, C. (2015, January 7\u201313). P-cnn: Pose-based cnn features for action recognition. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.368"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Pinz, A., and Zisserman, A. (2016, January 27\u201330). Convolutional Two-Stream Network Fusion for Video Action Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.213"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Weinzaepfel, P., Harchaoui, Z., and Schmid, C. (2015, January 7\u201313). Learning to track for spatio-temporal action localization. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.362"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Goyal, R., Kahou, S.E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., and Mueller-Freitag, M. (2017, January 22\u201329). The \u201cSomething Something\u201d Video Database for Learning and Evaluating Visual Common Sense. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.622"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., and Serre, T. (2011, January 6\u201313). HMDB: A large video database for human motion recognition. Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"ref_37","unstructured":"Soomro, K., Roshan Zamir, A., and Shah, M. (2012). UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild. arXiv."},{"key":"ref_38","unstructured":"Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., and Natsev, P. (2017). The Kinetics Human Action Video Dataset. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"502","DOI":"10.1109\/TPAMI.2019.2901464","article-title":"Moments in time dataset: One million videos for event understanding","volume":"42","author":"Monfort","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Caba Heilbron, F., Escorcia, V., Ghanem, B., and Carlos Niebles, J. (2015, January 7\u201312). Activitynet: A large-scale video benchmark for human activity understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Wang, F., Jiang, M., Qian, C., Yang, S., Li, C., Zhang, H., Wang, X., and Tang, X. (2017, January 21\u201326). Residual Attention Network for Image Classification. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.683"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Hamprecht, F.A., Schn\u00f6rr, C., and J\u00e4hne, B. (2007). A Duality Based Approach for Realtime TV-L1 Optical Flow. Pattern Recognition, Springer.","DOI":"10.1007\/978-3-540-74936-3"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/14\/4\/639\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:40:59Z","timestamp":1760136059000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/14\/4\/639"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,22]]},"references-count":42,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2022,4]]}},"alternative-id":["sym14040639"],"URL":"https:\/\/doi.org\/10.3390\/sym14040639","relation":{},"ISSN":["2073-8994"],"issn-type":[{"type":"electronic","value":"2073-8994"}],"subject":[],"published":{"date-parts":[[2022,3,22]]}}}