{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T04:38:06Z","timestamp":1787027886794,"version":"build-2736575974"},"reference-count":45,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T00:00:00Z","timestamp":1760313600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272236"],"award-info":[{"award-number":["62272236"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62376128"],"award-info":[{"award-number":["62376128"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100004608","name":"Natural Science Foundation of Jiangsu Province","doi-asserted-by":"crossref","award":["BK20201136"],"award-info":[{"award-number":["BK20201136"]}],"id":[{"id":"10.13039\/501100004608","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100004608","name":"Natural Science Foundation of Jiangsu Province","doi-asserted-by":"crossref","award":["BK20191401"],"award-info":[{"award-number":["BK20191401"]}],"id":[{"id":"10.13039\/501100004608","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>Dynamic hand gesture recognition based on computer vision aims at enabling computers to understand the semantic meaning conveyed by hand gestures in videos. Existing methods predominately rely on spatiotemporal attention mechanisms to extract hand motion features in a large spatiotemporal scope. However, they cannot accurately focus on the moving hand region for hand feature extraction because frame sequences contain a substantial amount of redundant information. Although multimodal techniques can extract a wider variety of hand features, they are less successful at utilizing information interactions between various modalities for accurate feature extraction. To address these challenges, this study proposes a multimodal hand gesture recognition model combining inter-frame motion and shared attention weights. By jointly using an inter-frame motion attention (IFMA) mechanism and adaptive down-sampling (ADS), the spatiotemporal search scope can be effectively narrowed down to the hand-related regions based on the characteristic of hands exhibiting obvious movements. The proposed inter-modal attention weight (IMAW) loss enables RGB and Depth modalities to share attention, allowing each to adjust its distribution based on the other. Experimental results on the EgoGesture, NVGesture, and Jester datasets demonstrate the superiority of our proposed model over existing state-of-the-art methods in terms of hand motion feature extraction and hand gesture recognition accuracy.<\/jats:p>","DOI":"10.3390\/computers14100432","type":"journal-article","created":{"date-parts":[[2025,10,15]],"date-time":"2025-10-15T07:17:52Z","timestamp":1760512672000},"page":"432","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["A Novel Multimodal Hand Gesture Recognition Model Using Combined Approach of Inter-Frame Motion and Shared Attention Weights"],"prefix":"10.3390","volume":"14","author":[{"given":"Xiaorui","family":"Zhang","sequence":"first","affiliation":[{"name":"College of Computer and Information Engineering, Nanjing Tech University, Nanjing 211816, China"},{"name":"College of Electronic and Information Engineering, Nanjing University of Information Science and Technology, Nanjing 210044, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuaitong","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science, Nanjing University of Information Science and Technology, Nanjing 210044, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-9120-6066","authenticated-orcid":false,"given":"Xianglong","family":"Zeng","sequence":"additional","affiliation":[{"name":"College of Automation Engineering, Nanjing University of Aeronautics and Astronautics, Nanjing 210016, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peisen","family":"Lu","sequence":"additional","affiliation":[{"name":"School of Computer Science, Nanjing University of Information Science and Technology, Nanjing 210044, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wei","family":"Sun","sequence":"additional","affiliation":[{"name":"College of Automation, Nanjing University of Information Science and Technology, Nanjing 210044, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,10,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"125929","DOI":"10.1016\/j.eswa.2024.125929","article-title":"A Comparative Study of Advanced Technologies and Methods in Hand Gesture Analysis and Recognition Systems","volume":"266","author":"Rahman","year":"2025","journal-title":"Expert Syst. Appl."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"58","DOI":"10.1016\/j.neucom.2022.12.022","article-title":"Improving Dynamic Gesture Recognition in Untrimmed Videos by an Online Lightweight Framework and a New Gesture Dataset ZJUGesture","volume":"523","author":"Xu","year":"2023","journal-title":"Neurocomputing"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1581","DOI":"10.1007\/s40747-023-01173-6","article-title":"Computer Vision-Based Hand Gesture Recognition for Human-Robot Interaction: A Review","volume":"10","author":"Qi","year":"2024","journal-title":"Complex Intell. Syst."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"84","DOI":"10.3724\/SP.J.2096-5796.2018.0006","article-title":"Gesture Interaction in Virtual Reality","volume":"1","author":"Yang","year":"2019","journal-title":"Virtual Real. Intell. Hardw."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"115657","DOI":"10.1016\/j.eswa.2021.115657","article-title":"Vision-Based Hand Gesture Recognition Using Deep Learning for the Interpretation of Sign Language","volume":"182","author":"Sharma","year":"2021","journal-title":"Expert Syst. Appl."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"143599","DOI":"10.1109\/ACCESS.2024.3421992","article-title":"A Systematic Review of Hand Gesture Recognition: An Update from 2018 to 2024","volume":"12","author":"Hashi","year":"2024","journal-title":"IEEE Access"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Xing, Z., Dai, Q., Hu, H., Chen, J., Wu, Z., and Jiang, Y.-G. (2023, January 17\u201324). Svformer: Semi-Supervised Video Transformer for Action Recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01804"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"109455","DOI":"10.1016\/j.patcog.2023.109455","article-title":"Relation-Mining Self-Attention Network for Skeleton-Based Human Action Recognition","volume":"139","author":"Gedamu","year":"2023","journal-title":"Pattern Recognit."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1038\/s41746-020-00376-2","article-title":"Deep Learning-Enabled Medical Computer Vision","volume":"4","author":"Esteva","year":"2021","journal-title":"npj Digit. Med."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"2226","DOI":"10.1109\/TCSS.2022.3184420","article-title":"A Local Spatial\u2013Temporal Synchronous Network to Dynamic Gesture Recognition","volume":"10","author":"Zhao","year":"2022","journal-title":"IEEE Trans. Comput. Soc. Syst."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"100379","DOI":"10.1016\/j.cosrev.2021.100379","article-title":"A Survey on Deep Learning and Its Applications","volume":"40","author":"Dong","year":"2021","journal-title":"Comput. Sci. Rev."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"113794","DOI":"10.1016\/j.eswa.2020.113794","article-title":"Sign Language Recognition: A Deep Survey","volume":"164","author":"Rastgoo","year":"2021","journal-title":"Expert Syst. Appl."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"48","DOI":"10.1016\/j.neucom.2021.03.091","article-title":"A Review on the Attention Mechanism of Deep Learning","volume":"452","author":"Niu","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Saini, M., Fatemi, M., and Alizad, A. (2024). Fast Inter-Frame Motion Correction in Contrast-Free Ultrasound Quantitative Microvasculature Imaging Using Deep Learning. Sci. Rep., 14.","DOI":"10.1038\/s41598-024-77610-4"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1711","DOI":"10.1007\/s11263-024-02258-6","article-title":"Facial Action Unit Detection by Adaptively Constraining Self-Attention and Causally Deconfounding Sample","volume":"133","author":"Shao","year":"2025","journal-title":"Int. J. Comput. Vis."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Wang, Y., Yang, G., Li, S., Li, Y., He, L., and Liu, D. (2023). Arrhythmia Classification Algorithm Based on Multi-Head Self-Attention Mechanism. Biomed. Signal Process. Control, 79.","DOI":"10.1016\/j.bspc.2022.104206"},{"key":"ref_17","first-page":"93","article-title":"Deep Learning Attention Mechanism in Medical Image Analysis: Basics and Beyonds","volume":"2","author":"Li","year":"2023","journal-title":"Int. J. Netw. Dyn. Intell."},{"key":"ref_18","unstructured":"Chen, Y., Zhao, L., Peng, X., Yuan, J., and Metaxas, D.N. (2019). Construct Dynamic Graphs for Hand Gesture Recognition via Spatial-Temporal Attention. arXiv."},{"key":"ref_19","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (December, January 30). Decoupled Spatial-Temporal Attention Network for Skeleton-Based Action-Gesture Recognition. Proceedings of the Asian Conference on Computer Vision, Kyoto, Japan."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Miah, A.S.M., Hasan, M.A.M., Shin, J., Okuyama, Y., and Tomioka, Y. (2023). Multistage Spatial Attention-Based Neural Network for Hand Gesture Recognition. Computers, 12.","DOI":"10.3390\/computers12010013"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2433","DOI":"10.1007\/s00371-020-01955-w","article-title":"STA-GCN: Two-Stream Graph Convolutional Network with Spatial\u2013Temporal Attention for Hand Gesture Recognition","volume":"36","author":"Zhang","year":"2020","journal-title":"Vis. Comput."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"2368","DOI":"10.1109\/TITS.2014.2337331","article-title":"Hand Gesture Recognition in Real Time for Automotive Interfaces: A Multimodal Vision-Based Approach and Evaluations","volume":"15","author":"Trivedi","year":"2014","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Miao, Q., Li, Y., Ouyang, W., Ma, Z., Xu, X., Shi, W., and Cao, X. (2017, January 22\u201329). Multimodal Gesture Recognition Based on the Resc3d Network. Proceedings of the IEEE International Conference on Computer Vision Workshops, Venice, Italy.","DOI":"10.1109\/ICCVW.2017.360"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"671","DOI":"10.32604\/csse.2023.035119","article-title":"Multimodal Spatiotemporal Feature Map for Dynamic Gesture Recognition","volume":"46","author":"Zhang","year":"2023","journal-title":"Comput. Syst. Sci. Eng."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"110","DOI":"10.1109\/JAS.2020.1003465","article-title":"Dynamic Hand Gesture Recognition Based on Short-Term Sampling Neural Networks","volume":"8","author":"Zhang","year":"2020","journal-title":"IEEECAA J. Autom. Sin."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"112829","DOI":"10.1016\/j.eswa.2019.112829","article-title":"MultiD-CNN: A Multi-Dimensional Feature Learning Approach Based on Deep Convolutional Networks for Gesture Recognition in RGB-D Image Sequences","volume":"139","author":"Elboushaki","year":"2020","journal-title":"Expert Syst. Appl."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"5626","DOI":"10.1109\/TIP.2021.3087348","article-title":"Searching Multi-Rate and Multi-Modal Temporal Enhanced Networks for Gesture Recognition","volume":"30","author":"Yu","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"7689","DOI":"10.1109\/TIP.2021.3108349","article-title":"TMMF: Temporal Multi-Modal Fusion for Single-Stage Continuous Gesture Recognition","volume":"30","author":"Gammulle","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"107356","DOI":"10.1016\/j.patcog.2020.107356","article-title":"SGM-Net: Skeleton-Guided Multimodal Network for Action Recognition","volume":"104","author":"Li","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1038","DOI":"10.1109\/TMM.2018.2808769","article-title":"EgoGesture: A New Dataset and Benchmark for Egocentric Hand Gesture Recognition","volume":"20","author":"Zhang","year":"2018","journal-title":"IEEE Trans. Multimed."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Molchanov, P., Yang, X., Gupta, S., Kim, K., Tyree, S., and Kautz, J. (2016, January 27\u201330). Online Detection and Classification of Dynamic Hand Gestures with Recurrent 3d Convolutional Neural Network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.456"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Materzynska, J., Berger, G., Bax, I., and Memisevic, R. (2019, January 27\u201328). The Jester Dataset: A Large-Scale Video Dataset of Human Gestures. Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops, Seoul, Republic of Korea.","DOI":"10.1109\/ICCVW.2019.00349"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_34","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention Is All You Need. Adv. Neural Inf. Process. Syst., 30."},{"key":"ref_35","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., and He, K. (November, January 27). Slowfast Networks for Video Recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_36","unstructured":"Simonyan, K., and Zisserman, A. (2015). Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1007\/978-3-319-46484-8_2","article-title":"Temporal Segment Networks: Towards Good Practices for Deep Action Recognition","volume":"Volume 9912","author":"Leibe","year":"2016","journal-title":"Computer Vision\u2014ECCV 2016"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning Spatiotemporal Features with 3d Convolutional Networks. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Hara, K., Kataoka, H., and Satoh, Y. (2017, January 22\u201329). Learning Spatio-Temporal Features with 3d Residual Networks for Action Recognition. Proceedings of the IEEE International Conference on Computer Vision Workshops, Venice, Italy.","DOI":"10.1109\/ICCVW.2017.373"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Carreira, J., and Zisserman, A. (2017, January 21\u201326). Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.502"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Abavisani, M., Joze, H.R.V., and Patel, V.M. (2019, January 15\u201320). Improving the Performance of Unimodal Dynamic Hand-Gesture Recognition with Multimodal Training. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00126"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Li, Y., Ji, B., Shi, X., Zhang, J., Kang, B., and Wang, L. (2020, January 13\u201319). Tea: Temporal Excitation and Aggregation for Action Recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00099"},{"key":"ref_43","unstructured":"Lin, J., Gan, C., and Han, S. (November, January 27). Tsm: Temporal Shift Module for Efficient Video Understanding. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C. (2020, January 13\u201319). X3d: Expanding Architectures for Efficient Video Recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00028"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Wang, Z., She, Q., and Smolic, A. (2021, January 20\u201325). Action-Net: Multipath Excitation for Action Recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01301"}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/14\/10\/432\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,16]],"date-time":"2025-10-16T04:40:54Z","timestamp":1760589654000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/14\/10\/432"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,13]]},"references-count":45,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2025,10]]}},"alternative-id":["computers14100432"],"URL":"https:\/\/doi.org\/10.3390\/computers14100432","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,13]]}}}