{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T18:31:25Z","timestamp":1783103485425,"version":"3.54.6"},"reference-count":38,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T00:00:00Z","timestamp":1778803200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272207"],"award-info":[{"award-number":["62272207"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Jiangxi Natural Science Foundation Project","award":["20224ACB202009"],"award-info":[{"award-number":["20224ACB202009"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>Dynamic Hand Gesture Recognition (DHGR) aims to detect dynamic hand movements by leveraging the features and continuity of video frames. Existing methods mainly utilize backbone networks to extract latent features from individual video frames and sequence modeling through a Transformer. However, the hand usually occupies a relatively small proportion in the video, resulting in a large amount of invalid information in the extracted features, which affects the model\u2019s robustness and subsequent temporal modeling performance. Moreover, the traditional Transformer structure has a high time complexity, which affects the model\u2019s operational efficiency. To address these issues, we propose a novel data preprocessing and data fusion approach. It filters the hand contour using the motion vector of video coding and extracts features from RGB images and contour images through a dual-stream network. Additionally, a Gated-MLP GCN (GM-GCN) fusion module is proposed to fully fuse the dual-stream features. Meanwhile, we developed an Efficient Multi-scale Recurrent Attention (EMRA) module as our temporal modeling network, which adopts a recurrent structure similar to an RNN for attention calculation, enabling efficient parallel training of the model. Moreover, it better captures the details and dynamic changes of gestures through wavelet transform and multi-scale pooling strategies. Extensive experiments demonstrate that our proposed framework achieves highly competitive results on key benchmarks (e.g., 83.87% accuracy on NVGesture) while ensuring computational efficiency, reducing MACs by 28% compared to the standard Transformer.<\/jats:p>","DOI":"10.1145\/3797044","type":"journal-article","created":{"date-parts":[[2026,3,17]],"date-time":"2026-03-17T20:19:13Z","timestamp":1773778753000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["MDRA: A Motion-guided Dual-stream Recurrent Attention Framework for Dynamic Hand Gesture Recognition"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0629-165X","authenticated-orcid":false,"given":"Guoqiong","family":"Liao","sequence":"first","affiliation":[{"name":"Modern Industry School of Virtual Reality (VR), Jiangxi University of Finance and Economics, Nanchang, China and Jiangxi Tourism and Commerce Vocational College, Nanchang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-3642-0451","authenticated-orcid":false,"given":"Longjie","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Computer and Artificial Intelligence, Jiangxi University of Finance and Economics, Nanchang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-9721-5740","authenticated-orcid":false,"given":"Kefan","family":"Chen","sequence":"additional","affiliation":[{"name":"Modern Industry School of Virtual Reality (VR), Jiangxi University of Finance and Economics, Nanchang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1706-525X","authenticated-orcid":false,"given":"Yong","family":"Gu","sequence":"additional","affiliation":[{"name":"School of Internet of Things and Artificial Intelligence, Jiangxi University of Finance and Economics, Nanchang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-1187-3178","authenticated-orcid":false,"given":"Tao","family":"Zhu","sequence":"additional","affiliation":[{"name":"Modern Industry School of Virtual Reality (VR), Jiangxi University of Finance and Economics, Nanchang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,5,15]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"623","volume-title":"Proceedings of the International Conference on 3D Vision","author":"D\u2019Eusanio A.","year":"2020","unstructured":"A. D\u2019Eusanio, A. Simoni, S. Pini, G. Borghi, R. Vezzani, and R. Cucchiara. 2020. A transformer-based network for dynamic hand gesture recognition. In Proceedings of the International Conference on 3D Vision. IEEE, 623\u2013632."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW63382.2024.00254"},{"key":"e_1_3_1_4_2","first-page":"6156","volume-title":"Proceedings of the IEEE Winter Conference on Applications of Computer Vision","author":"Garg M.","year":"2025","unstructured":"M. Garg, D. Ghosh, and P. M. Pradhan. 2025. ConvMixFormer\u2014A resource-efficient convolution mixer for transformer-based dynamic hand gesture recognition. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision. IEEE, 6156\u20136166."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.291"},{"key":"e_1_3_1_7_2","first-page":"7983","volume-title":"Proceedings of the International Conference on Robotics and Automation","author":"Chang J.-Y.","year":"2019","unstructured":"J.-Y. Chang, A. Tejero-de Pablos, and T. Harada. 2019. Improved optical flow for gesture-based human-robot interaction. In Proceedings of the International Conference on Robotics and Automation. IEEE, 7983\u20137989."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00718"},{"issue":"5","key":"e_1_3_1_9_2","doi-asserted-by":"crossref","first-page":"e12490","DOI":"10.1111\/exsy.12490","article-title":"Hand gesture recognition using multimodal data fusion and multiscale parallel convolutional neural network for human\u2013robot interaction","volume":"38","author":"Gao Q.","year":"2021","unstructured":"Q. Gao, J. Liu, and Z. Ju. 2021. Hand gesture recognition using multimodal data fusion and multiscale parallel convolutional neural network for human\u2013robot interaction. Expert Systems 38, 5 (2021), e12490.","journal-title":"Expert Systems"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cag.2021.04.017"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.456"},{"key":"e_1_3_1_12_2","first-page":"560","volume-title":"Proceedings of the International Conference on Image Analysis and Processing","author":"Manganaro F.","year":"2019","unstructured":"F. Manganaro, S. Pini, G. Borghi, R. Vezzani, and R. Cucchiara. 2019. Hand gestures for the human-car interaction: The Briareo dataset. In Proceedings of the International Conference on Image Analysis and Processing. Springer, 560\u2013571."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2808769"},{"key":"e_1_3_1_14_2","first-page":"3763","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Cao C.","year":"2017","unstructured":"C. Cao, Y. Zhang, Y. Wu, H. Lu, and J. Cheng. 2017. Egocentric gesture recognition using recurrent 3D convolutional neural networks with spatiotemporal transformer modules. In Proceedings of the IEEE International Conference on Computer Vision, 3763\u20133771."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.29322\/IJSRP.9.10.2019.p9420"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.12.022"},{"key":"e_1_3_1_19_2","first-page":"922","volume-title":"Proceedings of the International Conference on Intelligent Robots and Systems","author":"Maturana D.","year":"2015","unstructured":"D. Maturana and S. Scherer. 2015. VoxNet: A 3D convolutional neural network for real-time object recognition. In Proceedings of the International Conference on Intelligent Robots and Systems, 922\u2013928."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3028207"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.physd.2019.132306"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2024.3365274"},{"issue":"6","key":"e_1_3_1_23_2","doi-asserted-by":"crossref","first-page":"16275","DOI":"10.1007\/s11042-023-16130-1","article-title":"Real-time continuous detection and recognition of dynamic hand gestures in untrimmed sequences based on end-to-end architecture with 3D DenseNet and LSTM","volume":"83","author":"Lu Z.","year":"2024","unstructured":"Z. Lu, S. Qin, P. Lv, L. Sun, and B. Tang. 2024. Real-time continuous detection and recognition of dynamic hand gestures in untrimmed sequences based on end-to-end architecture with 3D DenseNet and LSTM. Multimedia Tools and Applications 83, 6 (2024), 16275\u201316312.","journal-title":"Multimedia Tools and Applications"},{"issue":"8","key":"e_1_3_1_24_2","first-page":"328","article-title":"Real-time LSTM-based multi-dimensional features gesture recognition","volume":"48","year":"2021","unstructured":"L. Liu, and H. -Y. Pu. 2021. Real-time LSTM-based multi-dimensional features gesture recognition. Computer Science 48, 8 (2021), 328\u2013333","journal-title":"Computer Science"},{"key":"e_1_3_1_25_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani A.","year":"2017","unstructured":"A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, \u0141. Kaiser, and I. Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"issue":"4","key":"e_1_3_1_26_2","doi-asserted-by":"crossref","first-page":"2041","DOI":"10.3390\/app12042041","article-title":"Content-adaptive and attention-based network for hand gesture recognition","volume":"12","author":"Cao Z.","year":"2022","unstructured":"Z. Cao, Y. Li, and B.-S. Shin. 2022. Content-adaptive and attention-based network for hand gesture recognition. Applied Sciences 12, 4 (2022), 2041.","journal-title":"Applied Sciences"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2023.107288"},{"key":"e_1_3_1_28_2","first-page":"10012","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Liu Z.","year":"2021","unstructured":"Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 10012\u201310022."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2023.3241857"},{"key":"e_1_3_1_30_2","first-page":"12021","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Chen J.","year":"2023","unstructured":"J. Chen, S-h Kao, H. He, W. Zhuo, S. Wen, C.-H. Lee, and S.-H. G. Chan. 2023. Run, don\u2019t walk: Chasing higher flops for faster neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 12021\u201312031."},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00677"},{"key":"e_1_3_1_33_2","first-page":"15","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Garg M.","year":"2024","unstructured":"M. Garg, D. Ghosh, and P. M. Pradhan. 2024. MVTN: A multiscale video transformer network for hand gesture recognition. In Proceedings of the European Conference on Computer Vision. Springer, 15\u201333."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0846-5"},{"key":"e_1_3_1_35_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition","author":"K\u00f6p\u00fckl\u00fc O.","year":"2019","unstructured":"O. K\u00f6p\u00fckl\u00fc, A. Gunduz, N. Kose, and G. Rigoll. 2019. Real-time hand gesture detection and classification using convolutional neural networks. In Proceedings of the IEEE International Conference on Automatic Face and Gesture Recognition. IEEE, 1\u20138."},{"key":"e_1_3_1_36_2","first-page":"1165","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Abavisani M.","year":"2019","unstructured":"M. Abavisani, H. R. V. Joze, and V. M. Patel. 2019. Improving the performance of unimodal dynamic hand-gesture recognition with multimodal training. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1165\u20131174."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3087348"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.3390\/informatics7030031"},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","first-page":"230","DOI":"10.1007\/978-3-030-60639-8_20","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang Z.","year":"2020","unstructured":"Z. Wang, Q. She, T. Chalasani, and A. Smolic. 2020. CatNet: Class incremental 3D ConvNets for lifelong egocentric gesture recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 230\u2013231."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3797044","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T07:52:37Z","timestamp":1778831557000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3797044"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,15]]},"references-count":38,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3797044"],"URL":"https:\/\/doi.org\/10.1145\/3797044","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,15]]},"assertion":[{"value":"2025-09-07","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-18","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}