{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T14:12:50Z","timestamp":1760710370217,"version":"build-2065373602"},"reference-count":43,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2020,9,24]],"date-time":"2020-09-24T00:00:00Z","timestamp":1600905600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Traditional convolution neural networks have achieved great success in human action recognition. However, it is challenging to establish effective associations between different human bone nodes to capture detailed information. In this paper, we propose a dual attention-guided multiscale dynamic aggregate graph convolution neural network (DAG-GCN) for skeleton-based human action recognition. Our goal is to explore the best correlation and determine high-level semantic features. First, a multiscale dynamic aggregate GCN module is used to capture important semantic information and to establish dependence relationships for different bone nodes. Second, the higher level semantic feature is further refined, and the semantic relevance is emphasized through a dual attention guidance module. In addition, we exploit the relationship of joints hierarchically and the spatial temporal correlations through two modules. Experiments with the DAG-GCN method result in good performance on the NTU-60-RGB+D and NTU-120-RGB+D datasets. The accuracy is 95.76% and 90.01%, respectively, for the cross (X)-View and X-Subon the NTU60dataset.<\/jats:p>","DOI":"10.3390\/sym12101589","type":"journal-article","created":{"date-parts":[[2020,9,25]],"date-time":"2020-09-25T01:39:33Z","timestamp":1600997973000},"page":"1589","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Dual Attention-Guided Multiscale Dynamic Aggregate Graph Convolutional Networks for Skeleton-Based Human Action Recognition"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3673-356X","authenticated-orcid":false,"given":"Zeyuan","family":"Hu","sequence":"first","affiliation":[{"name":"Department of Information Communication Engineering, Tongmyong University, Busan 48520, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eung-Joo","family":"Lee","sequence":"additional","affiliation":[{"name":"Department of Information Communication Engineering, Tongmyong University, Busan 48520, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,9,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Kerdvibulvech, C. (2019, January 26\u201331). A Review of Augmented Reality-Based Human-Computer Interaction Applications of Gesture-Based Interaction. Proceedings of the International Conference on Human-Computer Interaction, Orlando, FL, USA.","DOI":"10.1007\/978-3-030-30033-3_18"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"50400","DOI":"10.1109\/ACCESS.2019.2905641","article-title":"A Video Representation Method Based on Multi-view Structure Preserving Embedding for Action Retrieval","volume":"7","author":"Zhang","year":"2019","journal-title":"IEEE Access"},{"key":"ref_3","first-page":"142","article-title":"An end-to-end deep learning model for human activity recognition from highly sparse body sensor data in Internet of Medical Things environment","volume":"10","author":"Hassan","year":"2008","journal-title":"J. Supercomput."},{"key":"ref_4","first-page":"1","article-title":"Open Pose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields","volume":"99","author":"Cao","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Song, S., Lan, C., Xing, J., Zeng, W., and Liu, J. (2018, January 23\u201327). Skeleton-Indexed Deep Multi-Modal Feature Learning for High Performance Human Action Recognition. Proceedings of the 2018 IEEE International Conference on Multimedia and Expo (ICME), San Diego, CA, USA.","DOI":"10.1109\/ICME.2018.8486486"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1016\/j.cviu.2017.01.011","article-title":"Space-time representation of people based on 3D skeletal data","volume":"158","author":"Han","year":"2017","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Si, C., Jing, Y., Wang, W., Wang, L., and Tan, T. (2018, January 8\u201314). Skeleton-Based Action Recognition with Spatial Reasoning and Temporal Stack Learning. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01246-5_7"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zhang, P., Lan, C., Xing, J., Zeng, W., Xue, J., and Zheng, N. (2017, January 22\u201329). View adaptive recurrent neural networks for high performance human action recognition from skeleton data. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.233"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Huynh-The, T., Hua, C.H., and Kim, D.S. (2019, January 11\u201313). Learning Action Images Using Deep Convolutional Neural Networks For 3D Action Recognition. Proceedings of the IEEE Sensors Applications Symposium (SAS), Sophia Antipolis, France.","DOI":"10.1109\/SAS.2019.8705977"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Fan, H., Luo, C., Zeng, C., Ferianc, M., Que, Z., Liu, S., Niu, X., and Luk, W. (2019, January 15\u201317). F-E3D: FPGA-based Acceleration of an Efficient 3D Convolutional Neural Network for Human Action Recognition. Proceedings of the IEEE 30th International Conference on Application-Specific Systems, Architectures and Processors (ASAP), New York, NY, USA.","DOI":"10.1109\/ASAP.2019.00-44"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Wu, H., Ma, X., and Li, Y. (2019). Hierarchical dynamic depth projected difference images\u2013based action recognition in videos with convolutional neural networks. Int. J. Adv. Robot. Syst., 16.","DOI":"10.1177\/1729881418825093"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1963","DOI":"10.1109\/TPAMI.2019.2896631","article-title":"View Adaptive Neural Networks for High Performance Skeleton-based Human Action Recognition","volume":"41","author":"Zhang","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Cho, S., Maqbool, M., Liu, F., and Foroosh, H. (2019). Self-Attention Network for Skeleton-based Human Action Recognition. arXiv.","DOI":"10.1109\/WACV45572.2020.9093639"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"47220","DOI":"10.1109\/ACCESS.2020.2979549","article-title":"An End to End Framework with Adaptive Spatio-Temporal Attention Module for Human Action Recognition","volume":"8","author":"Liu","year":"2020","journal-title":"IEEE Access"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., and Lin, D. (2018). Spatial temporal graph convolutional networks for skeleton-based action recognition. arXiv.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"043032","DOI":"10.1117\/1.JEI.28.4.043032","article-title":"Attention module-based spatial-temporal graph convolutional networks for skeleton-based action recognition","volume":"28","author":"Kong","year":"2019","journal-title":"J. Electron. Imaging"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (2019, January 15\u201321). Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01230"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Liu, R., Xu, C., Zhang, T., Zhao, W., Cui, Z., and Yang, J. (2019, January 14\u201319). Si-GCN: Structure-induced Graph Convolution Network for Skeleton-based Action Recognition. Proceedings of the 2019 International Joint Conference on Neural Networks (IJCNN), Budapest, Hungary.","DOI":"10.1109\/IJCNN.2019.8851767"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Ding, X., Yang, K., and Chen, W. (2020, January 6\u20139). A Semantics-Guided Graph Convolutional Network for Skeleton-Based Action Recognition. Proceedings of the 2020 the 4th International Conference on Innovation in Artificial Intelligence, Xiamen, China.","DOI":"10.1145\/3390557.3394129"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Ke, Q., Bennamoun, M., An, S., Sohel, F., and Boussaid, F. (2017, January 21\u201326). A new representation of skeleton sequences for 3d action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.486"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Du, Y., Fu, Y., and Wang, L. (2015, January 3\u20136). Skeleton based action recognition with convolutional neural network. Proceedings of the 2015 3rd IAPR Asian Conference on Pattern Recognition (ACPR), Kuala Lumpur, Malaysia.","DOI":"10.1109\/ACPR.2015.7486569"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Shahroudy, A., Liu, J., Ng, T.T., and Wang, G. (2016, January 27\u201330). Ntu rgb+ d: A large scale dataset for 3d human activity analysis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.115"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Liu, J., Shahroudy, A., Xu, D., and Wang, G. (2016). Spatio-Temporal LSTM with Trust Gates for 3D Human Action Recognition. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46487-9_50"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"224","DOI":"10.1016\/j.neucom.2018.10.095","article-title":"Correlational Convolutional LSTM for Human Action Recognition","volume":"396","author":"Majd","year":"2019","journal-title":"Neurocomputing"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Gammulle, H., Denman, S., Sridharan, S., and Fookes, C. (2017, January 24\u201331). Two Stream LSTM: A Deep Fusion Framework for Human Action Recognition. Proceedings of the 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), Santa Rosa, CA, USA.","DOI":"10.1109\/WACV.2017.27"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"304","DOI":"10.1016\/j.neucom.2020.06.032","article-title":"Human action recognition using convolutional LSTM and fully-connected LSTM with different attentions","volume":"410","author":"Zhang","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Xiong, W., Wu, L., Alleva, F., Droppo, J., Huang, X., and Stolcke, A. (2018, January 15\u201320). The Microsoft 2017 Conversational Speech Recognition System. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461870"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., Gupta, A., and He, K. (2018, January 18\u201323). Non-local Neural Networks. Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"ref_29","first-page":"28","article-title":"Improved human action recognition approach based on two-stream convolutional neural network model","volume":"6","author":"Liu","year":"2020","journal-title":"Vis. Comput."},{"key":"ref_30","unstructured":"Torpey, D., and Celik, T. (2020). Human Action Recognition using Local Two-Stream Convolution Neural Network Features and Support Vector Machines. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Tang, Y., Tian, Y., Lu, J., Li, P., and Zhou, J. (2018, January 18\u201322). Deep Progressive Reinforcement Learning for Skeleton-Based Action Recognition. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00558"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"97757","DOI":"10.1109\/ACCESS.2020.2996779","article-title":"Multi-stream and Enhanced Spatial-temporal Graph Convolution Network for Skeleton-based Action Recognition","volume":"8","author":"Li","year":"2020","journal-title":"IEEE Access"},{"key":"ref_33","unstructured":"Shiraki, K., Hirakawa, T., Yamashita, T., and Fujiyoshi, H. Acquisition of Optimal Connection Patterns for Skeleton-based Action Recognition with Graph Convolutional Networks. Proceedings of the 15th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications\u2013Volume 5: VISAPP."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1016\/j.neucom.2020.04.145","article-title":"Multimodal Graph Convolutional Networks for High Quality Content Recognition","volume":"412","author":"Wang","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"621","DOI":"10.1007\/s00371-019-01644-3","article-title":"Skeleton-based action recognition by part-aware graph convolutional networks","volume":"36","author":"Qin","year":"2019","journal-title":"Vis. Comput."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"3499","DOI":"10.3390\/s20123499","article-title":"Centrality Graph Convolutional Networks for Skeleton-based Action Recognition","volume":"20","author":"Yang","year":"2020","journal-title":"Sensors"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Yang, K., Ding, X., and Chen, W. (2018, January 23\u201325). A Graph-Enhanced Convolution Network with Attention Gate for Skeleton Based Action Recognition. Proceedings of the ICCPR \u201919: 2019 8th International Conference on Computing and Pattern Recognition, Beijing, China.","DOI":"10.1145\/3373509.3373531"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Rashid, M., Kjellstrm, H., and Lee, Y.J. (2020, January 1\u20135). Action Graphs: Weakly-supervised Action Localization with Graph Convolution Networks. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), Snowmass Village, CO, USA.","DOI":"10.1109\/WACV45572.2020.9093404"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Lee, J., Jung, Y., and Kim, H. (2020). Dual Attention in Time and Frequency Domain for Voice Activity Detection. arXiv.","DOI":"10.21437\/Interspeech.2020-997"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., and Lu, H. (2019, January 15\u201321). Dual Attention Network for Scene Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00326"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Zhang, P., Xue, J., Lan, C., Zeng, W., Gao, Z., and Zheng, N. (2018, January 8\u201314). Adding attentiveness to the neurons in recurrent neural networks. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01240-3_9"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Li, M., Chen, S., Chen, X., Zhang, Y., Wang, Y., and Tian, Q. (2019, January 15\u201320). Actional-structural graph convolutional networks for skeleton-based action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00371"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Gao, X., Hu, W., Tang, J., Liu, J., and Guo, Z. (2019, January 21\u201325). Optimized skeleton-based action recognition via sparsified graph regression. Proceedings of the 27th ACM International Conference on Multimedia, Nice, France.","DOI":"10.1145\/3343031.3351170"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/12\/10\/1589\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:13:23Z","timestamp":1760177603000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/12\/10\/1589"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,24]]},"references-count":43,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2020,10]]}},"alternative-id":["sym12101589"],"URL":"https:\/\/doi.org\/10.3390\/sym12101589","relation":{},"ISSN":["2073-8994"],"issn-type":[{"type":"electronic","value":"2073-8994"}],"subject":[],"published":{"date-parts":[[2020,9,24]]}}}