{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,16]],"date-time":"2026-02-16T15:53:27Z","timestamp":1771257207195,"version":"3.50.1"},"reference-count":40,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2023,6,7]],"date-time":"2023-06-07T00:00:00Z","timestamp":1686096000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Research and Application of edge computing Technology Based on TinyML","award":["KYP0222010"],"award-info":[{"award-number":["KYP0222010"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Graph convolutional networks are widely used in skeleton-based action recognition because of their good fitting ability to non-Euclidean data. While conventional multi-scale temporal convolution uses several fixed-size convolution kernels or dilation rates at each layer of the network, we argue that different layers and datasets require different receptive fields. We use multi-scale adaptive convolution kernels and dilation rates to optimize traditional multi-scale temporal convolution with a simple and effective self attention mechanism, allowing different network layers to adaptively select convolution kernels of different sizes and dilation rates instead of being fixed and unchanged. Besides, the effective receptive field of the simple residual connection is not large, and there is a great deal of redundancy in the deep residual network, which will lead to the loss of context when aggregating spatio-temporal information. This article introduces a feature fusion mechanism that replaces the residual connection between initial features and temporal module outputs, effectively solving the problems of context aggregation and initial feature fusion. We propose a multi-modality adaptive feature fusion framework (MMAFF) to simultaneously increase the receptive field in both spatial and temporal dimensions. Concretely, we input the features extracted by the spatial module into the adaptive temporal fusion module to simultaneously extract multi-scale skeleton features in both spatial and temporal parts. In addition, based on the current multi-stream approach, we use the limb stream to uniformly process correlated data from multiple modalities. Extensive experiments show that our model obtains competitive results with state-of-the-art methods on the NTU-RGB+D 60 and NTU-RGB+D 120 datasets.<\/jats:p>","DOI":"10.3390\/s23125414","type":"journal-article","created":{"date-parts":[[2023,6,8]],"date-time":"2023-06-08T02:02:28Z","timestamp":1686189748000},"page":"5414","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Multi-Modality Adaptive Feature Fusion Graph Convolutional Network for Skeleton-Based Action Recognition"],"prefix":"10.3390","volume":"23","author":[{"given":"Haiping","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Computer Science, Hangzhou Dianzi University, Hangzhou 310005, China"},{"name":"School of Information Engineering, Hangzhou Dianzi University, Hangzhou 310005, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-9342-7474","authenticated-orcid":false,"given":"Xinhao","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Electronics and Information, Hangzhou Dianzi University, Hangzhou 310005, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dongjin","family":"Yu","sequence":"additional","affiliation":[{"name":"School of Computer Science, Hangzhou Dianzi University, Hangzhou 310005, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liming","family":"Guan","sequence":"additional","affiliation":[{"name":"School of Information Engineering, Hangzhou Dianzi University, Hangzhou 310005, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2152-0446","authenticated-orcid":false,"given":"Dongjing","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Hangzhou Dianzi University, Hangzhou 310005, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fuxing","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Electronics and Information, Hangzhou Dianzi University, Hangzhou 310005, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-0283-2407","authenticated-orcid":false,"given":"Wanjun","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Information Engineering, Hangzhou Dianzi University, Hangzhou 310005, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,6,7]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., and Lin, D. (2018, January 2\u20137). Spatial temporal graph convolutional networks for skeleton-based action recognition. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, FL, USA.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Li, M., Chen, S., Chen, X., Zhang, Y., Wang, Y., and Tian, Q. (2019, January 15\u201320). Actional-structural graph convolutional networks for skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00371"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (2019, January 15\u201320). Two-stream adaptive graph convolutional networks for skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01230"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Chen, Y., Zhang, Z., Yuan, C., Li, B., Deng, Y., and Hu, W. (2021, January 11\u201317). Channel-wise topology refinement graph convolution for skeleton-based action recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.01311"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (2019, January 15\u201320). Skeleton-based action recognition with directed graph neural networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00810"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Liu, Z., Zhang, H., Chen, Z., Wang, Z., and Ouyang, W. (2020, January 13\u201319). Disentangling and unifying graph convolutions for skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00022"},{"key":"ref_7","unstructured":"Yu, F., and Koltun, V. (2015). Multi-scale context aggregation by dilated convolutions. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Chi, H.g., Ha, M.H., Chi, S., Lee, S.W., Huang, Q., and Ramani, K. (2022, January 18\u201324). Infogcn: Representation learning for human skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01955"},{"key":"ref_9","unstructured":"Lee, J., Lee, M., Lee, D., and Lee, S. (2022). Hierarchically Decomposed Graph Convolutional Networks for Skeleton-Based Action Recognition. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Ke, L., Peng, K.C., and Lyu, S. (2022, January 20\u201327). Towards to-at spatio-temporal focus for skeleton-based action recognition. Proceedings of the AAAI Conference on Artificial Intelligence, Montreal, BC, Canada.","DOI":"10.1609\/aaai.v36i1.19998"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Cheng, K., Zhang, Y., He, X., Chen, W., Cheng, J., and Lu, H. (2020, January 13\u201319). Skeleton-based action recognition with shift graph convolutional network. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00026"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Dong, J., Gao, Y., Lee, H.J., Zhou, H., Yao, Y., Fang, Z., and Huang, B. (2020). Action recognition based on the fusion of graph convolutional networks with high order features. Appl. Sci., 10.","DOI":"10.3390\/app10041482"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Qin, Z., Liu, Y., Ji, P., Kim, D., Wang, L., McKay, B., Anwar, S., and Gedeon, T. (2021). Fusing higher-order features in graph neural networks for skeleton-based action recognition. arXiv.","DOI":"10.1109\/TNNLS.2022.3201518"},{"key":"ref_14","unstructured":"Trivedi, N., and Sarvadevabhatla, R.K. (2023). Proceedings, Part V, Proceedings of the Computer Vision\u2014ECCV 2022 Workshops, Tel Aviv, Israel, 23\u201327 October 2022, Springer."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Song, Y.F., Zhang, Z., Shan, C., and Wang, L. (2020, January 12\u201316). Stronger, faster and more explainable: A graph convolutional baseline for skeleton-based action recognition. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3413802"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Ch\u00e9ron, G., Laptev, I., and Schmid, C. (2015, January 7\u201313). P-cnn: Pose-based cnn features for action recognition. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.368"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"346","DOI":"10.1016\/j.patcog.2017.02.030","article-title":"Enhanced skeleton visualization for view invariant human action recognition","volume":"68","author":"Liu","year":"2017","journal-title":"Pattern Recognit."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Liu, J., Wang, G., Hu, P., Duan, L.Y., and Kot, A.C. (2017, January 21\u201326). Global context-aware attention lstm networks for 3d action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.391"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Duan, H., Zhao, Y., Chen, K., Lin, D., and Dai, B. (2022, January 18\u201324). Revisiting skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00298"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Song, S., Lan, C., Xing, J., Zeng, W., and Liu, J. (2017, January 4\u20139). An end-to-end spatio-temporal attention model for human action recognition from skeleton data. Proceedings of the AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.11212"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"103219","DOI":"10.1016\/j.cviu.2021.103219","article-title":"Skeleton-based action recognition via spatial and temporal transformer networks","volume":"208","author":"Plizzari","year":"2021","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_22","unstructured":"Cheng, K., Zhang, Y., Cao, C., Shi, L., Cheng, J., and Lu, H. (2020). Proceedings, Part XXIV 16, Proceedings of the Computer Vision\u2014ECCV 2020: 16th European Conference, Glasgow, UK, 23\u201328 August 2020, Springer."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Ye, F., Pu, S., Zhong, Q., Li, C., Xie, D., and Tang, H. (2020, January 16\u201318). Dynamic gcn: Context-enriched topology learning for skeleton-based action recognition. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3413941"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Qiu, H., Hou, B., Ren, B., and Zhang, X. (2022). Spatio-temporal tuples transformer for skeleton-based action recognition. arXiv.","DOI":"10.1016\/j.neucom.2022.10.084"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1474","DOI":"10.1109\/TPAMI.2022.3157033","article-title":"Constructing stronger and faster baselines for skeleton-based action recognition","volume":"45","author":"Song","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zhou, S.B., Chen, R.R., Jiang, X.Q., and Pan, F. (2023). 2s-GATCN: Two-Stream Graph Attentional Convolutional Networks for Skeleton-Based Action Recognition. Electronics, 12.","DOI":"10.3390\/electronics12071711"},{"key":"ref_27","unstructured":"Wang, S., Zhang, Y., Wei, F., Wang, K., Zhao, M., and Jiang, Y. (2022). Skeleton-based Action Recognition via Temporal-Channel Aggregation. arXiv."},{"key":"ref_28","unstructured":"Xu, K., Ye, F., Zhong, Q., and Xie, D. (March, January 22). Topology-aware convolutional neural network for efficient skeleton-based action recognition. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"9532","DOI":"10.1109\/TIP.2020.3028207","article-title":"Skeleton-based action recognition with multi-stream adaptive graph convolutional networks","volume":"29","author":"Shi","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Wang, Q., Wu, B., Zhu, P., Li, P., Zuo, W., and Hu, Q. (2020, January 13\u201319). ECA-Net: Efficient channel attention for deep convolutional neural networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01155"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Dai, Y., Gieseke, F., Oehmcke, S., Wu, Y., and Barnard, K. (2021, January 5\u20139). Attentional feature fusion. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Virtual.","DOI":"10.1109\/WACV48630.2021.00360"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Shahroudy, A., Liu, J., Ng, T.T., and Wang, G. (2016, January 27\u201330). Ntu rgb+ d: A large scale dataset for 3d human activity analysis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.115"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"2684","DOI":"10.1109\/TPAMI.2019.2916873","article-title":"Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding","volume":"42","author":"Liu","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zhang, P., Lan, C., Zeng, W., Xing, J., Xue, J., and Zheng, N. (2020, January 13\u201319). Semantics-guided neural networks for efficient skeleton-based human action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00119"},{"key":"ref_35","unstructured":"Korban, M., and Li, X. (2020). Proceedings, Part XX 16, Proceedings of the Computer Vision\u2014ECCV 2020: 16th European Conference, Glasgow, UK, 23\u201328 August 2020, Springer."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Chen, Z., Li, S., Yang, B., Li, Q., and Liu, H. (2021, January 2\u20139). Multi-scale spatial temporal graph convolutional network for skeleton-based action recognition. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual.","DOI":"10.1609\/aaai.v35i2.16197"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Duan, H., Wang, J., Chen, K., and Lin, D. (2022, January 10\u201314). Pyskl: Towards good practices for skeleton action recognition. Proceedings of the 30th ACM International Conference on Multimedia, Lisboa, Portugal.","DOI":"10.1145\/3503161.3548546"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Zhou, H., Liu, Q., and Wang, Y. (2023). Learning Discriminative Representations for Skeleton Based Action Recognition. arXiv.","DOI":"10.1109\/CVPR52729.2023.01022"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"109231","DOI":"10.1016\/j.patcog.2022.109231","article-title":"SpatioTemporal focus for skeleton-based action recognition","volume":"136","author":"Wu","year":"2023","journal-title":"Pattern Recognit."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"109455","DOI":"10.1016\/j.patcog.2023.109455","article-title":"Relation-mining self-attention network for skeleton-based human action recognition","volume":"139","author":"Gedamu","year":"2023","journal-title":"Pattern Recognit."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/12\/5414\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:50:18Z","timestamp":1760125818000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/12\/5414"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,7]]},"references-count":40,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2023,6]]}},"alternative-id":["s23125414"],"URL":"https:\/\/doi.org\/10.3390\/s23125414","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,7]]}}}