{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,5]],"date-time":"2025-11-05T14:01:37Z","timestamp":1762351297611,"version":"build-2065373602"},"reference-count":45,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2020,10,6]],"date-time":"2020-10-06T00:00:00Z","timestamp":1601942400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Recent progress on skeleton-based action recognition has been substantial, benefiting mostly from the explosive development of Graph Convolutional Networks (GCN). However, prevailing GCN-based methods may not effectively capture the global co-occurrence features among joints and the local spatial structure features composed of adjacent bones. They also ignore the effect of channels unrelated to action recognition on model performance. Accordingly, to address these issues, we propose a Global Co-occurrence feature and Local Spatial feature learning model (GCLS) consisting of two branches. The first branch, based on the Vertex Attention Mechanism branch (VAM-branch), captures the global co-occurrence feature of actions effectively; the second, based on the Cross-kernel Feature Fusion branch (CFF-branch), extracts local spatial structure features composed of adjacent bones and restrains the channels unrelated to action recognition. Extensive experiments on two large-scale datasets, NTU-RGB+D and Kinetics, demonstrate that GCLS achieves the best performance when compared to the mainstream approaches.<\/jats:p>","DOI":"10.3390\/e22101135","type":"journal-article","created":{"date-parts":[[2020,10,7]],"date-time":"2020-10-07T09:12:58Z","timestamp":1602061978000},"page":"1135","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Global Co-Occurrence Feature and Local Spatial Feature Learning for Skeleton-Based Action Recognition"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4359-1045","authenticated-orcid":false,"given":"Jun","family":"Xie","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6479-3092","authenticated-orcid":false,"given":"Wentian","family":"Xin","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6882-9566","authenticated-orcid":false,"given":"Ruyi","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qiguang","family":"Miao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lijie","family":"Sheng","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liang","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xuesong","family":"Gao","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Digital Multimedia Technology, Hisense Co., Ltd., Qingdao 266071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,10,6]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Gaur, U., Zhu, Y., Song, B., and Roy-Chowdhury, A. (2011, January 6\u201313). A \u201cstring of feature graphs\u201d model for recognition of complex activities in natural videos. Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126548"},{"key":"ref_2","first-page":"1","article-title":"Approaches and applications of virtual reality and gesture recognition: A review","volume":"8","author":"Sudha","year":"2017","journal-title":"IJACI"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1272","DOI":"10.1109\/JPROC.2002.801449","article-title":"Integrating perceptual and cognitive modeling for adaptive and intelligent human-computer interaction","volume":"90","author":"Duric","year":"2002","journal-title":"Proc. IEEE"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., and Lin, D. (2018). Spatial temporal graph convolutional networks for skeleton-based action recognition. arXiv.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Li, Q., Han, Z., and Wu, X.M. (2018). Deeper insights into graph convolutional networks for semi-supervised learning. arXiv.","DOI":"10.1609\/aaai.v32i1.11604"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Li, M., Chen, S., Chen, X., Zhang, Y., Wang, Y., and Tian, Q. (2019). Actional-Structural Graph Convolutional Networks for Skeleton-Based Action Recognition. CVPR, IEEE.","DOI":"10.1109\/CVPR.2019.00371"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (2019). Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition. CVPR, IEEE.","DOI":"10.1109\/CVPR.2019.01230"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Li, C., Zhong, Q., Xie, D., and Pu, S. (2018). Co-occurrence feature learning from skeleton data for action recognition and detection with hierarchical aggregation. arXiv.","DOI":"10.24963\/ijcai.2018\/109"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Vemulapalli, R., Arrate, F., and Chellappa, R. (2014, January 23\u201328). Human action recognition by representing 3d skeletons as points in a lie group. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.82"},{"key":"ref_10","unstructured":"Wang, J., Liu, Z., Wu, Y., and Yuan, J. (2012, January 16\u201321). Mining actionlet ensemble for action recognition with depth cameras. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA."},{"key":"ref_11","unstructured":"Hussein, M.E., Torki, M., Gowayyed, M.A., and El-Saban, M. (2013, January 3\u20139). Human action recognition using a temporal hierarchy of covariance descriptors on 3d joint locations. Proceedings of the Twenty-Third International Joint Conference on Artificial Intelligence, Beijing, China."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Song, S., Lan, C., Xing, J., Zeng, W., and Liu, J. (2018). An end-to-end spatio-temporal attention model for human action recognition from skeleton data. arXiv.","DOI":"10.1609\/aaai.v31i1.11212"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zhu, W., Lan, C., Xing, J., Zeng, W., Li, Y., Shen, L., and Xie, X. (2016). Co-occurrence feature learning for skeleton based action recognition using regularized deep lstm networks. arXiv.","DOI":"10.1609\/aaai.v30i1.10451"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Liu, J., Shahroudy, A., Xu, D., and Wang, G. (2016). Spatio-temporal lstm with trust gates for 3d human action recognition. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46487-9_50"},{"key":"ref_15","unstructured":"Du, Y., Wang, W., and Wang, L. (2015, January 7\u201315). Hierarchical recurrent neural network for skeleton based action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_16","unstructured":"Li, L., Zheng, W., Zhang, Z., Huang, Y., and Wang, L. (2018). Skeleton-Based Relational Modeling for Action Recognition. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Shahroudy, A., Liu, J., Ng, T.T., and Wang, G. (2016, January 27\u201330). NTU RGB+D: A Large Scale Dataset for 3d Human Activity Analysis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.115"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"3247","DOI":"10.1109\/TCSVT.2018.2879913","article-title":"Skeleton-Based Action Recognition with Gated Convolutional Neural Networks","volume":"29","author":"Cao","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Sedmidubsky, J., and Zezula, P. (2019, January 9\u201311). Augmenting Spatio-Temporal Human Motion Data for Effective 3D Action Recognition. Proceedings of the 2019 IEEE International Symposium on Multimedia (ISM), San Diego, CA, USA.","DOI":"10.1109\/ISM46123.2019.00044"},{"key":"ref_20","first-page":"90","article-title":"Conflux LSTMs Network: A Novel Approach for Multi-View Action Recognition","volume":"414","author":"Ullah","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Ullah, A., Muhammad, K., and Hussain, T. (2020). Deep LSTM-Based Sequence Learning Approaches for Action and Activity Recognition. Deep Learning in Computer Vision: Principles and Applications, CRC Press.","DOI":"10.1201\/9781351003827-5"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"346","DOI":"10.1016\/j.patcog.2017.02.030","article-title":"Enhanced skeleton visualization for view invariant human action recognition","volume":"68","author":"Liu","year":"2017","journal-title":"Pattern Recognit."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Ke, Q., Bennamoun, M., An, S., Sohel, F., and Boussaid, F. (2017, January 21\u201326). A New Representation of Skeleton Sequences for 3d Action Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.486"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Kim, T.S., and Reiter, A. (2017, January 21\u201326). Interpretable 3d human action analysis with temporal convolutional networks. Proceedings of the 2017 IEEE conference on Computer Vvision and Pattern Recognition Workshops (CVPRW), Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.207"},{"key":"ref_25","unstructured":"Liu, H., Tu, J., and Liu, M. (2017). Two-Stream 3d Convolutional Neural Network for Skeleton-Based Action Recognition. arXiv."},{"key":"ref_26","unstructured":"Li, B., Dai, Y., Cheng, X., Chen, H., Lin, Y., and He, M. (2017, January 10\u201314). Skeleton based action recognition using translation-scale invariant image mapping and multi-scale deep CNN. Proceedings of the 2017 IEEE International Conference on Multimedia & Expo Workshops (ICMEW), Hong Kong, China."},{"key":"ref_27","unstructured":"Li, C., Zhong, Q., Xie, D., and Pu, S. (2017, January 10\u201314). Skeleton-based action recognition with convolutional neural networks. Proceedings of the 2017 IEEE International Conference on Multimedia & Expo Workshops (ICMEW), Hong Kong, China."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Du, Y., Fu, Y., and Wang, L. (2015, January 3\u20136). Skeleton based action recognition with convolutional neural network. Proceedings of the 2015 3rd IAPR Asian Conference on Pattern Recognition, Kuala Lumpur, Malaysia.","DOI":"10.1109\/ACPR.2015.7486569"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"12073","DOI":"10.1007\/s11042-017-4859-7","article-title":"Effective and efficient similarity searching in motion capture data","volume":"77","author":"Sedmidubsky","year":"2018","journal-title":"Multimed. Tools Appl."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"386","DOI":"10.1016\/j.future.2019.01.029","article-title":"Action recognition using optimized deep autoencoder and CNN for surveillance data streams of non-stationary environments","volume":"96","author":"Ullah","year":"2019","journal-title":"Future Gener. Comput. Syst."},{"key":"ref_31","unstructured":"Veli\u010dkovi\u0107, P., Cucurull, G., Casanova, A., Romero, A., Lio, P., and Bengio, Y. (2017). Graph attention networks. arXiv."},{"key":"ref_32","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, MIT Press."},{"key":"ref_33","unstructured":"Sankar, A., Wu, Y., Gou, L., Zhang, W., and Yang, H. (2019). Dynamic graph representation learning via self-attention networks. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wang, Q., Wu, B., Zhu, P., Li, P., Zuo, W., and Hu, Q. (2019, January 16\u201320). ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR42600.2020.01155"},{"key":"ref_35","unstructured":"Kipf, T.N., and Welling, M. (2016). Semi-supervised classification with graph convolutional networks. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"83","DOI":"10.1007\/s00530-019-00635-7","article-title":"Recent evolution of modern datasets for human activity recognition: A deep survey","volume":"26","author":"Singh","year":"2020","journal-title":"Multimedia Syst."},{"key":"ref_37","unstructured":"Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Fabio, V., Tim, G., Trevor, B., and Paul, N. (2017). The Kinetics Human Action Video Dataset. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Cao, Z., Simon, T., Wei, S.E., and Sheikh, Y. (2017, January 21\u201326). Realtime multiperson 2d pose estimation using part affinity fields. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.143"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Zhang, P., Lan, C., Xing, J., Zeng, W., Xue, J., and Zheng, N. (2017, January 21\u201326). View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition From Skeleton Data. Proceedings of the IEEE International Conference on Computer Vision, Honolulu, HI, USA.","DOI":"10.1109\/ICCV.2017.233"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Li, S., Li, W., Cook, C., Zhu, C., and Gao, Y. (2018, January 18\u201322). Independently recurrent neural network (indrnn): Building A longer and deeper RNN. Proceedings of the IEEE Conference on Computer Vvision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00572"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (2019, January 16\u201320). Skeleton-based action recognition with directed graph neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00810"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Peng, W., Hong, X., Chen, H., and Zhao, G. (2020). Learning graph convolutional network for skeleton-based human action recognition by neural searching. AAAI, MIT Press.","DOI":"10.1609\/aaai.v34i03.5652"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Cheng, K., Zhang, Y., He, X., Chen, W., Cheng, J., and Lu, H. (2020, January 13\u201319). Skeleton-based action recognition with shift graph convolutional network. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00026"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Fernando, B., Gavves, E., Oramas, J.M., Ghodrati, A., and Tuytelaars, T. (2015, January 7\u201312). Modeling video evolution for action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299176"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Tang, Y., Tian, Y., Lu, J., Li, P., and Zhou, J. (2018, January 18\u201322). Deep progressive reinforcement learning for skeleton-based action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00558"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/22\/10\/1135\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:17:00Z","timestamp":1760177820000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/22\/10\/1135"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,6]]},"references-count":45,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2020,10]]}},"alternative-id":["e22101135"],"URL":"https:\/\/doi.org\/10.3390\/e22101135","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2020,10,6]]}}}