{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T22:18:29Z","timestamp":1784585909139,"version":"3.55.0"},"reference-count":37,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2023,5,14]],"date-time":"2023-05-14T00:00:00Z","timestamp":1684022400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2021ZD0113601"],"award-info":[{"award-number":["2021ZD0113601"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["SGH22Y1233"],"award-info":[{"award-number":["SGH22Y1233"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2022ZDLSF07-07"],"award-info":[{"award-number":["2022ZDLSF07-07"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Educational Science Foundation of Shaanxi Province of China","award":["2021ZD0113601"],"award-info":[{"award-number":["2021ZD0113601"]}]},{"name":"Educational Science Foundation of Shaanxi Province of China","award":["SGH22Y1233"],"award-info":[{"award-number":["SGH22Y1233"]}]},{"name":"Educational Science Foundation of Shaanxi Province of China","award":["2022ZDLSF07-07"],"award-info":[{"award-number":["2022ZDLSF07-07"]}]},{"name":"Shaanxi Province Key Research and Development Program","award":["2021ZD0113601"],"award-info":[{"award-number":["2021ZD0113601"]}]},{"name":"Shaanxi Province Key Research and Development Program","award":["SGH22Y1233"],"award-info":[{"award-number":["SGH22Y1233"]}]},{"name":"Shaanxi Province Key Research and Development Program","award":["2022ZDLSF07-07"],"award-info":[{"award-number":["2022ZDLSF07-07"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Fitness yoga is now a popular form of national fitness and sportive physical therapy. At present, Microsoft Kinect, a depth sensor, and other applications are widely used to monitor and guide yoga performance, but they are inconvenient to use and still a little expensive. To solve these problems, we propose spatial\u2013temporal self-attention enhanced graph convolutional networks (STSAE-GCNs) that can analyze RGB yoga video data captured by cameras or smartphones. In the STSAE-GCN, we build a spatial\u2013temporal self-attention module (STSAM), which can effectively enhance the spatial\u2013temporal expression ability of the model and improve the performance of the proposed model. The STSAM has the characteristics of plug-and-play so that it can be applied in other skeleton-based action recognition methods and improve their performance. To prove the effectiveness of the proposed model in recognizing fitness yoga actions, we collected 960 fitness yoga action video clips in 10 action classes and built the dataset Yoga10. The recognition accuracy of the model on Yoga10 achieves 93.83%, outperforming the state-of-the-art methods, which proves that this model can better recognize fitness yoga actions and help students learn fitness yoga independently.<\/jats:p>","DOI":"10.3390\/s23104741","type":"journal-article","created":{"date-parts":[[2023,5,15]],"date-time":"2023-05-15T08:28:56Z","timestamp":1684139336000},"page":"4741","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":12,"title":["Spatial\u2013Temporal Self-Attention Enhanced Graph Convolutional Networks for Fitness Yoga Action Recognition"],"prefix":"10.3390","volume":"23","author":[{"given":"Guixiang","family":"Wei","sequence":"first","affiliation":[{"name":"School of Sports Center, Xi\u2019an Jiaotong University, Xi\u2019an 710000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huijian","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Software Engineering, Xi\u2019an Jiaotong University, Xi\u2019an 710000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Liping","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Sports Center, Xi\u2019an Jiaotong University, Xi\u2019an 710000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianji","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Artificial Intelligence and Robotics, Xi\u2019an Jiaotong University, Xi\u2019an 710000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,5,14]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"976","DOI":"10.1016\/j.imavis.2009.11.014","article-title":"A survey on vision-based human action recognition","volume":"28","author":"Poppe","year":"2010","journal-title":"Image Vis. Comput."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"224","DOI":"10.1016\/j.cviu.2010.10.002","article-title":"A survey of vision-based methods for action representation, segmentation and recognition","volume":"115","author":"Weinland","year":"2011","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"16387","DOI":"10.1007\/s00521-018-3951-x","article-title":"Human activity recognition via optical flow: Decomposing activities into basic actions","volume":"32","author":"Ladjailia","year":"2020","journal-title":"Neural Comput. Appl."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Lin, J., Gan, C., and Han, S. (2019, January 27\u201328). TSM: Temporal shift module for efficient video understanding. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00718"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Li, Y., Ji, B., Shi, X., Zhang, J., Kang, B., and Wang, L. (2020, January 13\u201319). TEA: Temporal excitation and aggregation for action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00099"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Wang, Z., She, Q., and Smolic, A. (2021, January 20\u201325). Action-net: Multipath excitation for action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01301"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., and He, K. (2019, January 27\u201328). Slowfast networks for video recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00630"},{"key":"ref_8","first-page":"381","article-title":"Two-stream convolutional networks for action recognition in videos","volume":"27","author":"Simonyan","year":"2014","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Pinz, A., and Zisserman, A. (2016, January 27\u201330). Convolutional two-stream network fusion for video action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.213"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhu, W., Lan, C., Xing, J., Zeng, W., Li, Y., Shen, L., and Xie, X. (2016, January 12\u201317). Co-occurrence feature learning for skeleton based action recognition using regularized deep LSTM networks. Proceedings of the AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA.","DOI":"10.1609\/aaai.v30i1.10451"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Li, M., Chen, S., Chen, X., Zhang, Y., Wang, Y., and Tian, Q. (2019, January 15\u201320). Actional-structural graph convolutional networks for skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00371"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (2019, January 15\u201320). Two-stream adaptive graph convolutional networks for skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01230"},{"key":"ref_13","unstructured":"Li, B., Dai, Y., Cheng, X., Chen, H., Lin, Y., and He, M. (2017, January 10\u201314). Skeleton based action recognition using translation-scale invariant image mapping and multi-scale deep CNN. Proceedings of the 2017 IEEE International Conference on Multimedia & Expo Workshops (ICMEW), Hong Kong, China."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., and Lin, D. (2018, January 2\u20137). Spatial temporal graph convolutional networks for skeleton-based action recognition. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1963","DOI":"10.1109\/TPAMI.2019.2896631","article-title":"View adaptive neural networks for high performance skeleton-based human action recognition","volume":"41","author":"Zhang","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Lee, I., Kim, D., Kang, S., and Lee, S. (2017, January 22\u201329). Ensemble deep learning for skeleton-based action recognition using temporal sliding lstm networks. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.115"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"50788","DOI":"10.1109\/ACCESS.2018.2869751","article-title":"Skeleton feature fusion based on multi-stream LSTM for action recognition","volume":"6","author":"Wang","year":"2018","journal-title":"IEEE Access"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Chen, Y., Zhang, Z., Yuan, C., Li, B., Deng, Y., and Hu, W. (2021, January 11\u201317). Channel-wise topology refinement graph convolution for skeleton-based action recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.01311"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"9532","DOI":"10.1109\/TIP.2020.3028207","article-title":"Skeleton-based action recognition with multi-stream adaptive graph convolutional networks","volume":"29","author":"Shi","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Liu, Z., Zhang, H., Chen, Z., Wang, Z., and Ouyang, W. (2020, January 13\u201319). Disentangling and unifying graph convolutions for skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00022"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Duan, H., Zhao, Y., Chen, K., Lin, D., and Dai, B. (2022, January 18\u201324). Revisiting skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00298"},{"key":"ref_22","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention is all you need. Adv. Neural Inf. Process. Syst., 30."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Strudel, R., Garcia, R., Laptev, I., and Schmid, C. (2021, January 11\u201317). Segmenter: Transformer for semantic segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00717"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020, January 23\u201328). End-to-end object detection with transformers. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Malik, N.u.R., Sheikh, U.U., Abu-Bakar, S.A.R., and Channa, A. (2023). Multi-View Human Action Recognition Using Skeleton Based-FineKNN with Extraneous Frame Scrapping Technique. Sensors, 23.","DOI":"10.3390\/s23052745"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Duan, H., Wang, J., Chen, K., and Lin, D. (2022). PYSKL: Towards Good Practices for Skeleton Action Recognition. arXiv.","DOI":"10.1145\/3503161.3548546"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"11183","DOI":"10.1109\/ACCESS.2023.3240769","article-title":"A Survey on Yogic Posture Recognition","volume":"11","author":"Rajendran","year":"2023","journal-title":"IEEE Access"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3022729","article-title":"Design and real-world evaluation of Eyes-Free yoga: An Exergame for blind and Low-Vision exercise","volume":"9","author":"Rector","year":"2017","journal-title":"Acm Trans. Access. Comput. (Taccess)"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Wanjun, Y., Chong, C., and Rui, C. (2023, January 29\u201331). Yoga action recognition based on STF-ResNet. Proceedings of the 2023 IEEE 3rd International Conference on Power, Electronics and Computer Applications (ICPECA), Shenyang, China.","DOI":"10.1109\/ICPECA56706.2023.10076099"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"23969","DOI":"10.1007\/s11042-018-5721-2","article-title":"Computer-assisted yoga training system","volume":"77","author":"Chen","year":"2018","journal-title":"Multimed. Tools Appl."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Trejo, E.W., and Yuan, P. (2018, January 23\u201325). Recognition of Yoga poses through an interactive system with Kinect device. Proceedings of the 2018 2nd International Conference on Robotics and Automation Sciences (ICRAS), Wuhan, China.","DOI":"10.1109\/ICRAS.2018.8443267"},{"key":"ref_32","unstructured":"Jin, X., Yao, Y., Jiang, Q., Huang, X., Zhang, J., Zhang, X., and Zhang, K. (2015, January 18\u201321). Virtual personal trainer via the kinect sensor. Proceedings of the 2015 IEEE 16th International Conference on Communication Technology (ICCT), Hangzhou, China."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Chen, H.T., He, Y.Z., Hsu, C.C., Chou, C.L., Lee, S.Y., and Lin, B.S.P. (2014, January 6\u201310). Yoga posture recognition for self-training. Proceedings of the International Conference on Multimedia Modeling, Dublin, Ireland.","DOI":"10.1007\/978-3-319-04114-8_42"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Chen, H.T., He, Y.Z., Chou, C.L., Lee, S.Y., Lin, B.S.P., and Yu, J.Y. (2013, January 15\u201319). Computer-assisted self-training system for sports exercise using kinects. Proceedings of the 2013 IEEE International Conference on Multimedia and Expo Workshops (ICMEW), San Jose, CA, USA.","DOI":"10.1109\/ICMEW.2013.6618307"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Sun, K., Xiao, B., Liu, D., and Wang, J. (2019, January 15\u201320). Deep high-resolution representation learning for human pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00584"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Si, C., Chen, W., Wang, W., Wang, L., and Tan, T. (2019, January 15\u201320). An attention enhanced graph convolutional lstm network for skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00132"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/10\/4741\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:34:42Z","timestamp":1760124882000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/10\/4741"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,14]]},"references-count":37,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2023,5]]}},"alternative-id":["s23104741"],"URL":"https:\/\/doi.org\/10.3390\/s23104741","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,14]]}}}