{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,27]],"date-time":"2026-07-27T14:03:39Z","timestamp":1785161019842,"version":"3.55.0"},"reference-count":30,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2023,1,2]],"date-time":"2023-01-02T00:00:00Z","timestamp":1672617600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Systems"],"abstract":"<jats:p>Video-based scoring using neural networks is a very important means for evaluating many sports, especially figure skating. Although many methods for evaluating action quality have been proposed, there is no uniform conclusion on the best feature extractor and clip length for the existing methods. Furthermore, during the feature aggregation stage, these methods cannot accurately locate the target information. To address these tasks, firstly, we systematically compare the effects of the figure skating model with three different feature extractors (C3D, I3D, R3D) and four different segment lengths (5, 8, 16, 32). Secondly, we propose a Multi-Scale Location Attention Module (MS-LAM) to capture the location information of athletes in different video frames. Finally, we present a novel Multi-scale Location Attentive Long Short-Term Memory (MLA-LSTM), which can efficiently learn local and global sequence information in each video. In addition, our proposed model has been validated on the Fis-V and MIT-Skate datasets. The experimental results show that I3D and 32 frames per second are the best feature extractor and clip length for video scoring tasks. In addition, our model outperforms the current state-of-the-art method hybrid dynAmic-statiC conText-aware attentION NETwork (ACTION-NET), especially on MIT-Skate (by 0.069 on Spearman\u2019s rank correlation). In addition, it achieves average improvements of 0.059 on Fis-V compared with Multi-scale convolutional skip Self-attentive LSTM Module (MS-LSTM). It demonstrates the effectiveness of our models in learning to score figure skating videos.<\/jats:p>","DOI":"10.3390\/systems11010021","type":"journal-article","created":{"date-parts":[[2023,1,3]],"date-time":"2023-01-03T02:05:53Z","timestamp":1672711553000},"page":"21","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["MLA-LSTM: A Local and Global Location Attention LSTM Learning Model for Scoring Figure Skating"],"prefix":"10.3390","volume":"11","author":[{"given":"Chaoyu","family":"Han","sequence":"first","affiliation":[{"name":"Physics and Electronic Information Engineering, Zhejiang Normal University, Jinhua 321004, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fangyao","family":"Shen","sequence":"additional","affiliation":[{"name":"Mathematics and Computer Science, Zhejiang Normal University, Jinhua 321004, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lina","family":"Chen","sequence":"additional","affiliation":[{"name":"Physics and Electronic Information Engineering, Zhejiang Normal University, Jinhua 321004, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaoyi","family":"Lian","sequence":"additional","affiliation":[{"name":"Physics and Electronic Information Engineering, Zhejiang Normal University, Jinhua 321004, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongjie","family":"Gou","sequence":"additional","affiliation":[{"name":"Computer Science and Technology, Harbin Institute of Technology, Harbin 150006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hong","family":"Gao","sequence":"additional","affiliation":[{"name":"Computer Science and Technology, Harbin Institute of Technology, Harbin 150006, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,1,2]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Parmar, P., and Morris, B.T. (2019, January 15\u201320). What and how well you performed? A multitask learning approach to action quality assessment. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00039"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Tang, Y., Ni, Z., Zhou, J., Zhang, D., Lu, J., Wu, Y., and Zhou, J. (2020, January 13\u201319). Uncertainty-aware score distribution learning for action quality assessment. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00986"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Li, Y., Chai, X., and Chen, X. (2018). End-to-end learning for action quality assessment. Pacific Rim Conference on Multimedia, Springer.","DOI":"10.1007\/978-3-030-00767-6_12"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Zeng, L.A., Hong, F.T., Zheng, W.S., Yu, Q.Z., Zeng, W., Wang, Y.W., and Lai, J.H. (2020, January 12). Hybrid dynamic-static context-aware attention network for action assessment in long videos. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3413560"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2846","DOI":"10.1007\/s11263-021-01486-4","article-title":"SportsCap: Monocular 3D human motion capture and fine-grained understanding in challenging sports videos","volume":"129","author":"Chen","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Zuo, K., and Su, X. (2022). Three-Dimensional Action Recognition for Basketball Teaching Coupled with Deep Neural Network. Electronics, 11.","DOI":"10.3390\/electronics11223797"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning spatiotemporal features with 3d convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Carreira, J., and Zisserman, A. (2017, January 21\u201326). Quo vadis, action recognition? A new model and the kinetics dataset. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.502"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Hara, K., Kataoka, H., and Satoh, Y. (2017, January 22\u201329). Learning spatio-temporal features with 3d residual networks for action recognition. Proceedings of the IEEE International Conference on Computer Vision Workshops, Venice, Italy.","DOI":"10.1109\/ICCVW.2017.373"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Pirsiavash, H., Vondrick, C., and Torralba, A. (2014). Assessing the quality of actions. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10599-4_36"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Parmar, P., and Tran Morris, B. (2017, January 21\u201326). Learning to score olympic events. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.16"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Le, Q.V., Zou, W.Y., Yeung, S.Y., and Ng, A.Y. (2011, January 20\u201325). Learning hierarchical invariant spatio-temporal features for action recognition with independent subspace analysis. Proceedings of the CVPR 2011, Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995496"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"4578","DOI":"10.1109\/TCSVT.2019.2927118","article-title":"Learning to score figure skating sport videos","volume":"30","author":"Xu","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_14","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., and He, K. (November, January 27). Slowfast networks for video recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Wang, L., Tong, Z., Ji, B., and Wu, G. (2021, January 20\u201325). Tdn: Temporal difference networks for efficient action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00193"},{"key":"ref_16","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., and Paluri, M. (2018, January 18\u201323). A closer look at spatiotemporal convolutions for action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00675"},{"key":"ref_18","first-page":"101919","article-title":"WilDect-YOLO: An efficient and robust computer vision-based accurate object localization model for automated endangered wildlife detection","volume":"2022","author":"Roy","year":"2022","journal-title":"Ecol. Inform."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"8448","DOI":"10.1007\/s10489-021-02893-3","article-title":"RSOD: Real-time small object detection algorithm in UAV-based traffic monitoring","volume":"52","author":"Sun","year":"2022","journal-title":"Appl. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Doughty, H., Mayol-Cuevas, W., and Damen, D. (2019, January 15\u201320). The pros and cons: Rank-aware temporal attention for skill determination in long videos. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00805"},{"key":"ref_22","unstructured":"Nakano, T., Sakata, A., and Kishimoto, A. (2020). Estimating blink probability for highlight detection in figure skating videos. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1575","DOI":"10.1007\/s11760-021-01890-w","article-title":"Temporal attention learning for action quality assessment in sports video","volume":"15","author":"Lei","year":"2021","journal-title":"Signal Image Video Process."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Xu, A., Zeng, L.A., and Zheng, W.S. (2022, January 19\u201320). Likert Scoring With Grade Decoupling for Long-Term Action Assessment. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00323"},{"key":"ref_25","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 7\u20139). Batch normalization: Accelerating deep network training by reducing internal covariate shift. Proceedings of the International Conference on Machine Learning, PMLR, Lille, France."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., and Fei-Fei, L. (2014, January 23\u201328). Large-scale video classification with convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.223"},{"key":"ref_27","unstructured":"Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., and Natsev, P. (2017). The kinetics human action video dataset. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. (2017, January 22\u201329). Grad-cam: Visual explanations from deep networks via gradient-based localization. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.74"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Sahoo, J.P., Prakash, A.J., P\u0142awiak, P., and Samantray, S. (2022). Real-Time Hand Gesture Recognition Using Fine-Tuned Convolutional Neural Network. Sensors, 22.","DOI":"10.3390\/s22030706"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Shao, D., Zhao, Y., Dai, B., and Lin, D. (2020, January 13\u201319). Finegym: A hierarchical video dataset for fine-grained action understanding. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00269"}],"container-title":["Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2079-8954\/11\/1\/21\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T17:56:37Z","timestamp":1760118997000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2079-8954\/11\/1\/21"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,2]]},"references-count":30,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,1]]}},"alternative-id":["systems11010021"],"URL":"https:\/\/doi.org\/10.3390\/systems11010021","relation":{},"ISSN":["2079-8954"],"issn-type":[{"value":"2079-8954","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,2]]}}}