{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T14:28:09Z","timestamp":1781533689384,"version":"3.54.5"},"reference-count":36,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T00:00:00Z","timestamp":1779321600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Advanced Ocean Institute of Southeast University, Nantong","award":["GP20010002"],"award-info":[{"award-number":["GP20010002"]}]},{"award":["GP20010002"],"award-info":[{"award-number":["GP20010002"]}],"id":[{"id":"https:\/\/ror.org\/028khat13","id-type":"ROR","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Archery is a fine-grained skill sport in which small posture deviations can markedly affect performance, motivating the need for reliable automated technique assessment. However, most existing methods still focus on large-amplitude sports and cannot match coach-level nuance. To overcome these limitations, we introduce SEMA (Semantic Evidence-Driven Multimodal Assessment), a large language model (LLM)-based end-to-end system for fine-grained archery action quality assessment. Beyond score prediction and evaluation-text generation, SEMA further supports knowledge-grounded question answering and feedback generation through a hierarchical multi-source knowledge framework that integrates assessment outputs, structured coaching guidance, and general archery knowledge. Experimental results show that SEMA achieves strong performance on the novel AAV dataset, outperforming general-purpose VLMs and adapted prior AQA methods. In addition, we introduce the AAV (Archery Action Video) dataset, the first multimodal, fine-grained action quality assessment (AQA) dataset dedicated to archery, and release it publicly to the community. This dataset addresses a critical gap in current benchmarks for assessing archery action quality and intelligent archery training.<\/jats:p>","DOI":"10.3390\/info17050511","type":"journal-article","created":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T13:53:24Z","timestamp":1779371604000},"page":"511","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["A Modular Approach to Automated Archery Coaching for Action Quality Assessment and Feedback Generation Using Large Language Models"],"prefix":"10.3390","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-7369-1436","authenticated-orcid":false,"given":"Yunyixuan","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Automation, Southeast University, Nanjing 210096, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3075-0861","authenticated-orcid":false,"given":"Haoran","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Automation, Southeast University, Nanjing 210096, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Binrong","family":"Zhu","sequence":"additional","affiliation":[{"name":"School of Automation, Southeast University, Nanjing 210096, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-5904-0857","authenticated-orcid":false,"given":"Xiaozhi","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Physical Education, Southeast University, Nanjing 210096, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0953-6501","authenticated-orcid":false,"given":"Siyu","family":"Xia","sequence":"additional","affiliation":[{"name":"School of Automation, Southeast University, Nanjing 210096, China"},{"name":"Advanced Ocean Institute of Southeast University, Southeast University, Nantong 226010, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,5,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Lei, Q., Zhang, H.-B., Du, J.-X., Hsiao, T.-C., and Chen, C.-C. (2020). Learning Effective Skeletal Representations on RGB Video for Fine-Grained Human Action Quality Assessment. Electronics, 9.","DOI":"10.3390\/electronics9040568"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Lin, F., Huang, J., Chen, Z., Zhu, K., and Feng, C. (2025). Enhancing Long-Term Action Quality Assessment: A Dual-Modality Dataset and Causal Cross-Modal Framework for Trampoline Gymnastics. Sensors, 25.","DOI":"10.3390\/s25185824"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Ji, X., Miller, J., Gao, X., Al Tamimi, Z., Arzalluz, I., and Piovesan, D. (2024). An Ergonomics Analysis of Archers through Motion Tracking to Prevent Injuries and Improve Performance. Sensors, 24.","DOI":"10.3390\/s24061862"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Seino, T., Saito, N., Ogawa, T., Asamizu, S., and Haseyama, M. (2025). Expert Comment Generation Considering Sports Skill Level Using a Large Multimodal Model with Video and Spatial-Temporal Motion Features. Sensors, 25.","DOI":"10.3390\/s25020447"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Sun, W., and Zhai, G. (2025). A Perspective on Quality Evaluation for AI-Generated Videos. Sensors, 25.","DOI":"10.3390\/s25154668"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Pirsiavash, H., Vondrick, C., and Torralba, A. (2014, January 6\u201312). Assessing the Quality of Actions. Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10599-4_36"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Parmar, P., and Morris, B.T. (2017, January 21\u201326). Learning to Score Olympic Events. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.16"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"4578","DOI":"10.1109\/TCSVT.2019.2927118","article-title":"Learning to Score Figure Skating Sport Videos","volume":"30","author":"Xu","year":"2020","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., and Lin, D. (2018, January 2\u20137). Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Pan, J.-H., Gao, J., and Zheng, W.-S. (November, January 27). Action Assessment by Joint Relation Graphs. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00643"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Tang, Y., Ni, Z., Zhou, J., Zhang, D., Lu, J., Wu, Y., and Zhou, J. (2020, January 13\u201319). Uncertainty-Aware Score Distribution Learning for Action Quality Assessment. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00986"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Yu, X., Rao, Y., Zhao, W., Lu, J., and Zhou, J. (2021, January 10\u201317). Group-aware Contrastive Regression for Action Quality Assessment. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00782"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Xu, H., Ke, X., Wu, H., Xu, R., Li, Y., and Guo, W. (2025, January 11\u201315). Language-Guided Audio-Visual Learning for Long-Term Sports Assessment. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.02232"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Gao, R., Liu, X., Hu, Z., Xing, B., Xia, B., Yu, Z., and K\u00e4lvi\u00e4inen, H. (2025, January 11\u201315). FSBench: A Figure Skating Benchmark for Advancing Artistic Sports Understanding. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.01269"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Huang, B., Wang, X., Chen, H., Song, Z., and Zhu, W. (2024, January 16\u201322). VTimeLLM: Empower LLM to Grasp Video Moments. Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01353"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Deng, A., Gao, Z., Choudhuri, A., Planche, B., Zheng, M., Wang, B., Chen, T., Chen, C., and Wu, Z. (2025, January 11\u201315). Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.01285"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zeng, W., Gao, D., Shou, M.Z., and Ng, H.T. (2025, January 19\u201323). Factorized Learning for Temporally Grounded Video-Language Models. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Honolulu, HI, USA. Available online: https:\/\/openaccess.thecvf.com\/content\/ICCV2025\/html\/Zeng_Factorized_Learning_for_Temporally_Grounded_Video-Language_Models_ICCV_2025_paper.html.","DOI":"10.1109\/ICCV51701.2025.01923"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., and Fei-Fei, L. (2014, January 23\u201328). Large-Scale Video Classification with Convolutional Neural Networks. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.223"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Safdarnejad, S.M., Liu, X., Udpa, L., Andrus, B., Wood, J., and Craven, D. (2015, January 4\u20138). Sports Videos in the Wild (SVW): A Video Dataset for Sports Analysis. Proceedings of the 2015 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG), Ljubljana, Slovenia.","DOI":"10.1109\/FG.2015.7163105"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Li, Y., Chen, L., He, R., Wang, Z., Wu, G., and Wang, L. (2021, January 10\u201317). MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports Actions. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01328"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Cui, Y., Zeng, C., Zhao, X., Yang, Y., Wu, G., and Wang, L. (2023, January 1\u20136). SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports Scenes. Proceedings of the 2023 IEEE\/CVF International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.00910"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Xu, J., Zhao, G., Yin, S., Zhou, W., and Peng, Y. (2024, January 16\u201322). FineSports: A Multi-Person Hierarchical Sports Video Dataset for Fine-Grained Action Understanding. Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02057"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Scott, A., Uchida, I., Ding, N., Umemoto, R., Bunker, R., Kobayashi, R., Koyama, T., Onishi, M., Kameda, Y., and Fujii, K. (2024, January 17\u201318). TeamTrack: A Dataset for Multi-Sport Multi-Object Tracking in Full-pitch Videos. Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA.","DOI":"10.1109\/CVPRW63382.2024.00340"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Xu, J., Rao, Y., Yu, X., Chen, G., Zhou, J., and Lu, J. (2022, January 18\u201324). FineDiving: A Fine-grained Dataset for Procedure-aware Action Quality Assessment. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00296"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Xu, J., Yin, S., Zhao, G., Wang, Z., and Peng, Y. (2024, January 16\u201322). FineParser: A Fine-Grained Spatio-Temporal Action Parser for Human-Centric Action Quality Assessment. Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01386"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Parmar, P., and Morris, B.T. (2019, January 15\u201320). What and How Well You Performed? A Multitask Learning Approach to Action Quality Assessment. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00039"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Parmar, P., and Morris, B. (2019, January 7\u201311). Action Quality Assessment Across Multiple Actions. Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa Village, HI, USA.","DOI":"10.1109\/WACV.2019.00161"},{"key":"ref_28","unstructured":"Qi, M., Wu, Y., Zhang, X., and Ma, H. (2025). Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"516","DOI":"10.1080\/15438627.2021.1917402","article-title":"Body Proportions According to Stature Groups in Elite Athletes","volume":"30","author":"Porta","year":"2022","journal-title":"Res. Sport. Med."},{"key":"ref_30","unstructured":"Jiang, T., Lu, P., Zhang, L., Ma, N., Han, R., Lyu, C., Li, Y., and Chen, K. (2023). RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"2171","DOI":"10.1109\/TPAMI.2023.3330794","article-title":"Temporal Action Localization in the Deep Learning Era: A Survey","volume":"46","author":"Wang","year":"2024","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wang, H., Xu, Z., Cheng, Y., Diao, S., Zhou, Y., Cao, Y., Wang, Q., Ge, W., and Huang, L. (2025, January 4\u20139). Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models. Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China.","DOI":"10.18653\/v1\/2025.findings-emnlp.50"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"107299","DOI":"10.1016\/j.sigpro.2019.107299","article-title":"Selective Review of Offline Change Point Detection Methods","volume":"167","author":"Truong","year":"2020","journal-title":"Signal Process."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Carreira, J., and Zisserman, A. (2017, January 21\u201326). Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.502"},{"key":"ref_35","unstructured":"Zhang, T., Kishore, V., Wu, F., Weinberger, K.Q., and Artzi, Y. (2019). BERTScore: Evaluating Text Generation with BERT. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"561","DOI":"10.3233\/IDA-2007-11508","article-title":"Toward Accurate Dynamic Time Warping in Linear Time and Space","volume":"11","author":"Salvador","year":"2007","journal-title":"Intell. Data Anal."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/17\/5\/511\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T14:07:24Z","timestamp":1779372444000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/17\/5\/511"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,21]]},"references-count":36,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2026,5]]}},"alternative-id":["info17050511"],"URL":"https:\/\/doi.org\/10.3390\/info17050511","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,21]]}}}