{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T14:11:02Z","timestamp":1779372662025,"version":"3.53.1"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T00:00:00Z","timestamp":1779321600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62471013"],"award-info":[{"award-number":["62471013"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Beijing Natural Science Foundation","award":["L247025"],"award-info":[{"award-number":["L247025"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Graph networks face substantial challenges in handing large-scale graph-structured data. Although graph convolutional networks (GCNs) have been widely applied to group activity recognition, they still struggle with high graph structure complexity and semantic gap, especially in complex scenarios. To address these limitations, we harness large language models (LLMs) and multimodal learning to enhance semantic understanding as well as bridge the gap between graph structures and contextual meanings. For the simultaneous optimization of model efficiency and recognition accuracy in group activities, we propose a tri-optimization graph convolutional network with LLM (ToGCN-LLM) from the perspective of graph structure knowledge distillation. (1) To tackle high complexity of teacher network in GCN, we adopt a sparsification strategy to prune irrelevant edges and nodes, reducing computational overhead and enhancing training efficiency. (2) To mitigate information redundancy during knowledge distillation, we design a hierarchical optimization module combined with a hierarchical sampling mechanism, exploiting graph hierarchical structure and adjacency relationships to improve knowledge transfer efficiency. (3) Considering the student network\u2019s varying learning performance across training stages and the limitations of fixed learning strategies, we introduce a dynamic adaptive weight decay mechanism to achieve fine-grained convergence under different gradient updates, thereby boosting overall recognition accuracy. (4) We use the Qwen LLM to extract text description tokens, which are fused with the student GCN\u2019s last-layer features to enable multi-model, multi-optimization learning for group activity recognition. Six experiments on CAD, CAED, and BJUT-GAD dataset demonstrate that our ToGCN-LLM achieves competitive MPCA scores of 94.89%, 93.61%, and 95.77%, respectively.<\/jats:p>","DOI":"10.1145\/3811826","type":"journal-article","created":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T13:57:03Z","timestamp":1777125423000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["ToGCN-LLM: Tri-optimization Graph Convolutional Network with LLM for Group Activity Recognition"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-7890-5814","authenticated-orcid":false,"given":"Junpeng","family":"Kang","sequence":"first","affiliation":[{"name":"School of Information Science and Technology, Beijing University of Technology, Beijing, China and School of Artificial Intelligence and Language Sciences, Beijing International Studies University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1290-0738","authenticated-orcid":false,"given":"Jing","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Beijing University of Technology, Beijing, China and Beijing Key Laboratory of Computational Intelligence and Intelligent System, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-8775-0147","authenticated-orcid":false,"given":"Yufei","family":"Feng","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Beijing University of Technology, Beijing,China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9937-2669","authenticated-orcid":false,"given":"Li","family":"Zhuo","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Beijing University of Technology, Beijing, China and Beijing Key Laboratory of Computational Intelligence and Intelligent System, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,5,21]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477141"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539608"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3641289"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3032189"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3104182"},{"key":"e_1_3_1_7_2","first-page":"1282","volume-title":"Proceedings of the IEEE 12th International Conference on Computer Vision Workshops","author":"Choi Wongun","year":"2009","unstructured":"Wongun Choi, Khuram Shahid, and Silvio Savarese. 2009. What are they doing? Collective activity classification using spatio-temporal relationship among people. In Proceedings of the IEEE 12th International Conference on Computer Vision Workshops, 1282\u20131289."},{"key":"e_1_3_1_8_2","first-page":"3273","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Choi Wongun","year":"2011","unstructured":"Wongun Choi, Khuram Shahid, and Silvio Savarese. 2011. Learning context for collective activity recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3273\u20133280."},{"key":"e_1_3_1_9_2","first-page":"122","volume-title":"Proceedings of 31st International Symposium on High-Performance Parallel and Distributed Computing","author":"Fu Qiang","year":"2022","unstructured":"Qiang Fu, Yuede Ji, and H. Howie Huang. 2022. TLPGNN: A lightweight two-level parallelism paradigm for graph neural network computation on GPU. In Proceedings of 31st International Symposium on High-Performance Parallel and Distributed Computing, 122\u2013134."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01453-z"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00100"},{"key":"e_1_3_1_12_2","first-page":"534","volume-title":"Proceedings of 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"He Huarui","year":"2022","unstructured":"Huarui He, Jie Wang, Zhanqiu Zhang, and Feng Wu. 2022. Compressing deep graph neural networks via adversarial knowledge distillation. In Proceedings of 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 534\u2013544."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-023-45149-5"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2024.111104"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i4.25553"},{"key":"e_1_3_1_16_2","first-page":"1","volume-title":"Proceedings of the 2021 IEEE International Conference on Consumer Electronics-Asia","author":"Jang Sungjun","year":"2021","unstructured":"Sungjun Jang, Heansung Lee, Suhwan Cho, Sungmin Woo, and Sangyoun Lee. 2021. Ghost graph convolutional network for skeleton-based action recognition. In Proceedings of the 2021 IEEE International Conference on Consumer Electronics-Asia, 1\u20134."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2024.108412"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2024.106207"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2025.110634"},{"issue":"6","key":"e_1_3_1_20_2","doi-asserted-by":"crossref","first-page":"368","DOI":"10.1007\/s10489-024-06017-5","article-title":"RWGCN: Random walk graph convolutional network for group activity recognition","volume":"55","author":"Kang Junpeng","year":"2025","unstructured":"Junpeng Kang, Jing Zhang, Lin Chen, Hui Zhang, and Li Zhuo. 2025. RWGCN: Random walk graph convolutional network for group activity recognition. Applied Intelligence 55, 6 (2025), 368\u2013386.","journal-title":"Applied Intelligence"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.lindif.2023.102274"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-023-15830-y"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3341105.3373906"},{"key":"e_1_3_1_24_2","first-page":"6545","volume-title":"Proceedings of 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Li Jia","year":"2024","unstructured":"Jia Li, Xiangguo Sun, Yuhan Li, Zhixun Li, Hong Cheng, and Jeffrey Xu Yu. 2024. Graph intelligence with large language models and prompt learning. In Proceedings of 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 6545\u20136554."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3583780.3615017"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-021-04140-5"},{"issue":"2","key":"e_1_3_1_27_2","first-page":"524","article-title":"GAIM: Graph attention interaction model for collective activity recognition","volume":"22","author":"Lu Lihua","year":"2019","unstructured":"Lihua Lu, Yao Lu, Ruizhe Yu, Huijun Di, Lin Zhang, and Shunzhou Wang. 2019. GAIM: Graph attention interaction model for collective activity recognition. IEEE Transactions on Multimedia 22, 2 (2019), 524\u2013539.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.compag.2024.109294"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2025.103259"},{"key":"e_1_3_1_30_2","first-page":"1819","volume-title":"IEEE Transactions on Multimedia","author":"Tu Zhigang","year":"2022","unstructured":"Zhigang Tu, Jiaxu Zhang, Hongyan Li, Yujin Chen, and Junsong Yuan. 2022. Joint-bone fusion graph convolutional network for semi-supervised skeleton action recognition. IEEE Transactions on Multimedia 25 (2022), 1819\u20131831."},{"issue":"6","key":"e_1_3_1_31_2","doi-asserted-by":"crossref","first-page":"1096","DOI":"10.1049\/cje.2021.07.027","article-title":"Porn streamer recognition in live video based on multimodal knowledge distillation","volume":"30","author":"Wang Liyuan","year":"2021","unstructured":"Liyuan Wang, Jing Zhang, Jiacheng Yao, and Li Zhuo. 2021. Porn streamer recognition in live video based on multimodal knowledge distillation. Chinese Journal of Electronics 30, 6 (2021), 1096\u20131102.","journal-title":"Chinese Journal of Electronics"},{"key":"e_1_3_1_32_2","first-page":"5338","volume-title":"Advances in Neural Information Processing Systems","volume":"37","author":"Wu Xixi","year":"2024","unstructured":"Xixi Wu, Yifei Shen, Caihua Shan, Kaitao Song, Siwei Wang, Bohang Zhang, Jiarui Feng, Hong Cheng, Wei Chen, Yun Xiong, et al. 2024. Can graph learning improve planning in LLM-based agents? In Advances in Neural Information Processing Systems, Vol. 37, 5338\u20135383."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-024-4467-3"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2024.3362140"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2024.112453"},{"issue":"12","key":"e_1_3_1_36_2","doi-asserted-by":"crossref","first-page":"7574","DOI":"10.1109\/TNNLS.2021.3085567","article-title":"Position-aware participation-contributed temporal dynamic model for group activity recognition","volume":"33","author":"Yan Rui","year":"2021","unstructured":"Rui Yan, Xiangbo Shu, Chengcheng Yuan, Qi Tian, and Jinhui Tang. 2021. Position-aware participation-contributed temporal dynamic model for group activity recognition. IEEE Transactions on Neural Networks and Learning Systems 33, 12 (2021), 7574\u20137588.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3034233"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3450068"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2025.3616350"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3367412"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00710"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3193574"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3386777"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2025.3562877"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.4018\/IJSWIS.327353"},{"key":"e_1_3_1_46_2","first-page":"1","volume-title":"Proceedings of the 2024 International Joint Conference on Neural Networks","author":"Zhang Yimeng","year":"2024","unstructured":"Yimeng Zhang, Yang Yang, and Xuehao Gao. 2024. Lightweight graph convolutional network for efficient skeleton based action recognition. In Proceedings of the 2024 International Joint Conference on Neural Networks, 1\u20138."},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3719341"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547844"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.aiopen.2021.01.001"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00200"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3811826","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T13:32:41Z","timestamp":1779370361000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3811826"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,21]]},"references-count":49,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3811826"],"URL":"https:\/\/doi.org\/10.1145\/3811826","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,21]]},"assertion":[{"value":"2025-11-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-14","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}