{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T14:21:32Z","timestamp":1780928492686,"version":"3.54.1"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T00:00:00Z","timestamp":1780876800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"Beijing Natural Science Foundation","award":["L257022"],"award-info":[{"award-number":["L257022"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62441232"],"award-info":[{"award-number":["62441232"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>In untrimmed video data of indoor scenes, actions exhibit complex temporal relationships, such as co-occurrence and compositional dependencies, and multi-granularity semantic hierarchies. Detecting actions under these intricate relationships is a challenging task. Many existing methods primarily rely on complex temporal annotations to address the problem of multi-granularity action detection. However, this approach often lacks explicit modeling of hierarchical dependencies between actions. This oversight results in limited capability to handle multi-granularity action detection across complex temporal scales. To address the limitations, this article proposes a novel multi-granular action detection method based on highlighted a set of categories hierarchical prior knowledge graph, which is first introduced in this article. For representing the hierarchical prior knowledge graph, a Quantified Action Hierarchical Prior Construction (QAHC) approach is proposed, through which the semantics of action categories are integrated with the actions\u2019 temporal patterns to express the hierarchical dependencies and logical relationships among actions both qualitatively and quantitatively in the form of a directed acyclic graph. Based on the hierarchical prior knowledge graph, we design a Prior Highlight Guided Action Detection (PHAD) network. It employs a prior highlighted mechanism to select sufficient and necessary prior knowledge in real-time to guide the efficient detection of multi-granularity actions. Experimental results show that our method achieves excellent accuracy on the Charades (+1.49) and STS (+1.12) datasets.<\/jats:p>","DOI":"10.1145\/3815785","type":"journal-article","created":{"date-parts":[[2026,5,16]],"date-time":"2026-05-16T14:19:31Z","timestamp":1778941171000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Multi-Granular Action Detection Based on Highlighted Categories Hierarchical Prior Knowledge"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-4355-3169","authenticated-orcid":false,"given":"Xiaochen","family":"Wang","sequence":"first","affiliation":[{"name":"School of Information Science and Technology, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7722-7172","authenticated-orcid":false,"given":"Dehui","family":"Kong","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5583-8260","authenticated-orcid":false,"given":"Jinghua","family":"Li","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7962-4091","authenticated-orcid":false,"given":"Jing","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3121-1823","authenticated-orcid":false,"given":"Baocai","family":"Yin","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,8]]},"reference":[{"key":"e_1_3_1_2_2","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Piergiovanni A. J.","year":"2019","unstructured":"A. J. Piergiovanni and Michael S. Ryoo. 2019. Temporal gaussian mixture layer for videos. In Proceedings of the International Conference on Machine Learning. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:52903497"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3633516"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01063"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"e_1_3_1_6_2","volume-title":"Proceedings of the British Machine Vision Conference","author":"Dai Rui","year":"2021","unstructured":"Rui Dai, Srijan Das, and Fran\u00e7ois Br\u00e9mond. 2021. CTRN: Class-temporal relational network for action detection. In Proceedings of the British Machine Vision Conference. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:239885274"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01941"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV48630.2021.00301"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2022.3169976"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01181"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/945546.945547"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00028"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV51458.2022.00020"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.622"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW59228.2023.00379"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2022.3207511"},{"key":"e_1_3_1_17_2","unstructured":"Ashraful Islam Chengjiang Long and Richard J. Radke. 2021. A hybrid attention mechanism for weakly-supervised temporal action localization. arXiv:2101.00545. Retrieved from https:\/\/arxiv.org\/abs\/2101.00545"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00828"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2023.109595"},{"key":"e_1_3_1_20_2","first-page":"2373","volume-title":"Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Lee Pilhyeon","year":"2023","unstructured":"Pilhyeon Lee, Taeoh Kim, Minho Shim, Dongyoon Wee, and Hyeran Byun. 2023. Decomposed cross-modal distillation for RGB-based temporal action detection. In Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2373\u20132383. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:257833816"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3124671"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3712598"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP49359.2023.10222742"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3721433"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-022-0332-2"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3712594"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3100842"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548257"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00054"},{"key":"e_1_3_1_30_2","first-page":"5304","volume-title":"Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Piergiovanni A. J.","year":"2017","unstructured":"A. J. Piergiovanni and Michael S. Ryoo. 2017. Learning latent super-events to detect multiple activities in videos. In Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 5304\u20135313. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:4451203"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.plrev.2023.05.012"},{"key":"e_1_3_1_32_2","doi-asserted-by":"crossref","first-page":"19070","DOI":"10.1109\/CVPR52729.2023.01828","volume-title":"Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Ryoo Michael S.","year":"2023","unstructured":"Michael S. Ryoo, Keerthana Gopalakrishnan, Kumara Kahatapitiya, Ted Xiao, Kanishka Rao, Austin Stone, Yao Lu, Julian Ibarz, and Anurag Arnab. 2023. Token turing machines. In Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19070\u201319081. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:253553171"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV51458.2022.00022"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00269"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN52387.2021.9533426"},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","first-page":"510","DOI":"10.1007\/978-3-319-46448-0_31","volume-title":"Computer Vision\u2014ECCV 2016","author":"Sigurdsson A.","year":"2016","unstructured":"A. Sigurdsson, Gunnar G\u00fcl Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016. Hollywood in homes: Crowdsourcing data collection for activity understanding. In Computer Vision\u2014ECCV 2016. Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling (Eds.), Springer, 510\u2013526."},{"key":"e_1_3_1_37_2","unstructured":"Arkaprava Sinha Monish Soundar Raj Pu Wang Ahmed Helmy and Srijan Das. 2025. MS-Temba: Multi-scale temporal mamba for efficient temporal action detection. arXiv:2501.06138. Retrieved from https:\/\/arxiv.org\/abs\/2501.06138"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.5555\/3294996.3295163"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01327"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00151"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.5555\/3157382.3157504"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3098839"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2025.3595145"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-023-01917-4"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01727"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01932"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN60899.2024.10651373"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00628"},{"key":"e_1_3_1_49_2","first-page":"5794","volume-title":"Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV)","author":"Xu Huijuan","year":"2017","unstructured":"Huijuan Xu, Abir Das, and Kate Saenko. 2017. R-C3D: Region convolutional 3D network for temporal activity detection. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), 5794\u20135803. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:10140667"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2025.3588710"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2023.3329374"},{"key":"e_1_3_1_52_2","first-page":"18570","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Yang Min","year":"2024","unstructured":"Min Yang, Huan Gao, Ping Guo, and Limin Wang. 2024. Adapting short-term transformers for action detection in untrimmed videos. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18570\u201318579."},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2025.3632229"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02208"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3386553"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2025.3562877"},{"key":"e_1_3_1_57_2","doi-asserted-by":"crossref","first-page":"297","DOI":"10.1007\/978-3-031-19772-7_18","volume-title":"Computer Vision\u2014ECCV 2022","author":"Zheng Sipeng","year":"2022","unstructured":"Sipeng Zheng, Shizhe Chen, and Qin Jin. 2022. Few-shot action recognition with hierarchical matching and contrastive learning. In Computer Vision\u2014ECCV 2022. Shai Avidan, Gabriel Brostow, Moustapha Ciss\u00e9, Giovanni Maria Farinella, and Tal Hassner (Eds.), Springer Nature, 297\u2013313."},{"key":"e_1_3_1_58_2","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Zhou Jie","year":"2020","unstructured":"Jie Zhou, Chunping Ma, Dingkun Long, Guangwei Xu, Ning Ding, Haoyu Zhang, Pengjun Xie, and Gongshen Liu. 2020. Hierarchy-aware global model for hierarchical text classification. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:220044881"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00200"},{"key":"e_1_3_1_60_2","first-page":"18559","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhu Yuhan","year":"2024","unstructured":"Yuhan Zhu, Guozhen Zhang, Jing Tan, Gangshan Wu, and Limin Wang. 2024. Dual DETRs for multi-label temporal action detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18559\u201318569."},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3387946"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3815785","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T13:34:11Z","timestamp":1780925651000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3815785"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,8]]},"references-count":60,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3815785"],"URL":"https:\/\/doi.org\/10.1145\/3815785","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,8]]},"assertion":[{"value":"2025-09-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-29","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}