{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T19:54:26Z","timestamp":1783454066922,"version":"3.55.0"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T00:00:00Z","timestamp":1783382400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62471013, 61971016"],"award-info":[{"award-number":["62471013, 61971016"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Beijing Municipal Education Commission Cooperation Beijing Natural Science Foundation","award":["KZ201910005007"],"award-info":[{"award-number":["KZ201910005007"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>With the continuous advancement of the livestreaming industry, streamers as content producers pose significant challenges to the timeliness of regulatory response mechanisms, emerging as a critical weak link in cyberspace governance. Human-object interaction (HOI) detection plays a pivotal role in understanding multimodal livestreaming videos. In mainstream healthy online ecosystems, normal HOI categories dominate, while rare ones are extremely scarce, i.e., long-tail distribution that hinders HOI models from effectively detecting streamer violations. Driven by the transformative potential of foundation models (FMs) in multimodal video understanding, we propose a knowledge-driven memory network (KdM-Net) for long-tailed HOI in livestreaming, leveraging the extensive capabilities of the contrastive language-image pretraining (CLIP) model. First, human-object (HO) pairs are generated and modeled using general object detector\/tracker. After converting each HOI label into a short sentence description, text embeddings are extracted via the CLIP text encoder to initialize classifier weights. Notably, we introduce a visual-textual knowledge transfer strategy to align visual and text features, complementing for the multimodal knowledge deficit of rare categories that plague long-tailed HOI distributions. Finally, a knowledge-driven memory module is designed to dynamically assign adaptive weights and attention to interaction categories based on their long-tailed distribution characteristics, mitigating model forgetting tail data and enhancing HOI detection performance. Experimental results demonstrate that KdM-Net achieves HOI detection accuracies of 37.33@full, 50.63%@non-rare, and 27.14%@rare on the publicly available VidHOI dataset, and 45.41%@full, 61.77%@non-rare, and 31.95%@rare on the self-built BJUT-HOI dataset. These fundings validate the generalization power of our KdM-Net for HOI detection and its competitiveness in livestreaming scenarios.<\/jats:p>","DOI":"10.1145\/3817607","type":"journal-article","created":{"date-parts":[[2026,5,23]],"date-time":"2026-05-23T13:21:13Z","timestamp":1779542473000},"page":"1-28","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["KdM-Net: Knowledge-Driven Memory Network for Long-Tailed Human-Object Interaction in Livestreaming"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-0032-0745","authenticated-orcid":false,"given":"Menghui","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Integrated Circuits, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1290-0738","authenticated-orcid":false,"given":"Jing","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Integrated Circuits, Beijing University of Technology, Beijing, China and Beijing Key Laboratory of Computational Intelligence and Intelligent System, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5234-7606","authenticated-orcid":false,"given":"Lin","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Integrated Circuits, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9937-2669","authenticated-orcid":false,"given":"Li","family":"Zhuo","sequence":"additional","affiliation":[{"name":"School of Integrated Circuits, Beijing University of Technology, Beijing, China and Beijing Key Laboratory of Computational Intelligence and Intelligent System, Beijing University of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,7]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV56688.2023.00507"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_1_5_2","unstructured":"BJUT-AIVBD. n.d. BJUT-HOI Dataset. Retrieved from https:\/\/github.com\/BJUT-AIVBD\/BJUT-HOI-dataset"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3674980"},{"key":"e_1_3_1_8_2","unstructured":"Xinlei Chen Hao Fang Tsung-Yi Lin Ramakrishna Vedantam Saurabh Gupta Piotr Dollar and C. Lawrence Zitnick. 2015. Microsoft COCO captions: Data collection and evaluation server. arXiv:1504.00325. Retrieved from https:\/\/arxiv.org\/abs\/1504.00325"},{"key":"e_1_3_1_9_2","unstructured":"China Daily. 2022. China to Tighten Regulation of Livestreaming Short Videos. Retrieved May 18 2025 from https:\/\/global.chinadaily.com.cn\/a\/202203\/17\/WS62333dc7a310fd2b29e51976.html"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3463944.3469097"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00544"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01606"},{"key":"e_1_3_1_13_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00056"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","unstructured":"Glenn Jocher. 2020. Ultralytics YOLOv5 Version 7.0. Zenodo. DOI: 10.5281\/zenodo.3908559","DOI":"10.5281\/zenodo.3908559"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-92591-7_14"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58555-6_30"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913478446"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00596"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00056"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01949"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i2.20075"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.126243"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3663668"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2023.103741"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02251"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01240-3_25"},{"key":"e_1_3_1_29_2","unstructured":"Shuai Shao Zijian Zhao Boxun Li Tete Xiao Gang Yu Xiangyu Zhang and Jian Sun. 2018. CrowdHuman: A benchmark for detecting human in a crowd. arXiv:1805.00123. Retrieved from https:\/\/arxiv.org\/abs\/1805.00123"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3689638"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413778"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01027"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICNISC57059.2022.00050"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.5555\/3600270.3601966"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20053-3_38"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3603253"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00101"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-73013-9_23"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2017.8296962"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00286"},{"key":"e_1_3_1_41_2","first-page":"2048","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In Proceedings of the International Conference on Machine Learning, 2048\u20132057."},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3369699"},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","unstructured":"Menghui Zhang Jing Zhang Lin Chen and Li Zhuo. 2025. Prototype embedding optimization for human-object interaction detection in livestreaming. Retrieved from https:\/\/arxiv.org\/abs\/2505.22011","DOI":"10.1109\/MMSP64401.2025.11324104"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3283879"},{"key":"e_1_3_1_45_2","unstructured":"Wayne Xin Zhao Kun Zhou Junyi Li Tianyi Tang Xiaolei Wang Yupeng Hou Yingqian Min Beichen Zhang Junjie Zhang Zican Dong et al. 2023. A survey of large language models. arXiv:2303.18223. Retrieved from https:\/\/arxiv.org\/abs\/2303.18223"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3817607","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T18:40:17Z","timestamp":1783449617000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3817607"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,7]]},"references-count":44,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3817607"],"URL":"https:\/\/doi.org\/10.1145\/3817607","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,7]]},"assertion":[{"value":"2025-11-11","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-16","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-07-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}