{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:54:41Z","timestamp":1782312881387,"version":"3.54.5"},"reference-count":39,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T00:00:00Z","timestamp":1782259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62476154"],"award-info":[{"award-number":["62476154"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Taishan Scholar Project of Shandong Province","award":["tsqn202306066"],"award-info":[{"award-number":["tsqn202306066"]}]},{"name":"Teacher Visiting and Training Fund for Ordinary Undergraduate Universities in Shandong Province"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>\n                    Spatio-temporal video grounding (STVG) aims to precisely locate a spatio-temporal tube in an untrimmed video corresponding to a given language description. Many existing methods decouple spatial and temporal grounding as separate tasks, missing the strong interdependencies between the two, which are crucial for accurately aligning spatial regions (such as objects) with their motion over time. Thus, to enhance spatio-temporal associations, we introduce a new Prior-Driven Transformer Network (PDTNet) with predicted temporal boundaries as priors to guide object bounding boxes for improved spatial grounding over time. Firstly, PDTNet employs a temporal prior, termed reference query, to enhance discriminability between language-related and language-irrelevant visual content, improving temporal boundary localization. Further, the context within predicted temporal boundaries serves as another prior knowledge to modulate spatial features. We also introduce a prediction-aware Gaussian prior to precise object localization. This ensures consistent tube construction and accurate object localization. Extensive experiments on STVG benchmarks validate the effectiveness of PDTNet. Code is available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/tongzhang111\/PDTNet\">https:\/\/github.com\/tongzhang111\/PDTNet<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3769088","type":"journal-article","created":{"date-parts":[[2025,9,29]],"date-time":"2025-09-29T15:52:04Z","timestamp":1759161124000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Learning Prediction-aware Prior in Transformer Network for Accurate Spatio-Temporal Video Grounding"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-8143-244X","authenticated-orcid":false,"given":"Xin","family":"Wang","sequence":"first","affiliation":[{"name":"School of Sciences, Shandong Jianzhu University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8163-3050","authenticated-orcid":false,"given":"Tong","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3948-4471","authenticated-orcid":false,"given":"Yongshun","family":"Gong","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8554-7827","authenticated-orcid":false,"given":"Jialin","family":"Gao","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8926-7833","authenticated-orcid":false,"given":"Yanyu","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3925-1997","authenticated-orcid":false,"given":"Jin","family":"Qian","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9644-9723","authenticated-orcid":false,"given":"Xiushan","family":"Nie","sequence":"additional","affiliation":[{"name":"School of Computer and Artificial Intelligence, Shandong Jianzhu University, Jinan, China and Shandong Yunhai Guochuang Cloud Computing Equipment Industry Innovation Co., Ltd., Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8262-8883","authenticated-orcid":false,"given":"Lizhen","family":"Cui","sequence":"additional","affiliation":[{"name":"School of Software, Shandong University, Jinan, China  and Joint SDU-NTU Centre for Artificial Intelligence Research (C-FAIR), Shandong University, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5715-7154","authenticated-orcid":false,"given":"Chengqi","family":"Zhang","sequence":"additional","affiliation":[{"name":"The Hong Kong Polytechnic University, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1723"},{"issue":"2","key":"e_1_3_1_3_2","first-page":"1747","article-title":"Learning graph convolutional networks based on quantum vertex information propagation","volume":"35","author":"Bai Lu","year":"2023","unstructured":"Lu Bai, Yuhang Jiao, Lixin Cui, Luca Rossi, Yue Wang, Philip S. Yu, and Edwin R. Hancock. 2023. Learning graph convolutional networks based on quantum vertex information propagation. IEEE Transactions on Knowledge and Data Engineering 35, 2 (2023), 1747\u20131760.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00493"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2023.109877"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2024.111099"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01735"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-2117"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00180"},{"key":"e_1_3_1_11_2","first-page":"5583","volume-title":"Proceedings of International Conference on Machine Learning","author":"Kim Wonjae","year":"2021","unstructured":"Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. In Proceedings of International Conference on Machine Learning. PMLR, 5583\u20135594."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.353"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02212"},{"issue":"2","key":"e_1_3_1_14_2","first-page":"1","article-title":"Dynamic multimodal fusion via meta-learning towards micro-video recommendation","volume":"42","author":"Liu Han","year":"2023","unstructured":"Han Liu, Yinwei Wei, Fan Liu, Wenjie Wang, Liqiang Nie, and Tat-Seng Chua. 2023. Dynamic multimodal fusion via meta-learning towards micro-video recommendation. ACM Transactions on Information Systems 42, 2 (2023), 1\u201326.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_1_15_2","unstructured":"Han Liu Yinwei Wei Xuemeng Song Weili Guan Yuan-Fang Li and Liqiang Nie. 2024. Mmgrec: Multimodal generative recommendation with transformer model. arXiv:2404.16555. Retrieved from https:\/\/arxiv.org\/abs\/2404.16555"},{"issue":"6","key":"e_1_3_1_16_2","first-page":"5977","article-title":"HS-GCN: Hamming spatial graph convolutional networks for recommendation","volume":"35","author":"Liu Han","year":"2022","unstructured":"Han Liu, Yinwei Wei, Jianhua Yin, and Liqiang Nie. 2022. HS-GCN: Hamming spatial graph convolutional networks for recommendation. IEEE Transactions on Knowledge and Data Engineering 35, 6 (2022), 5977\u20135990.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.520"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2024.110819"},{"key":"e_1_3_1_19_2","volume-title":"Proceedings of International Conference on Learning Representations","author":"Liu Shilong","year":"2022","unstructured":"Shilong Liu, Feng Li, Hao Zhang, Xiao Yang, Xianbiao Qi, Hang Su, Jun Zhu, and Lei Zhang. 2022. DAB-DETR: Dynamic anchor boxes are better queries for DETR. In Proceedings of International Conference on Learning Representations."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00205"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01491"},{"key":"e_1_3_1_22_2","article-title":"Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In","volume":"32","author":"Lu Jiasen","year":"2019","unstructured":"Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Advances in Neural Information Processing Systems, Vol. 32.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_23_2","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks. In","volume":"28","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems, Vol. 28.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00075"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093328"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00156"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3250518"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3085907"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00434"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01789"},{"key":"e_1_3_1_31_2","article-title":"HC-GAE: The hierarchical cluster-based graph auto-encoder for graph representation learning","author":"Xu Zhuo","year":"2024","unstructured":"Zhuo Xu, Lu Bai, Lixin Cui, Ming Li, Yue Wang, and Edwin R. Hancock. 2024. HC-GAE: The hierarchical cluster-based graph auto-encoder for graph representation learning. In Advances in Neural Information Processing Systems.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2024.110973"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.162"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01595"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00478"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6984"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3387696"},{"key":"e_1_3_1_38_2","article-title":"SG-Net: Syntax guided transformer for language representation","author":"Zhang Zhuosheng","year":"2020","unstructured":"Zhuosheng Zhang, Yuwei Wu, Junru Zhou, Sufeng Duan, Hai Zhao, and Rui Wang. 2020. SG-Net: Syntax guided transformer for language representation. IEEE Transactions on Pattern Analysis and Machine Intelligence (2020).","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2020\/149"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01068"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3769088","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:42:21Z","timestamp":1782312141000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3769088"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":39,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3769088"],"URL":"https:\/\/doi.org\/10.1145\/3769088","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-04-10","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-11","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}