{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,16]],"date-time":"2026-05-16T16:17:11Z","timestamp":1778948231751,"version":"3.51.4"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62236003, 62376140, 62476071, 625B2065, U24A20328, and U23A20315"],"award-info":[{"award-number":["62236003, 62376140, 62476071, 625B2065, U24A20328, and U23A20315"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100021171","name":"GuangDong Basic and Applied Basic Research Foundation","doi-asserted-by":"crossref","award":["2025A1515011732"],"award-info":[{"award-number":["2025A1515011732"]}],"id":[{"id":"10.13039\/501100021171","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Beijing Natural Science Foundation","award":["4262074 and L254018"],"award-info":[{"award-number":["4262074 and L254018"]}]},{"name":"Science and Technology Innovation Program for Distinguished Young Scholars of Shandong Province Higher Education Institutions","award":["2023KJ128"],"award-info":[{"award-number":["2023KJ128"]}]},{"name":"Major Key Project of Pengcheng Laboratory","award":["PCL2025A14-4"],"award-info":[{"award-number":["PCL2025A14-4"]}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["HIT.DZJJ.2025048"],"award-info":[{"award-number":["HIT.DZJJ.2025048"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>Unlike conventional visual question answering, video-grounded dialog requires a deep understanding of both the dialog history and the video content to generate accurate responses. Although existing methods have achieved promising results, they still struggle with progressively comprehending complex dialog history and effectively integrating video information. To address these challenges, we propose an iterative search and reasoning framework composed of a textual encoder, a visual encoder, and a generator. Specifically, the textual encoder adopts a path search and aggregation strategy to identify key cues in the dialog history that are essential for understanding the current question. Meanwhile, the visual encoder employs an iterative reasoning network to extract and highlight critical visual evidence from the video, thereby enabling more comprehensive visual understanding. Finally, we use a pretrained GPT-2 model as the answer generator to transform the discovered latent cues into coherent and contextually appropriate responses. Extensive experiments on three public datasets demonstrate the effectiveness and generalizability of the proposed framework.<\/jats:p>","DOI":"10.1145\/3808220","type":"journal-article","created":{"date-parts":[[2026,4,15]],"date-time":"2026-04-15T14:52:08Z","timestamp":1776264728000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog"],"prefix":"10.1145","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3896-170X","authenticated-orcid":false,"given":"Haoyu","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Harbin Institute of Technology Shenzhen, Shenzhen, China and Peng Cheng Laboratory, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1582-5764","authenticated-orcid":false,"given":"Meng","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Shandong Jianzhu University, Jinan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-0158-322X","authenticated-orcid":false,"given":"Yisen","family":"Feng","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Harbin Institute of Technology Shenzhen, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2197-9038","authenticated-orcid":false,"given":"Yaowei","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Harbin Institute of Technology Shenzhen, Shenzhen, China and Peng Cheng Laboratory, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5658-5509","authenticated-orcid":false,"given":"Weili","family":"Guan","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Harbin Institute of Technology Shenzhen, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1476-0273","authenticated-orcid":false,"given":"Liqiang","family":"Nie","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Harbin Institute of Technology Shenzhen, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,5,11]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"AAAI Workshop","author":"Alamri Huda","year":"2018","unstructured":"Huda Alamri, Chiori Hori, Tim K. Marks, Dhruv Batra, and Devi Parikh. 2018. Audio visual scene-aware dialog (AVSD) track for natural language generation in DSTC7. In AAAI Workshop."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2014.10.010"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01757"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3606368"},{"key":"e_1_3_2_6_2","first-page":"753","volume-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing","volume":"32","author":"Chen Zhe","year":"2023","unstructured":"Zhe Chen, Hongcheng Liu, and Yu Wang. 2023. DialogMCF: Multimodal context flow for audio visual scene-aware dialog. IEEE\/ACM Transactions on Audio, Speech, and Language Processing 32 (2003), 753\u2013764."},{"key":"e_1_3_2_7_2","volume-title":"AAAI Workshop","author":"Chu Yun-Wei","year":"2020","unstructured":"Yun-Wei Chu, Kuan-Yen Lin, Chao-Chun Hsu, and Lun-Wei Ku. 2020. Multi-step joint-modality attention network for scene-aware dialogue system. In AAAI Workshop."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3584701"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02253"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i2.16231"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3085755"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01007"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8682583"},{"key":"e_1_3_2_14_2","volume-title":"AAAI Workshop","author":"Hori Chiori","year":"2020","unstructured":"Chiori Hori, Anoop Cherian, Takaaki Hori, and Tim K. Marks. 2020. Audio visual scene-aware dialog (AVSD) track for natural language generation in DSTC8. In AAAI Workshop."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i2.16273"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3597609"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.247"},{"key":"e_1_3_2_18_2","volume-title":"AAAI Workshop","author":"Le Hung","year":"2020","unstructured":"Hung Le and Nancy F. Chen. 2020. Multimodal transformer with pointer network for the DSTC8 AVSD challenge. In AAAI Workshop."},{"key":"e_1_3_2_19_2","volume-title":"ICLR","author":"Le Hung","year":"2021","unstructured":"Hung Le, Nancy F. Chen, and Steven C. H. Hoi. 2021. Learning reasoning paths over semantic graphs for video-grounded dialogues. In ICLR."},{"key":"e_1_3_2_20_2","volume-title":"AAAI Workshop","author":"Le Hung","year":"2019","unstructured":"Hung Le, S. Hoi, Doyen Sahoo, and N. Chen. 2019. End-to-end multimodal dialog systems with hierarchical multimodal attention on video features. In AAAI Workshop."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.518"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1564"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.145"},{"key":"e_1_3_2_24_2","volume-title":"AAAI Workshop","author":"Lee Hwanhee","year":"2020","unstructured":"Hwanhee Lee, Seunghyun Yoon, Franck Dernoncourt, Doo Soon Kim, Trung Bui, and Kyomin Jung. 2020. DSTC8-AVSD: Multimodal semantic transformer network with retrieval style word generator. In AAAI Workshop."},{"key":"e_1_3_2_25_2","unstructured":"Bo Li Yuanhan Zhang Liangyu Chen Jinghao Wang Jingkang Yang and Ziwei Liu. 2023. Otter: A multi-modal model with in-context instruction tuning. arXiv:2305.03726. Retrieved from https:\/\/arxiv.org\/abs\/2305.03726"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00205"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3498557"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3065823"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2890628"},{"key":"e_1_3_2_30_2","unstructured":"Bin Lin Bin Zhu Yang Ye Munan Ning Peng Jin and Li Yuan. 2023. Video-LLaVA: Learning united visual representation by alignment before projection. arXiv:2311.10122. Retrieved from https:\/\/arxiv.org\/abs\/2311.10122"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01742"},{"key":"e_1_3_2_32_2","volume-title":"AAAI Workshop","author":"Lin Kuan-Yen","year":"2019","unstructured":"Kuan-Yen Lin, Chao-Chun Hsu, Yun-Nung Chen, and Lun-Wei Ku. 2019. Entropy-enhanced multimodal attention model for scene-aware dialogue generation. In AAAI Workshop."},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2021.3105280"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.3210003"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58586-0_14"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00684"},{"key":"e_1_3_2_37_2","first-page":"1","volume-title":"NeurIPS","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In NeurIPS, 1\u20139."},{"key":"e_1_3_2_38_2","volume-title":"AAAI Workshop","author":"Sanabria Ramon","year":"2019","unstructured":"Ramon Sanabria, Shruti Palaskar, and Florian Metze. 2019. CMU Sinbad\u2019s submission for the DSTC7 AVSD challenge. In AAAI Workshop."},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00214"},{"key":"e_1_3_2_40_2","first-page":"17959","volume-title":"CVPR","author":"Hongsuck Seo Paul","year":"2022","unstructured":"Paul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, and Cordelia Schmid. 2022. End-to-end generative pretraining for multimodal video captioning. In CVPR, 17959\u201317968."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_31"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2019.2963282"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3680774"},{"key":"e_1_3_2_44_2","unstructured":"Shuhe Wang Yuxian Meng Xiaofei Sun Fei Wu Rongbin Ouyang Rui Yan Tianwei Zhang and Jiwei Li. 2021. Modeling text-visual mutual dependency for multi-modal dialog generation. arXiv:2105.14445. Retrieved from https:\/\/arxiv.org\/abs\/2105.14445"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.269"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.276"},{"key":"e_1_3_2_47_2","unstructured":"Qianlong Xiang Miao Zhang Yuzhang Shang Jianlong Wu Yan Yan and Liqiang Nie. 2024. DKDM: Data-free knowledge distillation for diffusion models with any architecture. arXiv:2409.03550. Retrieved from https:\/\/arxiv.org\/abs\/2409.03550"},{"key":"e_1_3_2_48_2","volume-title":"AAAI Workshop","author":"Xie Huiyuan","year":"2020","unstructured":"Huiyuan Xie and Ignacio Iacobacci. 2020. Audio visual scene-aware dialog system using dynamic memory networks. In AAAI Workshop."},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.dialdoc-1.2"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i3.20215"},{"key":"e_1_3_2_51_2","volume-title":"AAAI Workshop","author":"Yeh Yi-Ting","year":"2019","unstructured":"Yi-Ting Yeh, Tzu-Chuan Lin, Hsiao-Hua Cheng, Yu-Hsuan Deng, Shang-Yu Su, and Yun-Nung Chen. 2019. Reactive multi-stage feature fusion for multimodal dialogue modeling. In AAAI Workshop."},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2019.2938015"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2021.3054769"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.280"},{"key":"e_1_3_2_55_2","first-page":"810","volume-title":"ICCV","author":"Yu Xiaodong","year":"2011","unstructured":"Xiaodong Yu, Cornelia Ferm\u00fcller, Ching Lik Teo, Yezhou Yang, and Yiannis Aloimonos. 2011. Active scene recognition with vision and language. In ICCV, 810\u2013817."},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v40i15.38244"},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","unstructured":"Hang Zhang Xin Li and Lidong Bing. 2023. Video-LLaMA: An instruction-tuned audio-visual language model for video understanding. arXiv:2306.02858. Retrieved from https:\/\/arxiv.org\/abs\/2306.02858","DOI":"10.18653\/v1\/2023.emnlp-demo.49"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP46576.2022.9897613"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475234"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3312302"},{"key":"e_1_3_2_61_2","first-page":"1","volume-title":"Advances in Neural Information Processing Systems","author":"Zhang Haoyu","year":"2025","unstructured":"Haoyu Zhang, Meng Liu, Zaijing Li, Haokun Wen, Weili Guan, Yaowei Wang, and Liqiang Nie. 2025. Spatial understanding from videos: Structured prompts meet simulation data. In Advances in Neural Information Processing Systems, 1\u201315."},{"key":"e_1_3_2_62_2","first-page":"59310","volume-title":"ICML","author":"Zhang Haoyu","year":"2024","unstructured":"Haoyu Zhang, Meng Liu, Zixin Liu, Xuemeng Song, Yaowei Wang, and Liqiang Nie. 2024. Multi-factor adaptive vision selection for egocentric video question answering. In ICML, 59310\u201359328."},{"issue":"1","key":"e_1_3_2_63_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3522763","article-title":"A static and dynamic attention framework for multi turn dialogue generation","volume":"41","author":"Zhang Weinan","year":"2023","unstructured":"Weinan Zhang, Yiming Cui, Kaiyan Zhang, Yifa Wang, Qingfu Zhu, Lingzhi Li, and Ting Liu. 2023. A static and dynamic attention framework for multi turn dialogue generation. ACM Transactions on Information Systems 41, 1 (2023), 1\u201330.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2021.3119969"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2021.3124640"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-emnlp.442"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3808220","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,16]],"date-time":"2026-05-16T15:43:36Z","timestamp":1778946216000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3808220"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,11]]},"references-count":65,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3808220"],"URL":"https:\/\/doi.org\/10.1145\/3808220","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,11]]},"assertion":[{"value":"2024-08-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-04","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}