{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T18:08:49Z","timestamp":1784138929426,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":77,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,7,19]],"date-time":"2026-07-19T00:00:00Z","timestamp":1784419200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"National Natural Science Foundation of China","award":["No.62502404"],"award-info":[{"award-number":["No.62502404"]}]},{"name":"Anhui Natural Science Foundation","award":["No.2508085ZD006"],"award-info":[{"award-number":["No.2508085ZD006"]}]},{"name":"National Natural Science Foundation of China","award":["No.U22B2059"],"award-info":[{"award-number":["No.U22B2059"]}]},{"name":"Hong Kong Research Grants Council Research Impact Fund","award":["No.R1015-23"],"award-info":[{"award-number":["No.R1015-23"]}]},{"name":"Hong Kong Research Grants Council Collaborative Research Fund","award":["No.C1043-24GF"],"award-info":[{"award-number":["No.C1043-24GF"]}]},{"name":"Hong Kong Research Grants Council General Research Fund","award":["No.11218325"],"award-info":[{"award-number":["No.11218325"]}]},{"name":"Institute of Digital Medicine of City University of Hong Kong","award":["No.9229503"],"award-info":[{"award-number":["No.9229503"]}]},{"name":"Huawei Innovation Research Program","award":["-"],"award-info":[{"award-number":["-"]}]},{"name":"Tencent Rhino-Bird Focused Research Program","award":["-"],"award-info":[{"award-number":["-"]}]},{"name":"Tencent University Cooperation Project","award":["-"],"award-info":[{"award-number":["-"]}]},{"name":"CCF-Didi Gaia Scholars Research Fund","award":["-"],"award-info":[{"award-number":["-"]}]},{"name":"CCF-Kuaishou Large Model Explorer Fund","award":["No.2025008"],"award-info":[{"award-number":["No.2025008"]}]},{"name":"Kuaishou University Cooperation Project","award":["-"],"award-info":[{"award-number":["-"]}]},{"name":"Bytedance","award":["-"],"award-info":[{"award-number":["-"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,7,20]]},"DOI":"10.1145\/3805712.3809599","type":"proceedings-article","created":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T14:28:19Z","timestamp":1783693699000},"page":"2084-2095","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["ProEchoMem: Enhancing Long Video Understanding via Multi-Trace Probe-Echo Memory"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3971-9907","authenticated-orcid":false,"given":"Derong","family":"Xu","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, Hefei, China and City University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-7204-8399","authenticated-orcid":false,"given":"Yanxin","family":"Chen","sequence":"additional","affiliation":[{"name":"City University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5976-0707","authenticated-orcid":false,"given":"Wanyu","family":"Wang","sequence":"additional","affiliation":[{"name":"City University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4712-3676","authenticated-orcid":false,"given":"Pengyue","family":"Jia","sequence":"additional","affiliation":[{"name":"City University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-2579-8783","authenticated-orcid":false,"given":"Chao","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China and City University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0073-0172","authenticated-orcid":false,"given":"Maolin","family":"Wang","sequence":"additional","affiliation":[{"name":"City University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9594-1919","authenticated-orcid":false,"given":"Yiqi","family":"Wang","sequence":"additional","affiliation":[{"name":"Michigan State University, East Lansing, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5721-0293","authenticated-orcid":false,"given":"Jipeng","family":"Qiang","sequence":"additional","affiliation":[{"name":"Yangzhou University, Yangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4450-2251","authenticated-orcid":false,"given":"Xuetao","family":"Wei","sequence":"additional","affiliation":[{"name":"Southern University of Science and Technology, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1395-261X","authenticated-orcid":false,"given":"Hongzhi","family":"Yin","sequence":"additional","affiliation":[{"name":"University of Queensland, Brisbane, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4246-5386","authenticated-orcid":false,"given":"Tong","family":"Xu","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2926-4416","authenticated-orcid":false,"given":"Xiangyu","family":"Zhao","sequence":"additional","affiliation":[{"name":"City University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,19]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al.","author":"Achiam Josh","year":"2023","unstructured":"Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al., 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3627673.3680088"},{"key":"e_1_3_2_1_3_1","unstructured":"Shuai Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge Sibo Song Kai Dang Peng Wang Shijie Wang Jun Tang et al. 2025. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923 (2025)."},{"key":"e_1_3_2_1_4_1","volume-title":"Shikra: Unleashing multimodal llm's referential dialogue magic. arXiv preprint arXiv:2306.15195","author":"Chen Keqin","year":"2023","unstructured":"Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. 2023. Shikra: Unleashing multimodal llm's referential dialogue magic. arXiv preprint arXiv:2306.15195 (2023)."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3730184"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3729896"},{"key":"e_1_3_2_1_7_1","volume-title":"Longvila: Scaling long-context visual language models for long videos. arXiv preprint arXiv:2408.10188","author":"Chen Yukang","year":"2024","unstructured":"Yukang Chen, Fuzhao Xue, Dacheng Li, Qinghao Hu, Ligeng Zhu, Xiuyu Li, Yunhao Fang, Haotian Tang, Shang Yang, Zhijian Liu, et al., 2024. Longvila: Scaling long-context visual language models for long videos. arXiv preprint arXiv:2408.10188 (2024)."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3729899"},{"key":"e_1_3_2_1_9_1","volume-title":"Mem0: Building production-ready ai agents with scalable long-term memory. arXiv preprint arXiv:2504.19413","author":"Chhikara Prateek","year":"2025","unstructured":"Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. 2025. Mem0: Building production-ready ai agents with scalable long-term memory. arXiv preprint arXiv:2504.19413 (2025)."},{"key":"e_1_3_2_1_10_1","volume-title":"Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2156-2166","author":"Dong Xingning","year":"2024","unstructured":"Xingning Dong, Zipeng Feng, Chunluan Zhou, Xuzheng Yu, Ming Yang, and Qingpei Guo. 2024. M2-RAAP: A multi-modal recipe for advancing adaptation-based pre-training towards effective and efficient zero-shot video-text retrieval. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2156-2166."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3730107"},{"key":"e_1_3_2_1_12_1","volume-title":"Robert Osazuwa Ness, and Jonathan Larson","author":"Edge Darren","year":"2024","unstructured":"Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2024. From local to global: A graph rag approach to query-focused summarization. arXiv preprint arXiv:2404.16130 (2024)."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3637528.3671470"},{"key":"e_1_3_2_1_14_1","volume-title":"European Conference on Computer Vision. Springer, 75-92","author":"Fan Yue","year":"2024","unstructured":"Yue Fan, Xiaojian Ma, Rujie Wu, Yuntao Du, Jiaqi Li, Zhi Gao, and Qing Li. 2024b. Videoagent: A memory-augmented multimodal agent for video understanding. In European Conference on Computer Vision. Springer, 75-92."},{"key":"e_1_3_2_1_15_1","volume-title":"Lightmem: Lightweight and efficient memory-augmented generation. arXiv preprint arXiv:2510.18866","author":"Fang Jizhan","year":"2025","unstructured":"Jizhan Fang, Xinle Deng, Haoming Xu, Ziyan Jiang, Yuqi Tang, Ziwen Xu, Shumin Deng, Yunzhi Yao, Mengru Wang, Shuofei Qiao, et al., 2025. Lightmem: Lightweight and efficient memory-augmented generation. arXiv preprint arXiv:2510.18866 (2025)."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01457"},{"key":"e_1_3_2_1_17_1","volume-title":"Lightrag: Simple and fast retrieval-augmented generation.","author":"Guo Zirui","year":"2024","unstructured":"Zirui Guo, Lianghao Xia, Yanhua Yu, Tu Ao, and Chao Huang. 2024. Lightrag: Simple and fast retrieval-augmented generation. (2024)."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-1902"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3730015"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3730083"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.3758\/BF03202365"},{"key":"e_1_3_2_1_22_1","volume-title":"Schema abstraction'' in a multiple-trace memory model. Psychological review","author":"Hintzman Douglas L","year":"1986","unstructured":"Douglas L Hintzman. 1986. ''Schema abstraction'' in a multiple-trace memory model. Psychological review, Vol. 93, 4 (1986), 411."},{"key":"e_1_3_2_1_23_1","volume-title":"Judgments of frequency and recognition memory in a multiple-trace memory model. Psychological review","author":"Hintzman Douglas L","year":"1988","unstructured":"Douglas L Hintzman. 1988. Judgments of frequency and recognition memory in a multiple-trace memory model. Psychological review, Vol. 95, 4 (1988), 528."},{"key":"e_1_3_2_1_24_1","unstructured":"Wenyi Hong Weihan Wang Ming Ding Wenmeng Yu Qingsong Lv Yan Wang Yean Cheng Shiyu Huang Junhui Ji Zhao Xue et al. 2024. Cogvlm2: Visual language models for image and video understanding. arXiv preprint arXiv:2408.16500 (2024)."},{"key":"e_1_3_2_1_25_1","volume-title":"Unsupervised dense information retrieval with contrastive learning. arXiv preprint arXiv:2112.09118","author":"Izacard Gautier","year":"2021","unstructured":"Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2021. Unsupervised dense information retrieval with contrastive learning. arXiv preprint arXiv:2112.09118 (2021)."},{"key":"e_1_3_2_1_26_1","volume-title":"Videorag: Retrieval-augmented generation over video corpus. arXiv preprint arXiv:2501.05874","author":"Jeong Soyeong","year":"2025","unstructured":"Soyeong Jeong, Kangsan Kim, Jinheon Baek, and Sung Ju Hwang. 2025. Videorag: Retrieval-augmented generation over video corpus. arXiv preprint arXiv:2501.05874 (2025)."},{"key":"e_1_3_2_1_27_1","volume-title":"VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management. arXiv preprint arXiv:2512.04540","author":"Jin Hongbo","year":"2025","unstructured":"Hongbo Jin, Qingyuan Wang, Wenhao Zhang, Yang Liu, and Sijie Cheng. 2025. VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management. arXiv preprint arXiv:2512.04540 (2025)."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01300"},{"key":"e_1_3_2_1_29_1","volume-title":"Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih.","author":"Karpukhin Vladimir","year":"2020","unstructured":"Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick SH Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering.. In EMNLP (1). 6769-6781."},{"key":"e_1_3_2_1_30_1","volume-title":"Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326","author":"Li Bo","year":"2024","unstructured":"Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, et al., 2024b. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326 (2024)."},{"key":"e_1_3_2_1_31_1","volume-title":"European Conference on Computer Vision. Springer, 323-340","author":"Li Yanwei","year":"2024","unstructured":"Yanwei Li, Chengyao Wang, and Jiaya Jia. 2024a. Llama-vid: An image is worth 2 tokens in large language models. In European Conference on Computer Vision. Springer, 323-340."},{"key":"e_1_3_2_1_32_1","volume-title":"Towards general text embeddings with multi-stage contrastive learning. arXiv preprint arXiv:2308.03281","author":"Li Zehan","year":"2023","unstructured":"Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023. Towards general text embeddings with multi-stage contrastive learning. arXiv preprint arXiv:2308.03281 (2023)."},{"key":"e_1_3_2_1_33_1","volume-title":"Hippomm: Hippocampal-inspired multimodal memory for long audiovisual event understanding. arXiv preprint arXiv:2504.10739","author":"Lin Yueqian","year":"2025","unstructured":"Yueqian Lin, Qinsi Wang, Hancheng Ye, Yuzhe Fu, Hai Li, Yiran Chen, et al., 2025. Hippomm: Hippocampal-inspired multimodal memory for long audiovisual event understanding. arXiv preprint arXiv:2504.10739 (2025)."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3729984"},{"key":"e_1_3_2_1_35_1","volume-title":"Visual instruction tuning. Advances in neural information processing systems","author":"Liu Haotian","year":"2023","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. Advances in neural information processing systems, Vol. 36 (2023), 34892-34916."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3730066"},{"key":"e_1_3_2_1_37_1","volume-title":"listening, remembering, and reasoning: A multimodal agent with long-term memory. arXiv preprint arXiv:2508.09736","author":"Long Lin","year":"2025","unstructured":"Lin Long, Yichen He, Wentao Ye, Yiyuan Pan, Yuan Lin, Hang Li, Junbo Zhao, and Wei Li. 2025. Seeing, listening, remembering, and reasoning: A multimodal agent with long-term memory. arXiv preprint arXiv:2508.09736 (2025)."},{"key":"e_1_3_2_1_38_1","volume-title":"Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension. arXiv preprint arXiv:2411.13093","author":"Luo Yongdong","year":"2024","unstructured":"Yongdong Luo, Xiawu Zheng, Xiao Yang, Guilin Li, Haojia Lin, Jinfa Huang, Jiayi Ji, Fei Chao, Jiebo Luo, and Rongrong Ji. 2024. Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension. arXiv preprint arXiv:2411.13093 (2024)."},{"key":"e_1_3_2_1_39_1","volume-title":"Drvideo: Document retrieval based long video understanding. arXiv preprint arXiv:2406.12846","author":"Ma Ziyu","year":"2024","unstructured":"Ziyu Ma, Chenhui Gou, Hengcan Shi, Bin Sun, Shutao Li, Hamid Rezatofighi, and Jianfei Cai. 2024. Drvideo: Document retrieval based long video understanding. arXiv preprint arXiv:2406.12846 (2024)."},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3729945"},{"key":"e_1_3_2_1_41_1","volume-title":"Video-xl-2: Towards very long-video understanding through task-aware kv sparsification. arXiv preprint arXiv:2506.19225","author":"Qin Minghao","year":"2025","unstructured":"Minghao Qin, Xiangrui Liu, Zhengyang Liang, Yan Shu, Huaying Yuan, Juenjie Zhou, Shitao Xiao, Bo Zhao, and Zheng Liu. 2025a. Video-xl-2: Towards very long-video understanding through task-aware kv sparsification. arXiv preprint arXiv:2506.19225 (2025)."},{"key":"e_1_3_2_1_42_1","volume-title":"International conference on machine learning. PMLR, 28492-28518","author":"Radford Alec","year":"2023","unstructured":"Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International conference on machine learning. PMLR, 28492-28518."},{"key":"e_1_3_2_1_43_1","volume-title":"VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos. arXiv preprint arXiv:2502.01549","author":"Ren Xubin","year":"2025","unstructured":"Xubin Ren, Lingrui Xu, Long Xia, Shuaiqiang Wang, Dawei Yin, and Chao Huang. 2025. VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos. arXiv preprint arXiv:2502.01549 (2025)."},{"key":"e_1_3_2_1_44_1","volume-title":"Colbertv2: Effective and efficient retrieval via lightweight late interaction. arXiv preprint arXiv:2112.01488","author":"Santhanam Keshav","year":"2021","unstructured":"Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2021. Colbertv2: Effective and efficient retrieval via lightweight late interaction. arXiv preprint arXiv:2112.01488 (2021)."},{"key":"e_1_3_2_1_45_1","volume-title":"LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding. In Forty-second International Conference on Machine Learning.","author":"Shen Xiaoqian","unstructured":"Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu, Jun Chen, Chenchen Zhu, Zechun Liu, Fanyi Xiao, Balakrishnan Varadarajan, Florian Bordes, et al., [n.d.]. LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding. In Forty-second International Conference on Machine Learning."},{"key":"e_1_3_2_1_46_1","volume-title":"Longvu: Spatiotemporal adaptive compression for long video-language understanding. arXiv preprint arXiv:2410.17434","author":"Shen Xiaoqian","year":"2024","unstructured":"Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu, Jun Chen, Chenchen Zhu, Zechun Liu, Fanyi Xiao, Balakrishnan Varadarajan, Florian Bordes, et al., 2024. Longvu: Spatiotemporal adaptive compression for long video-language understanding. arXiv preprint arXiv:2410.17434 (2024)."},{"key":"e_1_3_2_1_47_1","volume-title":"Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding. arXiv preprint arXiv:2510.14032","author":"Shen Xiaoqian","year":"2025","unstructured":"Xiaoqian Shen, Wenxuan Zhang, Jun Chen, and Mohamed Elhoseiny. 2025. Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding. arXiv preprint arXiv:2510.14032 (2025)."},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02436"},{"key":"e_1_3_2_1_49_1","volume-title":"Mpnet: Masked and permuted pre-training for language understanding. Advances in neural information processing systems","author":"Song Kaitao","year":"2020","unstructured":"Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020. Mpnet: Masked and permuted pre-training for language understanding. Advances in neural information processing systems, Vol. 33 (2020), 16857-16867."},{"key":"e_1_3_2_1_50_1","unstructured":"Yunlong Tang Jing Bi Siting Xu Luchuan Song Susan Liang Teng Wang Daoan Zhang Jie An Jingyang Lin Rongyi Zhu et al. 2025. Video understanding with large language models: A survey. IEEE Transactions on Circuits and Systems for Video Technology (2025)."},{"key":"e_1_3_2_1_51_1","volume-title":"Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning. arXiv preprint arXiv:2506.13654","author":"Tian Shulin","year":"2025","unstructured":"Shulin Tian, Ruiqi Wang, Hongming Guo, Penghao Wu, Yuhao Dong, Xiuying Wang, Jingkang Yang, Hao Zhang, Hongyuan Zhu, and Ziwei Liu. 2025. Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning. arXiv preprint arXiv:2506.13654 (2025)."},{"key":"e_1_3_2_1_52_1","volume-title":"Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. arXiv preprint arXiv:2212.10509","author":"Trivedi Harsh","year":"2022","unstructured":"Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. arXiv preprint arXiv:2212.10509 (2022)."},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.01935"},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.02131"},{"key":"e_1_3_2_1_55_1","volume-title":"European Conference on Computer Vision. Springer, 58-76","author":"Wang Xiaohan","year":"2024","unstructured":"Xiaohan Wang, Yuhui Zhang, Orr Zohar, and Serena Yeung-Levy. 2024. Videoagent: Long-form video understanding with large language model as agent. In European Conference on Computer Vision. Springer, 58-76."},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"crossref","unstructured":"Ye Wang Ziheng Wang Boshen Xu Yang Du Kejun Lin Zihan Xiao Zihao Yue Jianzhong Ju Liang Zhang Dingyi Yang et al. 2025c. Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding. arXiv preprint arXiv:2503.13377 (2025).","DOI":"10.1109\/ICASSP49660.2025.10888314"},{"key":"e_1_3_2_1_57_1","volume-title":"Episodic memory representation for long-form video understanding. arXiv preprint arXiv:2508.09486","author":"Wang Yun","year":"2025","unstructured":"Yun Wang, Long Zhang, Jingren Liu, Jiaqi Yan, Zhanjie Zhang, Jiahao Zheng, Xun Yang, Dapeng Wu, Xiangyu Chen, and Xuelong Li. 2025d. Episodic memory representation for long-form video understanding. arXiv preprint arXiv:2508.09486 (2025)."},{"key":"e_1_3_2_1_58_1","volume-title":"MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding. arXiv preprint arXiv:2510.07915","author":"Wu Peiran","year":"2025","unstructured":"Peiran Wu, Zhuorui Yu, Yunze Liu, Chi-Hao Wu, Enmin Zhou, and Junxiao Shen. 2025. MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding. arXiv preprint arXiv:2510.07915 (2025)."},{"key":"e_1_3_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657878"},{"key":"e_1_3_2_1_60_1","unstructured":"Derong Xu Yi Wen Pengyue Jia Yingyi Zhang Yichao Wang Huifeng Guo Ruiming Tang Xiangyu Zhao Enhong Chen Tong Xu et al. 2025b. Towards Multi-Granularity Memory Association and Selection for Long-Term Conversational Agents. arXiv preprint arXiv:2505.19549 (2025)."},{"key":"e_1_3_2_1_61_1","volume-title":"The Fourteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=i2yIvZARnG","author":"Xu Derong","year":"2026","unstructured":"Derong Xu, Yi Wen, Pengyue Jia, Yingyi Zhang, Wenlin Zhang, Yichao Wang, Huifeng Guo, Ruiming Tang, Xiangyu Zhao, Enhong Chen, and Tong Xu. 2026. From Single to Multi-Granularity: Toward Long-Term Memory Association and Selection of Conversational Agents. In The Fourteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=i2yIvZARnG"},{"key":"e_1_3_2_1_62_1","volume-title":"See Kiong Ng, and Jiashi Feng","author":"Xu Lin","year":"2024","unstructured":"Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng. 2024. Pllava: Parameter-free llava extension from images to videos for video dense captioning. arXiv preprint arXiv:2404.16994 (2024)."},{"key":"e_1_3_2_1_63_1","volume-title":"A-mem: Agentic memory for llm agents. arXiv preprint arXiv:2502.12110","author":"Xu Wujiang","year":"2025","unstructured":"Wujiang Xu, Kai Mei, Hang Gao, Juntao Tan, Zujie Liang, and Yongfeng Zhang. 2025a. A-mem: Agentic memory for llm agents. arXiv preprint arXiv:2502.12110 (2025)."},{"key":"e_1_3_2_1_64_1","volume-title":"E-vrag: Enhancing long video understanding with resource-efficient retrieval augmented generation. arXiv preprint arXiv:2508.01546","author":"Xu Zeyu","year":"2025","unstructured":"Zeyu Xu, Junkang Zhang, Qiang Wang, and Yi Liu. 2025c. E-vrag: Enhancing long video understanding with resource-efficient retrieval augmented generation. arXiv preprint arXiv:2508.01546 (2025)."},{"key":"e_1_3_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02690"},{"key":"e_1_3_2_1_66_1","volume-title":"WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning. arXiv preprint arXiv:2512.02425","author":"Yeo Woongyeong","year":"2025","unstructured":"Woongyeong Yeo, Kangsan Kim, Jaehong Yoon, and Sung Ju Hwang. 2025. WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning. arXiv preprint arXiv:2512.02425 (2025)."},{"key":"e_1_3_2_1_67_1","volume-title":"VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding. arXiv preprint arXiv:2512.12360","author":"Yin Yufei","year":"2025","unstructured":"Yufei Yin, Qianke Meng, Minghao Chen, Jiajun Ding, Zhenwei Shao, and Zhou Yu. 2025. VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding. arXiv preprint arXiv:2512.12360 (2025)."},{"key":"e_1_3_2_1_68_1","volume-title":"Vismem: Latent vision memory unlocks potential of vision-language models. arXiv preprint arXiv:2511.11007","author":"Yu Xinlei","year":"2025","unstructured":"Xinlei Yu, Chengming Xu, Guibin Zhang, Zhangquan Chen, Yudong Zhang, Yongbo He, Peng-Tao Jiang, Jiangning Zhang, Xiaobin Hu, and Shuicheng Yan. 2025. Vismem: Latent vision memory unlocks potential of vision-language models. arXiv preprint arXiv:2511.11007 (2025)."},{"key":"e_1_3_2_1_69_1","volume-title":"Memory-enhanced retrieval augmentation for long video understanding. arXiv preprint arXiv:2503.09149","author":"Yuan Huaying","year":"2025","unstructured":"Huaying Yuan, Zheng Liu, Minghao Qin, Hongjin Qian, Yan Shu, Zhicheng Dou, Ji-Rong Wen, and Nicu Sebe. 2025. Memory-enhanced retrieval augmentation for long video understanding. arXiv preprint arXiv:2503.09149 (2025)."},{"key":"e_1_3_2_1_70_1","unstructured":"Boqiang Zhang Kehan Li Zesen Cheng Zhiqiang Hu Yuqian Yuan Guanzheng Chen Sicong Leng Yuming Jiang Hang Zhang Xin Li et al. 2025. Videollama 3: Frontier multimodal foundation models for image and video understanding. arXiv preprint arXiv:2501.13106 (2025)."},{"key":"e_1_3_2_1_71_1","volume-title":"2026 c. MemSearch-o1: Empowering Large Language Models with Reasoning-Aligned Memory Growth in Agentic Search. arXiv preprint arXiv:2604.17265","author":"Zhang Sheng","year":"2026","unstructured":"Sheng Zhang, Junyi Li, Yingyi Zhang, Pengyue Jia, Yichao Wang, Xiaowei Qian, Wenlin Zhang, Maolin Wang, Yong Liu, and Xiangyu Zhao. 2026 c. MemSearch-o1: Empowering Large Language Models with Reasoning-Aligned Memory Growth in Agentic Search. arXiv preprint arXiv:2604.17265 (2026)."},{"key":"e_1_3_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v40i19.38679"},{"key":"e_1_3_2_1_73_1","volume-title":"The Fourteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=f7p0F2X6XN","author":"Zhang Yingyi","year":"2026","unstructured":"Yingyi Zhang, Junyi Li, Wenlin Zhang, Pengyue Jia, Xianneng Li, Yichao Wang, Derong Xu, Yi Wen, Huifeng Guo, Yong Liu, and Xiangyu Zhao. 2026 b. Evoking User Memory: Personalizing LLM via Recollection-Familiarity Adaptive Retrieval. In The Fourteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=f7p0F2X6XN"},{"key":"e_1_3_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3729936"},{"key":"e_1_3_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657929"},{"key":"e_1_3_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i17.29946"},{"key":"e_1_3_2_1_77_1","volume-title":"Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592","author":"Zhu Deyao","year":"2023","unstructured":"Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592 (2023)."}],"event":{"name":"SIGIR '26: The 49th International ACM SIGIR Conference on Research and Development in Information Retrieval","location":"Melbourne VIC Australia","sponsor":["SIGIR ACM Special Interest Group on Information Retrieval"]},"container-title":["Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval"],"original-title":[],"deposited":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T17:24:54Z","timestamp":1784136294000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3805712.3809599"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,19]]},"references-count":77,"alternative-id":["10.1145\/3805712.3809599","10.1145\/3805712"],"URL":"https:\/\/doi.org\/10.1145\/3805712.3809599","relation":{},"subject":[],"published":{"date-parts":[[2026,7,19]]},"assertion":[{"value":"2026-07-19","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}