{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T08:09:39Z","timestamp":1765008579902,"version":"3.46.0"},"publisher-location":"New York, NY, USA","reference-count":45,"publisher":"ACM","funder":[{"name":"Japan Society for the Promotion of Science","award":["23K24868, 25K00161"],"award-info":[{"award-number":["23K24868, 25K00161"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,12,9]]},"DOI":"10.1145\/3743093.3770950","type":"proceedings-article","created":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T08:06:16Z","timestamp":1765008376000},"page":"1-7","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-3649-8651","authenticated-orcid":false,"given":"Junan","family":"Chen","sequence":"first","affiliation":[{"name":"Nagoya University, Nagoya, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8976-2922","authenticated-orcid":false,"given":"Trung Thanh","family":"Nguyen","sequence":"additional","affiliation":[{"name":"Nagoya University, Nagoya, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3041-4330","authenticated-orcid":false,"given":"Takahiro","family":"Komamizu","sequence":"additional","affiliation":[{"name":"Nagoya University, Nagoya, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3942-9296","authenticated-orcid":false,"given":"Ichiro","family":"Ide","sequence":"additional","affiliation":[{"name":"Nagoya University, Nagoya, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,12,6]]},"reference":[{"key":"e_1_3_3_2_2_2","unstructured":"Shuai Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge Sibo Song Kai Dang Peng Wang Shijie Wang Jun Tang Humen Zhong Yuanzhi Zhu Mingkun Yang Zhaohai Li Jianqiang Wan Pengfei Wang Wei Ding Zheren Fu Yiheng Xu Jiabo Ye Xi Zhang Tianbao Xie Zesen Cheng Hang Zhang Zhibo Yang Haiyang Xu and Junyang Lin. 2025. Qwen2.5-VL Technical Report. Computing Research Repository arXiv Preprint arXiv:https:\/\/arXiv.org\/abs\/2502.13923 (2025)."},{"key":"e_1_3_3_2_3_2","first-page":"65","volume-title":"Proc. ACL2005 Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization","author":"Banerjee Satanjeev","year":"2005","unstructured":"Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proc. ACL2005 Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization. 65\u201372."},{"key":"e_1_3_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"e_1_3_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.5555\/2002472.2002497"},{"key":"e_1_3_3_2_6_2","unstructured":"Shoufa Chen Chongjian Ge Zhan Tong Jiangliu Wang Yibing Song Jue Wang and Ping Luo. 2022. AdaptFormer: Adapting vision Transformers for scalable visual recognition. Advances in Neural Information Processing Systems 35 (2022) 16664\u201316678."},{"key":"e_1_3_3_2_7_2","unstructured":"Sihan Chen Handong Li Qunbo Wang Zijia Zhao Mingzhen Sun Xinxin Zhu and Jing Liu. 2023. VAST: A Vision-Audio-Subtitle-Text omni-modality foundation model and dataset. Advances in Neural Information Processing Systems 36 (2023) 72842\u201372866."},{"key":"e_1_3_3_2_8_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et\u00a0al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. Computing Research Repository arXiv Preprint arXiv:https:\/\/arXiv.org\/abs\/2010.11929 (2020)."},{"key":"e_1_3_3_2_9_2","doi-asserted-by":"crossref","unstructured":"Luciano Floridi and Massimo Chiriatti. 2020. GPT-3: Its nature scope limits and consequences. Minds and Machines 30 (2020) 681\u2013694.","DOI":"10.1007\/s11023-020-09548-1"},{"key":"e_1_3_3_2_10_2","first-page":"2790","volume-title":"Proc. 36th International Conference on Machine Learning","author":"Houlsby Neil","year":"2019","unstructured":"Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De\u00a0Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proc. 36th International Conference on Machine Learning. 2790\u20132799."},{"key":"e_1_3_3_2_11_2","first-page":"3","volume-title":"Proc. 10th International Conference on Learning Representations","author":"Hu Edward\u00a0J","year":"2022","unstructured":"Edward\u00a0J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of large language models. In Proc. 10th International Conference on Learning Representations. 3\u201316."},{"key":"e_1_3_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01353"},{"key":"e_1_3_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01318"},{"key":"e_1_3_3_2_14_2","first-page":"9837","volume-title":"Proc. 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Bing","year":"2021","unstructured":"Bing Li, Chunfeng Yuan, Weiming Hu, Ying Deng, Ying Shan, Zhongang Qi, and Ziqi Zhang. 2021. Open-book video captioning with retrieve-copy-generate network. In Proc. 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 9837\u20139846."},{"key":"e_1_3_3_2_15_2","first-page":"19730","volume-title":"Proc. 40th International Conference on Machine Learning","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping Language-Image Pre-training with frozen image encoders and large language models. In Proc. 40th International Conference on Machine Learning. 19730\u201319742."},{"key":"e_1_3_3_2_16_2","unstructured":"Junnan Li Ramprasaath Selvaraju Akhilesh Gotmare Shafiq Joty Caiming Xiong and Steven Chu\u00a0Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. Advances in Neural Information Processing Systems 34 (2021) 9694\u20139705."},{"key":"e_1_3_3_2_17_2","unstructured":"Kunchang Li Yinan He Yi Wang Yizhuo Li Wenhai Wang Ping Luo Yali Wang Limin Wang and Yu Qiao. 2023. VideoChat: Chat-centric video understanding. Computing Research Repository arXiv Preprint arXiv:https:\/\/arXiv.org\/abs\/2305.06355 (2023)."},{"key":"e_1_3_3_2_18_2","first-page":"74","volume-title":"Proc. ACL2004 Workshop on Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Proc. ACL2004 Workshop on Text Summarization Branches Out. 74\u201381."},{"key":"e_1_3_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01742"},{"key":"e_1_3_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_3_2_21_2","unstructured":"Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. Computing Research Repository arXiv Preprint arXiv:https:\/\/arXiv.org\/abs\/1711.05101 (2017)."},{"key":"e_1_3_3_2_22_2","doi-asserted-by":"crossref","unstructured":"Huaishao Luo Lei Ji Ming Zhong Yang Chen Wen Lei Nan Duan and Tianrui Li. 2022. CLIP4Clip: An empirical study of CLIP for end to end video Clip retrieval and captioning. Neurocomputing 508 (2022) 293\u2013304.","DOI":"10.1016\/j.neucom.2022.07.028"},{"key":"e_1_3_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00529"},{"key":"e_1_3_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01088"},{"key":"e_1_3_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.497"},{"key":"e_1_3_3_2_26_2","first-page":"311","volume-title":"Proc. 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proc. 40th Annual Meeting of the Association for Computational Linguistics. 311\u2013318."},{"key":"e_1_3_3_2_27_2","first-page":"8748","volume-title":"Proc. 38th International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong\u00a0Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proc. 38th International Conference on Machine Learning. 8748\u20138763."},{"key":"e_1_3_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298940"},{"key":"e_1_3_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.61"},{"key":"e_1_3_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.626"},{"key":"e_1_3_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01743"},{"key":"e_1_3_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413498"},{"key":"e_1_3_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00516"},{"key":"e_1_3_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3479207"},{"key":"e_1_3_3_2_35_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et\u00a0al. 2023. Llama 2: Open foundation and fine-tuned chat models. Computing Research Repository arXiv Preprint arXiv:https:\/\/arXiv.org\/abs\/2307.09288 (2023)."},{"key":"e_1_3_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_3_3_2_37_2","first-page":"38728","volume-title":"Proc. 40th International Conference on Machine Learning","author":"Xu Haiyang","year":"2023","unstructured":"Haiyang Xu, Qinghao Ye, Ming Yan, Yaya Shi, Jiabo Ye, Yuanhong Xu, Chenliang Li, Bin Bi, Qi Qian, Wei Wang, Guohai Xu, Ji Zhang, Songfang Huang, Fei Huang, and Jingren Zhou. 2023. mPLUG-2: A modularized multi-modal foundation model across text, image and video. In Proc. 40th International Conference on Machine Learning. 38728\u201338748."},{"key":"e_1_3_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.571"},{"key":"e_1_3_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01032"},{"key":"e_1_3_3_2_40_2","first-page":"1","volume-title":"Proc. 11th International Conference on Learning Representations","author":"Yang Taojiannan","year":"2023","unstructured":"Taojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang, Chen Chen, and Mu Li. 2023. AIM: Adapting Image Models for efficient video action recognition. In Proc. 11th International Conference on Learning Representations. 1\u201314."},{"key":"e_1_3_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01413"},{"key":"e_1_3_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01869"},{"key":"e_1_3_3_2_43_2","unstructured":"Shukang Yin Chaoyou Fu Sirui Zhao Ke Li Xing Sun Tong Xu and Enhong Chen. 2023. A survey on multimodal large language models. Computing Research Repository arXiv Preprint arXiv:https:\/\/arXiv.org\/abs\/2306.13549 (2023)."},{"key":"e_1_3_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.1137"},{"key":"e_1_3_3_2_45_2","unstructured":"Yuexiang Zhai Shengbang Tong Xiao Li Mu Cai Qing Qu Yong\u00a0Jae Lee and Yi Ma. 2023. Investigating the catastrophic forgetting in multimodal large language models. Computing Research Repository arXiv Preprint arXiv:https:\/\/arXiv.org\/abs\/2309.10313 (2023)."},{"key":"e_1_3_3_2_46_2","unstructured":"Renrui Zhang Jiaming Han Chris Liu Peng Gao Aojun Zhou Xiangfei Hu Shilin Yan Pan Lu Hongsheng Li and Yu Qiao. 2023. LLaMA-Adapter: Efficient fine-tuning of language models with zero-init attention. Computing Research Repository arXiv Preprint arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2303.16199 (2023)."}],"event":{"name":"MMAsia '25: ACM Multimedia Asia","location":"Kuala Lumpur Malaysia","acronym":"MMAsia '25","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 7th ACM International Conference on Multimedia in Asia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3743093.3770950","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T08:06:52Z","timestamp":1765008412000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3743093.3770950"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,6]]},"references-count":45,"alternative-id":["10.1145\/3743093.3770950","10.1145\/3743093"],"URL":"https:\/\/doi.org\/10.1145\/3743093.3770950","relation":{},"subject":[],"published":{"date-parts":[[2025,12,6]]},"assertion":[{"value":"2025-12-06","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}