{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T15:36:50Z","timestamp":1779377810583,"version":"3.53.1"},"reference-count":429,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2025,3,14]],"date-time":"2025-03-14T00:00:00Z","timestamp":1741910400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Science Fund for Distinguished Young Scholars","award":["62025205"],"award-info":[{"award-number":["62025205"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62032020, 62432007, 62272390"],"award-info":[{"award-number":["62032020, 62432007, 62272390"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2025,5,31]]},"abstract":"<jats:p>The intelligent dialogue system, aiming at communicating with humans harmoniously with natural language, is brilliant for promoting the advancement of human-machine interaction in the era of artificial intelligence. With the gradually complex human-computer interaction requirements, it is difficult for traditional text-based dialogue system to meet the demands for more vivid and convenient interaction. Consequently, Visual-Context Augmented Dialogue (VAD) System, which has the potential to communicate with humans by perceiving and understanding multimodal information (i.e., visual context in images or videos, textual dialogue history), has become a predominant research paradigm. Benefiting from the consistency and complementarity between visual and textual context, VAD possesses the potential to generate engaging and context-aware responses. To depict the development of VAD, we first characterize the concept model of VAD and then present its generic system architecture to illustrate the system workflow, followed by a summary of multimodal fusion techniques. Subsequently, several research challenges and representative works are investigated, followed by the summary of authoritative benchmarks and real-world application of VAD. We conclude this article by putting forward some open issues and promising research trends for VAD, e.g., the cognitive mechanisms of human-machine dialogue under cross-modal dialogue context, mobile and lightweight deployment of VAD.<\/jats:p>","DOI":"10.1145\/3715098","type":"journal-article","created":{"date-parts":[[2025,1,28]],"date-time":"2025-01-28T15:44:14Z","timestamp":1738079054000},"page":"1-59","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Enabling Harmonious Human-Machine Interaction with Visual-Context Augmented Dialogue System: A Review"],"prefix":"10.1145","volume":"43","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6959-7237","authenticated-orcid":false,"given":"Hao","family":"Wang","sequence":"first","affiliation":[{"name":"School of Computer Science, Northwestern Polytechnical University, Xi\u2019an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6097-2467","authenticated-orcid":false,"given":"Bin","family":"Guo","sequence":"additional","affiliation":[{"name":"School of Computer Science, Northwestern Polytechnical University, Xi\u2019an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-5847-9280","authenticated-orcid":false,"given":"Yating","family":"Zeng","sequence":"additional","affiliation":[{"name":"School of Computer Science, Northwestern Polytechnical University, Xi\u2019an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4607-1440","authenticated-orcid":false,"given":"Mengqi","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Computer Science, Northwestern Polytechnical University, Xi\u2019an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9051-5865","authenticated-orcid":false,"given":"Yasan","family":"Ding","sequence":"additional","affiliation":[{"name":"School of Computer Science, Northwestern Polytechnical University, Xi\u2019an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6411-4486","authenticated-orcid":false,"given":"Ying","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Northwestern Polytechnical University, Xi\u2019an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4149-839X","authenticated-orcid":false,"given":"Lina","family":"Yao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, University of New South Wales, Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9905-3238","authenticated-orcid":false,"given":"Zhiwen","family":"Yu","sequence":"additional","affiliation":[{"name":"School of Computer Science, Northwestern Polytechnical University, Xi\u2019an, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,3,14]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et al. 2023. Gpt-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_2_3_2","first-page":"3947","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Agrawal Garima","year":"2024","unstructured":"Garima Agrawal, Tharindu Kumarage, Zeyad Alghamdi, and Huan Liu. 2024. Can knowledge graphs reduce hallucinations in LLMs? A survey. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3947\u20133960."},{"key":"e_1_3_2_4_2","first-page":"225","volume-title":"Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop","author":"Ahn Janice","year":"2024","unstructured":"Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin. 2024. Large language models for mathematical reasoning: Progresses and challenges. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, 225\u2013237."},{"key":"e_1_3_2_5_2","first-page":"7558","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Alamri Huda","year":"2019","unstructured":"Huda Alamri, Vincent Cartillier, Abhishek Das, Jue Wang, Anoop Cherian, Irfan Essa, Dhruv Batra, Tim K. Marks, Chiori Hori, Peter Anderson, et al. 2019. Audio visual scene-aware dialog. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 7558\u20137567."},{"key":"e_1_3_2_6_2","first-page":"23716","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Alayrac Jean-Baptiste","year":"2022","unstructured":"Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: A visual language model for few-shot learning. In Proceedings of the Advances in Neural Information Processing Systems, 23716\u201323736."},{"key":"e_1_3_2_7_2","first-page":"39","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Andreas Jacob","year":"2016","unstructured":"Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016. Neural module networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 39\u201348."},{"key":"e_1_3_2_8_2","first-page":"6836","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Arnab Anurag","year":"2021","unstructured":"Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lu\u010di\u0107, and Cordelia Schmid. 2021. Vivit: A video vision transformer. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 6836\u20136846."},{"key":"e_1_3_2_9_2","unstructured":"Anas Awadalla Irena Gao Josh Gardner Jack Hessel Yusuf Hanafy Wanrong Zhu Kalyani Marathe Yonatan Bitton Samir Gadre Shiori Sagawa et al. 2023. OpenFlamingo: An open-source framework for training large autoregressive vision-language models. arXiv:2308.01390. Retrieved from https:\/\/arxiv.org\/abs\/2308.01390"},{"key":"e_1_3_2_10_2","unstructured":"Jinze Bai Shuai Bai Yunfei Chu Zeyu Cui Kai Dang Xiaodong Deng Yang Fan Wenbin Ge Yu Han Fei Huang et al. 2023. Qwen technical report. arXiv:2309.16609. Retrieved from https:\/\/arxiv.org\/abs\/2309.16609"},{"issue":"2","key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"423","DOI":"10.1109\/TPAMI.2018.2798607","article-title":"Multimodal machine learning: A survey and taxonomy","volume":"41","author":"Baltru\u0161aitis Tadas","year":"2018","unstructured":"Tadas Baltru\u0161aitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2018. Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence 41, 2 (2018), 423\u2013443.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_12_2","first-page":"354","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Bhardwaj Shweta","year":"2019","unstructured":"Shweta Bhardwaj, Mukundhan Srinivasan, and Mitesh M. Khapra. 2019. Efficient video classification using fewer frames. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 354\u2013363."},{"issue":"1","key":"e_1_3_2_13_2","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1016\/S0169-7552(98)00110-X","article-title":"The anatomy of a large-scale hypertextual web search engine","volume":"30","author":"Brin Sergey","year":"1998","unstructured":"Sergey Brin and Lawrence Page. 1998. The anatomy of a large-scale hypertextual web search engine. Computer Networks and ISDN Systems 30, 1\u20137 (1998), 107\u2013117.","journal-title":"Computer Networks and ISDN Systems"},{"key":"e_1_3_2_14_2","first-page":"1059","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Brock Andy","year":"2021","unstructured":"Andy Brock, Soham De, Samuel L. Smith, and Karen Simonyan. 2021. High-performance large-scale image recognition without normalization. In Proceedings of the International Conference on Machine Learning. PMLR, 1059\u20131071."},{"key":"e_1_3_2_15_2","first-page":"1877","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Proceedings of the Advances in Neural Information Processing Systems, 1877\u20131901."},{"key":"e_1_3_2_16_2","first-page":"5188","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Bulat Adrian","year":"2021","unstructured":"Adrian Bulat and Georgios Tzimiropoulos. 2021. Bit-Mixer: Mixed-precision networks with runtime bit-width selection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 5188\u20135197."},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","first-page":"13590","DOI":"10.18653\/v1\/2024.findings-acl.807","article-title":"The revolution of multimodal large language models: A survey","author":"Caffagni Davide","year":"2024","unstructured":"Davide Caffagni, Federico Cocchi, Luca Barsellotti, Nicholas Moratelli, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, and Rita Cucchiara. 2024. The revolution of multimodal large language models: A survey. In Findings of the Association for Computational Linguistics (ACL \u201924), 13590\u201313618.","journal-title":"Findings of the Association for Computational Linguistics (ACL \u201924)"},{"key":"e_1_3_2_18_2","first-page":"1818","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Caffagni Davide","year":"2024","unstructured":"Davide Caffagni, Federico Cocchi, Nicholas Moratelli, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2024. Wiki-LLaVA: Hierarchical retrieval-augmented generation for multimodal LLMs. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1818\u20131826."},{"key":"e_1_3_2_19_2","first-page":"6299","volume-title":"In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Carreira Joao","year":"2017","unstructured":"Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? A new model and the kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6299\u20136308."},{"issue":"3","key":"e_1_3_2_20_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3641289","article-title":"A survey on evaluation of large language models","volume":"15","author":"Chang Yupeng","year":"2024","unstructured":"Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology 15, 3 (2024), 1\u201345.","journal-title":"ACM Transactions on Intelligent Systems and Technology"},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","first-page":"520","DOI":"10.1109\/ICME.2019.00096","volume-title":"Proceedings of the 2019 IEEE International Conference on Multimedia and Expo (ICME \u201919)","author":"Chang Yen Wei","year":"2019","unstructured":"Yen Wei Chang and Wen-Hsiao Peng. 2019. Learning goal-oriented visual dialog agents: Imitating and surpassing analytic experts. In Proceedings of the 2019 IEEE International Conference on Multimedia and Expo (ICME \u201919). IEEE Computer Society, 520\u2013525."},{"issue":"5","key":"e_1_3_2_22_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3465055","article-title":"An attentive survey of attention models","volume":"12","author":"Chaudhari Sneha","year":"2021","unstructured":"Sneha Chaudhari, Varun Mithal, Gungor Polatkan, and Rohan Ramanath. 2021. An attentive survey of attention models. ACM Transactions on Intelligent Systems and Technology 12, 5 (2021), 1\u201332.","journal-title":"ACM Transactions on Intelligent Systems and Technology"},{"key":"e_1_3_2_23_2","first-page":"18103","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Cheng","year":"2022","unstructured":"Cheng Chen, Zhenshan Tan, Qingrong Cheng, Xin Jiang, Qun Liu, Yudong Zhu, and Xiaodong Gu. 2022. UTC: A unified transformer with inter-task contrastive learning for visual dialog. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 18103\u201318112."},{"key":"e_1_3_2_24_2","first-page":"17745","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"38","author":"Chen Delong","year":"2024","unstructured":"Delong Chen, Jianfeng Liu, Wenliang Dai, and Baoyuan Wang. 2024. Visual instruction tuning with polite flamingo. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 17745\u201317753."},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","first-page":"230","DOI":"10.18653\/v1\/2021.findings-acl.20","volume-title":"Findings of the Association for Computational Linguistics (ACL-IJCNLP \u201921)","author":"Chen Feilong","year":"2021","unstructured":"Feilong Chen, Xiuyi Chen, Fandong Meng, Peng Li, and Jie Zhou. 2021. GoG: Relation-aware graph-over-graph network for visual dialog. In Findings of the Association for Computational Linguistics (ACL-IJCNLP \u201921), 230\u2013243."},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","first-page":"436","DOI":"10.18653\/v1\/2021.findings-acl.38","volume-title":"Findings of the Association for Computational Linguistics (ACL-IJCNLP \u201921)","author":"Chen Feilong","year":"2021","unstructured":"Feilong Chen, Fandong Meng, Xiuyi Chen, Peng Li, and Jie Zhou. 2021. Multimodal incremental transformer with visual grounding for visual dialogue generation. In Findings of the Association for Computational Linguistics (ACL-IJCNLP \u201921), 436\u2013446."},{"key":"e_1_3_2_27_2","first-page":"7504","volume-title":"In Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Chen Feilong","year":"2020","unstructured":"Feilong Chen, Fandong Meng, Jiaming Xu, Peng Li, Bo Xu, and Jie Zhou. 2020. DMRM: A dual-channel multi-hop reasoning model for visual dialog. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34, 7504\u20137511."},{"issue":"1","key":"e_1_3_2_28_2","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1007\/s11633-022-1369-5","article-title":"VLP: A survey on vision-language pre-training","volume":"20","author":"Chen Fei-Long","year":"2023","unstructured":"Fei-Long Chen, Du-Zhen Zhang, Ming-Lun Han, Xiu-Yi Chen, Jing Shi, Shuang Xu, and Bo Xu. 2023. VLP: A survey on vision-language pre-training. Machine Intelligence Research 20, 1 (2023), 38\u201356.","journal-title":"Machine Intelligence Research"},{"key":"e_1_3_2_29_2","first-page":"742","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Chen Guobin","year":"2017","unstructured":"Guobin Chen, Wongun Choi, Xiang Yu, Tony Han, and Manmohan Chandraker. 2017. Learning efficient object detection models with knowledge distillation. In Proceedings of the Advances in Neural Information Processing Systems, 742\u2013751."},{"key":"e_1_3_2_30_2","first-page":"26540","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Gongwei","year":"2024","unstructured":"Gongwei Chen, Leyang Shen, Rui Shao, Xiang Deng, and Liqiang Nie. 2024. Lion: Empowering multimodal large language model with dual-level visual knowledge. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 26540\u201326550."},{"key":"e_1_3_2_31_2","unstructured":"Guiming Hardy Chen Shunian Chen Ruifei Zhang Junying Chen Xiangbo Wu Zhiyi Zhang Zhihong Chen Jianquan Li Xiang Wan and Benyou Wang. 2024. Allava: Harnessing gpt4v-synthesized data for a lite vision-language model. arXiv:2402.11684. Retrieved from https:\/\/arxiv.org\/abs\/2402.11684"},{"issue":"2","key":"e_1_3_2_32_2","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1145\/3166054.3166058","article-title":"A survey on dialogue systems: Recent advances and new frontiers","volume":"19","author":"Chen Hongshen","year":"2017","unstructured":"Hongshen Chen, Xiaorui Liu, Dawei Yin, and Jiliang Tang. 2017. A survey on dialogue systems: Recent advances and new frontiers. ACM SIGKDD Explorations Newsletter 19, 2 (2017), 25\u201335.","journal-title":"ACM SIGKDD Explorations Newsletter"},{"key":"e_1_3_2_33_2","unstructured":"Jun Chen Deyao Zhu Xiaoqian Shen Xiang Li Zechun Liu Pengchuan Zhang Raghuraman Krishnamoorthi Vikas Chandra Yunyang Xiong and Mohamed Elhoseiny. 2023. Minigpt-v2: Large language model as a unified interface for vision-language multi-task learning. arXiv:2310.09478. Retrieved from https:\/\/arxiv.org\/abs\/2310.09478"},{"key":"e_1_3_2_34_2","unstructured":"Wei-Ge Chen Irina Spiridonova Jianwei Yang Jianfeng Gao and Chunyuan Li. 2023. Llava-interactive: An all-in-one demo for image chat segmentation generation and editing. arXiv:2311.00571. Retrieved from https:\/\/arxiv.org\/abs\/2311.00571"},{"key":"e_1_3_2_35_2","unstructured":"Xi Chen Josip Djolonga Piotr Padlewski Basil Mustafa Soravit Changpinyo Jialin Wu Carlos Riquelme Ruiz Sebastian Goodman Xiao Wang Yi Tay et al. 2023. Pali-x: On scaling up a multilingual vision and language model. arXiv:2305.18565. Retrieved from https:\/\/arxiv.org\/abs\/2305.18565"},{"issue":"2","key":"e_1_3_2_36_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3606368","article-title":"Multimodal dialog systems with dual knowledge-enhanced generative pretrained language model","volume":"42","author":"Chen Xiaolin","year":"2023","unstructured":"Xiaolin Chen, Xuemeng Song, Liqiang Jing, Shuo Li, Linmei Hu, and Liqiang Nie. 2023. Multimodal dialog systems with dual knowledge-enhanced generative pretrained language model. ACM Transactions on Information Systems 42, 2 (2023), 1\u201325.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_2_37_2","first-page":"1","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Chen Xi","year":"2023","unstructured":"Xi Chen, Xiao Wang, Soravit Changpinyo, A. J. Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. 2023. PaLI: A jointly-scaled multilingual language-image model. In Proceedings of the 11th International Conference on Learning Representations, 1\u201333."},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","unstructured":"Zhe Chen Weiyun Wang Hao Tian Shenglong Ye Zhangwei Gao Erfei Cui Wenwen Tong Kongzhi Hu Jiapeng Luo Zheng Ma et al. 2024. How far are we to gpt-4v? Closing the gap to commercial multimodal models with open-source suites. arXiv:2404.16821. Retrieved from https:\/\/arxiv.org\/abs\/2404.16821","DOI":"10.1007\/s11432-024-4231-5"},{"key":"e_1_3_2_39_2","first-page":"24185","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Zhe","year":"2024","unstructured":"Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2024. InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 24185\u201324198."},{"key":"e_1_3_2_40_2","first-page":"33","volume-title":"Proceedings of the International Conference on Advanced Data Mining and Applications","author":"Cheng Fenghua","year":"2024","unstructured":"Fenghua Cheng, Xue Li, Haoyang Wu, Jiangcheng Sang, and Wenqi Zhao. 2024. Recent advances on multi-modal dialogue systems: A survey. In Proceedings of the International Conference on Advanced Data Mining and Applications. Springer, 33\u201347."},{"key":"e_1_3_2_41_2","doi-asserted-by":"crossref","first-page":"13877","DOI":"10.18653\/v1\/2023.emnlp-main.856","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Cheng Siyuan","year":"2023","unstructured":"Siyuan Cheng, Bozhong Tian, Qingbin Liu, Xi Chen, Yongheng Wang, Huajun Chen, and Ningyu Zhang. 2023. Can we edit multimodal large language models? In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13877\u201313888."},{"key":"e_1_3_2_42_2","unstructured":"Zesen Cheng Sicong Leng Hang Zhang Yifei Xin Xin Li Guanzheng Chen Yongxin Zhu Wenqi Zhang Ziyang Luo Deli Zhao et al. 2024. VideoLLaMA 2: Advancing spatial-temporal modeling and audio understanding in video-LLMs. arXiv:2406.07476. Retrieved from https:\/\/arxiv.org\/abs\/2406.07476"},{"key":"e_1_3_2_43_2","first-page":"2818","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Cherti Mehdi","year":"2023","unstructured":"Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev. 2023. Reproducible scaling laws for contrastive language-image learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2818\u20132829."},{"key":"e_1_3_2_44_2","unstructured":"Wei-Lin Chiang Zhuohan Li Zi Lin Ying Sheng Zhanghao Wu Hao Zhang Lianmin Zheng Siyuan Zhuang Yonghao Zhuang Joseph E. Gonzalez et al. 2023. Vicuna: An Open-Source Chatbot Impressing Gpt-4 with 90%* Chatgpt Quality. Retrieved April 14 2023 from https:\/\/vicuna.lmsys.org"},{"issue":"240","key":"e_1_3_2_45_2","first-page":"1","article-title":"Palm: Scaling language modeling with pathways","volume":"24","author":"Chowdhery Aakanksha","year":"2023","unstructured":"Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research 24, 240 (2023), 1\u2013113.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_46_2","unstructured":"Xiangxiang Chu Limeng Qiao Xinyang Lin Shuang Xu Yang Yiming Hu Fei Wei Xinyu Zhang Bo Zhang Xiaolin Wei et al. 2023. MobileVLM: A fast reproducible and strong vision language assistant for mobile devices. arXiv:2312.16886. Retrieved from https:\/\/arxiv.org\/abs\/2312.16886"},{"key":"e_1_3_2_47_2","unstructured":"Xiangxiang Chu Limeng Qiao Xinyu Zhang Shuang Xu Fei Wei Yang Yang Xiaofei Sun Yiming Hu Xinyang Lin Bo Zhang et al. 2024. MobileVLM v2: Faster and stronger baseline for vision language model. arXiv:2402.03766. Retrieved from https:\/\/arxiv.org\/abs\/2402.03766"},{"key":"e_1_3_2_48_2","doi-asserted-by":"crossref","first-page":"2456","DOI":"10.1109\/TASLP.2021.3065852","article-title":"End-to-end recurrent cross-modality attention for video dialogue","volume":"29","author":"Chu Yun-Wei","year":"2021","unstructured":"Yun-Wei Chu, Kuan-Yen Lin, Chao-Chun Hsu, and Lun-Wei Ku. 2021. End-to-end recurrent cross-modality attention for video dialogue. IEEE\/ACM Transactions on Audio, Speech, and Language Processing 29 (2021), 2456\u20132464.","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"issue":"70","key":"e_1_3_2_49_2","first-page":"1","article-title":"Scaling instruction-finetuned language models","volume":"25","author":"Chung Hyung Won","year":"2024","unstructured":"Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2024. Scaling instruction-finetuned language models. Journal of Machine Learning Research 25, 70 (2024), 1\u201353.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_50_2","first-page":"16372","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Cong Yuren","year":"2021","unstructured":"Yuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn, and Michael Ying Yang. 2021. Spatial-temporal transformer for dynamic scene graph generation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 16372\u201316382."},{"key":"e_1_3_2_51_2","first-page":"49250","volume-title":"Proceedings of the 37th Conference on Neural Information Processing Systems","author":"Dai Wenliang","year":"2023","unstructured":"Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. In Proceedings of the 37th Conference on Neural Information Processing Systems, 49250\u201349267."},{"key":"e_1_3_2_52_2","doi-asserted-by":"crossref","first-page":"90","DOI":"10.1016\/j.cviu.2017.10.001","article-title":"Human attention in visual question answering: Do humans and deep networks look at the same regions?","volume":"163","author":"Das Abhishek","year":"2017","unstructured":"Abhishek Das, Harsh Agrawal, Larry Zitnick, Devi Parikh, and Dhruv Batra. 2017. Human attention in visual question answering: Do humans and deep networks look at the same regions? Computer Vision and Image Understanding 163 (2017), 90\u2013100.","journal-title":"Computer Vision and Image Understanding"},{"key":"e_1_3_2_53_2","first-page":"326","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Das Abhishek","year":"2017","unstructured":"Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, Jos\u00e9 M. F. Moura, Devi Parikh, and Dhruv Batra. 2017. Visual dialog. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 326\u2013335."},{"key":"e_1_3_2_54_2","first-page":"2951","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Das Abhishek","year":"2017","unstructured":"Abhishek Das, Satwik Kottur, Jos\u00e9 M. F. Moura, Stefan Lee, and Dhruv Batra. 2017. Learning cooperative visual dialog agents with deep reinforcement learning. In Proceedings of the IEEE International Conference on Computer Vision, 2951\u20132960."},{"issue":"8","key":"e_1_3_2_55_2","doi-asserted-by":"crossref","first-page":"6037","DOI":"10.1007\/s10462-022-10148-x","article-title":"Attention, please! A survey of neural attention models in deep learning","volume":"55","author":"Correia Alana de Santana","year":"2022","unstructured":"Alana de Santana Correia and Esther Luna Colombini. 2022. Attention, please! A survey of neural attention models in deep learning. Artificial Intelligence Review 55, 8 (2022), 6037\u20136124.","journal-title":"Artificial Intelligence Review"},{"key":"e_1_3_2_56_2","first-page":"5503","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"De Vries Harm","year":"2017","unstructured":"Harm De Vries, Florian Strub, Sarath Chandar, Olivier Pietquin, Hugo Larochelle, and Aaron Courville. 2017. Guesswhat?! Visual object discovery through multi-modal dialogue. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5503\u20135512."},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","first-page":"6448","DOI":"10.1145\/3637528.3671474","volume-title":"Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Deldjoo Yashar","year":"2024","unstructured":"Yashar Deldjoo, Zhankui He, Julian McAuley, Anton Korikov, Scott Sanner, Arnau Ramisa, Ren\u00e9 Vidal, Maheswaran Sathiamoorthy, Atoosa Kasirzadeh, and Silvia Milano. 2024. A review of modern recommender systems using generative models (Gen-RecSys). In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 6448\u20136458."},{"key":"e_1_3_2_58_2","doi-asserted-by":"crossref","unstructured":"Yashar Deldjoo Zhankui He Julian McAuley Anton Korikov Scott Sanner Arnau Ramisa Rene Vidal Maheswaran Sathiamoorthy Atoosa Kasrizadeh Silvia Milano et al. 2024. Recommendation with generative models. arXiv:2409.15173. Retrieved from https:\/\/arxiv.org\/abs\/2409.15173","DOI":"10.1145\/3701551.3703485"},{"key":"e_1_3_2_59_2","first-page":"485","volume-title":"Proceedings of the IEEE","volume":"108","author":"Deng Lei","year":"2020","unstructured":"Lei Deng, Guoqi Li, Song Han, Luping Shi, and Yuan Xie. 2020. Model compression and hardware acceleration for neural networks: A comprehensive survey. Proceedings of the IEEE 108, 4 (2020), 485\u2013532."},{"key":"e_1_3_2_60_2","doi-asserted-by":"crossref","first-page":"755","DOI":"10.1007\/s10462-020-09866-x","article-title":"Survey on evaluation methods for dialogue systems","volume":"54","author":"Deriu Jan","year":"2021","unstructured":"Jan Deriu, Alvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak. 2021. Survey on evaluation methods for dialogue systems. Artificial Intelligence Review 54 (2021), 755\u2013810.","journal-title":"Artificial Intelligence Review"},{"key":"e_1_3_2_61_2","first-page":"10088","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Dettmers Tim","year":"2024","unstructured":"Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2024. Qlora: Efficient finetuning of quantized LLMs. In Proceedings of the Advances in Neural Information Processing Systems, 10088\u201310115."},{"key":"e_1_3_2_62_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4171\u20134186."},{"key":"e_1_3_2_63_2","first-page":"3563","article-title":"Chain-of-verification reduces hallucination in large language models","author":"Dhuliawala Shehzaad","year":"2024","unstructured":"Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason E. Weston. 2024. Chain-of-verification reduces hallucination in large language models. In Findings of the Association for Computational Linguistics: ACL 2024, 3563\u20133578.","journal-title":"Findings of the Association for Computational Linguistics: ACL 2024"},{"key":"e_1_3_2_64_2","doi-asserted-by":"crossref","first-page":"138","DOI":"10.3115\/1289189.1289273","volume-title":"Proceedings of the 2nd International Conference on Human Language Technology Research","author":"Doddington George","year":"2002","unstructured":"George Doddington. 2002. Automatic evaluation of machine translation quality using n-gram co-occurrence statistics. In Proceedings of the 2nd International Conference on Human Language Technology Research, 138\u2013145."},{"key":"e_1_3_2_65_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2020","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations, 1\u201321."},{"key":"e_1_3_2_66_2","first-page":"320","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics","author":"Du Zhengxiao","year":"2022","unstructured":"Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. GLM: General language model pretraining with autoregressive blank infilling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 320\u2013335."},{"issue":"2","key":"e_1_3_2_67_2","doi-asserted-by":"crossref","first-page":"230","DOI":"10.1109\/TETCI.2022.3141105","article-title":"A survey of embodied AI: From simulators to research tasks","volume":"6","author":"Duan Jiafei","year":"2022","unstructured":"Jiafei Duan, Samson Yu, Hui Li Tan, Hongyuan Zhu, and Cheston Tan. 2022. A survey of embodied AI: From simulators to research tasks. IEEE Transactions on Emerging Topics in Computational Intelligence 6, 2 (2022), 230\u2013244.","journal-title":"IEEE Transactions on Emerging Topics in Computational Intelligence"},{"issue":"3","key":"e_1_3_2_68_2","first-page":"211","article-title":"The algorithmic foundations of differential privacy","volume":"9","author":"Dwork Cynthia","year":"2014","unstructured":"Cynthia Dwork and Aaron Roth. 2014. The algorithmic foundations of differential privacy. Foundations and Trends\u00ae in Theoretical Computer Science 9, 3\u20134 (2014), 211\u2013407.","journal-title":"Foundations and Trends\u00ae in Theoretical Computer Science"},{"key":"e_1_3_2_69_2","first-page":"16091","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Fang Gongfan","year":"2023","unstructured":"Gongfan Fang, Xinyin Ma, Mingli Song, Michael Bi Mi, and Xinchao Wang. 2023. Depgraph: Towards any structural pruning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 16091\u201316101."},{"key":"e_1_3_2_70_2","first-page":"19358","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Fang Yuxin","year":"2023","unstructured":"Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. 2023. Eva: Exploring the limits of masked visual representation learning at scale. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 19358\u201319369."},{"key":"e_1_3_2_71_2","unstructured":"Jiajun Fei Dian Li Zhidong Deng Zekun Wang Gang Liu and Hui Wang. 2024. Video-CCAM: Enhancing video-language understanding with causal cross-attention masks for short and long videos. arXiv:2408.14023. Retrieved from https:\/\/arxiv.org\/abs\/2408.14023"},{"key":"e_1_3_2_72_2","first-page":"203","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Feichtenhofer Christoph","year":"2020","unstructured":"Christoph Feichtenhofer. 2020. X3d: Expanding architectures for efficient video recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 203\u2013213."},{"key":"e_1_3_2_73_2","first-page":"70757","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Feng Guhao","year":"2024","unstructured":"Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang. 2024. Towards revealing the mystery behind chain of thought: A theoretical perspective. In Proceedings of the Advances in Neural Information Processing Systems, 70757\u201370798."},{"key":"e_1_3_2_74_2","first-page":"7348","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Feng Jiazhan","year":"2023","unstructured":"Jiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao, Yaming Yang, Chongyang Tao, Dongyan Zhao, and Qingwei Lin. 2023. MMDialog: A large-scale multi-turn dialogue dataset towards multi-modal open-domain conversation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 7348\u20137363."},{"key":"e_1_3_2_75_2","first-page":"168","volume-title":"In Proceedings of the NIPS, Modern Machine Learning and Natural Language Processing Workshop","volume":"2","author":"Forgues Gabriel","year":"2014","unstructured":"Gabriel Forgues, Joelle Pineau, Jean-Marie Larchev\u00eaque, and R\u00e9al Tremblay. 2014. Bootstrapping dialog systems with word embeddings. In Proceedings of the NIPS, Modern Machine Learning and Natural Language Processing Workshop, Vol. 2, 168."},{"key":"e_1_3_2_76_2","unstructured":"Chaoyou Fu Peixian Chen Yunhang Shen Yulei Qin Mengdan Zhang Xu Lin Jinrui Yang Xiawu Zheng Ke Li Xing Sun et al. 2023. MME: A comprehensive evaluation benchmark for multimodal large language models. arXiv:2306.13394. Retrieved from https:\/\/arxiv.org\/abs\/2306.13394"},{"key":"e_1_3_2_77_2","unstructured":"Chaoyou Fu Yuhan Dai Yondong Luo Lei Li Shuhuai Ren Renrui Zhang Zihan Wang Chenyu Zhou Yunhang Shen Mengdan Zhang et al. 2024. Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis. arXiv:2405.21075. Retrieved from https:\/\/arxiv.org\/abs\/2405.21075"},{"key":"e_1_3_2_78_2","unstructured":"Chaoyou Fu Haojia Lin Zuwei Long Yunhang Shen Meng Zhao Yifan Zhang Xiong Wang Di Yin Long Ma Xiawu Zheng et al. 2024. Vita: Towards open-source interactive omni multimodal LLM. arXiv:2408.05211. Retrieved from https:\/\/arxiv.org\/abs\/2408.05211"},{"key":"e_1_3_2_79_2","unstructured":"Chaoyou Fu Yi-Fan Zhang Shukang Yin Bo Li Xinyu Fang Sirui Zhao Haodong Duan Xing Sun Ziwei Liu Liang Wang et al. 2024. MME-survey: A comprehensive survey on evaluation of multimodal LLMs. arXiv:2411.15296. Retrieved from https:\/\/arxiv.org\/abs\/2411.15296"},{"key":"e_1_3_2_80_2","first-page":"1579","volume-title":"Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Fu Tsu-Jui","year":"2019","unstructured":"Tsu-Jui Fu, Shao-Heng Tai, and Hwann-Tzong Chen. 2019. Attentive and adversarial learning for video summarization. In Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 1579\u20131587."},{"key":"e_1_3_2_81_2","first-page":"12799","volume-title":"In Proceedings of the AAAI Conference on Artificial Intelligence","volume":"37","author":"Fu Zihao","year":"2023","unstructured":"Zihao Fu, Haoran Yang, Anthony Man-Cho So, Wai Lam, Lidong Bing, and Nigel Collier. 2023. On the effectiveness of parameter-efficient fine-tuning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, 12799\u201312807."},{"key":"e_1_3_2_82_2","doi-asserted-by":"crossref","first-page":"6463","DOI":"10.18653\/v1\/P19-1648","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Gan Zhe","year":"2019","unstructured":"Zhe Gan, Yu Cheng, Ahmed Kholy, Linjie Li, Jingjing Liu, and Jianfeng Gao. 2019. Multi-step reasoning via recurrent dual attention for visual dialog. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 6463\u20136474."},{"issue":"3","key":"e_1_3_2_83_2","doi-asserted-by":"crossref","first-page":"163","DOI":"10.1561\/0600000105","article-title":"Vision-language pre-training: Basics, recent advances, and future trends","volume":"14","author":"Gan Zhe","year":"2022","unstructured":"Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang, Zicheng Liu, and Jianfeng Gao. 2022. Vision-language pre-training: Basics, recent advances, and future trends. Foundations and Trends\u00ae in Computer Graphics and Vision 14, 3\u20134 (2022), 163\u2013352.","journal-title":"Foundations and Trends\u00ae in Computer Graphics and Vision"},{"key":"e_1_3_2_84_2","unstructured":"Peng Gao Jiaming Han Renrui Zhang Ziyi Lin Shijie Geng Aojun Zhou Wei Zhang Pan Lu Conghui He Xiangyu Yue et al. 2023. Llama-adapter v2: Parameter-efficient visual instruction model. arXiv:2304.15010. Retrieved from https:\/\/arxiv.org\/abs\/2304.15010"},{"key":"e_1_3_2_85_2","first-page":"32400","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Gao Peng","year":"2024","unstructured":"Peng Gao, Renrui Zhang, Chris Liu, Longtian Qiu, Siyuan Huang, Weifeng Lin, Shitian Zhao, Shijie Geng, Ziyi Lin, Peng Jin, et al. 2024. SPHINX-X: Scaling data and parameters for a family of multi-modal large language models. In Proceedings of the 41st International Conference on Machine Learning, 32400\u201332420."},{"key":"e_1_3_2_86_2","first-page":"1","volume-title":"Are we Bayesian referring expression generatorsProceedings of the CogSci Workshop on the Production of Referring Expressions","author":"Gatt Albert","year":"2013","unstructured":"Albert Gatt, Roger P. G. van Gompel, Kees van Deemter, and Emiel Krahmer. 2013. Are we Bayesian referring expression generators. In Proceedings of the CogSci Workshop on the Production of Referring Expressions, 1\u20136."},{"key":"e_1_3_2_87_2","first-page":"1415","volume-title":"In Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Geng Shijie","year":"2021","unstructured":"Shijie Geng, Peng Gao, Moitreya Chatterjee, Chiori Hori, Jonathan Le Roux, Yongfeng Zhang, Hongsheng Li, and Anoop Cherian. 2021. Dynamic graph representation learning for video dialog via multi-modal shuffled transformers. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35, 1415\u20131423."},{"key":"e_1_3_2_88_2","doi-asserted-by":"crossref","first-page":"291","DOI":"10.1201\/9781003162810-13","volume-title":"Low-Power Computer Vision","author":"Gholami Amir","year":"2022","unstructured":"Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. 2022. A survey of quantization methods for efficient neural network inference. In Low-Power Computer Vision. Chapman and Hall\/CRC, 291\u2013326."},{"key":"e_1_3_2_89_2","first-page":"15180","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Girdhar Rohit","year":"2023","unstructured":"Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. Imagebind: One embedding space to bind them all. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 15180\u201315190."},{"key":"e_1_3_2_90_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Girdhar Rohit","year":"2019","unstructured":"Rohit Girdhar and Deva Ramanan. 2019. CATER: A diagnostic dataset for compositional actions & temporal reasoning. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_91_2","unstructured":"Team G. L. M. Aohan Zeng Bin Xu Bowen Wang Chenhui Zhang Da Yin Diego Rojas Guanyu Feng Hanlin Zhao Hanyu Lai et al. 2024. ChatGLM: A family of large language models from GLM-130B to GLM-4 all tools. arXiv:2406.12793. Retrieved from https:\/\/arxiv.org\/abs\/2406.12793"},{"issue":"6","key":"e_1_3_2_92_2","doi-asserted-by":"crossref","first-page":"1789","DOI":"10.1007\/s11263-021-01453-z","article-title":"Knowledge distillation: A survey","volume":"129","author":"Gou Jianping","year":"2021","unstructured":"Jianping Gou, Baosheng Yu, Stephen J. Maybank, and Dacheng Tao. 2021. Knowledge distillation: A survey. International Journal of Computer Vision 129, 6 (2021), 1789\u20131819.","journal-title":"International Journal of Computer Vision"},{"key":"e_1_3_2_93_2","first-page":"43","article-title":"Logic and conversation","volume":"3","author":"Grice Herbert P.","year":"1975","unstructured":"Herbert P. Grice. 1975. Logic and conversation. Syntax and Semantics 3 (1975), 43\u201358.","journal-title":"Syntax and Semantics"},{"key":"e_1_3_2_94_2","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1016\/j.patcog.2017.10.013","article-title":"Recent advances in convolutional neural networks","volume":"77","author":"Gu Jiuxiang","year":"2018","unstructured":"Jiuxiang Gu, Zhenhua Wang, Jason Kuen, Lianyang Ma, Amir Shahroudy, Bing Shuai, Ting Liu, Xingxing Wang, Gang Wang, Jianfei Cai, et al. 2018. Recent advances in convolutional neural networks. Pattern Recognition 77 (2018), 354\u2013377.","journal-title":"Pattern Recognition"},{"key":"e_1_3_2_95_2","first-page":"1969","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Gu Jiuxiang","year":"2019","unstructured":"Jiuxiang Gu, Handong Zhao, Zhe Lin, Sheng Li, Jianfei Cai, and Mingyang Ling. 2019. Scene graph generation with external knowledge and image reconstruction. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1969\u20131978."},{"key":"e_1_3_2_96_2","first-page":"1","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Gu Yuxian","year":"2024","unstructured":"Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2024. MiniLLM: Knowledge distillation of large language models. In Proceedings of the 12th International Conference on Learning Representations, 1\u201324."},{"issue":"2","key":"e_1_3_2_97_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3439816","article-title":"Conditional text generation for harmonious human-machine interaction","volume":"12","author":"Guo Bin","year":"2021","unstructured":"Bin Guo, Hao Wang, Yasan Ding, Wei Wu, Shaoyang Hao, Yueqi Sun, and Zhiwen Yu. 2021. Conditional text generation for harmonious human-machine interaction. ACM Transactions on Intelligent Systems and Technology 12, 2 (2021), 1\u201350.","journal-title":"ACM Transactions on Intelligent Systems and Technology"},{"key":"e_1_3_2_98_2","first-page":"4989","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI)","author":"Guo Dan","year":"2019","unstructured":"Dan Guo, Hui Wang, and Meng Wang. 2019. Dual visual attention network for visual dialog. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), 4989\u20134995."},{"issue":"10","key":"e_1_3_2_99_2","first-page":"6056","article-title":"Context-aware graph inference with knowledge distillation for visual dialog","volume":"44","author":"Guo Dan","year":"2021","unstructured":"Dan Guo, Hui Wang, and Meng Wang. 2021. Context-aware graph inference with knowledge distillation for visual dialog. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 10 (2021), 6056\u20136073.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_100_2","doi-asserted-by":"crossref","first-page":"6655","DOI":"10.1109\/TIP.2020.2992888","article-title":"Textual-visual reference-aware attention network for visual dialog","volume":"29","author":"Guo Dan","year":"2020","unstructured":"Dan Guo, Hui Wang, Shuhui Wang, and Meng Wang. 2020. Textual-visual reference-aware attention network for visual dialog. IEEE Transactions on Image Processing 29 (2020), 6655\u20136666.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_2_101_2","first-page":"10055","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Guo Dan","year":"2020","unstructured":"Dan Guo, Hui Wang, Hanwang Zhang, Zheng-Jun Zha, and Meng Wang. 2020. Iterative context-aware graph inference for visual dialog. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 10055\u201310064."},{"key":"e_1_3_2_102_2","first-page":"10434","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Guo Dalu","year":"2019","unstructured":"Dalu Guo, Chang Xu, and Dacheng Tao. 2019. Image-question-answer synergistic network for visual dialog. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 10434\u201310443."},{"key":"e_1_3_2_103_2","first-page":"1508","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Guo Jinyang","year":"2020","unstructured":"Jinyang Guo, Wanli Ouyang, and Dong Xu. 2020. Multi-dimensional pruning: A unified framework for model compression. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1508\u20131517."},{"issue":"1","key":"e_1_3_2_104_2","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1109\/TPAMI.2022.3152247","article-title":"A survey on vision transformer","volume":"45","author":"Han Kai","year":"2022","unstructured":"Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al. 2022. A survey on vision transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 1 (2022), 87\u2013110.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_105_2","unstructured":"Zeyu Han Chao Gao Jinyang Liu and Sai Qian Zhang. 2024. Parameter-efficient fine-tuning for large models: A comprehensive survey. arXiv:2403.14608. Retrieved from https:\/\/arxiv.org\/abs\/2403.14608"},{"key":"e_1_3_2_106_2","first-page":"6546","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Hara Kensho","year":"2018","unstructured":"Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. 2018. Can spatiotemporal 3D CNNs retrace the history of 2D CNNs and ImageNet? In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6546\u20136555."},{"key":"e_1_3_2_107_2","first-page":"770","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770\u2013778."},{"key":"e_1_3_2_108_2","first-page":"784","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"He Yihui","year":"2018","unstructured":"Yihui He, Ji Lin, Zhijian Liu, Hanrui Wang, Li-Jia Li, and Song Han. 2018. AMC: AutoML for model compression and acceleration on mobile devices. In Proceedings of the European Conference on Computer Vision (ECCV), 784\u2013800."},{"issue":"8","key":"e_1_3_2_109_2","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter Sepp","year":"1997","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long short-term memory. Neural Computation 9, 8 (1997), 1735\u20131780.","journal-title":"Neural Computation"},{"key":"e_1_3_2_110_2","first-page":"30016","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems","author":"Hoffmann Jordan","year":"2022","unstructured":"Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022. Training compute-optimal large language models. In Proceedings of the 36th International Conference on Neural Information Processing Systems, 30016\u201330030."},{"key":"e_1_3_2_111_2","unstructured":"Wenyi Hong Weihan Wang Ming Ding Wenmeng Yu Qingsong Lv Yan Wang Yean Cheng Shiyu Huang Junhui Ji Zhao Xue et al. 2024. Cogvlm2: Visual language models for image and video understanding. arXiv:2408.16500. Retrieved from https:\/\/arxiv.org\/abs\/2408.16500"},{"key":"e_1_3_2_112_2","first-page":"2352","volume-title":"Proceedings of the 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201919)","author":"Hori Chiori","year":"2019","unstructured":"Chiori Hori, Huda Alamri, Jue Wang, Gordon Wichern, Takaaki Hori, Anoop Cherian, Tim K. Marks, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das, et al. 2019. End-to-end audio visual scene-aware dialog using multimodal attention-based video features. In Proceedings of the 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201919). IEEE, 2352\u20132356."},{"key":"e_1_3_2_113_2","unstructured":"Andrew G. Howard Menglong Zhu Bo Chen Dmitry Kalenichenko Weijun Wang Tobias Weyand Marco Andreetto and Hartwig Adam. 2017. MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv:1704.04861. Retrieved from https:\/\/arxiv.org\/abs\/1704.04861"},{"key":"e_1_3_2_114_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Hu Edward J.","year":"2022","unstructured":"Edward J. Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations, 1\u201313."},{"key":"e_1_3_2_115_2","first-page":"10294","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Hu Ronghang","year":"2019","unstructured":"Ronghang Hu, Anna Rohrbach, Trevor Darrell, and Kate Saenko. 2019. Language-conditioned graph networks for relational reasoning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 10294\u201310303."},{"key":"e_1_3_2_116_2","first-page":"2256","volume-title":"In Proceedings of the AAAI Conference on Artificial Intelligence","volume":"38","author":"Hu Wenbo","year":"2024","unstructured":"Wenbo Hu, Yifan Xu, Yi Li, Weiyue Li, Zeyuan Chen, and Zhuowen Tu. 2024. Bliva: A simple multimodal LLM for better handling of text-rich visual questions. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 2256\u20132264."},{"key":"e_1_3_2_117_2","first-page":"23369","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Hu Ziniu","year":"2023","unstructured":"Ziniu Hu, Ahmet Iscen, Chen Sun, Zirui Wang, Kai-Wei Chang, Yizhou Sun, Cordelia Schmid, David A. Ross, and Alireza Fathi. 2023. Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge memory. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 23369\u201323379."},{"key":"e_1_3_2_118_2","first-page":"5254","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Hu Zhiqiang","year":"2023","unstructured":"Zhiqiang Hu, Lei Wang, Yihuai Lan, Wanyu Xu, Ee-Peng Lim, Lidong Bing, Xing Xu, Soujanya Poria, and Roy Lee. 2023. LLM-Adapters: An adapter family for parameter-efficient fine-tuning of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 5254\u20135276."},{"key":"e_1_3_2_119_2","first-page":"14271","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Huang Bin","year":"2024","unstructured":"Bin Huang, Xin Wang, Hong Chen, Zihan Song, and Wenwu Zhu. 2024. VtimeLLM: Empower LLM to grasp video moments. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 14271\u201314280."},{"issue":"2","key":"e_1_3_2_120_2","doi-asserted-by":"crossref","first-page":"577","DOI":"10.1109\/TCSVT.2019.2890899","article-title":"A novel key-frames selection framework for comprehensive video summarization","volume":"30","author":"Huang Cheng","year":"2019","unstructured":"Cheng Huang and Hongmei Wang. 2019. A novel key-frames selection framework for comprehensive video summarization. IEEE Transactions on Circuits and Systems for Video Technology 30, 2 (2019), 577\u2013589.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_2_121_2","first-page":"26135","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Huang Hanzhuo","year":"2024","unstructured":"Hanzhuo Huang, Yufan Feng, Cheng Shi, Lan Xu, Jingyi Yu, and Sibei Yang. 2024. Free-bloom: Zero-shot text-to-video generator with LLM director and LDM animator. In Proceedings of the Advances in Neural Information Processing Systems, 26135\u201326158."},{"issue":"3","key":"e_1_3_2_122_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3383123","article-title":"Challenges in building intelligent open-domain dialog systems","volume":"38","author":"Huang Minlie","year":"2020","unstructured":"Minlie Huang, Xiaoyan Zhu, and Jianfeng Gao. 2020. Challenges in building intelligent open-domain dialog systems. ACM Transactions on Information Systems 38, 3 (2020), 1\u201332.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_2_123_2","first-page":"72096","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Huang Shaohan","year":"2023","unstructured":"Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, et al. 2023. Language is not all you need: Aligning perception with language models. In Proceedings of the Advances in Neural Information Processing Systems, 72096\u201372109."},{"key":"e_1_3_2_124_2","first-page":"33716","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Huang Tao","year":"2022","unstructured":"Tao Huang, Shan You, Fei Wang, Chen Qian, and Chang Xu. 2022. Knowledge distillation from a stronger teacher. In Proceedings of the Advances in Neural Information Processing Systems, 33716\u201333727."},{"key":"e_1_3_2_125_2","volume-title":"Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, 1","author":"Huber Bernd","year":"2018","unstructured":"Bernd Huber, Daniel McDuff, Chris Brockett, Michel Galley, and Bill Dolan. 2018. Emotional dialogue generation using image-grounded language models. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, 1\u201312."},{"key":"e_1_3_2_126_2","doi-asserted-by":"crossref","first-page":"107567","DOI":"10.1016\/j.patcog.2020.107567","article-title":"A comprehensive survey of multi-view video summarization","volume":"109","author":"Hussain Tanveer","year":"2021","unstructured":"Tanveer Hussain, Khan Muhammad, Weiping Ding, Jaime Lloret, Sung Wook Baik, and Victor Hugo C. de Albuquerque. 2021. A comprehensive survey of multi-view video summarization. Pattern Recognition 109 (2021), 107567.","journal-title":"Pattern Recognition"},{"key":"e_1_3_2_127_2","unstructured":"Forrest N. Iandola Song Han Matthew W. Moskewicz Khalid Ashraf William J. Dally and Kurt Keutzer. 2016. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5 MB model size. arXiv:1602.07360. Retrieved from https:\/\/arxiv.org\/abs\/1602.07360"},{"key":"e_1_3_2_128_2","first-page":"37","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Imani Shima","year":"2023","unstructured":"Shima Imani, Liang Du, and Harsh Shrivastava. 2023. MathPrompter: Mathematical reasoning using large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 37\u201342."},{"key":"e_1_3_2_129_2","first-page":"4651","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Jaegle Andrew","year":"2021","unstructured":"Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. 2021. Perceiver: General perception with iterative attention. In Proceedings of the International Conference on Machine Learning. PMLR, 4651\u20134664."},{"issue":"1","key":"e_1_3_2_130_2","doi-asserted-by":"crossref","first-page":"116","DOI":"10.1016\/j.cviu.2006.10.019","article-title":"Multimodal human\u2013computer interaction: A survey","volume":"108","author":"Jaimes Alejandro","year":"2007","unstructured":"Alejandro Jaimes and Nicu Sebe. 2007. Multimodal human\u2013computer interaction: A survey. Computer Vision and Image Understanding 108, 1\u20132 (2007), 116\u2013134.","journal-title":"Computer Vision and Image Understanding"},{"key":"e_1_3_2_131_2","first-page":"27992","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Jain Jitesh","year":"2024","unstructured":"Jitesh Jain, Jianwei Yang, and Humphrey Shi. 2024. Vcoder: Versatile vision encoders for multimodal large language models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 27992\u201328002."},{"key":"e_1_3_2_132_2","first-page":"5754","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Jain Unnat","year":"2018","unstructured":"Unnat Jain, Svetlana Lazebnik, and Alexander G. Schwing. 2018. Two can play this game: Visual dialog with discriminative question generation and answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5754\u20135763."},{"issue":"13","key":"e_1_3_2_133_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3584700","article-title":"A survey on multi-modal summarization","volume":"55","author":"Jangra Anubhav","year":"2023","unstructured":"Anubhav Jangra, Sourajit Mukherjee, Adam Jatowt, Sriparna Saha, and Mohammad Hasanuzzaman. 2023. A survey on multi-modal summarization. ACM Computing Surveys 55, 13s (2023), 1\u201336.","journal-title":"ACM Computing Surveys"},{"key":"e_1_3_2_134_2","first-page":"1655","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Ji Jiayi","year":"2021","unstructured":"Jiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen, Gen Luo, Yongjian Wu, Yue Gao, and Rongrong Ji. 2021. Improving image captioning by leveraging intra-and inter-layer global representation in transformer network. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35, 1655\u20131663."},{"issue":"1","key":"e_1_3_2_135_2","first-page":"221","article-title":"3D convolutional neural networks for human action recognition","volume":"35","author":"Ji Shuiwang","year":"2012","unstructured":"Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu. 2012. 3D convolutional neural networks for human action recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 35, 1 (2012), 221\u2013231.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"12","key":"e_1_3_2_136_2","first-page":"1","article-title":"Survey of hallucination in natural language generation","volume":"55","author":"Ji Ziwei","year":"2023","unstructured":"Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys 55, 12 (2023), 1\u201338.","journal-title":"ACM Computing Surveys"},{"issue":"6","key":"e_1_3_2_137_2","first-page":"1709","article-title":"Video summarization with attention-based encoder\u2013decoder networks","volume":"30","author":"Ji Zhong","year":"2019","unstructured":"Zhong Ji, Kailin Xiong, Yanwei Pang, and Xuelong Li. 2019. Video summarization with attention-based encoder\u2013decoder networks. IEEE Transactions on Circuits and Systems for Video Technology 30, 6 (2019), 1709\u20131717.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_2_138_2","first-page":"57","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Jian Yiren","year":"2024","unstructured":"Yiren Jian, Chongyang Gao, and Soroush Vosoughi. 2024. Bootstrapping vision-language learning with decoupled language pre-training. In Proceedings of the Advances in Neural Information Processing Systems, 57\u201372."},{"key":"e_1_3_2_139_2","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Antoine Roux Arthur Mensch Blanche Savary Chris Bamford Devendra Singh Chaplot Diego de las Casas Emma Bou Hanna Florian Bressand et al. 2024. Mixtral of experts. arXiv:2401.04088. Retrieved from https:\/\/arxiv.org\/abs\/2401.04088"},{"key":"e_1_3_2_140_2","unstructured":"Dongzhi Jiang Renrui Zhang Ziyu Guo Yanmin Wu Jiayi Lei Pengshuo Qiu Pan Lu Zehui Chen Guanglu Song Peng Gao et al. 2024. MMSearch: Benchmarking the potential of large models as multi-modal search engines. arXiv:2409.12959. Retrieved from https:\/\/arxiv.org\/abs\/2409.12959"},{"key":"e_1_3_2_141_2","doi-asserted-by":"crossref","first-page":"1265","DOI":"10.1145\/3394171.3413826","volume-title":"Proceedings of the 28th ACM International Conference on Multimedia","author":"Jiang Xiaoze","year":"2020","unstructured":"Xiaoze Jiang, Siyi Du, Zengchang Qin, Yajing Sun, and Jing Yu. 2020. KBGN: Knowledge-bridge graph network for adaptive vision-text reasoning in visual dialogue. In Proceedings of the 28th ACM International Conference on Multimedia, 1265\u20131273."},{"key":"e_1_3_2_142_2","first-page":"11125","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Jiang Xiaoze","year":"2020","unstructured":"Xiaoze Jiang, Jing Yu, Zengchang Qin, Yingying Zhuang, Xingxing Zhang, Yue Hu, and Qi Wu. 2020. DualVD: An adaptive dual encoding model for deep visual understanding in visual dialogue. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34, 11125\u201311132."},{"key":"e_1_3_2_143_2","first-page":"13700","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Jin Peng","year":"2024","unstructured":"Peng Jin, Ryuichi Takanobu, Wancai Zhang, Xiaochun Cao, and Li Yuan. 2024. Chat-UniVi: Unified visual representation empowers large language models with image and video understanding. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 13700\u201313710."},{"key":"e_1_3_2_144_2","first-page":"2901","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Johnson Justin","year":"2017","unstructured":"Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2901\u20132910."},{"key":"e_1_3_2_145_2","first-page":"3668","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Johnson Justin","year":"2015","unstructured":"Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David Shamma, Michael Bernstein, and Li Fei-Fei. 2015. Image retrieval using scene graphs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 3668\u20133678."},{"key":"e_1_3_2_146_2","volume-title":"Thinking, Fast and Slow","author":"Kahneman Daniel","year":"2011","unstructured":"Daniel Kahneman. 2011. Thinking, Fast and Slow. Farrar, Straus and Giroux."},{"key":"e_1_3_2_147_2","first-page":"6746","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Kang Gi-Cheon","year":"2023","unstructured":"Gi-Cheon Kang, Sungdong Kim, Jin-Hwa Kim, Donghyun Kwak, and Byoung-Tak Zhang. 2023. The dialog must go on: Improving visual dialog via generative self-training. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 6746\u20136756."},{"key":"e_1_3_2_148_2","doi-asserted-by":"crossref","first-page":"2024","DOI":"10.18653\/v1\/D19-1209","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Kang Gi-Cheon","year":"2019","unstructured":"Gi-Cheon Kang, Jaeseo Lim, and Byoung-Tak Zhang. 2019. Dual attention networks for visual reference resolution in visual dialog. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2024\u20132033."},{"key":"e_1_3_2_149_2","doi-asserted-by":"crossref","first-page":"327","DOI":"10.18653\/v1\/2021.findings-emnlp.31","volume-title":"Findings of the Association for Computational Linguistics (EMNLP \u201921)","author":"Kang Gi-Cheon","year":"2021","unstructured":"Gi-Cheon Kang, Junseok Park, Hwaran Lee, Byoung-Tak Zhang, and Jin-Hwa Kim. 2021. Reasoning visual dialog with sparse graph learning and knowledge transfer. In Findings of the Association for Computational Linguistics (EMNLP \u201921), 327\u2013339."},{"issue":"1","key":"e_1_3_2_150_2","doi-asserted-by":"crossref","first-page":"615","DOI":"10.1145\/3093337.3037698","article-title":"Neurosurgeon: Collaborative intelligence between the cloud and mobile edge","volume":"45","author":"Kang Yiping","year":"2017","unstructured":"Yiping Kang, Johann Hauswald, Cao Gao, Austin Rovinski, Trevor Mudge, Jason Mars, and Lingjia Tang. 2017. Neurosurgeon: Collaborative intelligence between the cloud and mobile edge. ACM SIGARCH Computer Architecture News 45, 1 (2017), 615\u2013629.","journal-title":"ACM SIGARCH Computer Architecture News"},{"key":"e_1_3_2_151_2","unstructured":"Anjuli Kannan and Oriol Vinyals. 2017. Adversarial evaluation of dialogue models. arXiv:1701.08198. Retrieved from https:\/\/arxiv.org\/abs\/1701.08198"},{"key":"e_1_3_2_152_2","unstructured":"Jared Kaplan Sam McCandlish Tom Henighan Tom B. Brown Benjamin Chess Rewon Child Scott Gray Alec Radford Jeffrey Wu and Dario Amodei. 2020. Scaling laws for neural language models. arXiv:2001.08361. Retrieved from https:\/\/arxiv.org\/abs\/2001.08361"},{"issue":"8","key":"e_1_3_2_153_2","doi-asserted-by":"crossref","first-page":"5455","DOI":"10.1007\/s10462-020-09825-6","article-title":"A survey of the recent architectures of deep convolutional neural networks","volume":"53","author":"Khan Asifullah","year":"2020","unstructured":"Asifullah Khan, Anabia Sohail, Umme Zahoora, and Aqsa Saeed Qureshi. 2020. A survey of the recent architectures of deep convolutional neural networks. Artificial Intelligence Review 53, 8 (2020), 5455\u20135516.","journal-title":"Artificial Intelligence Review"},{"issue":"10","key":"e_1_3_2_154_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3505244","article-title":"Transformers in vision: A survey","volume":"54","author":"Khan Salman","year":"2022","unstructured":"Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. 2022. Transformers in vision: A survey. ACM Computing Surveys 54, 10s (2022), 1\u201341.","journal-title":"ACM Computing Surveys"},{"key":"e_1_3_2_155_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Kim Byeongchang","year":"2019","unstructured":"Byeongchang Kim, Jaewoo Ahn, and Gunhee Kim. 2019. Sequential latent knowledge selection for knowledge-grounded dialogue. In Proceedings of the International Conference on Learning Representations, 1\u201314."},{"key":"e_1_3_2_156_2","first-page":"1789","volume-title":"In Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Kim Junyeong","year":"2021","unstructured":"Junyeong Kim, Sunjae Yoon, Dahyun Kim, and Chang D. Yoo. 2021. Structured co-reference graph attention for video-grounded dialogue. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35, 1789\u20131797."},{"key":"e_1_3_2_157_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Kipf Thomas N.","year":"2017","unstructured":"Thomas N. Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In Proceedings of the International Conference on Learning Representations, 1\u201314."},{"key":"e_1_3_2_158_2","first-page":"17283","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Koh Jing Yu","year":"2023","unstructured":"Jing Yu Koh, Ruslan Salakhutdinov, and Daniel Fried. 2023. Grounding language models to images for multimodal inputs and outputs. In Proceedings of the International Conference on Machine Learning. PMLR, 17283\u201317300."},{"key":"e_1_3_2_159_2","first-page":"16020","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Kondratyuk Dan","year":"2021","unstructured":"Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew Brown, and Boqing Gong. 2021. MoviNets: Mobile video networks for efficient video recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 16020\u201316030."},{"key":"e_1_3_2_160_2","first-page":"371","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track","author":"Kong Yilun","year":"2024","unstructured":"Yilun Kong, Jingqing Ruan, YiHong Chen, Bin Zhang, Tianpeng Bao, Hangyu Mao, Ziyue Li, Xingyu Zeng, Rui Zhao, Xueqian Wang, et al. 2024. TPTU-v2: Boosting task planning and tool usage of large language model-based agents in real-world systems. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, 371\u2013385."},{"key":"e_1_3_2_161_2","first-page":"153","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Kottur Satwik","year":"2018","unstructured":"Satwik Kottur, Jos\u00e9 M. F. Moura, Devi Parikh, Dhruv Batra, and Marcus Rohrbach. 2018. Visual coreference resolution in visual dialog using neural module networks. In Proceedings of the European Conference on Computer Vision (ECCV), 153\u2013169."},{"key":"e_1_3_2_162_2","first-page":"582","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Kottur Satwik","year":"2019","unstructured":"Satwik Kottur, Jos\u00e9 M. F. Moura, Devi Parikh, Dhruv Batra, and Marcus Rohrbach. 2019. CLEVR-Dialog: A diagnostic dataset for multi-round reasoning in visual dialog. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 582\u2013595."},{"key":"e_1_3_2_163_2","doi-asserted-by":"crossref","first-page":"5260","DOI":"10.1145\/3637528.3671573","volume-title":"Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Kuang Weirui","year":"2024","unstructured":"Weirui Kuang, Bingchen Qian, Zitao Li, Daoyuan Chen, Dawei Gao, Xuchen Pan, Yuexiang Xie, Yaliang Li, Bolin Ding, and Jingren Zhou. 2024. Federatedscope-LLM: A comprehensive package for fine-tuning large language models in federated learning. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 5260\u20135271."},{"issue":"25","key":"e_1_3_2_164_2","doi-asserted-by":"crossref","first-page":"3559","DOI":"10.1016\/S0042-6989(01)00102-X","article-title":"In what ways do eye movements contribute to everyday activities?","volume":"41","author":"Land Michael F.","year":"2001","unstructured":"Michael F. Land and Mary Hayhoe. 2001. In what ways do eye movements contribute to everyday activities? Vision Research 41, 25\u201326 (2001), 3559\u20133565.","journal-title":"Vision Research"},{"key":"e_1_3_2_165_2","first-page":"71683","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Lauren\u00e7on Hugo","year":"2024","unstructured":"Hugo Lauren\u00e7on, Lucile Saulnier, L\u00e9o Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander Rush, Douwe Kiela, et al. 2024. Obelics: An open web-scale filtered dataset of interleaved image-text documents. In Proceedings of the Advances in Neural Information Processing Systems, 71683\u201371702."},{"key":"e_1_3_2_166_2","unstructured":"Hugo Lauren\u00e7on L\u00e9o Tronchon Matthieu Cord and Victor Sanh. 2024. What matters when building vision-language models? arXiv:2405.02246. Retrieved from https:\/\/arxiv.org\/abs\/2405.02246"},{"key":"e_1_3_2_167_2","first-page":"228","volume-title":"Proceedings of the 2nd Workshop on Statistical Machine Translation","author":"Lavie Alon","year":"2007","unstructured":"Alon Lavie and Abhaya Agarwal. 2007. METEOR: An automatic metric for MT evaluation with high levels of correlation with human judgments. In Proceedings of the 2nd Workshop on Statistical Machine Translation, 228\u2013231."},{"key":"e_1_3_2_168_2","first-page":"3377","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Le Hung","year":"2022","unstructured":"Hung Le, Nancy Chen, and Steven Hoi. 2022. VGNMN: Video-grounded neural module networks for video-grounded dialogue systems. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 3377\u20133393."},{"key":"e_1_3_2_169_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Le Hung","year":"2020","unstructured":"Hung Le, Nancy F. Chen, and Steven Hoi. 2020. Learning reasoning paths over semantic graphs for video-grounded dialogues. In Proceedings of the International Conference on Learning Representations, 1\u201318."},{"key":"e_1_3_2_170_2","first-page":"5612","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Le Hung","year":"2019","unstructured":"Hung Le, Doyen Sahoo, Nancy Chen, and Steven Hoi. 2019. Multimodal transformer networks for end-to-end video-grounded dialogue systems. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 5612\u20135623."},{"key":"e_1_3_2_171_2","first-page":"1846","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Le Hung","year":"2020","unstructured":"Hung Le, Doyen Sahoo, Nancy Chen, and Steven C. H. Hoi. 2020. BiST: Bi-directional spatio-temporal reasoning for video-grounded dialogues. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 1846\u20131859."},{"key":"e_1_3_2_172_2","first-page":"5651","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing","author":"Le Hung","year":"2021","unstructured":"Hung Le, Chinnadhurai Sankar, Seungwhan Moon, Ahmad Beirami, Alborz Geramifard, and Satwik Kottur. 2021. DVD: A diagnostic dataset for multi-step reasoning in video grounded dialogue. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, 5651\u20135665."},{"key":"e_1_3_2_173_2","doi-asserted-by":"crossref","first-page":"420","DOI":"10.1016\/j.neunet.2021.05.034","article-title":"QTTNet: Quantized tensor train neural networks for 3D object and video recognition","volume":"141","author":"Lee Donghyun","year":"2021","unstructured":"Donghyun Lee, Dingheng Wang, Yukuan Yang, Lei Deng, Guangshe Zhao, and Guoqi Li. 2021. QTTNet: Quantized tensor train neural networks for 3D object and video recognition. Neural Networks 141 (2021), 420\u2013432.","journal-title":"Neural Networks"},{"key":"e_1_3_2_174_2","first-page":"391","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Lee Seongyun","year":"2024","unstructured":"Seongyun Lee, Sue Park, Yongrae Jo, and Minjoon Seo. 2024. Volcano: Mitigating multimodal hallucination through self-feedback guided revision. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 391\u2013404."},{"key":"e_1_3_2_175_2","unstructured":"Bohao Li Yuying Ge Yi Chen Yixiao Ge Ruimao Zhang and Ying Shan. 2024. Seed-bench-2-plus: Benchmarking multimodal large language models with text-rich visual comprehension. arXiv:2404.16790. Retrieved from https:\/\/arxiv.org\/abs\/2404.16790"},{"key":"e_1_3_2_176_2","unstructured":"Bohao Li Yuying Ge Yixiao Ge Guangzhi Wang Rui Wang Ruimao Zhang and Ying Shan. 2023. SEED-Bench-2: Benchmarking multimodal large language models. arXiv:2311.17092. Retrieved from https:\/\/arxiv.org\/abs\/2311.17092"},{"key":"e_1_3_2_177_2","first-page":"13299","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Bohao","year":"2024","unstructured":"Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. 2024. SEED-Bench: Benchmarking multimodal large language models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 13299\u201313308."},{"key":"e_1_3_2_178_2","unstructured":"Bo Li Yuanhan Zhang Liangyu Chen Jinghao Wang Jingkang Yang and Ziwei Liu. 2023. Otter: A multi-modal model with in-context instruction tuning. arXiv:2305.03726. Retrieved from https:\/\/arxiv.org\/abs\/2305.03726"},{"key":"e_1_3_2_179_2","unstructured":"Bo Li Yuanhan Zhang Dong Guo Renrui Zhang Feng Li Hao Zhang Kaichen Zhang Yanwei Li Ziwei Liu and Chunyuan Li. 2024. Llava-onevision: Easy visual task transfer. arXiv:2408.03326. Retrieved from https:\/\/arxiv.org\/abs\/2408.03326"},{"key":"e_1_3_2_180_2","first-page":"3490","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics","author":"Li Chaofan","year":"2024","unstructured":"Chaofan Li, Zheng Liu, Shitao Xiao, Yingxia Shao, and Defu Lian. 2024. Llama2vec: Unsupervised adaptation of large language models for dense retrieval. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 3490\u20133500."},{"key":"e_1_3_2_181_2","unstructured":"Dongxu Li Yudong Liu Haoning Wu Yue Wang Zhiqi Shen Bowen Qu Xinyao Niu Guoyin Wang Bei Chen and Junnan Li. 2024. Aria: An open multimodal native mixture-of-experts model. arXiv:2410.05993. Retrieved from https:\/\/arxiv.org\/abs\/2410.05993"},{"key":"e_1_3_2_182_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Li Hao","year":"2016","unstructured":"Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf. 2016. Pruning filters for efficient ConvNets. In Proceedings of the International Conference on Learning Representations, 1\u201313."},{"key":"e_1_3_2_183_2","first-page":"5584","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Li Haopeng","year":"2023","unstructured":"Haopeng Li, Qiuhong Ke, Mingming Gong, and Tom Drummond. 2023. Progressive video summarization via multimodal self-supervised learning. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, 5584\u20135593."},{"key":"e_1_3_2_184_2","first-page":"4952","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Li Hongxin","year":"2024","unstructured":"Hongxin Li, Jingran Su, Yuntao Chen, Qing Li, and Zhao-Xiang Zhang. 2024. SheetCopilot: Bringing software productivity to the next level through large language models. In Proceedings of the Advances in Neural Information Processing Systems., 4952\u20134984."},{"key":"e_1_3_2_185_2","first-page":"19730","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the International Conference on Machine Learning. PMLR, 19730\u201319742."},{"key":"e_1_3_2_186_2","first-page":"12888","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine Learning. PMLR, 12888\u201312900."},{"key":"e_1_3_2_187_2","unstructured":"KunChang Li Yinan He Yi Wang Yizhuo Li Wenhai Wang Ping Luo Yali Wang Limin Wang and Yu Qiao. 2023. Videochat: Chat-centric video understanding. arXiv:2305.06355. Retrieved from https:\/\/arxiv.org\/abs\/2305.06355"},{"key":"e_1_3_2_188_2","first-page":"22195","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Kunchang","year":"2024","unstructured":"Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. 2024. Mvbench: A comprehensive multi-modal video understanding benchmark. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 22195\u201322206."},{"key":"e_1_3_2_189_2","first-page":"10313","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Li Linjie","year":"2019","unstructured":"Linjie Li, Zhe Gan, Yu Cheng, and Jingjing Liu. 2019. Relation-aware graph attention network for visual question answering. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 10313\u201310322."},{"key":"e_1_3_2_190_2","first-page":"8475","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Li Linxiao","year":"2020","unstructured":"Linxiao Li, Can Xu, Wei Wu, Yufan Zhao, Xueliang Zhao, and Chongyang Tao. 2020. Zero-resource knowledge-grounded dialogue generation. In Proceedings of the Advances in Neural Information Processing Systems, 8475\u20138485."},{"key":"e_1_3_2_191_2","doi-asserted-by":"crossref","first-page":"1348","DOI":"10.1145\/3583780.3615017","volume-title":"Proceedings of the 32nd ACM International Conference on Information and Knowledge Management","author":"Li Lei","year":"2023","unstructured":"Lei Li, Yongfeng Zhang, and Li Chen. 2023. Prompt distillation for efficient LLM-based recommendation. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management, 1348\u20131357."},{"key":"e_1_3_2_192_2","doi-asserted-by":"crossref","first-page":"107677","DOI":"10.1016\/j.patcog.2020.107677","article-title":"Exploring global diverse attention via pairwise temporal relation for video summarization","volume":"111","author":"Li Ping","year":"2021","unstructured":"Ping Li, Qinghao Ye, Luming Zhang, Li Yuan, Xianghua Xu, and Ling Shao. 2021. Exploring global diverse attention via pairwise temporal relation for video summarization. Pattern Recognition 111 (2021), 107677.","journal-title":"Pattern Recognition"},{"key":"e_1_3_2_193_2","first-page":"2845","volume-title":"Proceedings of the 25th International Joint Conference on Artificial Intelligence","author":"Li Xiang","year":"2016","unstructured":"Xiang Li, Lili Mou, Rui Yan, and Ming Zhang. 2016. StalemateBreaker: A proactive content-introducing approach to automatic human-computer conversation. In Proceedings of the 25th International Joint Conference on Artificial Intelligence, 2845\u20132851."},{"key":"e_1_3_2_194_2","doi-asserted-by":"crossref","first-page":"10952","DOI":"10.1109\/TMM.2024.3428317","article-title":"Lmeye: An interactive perception network for large language models","volume":"26","author":"Li Yunxin","year":"2024","unstructured":"Yunxin Li, Baotian Hu, Xinyu Chen, Lin Ma, Yong Xu, and Min Zhang. 2024. Lmeye: An interactive perception network for large language models. IEEE Transactions on Multimedia 26 (2024), 10952\u201310964.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_195_2","first-page":"986","volume-title":"Proceedings of the 8th International Joint Conference on Natural Language Processing","author":"Li Yanran","year":"2017","unstructured":"Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. DailyDialog: A manually labelled multi-turn dialogue dataset. In Proceedings of the 8th International Joint Conference on Natural Language Processing, 986\u2013995."},{"key":"e_1_3_2_196_2","first-page":"323","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Li Yanwei","year":"2025","unstructured":"Yanwei Li, Chengyao Wang, and Jiaya Jia. 2025. Llama-vid: An image is worth 2 tokens in large language models. In Proceedings of the European Conference on Computer Vision. Springer, 323\u2013340."},{"key":"e_1_3_2_197_2","unstructured":"Yanda Li Chi Zhang Gang Yu Zhibin Wang Bin Fu Guosheng Lin Chunhua Shen Ling Chen and Yunchao Wei. 2023. Stablellava: Enhanced visual instruction tuning with synthesized image-dialogue data. arXiv:2308.10253. Retrieved from https:\/\/arxiv.org\/abs\/2308.10253"},{"issue":"6","key":"e_1_3_2_198_2","doi-asserted-by":"crossref","first-page":"6939","DOI":"10.1007\/s11042-018-6445-z","article-title":"A survey on sentiment analysis and opinion mining for social multimedia","volume":"78","author":"Li Zuhe","year":"2019","unstructured":"Zuhe Li, Yangyu Fan, Bin Jiang, Tao Lei, and Weihua Liu. 2019. A survey on sentiment analysis and opinion mining for social multimedia. Multimedia Tools and Applications 78, 6 (2019), 6939\u20136967.","journal-title":"Multimedia Tools and Applications"},{"key":"e_1_3_2_199_2","first-page":"17065","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Li Zhikai","year":"2023","unstructured":"Zhikai Li and Qingyi Gu. 2023. I-vit: Integer-only quantization for efficient vision transformer inference. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 17065\u201317075."},{"issue":"10","key":"e_1_3_2_200_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3656580","article-title":"Foundations & trends in multimodal machine learning: Principles, challenges, and open questions","volume":"56","author":"Liang Paul Pu","year":"2024","unstructured":"Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2024. Foundations & trends in multimodal machine learning: Principles, challenges, and open questions. ACM Computing Surveys 56, 10 (2024), 1\u201342.","journal-title":"ACM Computing Surveys"},{"key":"e_1_3_2_201_2","doi-asserted-by":"crossref","first-page":"0063","DOI":"10.34133\/icomputing.0063","article-title":"TaskMatrix.AI: Completing tasks by connecting foundation models with millions of APIs","volume":"3","author":"Liang Yaobo","year":"2024","unstructured":"Yaobo Liang, Chenfei Wu, Ting Song, Wenshan Wu, Yan Xia, Yu Liu, Yang Ou, Shuai Lu, Lei Ji, Shaoguang Mao, et al. 2024. TaskMatrix.AI: Completing tasks by connecting foundation models with millions of APIs. Intelligent Computing 3 (2024), 0063.","journal-title":"Intelligent Computing"},{"key":"e_1_3_2_202_2","doi-asserted-by":"crossref","first-page":"675","DOI":"10.1145\/3404835.3462970","volume-title":"Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Liao Lizi","year":"2021","unstructured":"Lizi Liao, Le Hong Long, Zheng Zhang, Minlie Huang, and Tat-Seng Chua. 2021. MMConv: An environment for multimodal conversational search across multiple domains. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 675\u2013684."},{"key":"e_1_3_2_203_2","doi-asserted-by":"crossref","first-page":"801","DOI":"10.1145\/3240508.3240605","volume-title":"Proceedings of the 26th ACM International Conference on Multimedia","author":"Liao Lizi","year":"2018","unstructured":"Lizi Liao, Yunshan Ma, Xiangnan He, Richang Hong, and Tat-seng Chua. 2018. Knowledge-aware multimodal dialogue systems. In Proceedings of the 26th ACM International Conference on Multimedia, 801\u2013809."},{"key":"e_1_3_2_204_2","unstructured":"Bin Lin Zhenyu Tang Yang Ye Jiaxi Cui Bin Zhu Peng Jin Junwu Zhang Munan Ning and Li Yuan. 2024. Moe-llava: Mixture of experts for large vision-language models. arXiv:2401.15947. Retrieved from https:\/\/arxiv.org\/abs\/2401.15947"},{"key":"e_1_3_2_205_2","unstructured":"Bin Lin Bin Zhu Yang Ye Munan Ning Peng Jin and Li Yuan. 2023. Video-llava: Learning united visual representation by alignment before projection. arXiv:2311.10122. Retrieved from https:\/\/arxiv.org\/abs\/2311.10122"},{"key":"e_1_3_2_206_2","first-page":"150","volume-title":"Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics","author":"Lin Chin-Yew","year":"2003","unstructured":"Chin-Yew Lin and Eduard Hovy. 2003. Automatic evaluation of summaries using n-gram co-occurrence statistics. In Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics, 150\u2013157."},{"key":"e_1_3_2_207_2","first-page":"87","volume-title":"Proceedings of Machine Learning and Systems","volume":"6","author":"Lin Ji","year":"2024","unstructured":"Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024. AWQ: Activation-aware weight quantization for on-Device LLM compression and acceleration. Proceedings of Machine Learning and Systems 6 (2024), 87\u2013100."},{"key":"e_1_3_2_208_2","first-page":"26689","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lin Ji","year":"2024","unstructured":"Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han. 2024. Vila: On pre-training for visual language models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 26689\u201326699."},{"key":"e_1_3_2_209_2","first-page":"740","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In Proceedings of the European Conference on Computer Vision. Springer, 740\u2013755."},{"key":"e_1_3_2_210_2","first-page":"32400","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Liu Dongyang","year":"2024","unstructured":"Dongyang Liu, Renrui Zhang, Longtian Qiu, Siyuan Huang, Weifeng Lin, Shitian Zhao, Shijie Geng, Ziyi Lin, Peng Jin, Kaipeng Zhang, et al. 2024. SPHINX-X: Scaling data and parameters for a family of multi-modal large language models. In Proceedings of the 41st International Conference on Machine Learning, 32400\u201332420."},{"key":"e_1_3_2_211_2","first-page":"26296","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu Haotian","year":"2024","unstructured":"Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved baselines with visual instruction tuning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 26296\u201326306."},{"key":"e_1_3_2_212_2","first-page":"49250","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Liu Haotian","year":"2024","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. In Proceedings of the Advances in Neural Information Processing Systems, 49250\u201349267."},{"key":"e_1_3_2_213_2","first-page":"4876","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Liu Ning","year":"2020","unstructured":"Ning Liu, Xiaolong Ma, Zhiyuan Xu, Yanzhi Wang, Jian Tang, and Jieping Ye. 2020. Autocompress: An automatic DNN structured pruning framework for ultra-high compression rates. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34, 4876\u20134883."},{"key":"e_1_3_2_214_2","first-page":"6566","volume-title":"Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Liu Qijiong","year":"2024","unstructured":"Qijiong Liu, Jieming Zhu, Yanting Yang, Quanyu Dai, Zhaocheng Du, Xiao-Ming Wu, Zhou Zhao, Rui Zhang, and Zhenhua Dong. 2024. Multimodal pretraining, adaptation, and generation for recommendation: A survey. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 6566\u20136576."},{"issue":"1","key":"e_1_3_2_215_2","first-page":"389","article-title":"Enabling resource-efficient AIoT system with cross-level optimization: A survey","volume":"26","author":"Liu Sicong","year":"2023","unstructured":"Sicong Liu, Bin Guo, Cheng Fang, Ziqi Wang, Shiyan Luo, Zimu Zhou, and Zhiwen Yu. 2023. Enabling resource-efficient AIoT system with cross-level optimization: A survey. IEEE Communications Surveys & Tutorials 26, 1 (2023), 389\u2013427.","journal-title":"IEEE Communications Surveys & Tutorials"},{"key":"e_1_3_2_216_2","doi-asserted-by":"crossref","first-page":"1573","DOI":"10.1109\/TIP.2022.3143699","article-title":"Video summarization through reinforcement learning with a 3D spatio-temporal U-net","volume":"31","author":"Liu Tianrui","year":"2022","unstructured":"Tianrui Liu, Qingjie Meng, Jun-Jie Huang, Athanasios Vlontzos, Daniel Rueckert, and Bernhard Kainz. 2022. Video summarization through reinforcement learning with a 3D spatio-temporal U-net. IEEE Transactions on Image Processing 31 (2022), 1573\u20131586.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_2_217_2","unstructured":"Yang Liu Weixing Chen Yongjie Bai Xiaodan Liang Guanbin Li Wen Gao and Liang Lin. 2024. Aligning cyber space with physical world: A comprehensive survey on embodied AI. arXiv:2407.06886. Retrieved from https:\/\/arxiv.org\/abs\/2407.06886"},{"issue":"10","key":"e_1_3_2_218_2","doi-asserted-by":"crossref","first-page":"11624","DOI":"10.1109\/TPAMI.2023.3284038","article-title":"Cross-modal causal relational reasoning for event-level visual question answering","volume":"45","author":"Liu Yang","year":"2023","unstructured":"Yang Liu, Guanbin Li, and Liang Lin. 2023. Cross-modal causal relational reasoning for event-level visual question answering. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 10 (2023), 11624\u201311641.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_219_2","unstructured":"Yuanxin Liu Shicheng Li Yi Liu Yuxiang Wang Shuhuai Ren Lei Li Sishuo Chen Xu Sun and Lu Hou. 2024. Tempcompass: Do video LLMs really understand videos? arXiv:2403.00476. Retrieved from https:\/\/arxiv.org\/abs\/2403.00476"},{"key":"e_1_3_2_220_2","first-page":"4861","volume-title":"Findings of the Association for Computational Linguistics (EMNLP \u201920)","author":"Liu Yafei","year":"2020","unstructured":"Yafei Liu, Hongjin Qian, Hengpeng Xu, and Jinmao Wei. 2020. Speaker or listener? The role of a dialog agent. In Findings of the Association for Computational Linguistics (EMNLP \u201920), 4861\u20134869."},{"key":"e_1_3_2_221_2","unstructured":"Zuyan Liu Yuhao Dong Ziwei Liu Winston Hu Jiwen Lu and Yongming Rao. 2024. Oryx MLLM: On-demand spatial-temporal understanding at arbitrary resolution. arXiv:2409.12961. Retrieved from https:\/\/arxiv.org\/abs\/2409.12961"},{"key":"e_1_3_2_222_2","first-page":"1116","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics","author":"Lowe Ryan","year":"2017","unstructured":"Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017. Towards an automatic Turing test: Learning to evaluate dialogue responses. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, 1116\u20131126."},{"key":"e_1_3_2_223_2","unstructured":"Haoyu Lu Wen Liu Bo Zhang Bingxuan Wang Kai Dong Bo Liu Jingxiang Sun Tongzheng Ren Zhuoshu Li Yaofeng Sun et al. 2024. DeepSeek-VL: Towards real-world vision-language understanding. arXiv:2403.05525. Retrieved from https:\/\/arxiv.org\/abs\/2403.05525"},{"issue":"1","key":"e_1_3_2_224_2","first-page":"1","article-title":"Chinese image captioning via fuzzy attention-based DenseNet-BiLSTM","volume":"17","author":"Lu Huimin","year":"2021","unstructured":"Huimin Lu, Rui Yang, Zhenrong Deng, Yonglin Zhang, Guangwei Gao, and Rushi Lan. 2021. Chinese image captioning via fuzzy attention-based DenseNet-BiLSTM. ACM Transactions on Multimedia Computing, Communications, and Applications 17, 1s (2021), 1\u201318.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_2_225_2","first-page":"313","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Lu Jiasen","year":"2017","unstructured":"Jiasen Lu, Anitha Kannan, Jianwei Yang, Devi Parikh, and Dhruv Batra. 2017. Best of both worlds: Transferring knowledge from discriminative learning to a generative visual dialog model. In Proceedings of the Advances in Neural Information Processing Systems, 313\u2013323."},{"key":"e_1_3_2_226_2","first-page":"1","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Lu Pan","year":"2024","unstructured":"Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. 2024. MathVista: Evaluating mathematical reasoning of foundation models in visual contexts. In Proceedings of the 12th International Conference on Learning Representations, 1\u201317."},{"key":"e_1_3_2_227_2","first-page":"29615","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Luo Gen","year":"2024","unstructured":"Gen Luo, Yiyi Zhou, Tianhe Ren, Shengxin Chen, Xiaoshuai Sun, and Rongrong Ji. 2024. Cheap and quick: Efficient vision-language instruction tuning for large language models. In Proceedings of the Advances in Neural Information Processing Systems, 29615\u201329627."},{"key":"e_1_3_2_228_2","unstructured":"Ruipu Luo Ziwang Zhao Min Yang Junwei Dong Da Li Pengcheng Lu Tao Wang Linmei Hu Minghui Qiu and Zhongyu Wei. 2023. Valley: Video assistant with large language model enhanced ability. arXiv:2306.07207. Retrieved from https:\/\/arxiv.org\/abs\/2306.07207"},{"issue":"6","key":"e_1_3_2_229_2","first-page":"1683","article-title":"Image and video compression with neural networks: A review","volume":"30","author":"Ma Siwei","year":"2019","unstructured":"Siwei Ma, Xinfeng Zhang, Chuanmin Jia, Zhenghui Zhao, Shiqi Wang, and Shanshe Wang. 2019. Image and video compression with neural networks: A review. IEEE Transactions on Circuits and Systems for Video Technology 30, 6 (2019), 1683\u20131698.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_2_230_2","first-page":"21702","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Ma Xinyin","year":"2023","unstructured":"Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2023. LLM-pruner: On the structural pruning of large language models. In Proceedings of the Advances in Neural Information Processing Systems, 21702\u201321720."},{"key":"e_1_3_2_231_2","doi-asserted-by":"crossref","first-page":"50","DOI":"10.1016\/j.inffus.2020.06.011","article-title":"A survey on empathetic dialogue systems","volume":"64","author":"Ma Yukun","year":"2020","unstructured":"Yukun Ma, Khanh Linh Nguyen, Frank Z. Xing, and Erik Cambria. 2020. A survey on empathetic dialogue systems. Information Fusion 64 (2020), 50\u201370.","journal-title":"Information Fusion"},{"key":"e_1_3_2_232_2","first-page":"12585","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics","author":"Maaz Muhammad","year":"2024","unstructured":"Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. 2024. Video-ChatGPT: Towards detailed video understanding via large vision and language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 12585\u201312602."},{"key":"e_1_3_2_233_2","first-page":"16488","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Majumdar Arjun","year":"2024","unstructured":"Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul Mcvay, Oleksandr Maksymets, Sergio Arnaud, et al. 2024. OpenEQA: Embodied question answering in the era of foundation models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 16488\u201316498."},{"key":"e_1_3_2_234_2","first-page":"46212","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Mangalam Karttikeya","year":"2023","unstructured":"Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. 2023. EgoSchema: A diagnostic benchmark for very long-form video language understanding. In Proceedings of the Advances in Neural Information Processing Systems, 46212\u201346244."},{"issue":"3","key":"e_1_3_2_235_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3617833","article-title":"Multimodality representation learning: A survey on evolution, pretraining and its applications","volume":"20","author":"Manzoor Muhammad Arslan","year":"2023","unstructured":"Muhammad Arslan Manzoor, Sarah Albarri, Ziting Xian, Zaiqiao Meng, Preslav Nakov, and Shangsong Liang. 2023. Multimodality representation learning: A survey on evolution, pretraining and its applications. ACM Transactions on Multimedia Computing, Communications and Applications 20, 3 (2023), 1\u201334.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_236_2","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/9780262514620.001.0001","volume-title":"Vision: A Computational Investigation into the Human Representation and Processing of Visual Information","author":"Marr David","year":"2010","unstructured":"David Marr. 2010. Vision: A Computational Investigation into the Human Representation and Processing of Visual Information. MIT Press."},{"key":"e_1_3_2_237_2","first-page":"6097","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Massiceti Daniela","year":"2018","unstructured":"Daniela Massiceti, N. Siddharth, Puneet K. Dokania, and Philip H. S. Torr. 2018. FlipDial: A generative model for two-way visual dialogue. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6097\u20136105."},{"key":"e_1_3_2_238_2","doi-asserted-by":"crossref","first-page":"681","DOI":"10.18653\/v1\/2020.acl-main.64","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Mehri Shikib","year":"2020","unstructured":"Shikib Mehri and Maxine Eskenazi. 2020. USR: An unsupervised and reference free evaluation metric for dialog generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 681\u2013707."},{"key":"e_1_3_2_239_2","unstructured":"Yuxian Meng Shuhe Wang Qinghong Han Xiaofei Sun Fei Wu Rui Yan and Jiwei Li. 2020. Openvidial: A large-scale open-domain dialogue dataset with visual contexts. arXiv:2012.15015. Retrieved from https:\/\/arxiv.org\/abs\/2012.15015"},{"issue":"12","key":"e_1_3_2_240_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3578938","article-title":"Efficient deep learning: A survey on making deep learning models smaller, faster, and better","volume":"55","author":"Menghani Gaurav","year":"2023","unstructured":"Gaurav Menghani. 2023. Efficient deep learning: A survey on making deep learning models smaller, faster, and better. ACM Computing Surveys 55, 12 (2023), 1\u201337.","journal-title":"ACM Computing Surveys"},{"issue":"2","key":"e_1_3_2_241_2","first-page":"1","article-title":"Recent advances in natural language processing via large pre-trained language models: A survey","volume":"56","author":"Min Bonan","year":"2023","unstructured":"Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth. 2023. Recent advances in natural language processing via large pre-trained language models: A survey. ACM Computing Surveys 56, 2 (2023), 1\u201340.","journal-title":"ACM Computing Surveys"},{"key":"e_1_3_2_242_2","first-page":"10794","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Minnen David","year":"2018","unstructured":"David Minnen, Johannes Ball\u00e9, and George D. Toderici. 2018. Joint autoregressive and hierarchical priors for learned image compression. In Proceedings of the Advances in Neural Information Processing Systems, 10794\u201310803."},{"key":"e_1_3_2_243_2","first-page":"25081","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Mu Yao","year":"2024","unstructured":"Yao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang, Mingyu Ding, Jun Jin, Bin Wang, Jifeng Dai, Yu Qiao, and Ping Luo. 2024. EmbodiedGPT: Vision-language pre-training via embodied chain of thought. In Proceedings of the Advances in Neural Information Processing Systems, 25081\u201325094."},{"key":"e_1_3_2_244_2","first-page":"3573","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Mullapudi Ravi Teja","year":"2019","unstructured":"Ravi Teja Mullapudi, Steven Chen, Keyi Zhang, Deva Ramanan, and Kayvon Fatahalian. 2019. Online model distillation for efficient video inference. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 3573\u20133582."},{"key":"e_1_3_2_245_2","doi-asserted-by":"crossref","first-page":"1449","DOI":"10.18653\/v1\/D19-1152","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Murahari Vishvak","year":"2019","unstructured":"Vishvak Murahari, Prithvijit Chattopadhyay, Dhruv Batra, Devi Parikh, and Abhishek Das. 2019. Improving generative visual dialog by answering diverse questions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 1449\u20131454."},{"key":"e_1_3_2_246_2","first-page":"6087","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Nguyen Duy-Kien","year":"2018","unstructured":"Duy-Kien Nguyen and Takayuki Okatani. 2018. Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6087\u20136096."},{"key":"e_1_3_2_247_2","first-page":"223","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Nguyen Van-Quang","year":"2020","unstructured":"Van-Quang Nguyen, Masanori Suganuma, and Takayuki Okatani. 2020. Efficient attention mechanism for visual dialog that can handle all the interactions between multiple inputs. In Proceedings of the European Conference on Computer Vision. Springer, 223\u2013240."},{"issue":"4","key":"e_1_3_2_248_2","doi-asserted-by":"crossref","first-page":"3055","DOI":"10.1007\/s10462-022-10248-8","article-title":"Recent advances in deep learning based dialogue systems: A systematic survey","volume":"56","author":"Ni Jinjie","year":"2023","unstructured":"Jinjie Ni, Tom Young, Vlad Pandelea, Fuzhao Xue, and Erik Cambria. 2023. Recent advances in deep learning based dialogue systems: A systematic survey. Artificial Intelligence Review 56, 4 (2023), 3055\u20133155.","journal-title":"Artificial Intelligence Review"},{"key":"e_1_3_2_249_2","unstructured":"Munan Ning Bin Zhu Yujia Xie Bin Lin Jiaxi Cui Lu Yuan Dongdong Chen and Li Yuan. 2023. Video-bench: A comprehensive benchmark and toolkit for evaluating video-based large language models. arXiv:2311.16103. Retrieved from https:\/\/arxiv.org\/abs\/2311.16103"},{"key":"e_1_3_2_250_2","first-page":"6679","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Niu Yulei","year":"2019","unstructured":"Yulei Niu, Hanwang Zhang, Manli Zhang, Jianhong Zhang, Zhiwu Lu, and Ji-Rong Wen. 2019. Recursive visual attention in visual dialog. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 6679\u20136688."},{"key":"e_1_3_2_251_2","first-page":"2856","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ost Julian","year":"2021","unstructured":"Julian Ost, Fahim Mannan, Nils Thuerey, Julian Knodt, and Felix Heide. 2021. Neural scene graphs for dynamic scenes. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2856\u20132865."},{"issue":"9","key":"e_1_3_2_252_2","doi-asserted-by":"crossref","first-page":"4177","DOI":"10.1016\/j.eswa.2015.01.041","article-title":"Visual privacy protection methods: A survey","volume":"42","author":"Padilla-L\u00f3pez Jos\u00e9 Ram\u00f3n","year":"2015","unstructured":"Jos\u00e9 Ram\u00f3n Padilla-L\u00f3pez, Alexandros Andre Chaaraoui, and Francisco Fl\u00f3rez-Revuelta. 2015. Visual privacy protection methods: A survey. Expert Systems with Applications 42, 9 (2015), 4177\u20134195.","journal-title":"Expert Systems with Applications"},{"key":"e_1_3_2_253_2","doi-asserted-by":"crossref","DOI":"10.4324\/9781315798868","volume-title":"Imagery and Verbal Processes","author":"Paivio Allan","year":"2013","unstructured":"Allan Paivio. 2013. Imagery and Verbal Processes. Psychology Press."},{"key":"e_1_3_2_254_2","first-page":"1","article-title":"Unifying large language models and knowledge graphs: A roadmap","volume":"01","author":"Pan Shirui","year":"2024","unstructured":"Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu. 2024. Unifying large language models and knowledge graphs: A roadmap. IEEE Transactions on Knowledge and Data Engineering 01 (2024), 1\u201320.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_3_2_255_2","unstructured":"Artemis Panagopoulou Le Xue Ning Yu Junnan Li Dongxu Li Shafiq Joty Ran Xu Silvio Savarese Caiming Xiong and Juan Carlos Niebles. 2023. X-InstructBLIP: A framework for aligning x-modal instruction-aware representations to LLMs and emergent cross-modal reasoning. arXiv:2311.18799. Retrieved from https:\/\/arxiv.org\/abs\/2311.18799"},{"key":"e_1_3_2_256_2","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311\u2013318."},{"key":"e_1_3_2_257_2","volume-title":"In Proceedings of the DSTC7 at AAAI2019 Workshop","author":"Pasunuru Ramakanth","year":"2019","unstructured":"Ramakanth Pasunuru and Mohit Bansal. 2019. Dstc7-avsd: Scene-aware video-dialogue systems with dual attention. In Proceedings of the DSTC7 at AAAI2019 Workshop."},{"key":"e_1_3_2_258_2","unstructured":"Baolin Peng Chunyuan Li Pengcheng He Michel Galley and Jianfeng Gao. 2023. Instruction tuning with gpt-4. arXiv:2304.03277. Retrieved from https:\/\/arxiv.org\/abs\/2304.03277"},{"issue":"9","key":"e_1_3_2_259_2","doi-asserted-by":"crossref","first-page":"2372","DOI":"10.1109\/TCSVT.2017.2705068","article-title":"An overview of cross-media retrieval: Concepts, methodologies, benchmarks, and challenges","volume":"28","author":"Peng Yuxin","year":"2017","unstructured":"Yuxin Peng, Xin Huang, and Yunzhen Zhao. 2017. An overview of cross-media retrieval: Concepts, methodologies, benchmarks, and challenges. IEEE Transactions on Circuits and Systems for Video Technology 28, 9 (2017), 2372\u20132385.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_2_260_2","unstructured":"Zhiliang Peng Wenhui Wang Li Dong Yaru Hao Shaohan Huang Shuming Ma and Furu Wei. 2023. Kosmos-2: Grounding multimodal large language models to the world. arXiv:2306.14824. Retrieved from https:\/\/arxiv.org\/abs\/2306.14824"},{"key":"e_1_3_2_261_2","first-page":"10860","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Qi Jiaxin","year":"2020","unstructured":"Jiaxin Qi, Yulei Niu, Jianqiang Huang, and Hanwang Zhang. 2020. Two causal principles for improving visual dialog. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 10860\u201310869."},{"issue":"3","key":"e_1_3_2_262_2","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1145\/568513.568514","article-title":"Multimodal human discourse: Gesture and speech","volume":"9","author":"Quek Francis","year":"2002","unstructured":"Francis Quek, David McNeill, Robert Bryll, Susan Duncan, Xin-Feng Ma, Cemil Kirbas, Karl E. McCullough, and Rashid Ansari. 2002. Multimodal human discourse: Gesture and speech. ACM Transactions on Computer-Human Interaction 9, 3 (2002), 171\u2013193.","journal-title":"ACM Transactions on Computer-Human Interaction"},{"key":"e_1_3_2_263_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"issue":"8","key":"e_1_3_2_264_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"issue":"11","key":"e_1_3_2_265_2","doi-asserted-by":"crossref","first-page":"13293","DOI":"10.1007\/s10462-023-10414-6","article-title":"Video description: A comprehensive survey of deep learning approaches","volume":"56","author":"Rafiq Ghazala","year":"2023","unstructured":"Ghazala Rafiq, Muhammad Rafiq, and Gyu Sang Choi. 2023. Video description: A comprehensive survey of deep learning approaches. Artificial Intelligence Review 56, 11 (2023), 13293\u201313372.","journal-title":"Artificial Intelligence Review"},{"key":"e_1_3_2_266_2","first-page":"1653","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Rahman Tanzila","year":"2021","unstructured":"Tanzila Rahman, Shih-Han Chou, Leonid Sigal, and Giuseppe Carenini. 2021. An improved attention for visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1653\u20131662."},{"key":"e_1_3_2_267_2","doi-asserted-by":"crossref","first-page":"251","DOI":"10.1145\/1873951.1873987","volume-title":"Proceedings of the 18th ACM International Conference on Multimedia","author":"Rasiwasia Nikhil","year":"2010","unstructured":"Nikhil Rasiwasia, Jose Costa Pereira, Emanuele Coviello, Gabriel Doyle, Gert R. G. Lanckriet, Roger Levy, and Nuno Vasconcelos. 2010. A new approach to cross-modal multimedia retrieval. In Proceedings of the 18th ACM International Conference on Multimedia, 251\u2013260."},{"issue":"9","key":"e_1_3_2_268_2","doi-asserted-by":"crossref","first-page":"2352","DOI":"10.1162\/neco_a_00990","article-title":"Deep convolutional neural networks for image classification: A comprehensive review","volume":"29","author":"Rawat Waseem","year":"2017","unstructured":"Waseem Rawat and Zenghui Wang. 2017. Deep convolutional neural networks for image classification: A comprehensive review. Neural Computation 29, 9 (2017), 2352\u20132449.","journal-title":"Neural Computation"},{"key":"e_1_3_2_269_2","unstructured":"Machel Reid Nikolay Savinov Denis Teplyashin Dmitry Lepikhin Timothy Lillicrap Jean-baptiste Alayrac Radu Soricut Angeliki Lazaridou Orhan Firat Julian Schrittwieser et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv:2403.05530. Retrieved from https:\/\/arxiv.org\/abs\/2403.05530"},{"issue":"5","key":"e_1_3_2_270_2","doi-asserted-by":"crossref","first-page":"5031","DOI":"10.1109\/TVT.2019.2904244","article-title":"Collaborative cloud and edge computing for latency minimization","volume":"68","author":"Ren Jinke","year":"2019","unstructured":"Jinke Ren, Guanding Yu, Yinghui He, and Geoffrey Ye Li. 2019. Collaborative cloud and edge computing for latency minimization. IEEE Transactions on Vehicular Technology 68, 5 (2019), 5031\u20135044.","journal-title":"IEEE Transactions on Vehicular Technology"},{"key":"e_1_3_2_271_2","first-page":"91","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Proceedings of the Advances in Neural Information Processing Systems, 91\u201399."},{"key":"e_1_3_2_272_2","first-page":"14313","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ren Shuhuai","year":"2024","unstructured":"Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, and Lu Hou. 2024. Timechat: A time-sensitive multimodal large language model for long video understanding. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 14313\u201314323."},{"issue":"1","key":"e_1_3_2_273_2","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1080\/135062800394667","article-title":"The dynamic representation of scenes","volume":"7","author":"Rensink Ronald A.","year":"2000","unstructured":"Ronald A. Rensink. 2000. The dynamic representation of scenes. Visual Cognition 7, 1\u20133 (2000), 17\u201342.","journal-title":"Visual Cognition"},{"key":"e_1_3_2_274_2","first-page":"6036","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Ristani Ergys","year":"2018","unstructured":"Ergys Ristani and Carlo Tomasi. 2018. Features for multi-target multi-camera tracking and re-identification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6036\u20136046."},{"issue":"2","key":"e_1_3_2_275_2","doi-asserted-by":"crossref","first-page":"123","DOI":"10.1080\/09540099550039318","article-title":"Catastrophic forgetting, rehearsal and pseudorehearsal","volume":"7","author":"Robins Anthony","year":"1995","unstructured":"Anthony Robins. 1995. Catastrophic forgetting, rehearsal and pseudorehearsal. Connection Science 7, 2 (1995), 123\u2013146.","journal-title":"Connection Science"},{"key":"e_1_3_2_276_2","first-page":"675","volume-title":"Proceedings of the 11th International Conference on Intelligent Tutoring Systems (ITS \u201912)","author":"Rus Vasile","year":"2012","unstructured":"Vasile Rus and Mihai Lintean. 2012. An optimal assessment of natural language student input using word-to-word similarity metrics. In Proceedings of the 11th International Conference on Intelligent Tutoring Systems (ITS \u201912). Springer, 675\u2013676."},{"key":"e_1_3_2_277_2","doi-asserted-by":"crossref","first-page":"106584","DOI":"10.1016\/j.engappai.2023.106584","article-title":"Domain adaptation assisted automatic real-time human-based video summarization","volume":"124","author":"Sabha Ambreen","year":"2023","unstructured":"Ambreen Sabha and Arvind Selwal. 2023. Domain adaptation assisted automatic real-time human-based video summarization. Engineering Applications of Artificial Intelligence 124 (2023), 106584.","journal-title":"Engineering Applications of Artificial Intelligence"},{"issue":"11","key":"e_1_3_2_278_2","doi-asserted-by":"crossref","first-page":"12347","DOI":"10.1007\/s10462-023-10444-0","article-title":"Video summarization using deep learning techniques: A detailed analysis and investigation","volume":"56","author":"Saini Parul","year":"2023","unstructured":"Parul Saini, Krishan Kumar, Shamal Kashid, Ashray Saini, and Alok Negi. 2023. Video summarization using deep learning techniques: A detailed analysis and investigation. Artificial Intelligence Review 56, 11 (2023), 12347\u201312385.","journal-title":"Artificial Intelligence Review"},{"key":"e_1_3_2_279_2","volume-title":"Proceedings of the DSTC7 at AAAI2019 Workshop","volume":"6","author":"Sanabria Ramon","year":"2019","unstructured":"Ramon Sanabria, Shruti Palaskar, and Florian Metze. 2019. CMU Sinbad\u2019s submission for the DSTC7 AVSD challenge. In Proceedings of the DSTC7 at AAAI2019 Workshop, Vol. 6."},{"key":"e_1_3_2_280_2","first-page":"12548","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Schwartz Idan","year":"2019","unstructured":"Idan Schwartz, Alexander G. Schwing, and Tamir Hazan. 2019. A simple baseline for audio-visual scene-aware dialog. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 12548\u201312558."},{"key":"e_1_3_2_281_2","first-page":"2039","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Schwartz Idan","year":"2019","unstructured":"Idan Schwartz, Seunghak Yu, Tamir Hazan, and Alexander G. Schwing. 2019. Factor graph attention. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2039\u20132048."},{"key":"e_1_3_2_282_2","doi-asserted-by":"crossref","first-page":"7881","DOI":"10.18653\/v1\/2020.acl-main.704","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Sellam Thibault","year":"2020","unstructured":"Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 7881\u20137892."},{"key":"e_1_3_2_283_2","first-page":"3722","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Seo Paul Hongsuck","year":"2017","unstructured":"Paul Hongsuck Seo, Andreas Lehrmann, Bohyung Han, and Leonid Sigal. 2017. Visual reference resolution using attention memory for visual dialog. In Proceedings of the Advances in Neural Information Processing Systems, 3722\u20133732."},{"key":"e_1_3_2_284_2","first-page":"3776","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Serban Iulian","year":"2016","unstructured":"Iulian Serban, Alessandro Sordoni, Yoshua Bengio, Aaron Courville, and Joelle Pineau. 2016. Building end-to-end dialogue systems using generative hierarchical neural network models. In Proceedings of the AAAI Conference on Artificial Intelligence, 3776\u20133783."},{"issue":"2","key":"e_1_3_2_285_2","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1145\/3624724","article-title":"Talking about large language models","volume":"67","author":"Shanahan Murray","year":"2024","unstructured":"Murray Shanahan. 2024. Talking about large language models. Communications of the ACM 67, 2 (2024), 68\u201379.","journal-title":"Communications of the ACM"},{"key":"e_1_3_2_286_2","first-page":"1","volume-title":"Proceedings of the CHI Conference on Human Factors in Computing Systems","author":"Sharma Nikhil","year":"2024","unstructured":"Nikhil Sharma, Q. Vera Liao, and Ziang Xiao. 2024. Generative echo chamber? Effect of LLM-powered search systems on diverse information seeking. In Proceedings of the CHI Conference on Human Factors in Computing Systems, 1\u201317."},{"key":"e_1_3_2_287_2","first-page":"853","volume-title":"Proceedings of the IEEE","volume":"86","author":"Sharma Rajeev","year":"1998","unstructured":"Rajeev Sharma, Vladimir I. Pavlovic, and Thomas S. Huang. 1998. Toward multimodal human-computer interface. Proceedings of the IEEE 86, 5 (1998), 853\u2013869."},{"key":"e_1_3_2_288_2","doi-asserted-by":"crossref","first-page":"6442","DOI":"10.18653\/v1\/P19-1646","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Shukla Pushkar","year":"2019","unstructured":"Pushkar Shukla, Carlos Elmadjian, Richika Sharan, Vivek Kulkarni, Matthew Turk, and William Yang Wang. 2019. What should I ask? Using conversationally informative rewards for goal-oriented visual dialog. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 6442\u20136451."},{"key":"e_1_3_2_289_2","doi-asserted-by":"crossref","first-page":"2414","DOI":"10.18653\/v1\/2020.acl-main.219","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Shuster Kurt","year":"2020","unstructured":"Kurt Shuster, Samuel Humeau, Antoine Bordes, and Jason Weston. 2020. Image-Chat: Engaging grounded conversations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2414\u20132429."},{"key":"e_1_3_2_290_2","doi-asserted-by":"crossref","first-page":"2453","DOI":"10.18653\/v1\/2020.acl-main.222","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Shuster Kurt","year":"2020","unstructured":"Kurt Shuster, Da Ju, Stephen Roller, Emily Dinan, Y.-Lan Boureau, and Jason Weston. 2020. The dialogue Dodecathlon: Open-domain knowledge and image grounded conversational agents. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2453\u20132470."},{"key":"e_1_3_2_291_2","first-page":"510","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Sigurdsson Gunnar A.","year":"2016","unstructured":"Gunnar A. Sigurdsson, G\u00fcl Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016. Hollywood in homes: Crowdsourcing data collection for activity understanding. In Proceedings of the European Conference on Computer Vision. Springer, 510\u2013526."},{"key":"e_1_3_2_292_2","doi-asserted-by":"crossref","first-page":"1593","DOI":"10.1109\/ICRA46639.2022.9812166","volume-title":"Proceedings of the 2022 International Conference on Robotics and Automation (ICRA)","author":"Smith Laura","year":"2022","unstructured":"Laura Smith, J. Chase Kew, Xue Bin Peng, Sehoon Ha, Jie Tan, and Sergey Levine. 2022. Legged robots that keep on learning: Fine-tuning locomotion policies in the real world. In Proceedings of the 2022 International Conference on Robotics and Automation (ICRA). IEEE, 1593\u20131599."},{"key":"e_1_3_2_293_2","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1016\/j.imavis.2017.08.003","article-title":"A survey of multimodal sentiment analysis","volume":"65","author":"Soleymani Mohammad","year":"2017","unstructured":"Mohammad Soleymani, David Garcia, Brendan Jou, Bj\u00f6rn Schuller, Shih-Fu Chang, and Maja Pantic. 2017. A survey of multimodal sentiment analysis. Image and Vision Computing 65 (2017), 3\u201314.","journal-title":"Image and Vision Computing"},{"key":"e_1_3_2_294_2","first-page":"18221","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Song Enxin","year":"2024","unstructured":"Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al. 2024. MovieChat: From dense token to sparse memory for long video understanding. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 18221\u201318232."},{"key":"e_1_3_2_295_2","first-page":"167","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing","author":"Song Haoyu","year":"2021","unstructured":"Haoyu Song, Yan Wang, Kaiyan Zhang, Weinan Zhang, and Ting Liu. 2021. BoB: BERT over BERT for training persona-based dialogue models from limited personalized data. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, 167\u2013177."},{"issue":"4","key":"e_1_3_2_296_2","doi-asserted-by":"crossref","first-page":"521","DOI":"10.1162\/089120101753342653","article-title":"A machine learning approach to coreference resolution of noun phrases","volume":"27","author":"Soon Wee Meng","year":"2001","unstructured":"Wee Meng Soon, Hwee Tou Ng, and Daniel Chung Yong Lim. 2001. A machine learning approach to coreference resolution of noun phrases. Computational Linguistics 27, 4 (2001), 521\u2013544.","journal-title":"Computational Linguistics"},{"key":"e_1_3_2_297_2","first-page":"196","volume-title":"Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Sordoni Alessandro","year":"2015","unstructured":"Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, and William B. Dolan. 2015. A neural network approach to context-sensitive generation of conversational responses. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 196\u2013205."},{"key":"e_1_3_2_298_2","first-page":"4444","volume-title":"In Proceedings of the 31st AAAI Conference on Artificial Intelligence","author":"Speer Robyn","year":"2017","unstructured":"Robyn Speer, Joshua Chin, and Catherine Havasi. 2017. ConceptNet 5.5: An open multilingual graph of general knowledge. In Proceedings of the 31st AAAI Conference on Artificial Intelligence, 4444\u20134451."},{"key":"e_1_3_2_299_2","first-page":"2765","volume-title":"Proceedings of the 26th International Joint Conference on Artificial Intelligence","author":"Strub Florian","year":"2017","unstructured":"Florian Strub, Harm De Vries, Jeremie Mary, Bilal Piot, Aaron Courvile, and Olivier Pietquin. 2017. End-to-end optimization of goal-driven and visually grounded dialogue systems. In Proceedings of the 26th International Joint Conference on Artificial Intelligence, 2765\u20132771."},{"key":"e_1_3_2_300_2","doi-asserted-by":"crossref","first-page":"114466","DOI":"10.1016\/j.eswa.2020.114466","article-title":"A neural entity coreference resolution review","volume":"168","author":"Stylianou Nikolaos","year":"2021","unstructured":"Nikolaos Stylianou and Ioannis Vlahavas. 2021. A neural entity coreference resolution review. Expert Systems with Applications 168 (2021), 114466.","journal-title":"Expert Systems with Applications"},{"issue":"5","key":"e_1_3_2_301_2","doi-asserted-by":"crossref","first-page":"103008","DOI":"10.1016\/j.ipm.2022.103008","article-title":"HVLM: Exploring human-like visual cognition and language-memory network for visual dialog","volume":"59","author":"Sun Kaili","year":"2022","unstructured":"Kaili Sun, Chi Guo, Huyin Zhang, and Yuan Li. 2022. HVLM: Exploring human-like visual cognition and language-memory network for visual dialog. Information Processing & Management 59, 5 (2022), 103008.","journal-title":"Information Processing & Management"},{"key":"e_1_3_2_302_2","first-page":"1","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Sun Mingjie","year":"2024","unstructured":"Mingjie Sun, Zhuang Liu, Anna Bair, and J. Zico Kolter. 2024. A simple and effective pruning approach for large language models. In Proceedings of the 12th International Conference on Learning Representations, 1\u201323."},{"key":"e_1_3_2_303_2","first-page":"1","volume-title":"Proceedings of the 2020 57th ACM\/IEEE Design Automation Conference (DAC)","author":"Sun Mengshu","year":"2020","unstructured":"Mengshu Sun, Pu Zhao, Mehmet Gungor, Massoud Pedram, Miriam Leeser, and Xue Lin. 2020. 3D CNN acceleration on FPGA using hardware-aware pruning. In Proceedings of the 2020 57th ACM\/IEEE Design Automation Conference (DAC). IEEE, 1\u20136."},{"key":"e_1_3_2_304_2","first-page":"14454","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Sun Peize","year":"2021","unstructured":"Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, et al. 2021. Sparse R-CNN: End-to-end object detection with learnable proposals. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 14454\u201314463."},{"key":"e_1_3_2_305_2","first-page":"14398","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Sun Quan","year":"2024","unstructured":"Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiying Yu, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. 2024. Generative multimodal models are in-context learners. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 14398\u201314409."},{"key":"e_1_3_2_306_2","first-page":"7375","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Sun Ximeng","year":"2021","unstructured":"Ximeng Sun, Rameswar Panda, Chun-Fu Richard Chen, Aude Oliva, Rogerio Feris, and Kate Saenko. 2021. Dynamic network quantization for efficient video inference. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 7375\u20137385."},{"key":"e_1_3_2_307_2","unstructured":"Yunlong Tang Jing Bi Siting Xu Luchuan Song Susan Liang Teng Wang Daoan Zhang Jie An Jingyang Lin Rongyi Zhu et al. 2023. Video understanding with large language models: A survey. arXiv:2312.17432. Retrieved from https:\/\/arxiv.org\/abs\/2312.17432"},{"key":"e_1_3_2_308_2","first-page":"722","volume-title":"In Proceedings of the 32nd AAAI Conference on Artificial Intelligence","author":"Tao Chongyang","year":"2018","unstructured":"Chongyang Tao, Lili Mou, Dongyan Zhao, and Rui Yan. 2018. Ruber: An unsupervised method for automatic evaluation of open-domain dialog systems. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence, 722\u2013729."},{"key":"e_1_3_2_309_2","unstructured":"Rohan Taori Ishaan Gulrajani Tianyi Zhang Yann Dubois Xuechen Li Carlos Guestrin Percy Liang and Tatsunori B. Hashimoto. 2023. Stanford Alpaca: An Instruction-Following Llama Model. Retrieved from https:\/\/github. com\/tatsu-lab\/stanford_alpaca"},{"key":"e_1_3_2_310_2","first-page":"1","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Tay Yi","year":"2023","unstructured":"Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Steven Zheng, et al. 2023. UL2: Unifying language learning paradigms. In Proceedings of the 11th International Conference on Learning Representations, 1\u201333."},{"key":"e_1_3_2_311_2","first-page":"1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Teney Damien","year":"2017","unstructured":"Damien Teney, Lingqiao Liu, and Anton van Den Hengel. 2017. Graph-structured representations for visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1\u20139."},{"key":"e_1_3_2_312_2","doi-asserted-by":"crossref","first-page":"9330","DOI":"10.18653\/v1\/2021.emnlp-main.736","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Testoni Alberto","year":"2021","unstructured":"Alberto Testoni and Raffaella Bernardi. 2021. Looking for confirmations: An effective and human-like visual dialogue strategy. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 9330\u20139338."},{"key":"e_1_3_2_313_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et al. 2023. Llama: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_2_314_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_2_315_2","first-page":"4489","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Tran Du","year":"2015","unstructured":"Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015. Learning spatiotemporal features with 3D convolutional networks. In Proceedings of the IEEE International Conference on Computer Vision, 4489\u20134497."},{"key":"e_1_3_2_316_2","volume-title":"Computing Machinery and Intelligence","author":"Turing Alan M.","year":"2009","unstructured":"Alan M. Turing. 2009. Computing Machinery and Intelligence. Springer."},{"key":"e_1_3_2_317_2","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1016\/j.patrec.2013.07.003","article-title":"Multimodal interaction: A review","volume":"36","author":"Turk Matthew","year":"2014","unstructured":"Matthew Turk. 2014. Multimodal interaction: A review. Pattern Recognition Letters 36 (2014), 189\u2013195.","journal-title":"Pattern Recognition Letters"},{"key":"e_1_3_2_318_2","doi-asserted-by":"crossref","first-page":"355","DOI":"10.18653\/v1\/W19-8643","volume-title":"Proceedings of the 12th International Conference on Natural Language Generation","author":"Van Der Lee Chris","year":"2019","unstructured":"Chris Van Der Lee, Albert Gatt, Emiel Van Miltenburg, Sander Wubben, and Emiel Krahmer. 2019. Best practices for the human evaluation of automatically generated text. In Proceedings of the 12th International Conference on Natural Language Generation, 355\u2013368."},{"key":"e_1_3_2_319_2","first-page":"6000","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, 6000\u20136010."},{"key":"e_1_3_2_320_2","first-page":"4566","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Vedantam Ramakrishna","year":"2015","unstructured":"Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 4566\u20134575."},{"key":"e_1_3_2_321_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Veli\u010dkovi\u0107 Petar","year":"2018","unstructured":"Petar Veli\u010dkovi\u0107, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Li\u00f2, and Yoshua Bengio. 2018. Graph attention networks. In Proceedings of the International Conference on Learning Representations, 1\u201312."},{"key":"e_1_3_2_322_2","doi-asserted-by":"crossref","first-page":"699","DOI":"10.1145\/3640543.3645143","volume-title":"Proceedings of the 29th International Conference on Intelligent User Interfaces","author":"Wang Bryan","year":"2024","unstructured":"Bryan Wang, Yuliang Li, Zhaoyang Lv, Haijun Xia, Yan Xu, and Raj Sodhi. 2024. LAVE: LLM-powered agent assistance and language augmentation for video editing. In Proceedings of the 29th International Conference on Intelligent User Interfaces, 699\u2013714."},{"key":"e_1_3_2_323_2","first-page":"1","volume-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies","volume":"5","author":"Wang Hongli","year":"2021","unstructured":"Hongli Wang, Bin Guo, Jiaqi Liu, Sicong Liu, Yungang Wu, and Zhiwen Yu. 2021. Context-aware adaptive surgery: A fast and effective framework for adaptative model partition. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 5, 3 (2021), 1\u201322."},{"key":"e_1_3_2_324_2","doi-asserted-by":"crossref","first-page":"248","DOI":"10.1016\/j.neucom.2021.08.131","article-title":"Towards information-rich, logical dialogue systems with knowledge-enhanced neural models","volume":"465","author":"Wang Hao","year":"2021","unstructured":"Hao Wang, Bin Guo, Wei Wu, Sicong Liu, and Zhiwen Yu. 2021. Towards information-rich, logical dialogue systems with knowledge-enhanced neural models. Neurocomputing 465 (2021), 248\u2013264.","journal-title":"Neurocomputing"},{"key":"e_1_3_2_325_2","unstructured":"Hongru Wang Lingzhi Wang Yiming Du Liang Chen Jingyan Zhou Yufei Wang and Kam-Fai Wong. 2023. A survey of the evolution of language model-based dialogue systems. arXiv:2311.16789. Retrieved from https:\/\/arxiv.org\/abs\/2311.16789"},{"key":"e_1_3_2_326_2","first-page":"8455","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Hanqing","year":"2021","unstructured":"Hanqing Wang, Wenguan Wang, Wei Liang, Caiming Xiong, and Jianbing Shen. 2021. Structured scene memory for vision-language navigation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 8455\u20138464."},{"issue":"3","key":"e_1_3_2_327_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3209659","article-title":"Enabling live video analytics with a scalable and privacy-aware framework","volume":"14","author":"Wang Junjue","year":"2018","unstructured":"Junjue Wang, Brandon Amos, Anupam Das, Padmanabhan Pillai, Norman Sadeh, and Mahadev Satyanarayanan. 2018. Enabling live video analytics with a scalable and privacy-aware framework. ACM Transactions on Multimedia Computing, Communications, and Applications 14, 3s (2018), 1\u201324.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_2_328_2","volume-title":"Proceedings of the ICLR 2024 Workshop on Large Language Model (LLM) Agents","author":"Wang Junyang","year":"2024","unstructured":"Junyang Wang, Haiyang Xu, Jiabo Ye, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang. 2024. Mobile-Agent: Autonomous multi-modal mobile device agent with visual perception. In Proceedings of the ICLR 2024 Workshop on Large Language Model (LLM) Agents."},{"key":"e_1_3_2_329_2","unstructured":"Kaiye Wang Qiyue Yin Wei Wang Shu Wu and Liang Wang. 2016. A comprehensive survey on cross-modal retrieval. arXiv:1607.06215. Retrieved from https:\/\/arxiv.org\/abs\/1607.06215"},{"key":"e_1_3_2_330_2","first-page":"1","article-title":"A comprehensive survey of continual learning: Theory, method and application","volume":"01","author":"Wang Liyuan","year":"2024","unstructured":"Liyuan Wang, Xingxing Zhang, Hang Su, and Jun Zhu. 2024. A comprehensive survey of continual learning: Theory, method and application. IEEE Transactions on Pattern Analysis and Machine Intelligence 01 (2024), 1\u201320.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_331_2","doi-asserted-by":"crossref","first-page":"4747","DOI":"10.1109\/ICDE60146.2024.00361","volume-title":"Proceedings of the 2024 IEEE 40th International Conference on Data Engineering (ICDE)","author":"Wang Mengzhao","year":"2024","unstructured":"Mengzhao Wang, Xiangyu Ke, Xiaoliang Xu, Lu Chen, Yunjun Gao, Pinpin Huang, and Runkai Zhu. 2024. Must: An effective and scalable framework for multimodal search of target modality. In Proceedings of the 2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 4747\u20134759."},{"key":"e_1_3_2_332_2","unstructured":"Peng Wang Shuai Bai Sinan Tan Shijie Wang Zhihao Fan Jinze Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge et al. 2024. Qwen2-VL: Enhancing vision-language model\u2019s perception of the world at any resolution. arXiv:2409.12191. Retrieved from https:\/\/arxiv.org\/abs\/2409.12191"},{"key":"e_1_3_2_333_2","first-page":"1960","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Peng","year":"2019","unstructured":"Peng Wang, Qi Wu, Jiewei Cao, Chunhua Shen, Lianli Gao, and Anton van den Hengel. 2019. Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1960\u20131968."},{"key":"e_1_3_2_334_2","first-page":"19175","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Wenhui","year":"2023","unstructured":"Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. 2023. Image as a foreign language: Beit pretraining for vision and vision-language tasks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 19175\u201319186."},{"issue":"6","key":"e_1_3_2_335_2","doi-asserted-by":"crossref","first-page":"3239","DOI":"10.1109\/TPAMI.2021.3051099","article-title":"Salient object detection in the deep learning era: An in-depth survey","volume":"44","author":"Wang Wenguan","year":"2021","unstructured":"Wenguan Wang, Qiuxia Lai, Huazhu Fu, Jianbing Shen, Haibin Ling, and Ruigang Yang. 2021. Salient object detection in the deep learning era: An in-depth survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 6 (2021), 3239\u20133259.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_336_2","unstructured":"Weihan Wang Qingsong Lv Wenmeng Yu Wenyi Hong Ji Qi Yan Wang Junhui Ji Zhuoyi Yang Lei Zhao Xixuan Song et al. 2023. CogVLM: Visual expert for pretrained language models. arXiv:2311.03079. Retrieved from https:\/\/arxiv.org\/abs\/2311.03079"},{"issue":"4","key":"e_1_3_2_337_2","doi-asserted-by":"crossref","first-page":"447","DOI":"10.1007\/s11633-022-1410-8","article-title":"Large-scale multi-modal pre-trained models: A comprehensive survey","volume":"20","author":"Wang Xiao","year":"2023","unstructured":"Xiao Wang, Guangyao Chen, Guangwu Qian, Pengcheng Gao, Xiao-Yong Wei, Yaowei Wang, Yonghong Tian, and Wen Gao. . 2023. Large-scale multi-modal pre-trained models: A comprehensive survey. Machine Intelligence Research 20, 4 (2023), 447\u2013482.","journal-title":"Machine Intelligence Research"},{"key":"e_1_3_2_338_2","first-page":"26606","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Xinyu","year":"2024","unstructured":"Xinyu Wang, Bohan Zhuang, and Qi Wu. 2024. Modaverse: Efficiently transforming modalities with LLMs. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 26606\u201326616."},{"key":"e_1_3_2_339_2","first-page":"5036","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Wang Yuxuan","year":"2023","unstructured":"Yuxuan Wang, Zilong Zheng, Xueliang Zhao, Jinpeng Li, Yueqian Wang, and Dongyan Zhao. 2023. VSTAR: A video-grounded dialogue dataset for situated semantic understanding with scene and topic transitions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 5036\u20135048."},{"key":"e_1_3_2_340_2","first-page":"11293","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Zirui","year":"2019","unstructured":"Zirui Wang, Zihang Dai, Barnab\u00e1s P\u00f3czos, and Jaime Carbonell. 2019. Characterizing and avoiding negative transfer. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 11293\u201311302."},{"key":"e_1_3_2_341_2","article-title":"Emergent abilities of large language models","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022. Emergent abilities of large language models. Transactions on Machine Learning Research (2022).","journal-title":"Transactions on Machine Learning Research"},{"key":"e_1_3_2_342_2","first-page":"24824","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the Advances in Neural Information Processing Systems, 24824\u201324837."},{"key":"e_1_3_2_343_2","unstructured":"John Wieting Mohit Bansal Kevin Gimpel and Karen Livescu. 2015. Towards universal paraphrastic sentence embeddings. arXiv:1511.08198. Retrieved from https:\/\/arxiv.org\/abs\/1511.08198"},{"issue":"12","key":"e_1_3_2_344_2","doi-asserted-by":"crossref","first-page":"14257","DOI":"10.1007\/s10462-023-10489-1","article-title":"Dimensionality reduced training by pruning and freezing parts of a deep neural network: A survey","volume":"56","author":"Wimmer Paul","year":"2023","unstructured":"Paul Wimmer, Jens Mehnert, and Alexandru Paul Condurache. 2023. Dimensionality reduced training by pruning and freezing parts of a deep neural network: A survey. Artificial Intelligence Review 56, 12 (2023), 14257\u201314295.","journal-title":"Artificial Intelligence Review"},{"issue":"1","key":"e_1_3_2_345_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/0010-0285(72)90002-3","article-title":"Understanding natural language","volume":"3","author":"Winograd Terry","year":"1972","unstructured":"Terry Winograd. 1972. Understanding natural language. Cognitive Psychology 3, 1 (1972), 1\u2013191.","journal-title":"Cognitive Psychology"},{"issue":"4","key":"e_1_3_2_346_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3009906","article-title":"Computer vision and natural language processing: Recent approaches in multimedia and robotics","volume":"49","author":"Wiriyathammabhum Peratham","year":"2016","unstructured":"Peratham Wiriyathammabhum, Douglas Summers-Stay, Cornelia Ferm\u00fcller, and Yiannis Aloimonos. 2016. Computer vision and natural language processing: Recent approaches in multimedia and robotics. ACM Computing Surveys 49, 4 (2016), 1\u201344.","journal-title":"ACM Computing Surveys"},{"key":"e_1_3_2_347_2","series-title":"Datasets and Benchmarks Track (Round 2)","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems","author":"Wu Bo","year":"2021","unstructured":"Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B. Tenenbaum, and Chuang Gan. 2021. STAR: A benchmark for situated reasoning in real-world videos. In Proceedings of the Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)."},{"key":"e_1_3_2_348_2","unstructured":"Chenfei Wu Shengming Yin Weizhen Qi Xiaodong Wang Zecheng Tang and Nan Duan. 2023. Visual ChatGPT: Talking drawing and editing with visual foundation models. arXiv:2303.04671. Retrieved from https:\/\/arxiv.org\/abs\/2303.04671"},{"key":"e_1_3_2_349_2","first-page":"2247","volume-title":"Proceedings of the 2023 IEEE International Conference on Big Data (BigData)","author":"Wu Jiayang","year":"2023","unstructured":"Jiayang Wu, Wensheng Gan, Zefeng Chen, Shicheng Wan, and S. Yu Philip. 2023. Multimodal large language models: A survey. In Proceedings of the 2023 IEEE International Conference on Big Data (BigData). IEEE, 2247\u20132256."},{"key":"e_1_3_2_350_2","first-page":"68","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Wu Kan","year":"2022","unstructured":"Kan Wu, Jinnian Zhang, Houwen Peng, Mengchen Liu, Bin Xiao, Jianlong Fu, and Lu Yuan. 2022. Tinyvit: Fast pretraining distillation for small vision transformers. In Proceedings of the European Conference on Computer Vision. Springer, 68\u201385."},{"key":"e_1_3_2_351_2","first-page":"4840","volume-title":"Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Wu Lingfei","year":"2022","unstructured":"Lingfei Wu, Peng Cui, Jian Pei, Liang Zhao, and Xiaojie Guo. 2022. Graph neural networks: Foundation, frontiers and applications. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 4840\u20134841."},{"key":"e_1_3_2_352_2","first-page":"6106","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wu Qi","year":"2018","unstructured":"Qi Wu, Peng Wang, Chunhua Shen, Ian Reid, and Anton Van Den Hengel. 2018. Are you talking to me? Reasoned visual dialog generation through adversarial learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6106\u20136115."},{"key":"e_1_3_2_353_2","first-page":"53366","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Wu Shengqiong","year":"2024","unstructured":"Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. 2024. NExT-GPT: Any-to-any multimodal LLM. In Proceedings of the 41st International Conference on Machine Learning, 53366\u201353397."},{"key":"e_1_3_2_354_2","unstructured":"Tongtong Wu Linhao Luo Yuan-Fang Li Shirui Pan Thuy-Trang Vu and Gholamreza Haffari. 2024. Continual learning for large language models: A survey. arXiv:2402.01364. Retrieved from https:\/\/arxiv.org\/abs\/2402.01364"},{"issue":"1","key":"e_1_3_2_355_2","first-page":"4","article-title":"A comprehensive survey on graph neural networks","volume":"32","author":"Wu Zonghan","year":"2020","unstructured":"Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and S. Yu Philip. 2020. A comprehensive survey on graph neural networks. IEEE Transactions on Neural Networks and Learning Systems 32, 1 (2020), 4\u201324.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_2_356_2","first-page":"75392","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Xiang Jiannan","year":"2024","unstructured":"Jiannan Xiang, Tianhua Tao, Yi Gu, Tianmin Shu, Zirui Wang, Zichao Yang, and Zhiting Hu. 2024. Language models meet world models: Embodied experiences enhance language models. In Proceedings of the Advances in Neural Information Processing Systems, 75392\u201375412."},{"key":"e_1_3_2_357_2","first-page":"38087","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Xiao Guangxuan","year":"2023","unstructured":"Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. SmoothQuant: Accurate and efficient post-training quantization for large language models. In Proceedings of the International Conference on Machine Learning. PMLR, 38087\u201338099."},{"key":"e_1_3_2_358_2","first-page":"1","article-title":"When search engine services meet large language models: Visions and challenges","volume":"01","author":"Xiong Haoyi","year":"2024","unstructured":"Haoyi Xiong, Jiang Bian, Yuchen Li, Xuhong Li, Mengnan Du, Shuaiqiang Wang, Dawei Yin, and Sumi Helal. 2024. When search engine services meet large language models: Visions and challenges. IEEE Transactions on Services Computing 01 (2024), 1\u201323.","journal-title":"IEEE Transactions on Services Computing"},{"key":"e_1_3_2_359_2","first-page":"10566","volume-title":"In Proceedings of the AAAI Conference on Artificial Intelligence","volume":"37","author":"Xu Canwen","year":"2023","unstructured":"Canwen Xu and Julian McAuley. 2023. A survey on model compression and acceleration for pretrained language models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, 10566\u201310575."},{"key":"e_1_3_2_360_2","doi-asserted-by":"crossref","first-page":"6787","DOI":"10.18653\/v1\/2021.emnlp-main.544","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Xu Hu","year":"2021","unstructured":"Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021. VideoCLIP: Contrastive pre-training for zero-shot video-text understanding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 6787\u20136800."},{"issue":"10","key":"e_1_3_2_361_2","doi-asserted-by":"crossref","first-page":"12113","DOI":"10.1109\/TPAMI.2023.3275156","article-title":"Multimodal learning with transformers: A survey","volume":"45","author":"Xu Peng","year":"2023","unstructured":"Peng Xu, Xiatian Zhu, and David A. Clifton. 2023. Multimodal learning with transformers: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 10 (2023), 12113\u201312132.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_362_2","unstructured":"Zhiyuan Xu Kun Wu Junjie Wen Jinming Li Ning Liu Zhengping Che and Jian Tang. 2024. A survey on robotics with foundation models: Toward embodied AI. arXiv:2402.02385. Retrieved from https:\/\/arxiv.org\/abs\/2402.02385"},{"issue":"1","key":"e_1_3_2_363_2","first-page":"43","article-title":"Task-adaptive attention for image captioning","volume":"32","author":"Yan Chenggang","year":"2021","unstructured":"Chenggang Yan, Yiming Hao, Liang Li, Jian Yin, Anan Liu, Zhendong Mao, Zhenyu Chen, and Xingyu Gao. 2021. Task-adaptive attention for image captioning. IEEE Transactions on Circuits and Systems for Video technology 32, 1 (2021), 43\u201351.","journal-title":"IEEE Transactions on Circuits and Systems for Video technology"},{"key":"e_1_3_2_364_2","first-page":"10714","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yang Antoine","year":"2023","unstructured":"Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont-Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid. 2023. Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 10714\u201310726."},{"key":"e_1_3_2_365_2","unstructured":"An Yang Baosong Yang Binyuan Hui Bo Zheng Bowen Yu Chang Zhou Chengpeng Li Chengyuan Li Dayiheng Liu Fei Huang et al. 2024. Qwen2 technical report. arXiv:2407.10671. Retrieved from https:\/\/arxiv.org\/abs\/2407.10671"},{"key":"e_1_3_2_366_2","doi-asserted-by":"crossref","first-page":"3480","DOI":"10.1145\/3503161.3548291","volume-title":"Proceedings of the 30th ACM International Conference on Multimedia","author":"Yang Pinci","year":"2022","unstructured":"Pinci Yang, Xin Wang, Xuguang Duan, Hong Chen, Runze Hou, Cong Jin, and Wenwu Zhu. 2022. Avqa: A dataset for audio-visual question answering on videos. In Proceedings of the 30th ACM International Conference on Multimedia, 3480\u20133491."},{"key":"e_1_3_2_367_2","first-page":"2561","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Yang Tianhao","year":"2019","unstructured":"Tianhao Yang, Zheng-Jun Zha, and Hanwang Zhang. 2019. Making history matter: History-advantage sequence training for visual dialog. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2561\u20132569."},{"key":"e_1_3_2_368_2","first-page":"67","article-title":"A cognitive system for understanding human manipulation actions","volume":"3","author":"Yang Yezhou","year":"2014","unstructured":"Yezhou Yang, Anupam Guha, C. Fermuller, and Yiannis Aloimonos. 2014. A cognitive system for understanding human manipulation actions. Advances in Cognitive Sysytems 3 (2014), 67\u201386.","journal-title":"Advances in Cognitive Sysytems"},{"key":"e_1_3_2_369_2","first-page":"55976","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Yang Zongxin","year":"2024","unstructured":"Zongxin Yang, Guikun Chen, Xiaodi Li, Wenguan Wang, and Yi Yang. 2024. DoraemonGPT: Toward understanding dynamic scenes with large language models (exemplified as a video agent). In Proceedings of the 41st International Conference on Machine Learning, 55976\u201355997."},{"key":"e_1_3_2_370_2","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1016\/j.patrec.2018.05.018","article-title":"A review of convolutional-neural-network-based action recognition","volume":"118","author":"Yao Guangle","year":"2019","unstructured":"Guangle Yao, Tao Lei, and Jiandan Zhong. 2019. A review of convolutional-neural-network-based action recognition. Pattern Recognition Letters 118 (2019), 14\u201322.","journal-title":"Pattern Recognition Letters"},{"issue":"2","key":"e_1_3_2_371_2","doi-asserted-by":"crossref","first-page":"100211","DOI":"10.1016\/j.hcc.2024.100211","article-title":"A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly","volume":"4","author":"Yao Yifan","year":"2024","unstructured":"Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang. 2024. A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly. High-Confidence Computing 4, 2 (2024), 100211.","journal-title":"High-Confidence Computing"},{"key":"e_1_3_2_372_2","unstructured":"Yuan Yao Tianyu Yu Ao Zhang Chongyi Wang Junbo Cui Hongji Zhu Tianchi Cai Haoyu Li Weilin Zhao Zhihui He et al. 2024. MiniCPM-V: A gpt-4v level MLLM on your phone. arXiv:2408.01800. Retrieved from https:\/\/arxiv.org\/abs\/2408.01800"},{"key":"e_1_3_2_373_2","first-page":"27168","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Yao Zhewei","year":"2022","unstructured":"Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. 2022. ZeroQuant: Efficient and affordable post-training quantization for large-scale transformers. In Proceedings of the Advances in Neural Information Processing Systems, 27168\u201327183."},{"key":"e_1_3_2_374_2","unstructured":"Qinghao Ye Haiyang Xu Guohai Xu Jiabo Ye Ming Yan Yiyang Zhou Junyang Wang Anwen Hu Pengcheng Shi Yaya Shi et al. 2023. mPLUG-Owl: Modularization empowers large language models with multimodality. arXiv:2304.14178. Retrieved from https:\/\/arxiv.org\/abs\/2304.14178"},{"key":"e_1_3_2_375_2","first-page":"13040","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ye Qinghao","year":"2024","unstructured":"Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Anwen Hu, Haowei Liu, Qi Qian, Ji Zhang, and Fei Huang. 2024. mPLUG-Owl2: Revolutionizing multi-modal large language model with modality collaboration. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 13040\u201313051."},{"key":"e_1_3_2_376_2","unstructured":"Shukang Yin Chaoyou Fu Sirui Zhao Ke Li Xing Sun Tong Xu and Enhong Chen. 2023. A survey on multimodal large language models. arXiv:2306.13549. Retrieved from https:\/\/arxiv.org\/abs\/2306.13549"},{"issue":"6","key":"e_1_3_2_377_2","first-page":"1","article-title":"A comprehensive survey of privacy-preserving federated learning: A taxonomy, review, and future directions","volume":"54","author":"Yin Xuefei","year":"2021","unstructured":"Xuefei Yin, Yanming Zhu, and Jiankun Hu. 2021. A comprehensive survey of privacy-preserving federated learning: A taxonomy, review, and future directions. ACM Computing Surveys 54, 6 (2021), 1\u201336.","journal-title":"ACM Computing Surveys"},{"key":"e_1_3_2_378_2","first-page":"57116","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Ying Kaining","year":"2024","unstructured":"Kaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li, Han Lin, Yue Yang, Hao Zhang, Wenbo Zhang, Yuqi Lin, Shuo Liu, et al. 2024. MMT-Bench: A comprehensive multimodal benchmark for evaluating large vision-language models towards multitask AGI. In Proceedings of the 41st International Conference on Machine Learning, 57116\u201357198."},{"key":"e_1_3_2_379_2","doi-asserted-by":"crossref","first-page":"4182","DOI":"10.18653\/v1\/2022.emnlp-main.280","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Yoon Sunjae","year":"2022","unstructured":"Sunjae Yoon, Eunseop Yoon, Hee Suk Yoon, Junyeong Kim, and Chang Yoo. 2022. Information-theoretic text hallucination reduction for video-grounded dialogue. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 4182\u20134193."},{"key":"e_1_3_2_380_2","first-page":"1","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"You Haoxuan","year":"2024","unstructured":"Haoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du, Bowen Zhang, Zirui Wang, Liangliang Cao, Shih-Fu Chang, and Yinfei Yang. 2024. Ferret: Refer and ground anything anywhere at any granularity. In Proceedings of the 12th International Conference on Learning Representations, 1\u201328."},{"key":"e_1_3_2_381_2","unstructured":"Alex Young Bei Chen Chao Li Chengen Huang Ge Zhang Guanwei Zhang Heng Li Jiangcheng Zhu Jianqun Chen Jing Chang et al. 2024. Yi: Open foundation models by 01.AI. arXiv:2403.04652. Retrieved from https:\/\/arxiv.org\/abs\/2403.04652"},{"key":"e_1_3_2_382_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Yu Da","year":"2022","unstructured":"Da Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi, Huseyin A. Inan, Gautam Kamath, Janardhan Kulkarni, Yin Tat Lee, Andre Manoel, Lukas Wutschitz, et al. 2022. Differentially private fine-tuning of language models. In Proceedings of the International Conference on Learning Representations, 1\u201319."},{"key":"e_1_3_2_383_2","first-page":"1089","volume-title":"Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security","author":"Yu Hyunwoo","year":"2018","unstructured":"Hyunwoo Yu, Jaemin Lim, Kiyeon Kim, and Suk-Bok Lee. 2018. Pinto: Enabling video privacy for commodity IoT cameras. In Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security, 1089\u20131101."},{"key":"e_1_3_2_384_2","first-page":"220","article-title":"Learning dual encoding model for adaptive visual understanding in visual dialogue","volume":"30","author":"Yu Jing","year":"2020","unstructured":"Jing Yu, Xiaoze Jiang, Zengchang Qin, Weifeng Zhang, Yue Hu, and Qi Wu. 2020. Learning dual encoding model for adaptive visual understanding in visual dialogue. IEEE Transactions on Image Processing 30 (2020), 220\u2013233.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_2_385_2","first-page":"69","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Yu Licheng","year":"2016","unstructured":"Licheng Yu, Patrick Poirson, Shan Yang, Alexander C. Berg, and Tamara L. Berg. 2016. Modeling context in referring expressions. In Proceedings of the European Conference on Computer Vision. Springer, 69\u201385."},{"key":"e_1_3_2_386_2","first-page":"57730","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Yu Weihao","year":"2024","unstructured":"Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. 2024. MM-Vet: Evaluating large multimodal models for integrated capabilities. In Proceedings of the 41st International Conference on Machine Learning, 57730\u201357754."},{"key":"e_1_3_2_387_2","unstructured":"Weihao Yu Zhengyuan Yang Linfeng Ren Linjie Li Jianfeng Wang Kevin Lin Chung-Ching Lin Zicheng Liu Lijuan Wang and Xinchao Wang. 2024. MM-Vet v2: A challenging benchmark to evaluate large multimodal models for integrated capabilities. arXiv:2408.00765. Retrieved from https:\/\/arxiv.org\/abs\/2408.00765"},{"key":"e_1_3_2_388_2","doi-asserted-by":"crossref","first-page":"108540","DOI":"10.1016\/j.patcog.2022.108540","article-title":"VD-PCR: Improving visual dialog with pronoun coreference resolution","volume":"125","author":"Yu Xintong","year":"2022","unstructured":"Xintong Yu, Hongming Zhang, Ruixin Hong, Yangqiu Song, and Changshui Zhang. 2022. VD-PCR: Improving visual dialog with pronoun coreference resolution. Pattern Recognition 125 (2022), 108540.","journal-title":"Pattern Recognition"},{"key":"e_1_3_2_389_2","volume-title":"Proceedings of the 2nd Workshop on Advancing Neural Network Training: Computational Efficiency, Scalability, and Resource Optimization (WANT@ ICML \u201924)","author":"Yuan Zhengqing","year":"2024","unstructured":"Zhengqing Yuan, Zhaoxu Li, Weiran Huang, Yanfang Ye, and Lichao Sun. 2024. TinyGPT-V: Efficient multimodal large language model via small backbones. In Proceedings of the 2nd Workshop on Advancing Neural Network Training: Computational Efficiency, Scalability, and Resource Optimization (WANT@ ICML \u201924)."},{"key":"e_1_3_2_390_2","first-page":"7443","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Zan Daoguang","year":"2023","unstructured":"Daoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Wang Yongji, and Jian-Guang Lou. 2023. Large language models meet NL2Code: A survey. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 7443\u20137464."},{"key":"e_1_3_2_391_2","first-page":"1","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Zeng Aohan","year":"2023","unstructured":"Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. 2023. GLM-130B: An open bilingual pre-trained model. In Proceedings of the 11th International Conference on Learning Representations, 1\u201319."},{"key":"e_1_3_2_392_2","first-page":"11975","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhai Xiaohua","year":"2023","unstructured":"Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. Sigmoid loss for language image pre-training. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 11975\u201311986."},{"key":"e_1_3_2_393_2","volume-title":"Proceedings of the NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following","author":"Zhai Yuexiang","year":"2023","unstructured":"Yuexiang Zhai, Shengbang Tong, Xiao Li, Mu Cai, Qing Qu, Yong Jae Lee, and Yi Ma. 2023. Investigating the catastrophic forgetting in multimodal large language models. In Proceedings of the NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following."},{"key":"e_1_3_2_394_2","unstructured":"Jun Zhan Junqi Dai Jiasheng Ye Yunhua Zhou Dong Zhang Zhigeng Liu Xin Zhang Ruibin Yuan Ge Zhang Linyang Li et al. 2024. AnyGPT: Unified multimodal LLM with discrete sequence modeling. arXiv:2402.12226. Retrieved from https:\/\/arxiv.org\/abs\/2402.12226"},{"key":"e_1_3_2_395_2","doi-asserted-by":"crossref","first-page":"793","DOI":"10.1145\/3292500.3330961","volume-title":"Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining","author":"Zhang Chuxu","year":"2019","unstructured":"Chuxu Zhang, Dongjin Song, Chao Huang, Ananthram Swami, and Nitesh V. Chawla. 2019. Heterogeneous graph neural network. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 793\u2013803."},{"key":"e_1_3_2_396_2","first-page":"12401","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Zhang Duzhen","year":"2024","unstructured":"Duzhen Zhang, Yanhan Yu, Chenxing Li, Jiahua Dong, Dan Su, Chenhui Chu, and Dong Yu. 2024. MM-LLMs: Recent advances in multimodal large language models. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, 12401\u201312430."},{"issue":"6","key":"e_1_3_2_397_2","doi-asserted-by":"crossref","first-page":"4347","DOI":"10.1007\/s10462-021-10123-y","article-title":"Visual privacy attacks and defenses in deep learning: A survey","volume":"55","author":"Zhang Guangsheng","year":"2022","unstructured":"Guangsheng Zhang, Bo Liu, Tianqing Zhu, Andi Zhou, and Wanlei Zhou. 2022. Visual privacy attacks and defenses in deep learning: A survey. Artificial Intelligence Review 55, 6 (2022), 4347\u20134401.","journal-title":"Artificial Intelligence Review"},{"key":"e_1_3_2_398_2","first-page":"1025","volume-title":"Proceedings of the 28th International Joint Conference on Artificial Intelligence","author":"Zhang Heming","year":"2019","unstructured":"Heming Zhang, Shalini Ghosh, Larry Heck, Stephen Walsh, Junting Zhang, Jie Zhang, and C.-C. Jay Kuo. 2019. Generative visual dialogue system via weighted likelihood estimation. In Proceedings of the 28th International Joint Conference on Artificial Intelligence, 1025\u20131031."},{"key":"e_1_3_2_399_2","doi-asserted-by":"crossref","first-page":"543","DOI":"10.18653\/v1\/2023.emnlp-demo.49","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Zhang Hang","year":"2023","unstructured":"Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An instruction-tuned audio-visual language model for video understanding. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 543\u2013553."},{"key":"e_1_3_2_400_2","first-page":"1","article-title":"Vision-language models for vision tasks: A survey","volume":"01","author":"Zhang Jingyi","year":"2024","unstructured":"Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. 2024. Vision-language models for vision tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 01 (2024), 1\u201320.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_401_2","first-page":"766","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Zhang Ke","year":"2016","unstructured":"Ke Zhang, Wei-Lun Chao, Fei Sha, and Kristen Grauman. 2016. Video summarization with long short-term memory. In Proceedings of the European Conference on Computer Vision. Springer, 766\u2013782."},{"key":"e_1_3_2_402_2","unstructured":"Pan Zhang Xiaoyi Dong Yuhang Zang Yuhang Cao Rui Qian Lin Chen Qipeng Guo Haodong Duan Bin Wang Linke Ouyang et al. 2024. Internlm-xcomposer-2.5: A versatile large vision language model supporting long-contextual input and output. arXiv:2407.03320. Retrieved from https:\/\/arxiv.org\/abs\/2407.03320"},{"key":"e_1_3_2_403_2","unstructured":"Peiyuan Zhang Kaichen Zhang Bo Li Guangtao Zeng Jingkang Yang Yuanhan Zhang Ziyue Wang Haoran Tan Chunyuan Li and Ziwei Liu. 2024. Long context transfer from language to vision. arXiv:2406.16852. Retrieved from https:\/\/arxiv.org\/abs\/2406.16852"},{"key":"e_1_3_2_404_2","first-page":"26809","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Zhang Qingru","year":"2022","unstructured":"Qingru Zhang, Simiao Zuo, Chen Liang, Alexander Bukharin, Pengcheng He, Weizhu Chen, and Tuo Zhao. 2022. Platon: Pruning large transformer models with upper confidence bound of weight importance. In Proceedings of the International Conference on Machine Learning. PMLR, 26809\u201326823."},{"key":"e_1_3_2_405_2","first-page":"1","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Zhang Renrui","year":"2024","unstructured":"Renrui Zhang, Jiaming Han, Chris Liu, Aojun Zhou, Pan Lu, Yu Qiao, Hongsheng Li, and Peng Gao. 2024. LLaMA-adapter: Efficient fine-tuning of large language models with zero-initialized attention. In Proceedings of the 12th International Conference on Learning Representations, 1\u201330."},{"key":"e_1_3_2_406_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.inffus.2021.02.022","article-title":"Multimodal feature-wise co-attention method for visual question answering","volume":"73","author":"Zhang Sheng","year":"2021","unstructured":"Sheng Zhang, Min Chen, Jincai Chen, Fuhao Zou, Yuan-Fang Li, and Ping Lu. 2021. Multimodal feature-wise co-attention method for visual question answering. Information Fusion 73 (2021), 1\u201310.","journal-title":"Information Fusion"},{"key":"e_1_3_2_407_2","unstructured":"Shengyu Zhang Linfeng Dong Xiaoya Li Sen Zhang Xiaofei Sun Shuhe Wang Jiwei Li Runyi Hu Tianwei Zhang Fei Wu et al. 2023. Instruction tuning for large language models: A survey. arXiv:2308.10792. Retrieved from https:\/\/arxiv.org\/abs\/2308.10792"},{"key":"e_1_3_2_408_2","first-page":"4600","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Shunyu","year":"2022","unstructured":"Shunyu Zhang, Xiaoze Jiang, Zequn Yang, Tao Wan, and Zengchang Qin. 2022. Reasoning with multi-structure commonsense knowledge in visual dialog. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 4600\u20134609."},{"key":"e_1_3_2_409_2","unstructured":"Susan Zhang Stephen Roller Naman Goyal Mikel Artetxe Moya Chen Shuohui Chen Christopher Dewan Mona Diab Xian Li Xi Victoria Lin et al. 2022. OPT: Open pre-trained transformer language models. arXiv:2205.01068. Retrieved from https:\/\/arxiv.org\/abs\/2205.01068"},{"key":"e_1_3_2_410_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Zhang Tianyi","year":"2019","unstructured":"Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2019. BERTScore: Evaluating text generation with BERT. In Proceedings of the International Conference on Learning Representations, 1\u201314."},{"key":"e_1_3_2_411_2","first-page":"6848","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhang Xiangyu","year":"2018","unstructured":"Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. 2018. ShuffleNet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6848\u20136856."},{"key":"e_1_3_2_412_2","first-page":"1724","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Youcai","year":"2024","unstructured":"Youcai Zhang, Xinyu Huang, Jinyu Ma, Zhaoyang Li, Zhaochuan Luo, Yanchun Xie, Yuzhuo Qin, Tong Luo, Yaqian Li, Shilong Liu, et al. 2024. Recognize anything: A strong image tagging model. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1724\u20131732."},{"key":"e_1_3_2_413_2","first-page":"1","volume-title":"Proceedings of the ACM Turing Celebration Conference-China","author":"Zhang Yujia","year":"2019","unstructured":"Yujia Zhang, Michael Kampffmeyer, Xiaoguang Zhao, and Min Tan. 2019. DTR-GAN: Dilated temporal relational adversarial network for video summarization. In Proceedings of the ACM Turing Celebration Conference-China, 1\u20136."},{"key":"e_1_3_2_414_2","unstructured":"Yuanhan Zhang Bo Li Haotian Liu Yong Jae Lee Liangke Gui Di Fu Jiashi Feng Ziwei Liu and Chunyuan Li. 2024. LLaVA-NeXT: A Strong Zero-Shot Video Understanding Model. Retrieved from https:\/\/llava-vl.github.io\/blog\/2024-04-30-llava-next-video\/"},{"key":"e_1_3_2_415_2","unstructured":"Yue Zhang Yafu Li Leyang Cui Deng Cai Lemao Liu Tingchen Fu Xinting Huang Enbo Zhao Yu Zhang Yulong Chen et al. 2023. Siren\u2019s song in the AI ocean: A survey on hallucination in large language models. arXiv:2309.01219. Retrieved from https:\/\/arxiv.org\/abs\/2309.01219"},{"issue":"4","key":"e_1_3_2_416_2","doi-asserted-by":"crossref","first-page":"2775","DOI":"10.1109\/TCSVT.2023.3312325","article-title":"VSS-Net: Visual semantic self-mining network for video summarization","volume":"34","author":"Zhang Yunzuo","year":"2023","unstructured":"Yunzuo Zhang, Yameng Liu, Weili Kang, and Ran Tao. 2023. VSS-Net: Visual semantic self-mining network for video summarization. IEEE Transactions on Circuits and Systems for Video Technology 34, 4 (2024), 2775\u20132788.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_2_417_2","unstructured":"Yuanhan Zhang Jinming Wu Wei Li Bo Li Zejun Ma Ziwei Liu and Chunyuan Li. 2024. Video instruction tuning with synthetic data. arXiv:2410.02713. Retrieved from https:\/\/arxiv.org\/abs\/2410.02713"},{"key":"e_1_3_2_418_2","unstructured":"Yi-Fan Zhang Huanyu Zhang Haochen Tian Chaoyou Fu Shuangqing Zhang Junfei Wu Feng Li Kun Wang Qingsong Wen Zhang et al. 2024. MME-RealWorld: Could your multimodal LLM challenge high-resolution real-world scenarios that are difficult for humans? arXiv:2408.13257. Retrieved from https:\/\/arxiv.org\/abs\/2408.13257"},{"key":"e_1_3_2_419_2","first-page":"1","volume-title":"Proceedings of the 2021 IEEE International Conference on Multimedia and Expo (ICME)","author":"Zhao Lei","year":"2021","unstructured":"Lei Zhao, Lianli Gao, Yuyu Guo, Jingkuan Song, and Hengtao Shen. 2021. SKANet: Structured knowledge-aware network for visual dialog. In Proceedings of the 2021 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1\u20136."},{"key":"e_1_3_2_420_2","volume-title":"In Proceedings of the LaCATODA@ IJCAI","author":"Zhao Rui","year":"2018","unstructured":"Rui Zhao and Volker Tresp. 2018. Improving goal-oriented visual dialog agents via advanced recurrent nets with tempered policy gradient. In Proceedings of the LaCATODA@ IJCAI."},{"key":"e_1_3_2_421_2","first-page":"1","volume-title":"Proceedings of the 19th Annual SIGdial Meeting on Discourse and Dialogue","author":"Zhao Tiancheng","year":"2018","unstructured":"Tiancheng Zhao and Maxine Eskenazi. 2018. Zero-shot dialog generation with cross-domain latent actions. In Proceedings of the 19th Annual SIGdial Meeting on Discourse and Dialogue, 1\u201310."},{"issue":"4","key":"e_1_3_2_422_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3637870","article-title":"Dense text retrieval based on pretrained language models: A survey","volume":"42","author":"Zhao Wayne Xin","year":"2024","unstructured":"Wayne Xin Zhao, Jing Liu, Ruiyang Ren, and Ji-Rong Wen. 2024. Dense text retrieval based on pretrained language models: A survey. ACM Transactions on Information Systems 42, 4 (2024), 1\u201360.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_2_423_2","unstructured":"Wayne Xin Zhao Kun Zhou Junyi Li Tianyi Tang Xiaolei Wang Yupeng Hou Yingqian Min Beichen Zhang Junjie Zhang Zican Dong et al. 2023. A survey of large language models. arXiv:2303.18223. Retrieved from https:\/\/arxiv.org\/abs\/2303.18223"},{"key":"e_1_3_2_424_2","doi-asserted-by":"crossref","first-page":"5988","DOI":"10.18653\/v1\/2022.findings-emnlp.442","volume-title":"Findings of the Association for Computational Linguistics (EMNLP \u201922)","author":"Zhao Xueliang","year":"2022","unstructured":"Xueliang Zhao, Yuxuan Wang, Chongyang Tao, Chenshuo Wang, and Dongyan Zhao. 2022. Collaborative reasoning on multi-modal semantic graphs for video-grounded dialogue generation. In Findings of the Association for Computational Linguistics (EMNLP \u201922), 5988\u20135998."},{"key":"e_1_3_2_425_2","first-page":"196","volume-title":"Proceedings of Machine Learning and Systems","volume":"6","author":"Zhao Yilong","year":"2024","unstructured":"Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. 2024. Atom: Low-bit quantization for efficient and accurate LLM serving. Proceedings of Machine Learning and Systems 6 (2024), 196\u2013209."},{"key":"e_1_3_2_426_2","first-page":"6669","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zheng Zilong","year":"2019","unstructured":"Zilong Zheng, Wenguan Wang, Siyuan Qi, and Song-Chun Zhu. 2019. Reasoning visual dialogs with structural and partial observations. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 6669\u20136678."},{"key":"e_1_3_2_427_2","first-page":"7582","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Zhou Kaiyang","year":"2018","unstructured":"Kaiyang Zhou, Yu Qiao, and Tao Xiang. 2018. Deep reinforcement learning for unsupervised video summarization with diversity-representativeness reward. In Proceedings of the AAAI Conference on Artificial Intelligence, 7582\u20137589."},{"key":"e_1_3_2_428_2","doi-asserted-by":"crossref","first-page":"143","DOI":"10.18653\/v1\/D19-1014","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Zhou Mingyang","year":"2019","unstructured":"Mingyang Zhou, Josh Arnold, and Zhou Yu. 2019. Building task-oriented visual dialog systems through alternative optimization between dialog policy and language generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 143\u2013153."},{"key":"e_1_3_2_429_2","first-page":"1738","volume-title":"Proceedings of the IEEE","volume":"107","author":"Zhou Zhi","year":"2019","unstructured":"Zhi Zhou, Xu Chen, En Li, Liekang Zeng, Ke Luo, and Junshan Zhang. 2019. Edge intelligence: Paving the last mile of artificial intelligence with edge computing. Proceedings of the IEEE 107, 8 (2019), 1738\u20131762."},{"key":"e_1_3_2_430_2","first-page":"1","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Zhu Deyao","year":"2024","unstructured":"Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2024. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. In Proceedings of the 12th International Conference on Learning Representations, 1\u201317."}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715098","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3715098","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:56:53Z","timestamp":1750298213000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715098"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,14]]},"references-count":429,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,5,31]]}},"alternative-id":["10.1145\/3715098"],"URL":"https:\/\/doi.org\/10.1145\/3715098","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,14]]},"assertion":[{"value":"2023-11-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-01-06","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}