{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,6]],"date-time":"2026-02-06T04:45:46Z","timestamp":1770353146624,"version":"3.49.0"},"reference-count":121,"publisher":"Association for Computing Machinery (ACM)","issue":"8","license":[{"start":{"date-parts":[[2024,6,13]],"date-time":"2024-06-13T00:00:00Z","timestamp":1718236800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Program of China","award":["No. 2021ZD0111000 and No. 2021ZD0111004"],"award-info":[{"award-number":["No. 2021ZD0111000 and No. 2021ZD0111004"]}]},{"DOI":"10.13039\/501100003399","name":"Science and Technology Commission of Shanghai Municipality","doi-asserted-by":"crossref","award":["No. 21511100101, No. 22511105901, No. 19511120200, No. 21511100100, and No. 18DZ2270800"],"award-info":[{"award-number":["No. 21511100101, No. 22511105901, No. 19511120200, No. 21511100100, and No. 18DZ2270800"]}],"id":[{"id":"10.13039\/501100003399","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Innovation 2030 Major S&T Project of China","award":["No. 2020AAA0104200 and No. 2020AAA0104205"],"award-info":[{"award-number":["No. 2020AAA0104200 and No. 2020AAA0104205"]}]},{"name":"Open Fund of PDL, and Alibaba Group through Alibaba Research Intern Program"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,8,31]]},"abstract":"<jats:p>Referring expression comprehension aims to align natural language queries with visual scenes, which requires establishing fine-grained correspondence between vision and language. This has important applications in multi-modal reasoning systems. Existing methods typically use text-agnostic visual backbones to extract features independently without considering the specific text input. However, we argue that the extracted visual features can be inconsistent with the referring expression, which hurts multi-modal understanding. To address this, we first propose Query-modulated Refinement Network (QRNet) that leverages language guidance to guide visual feature extraction. However, it only focuses on the grounding task that can only provide coarse-grained annotations in the form of bounding box coordinates. The guidance for the visual backbone is indirect, and the inconsistent issue still exists. To this end, we further propose UniQRNet, a multi-task framework over the QRNet to learn referring expression grounding and segmentation jointly. The framework introduces a multi-task head that leverages fine-grained pixel-level supervision from the segmentation task to directly guide the intermediate layers of QRNet to learn text-consistent visual features. Besides, UniQRNet also includes a loss balance strategy that allows two types of supervision signals to cooperate and optimize the model together. We conduct the most comprehensive comparison experiment covering four major datasets, ten evaluation set and three evaluation metrics used in previous work. UniQRNet outperforms previous state-of-the-art methods by a large margin on both referring comprehensive grounding (1.8%~5.09%) and segmentation tasks (0.57%~5.56%). Ablation and analysis reveal that UniQRNet can improve the consistency of visual features with text input and can bring significant performance improvement.<\/jats:p>","DOI":"10.1145\/3660638","type":"journal-article","created":{"date-parts":[[2024,4,25]],"date-time":"2024-04-25T11:12:27Z","timestamp":1714043547000},"page":"1-28","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["UniQRNet: Unifying Referring Expression Grounding and Segmentation with QRNet"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-5451-8984","authenticated-orcid":false,"given":"Jiabo","family":"Ye","sequence":"first","affiliation":[{"name":"East China Normal University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8566-7012","authenticated-orcid":false,"given":"Junfeng","family":"Tian","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4959-8878","authenticated-orcid":false,"given":"Ming","family":"Yan","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9442-5912","authenticated-orcid":false,"given":"Haiyang","family":"Xu","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7977-5540","authenticated-orcid":false,"given":"Qinghao","family":"Ye","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0465-6712","authenticated-orcid":false,"given":"Yaya","family":"Shi","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, University of Science and Technology of China, Hefei, China and NLPR, Chinese Academy of Sciences Institute of Automation, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5453-9755","authenticated-orcid":false,"given":"Xiaoshan","family":"Yang","sequence":"additional","affiliation":[{"name":"NLPR, CASIA, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3363-570X","authenticated-orcid":false,"given":"Xuwu","family":"Wang","sequence":"additional","affiliation":[{"name":"Fudan University - Handan Campus, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3835-7975","authenticated-orcid":false,"given":"Ji","family":"Zhang","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4723-5486","authenticated-orcid":false,"given":"Liang","family":"He","sequence":"additional","affiliation":[{"name":"East China Normal University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-4110-8989","authenticated-orcid":false,"given":"Xin","family":"Lin","sequence":"additional","affiliation":[{"name":"East China Normal University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,6,13]]},"reference":[{"key":"e_1_3_3_2_2","article-title":"XCiT: Cross-covariance image transformers","author":"Ali Alaaeldin","year":"2021","unstructured":"Alaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski, Matthijs Douze, Armand Joulin, Ivan Laptev, Natalia Neverova, Gabriel Synnaeve, Jakob Verbeek et\u00a0al. 2021. XCiT: Cross-covariance image transformers. Proceedings of the NeurIPS.","journal-title":"Proceedings of the NeurIPS"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00387"},{"key":"e_1_3_3_5_2","unstructured":"John Arevalo Thamar Solorio Manuel Montes-y G\u00f3mez and Fabio A. Gonz\u00e1lez. 2017. Gated multimodal units for information fusion. Retrieved from https:\/\/arXiv:1702.01992"},{"key":"e_1_3_3_6_2","unstructured":"Jinze Bai Shuai Bai Shusheng Yang Shijie Wang Sinan Tan Peng Wang Junyang Lin Chang Zhou and Jingren Zhou. 2023. Qwen-VL: A versatile vision-language model for understanding localization text reading and beyond. Retrieved from https:\/\/arXiv:2308.12966"},{"key":"e_1_3_3_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00438"},{"key":"e_1_3_3_8_2","article-title":"Multimodal machine learning: A survey and taxonomy","author":"Baltru\u0161aitis Tadas","year":"2018","unstructured":"Tadas Baltru\u0161aitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2018. Multimodal machine learning: A survey and taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 41, 2 (2018), 423\u2013443.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_22"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00755"},{"key":"e_1_3_3_11_2","unstructured":"Keqin Chen Zhao Zhang Weili Zeng Richong Zhang Feng Zhu and Rui Zhao. 2023. Shikra: Unleashing multimodal LLM\u2019s referential dialogue magic. Retrieved from https:\/\/arXiv:2306.15195"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i2.16188"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00998"},{"key":"e_1_3_3_14_2","unstructured":"Xinpeng Chen Lin Ma Jingyuan Chen Zequn Jie Wei Liu and Jiebo Luo. 2018. Real-time referring expression comprehension by single-stage grounding network. Retrieved from https:\/\/arXiv:1812.03426"},{"key":"e_1_3_3_15_2","volume-title":"Proceedings of the BMVC","author":"Chen Yi Wen","year":"2020","unstructured":"Yi Wen Chen, Yi Hsuan Tsai, Tiantian Wang, Yen Yu Lin, and Ming Hsuan Yang. 2020. Referring expression object segmentation with caption-aware consistency. In Proceedings of the BMVC."},{"key":"e_1_3_3_16_2","doi-asserted-by":"crossref","unstructured":"Zesen Cheng Kehan Li Peng Jin Xiangyang Ji Li Yuan Chang Liu and Jie Chen. 2023. Parallel vertex diffusion for unified visual grounding. Retrieved from https:\/\/arXiv:2303.07216","DOI":"10.1609\/aaai.v38i2.27896"},{"key":"e_1_3_3_17_2","unstructured":"Wei-Lin Chiang Zhuohan Li Zi Lin Ying Sheng Zhanghao Wu Hao Zhang Lianmin Zheng Siyuan Zhuang Yonghao Zhuang Joseph E. Gonzalez Ion Stoica and Eric P. Xing. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. 2 3 (2023) 6. https:\/\/vicuna.lmsys.org (accessed 14 April 2023)."},{"key":"e_1_3_3_18_2","article-title":"Twins: Revisiting the design of spatial attention in vision transformers","author":"Chu Xiangxiang","year":"2021","unstructured":"Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. 2021. Twins: Revisiting the design of spatial attention in vision transformers. Proceedings of the NeurIPS (2021).","journal-title":"Proceedings of the NeurIPS"},{"key":"e_1_3_3_19_2","unstructured":"Xiangxiang Chu Bo Zhang Zhi Tian Xiaolin Wei and Huaxia Xia. 2021. Do we really need explicit position encodings for vision transformers? arXiv preprint arXiv:2102.10882 3 8 (2021)."},{"key":"e_1_3_3_20_2","volume-title":"Proceedings of the NeurIPS","author":"Vries Harm De","year":"2017","unstructured":"Harm De Vries, Florian Strub, Jeremie Mary, Hugo Larochelle, Olivier Pietquin, and Aaron C. Courville. 2017. Modulating early visual processing by language. In Proceedings of the NeurIPS."},{"key":"e_1_3_3_21_2","doi-asserted-by":"crossref","unstructured":"Jiajun Deng Zhengyuan Yang Tianlang Chen Wengang Zhou and Houqiang Li. 2021. TransVG: End-to-end visual grounding with transformers. Retrieved from https:\/\/arXiv:2104.08541","DOI":"10.1109\/ICCV48922.2021.00179"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3296823"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2021.104257"},{"key":"e_1_3_3_24_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. Retrieved from https:\/\/arXiv:1810.04805"},{"key":"e_1_3_3_25_2","unstructured":"Wonjae Kim Bokyung Son and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning PMLR 5583\u20135594."},{"key":"e_1_3_3_26_2","volume-title":"Proceedings of the ICLR","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16 \\(\\times\\) 16 words: Transformers for image recognition at scale. In Proceedings of the ICLR."},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICME52920.2022.9859880"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2009.03.008"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01525"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1044"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.201"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.169"},{"key":"e_1_3_3_33_2","unstructured":"Kai Han An Xiao Enhua Wu Jianyuan Guo Chunjing Xu and Yunhe Wang. 2021. Transformer in transformer. Retrieved from https:\/\/arXiv:2103.00112"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.322"},{"key":"e_1_3_3_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_3_36_2","volume-title":"Proceedings of the ECCV","author":"Ho Chih-Hui","year":"2022","unstructured":"Chih-Hui Ho, Srikar Appalaraju, Bhavan Jasani, R. Manmatha, and Nuno Vasconcelos. 2022. Yoro-lightweight end to end visual grounding. In Proceedings of the ECCV."},{"key":"e_1_3_3_37_2","article-title":"Learning to compose and reason with language tree structures for visual grounding","author":"Hong Richang","year":"2019","unstructured":"Richang Hong, Daqing Liu, Xiaoyu Mo, Xiangnan He, and Hanwang Zhang. 2019. Learning to compose and reason with language tree structures for visual grounding. IEEE Trans. Pattern Anal. Mach. Intell. 44, 2 (2019), 684\u2013696.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_3_38_2","doi-asserted-by":"crossref","unstructured":"Wenyi Hong Weihan Wang Qingsong Lv Jiazheng Xu Wenmeng Yu Junhui Ji Yan Wang Zihan Wang Yuxiao Dong Ming Ding et\u00a0al. 2023. CogAgent: A visual language model for GUI agents. Retrieved from https:\/\/arXiv:2312.08914","DOI":"10.1109\/CVPR52733.2024.01354"},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.470"},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_7"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00448"},{"key":"e_1_3_3_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01661"},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i1.19983"},{"key":"e_1_3_3_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01050"},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58607-2_4"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i7.25971"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00180"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1086"},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01761"},{"key":"e_1_3_3_50_2","unstructured":"Junnan Li Dongxu Li Silvio Savarese and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. Retrieved from https:\/\/arXiv:2301.12597"},{"key":"e_1_3_3_51_2","article-title":"Referring transformer: A one-step approach to multi-task visual grounding","author":"Li Muchen","year":"2021","unstructured":"Muchen Li and Leonid Sigal. 2021. Referring transformer: A one-step approach to multi-task visual grounding. Proceedings of the NeurIPS (2021).","journal-title":"Proceedings of the NeurIPS"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00602"},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01089"},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.324"},{"key":"e_1_3_3_55_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02259"},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.143"},{"key":"e_1_3_3_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00477"},{"key":"e_1_3_3_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01386"},{"key":"e_1_3_3_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00205"},{"key":"e_1_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6833"},{"key":"e_1_3_3_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01170"},{"key":"e_1_3_3_63_2","article-title":"Swin transformer: Hierarchical vision transformer using shifted windows","author":"Liu Ze","year":"2021","unstructured":"Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the ICCV.","journal-title":"Proceedings of the ICCV"},{"key":"e_1_3_3_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01005"},{"key":"e_1_3_3_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.9"},{"key":"e_1_3_3_66_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01252-6_39"},{"key":"e_1_3_3_67_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i15.17602"},{"key":"e_1_3_3_68_2","unstructured":"Arsha Nagrani Shan Yang Anurag Arnab Aren Jansen Cordelia Schmid and Chen Sun. 2021. Attention bottlenecks for multimodal fusion. Retrieved from https:\/\/arXiv:2107.00135"},{"key":"e_1_3_3_69_2","unstructured":"Zhiliang Peng Wenhui Wang Li Dong Yaru Hao Shaohan Huang Shuming Ma and Furu Wei. 2023. Kosmos-2: Grounding multimodal large language models to the world. Retrieved from https:\/\/arXiv:2306.14824"},{"key":"e_1_3_3_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.303"},{"key":"e_1_3_3_71_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19833-5_32"},{"key":"e_1_3_3_72_2","doi-asserted-by":"crossref","unstructured":"Hanoona Rasheed Muhammad Maaz Sahal Shaji Abdelrahman Shaker Salman Khan Hisham Cholakkal Rao M. Anwer Erix Xing Ming-Hsuan Yang and Fahad S. Khan. 2023. Glamm: Pixel grounding large multimodal model. Retrieved from https:\/\/arXiv:2311.03356","DOI":"10.1109\/CVPR52733.2024.01236"},{"key":"e_1_3_3_73_2","unstructured":"Joseph Redmon and Ali Farhadi. 2018. Yolov3: An incremental improvement. Retrieved from https:\/\/arXiv:1804.02767"},{"key":"e_1_3_3_74_2","unstructured":"Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. Retrieved from https:\/\/arXiv:1506.01497"},{"key":"e_1_3_3_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00075"},{"key":"e_1_3_3_76_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_49"},{"key":"e_1_3_3_77_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01231-1_3"},{"key":"e_1_3_3_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01384"},{"key":"e_1_3_3_79_2","volume-title":"Proceedings of the ICML","author":"Touvron Hugo","year":"2021","unstructured":"Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. 2021. Training data-efficient image transformers distillation through attention. In Proceedings of the ICML."},{"key":"e_1_3_3_80_2","volume-title":"Proceedings of the CVPR","author":"Joze Hamid Reza Vaezi","year":"2020","unstructured":"Hamid Reza Vaezi Joze, Amirreza Shaban, Michael L. Iuzzolino, and Kazuhito Koishida. 2020. MMTM: Multimodal transfer module for CNN fusion. In Proceedings of the CVPR."},{"key":"e_1_3_3_81_2","volume-title":"Proceedings of the NeurIPS","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the NeurIPS."},{"key":"e_1_3_3_82_2","volume-title":"Proceedings of the ACL","author":"Vogel Adam","year":"2010","unstructured":"Adam Vogel and Dan Jurafsky. 2010. Learning to follow navigational directions. In Proceedings of the ACL."},{"key":"e_1_3_3_83_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612322"},{"key":"e_1_3_3_84_2","article-title":"Learning two-branch neural networks for image-text matching tasks","author":"Wang Liwei","year":"2018","unstructured":"Liwei Wang, Yin Li, Jing Huang, and Svetlana Lazebnik. 2018. Learning two-branch neural networks for image-text matching tasks. IEEE Trans. Pattern Anal. Mach. Intell. 41, 2 (2018), 394\u2013407.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_3_85_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00206"},{"key":"e_1_3_3_86_2","doi-asserted-by":"publisher","DOI":"10.2737\/FPL-GTR-290"},{"key":"e_1_3_3_87_2","unstructured":"Wenhai Wang Zhe Chen Xiaokang Chen Jiannan Wu Xizhou Zhu Gang Zeng Ping Luo Tong Lu Jie Zhou Yu Qiao et\u00a0al. 2023. Visionllm: Large language model is also an open-ended decoder for vision-centric tasks. Retrieved from https:\/\/arXiv:2305.11175"},{"key":"e_1_3_3_88_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00061"},{"key":"e_1_3_3_89_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01014"},{"key":"e_1_3_3_90_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413971"},{"key":"e_1_3_3_91_2","doi-asserted-by":"publisher","DOI":"10.2737\/FPL-GTR-290"},{"key":"e_1_3_3_92_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"e_1_3_3_93_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00278"},{"key":"e_1_3_3_94_2","unstructured":"Liang Xie Jialie Shen Jungong Han Lei Zhu and Ling Shao. 2017. Dynamic multi-view hashing for online image retrieval. In IJCAI 78 (2017) 122."},{"key":"e_1_3_3_95_2","article-title":"Deep multi-view enhancement hashing for image retrieval","author":"Yan Chenggang","year":"2020","unstructured":"Chenggang Yan, Biao Gong, Yuxuan Wei, and Yue Gao. 2020. Deep multi-view enhancement hashing for image retrieval. IEEE Trans. Pattern Anal. Mach. Intell. 43, 4 (2020), 1445\u20131451.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_3_96_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00928"},{"key":"e_1_3_3_97_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00474"},{"key":"e_1_3_3_98_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00997"},{"key":"e_1_3_3_99_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58568-6_23"},{"key":"e_1_3_3_100_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00478"},{"key":"e_1_3_3_101_2","unstructured":"Zhengyuan Yang Linjie Li Jianfeng Wang Kevin Lin Ehsan Azarnasab Faisal Ahmed Zicheng Liu Ce Liu Michael Zeng and Lijuan Wang. 2023. Mm-react: Prompting chatgpt for multimodal reasoning and action. arXiv preprint arXiv:2303.11381."},{"key":"e_1_3_3_102_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01762"},{"key":"e_1_3_3_103_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475313"},{"key":"e_1_3_3_104_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01506"},{"key":"e_1_3_3_105_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00640"},{"key":"e_1_3_3_106_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.503"},{"key":"e_1_3_3_107_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.508"},{"key":"e_1_3_3_108_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00142"},{"key":"e_1_3_3_109_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_5"},{"key":"e_1_3_3_110_2","doi-asserted-by":"crossref","unstructured":"Li Yuan Yunpeng Chen Tao Wang Weihao Yu Yujun Shi Zihang Jiang Francis E. H. Tay Jiashi Feng and Shuicheng Yan. 2021. Tokens-to-token ViT: Training vision transformers from scratch on imagenet. Retrieved from https:\/\/arXiv:2101.11986","DOI":"10.1109\/ICCV48922.2021.00060"},{"key":"e_1_3_3_111_2","doi-asserted-by":"crossref","unstructured":"Amir Zadeh Minghai Chen Soujanya Poria Erik Cambria and Louis-Philippe Morency. 2017. Tensor fusion network for multimodal sentiment analysis. Retrieved from https:\/\/arXiv:1707.07250","DOI":"10.18653\/v1\/D17-1115"},{"key":"e_1_3_3_112_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01030"},{"key":"e_1_3_3_113_2","unstructured":"Chi Zhang Zhao Yang Jiaxuan Liu Yucheng Han Xin Chen Zebiao Huang Bin Fu and Gang Yu. 2023. AppAgent: Multimodal Agents as Smartphone Users. Retrieved from https:\/\/arxiv:2312.13771"},{"key":"e_1_3_3_114_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00437"},{"key":"e_1_3_3_115_2","article-title":"Multimodal marketing intent analysis for effective targeted advertising","author":"Zhang Lu","year":"2021","unstructured":"Lu Zhang, Jialie Shen, Jian Zhang, Jingsong Xu, Zhibin Li, Yazhou Yao, and Litao Yu. 2021. Multimodal marketing intent analysis for effective targeted advertising. IEEE Trans. Multimedia 24 (2021), 1830\u20131843.","journal-title":"IEEE Trans. Multimedia"},{"key":"e_1_3_3_116_2","article-title":"Language-guided navigation via cross-modal grounding and alternate adversarial learning","author":"Zhang Weixia","year":"2020","unstructured":"Weixia Zhang, Chao Ma, Qi Wu, and Xiaokang Yang. 2020. Language-guided navigation via cross-modal grounding and alternate adversarial learning. IEEE Trans. Circ. Syst. Video Technol. 31, 9 (2020), 3469\u20133481.","journal-title":"IEEE Trans. Circ. Syst. Video Technol."},{"key":"e_1_3_3_117_2","article-title":"CoupAlign: Coupling word-pixel with sentence-mask alignments for referring image segmentation","author":"Zhang Zicheng","year":"2022","unstructured":"Zicheng Zhang, Yi Zhu, Jianzhuang Liu, Xiaodan Liang, and Wei Ke. 2022. CoupAlign: Coupling word-pixel with sentence-mask alignments for referring image segmentation. Proceedings of the NeurIPS (2022).","journal-title":"Proceedings of the NeurIPS"},{"key":"e_1_3_3_118_2","article-title":"Word2Pix: Word to pixel cross-attention transformer in visual grounding","author":"Zhao Heng","year":"2022","unstructured":"Heng Zhao, Joey Tianyi Zhou, and Yew-Soon Ong. 2022. Word2Pix: Word to pixel cross-attention transformer in visual grounding. IEEE Trans. Neural Netw. Learn. Syst. (2022).","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"e_1_3_3_119_2","unstructured":"Yang Zhao Zhijie Lin Daquan Zhou Zilong Huang Jiashi Feng and Bingyi Kang. 2023. BuboGPT: Enabling visual grounding in multi-modal LLMs. Retrieved from https:\/\/arXiv:2307.08581"},{"key":"e_1_3_3_120_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19833-5_35"},{"key":"e_1_3_3_121_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00995"},{"key":"e_1_3_3_122_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.540"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3660638","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3660638","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:50:21Z","timestamp":1750287021000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3660638"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,13]]},"references-count":121,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2024,8,31]]}},"alternative-id":["10.1145\/3660638"],"URL":"https:\/\/doi.org\/10.1145\/3660638","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,13]]},"assertion":[{"value":"2023-08-11","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-17","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-06-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}