{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T18:12:02Z","timestamp":1767982322288,"version":"3.49.0"},"publisher-location":"New York, NY, USA","reference-count":47,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2021ZD0113004"],"award-info":[{"award-number":["2021ZD0113004"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3548417","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:43:01Z","timestamp":1665416581000},"page":"4336-4344","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Towards Further Comprehension on Referring Expression with Rationale"],"prefix":"10.1145","author":[{"given":"Rengang","family":"Li","sequence":"first","affiliation":[{"name":"Inspur Electronic Information Industry Co.,Ltd. &amp; State Key Laboratory of High-end Server &amp; Storage Technology, Jinan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Baoyu","family":"Fan","sequence":"additional","affiliation":[{"name":"Inspur Electronic Information Industry Co.,Ltd. &amp; State Key Laboratory of High-end Server &amp; Storage Technology, Jinan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaochuan","family":"Li","sequence":"additional","affiliation":[{"name":"Inspur Electronic Information Industry Co.,Ltd. &amp; State Key Laboratory of High-end Server &amp; Storage Technology, Jinan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Runze","family":"Zhang","sequence":"additional","affiliation":[{"name":"Inspur Electronic Information Industry Co.,Ltd. &amp; State Key Laboratory of High-end Server &amp; Storage Technology, Jinan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhenhua","family":"Guo","sequence":"additional","affiliation":[{"name":"Inspur Electronic Information Industry Co.,Ltd. &amp; State Key Laboratory of High-end Server &amp; Storage Technology, Jinan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kun","family":"Zhao","sequence":"additional","affiliation":[{"name":"Inspur Electronic Information Industry Co.,Ltd. &amp; State Key Laboratory of High-end Server &amp; Storage Technology, Jinan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yaqian","family":"Zhao","sequence":"additional","affiliation":[{"name":"Inspur Electronic Information Industry Co.,Ltd. &amp; State Key Laboratory of High-end Server &amp; Storage Technology, Jinan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weifeng","family":"Gong","sequence":"additional","affiliation":[{"name":"Inspur Electronic Information Industry Co.,Ltd. &amp; State Key Laboratory of High-end Server &amp; Storage Technology, Jinan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Endong","family":"Wang","sequence":"additional","affiliation":[{"name":"Inspur Electronic Information Industry Co.,Ltd. &amp; State Key Laboratory of High-end Server &amp; Storage Technology, Jinan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10514-018-9792-8"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"e_1_3_2_2_4_1","volume-title":"arXiv preprint arXiv:2106.02192","author":"Chandu Khyathi Raghavi","year":"2021","unstructured":"Khyathi Raghavi Chandu , Yonatan Bisk , and Alan W Black . 2021. Grounding'Grounding'in NLP. arXiv preprint arXiv:2106.02192 ( 2021 ). Khyathi Raghavi Chandu, Yonatan Bisk, and Alan W Black. 2021. Grounding'Grounding'in NLP. arXiv preprint arXiv:2106.02192 (2021)."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58577-8_7"},{"key":"e_1_3_2_2_6_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10086--10095","author":"Chen Zhenfang","year":"2020","unstructured":"Zhenfang Chen , Peng Wang , Lin Ma , Kwan-Yee K Wong , and Qi Wu . 2020 . Copsref: A new dataset and task on compositional referring expression comprehension . In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10086--10095 . Zhenfang Chen, Peng Wang, Lin Ma, Kwan-Yee K Wong, and Qi Wu. 2020. Copsref: A new dataset and task on compositional referring expression comprehension. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10086--10095."},{"key":"e_1_3_2_2_7_1","volume-title":"Unifying vision-and language tasks via text generation. arXiv preprint arXiv:2102.02779","author":"Cho Jaemin","year":"2021","unstructured":"Jaemin Cho , Jie Lei , Hao Tan , and Mohit Bansal . 2021. Unifying vision-and language tasks via text generation. arXiv preprint arXiv:2102.02779 ( 2021 ). Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. 2021. Unifying vision-and language tasks via text generation. arXiv preprint arXiv:2102.02779 (2021)."},{"key":"e_1_3_2_2_8_1","volume-title":"Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio.","author":"Cho Kyunghyun","year":"2014","unstructured":"Kyunghyun Cho , Bart Van Merri\u00ebnboer , Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 . Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078 (2014). Kyunghyun Cho, Bart Van Merri\u00ebnboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078 (2014)."},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00850"},{"key":"e_1_3_2_2_10_1","volume-title":"TransVG: End-to-End Visual Grounding with Transformers. arXiv preprint arXiv:2104.08541","author":"Deng Jiajun","year":"2021","unstructured":"Jiajun Deng , Zhengyuan Yang , Tianlang Chen , Wengang Zhou , and Houqiang Li. 2021. TransVG: End-to-End Visual Grounding with Transformers. arXiv preprint arXiv:2104.08541 ( 2021 ). Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li. 2021. TransVG: End-to-End Visual Grounding with Transformers. arXiv preprint arXiv:2104.08541 (2021)."},{"key":"e_1_3_2_2_11_1","volume-title":"Luc Van Gool, and Marie-Francine Moens","author":"Deruyttere Thierry","year":"2019","unstructured":"Thierry Deruyttere , Simon Vandenhende , Dusan Grujicic , Luc Van Gool, and Marie-Francine Moens . 2019 . Talk2car: Taking control of your self-driving car. arXiv preprint arXiv:1909.10838 (2019). Thierry Deruyttere, Simon Vandenhende, Dusan Grujicic, Luc Van Gool, and Marie-Francine Moens. 2019. Talk2car: Taking control of your self-driving car. arXiv preprint arXiv:1909.10838 (2019)."},{"key":"e_1_3_2_2_12_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_3_2_2_13_1","volume-title":"Understanding image and text simultaneously: a dual vision-language machine comprehension task. arXiv preprint arXiv:1612.07833","author":"Ding Nan","year":"2016","unstructured":"Nan Ding , Sebastian Goodman , Fei Sha , and Radu Soricut . 2016. Understanding image and text simultaneously: a dual vision-language machine comprehension task. arXiv preprint arXiv:1612.07833 ( 2016 ). Nan Ding, Sebastian Goodman, Fei Sha, and Radu Soricut. 2016. Understanding image and text simultaneously: a dual vision-language machine comprehension task. arXiv preprint arXiv:1612.07833 (2016)."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3414038"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2911066"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.470"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.493"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01661"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.215"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00180"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1086"},{"key":"e_1_3_2_2_23_1","volume-title":"Children's reading comprehension and assessment","author":"Kintsch Walter","unstructured":"Walter Kintsch and Eileen Kintsch . 2005. Comprehension . In Children's reading comprehension and assessment . Routledge , 89--110. Walter Kintsch and Eileen Kintsch. 2005. Comprehension. In Children's reading comprehension and assessment. Routledge, 89--110."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"crossref","unstructured":"Ranjay Krishna Yuke Zhu Oliver Groth Justin Johnson Kenji Hata Joshua Kravitz Stephanie Chen Yannis Kalantidis Li-Jia Li David A Shamma etal 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision 123 1 (2017) 32--73.  Ranjay Krishna Yuke Zhu Oliver Groth Justin Johnson Kenji Hata Joshua Kravitz Stephanie Chen Yannis Kalantidis Li-Jia Li David A Shamma et al. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision 123 1 (2017) 32--73.","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3124317"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3161076"},{"key":"e_1_3_2_2_27_1","volume-title":"Joint active learning with feature selection via cur matrix decomposition","author":"Li Changsheng","year":"2018","unstructured":"Changsheng Li , Xiangfeng Wang , Weishan Dong , Junchi Yan , Qingshan Liu , and Hongyuan Zha . 2018. Joint active learning with feature selection via cur matrix decomposition . IEEE transactions on pattern analysis and machine intelligence 41, 6 ( 2018 ), 1382--1396. Changsheng Li, Xiangfeng Wang, Weishan Dong, Junchi Yan, Qingshan Liu, and Hongyuan Zha. 2018. Joint active learning with feature selection via cur matrix decomposition. IEEE transactions on pattern analysis and machine intelligence 41, 6 (2018), 1382--1396."},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00431"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01045"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.9"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3042066"},{"key":"e_1_3_2_2_33_1","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei Ilya Sutskever etal 2019. Language models are unsupervised multitask learners. OpenAI blog 1 8 (2019) 9.  Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei Ilya Sutskever et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1 8 (2019) 9."},{"key":"e_1_3_2_2_34_1","volume-title":"Foil it! find one mismatch between image and language caption. arXiv preprint arXiv:1705.01359","author":"Shekhar Ravi","year":"2017","unstructured":"Ravi Shekhar , Sandro Pezzelle , Yauhen Klimovich , Aur\u00e9lie Herbelot , Moin Nabi , Enver Sangineto , and Raffaella Bernardi . 2017. Foil it! find one mismatch between image and language caption. arXiv preprint arXiv:1705.01359 ( 2017 ). Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aur\u00e9lie Herbelot, Moin Nabi, Enver Sangineto, and Raffaella Bernardi. 2017. Foil it! find one mismatch between image and language caption. arXiv preprint arXiv:1705.01359 (2017)."},{"key":"e_1_3_2_2_35_1","volume-title":"Language Grounding in Robots","author":"Spranger Michael","unstructured":"Michael Spranger and Simon Pauw . 2012. Dealing with perceptual deviation: vague semantics for spatial language and quantification . In Language Grounding in Robots . Springer , 173--192. Michael Spranger and Simon Pauw. 2012. Dealing with perceptual deviation: vague semantics for spatial language and quantification. In Language Grounding in Robots. Springer, 173--192."},{"key":"e_1_3_2_2_36_1","volume-title":"Evolving grounded communication for robots. Trends in cognitive sciences 7, 7","author":"Steels Luc","year":"2003","unstructured":"Luc Steels . 2003. Evolving grounded communication for robots. Trends in cognitive sciences 7, 7 ( 2003 ), 308--312. Luc Steels. 2003. Evolving grounded communication for robots. Trends in cognitive sciences 7, 7 (2003), 308--312."},{"key":"e_1_3_2_2_37_1","volume-title":"Proceedings of the fourth european conference on artificial life","volume":"97","author":"Steels Luc","year":"1997","unstructured":"Luc Steels and Paul Vogt . 1997 . Grounding adaptive language games in robotic agents . In Proceedings of the fourth european conference on artificial life , Vol. 97 . Luc Steels and Paul Vogt. 1997. Grounding adaptive language games in robotic agents. In Proceedings of the fourth european conference on artificial life, Vol. 97."},{"key":"e_1_3_2_2_38_1","volume-title":"Vl-bert: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530","author":"Su Weijie","year":"2019","unstructured":"Weijie Su , Xizhou Zhu , Yue Cao , Bin Li , Lewei Lu , Furu Wei , and Jifeng Dai . 2019 . Vl-bert: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530 (2019). Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2019. Vl-bert: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530 (2019)."},{"key":"e_1_3_2_2_39_1","volume-title":"Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490","author":"Tan Hao","year":"2019","unstructured":"Hao Tan and Mohit Bansal . 2019 . Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490 (2019). Hao Tan and Mohit Bansal. 2019. Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490 (2019)."},{"key":"e_1_3_2_2_40_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , \u0141ukasz Kaiser , and Illia Polosukhin . 2017. Attention is all you need. Advances in neural information processing systems 30 ( 2017 ). Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475340"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413905"},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58568-6_23"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00142"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_5"},{"key":"e_1_3_2_2_47_1","volume-title":"A real-time global inference network for onestage referring expression comprehension","author":"Zhou Yiyi","year":"2021","unstructured":"Yiyi Zhou , Rongrong Ji , Gen Luo , Xiaoshuai Sun , Jinsong Su , Xinghao Ding , Chia-Wen Lin , and Qi Tian . 2021. A real-time global inference network for onestage referring expression comprehension . IEEE Transactions on Neural Networks and Learning Systems ( 2021 ). Yiyi Zhou, Rongrong Ji, Gen Luo, Xiaoshuai Sun, Jinsong Su, Xinghao Ding, Chia-Wen Lin, and Qi Tian. 2021. A real-time global inference network for onestage referring expression comprehension. IEEE Transactions on Neural Networks and Learning Systems (2021)."}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548417","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3548417","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:17Z","timestamp":1750182557000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548417"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":47,"alternative-id":["10.1145\/3503161.3548417","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3548417","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}