{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,1]],"date-time":"2025-12-01T11:22:26Z","timestamp":1764588146791,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":56,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T00:00:00Z","timestamp":1602460800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Natural Science Foundation of China","award":["No. 61702037","No. 61773062"],"award-info":[{"award-number":["No. 61702037","No. 61773062"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,12]]},"DOI":"10.1145\/3394171.3413902","type":"proceedings-article","created":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T12:26:25Z","timestamp":1602505585000},"page":"4041-4050","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":28,"title":["Visual-Semantic Graph Matching for Visual Grounding"],"prefix":"10.1145","author":[{"given":"Chenchen","family":"Jing","sequence":"first","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuwei","family":"Wu","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mingtao","family":"Pei","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing , China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yao","family":"Hu","sequence":"additional","affiliation":[{"name":"Alibaba Youku Cognitive &amp;38; Intelligent Lab, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yunde","family":"Jia","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qi","family":"Wu","sequence":"additional","affiliation":[{"name":"University of Adelaide, Adelaide, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,10,12]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3674--3683","author":"Anderson Peter","key":"e_1_3_2_2_2_1","unstructured":"Peter Anderson , Qi Wu , Damien Teney , Jake Bruce , Mark Johnson , Niko S\u00fcnderhauf , Ian Reid , Stephen Gould , and Anton van den Hengel. 2018b. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3674--3683 . Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko S\u00fcnderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel. 2018b. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3674--3683."},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00438"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00808"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00430"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00190"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.337"},{"key":"e_1_3_2_2_9_1","volume-title":"Learning to Compose and Reason with Language Tree Structures for Visual Grounding","author":"Hong Richang","year":"2019","unstructured":"Richang Hong , Daqing Liu , Xiaoyu Mo , Xiangnan He , and Hanwang Zhang . 2019. Learning to Compose and Reason with Language Tree Structures for Visual Grounding . IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) ( 2019 ). Richang Hong, Daqing Liu, Xiaoyu Mo, Xiangnan He, and Hanwang Zhang. 2019. Learning to Compose and Reason with Language Tree Structures for Visual Grounding. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) (2019)."},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.01039"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.470"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.493"},{"key":"e_1_3_2_2_13_1","unstructured":"Drew Hudson and Christopher D Manning. 2019. Learning by abstraction: The neural state machine. In Advances in Neural Information Processing Systems (NeurIPS). 5901--5914.  Drew Hudson and Christopher D Manning. 2019. Learning by abstraction: The neural state machine. In Advances in Neural Information Processing Systems (NeurIPS). 5901--5914."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6776"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298990"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1086"},{"key":"e_1_3_2_2_17_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Jimmy Ba . 2014 . Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014). Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2005.20"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_2_21_1","volume-title":"2019 a. Referring Expression Grounding by Marginalizing Scene Graph Likelihood. arXiv preprint arXiv:1906.03561","author":"Liu Daqing","year":"2019","unstructured":"Daqing Liu , Hanwang Zhang , Zheng-Jun Zha , and Fanglin Wang . 2019 a. Referring Expression Grounding by Marginalizing Scene Graph Likelihood. arXiv preprint arXiv:1906.03561 ( 2019 ). Daqing Liu, Hanwang Zhang, Zheng-Jun Zha, and Fanglin Wang. 2019 a. Referring Expression Grounding by Marginalizing Scene Graph Likelihood. arXiv preprint arXiv:1906.03561 (2019)."},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00477"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1515"},{"key":"e_1_3_2_2_24_1","volume-title":"Learning Cross-modal Context Graph for Visual Grounding. In Thirty-Fourth AAAI Conference on Artificial Intelligence. 11645--11652","author":"Liu Yongfei","year":"2020","unstructured":"Yongfei Liu , Bo Wan , Xiaodan Zhu , and Xuming He . 2020 . Learning Cross-modal Context Graph for Visual Grounding. In Thirty-Fourth AAAI Conference on Artificial Intelligence. 11645--11652 . Yongfei Liu, Bo Wan, Xiaodan Zhu, and Xuming He. 2020. Learning Cross-modal Context Graph for Visual Grounding. In Thirty-Fourth AAAI Conference on Artificial Intelligence. 11645--11652."},{"key":"e_1_3_2_2_25_1","volume-title":"Paulo Oswaldo Boaventura-Netto, Peter Hahn, and Tania Querido.","author":"Loiola Eliane Maria","year":"2007","unstructured":"Eliane Maria Loiola , Nair Maria Maia de Abreu , Paulo Oswaldo Boaventura-Netto, Peter Hahn, and Tania Querido. 2007 . A survey for the quadratic assignment problem. European journal of operational research, Vol. 176 , 2 (2007), 657--690. Eliane Maria Loiola, Nair Maria Maia de Abreu, Paulo Oswaldo Boaventura-Netto, Peter Hahn, and Tania Querido. 2007. A survey for the quadratic assignment problem. European journal of operational research, Vol. 176, 2 (2007), 657--690."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.333"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.9"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_48"},{"key":"e_1_3_2_2_29_1","unstructured":"Adam Paszke Sam Gross Soumith Chintala Gregory Chanan Edward Yang Zachary DeVito Zeming Lin Alban Desmaison Luca Antiga and Adam Lerer. 2017. Automatic differentiation in pytorch. (2017).  Adam Paszke Sam Gross Soumith Chintala Gregory Chanan Edward Yang Zachary DeVito Zeming Lin Alban Desmaison Luca Antiga and Adam Lerer. 2017. Automatic differentiation in pytorch. (2017)."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1202"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01258-8_16"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.303"},{"key":"e_1_3_2_2_33_1","unstructured":"Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems (NeurIPS). 91--99.  Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems (NeurIPS). 91--99."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_49"},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/78.650093"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W15-2812"},{"key":"e_1_3_2_2_37_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.541"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_42"},{"volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 1960--1968","author":"Wang Peng","key":"e_1_3_2_2_41_1","unstructured":"Peng Wang , Qi Wu , Jiewei Cao , Chunhua Shen , Lianli Gao , and Anton van den Hengel. 2019 b. Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks . In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 1960--1968 . Peng Wang, Qi Wu, Jiewei Cao, Chunhua Shen, Lianli Gao, and Anton van den Hengel. 2019 b. Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 1960--1968."},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00315"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00679"},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2017.05.001"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00427"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00474"},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01094"},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2893066"},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00142"},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_5"},{"key":"e_1_3_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.375"},{"key":"e_1_3_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2018\/155"},{"key":"e_1_3_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00284"},{"key":"e_1_3_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00437"},{"key":"e_1_3_2_2_55_1","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 127--134","author":"Zhou Feng","year":"2012","unstructured":"Feng Zhou and Fernando De la Torre . 2012 . Factorized graph matching . In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 127--134 . Feng Zhou and Fernando De la Torre. 2012. Factorized graph matching. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 127--134."},{"volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 4252--4261","author":"Zhuang Bohan","key":"e_1_3_2_2_56_1","unstructured":"Bohan Zhuang , Qi Wu , Chunhua Shen , Ian Reid , and Anton van den Hengel. 2018. Parallel attention: A unified framework for visual object discovery through dialogs and queries . In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 4252--4261 . Bohan Zhuang, Qi Wu, Chunhua Shen, Ian Reid, and Anton van den Hengel. 2018. Parallel attention: A unified framework for visual object discovery through dialogs and queries. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 4252--4261."}],"event":{"name":"MM '20: The 28th ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Seattle WA USA","acronym":"MM '20"},"container-title":["Proceedings of the 28th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413902","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3394171.3413902","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:32:06Z","timestamp":1750195926000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413902"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,12]]},"references-count":56,"alternative-id":["10.1145\/3394171.3413902","10.1145\/3394171"],"URL":"https:\/\/doi.org\/10.1145\/3394171.3413902","relation":{},"subject":[],"published":{"date-parts":[[2020,10,12]]},"assertion":[{"value":"2020-10-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}