{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:17:11Z","timestamp":1750220231716,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":44,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,12,1]],"date-time":"2021-12-01T00:00:00Z","timestamp":1638316800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,12]]},"DOI":"10.1145\/3469877.3490592","type":"proceedings-article","created":{"date-parts":[[2022,1,10]],"date-time":"2022-01-10T18:27:21Z","timestamp":1641839241000},"page":"1-8","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Entity Relation Fusion for Real-Time One-Stage Referring Expression Comprehension"],"prefix":"10.1145","author":[{"given":"Hang","family":"Yu","sequence":"first","affiliation":[{"name":"Beihang University, CN"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weixin","family":"Li","sequence":"additional","affiliation":[{"name":"Beihang University, CN"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiankai","family":"Li","sequence":"additional","affiliation":[{"name":"Beihang University, CN"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ye","family":"Du","sequence":"additional","affiliation":[{"name":"Beihang University, CN"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,1,10]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI","author":"Chen Long","year":"2021","unstructured":"Long Chen , Wenbo Ma , Jun Xiao , Hanwang Zhang , and Shih-Fu Chang . 2021. Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression Grounding . In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021 , Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021. 1036\u20131044. Long Chen, Wenbo Ma, Jun Xiao, Hanwang Zhang, and Shih-Fu Chang. 2021. Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression Grounding. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021. 1036\u20131044."},{"unstructured":"Xinpeng Chen Lin Ma Jingyuan Chen Zequn Jie Wei Liu and Jiebo Luo. 2018. Real-Time Referring Expression Comprehension by Single-Stage Grounding Network. arxiv:1812.03426\u00a0[cs.CV]  Xinpeng Chen Lin Ma Jingyuan Chen Zequn Jie Wei Liu and Jiebo Luo. 2018. Real-Time Referring Expression Comprehension by Single-Stage Grounding Network. arxiv:1812.03426\u00a0[cs.CV]","key":"e_1_3_2_1_2_1"},{"key":"e_1_3_2_1_3_1","volume-title":"Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18)","author":"Cirik Volkan","year":"2018","unstructured":"Volkan Cirik , Taylor Berg-Kirkpatrick , and Louis-Philippe Morency . 2018 . Using Syntax to Ground Referring Expressions in Natural Images . In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18) , the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18). 6756\u20136764. Volkan Cirik, Taylor Berg-Kirkpatrick, and Louis-Philippe Morency. 2018. Using Syntax to Ground Referring Expressions in Natural Images. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18). 6756\u20136764."},{"key":"e_1_3_2_1_4_1","volume-title":"Deformable Convolutional Networks. In IEEE International Conference on Computer Vision, ICCV","author":"Dai Jifeng","year":"2017","unstructured":"Jifeng Dai , Haozhi Qi , Yuwen Xiong , Yi Li , Guodong Zhang , Han Hu , and Yichen Wei . 2017 . Deformable Convolutional Networks. In IEEE International Conference on Computer Vision, ICCV 2017. 764\u2013773. Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. 2017. Deformable Convolutional Networks. In IEEE International Conference on Computer Vision, ICCV 2017. 764\u2013773."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_5_1","DOI":"10.1109\/CVPR.2018.00808"},{"key":"e_1_3_2_1_6_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. (2019), 4171\u20134186. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. (2019), 4171\u20134186."},{"key":"e_1_3_2_1_7_1","volume-title":"CenterNet: Keypoint Triplets for Object Detection. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV","author":"Duan Kaiwen","year":"2019","unstructured":"Kaiwen Duan , Song Bai , Lingxi Xie , Honggang Qi , Qingming Huang , and Qi Tian . 2019 . CenterNet: Keypoint Triplets for Object Detection. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV 2019. 6568\u20136577. Kaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi, Qingming Huang, and Qi Tian. 2019. CenterNet: Keypoint Triplets for Object Detection. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV 2019. 6568\u20136577."},{"key":"e_1_3_2_1_8_1","volume-title":"Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. In 2014 IEEE Conference on Computer Vision and Pattern Recognition, CVPR","author":"Girshick B.","year":"2014","unstructured":"Ross\u00a0 B. Girshick , Jeff Donahue , Trevor Darrell , and Jitendra Malik . 2014 . Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. In 2014 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2014. 580\u2013587. Ross\u00a0B. Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. 2014. Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. In 2014 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2014. 580\u2013587."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_9_1","DOI":"10.1145\/3343031.3350943"},{"key":"e_1_3_2_1_10_1","volume-title":"Mask R-CNN. In IEEE International Conference on Computer Vision, ICCV","author":"He Kaiming","year":"2017","unstructured":"Kaiming He , Georgia Gkioxari , Piotr Doll\u00e1r , and Ross\u00a0 B. Girshick . 2017 . Mask R-CNN. In IEEE International Conference on Computer Vision, ICCV 2017. 2980\u20132988. Kaiming He, Georgia Gkioxari, Piotr Doll\u00e1r, and Ross\u00a0B. Girshick. 2017. Mask R-CNN. In IEEE International Conference on Computer Vision, ICCV 2017. 2980\u20132988."},{"key":"e_1_3_2_1_11_1","volume-title":"Modeling Relationships in Referential Expressions with Compositional Modular Networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR","author":"Hu Ronghang","year":"2017","unstructured":"Ronghang Hu , Marcus Rohrbach , Jacob Andreas , Trevor Darrell , and Kate Saenko . 2017 . Modeling Relationships in Referential Expressions with Compositional Modular Networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017. 4418\u20134427. Ronghang Hu, Marcus Rohrbach, Jacob Andreas, Trevor Darrell, and Kate Saenko. 2017. Modeling Relationships in Referential Expressions with Compositional Modular Networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017. 4418\u20134427."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_12_1","DOI":"10.1145\/3343031.3351065"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_13_1","DOI":"10.3115\/v1\/D14-1086"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_14_1","DOI":"10.1109\/CVPR42600.2020.01089"},{"key":"e_1_3_2_1_15_1","volume-title":"Feature Pyramid Networks for Object Detection. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR","author":"Lin Tsung-Yi","year":"2017","unstructured":"Tsung-Yi Lin , Piotr Doll\u00e1r , Ross\u00a0 B. Girshick , Kaiming He , Bharath Hariharan , and Serge\u00a0 J. Belongie . 2017 . Feature Pyramid Networks for Object Detection. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017. 936\u2013944. Tsung-Yi Lin, Piotr Doll\u00e1r, Ross\u00a0B. Girshick, Kaiming He, Bharath Hariharan, and Serge\u00a0J. Belongie. 2017. Feature Pyramid Networks for Object Detection. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017. 936\u2013944."},{"volume-title":"Computer Vision - ECCV 2014 - 13th European Conference(Lecture Notes in Computer Science, Vol.\u00a08693). 740\u2013755.","author":"Lin Tsung-Yi","unstructured":"Tsung-Yi Lin , Michael Maire , Serge\u00a0 J. Belongie , James Hays , Pietro Perona , Deva Ramanan , Piotr Doll\u00e1r , and C.\u00a0 Lawrence Zitnick . 2014. Microsoft COCO: Common Objects in Context . In Computer Vision - ECCV 2014 - 13th European Conference(Lecture Notes in Computer Science, Vol.\u00a08693). 740\u2013755. Tsung-Yi Lin, Michael Maire, Serge\u00a0J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C.\u00a0Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. In Computer Vision - ECCV 2014 - 13th European Conference(Lecture Notes in Computer Science, Vol.\u00a08693). 740\u2013755.","key":"e_1_3_2_1_16_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_17_1","DOI":"10.1109\/ICCV.2019.00477"},{"key":"e_1_3_2_1_18_1","volume-title":"SSD: Single Shot MultiBox Detector. In Computer Vision - ECCV 2016 - 14th European Conference(Lecture Notes in Computer Science, Vol.\u00a09905). 21\u201337.","author":"Liu Wei","year":"2016","unstructured":"Wei Liu , Dragomir Anguelov , Dumitru Erhan , Christian Szegedy , Scott\u00a0 E. Reed , Cheng-Yang Fu , and Alexander\u00a0 C. Berg . 2016 . SSD: Single Shot MultiBox Detector. In Computer Vision - ECCV 2016 - 14th European Conference(Lecture Notes in Computer Science, Vol.\u00a09905). 21\u201337. Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott\u00a0E. Reed, Cheng-Yang Fu, and Alexander\u00a0C. Berg. 2016. SSD: Single Shot MultiBox Detector. In Computer Vision - ECCV 2016 - 14th European Conference(Lecture Notes in Computer Science, Vol.\u00a09905). 21\u201337."},{"key":"e_1_3_2_1_19_1","volume-title":"Improving Referring Expression Grounding With Cross-Modal Attention-Guided Erasing. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019. 1950","author":"Liu Xihui","year":"2019","unstructured":"Xihui Liu , Zihao Wang , Jing Shao , Xiaogang Wang , and Hongsheng Li . 2019 . Improving Referring Expression Grounding With Cross-Modal Attention-Guided Erasing. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019. 1950 \u20131959. Xihui Liu, Zihao Wang, Jing Shao, Xiaogang Wang, and Hongsheng Li. 2019. Improving Referring Expression Grounding With Cross-Modal Attention-Guided Erasing. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019. 1950\u20131959."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_20_1","DOI":"10.1609\/aaai.v34i07.6833"},{"key":"e_1_3_2_1_21_1","volume-title":"Self-boosted Gesture Interactive System with ST-Net. In 2018 ACM Multimedia Conference on Multimedia Conference, MM","author":"Liu Zhengzhe","year":"2018","unstructured":"Zhengzhe Liu , Xiaojuan Qi , and Lei Pang . 2018 . Self-boosted Gesture Interactive System with ST-Net. In 2018 ACM Multimedia Conference on Multimedia Conference, MM 2018. 145\u2013153. Zhengzhe Liu, Xiaojuan Qi, and Lei Pang. 2018. Self-boosted Gesture Interactive System with ST-Net. In 2018 ACM Multimedia Conference on Multimedia Conference, MM 2018. 145\u2013153."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_22_1","DOI":"10.1145\/3343031.3350961"},{"key":"e_1_3_2_1_23_1","volume-title":"Generation and Comprehension of Unambiguous Object Descriptions. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR","author":"Mao Junhua","year":"2016","unstructured":"Junhua Mao , Jonathan Huang , Alexander Toshev , Oana Camburu , Alan\u00a0 L. Yuille , and Kevin Murphy . 2016 . Generation and Comprehension of Unambiguous Object Descriptions. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016. 11\u201320. Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan\u00a0L. Yuille, and Kevin Murphy. 2016. Generation and Comprehension of Unambiguous Object Descriptions. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016. 11\u201320."},{"volume-title":"Computer Vision - ECCV 2016 - 14th European Conference(Lecture Notes in Computer Science, Vol.\u00a09908). 792\u2013807.","author":"Nagaraja K.","unstructured":"Varun\u00a0 K. Nagaraja , Vlad\u00a0 I. Morariu , and Larry\u00a0 S. Davis . 2016. Modeling Context Between Objects for Referring Expression Understanding . In Computer Vision - ECCV 2016 - 14th European Conference(Lecture Notes in Computer Science, Vol.\u00a09908). 792\u2013807. Varun\u00a0K. Nagaraja, Vlad\u00a0I. Morariu, and Larry\u00a0S. Davis. 2016. Modeling Context Between Objects for Referring Expression Understanding. In Computer Vision - ECCV 2016 - 14th European Conference(Lecture Notes in Computer Science, Vol.\u00a09908). 792\u2013807.","key":"e_1_3_2_1_24_1"},{"key":"e_1_3_2_1_25_1","volume-title":"The 28th ACM International Conference on Multimedia, Virtual Event \/ Seattle. 4171\u20134180","author":"Qiu Heqian","year":"2020","unstructured":"Heqian Qiu , Hongliang Li , Qingbo Wu , Fanman Meng , Hengcan Shi , Taijin Zhao , and King\u00a0Ngi Ngan . 2020 . Language-Aware Fine-Grained Object Representation for Referring Expression Comprehension. In MM \u201920 : The 28th ACM International Conference on Multimedia, Virtual Event \/ Seattle. 4171\u20134180 . Heqian Qiu, Hongliang Li, Qingbo Wu, Fanman Meng, Hengcan Shi, Taijin Zhao, and King\u00a0Ngi Ngan. 2020. Language-Aware Fine-Grained Object Representation for Referring Expression Comprehension. In MM \u201920: The 28th ACM International Conference on Multimedia, Virtual Event \/ Seattle. 4171\u20134180."},{"unstructured":"Joseph Redmon and Ali Farhadi. 2018. YOLOv3: An Incremental Improvement. CoRR abs\/1804.02767(2018). arxiv:1804.02767  Joseph Redmon and Ali Farhadi. 2018. YOLOv3: An Incremental Improvement. CoRR abs\/1804.02767(2018). arxiv:1804.02767","key":"e_1_3_2_1_26_1"},{"unstructured":"Shaoqing Ren Kaiming He Ross\u00a0B. Girshick and Jian Sun. 2015. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. (2015) 91\u201399.  Shaoqing Ren Kaiming He Ross\u00a0B. Girshick and Jian Sun. 2015. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. (2015) 91\u201399.","key":"e_1_3_2_1_27_1"},{"key":"e_1_3_2_1_28_1","volume-title":"Amsterdam, The Netherlands","author":"Rohrbach Anna","year":"2016","unstructured":"Anna Rohrbach , Marcus Rohrbach , Ronghang Hu , Trevor Darrell , and Bernt Schiele . 2016 . Grounding of Textual Phrases in Images by Reconstruction. In Computer Vision - ECCV 2016 - 14th European Conference , Amsterdam, The Netherlands , October 11-14, 2016, Proceedings, Part I(Lecture Notes in Computer Science, Vol.\u00a09905). 817\u2013834. Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele. 2016. Grounding of Textual Phrases in Images by Reconstruction. In Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part I(Lecture Notes in Computer Science, Vol.\u00a09905). 817\u2013834."},{"key":"e_1_3_2_1_29_1","volume-title":"Zero-Shot Grounding of Objects From Natural Language Queries. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV","author":"Sadhu Arka","year":"2019","unstructured":"Arka Sadhu , Kan Chen , and Ram Nevatia . 2019 . Zero-Shot Grounding of Objects From Natural Language Queries. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV 2019. 4693\u20134702. Arka Sadhu, Kan Chen, and Ram Nevatia. 2019. Zero-Shot Grounding of Objects From Natural Language Queries. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV 2019. 4693\u20134702."},{"volume-title":"Computer Vision - ECCV 2018 - 15th European Conference(Lecture Notes in Computer Science, Vol.\u00a011210). 38\u201354.","author":"Shi Hengcan","unstructured":"Hengcan Shi , Hongliang Li , Fanman Meng , and Qingbo Wu. 2018. Key-Word-Aware Network for Referring Expression Image Segmentation . In Computer Vision - ECCV 2018 - 15th European Conference(Lecture Notes in Computer Science, Vol.\u00a011210). 38\u201354. Hengcan Shi, Hongliang Li, Fanman Meng, and Qingbo Wu. 2018. Key-Word-Aware Network for Referring Expression Image Segmentation. In Computer Vision - ECCV 2018 - 15th European Conference(Lecture Notes in Computer Science, Vol.\u00a011210). 38\u201354.","key":"e_1_3_2_1_30_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_31_1","DOI":"10.1109\/TMM.2020.2991504"},{"key":"e_1_3_2_1_32_1","volume-title":"Very Deep Convolutional Networks for Large-Scale Image Recognition. In 3rd International Conference on Learning Representations, ICLR","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman . 2015 . Very Deep Convolutional Networks for Large-Scale Image Recognition. In 3rd International Conference on Learning Representations, ICLR 2015. Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In 3rd International Conference on Learning Representations, ICLR 2015."},{"key":"e_1_3_2_1_33_1","volume-title":"Explore Multi-Step Reasoning in Video Question Answering. In 2018 ACM Multimedia Conference on Multimedia Conference, MM","author":"Song Xiaomeng","year":"2018","unstructured":"Xiaomeng Song , Yucheng Shi , Xin Chen , and Yahong Han . 2018 . Explore Multi-Step Reasoning in Video Question Answering. In 2018 ACM Multimedia Conference on Multimedia Conference, MM 2018. 239\u2013247. Xiaomeng Song, Yucheng Shi, Xin Chen, and Yahong Han. 2018. Explore Multi-Step Reasoning in Video Question Answering. In 2018 ACM Multimedia Conference on Multimedia Conference, MM 2018. 239\u2013247."},{"unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan\u00a0N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is All you Need. (2017) 5998\u20136008.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan\u00a0N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is All you Need. (2017) 5998\u20136008.","key":"e_1_3_2_1_34_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_35_1","DOI":"10.1109\/CVPR.2019.00206"},{"key":"e_1_3_2_1_36_1","volume-title":"Dynamic Graph Attention for Referring Expression Comprehension. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV","author":"Yang Sibei","year":"2019","unstructured":"Sibei Yang , Guanbin Li , and Yizhou Yu . 2019 . Dynamic Graph Attention for Referring Expression Comprehension. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV 2019. 4643\u20134652. Sibei Yang, Guanbin Li, and Yizhou Yu. 2019. Dynamic Graph Attention for Referring Expression Comprehension. In 2019 IEEE\/CVF International Conference on Computer Vision, ICCV 2019. 4643\u20134652."},{"key":"e_1_3_2_1_37_1","volume-title":"Graph-Structured Referring Expression Reasoning in the Wild. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR","author":"Yang Sibei","year":"2020","unstructured":"Sibei Yang , Guanbin Li , and Yizhou Yu . 2020 . Graph-Structured Referring Expression Reasoning in the Wild. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020. 9949\u20139958. Sibei Yang, Guanbin Li, and Yizhou Yu. 2020. Graph-Structured Referring Expression Reasoning in the Wild. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020. 9949\u20139958."},{"key":"e_1_3_2_1_38_1","volume-title":"Improving One-Stage Visual Grounding by Recursive Sub-query Construction. 12359","author":"Yang Zhengyuan","year":"2020","unstructured":"Zhengyuan Yang , Tianlang Chen , Liwei Wang , and Jiebo Luo . 2020. Improving One-Stage Visual Grounding by Recursive Sub-query Construction. 12359 ( 2020 ), 387\u2013404. Zhengyuan Yang, Tianlang Chen, Liwei Wang, and Jiebo Luo. 2020. Improving One-Stage Visual Grounding by Recursive Sub-query Construction. 12359 (2020), 387\u2013404."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_39_1","DOI":"10.1109\/ICCV.2019.00478"},{"key":"e_1_3_2_1_40_1","volume-title":"MAttNet: Modular Attention Network for Referring Expression Comprehension. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR","author":"Yu Licheng","year":"2018","unstructured":"Licheng Yu , Zhe Lin , Xiaohui Shen , Jimei Yang , Xin Lu , Mohit Bansal , and Tamara\u00a0 L. Berg . 2018 . MAttNet: Modular Attention Network for Referring Expression Comprehension. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018. 1307\u20131315. Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara\u00a0L. Berg. 2018. MAttNet: Modular Attention Network for Referring Expression Comprehension. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018. 1307\u20131315."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_41_1","DOI":"10.1109\/CVPR.2017.375"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_42_1","DOI":"10.1145\/3343031.3350932"},{"key":"e_1_3_2_1_43_1","volume-title":"Parallel Attention: A Unified Framework for Visual Object Discovery Through Dialogs and Queries. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR","author":"Zhuang Bohan","year":"2018","unstructured":"Bohan Zhuang , Qi Wu , Chunhua Shen , Ian\u00a0 D. Reid , and Anton van\u00a0den Hengel . 2018 . Parallel Attention: A Unified Framework for Visual Object Discovery Through Dialogs and Queries. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018. 4252\u20134261. Bohan Zhuang, Qi Wu, Chunhua Shen, Ian\u00a0D. Reid, and Anton van\u00a0den Hengel. 2018. Parallel Attention: A Unified Framework for Visual Object Discovery Through Dialogs and Queries. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018. 4252\u20134261."},{"key":"e_1_3_2_1_44_1","volume-title":"Edge Boxes: Locating Object Proposals from Edges. In Computer Vision - ECCV 2014 - 13th European Conference(Lecture Notes in Computer Science, Vol.\u00a08693). 391\u2013405.","author":"Zitnick Lawrence","year":"2014","unstructured":"C.\u00a0 Lawrence Zitnick and Piotr Doll\u00e1r . 2014 . Edge Boxes: Locating Object Proposals from Edges. In Computer Vision - ECCV 2014 - 13th European Conference(Lecture Notes in Computer Science, Vol.\u00a08693). 391\u2013405. C.\u00a0Lawrence Zitnick and Piotr Doll\u00e1r. 2014. Edge Boxes: Locating Object Proposals from Edges. In Computer Vision - ECCV 2014 - 13th European Conference(Lecture Notes in Computer Science, Vol.\u00a08693). 391\u2013405."}],"event":{"sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"acronym":"MMAsia '21","name":"MMAsia '21: ACM Multimedia Asia","location":"Gold Coast Australia"},"container-title":["ACM Multimedia Asia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3469877.3490592","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3469877.3490592","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:16Z","timestamp":1750188616000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3469877.3490592"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12]]},"references-count":44,"alternative-id":["10.1145\/3469877.3490592","10.1145\/3469877"],"URL":"https:\/\/doi.org\/10.1145\/3469877.3490592","relation":{},"subject":[],"published":{"date-parts":[[2021,12]]},"assertion":[{"value":"2022-01-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}