{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T16:41:58Z","timestamp":1783096918342,"version":"3.54.6"},"publisher-location":"New York, NY, USA","reference-count":69,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,8,24]],"date-time":"2021-08-24T00:00:00Z","timestamp":1629763200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Natural Science Foundation of China Projects","award":["U1936213"],"award-info":[{"award-number":["U1936213"]}]},{"DOI":"10.13039\/100012543","name":"Shanghai Science and Technology Development Foundation","doi-asserted-by":"publisher","award":["19511121204"],"award-info":[{"award-number":["19511121204"]}],"id":[{"id":"10.13039\/100012543","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,8,24]]},"DOI":"10.1145\/3460426.3463650","type":"proceedings-article","created":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T22:50:29Z","timestamp":1630536629000},"page":"82-91","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["HSGMP"],"prefix":"10.1145","author":[{"given":"Yu","family":"Duan","sequence":"first","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yun","family":"Xiong","sequence":"additional","affiliation":[{"name":"Fudan University &amp; Shanghai Institute for Advanced Communication and Data Science, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuwei","family":"Fu","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yangyong","family":"Zhu","sequence":"additional","affiliation":[{"name":"Fudan University &amp; Shanghai Institute for Advanced Communication and Data Science, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,9]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3289600.3290967"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6631"},{"key":"e_1_3_2_1_5_1","volume-title":"Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325","author":"Chen Xinlei","year":"2015","unstructured":"Xinlei Chen , Hao Fang , Tsung-Yi Lin , Ramakrishna Vedantam , Saurabh Gupta , Piotr Doll\u00e1r , and C Lawrence Zitnick . 2015. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325 ( 2015 ). Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Doll\u00e1r, and C Lawrence Zitnick. 2015. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325 (2015)."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.352"},{"key":"e_1_3_2_1_7_1","unstructured":"Andrea Detommaso. 2020. Semantics-Aware VQA a scene-graph-based approach to enable commonsense reasoning . Thesis.  Andrea Detommaso. 2020. Semantics-Aware VQA a scene-graph-based approach to enable commonsense reasoning . Thesis."},{"key":"e_1_3_2_1_8_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098036"},{"key":"e_1_3_2_1_10_1","volume-title":"Jamie Ryan Kiros, and Sanja Fidler","author":"Faghri Fartash","year":"2017","unstructured":"Fartash Faghri , David J Fleet , Jamie Ryan Kiros, and Sanja Fidler . 2017 . Vse Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2017. Vse"},{"key":"e_1_3_2_1_11_1","volume-title":"Improving visual-semantic embeddings with hard negatives. arXiv preprint arXiv:1707.05612","year":"2017","unstructured":": Improving visual-semantic embeddings with hard negatives. arXiv preprint arXiv:1707.05612 ( 2017 ). : Improving visual-semantic embeddings with hard negatives. arXiv preprint arXiv:1707.05612 (2017)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.476"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00750"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3372278.3390709"},{"key":"e_1_3_2_1_15_1","volume-title":"Scene Graph Reasoning for Visual Question Answering. arXiv preprint arXiv:2007.01072","author":"Hildebrandt Marcel","year":"2020","unstructured":"Marcel Hildebrandt , Hang Li , Rajat Koner , Volker Tresp , and Stephan G\u00fcnnemann . 2020. Scene Graph Reasoning for Visual Question Answering. arXiv preprint arXiv:2007.01072 ( 2020 ). Marcel Hildebrandt, Hang Li, Rajat Koner, Volker Tresp, and Stephan G\u00fcnnemann. 2020. Scene Graph Reasoning for Visual Question Answering. arXiv preprint arXiv:2007.01072 (2020)."},{"key":"e_1_3_2_1_16_1","volume-title":"Sofie Van Landeghem, and Adriane Boyd","author":"Honnibal Matthew","year":"2020","unstructured":"Matthew Honnibal , Ines Montani , Sofie Van Landeghem, and Adriane Boyd . 2020 . spaCy: Industrial-strength Natural Language Processing in Python . https:\/\/doi.org\/10.5281\/zenodo.1212303 10.5281\/zenodo.1212303 Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020. spaCy: Industrial-strength Natural Language Processing in Python. https:\/\/doi.org\/10.5281\/zenodo.1212303"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219965"},{"key":"e_1_3_2_1_18_1","volume-title":"Provable benefit of orthogonal initialization in optimizing deep linear networks. arXiv preprint arXiv:2001.05992","author":"Hu Wei","year":"2020","unstructured":"Wei Hu , Lechao Xiao , and Jeffrey Pennington . 2020. Provable benefit of orthogonal initialization in optimizing deep linear networks. arXiv preprint arXiv:2001.05992 ( 2020 ). Wei Hu, Lechao Xiao, and Jeffrey Pennington. 2020. Provable benefit of orthogonal initialization in optimizing deep linear networks. arXiv preprint arXiv:2001.05992 (2020)."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.767"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00645"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00686"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00133"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_3_2_1_24_1","volume-title":"Unifying visual-semantic embeddings with multimodal neural language models. arXiv preprint arXiv:1411.2539","author":"Kiros Ryan","year":"2014","unstructured":"Ryan Kiros , Ruslan Salakhutdinov , and Richard S Zemel . 2014. Unifying visual-semantic embeddings with multimodal neural language models. arXiv preprint arXiv:1411.2539 ( 2014 ). Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel. 2014. Unifying visual-semantic embeddings with multimodal neural language models. arXiv preprint arXiv:1411.2539 (2014)."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2396761.2398474"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_2_1_27_1","volume-title":"International Conference on Machine Learning . PMLR, 3734--3743","author":"Lee Junhyun","year":"2019","unstructured":"Junhyun Lee , Inyeop Lee , and Jaewoo Kang . 2019 . Self-attention graph pooling . In International Conference on Machine Learning . PMLR, 3734--3743 . Junhyun Lee, Inyeop Lee, and Jaewoo Kang. 2019. Self-attention graph pooling. In International Conference on Machine Learning . PMLR, 3734--3743."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01225-0_13"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00475"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.142"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.442"},{"key":"e_1_3_2_1_32_1","volume-title":"Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025","author":"Luong Minh-Thang","year":"2015","unstructured":"Minh-Thang Luong , Hieu Pham , and Christopher D Manning . 2015. Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025 ( 2015 ). Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015. Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025 (2015)."},{"key":"e_1_3_2_1_33_1","volume-title":"A Faster Pytorch Implementation of Faster R-CNN. https:\/\/github.com\/jwyang\/faster-rcnn.pytorch","author":"Parikh Jianwei Yang","year":"2017","unstructured":"Jianwei Yang Parikh , Jiasen Lu , Dhruv Batra , and Devi. 2017. A Faster Pytorch Implementation of Faster R-CNN. https:\/\/github.com\/jwyang\/faster-rcnn.pytorch ( 2017 ). Jianwei Yang Parikh, Jiasen Lu, Dhruv Batra, and Devi. 2017. A Faster Pytorch Implementation of Faster R-CNN. https:\/\/github.com\/jwyang\/faster-rcnn.pytorch (2017)."},{"key":"e_1_3_2_1_34_1","volume-title":"Faster r-cnn: Towards real-time object detection with region proposal networks. arXiv preprint arXiv:1506.01497","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren , Kaiming He , Ross Girshick , and Jian Sun . 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. arXiv preprint arXiv:1506.01497 ( 2015 ). Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. arXiv preprint arXiv:1506.01497 (2015)."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W15-2812"},{"key":"e_1_3_2_1_36_1","volume-title":"Workshop on Vision and Language (VL15)","author":"Schuster Sebastian","unstructured":"Sebastian Schuster , Ranjay Krishna , Angel Chang , Li Fei-Fei , and Christopher D. Manning . 2015b. Generating Semantically Precise Scene Graphs from Textual Descriptions for Improved Image Retrieval . In Workshop on Vision and Language (VL15) . Association for Computational Linguistics, Lisbon, Portugal. Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D. Manning. 2015b. Generating Semantically Precise Scene Graphs from Textual Descriptions for Improved Image Retrieval. In Workshop on Vision and Language (VL15) . Association for Computational Linguistics, Lisbon, Portugal."},{"key":"e_1_3_2_1_37_1","unstructured":"Richard Socher Danqi Chen Christopher D Manning and Andrew Ng. 2013. Reasoning with neural tensor networks for knowledge base completion. In Advances in neural information processing systems. Citeseer 926--934.  Richard Socher Danqi Chen Christopher D Manning and Andrew Ng. 2013. Reasoning with neural tensor networks for knowledge base completion. In Advances in neural information processing systems. Citeseer 926--934."},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00208"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/3402707.3402736"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/1557019.1557107"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00377"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00678"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.344"},{"key":"e_1_3_2_1_44_1","volume-title":"Attention is all you need. arXiv preprint arXiv:1706.03762","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , Lukasz Kaiser , and Illia Polosukhin . 2017. Attention is all you need. arXiv preprint arXiv:1706.03762 ( 2017 ). Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. arXiv preprint arXiv:1706.03762 (2017)."},{"key":"e_1_3_2_1_45_1","volume-title":"Graph attention networks. arXiv preprint arXiv:1710.10903","author":"Cucurull Guillem","year":"2017","unstructured":"Guillem Cucurull , Arantxa Casanova , Adriana Romero , Pietro Lio , and Yoshua Bengio . 2017. Graph attention networks. arXiv preprint arXiv:1710.10903 ( 2017 ). Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2017. Graph attention networks. arXiv preprint arXiv:1710.10903 (2017)."},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123326"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58520-4_15"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.541"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093614"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313562"},{"key":"e_1_3_2_1_51_1","volume-title":"Scene graph parsing as dependency parsing. arXiv preprint arXiv:1803.09189","author":"Wang Yu-Siang","year":"2018","unstructured":"Yu-Siang Wang , Chenxi Liu , Xiaohui Zeng , and Alan Yuille . 2018. Scene graph parsing as dependency parsing. arXiv preprint arXiv:1803.09189 ( 2018 ). Yu-Siang Wang, Chenxi Liu, Xiaohui Zeng, and Alan Yuille. 2018. Scene graph parsing as dependency parsing. arXiv preprint arXiv:1803.09189 (2018)."},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00586"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01095"},{"key":"e_1_3_2_1_54_1","volume-title":"Linknet: Relational embedding for scene graph. arXiv preprint arXiv:1811.06410","author":"Woo Sanghyun","year":"2018","unstructured":"Sanghyun Woo , Dahun Kim , Donghyeon Cho , and In So Kweon . 2018 . Linknet: Relational embedding for scene graph. arXiv preprint arXiv:1811.06410 (2018). Sanghyun Woo, Dahun Kim, Donghyeon Cho, and In So Kweon. 2018. Linknet: Relational embedding for scene graph. arXiv preprint arXiv:1811.06410 (2018)."},{"key":"e_1_3_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.330"},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00143"},{"key":"e_1_3_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_41"},{"key":"e_1_3_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01094"},{"key":"e_1_3_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_42"},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01219-9_20"},{"key":"e_1_3_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.503"},{"key":"e_1_3_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00166"},{"key":"e_1_3_2_1_63_1","volume-title":"Heterogeneous graph learning for visual commonsense reasoning. arXiv preprint arXiv:1910.11475","author":"Yu Weijiang","year":"2019","unstructured":"Weijiang Yu , Jingwen Zhou , Weihao Yu , Xiaodan Liang , and Nong Xiao . 2019. Heterogeneous graph learning for visual commonsense reasoning. arXiv preprint arXiv:1910.11475 ( 2019 ). Weijiang Yu, Jingwen Zhou, Weihao Yu, Xiaodan Liang, and Nong Xiao. 2019. Heterogeneous graph learning for visual commonsense reasoning. arXiv preprint arXiv:1910.11475 (2019)."},{"key":"e_1_3_2_1_64_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5831--5840","author":"Zellers Rowan","year":"2014","unstructured":"Rowan Zellers , Mark Yatskar , Sam Thomson , and Yejin Choi . 2014 . Neural motifs: Scene graph parsing with global context . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5831--5840 . Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi. 2014. Neural motifs: Scene graph parsing with global context. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5831--5840."},{"key":"e_1_3_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-93037-4_16"},{"key":"e_1_3_2_1_66_1","volume-title":"2020 b. Resnest: Split-attention networks. arXiv preprint arXiv:2004.08955","author":"Zhang Hang","year":"2020","unstructured":"Hang Zhang , Chongruo Wu , Zhongyue Zhang , Yi Zhu , Zhi Zhang , Haibin Lin , Yue Sun , Tong He , Jonas Mueller , and R Manmatha . 2020 b. Resnest: Split-attention networks. arXiv preprint arXiv:2004.08955 ( 2020 ). Hang Zhang, Chongruo Wu, Zhongyue Zhang, Yi Zhu, Zhi Zhang, Haibin Lin, Yue Sun, Tong He, Jonas Mueller, and R Manmatha. 2020 b. Resnest: Split-attention networks. arXiv preprint arXiv:2004.08955 (2020)."},{"key":"e_1_3_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00359"},{"key":"e_1_3_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_42"},{"key":"e_1_3_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58568-6_13"}],"event":{"name":"ICMR '21: International Conference on Multimedia Retrieval","location":"Taipei Taiwan","acronym":"ICMR '21","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 2021 International Conference on Multimedia Retrieval"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3460426.3463650","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3460426.3463650","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:17:04Z","timestamp":1750191424000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3460426.3463650"}},"subtitle":["Heterogeneous Scene Graph Message Passing for Cross-modal Retrieval"],"short-title":[],"issued":{"date-parts":[[2021,8,24]]},"references-count":69,"alternative-id":["10.1145\/3460426.3463650","10.1145\/3460426"],"URL":"https:\/\/doi.org\/10.1145\/3460426.3463650","relation":{},"subject":[],"published":{"date-parts":[[2021,8,24]]},"assertion":[{"value":"2021-09-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}