{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T17:17:06Z","timestamp":1777655826125,"version":"3.51.4"},"reference-count":77,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2023,5,9]],"date-time":"2023-05-09T00:00:00Z","timestamp":1683590400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,5,31]]},"abstract":"<jats:p>With the rapid development of information technology, image and text data have increased dramatically. Image and text matching techniques enable computers to understand information from both visual and text modalities and match them based on semantic content. Existing methods focus on visual and textual object co-occurrence statistics and learning coarse-level associations. However, the lack of intramodal semantic inference leads to the failure of fine-level association between modalities. Scene graphs can capture the interactions between visual and textual objects and model intramodal semantic associations, which are crucial for the understanding of scenes contained in images and text. In this article, we propose a novel scene graph semantic inference network (SGSIN) for image and text matching that effectively learns fine-level semantic information in vision and text to facilitate bridging cross-modal discrepancies. Specifically, we design two matching modules and construct scene graphs within each matching module for aggregating neighborhood information to refine the semantic representation of each object and achieve fine-level alignment of visual and textual modalities. We perform extended experiments in Flickr30K and MSCOCO and achieve state-of-the-art results, which validate the advantages of our proposed approach.<\/jats:p>","DOI":"10.1145\/3563390","type":"journal-article","created":{"date-parts":[[2022,9,14]],"date-time":"2022-09-14T13:30:14Z","timestamp":1663162214000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":32,"title":["Scene Graph Semantic Inference for Image and Text Matching"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2774-0511","authenticated-orcid":false,"given":"Jiaming","family":"Pei","sequence":"first","affiliation":[{"name":"School of Computer Science, University of Sydney, Sydney, NSW, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5575-0306","authenticated-orcid":false,"given":"Kaiyang","family":"Zhong","sequence":"additional","affiliation":[{"name":"School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics, Chengdu, Sichuan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2305-0492","authenticated-orcid":false,"given":"Zhi","family":"Yu","sequence":"additional","affiliation":[{"name":"School of Microelectronics and Communication Engineering, Chongqing University, Chongqing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9207-7424","authenticated-orcid":false,"given":"Lukun","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Intelligent equipment, Shandong University of Science and Technology, Qingdao, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3480-4851","authenticated-orcid":false,"given":"Kuruva","family":"Lakshmanna","sequence":"additional","affiliation":[{"name":"School of Information Technology and Engineering, Vellore Institute of Technology, Tamil Nadu, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,5,9]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-16-1395-1_17"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3447651"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV51458.2022.00254"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460426.3463615"},{"key":"e_1_3_1_8_2","article-title":"Structure-aware positional transformer for visible-infrared person re-identification","author":"Chen Cuiqun","year":"2022","unstructured":"Cuiqun Chen, Mang Ye, Meibin Qi, Jingjing Wu, Jianguo Jiang, and Chia-Wen Lin. 2022. Structure-aware positional transformer for visible-infrared person re-identification. IEEE Trans. Image Process. 31 (2022), 2352\u20132364.","journal-title":"IEEE Trans. Image Process."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00756"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2018.10.054"},{"issue":"1","key":"e_1_3_1_11_2","first-page":"1","article-title":"A new concept of electronic text based on semantic coding system for machine translation","volume":"21","author":"Eddine Meftah Mohammed Charaf","year":"2021","unstructured":"Meftah Mohammed Charaf Eddine. 2021. A new concept of electronic text based on semantic coding system for machine translation. ACM Trans. Asian Low-resour. Lang. Inf. 21, 1 (2021), 1\u201316.","journal-title":"ACM Trans. Asian Low-resour. Lang. Inf"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611976236.40"},{"key":"e_1_3_1_13_2","first-page":"935","volume-title":"Proceedings of the British Machine Vision Conference.","author":"Fartash Faghri","year":"2018","unstructured":"Faghri Fartash, D. Fleet, J. Kiros, and Sanja Fidler. 2018. VSE++: Improved visual semantic embeddings. In Proceedings of the British Machine Vision Conference.935\u2013943."},{"key":"e_1_3_1_14_2","article-title":"Devise: A deep visual-semantic embedding model","volume":"26","author":"Frome Andrea","year":"2013","unstructured":"Andrea Frome, Greg S. Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc\u2019Aurelio Ranzato, and Tomas Mikolov. 2013. Devise: A deep visual-semantic embedding model. Adv. Neural Inf. Process. Syst. 26 (2013).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00750"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3463031"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00645"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2883466"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00585"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2019.00311"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00133"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_3_1_23_2","article-title":"Deep fragment embeddings for bidirectional image sentence mapping","volume":"27","author":"Karpathy Andrej","year":"2014","unstructured":"Andrej Karpathy, Armand Joulin, and Li F. Fei-Fei. 2014. Deep fragment embeddings for bidirectional image sentence mapping. Adv. Neural Inf. Process. Syst. 27 (2014).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_1_24_2","unstructured":"Thomas N. Kipf and Max Welling. 2016. Semi-supervised classification with graph convolutional networks. International Conference on Learning Representations (ICLR\u201917) ."},{"key":"e_1_3_1_25_2","article-title":"Unifying visual-semantic embeddings with multimodal neural language models","author":"Kiros Ryan","year":"2014","unstructured":"Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel. 2014. Unifying visual-semantic embeddings with multimodal neural language models. arXiv preprint arXiv:1411.2539 (2014).","journal-title":"arXiv preprint arXiv:1411.2539"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299073"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_1_28_2","article-title":"Learning and integrating multi-level matching features for image-text retrieval","author":"Lan Hong","year":"2021","unstructured":"Hong Lan and Pufen Zhang. 2021. Learning and integrating multi-level matching features for image-text retrieval. IEEE Sig. Process. Lett. 29 (2021), 374\u2013378.","journal-title":"IEEE Sig. Process. Lett."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01225-0_13"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00475"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.209"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.142"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01280"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350869"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01093"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.442"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.301"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3432246"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240712"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.232"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.208"},{"key":"e_1_3_1_43_2","first-page":"3846","volume-title":"Proceedings of the International Joint Conferences on Artificial Intelligence","author":"Peng Yuxin","year":"2016","unstructured":"Yuxin Peng, Xin Huang, and Jinwei Qi. 2016. Cross-media shared representation by hierarchical learning with multiple deep networks. In Proceedings of the International Joint Conferences on Artificial Intelligence. 3846\u20133853."},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3284750"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.303"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462829"},{"key":"e_1_3_1_47_2","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"28","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. Adv. Neural Inf. Process. Syst. 28 (2015).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_1_48_2","article-title":"A simple neural network module for relational reasoning","volume":"30","author":"Santoro Adam","year":"2017","unstructured":"Adam Santoro, David Raposo, David G. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap. 2017. A simple neural network module for relational reasoning. Adv. Neural Inf. Process. Syst. 30 (2017).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W15-2812"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.344"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123326"},{"key":"e_1_3_1_52_2","first-page":"6596","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Wang Cheng","year":"2019","unstructured":"Cheng Wang and Mathias Niepert. 2019. State-regularized recurrent neural networks. In Proceedings of the International Conference on Machine Learning. PMLR, 6596\u20136606."},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2797921"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.541"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00206"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093614"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.5555\/3367471.3367568"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1037"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00586"},{"key":"e_1_3_1_60_2","volume-title":"Proceedings of the 22nd International Joint Conference on Artificial Intelligence","author":"Weston Jason","year":"2011","unstructured":"Jason Weston, Samy Bengio, and Nicolas Usunier. 2011. WSABIE: Scaling up to large vocabulary image annotation. In Proceedings of the 22nd International Joint Conference on Artificial Intelligence."},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00677"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jvcir.2021.103261"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.image.2021.116319"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350940"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2017.206"},{"issue":"1","key":"e_1_3_1_66_2","first-page":"1","article-title":"Event graph neural network for opinion target classification of microblog comments","volume":"21","author":"Xiang Yan","year":"2021","unstructured":"Yan Xiang, Zhengtao Yu, Junjun Guo, Yuxin Huang, and Yantuan Xian. 2021. Event graph neural network for opinion target classification of microblog comments. ACM Trans. Asian Low-resour. Lang. Inf. 21, 1 (2021), 1\u201313.","journal-title":"ACM Trans. Asian Low-resour. Lang. Inf"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.330"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298966"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01094"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_42"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00601"},{"key":"e_1_3_1_72_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00166"},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00611"},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475380"},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_42"},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01064"},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","DOI":"10.1145\/3383184"},{"key":"e_1_3_1_78_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_49"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3563390","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3563390","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:35Z","timestamp":1750182575000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3563390"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,9]]},"references-count":77,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2023,5,31]]}},"alternative-id":["10.1145\/3563390"],"URL":"https:\/\/doi.org\/10.1145\/3563390","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,9]]},"assertion":[{"value":"2022-04-06","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-09-08","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-05-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}