{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T16:08:44Z","timestamp":1783094924677,"version":"3.54.6"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,2,27]],"date-time":"2023-02-27T00:00:00Z","timestamp":1677456000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U21B2024, 62202327"],"award-info":[{"award-number":["U21B2024, 62202327"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"crossref","award":["2022M712369"],"award-info":[{"award-number":["2022M712369"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,7,31]]},"abstract":"<jats:p>Image\u2013text retrieval is a vital task in computer vision and has received growing attention, since it connects cross-modality data. It comes with the critical challenges of learning unified representations and eliminating the large gap between visual and textual domains. Over the past few decades, although many works have made significant progress in image\u2013text retrieval, they are still confronted with the challenge of incomplete text descriptions of images, i.e., how to fully learn the correlations between relevant region\u2013word pairs with semantic diversity. In this article, we propose a novel semantic completion and filtration\u00a0(SCAF) method to alleviate the above issue. Specifically, the text semantic completion module is presented to generate a complete semantic description of an image using multi-view text descriptions, guiding the model to explore the correlations of relevant region\u2013word pairs fully. Meanwhile, the adaptive structural semantic matching module is presented to filter irrelevant region\u2013word pairs by considering the relevance score of each region\u2013word pair, which facilitates the model to focus on learning the relevance of matching pairs. Extensive experiments show that our SCAF outperforms the existing methods on Flickr30K and MSCOCO datasets, which demonstrates the superiority of our proposed method.<\/jats:p>","DOI":"10.1145\/3572844","type":"journal-article","created":{"date-parts":[[2022,11,23]],"date-time":"2022-11-23T11:54:56Z","timestamp":1669204496000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["Semantic Completion and Filtration for Image\u2013Text Retrieval"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8238-8226","authenticated-orcid":false,"given":"Song","family":"Yang","sequence":"first","affiliation":[{"name":"Tianjin University; China and also with the Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7129-1456","authenticated-orcid":false,"given":"Qiang","family":"Li","sequence":"additional","affiliation":[{"name":"Tianjin University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9609-6120","authenticated-orcid":false,"given":"Wenhui","family":"Li","sequence":"additional","affiliation":[{"name":"Tianjin University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2227-207X","authenticated-orcid":false,"given":"Xuan-Ya","family":"Li","sequence":"additional","affiliation":[{"name":"Baidu Inc., Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7113-7838","authenticated-orcid":false,"given":"Ran","family":"Jin","sequence":"additional","affiliation":[{"name":"Zhejiang Wanli University, Ningbo, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2499-0948","authenticated-orcid":false,"given":"Bo","family":"Lv","sequence":"additional","affiliation":[{"name":"The 30th Research Institute of China Electronics Technology Group Corporation, ChengDu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7274-4525","authenticated-orcid":false,"given":"Rui","family":"Wang","sequence":"additional","affiliation":[{"name":"The 30th Research Institute of China Electronics Technology Group Corporation, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5755-9145","authenticated-orcid":false,"given":"Anan","family":"Liu","sequence":"additional","affiliation":[{"name":"Tianjin University; China and also with the Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,2,27]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_1_3_2","article-title":"Learning to paraphrase: An unsupervised approach using multiple-sequence alignment","volume":"0304006","author":"Barzilay Regina","year":"2003","unstructured":"Regina Barzilay and Lillian Lee. 2003. Learning to paraphrase: An unsupervised approach using multiple-sequence alignment. CoRR cs.CL\/0304006 (2003).","journal-title":"CoRR"},{"key":"e_1_3_1_4_2","first-page":"12466","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Biten Ali Furkan","year":"2019","unstructured":"Ali Furkan Biten, Llu\u00eds G\u00f3mez, Mar\u00e7al Rusi\u00f1ol, and Dimosthenis Karatzas. 2019. Good news, everyone! Context driven entity-aware captioning for news images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 12466\u201312475."},{"key":"e_1_3_1_5_2","first-page":"12652","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Chen Hui","year":"2020","unstructured":"Hui Chen, Guiguang Ding, Xudong Liu, Zijia Lin, Ji Liu, and Jungong Han. 2020. IMRAM: Iterative matching with recurrent attention memory for cross-modal image-text retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 12652\u201312660."},{"key":"e_1_3_1_6_2","article-title":"Empirical evaluation of gated recurrent neural networks on sequence modeling","volume":"1412","author":"Chung Junyoung","year":"2014","unstructured":"Junyoung Chung, \u00c7aglar G\u00fcl\u00e7ehre, KyungHyun Cho, and Yoshua Bengio. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. CoRR abs\/1412.3555 (2014).","journal-title":"CoRR"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548206"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.5555\/1608858.1608859"},{"key":"e_1_3_1_9_2","first-page":"12","volume-title":"Proceedings of the British Machine Vision Conference","author":"Faghri Fartash","year":"2018","unstructured":"Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2018. VSE++: Improving visual-semantic embeddings with hard negatives. In Proceedings of the British Machine Vision Conference. 12."},{"key":"e_1_3_1_10_2","first-page":"6514","volume-title":"Proceedings of the Conference of the Association for Computational Linguistics","author":"Fan Zhihao","year":"2019","unstructured":"Zhihao Fan, Zhongyu Wei, Siyuan Wang, and Xuanjing Huang. 2019. Bridging by word: Image grounded vocabulary construction for visual captioning. In Proceedings of the Conference of the Association for Computational Linguistics. 6514\u20136524."},{"key":"e_1_3_1_11_2","first-page":"322","volume-title":"Proceedings of the International Conference on Computational Linguistics, Proceedings of the Conference","author":"Filippova Katja","year":"2010","unstructured":"Katja Filippova. 2010. Multi-Sentence compression: Finding shortest paths in word graphs. In Proceedings of the International Conference on Computational Linguistics, Proceedings of the Conference. 322\u2013330."},{"key":"e_1_3_1_12_2","first-page":"2121","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems","author":"Frome Andrea","year":"2013","unstructured":"Andrea Frome, Gregory S. Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc\u2019Aurelio Ranzato, and Tom\u00e1s Mikolov. 2013. DeViSE: A deep visual-semantic embedding model. In Proceedings of the Annual Conference on Neural Information Processing Systems. 2121\u20132129."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00750"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2882225"},{"key":"e_1_3_1_16_2","first-page":"7254","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Huang Yan","year":"2017","unstructured":"Yan Huang, Wei Wang, and Liang Wang. 2017. Instance-Aware image and sentence matching with selective multimodal LSTM. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7254\u20137262."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00645"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2598339"},{"key":"e_1_3_1_19_2","first-page":"1889","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems","author":"Karpathy Andrej","year":"2014","unstructured":"Andrej Karpathy, Armand Joulin, and Fei-Fei Li. 2014. Deep fragment embeddings for bidirectional image sentence mapping. In Proceedings of the Annual Conference on Neural Information Processing Systems. 1889\u20131897."},{"key":"e_1_3_1_20_2","article-title":"Unifying visual-semantic embeddings with multimodal neural language models","volume":"1411","author":"Kiros Ryan","year":"2014","unstructured":"Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel. 2014. Unifying visual-semantic embeddings with multimodal neural language models. CoRR abs\/1411.2539 (2014).","journal-title":"CoRR"},{"key":"e_1_3_1_21_2","article-title":"Analyzing sentence fusion in abstractive summarization","volume":"1910","author":"Lebanoff Logan","year":"2019","unstructured":"Logan Lebanoff, John Muchovej, Franck Dernoncourt, Doo Soon Kim, Seokhwan Kim, Walter Chang, and Fei Liu. 2019. Analyzing sentence fusion in abstractive summarization. CoRR abs\/1910.00203 (2019).","journal-title":"CoRR"},{"key":"e_1_3_1_22_2","first-page":"212","volume-title":"Proceedings of the European Conference on Computer Vision","volume":"11208","author":"Lee Kuang-Huei","year":"2018","unstructured":"Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. 2018. Stacked cross attention for image-text matching. In Proceedings of the European Conference on Computer Vision, Vol. 11208. 212\u2013228."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3158546"},{"key":"e_1_3_1_24_2","first-page":"12783","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Li Yongzhi","year":"2020","unstructured":"Yongzhi Li, Duo Zhang, and Yadong Mu. 2020. Visual-Semantic matching by exploring high-order attention and distraction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 12783\u201312792."},{"key":"e_1_3_1_25_2","first-page":"740","volume-title":"Proceedings of the European Conference on Computer Vision","volume":"8693","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In Proceedings of the European Conference on Computer Vision, Vol. 8693. 740\u2013755."},{"key":"e_1_3_1_26_2","first-page":"3","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Liu Chunxiao","year":"2019","unstructured":"Chunxiao Liu, Zhendong Mao, An-An Liu, Tianzhu Zhang, Bin Wang, and Yongdong Zhang. 2019. Focus your attention: A bidirectional focal attention network for image-text matching. In Proceedings of the ACM International Conference on Multimedia. 3\u201311."},{"key":"e_1_3_1_27_2","first-page":"10918","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Liu Chunxiao","year":"2020","unstructured":"Chunxiao Liu, Zhendong Mao, Tianzhu Zhang, Hongtao Xie, Bin Wang, and Yongdong Zhang. 2020. Graph structured network for image-text matching. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 10918\u201310927."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2020.3037661"},{"key":"e_1_3_1_29_2","unstructured":"Xuejing Liu Liang Li Shuhui Wang Zheng-Jun Zha Zechao Li Qi Tian and Qingming Huang. 2022. Entity-enhanced adaptive reconstruction network for weakly supervised referring expression grounding. arXiv:2207.08386. Retrieved from https:\/\/arxiv.org\/abs\/2207.08386."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2017.07.001"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.301"},{"issue":"1","key":"e_1_3_1_32_2","first-page":"105","article-title":"PaddlePaddle: An open-source deep learning platform from industrial practice","volume":"1","author":"Ma Yanjun","year":"2019","unstructured":"Yanjun Ma, Dianhai Yu, Tian Wu, and Haifeng Wang. 2019. PaddlePaddle: An open-source deep learning platform from industrial practice. Front. Data Comput. 1, 1 (2019), 105\u2013115.","journal-title":"Front. Data Comput."},{"key":"e_1_3_1_33_2","first-page":"2156","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Nam Hyeonseob","year":"2017","unstructured":"Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim. 2017. Dual attention networks for multimodal reasoning and matching. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2156\u20132164."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2020.2974877"},{"key":"e_1_3_1_35_2","first-page":"1899","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Niu Zhenxing","year":"2017","unstructured":"Zhenxing Niu, Mo Zhou, Le Wang, Xinbo Gao, and Gang Hua. 2017. Hierarchical multimodal LSTM for dense visual-semantic embedding. In Proceedings of the IEEE International Conference on Computer Vision. 1899\u20131907."},{"key":"e_1_3_1_36_2","unstructured":"Paddlepaddle. 2019. PaddlePaddle: An Easy-to-use Easy-to-learn Deep Learning Platform. Retrieved from http:\/\/www.paddlepaddle.org\/."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0965-7"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00160"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_1_40_2","first-page":"5813","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Sarafianos Nikolaos","year":"2019","unstructured":"Nikolaos Sarafianos, Xiang Xu, and Ioannis A. Kakadiaris. 2019. Adversarial representation learning for text-to-image matching. In Proceedings of the IEEE International Conference on Computer Vision. 5813\u20135823."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.499"},{"key":"e_1_3_1_42_2","first-page":"11572","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang Hao","year":"2019","unstructured":"Hao Wang, Doyen Sahoo, Chenghao Liu, Ee-Peng Lim, and Steven C. H. Hoi. 2019. Learning cross-modal embeddings with adversarial networks for cooking recipes and food images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 11572\u201311581."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00695"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2797921"},{"key":"e_1_3_1_45_2","first-page":"1497","volume-title":"Proceedings of the IEEE Winter Conference on Applications of Computer Vision","author":"Wang Sijin","year":"2020","unstructured":"Sijin Wang, Ruiping Wang, Ziwei Yao, Shiguang Shan, and Xilin Chen. 2020. Cross-modal scene graph matching for relationship-aware image-text retrieval. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision. 1497\u20131506."},{"key":"e_1_3_1_46_2","first-page":"3792","volume-title":"Proceedings of the ACM International Joint Conference on Artificial Intelligence","author":"Wang Yaxiong","year":"2019","unstructured":"Yaxiong Wang, Hao Yang, Xueming Qian, Lin Ma, Jing Lu, Biao Li, and Xin Fan. 2019. Position focused attention network for image-text matching. In Proceedings of the ACM International Joint Conference on Artificial Intelligence. 3792\u20133798."},{"key":"e_1_3_1_47_2","first-page":"5763","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Wang Zihao","year":"2019","unstructured":"Zihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng, Junjie Yan, Xiaogang Wang, and Jing Shao. 2019. CAMP: Cross-Modal adaptive message passing for text-image retrieval. In Proceedings of the IEEE International Conference on Computer Vision. 5763\u20135772."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2020.3017344"},{"issue":"8","key":"e_1_3_1_49_2","doi-asserted-by":"crossref","first-page":"3118","DOI":"10.1109\/TCSVT.2020.3036860","article-title":"Noise augmented double-stream graph convolutional networks for image captioning","volume":"31","author":"Wu Lingxiang","year":"2021","unstructured":"Lingxiang Wu, Min Xu, Lei Sang, Ting Yao, and Tao Mei. 2021. Noise augmented double-stream graph convolutional networks for image captioning. IEEE Trans. Circuits Syst. Video Technol. 31, 8 (2021), 3118\u20133127.","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.2985540"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11280-018-0541-x"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00611"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2020.3047095"},{"key":"e_1_3_1_54_2","first-page":"3533","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhang Qi","year":"2020","unstructured":"Qi Zhang, Zhen Lei, Zhaoxiang Zhang, and Stan Z. Li. 2020. Context-Aware attention network for image-text retrieval. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3533\u20133542."},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01174"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3572844","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3572844","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:08:09Z","timestamp":1750183689000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3572844"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,27]]},"references-count":54,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,7,31]]}},"alternative-id":["10.1145\/3572844"],"URL":"https:\/\/doi.org\/10.1145\/3572844","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,27]]},"assertion":[{"value":"2022-04-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-11-20","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-27","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}