{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T23:17:57Z","timestamp":1784330277288,"version":"3.55.0"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2021,11,12]],"date-time":"2021-11-12T00:00:00Z","timestamp":1636675200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Intelligenza Artificiale per il Monitoraggio Visuale dei Siti Culturali","award":["CUP B15J19001040004"],"award-info":[{"award-number":["CUP B15J19001040004"]}]},{"name":"EC","award":["825619"],"award-info":[{"award-number":["825619"]}]},{"name":"AI4Media","award":["GA 951911"],"award-info":[{"award-number":["GA 951911"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2021,11,30]]},"abstract":"<jats:p>\n            Despite the evolution of deep-learning-based visual-textual processing systems, precise multi-modal matching remains a challenging task. In this work, we tackle the task of cross-modal retrieval through image-sentence matching based on word-region alignments, using supervision only at the global image-sentence level. Specifically, we present a novel approach called\n            <jats:italic>Transformer Encoder Reasoning and Alignment Network<\/jats:italic>\n            (TERAN). TERAN enforces a fine-grained match between the underlying components of images and sentences (i.e., image regions and words, respectively) to preserve the informative richness of both modalities. TERAN obtains state-of-the-art results on the image retrieval task on both MS-COCO and Flickr30k datasets. Moreover, on MS-COCO, it also outperforms current approaches on the sentence retrieval task.\n          <\/jats:p>\n          <jats:p>\n            Focusing on scalable cross-modal information retrieval, TERAN is designed to keep the visual and textual data pipelines well separated. Cross-attention links invalidate any chance to separately extract visual and textual features needed for the online search and the offline indexing steps in large-scale retrieval systems. In this respect, TERAN merges the information from the two domains only during the final alignment phase, immediately before the loss computation. We argue that the fine-grained alignments produced by TERAN pave the way toward the research for effective and efficient methods for large-scale cross-modal information retrieval. We compare the effectiveness of our approach against relevant state-of-the-art methods. On the MS-COCO 1K test set, we obtain an improvement of 5.7% and 3.5% respectively on the image and the sentence retrieval tasks on the Recall@1 metric. The code used for the experiments is publicly available on GitHub at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/mesnico\/TERAN\">https:\/\/github.com\/mesnico\/TERAN<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3451390","type":"journal-article","created":{"date-parts":[[2021,11,12]],"date-time":"2021-11-12T21:16:06Z","timestamp":1636751766000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":127,"title":["Fine-Grained Visual Textual Alignment for Cross-Modal Retrieval Using Transformer Encoders"],"prefix":"10.1145","volume":"17","author":[{"given":"Nicola","family":"Messina","sequence":"first","affiliation":[{"name":"ISTI-CNR, Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Giuseppe","family":"Amato","sequence":"additional","affiliation":[{"name":"ISTI-CNR, Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrea","family":"Esuli","sequence":"additional","affiliation":[{"name":"ISTI-CNR, Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fabrizio","family":"Falchi","sequence":"additional","affiliation":[{"name":"ISTI-CNR, Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Claudio","family":"Gennaro","sequence":"additional","affiliation":[{"name":"ISTI-CNR, Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"St\u00e9phane","family":"Marchand-Maillet","sequence":"additional","affiliation":[{"name":"VIPER Group\u2013University of Geneva, Geneva, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,11,12]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-017-9318-6"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01267"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","unstructured":"Yen-Chun Chen Linjie Li Licheng Yu Ahmed El Kholy Faisal Ahmed Zhe Gan Yu Cheng and Jingjing Liu. 2019. Uniter: Learning universal image-text representations. arXiv:1909.11740.","DOI":"10.1007\/978-3-030-58577-8_7"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00850"},{"key":"e_1_3_2_8_2","first-page":"4171","volume-title":"Proceedings of the 2019 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT\u201919)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT\u201919). 4171\u20134186."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.201"},{"key":"e_1_3_2_10_2","first-page":"12","volume-title":"Proceedings of the British Machine Vision Conference (BMVC\u201918)","author":"Faghri Fartash","year":"2018","unstructured":"Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2018. VSE++: Improving visual-semantic embeddings with hard negatives. In Proceedings of the British Machine Vision Conference (BMVC\u201918). 12."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00750"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.3390\/app10165516"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.93"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2882225"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00473"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00587"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.767"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00645"},{"key":"e_1_3_2_19_2","article-title":"Image and sentence matching via semantic concepts and order learning","author":"Huang Yan","year":"2018","unstructured":"Yan Huang, Qi Wu, Wei Wang, and Liang Wang. 2018. Image and sentence matching via semantic concepts and order learning. IEEE Transactions on Pattern Analysis and Machine Intelligence 42, 3 (2018), 636\u2013650.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_20_2","unstructured":"Zhicheng Huang Zhaoyang Zeng Bei Liu Dongmei Fu and Jianlong Fu. 2020. Pixel-BERT: Aligning image pixels with text by deep multi-modal transformers. arXiv:2004.00849."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2975594"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00585"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2020.2985716"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.325"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_3_2_26_2","first-page":"7482","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Kendall Alex","year":"2018","unstructured":"Alex Kendall, Yarin Gal, and Roberto Cipolla. 2018. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7482\u20137491."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299073"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01225-0_13"},{"key":"e_1_3_2_30_2","unstructured":"Kuang-Huei Lee Hamid Palangi Xi Chen Houdong Hu and Jianfeng Gao. 2019. Learning visual relation priors for image-text matching and image captioning with neural scene graph generators. arXiv:1909.09953."},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00475"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2896516"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_21"},{"key":"e_1_3_2_34_2","first-page":"Association for","volume-title":"Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out. Association for Computational Linguistics. 74\u201381."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_17"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350869"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01093"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.442"},{"key":"e_1_3_2_40_2","first-page":"13","volume-title":"Advances in Neural Information Processing Systems","author":"Lu Jiasen","year":"2019","unstructured":"Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Advances in Neural Information Processing Systems. 13\u201323."},{"key":"e_1_3_2_41_2","doi-asserted-by":"crossref","unstructured":"Sean MacAvaney Franco Maria Nardini Raffaele Perego Nicola Tonellotto Nazli Goharian and Ophir Frieder. 2020. Efficient document re-ranking for transformers by precomputing term representations. arXiv:2004.14255.","DOI":"10.1145\/3397271.3401093"},{"key":"e_1_3_2_42_2","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201918)","author":"Messina Nicola","year":"2018","unstructured":"Nicola Messina, Giuseppe Amato, Fabio Carrara, Fabrizio Falchi, and Claudio Gennaro. 2018. Learning relationship-aware visual features. In Proceedings of the European Conference on Computer Vision (ECCV\u201918)."},{"key":"e_1_3_2_43_2","first-page":"113","article-title":"Learning visual features for relational CBIR","author":"Messina Nicola","year":"2019","unstructured":"Nicola Messina, Giuseppe Amato, Fabio Carrara, Fabrizio Falchi, and Claudio Gennaro. 2019. Learning visual features for relational CBIR. International Journal of Multimedia Information Retrieval 9 (2019), 113\u2013124.","journal-title":"International Journal of Multimedia Information Retrieval"},{"key":"e_1_3_2_44_2","volume-title":"Proceedings of the International Conference on Pattern Recognition (ICPR\u201920)","author":"Messina Nicola","year":"2020","unstructured":"Nicola Messina, Fabrizio Falchi, Andrea Esuli, and Giuseppe Amato. 2020. Transformer reasoning network for image-text matching and retrieval. In Proceedings of the International Conference on Pattern Recognition (ICPR\u201920)."},{"key":"e_1_3_2_45_2","volume-title":"Proceedings of the 1st International Conference on Learning Representations (ICLR\u201913)","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. In Proceedings of the 1st International Conference on Learning Representations (ICLR\u201913)."},{"key":"e_1_3_2_46_2","unstructured":"Di Qi Lin Su Jia Song Edward Cui Taroon Bharti and Arun Sacheti. 2020. ImageBERT: Cross-modal pre-training with large-scale weak-supervised image-text data. arXiv:2001.07966."},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413961"},{"key":"e_1_3_2_48_2","first-page":"91","volume-title":"Advances in Neural Information Processing Systems","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems. 91\u201399."},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.131"},{"key":"e_1_3_2_50_2","unstructured":"Adam Santoro David Raposo David G. Barrett Mateusz Malinowski Razvan Pascanu Peter Battaglia and Timothy Lillicrap. 2017. A simple neural network module for relational reasoning. In Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS\u201917) . 4967\u20134976."},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00591"},{"key":"e_1_3_2_52_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Su Weijie","year":"2020","unstructured":"Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020. VL-BERT: Pre-training of generic visual-linguistic representations. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.344"},{"key":"e_1_3_2_54_2","first-page":"5998","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_2_55_2","volume-title":"Proceedings of the 4th International Conference on Learning Representations (ICLR\u201916)","author":"Vendrov Ivan","year":"2016","unstructured":"Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun. 2016. Order-embeddings of images and language. In Proceedings of the 4th International Conference on Learning Representations (ICLR\u201916)."},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240535"},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","unstructured":"Yaxiong Wang Hao Yang Xueming Qian Lin Ma Jing Lu Biao Li and Xin Fan. 2019. Position focused attention network for image-text matching. arXiv:1907.09748.","DOI":"10.24963\/ijcai.2019\/526"},{"key":"e_1_3_2_58_2","article-title":"Adversarial attentive multi-modal embedding learning for image-text matching","author":"Wei Kaimin","year":"2020","unstructured":"Kaimin Wei and Zhibo Zhou. 2020. Adversarial attentive multi-modal embedding learning for image-text matching. IEEE Access 8 (2020), 96237\u201396248.","journal-title":"IEEE Access"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01095"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350940"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.2967597"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_41"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01094"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_42"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00166"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.7005"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3451390","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3451390","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:17:30Z","timestamp":1750191450000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3451390"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,11,12]]},"references-count":65,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2021,11,30]]}},"alternative-id":["10.1145\/3451390"],"URL":"https:\/\/doi.org\/10.1145\/3451390","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,11,12]]},"assertion":[{"value":"2020-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-02-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-11-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}