{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:55:15Z","timestamp":1760147715219,"version":"build-2065373602"},"reference-count":31,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2023,2,25]],"date-time":"2023-02-25T00:00:00Z","timestamp":1677283200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Humanities and Social Sciences Project of Education Ministry","award":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"],"award-info":[{"award-number":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"]}]},{"DOI":"10.13039\/501100007129","name":"Natural Science Foundation of Shandong Province","doi-asserted-by":"publisher","award":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"],"award-info":[{"award-number":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"]}],"id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"publisher"}]},{"name":"NSFC-Zhejiang Joint Fund of the Integration of Informatization and Industrialization","award":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"],"award-info":[{"award-number":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"]}]},{"name":"Scientific Research Studio in Colleges and Universities of Ji\u2019nan City","award":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"],"award-info":[{"award-number":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"]}]},{"name":"Introduction and Education Plan of Young Creative Talents in Colleges and Universities of Shandong Province, Research Project of Undergraduate Teaching Reform in Shandong Province","award":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"],"award-info":[{"award-number":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"]}]},{"name":"Key Research and Development Project of Shandong Province","award":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"],"award-info":[{"award-number":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"]}]},{"name":"Innovation Team of Youth Innovation Science and Technology Plan in Colleges and Universities of Shandong Province","award":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"],"award-info":[{"award-number":["20YJA870013","ZR2019MF016","ZR2020MF037","U1909210","202228105","2021GXRC092","Z2020025","2019GSF109112","2020KJN007"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Image-text retrieval aims to search related results of one modality by querying another modality. As a fundamental and key problem in cross-modal retrieval, image-text retrieval is still a challenging problem owing to the complementary and imbalanced relationship between different modalities (i.e., Image and Text) and different granularities (i.e., Global-level and Local-level). However, existing works have not fully considered how to effectively mine and fuse the complementarities between images and texts at different granularities. Therefore, in this paper, we propose a hierarchical adaptive alignment network, whose contributions are as follows: (1) We propose a multi-level alignment network, which simultaneously mines global-level and local-level data, thereby enhancing the semantic association between images and texts. (2) We propose an adaptive weighted loss to flexibly optimize the image-text similarity with two stages in a unified framework. (3) We conduct extensive experiments on three public benchmark datasets (Corel 5K, Pascal Sentence, and Wiki) and compare them with eleven state-of-the-art methods. The experimental results thoroughly verify the effectiveness of our proposed method.<\/jats:p>","DOI":"10.3390\/s23052559","type":"journal-article","created":{"date-parts":[[2023,2,27]],"date-time":"2023-02-27T02:10:46Z","timestamp":1677463846000},"page":"2559","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["HAAN: Learning a Hierarchical Adaptive Alignment Network for Image-Text Retrieval"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7062-2660","authenticated-orcid":false,"given":"Shuhuai","family":"Wang","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Shandong University of Finance and Economics, Jinan 250014, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6846-5782","authenticated-orcid":false,"given":"Zheng","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Shandong University of Finance and Economics, Jinan 250014, China"},{"name":"Shandong Provincial Key Laboratory of Digital Media Technology, Jinan 250014, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0290-3399","authenticated-orcid":false,"given":"Xinlei","family":"Pei","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Shandong University of Finance and Economics, Jinan 250014, China"},{"name":"Shandong Provincial Key Laboratory of Digital Media Technology, Jinan 250014, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2293-2840","authenticated-orcid":false,"given":"Junhao","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Shandong University of Finance and Economics, Jinan 250014, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,2,25]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"2372","DOI":"10.1109\/TCSVT.2017.2705068","article-title":"An overview of cross-media retrieval: Concepts, methodologies, benchmarks, and challenges","volume":"28","author":"Peng","year":"2017","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_2","first-page":"423","article-title":"Multimodal machine learning: A survey and taxonomy","volume":"41","author":"Ahuja","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Gu, J., Cai, J., Joty, S.R., Niu, L., and Wang, G. (2018, January 18\u201322). Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00750"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"394","DOI":"10.1109\/TPAMI.2018.2797921","article-title":"Learning two-branch neural networks for image-text matching tasks","volume":"41","author":"Wang","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_5","unstructured":"Faghri, F., Fleet, D.J., Kiros, J.R., and Fidler, S. (2017). Vse++: Improving visual-semantic embeddings with hard negatives. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"2427","DOI":"10.1109\/TCSVT.2020.3017344","article-title":"CMPD: Using cross memory network with pair discrimination for image-text retrieval","volume":"31","author":"Wen","year":"2020","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Lee, K.H., Chen, X., Hua, G., Hu, H., and He, X. (2018, January 8\u201314). Stacked cross attention for image-text matching. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01225-0_13"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Liu, C., Mao, Z., Liu, A.A., Zhang, T., Wang, B., and Zhang, Y. (2019, January 21\u201325). Focus Your Attention: A Bidirectional Focal Attention Network for Image-Text Matching. Proceedings of the 27th ACM International Conference on Multimedia, Nice, France.","DOI":"10.1145\/3343031.3350869"},{"key":"ref_9","unstructured":"Li, K., Zhang, Y., Li, K., Li, Y., and Fu, Y. (November, January 27). Visual semantic reasoning for image-text matching. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhang, K., Mao, Z., Wang, Q., and Zhang, Y. (2022, January 18\u201324). Negative-Aware Attention Framework for Image-Text Matching. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01521"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Chen, H., Ding, G., Liu, X., Lin, Z., Liu, J., and Han, J. (2020, January 13\u201319). Imram: Iterative matching with recurrent attention memory for cross-modal image-text retrieval. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01267"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Schroff, F., Kalenichenko, D., and Philbin, J. (2015, January 7\u201312). Facenet: A unified embedding for face recognition and clustering. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1234","DOI":"10.1109\/TMM.2016.2646180","article-title":"Deep coupled metric learning for cross-modal matching","volume":"19","author":"Liong","year":"2016","journal-title":"IEEE Trans. Multimed."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"657","DOI":"10.1007\/s11280-018-0541-x","article-title":"Deep adversarial metric learning for cross-modal retrieval","volume":"22","author":"Xu","year":"2019","journal-title":"World Wide Web"},{"key":"ref_15","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Duygulu, P., Barnard, K., de Freitas, J.F., and Forsyth, D.A. (2002, January 28\u201331). Object recognition as machine translation: Learning a lexicon for a fixed image vocabulary. Proceedings of the Computer Vision\u2014ECCV 2002: 7th European Conference on Computer Vision, Copenhagen, Denmark.","DOI":"10.1007\/3-540-47979-1_7"},{"key":"ref_17","unstructured":"Rashtchian, C., Young, P., Hodosh, M., and Hockenmaier, J. (2010). Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon\u2019s Mechanical Turk, NAACL."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Rasiwasia, N., Costa Pereira, J., Coviello, E., Doyle, G., Lanckriet, G.R., Levy, R., and Vasconcelos, N. (2010, January 25\u201329). A new approach to cross-modal multimedia retrieval. Proceedings of the 18th ACM international Conference on Multimedia, Firenze, Italy.","DOI":"10.1145\/1873951.1873987"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Chen, H., Ding, G., Lin, Z., Zhao, S., and Han, J. (2019, January 21\u201325). Cross-modal image-text retrieval with semantic consistency. Proceedings of the 27th ACM International Conference on Multimedia, Nice, France.","DOI":"10.1145\/3343031.3351055"},{"key":"ref_20","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_21","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 7\u20139). Batch normalization: Accelerating deep network training by reducing internal covariate shift. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_22","unstructured":"Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017, January 4\u20139). Automatic differentiation in pytorch. Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"965","DOI":"10.1109\/TCSVT.2013.2276704","article-title":"Learning cross-media joint representation with sparse and semisupervised regularization","volume":"24","author":"Zhai","year":"2013","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Ballan, L., Uricchio, T., Seidenari, L., and Del Bimbo, A. (2014, January 1\u20134). A cross-media model for automatic image annotation. Proceedings of the International Conference on Multimedia Retrieval, Glasgow, UK.","DOI":"10.1145\/2578726.2578728"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"2010","DOI":"10.1109\/TPAMI.2015.2505311","article-title":"Joint feature selection and subspace learning for cross-modal retrieval","volume":"38","author":"Wang","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_26","first-page":"1247","article-title":"Deep canonical correlation analysis","volume":"28","author":"Andrew","year":"2013","journal-title":"Proc. Int. Conf. Mach. Learn."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"2728","DOI":"10.1109\/TIP.2019.2952085","article-title":"MAVA: Multi-level adaptive visual-textual alignment by cross-media bi-attention mechanism","volume":"29","author":"Peng","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_28","first-page":"1218","article-title":"Similarity reasoning and filtration for image-text matching","volume":"35","author":"Diao","year":"2021","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Liu, Y., Wu, J., Qu, L., Gan, T., Yin, J., and Nie, L. (2022). Self-supervised correlation learning for cross-modal retrieval. IEEE Trans. Multimed., e3152086.","DOI":"10.1109\/TMM.2022.3152086"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3499027","article-title":"Cross-modal graph matching network for image-text retrieval","volume":"18","author":"Cheng","year":"2022","journal-title":"ACM Trans. Multimed. Comput. Commun. Appl."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"641","DOI":"10.1109\/TPAMI.2022.3148470","article-title":"Image-Text Embedding Learning via Visual and Textual Semantic Reasoning","volume":"45","author":"Li","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/5\/2559\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:42:38Z","timestamp":1760121758000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/5\/2559"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,25]]},"references-count":31,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["s23052559"],"URL":"https:\/\/doi.org\/10.3390\/s23052559","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2023,2,25]]}}}