{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T04:37:04Z","timestamp":1777696624098,"version":"3.51.4"},"reference-count":34,"publisher":"SAGE Publications","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IDA"],"published-print":{"date-parts":[[2024,5,28]]},"abstract":"<jats:p>There are two mainstream strategies for image-text matching at present. The one, termed as joint embedding learning, aims to model the semantic information of both image and sentence in a shared feature subspace, which facilitates the measurement of semantic similarity but only focuses on global alignment relationship. To explore the local semantic relationship more fully, the other one, termed as metric learning, aims to learn a complex similarity function to directly output score of each image-text pair. However, it significantly suffers from more computation burden at retrieval stage. In this paper, we propose a hierarchically joint embedding model to incorporate the local semantic relationship into a joint embedding learning framework. The proposed method learns the shared local and global embedding spaces simultaneously, and models the joint local embedding space with respect to specific local similarity labels which are easy to access from the lexical information of corpus. Unlike the methods based on metric learning, we can prepare the fixed representations of both images and sentences by concatenating the normalized local and global representations, which makes it feasible to perform the efficient retrieval. And experiments show that the proposed model can achieve competitive performance when compared to the existing joint embedding learning models on two publicly available datasets Flickr30k and MS-COCO.<\/jats:p>","DOI":"10.3233\/ida-230214","type":"journal-article","created":{"date-parts":[[2023,9,26]],"date-time":"2023-09-26T13:37:54Z","timestamp":1695735474000},"page":"647-665","source":"Crossref","is-referenced-by-count":1,"title":["Learning hierarchical embedding space for image-text matching"],"prefix":"10.1177","volume":"28","author":[{"given":"Hao","family":"Sun","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaolin","family":"Qin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaojing","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","reference":[{"key":"10.3233\/IDA-230214_ref1","doi-asserted-by":"crossref","unstructured":"G. Li, N. Duan, Y. Fang, M. Gong and D. Jiang, Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol.\u00a034, 2020, pp.\u00a011336\u201311344.","DOI":"10.1609\/aaai.v34i07.6795"},{"key":"10.3233\/IDA-230214_ref2","doi-asserted-by":"crossref","unstructured":"W. Hong, K. Ji, J. Liu, J. Wang, J. Chen and W. Chu, GilBERT: Generative Vision-Language Pre-Training for Image-Text Retrieval, in: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2021, pp.\u00a01379\u20131388.","DOI":"10.1145\/3404835.3462838"},{"key":"10.3233\/IDA-230214_ref3","doi-asserted-by":"crossref","unstructured":"K. Li, Y. Zhang, K. Li, Y. Li and Y. Fu, Visual semantic reasoning for image-text matching, in: Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2019, pp.\u00a04654\u20134662.","DOI":"10.1109\/ICCV.2019.00475"},{"key":"10.3233\/IDA-230214_ref5","doi-asserted-by":"crossref","unstructured":"Y. Zhang and H. Lu, Deep cross-modal projection learning for image-text matching, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp.\u00a0686\u2013701.","DOI":"10.1007\/978-3-030-01246-5_42"},{"key":"10.3233\/IDA-230214_ref6","doi-asserted-by":"crossref","unstructured":"J. Gu, J. Cai, S.R. Joty, L. Niu and G. Wang, Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp.\u00a07181\u20137189.","DOI":"10.1109\/CVPR.2018.00750"},{"key":"10.3233\/IDA-230214_ref7","doi-asserted-by":"crossref","unstructured":"Y. Huang, Q. Wu, C. Song and L. Wang, Learning semantic concepts and order for image and sentence matching, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp.\u00a06163\u20136171.","DOI":"10.1109\/CVPR.2018.00645"},{"key":"10.3233\/IDA-230214_ref8","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.442"},{"key":"10.3233\/IDA-230214_ref9","doi-asserted-by":"crossref","unstructured":"N. Sarafianos, X. Xu and I.A. Kakadiaris, Adversarial representation learning for text-to-image matching, in: Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2019, pp.\u00a05814\u20135824.","DOI":"10.1109\/ICCV.2019.00591"},{"key":"10.3233\/IDA-230214_ref11","doi-asserted-by":"crossref","unstructured":"D. Semedo and J. Magalh\u00e3es, Cross-Modal Subspace Learning with Scheduled Adaptive Margin Constraints, in: Proceedings of the 27th ACM International Conference on Multimedia, 2019, pp.\u00a075\u201383.","DOI":"10.1145\/3343031.3351030"},{"issue":"2","key":"10.3233\/IDA-230214_ref12","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3383184","article-title":"Dual-path convolutional image-text embeddings with instance loss","volume":"16","author":"Zheng","year":"2020","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)"},{"issue":"2","key":"10.3233\/IDA-230214_ref13","doi-asserted-by":"crossref","first-page":"394","DOI":"10.1109\/TPAMI.2018.2797921","article-title":"Learning two-branch neural networks for image-text matching tasks","volume":"41","author":"Wang","year":"2018","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"10.3233\/IDA-230214_ref14","doi-asserted-by":"crossref","unstructured":"A. Eisenschtat and L. Wolf, Linking image and text with 2-way nets, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp.\u00a04601\u20134611.","DOI":"10.1109\/CVPR.2017.201"},{"key":"10.3233\/IDA-230214_ref15","doi-asserted-by":"crossref","unstructured":"P. Hu, L. Zhen, D. Peng and P. Liu, Scalable deep multimodal learning for cross-modal retrieval, in: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, 2019, pp.\u00a0635\u2013644.","DOI":"10.1145\/3331184.3331213"},{"key":"10.3233\/IDA-230214_ref16","doi-asserted-by":"crossref","unstructured":"T. Yu, Y. Yang, Y. Li, L. Liu, H. Fei and P. Li, Heterogeneous attention network for effective and efficient cross-modal retrieval, in: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2021, pp.\u00a01146\u20131156.","DOI":"10.1145\/3404835.3462924"},{"key":"10.3233\/IDA-230214_ref17","doi-asserted-by":"crossref","unstructured":"H. Nam, J.-W. Ha and J. Kim, Dual attention networks for multimodal reasoning and matching, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp.\u00a0299\u2013307.","DOI":"10.1109\/CVPR.2017.232"},{"key":"10.3233\/IDA-230214_ref18","doi-asserted-by":"crossref","unstructured":"K.-H. Lee, X. Chen, G. Hua, H. Hu and X. He, Stacked cross attention for image-text matching, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp.\u00a0201\u2013216.","DOI":"10.1007\/978-3-030-01225-0_13"},{"key":"10.3233\/IDA-230214_ref19","doi-asserted-by":"crossref","unstructured":"Q. Zhang, Z. Lei, Z. Zhang and S.Z. Li, Context-aware attention network for image-text retrieval, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp.\u00a03536\u20133545.","DOI":"10.1109\/CVPR42600.2020.00359"},{"key":"10.3233\/IDA-230214_ref20","doi-asserted-by":"crossref","unstructured":"Z. Ji, H. Wang, J. Han and Y. Pang, Saliency-guided attention network for image-sentence matching, in: Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2019, pp.\u00a05754\u20135763.","DOI":"10.1109\/ICCV.2019.00585"},{"key":"10.3233\/IDA-230214_ref21","doi-asserted-by":"crossref","unstructured":"H. Chen, G. Ding, X. Liu, Z. Lin, J. Liu and J. Han, Imram: Iterative matching with recurrent attention memory for cross-modal image-text retrieval, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp.\u00a012655\u201312663.","DOI":"10.1109\/CVPR42600.2020.01267"},{"key":"10.3233\/IDA-230214_ref22","doi-asserted-by":"crossref","unstructured":"L. Qu, M. Liu, J. Wu, Z. Gao and L. Nie, Dynamic modality interaction modeling for image-text retrieval, in: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2021, pp.\u00a01104\u20131113.","DOI":"10.1145\/3404835.3462829"},{"key":"10.3233\/IDA-230214_ref23","doi-asserted-by":"crossref","unstructured":"C. Liu, Z. Mao, T. Zhang, H. Xie, B. Wang and Y. Zhang, Graph structured network for image-text matching, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp.\u00a010921\u201310930.","DOI":"10.1109\/CVPR42600.2020.01093"},{"key":"10.3233\/IDA-230214_ref24","doi-asserted-by":"crossref","unstructured":"S. Wang, R. Wang, Z. Yao, S. Shan and X. Chen, Cross-modal scene graph matching for relationship-aware image-text retrieval, in: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, 2020, pp.\u00a01508\u20131517.","DOI":"10.1109\/WACV45572.2020.9093614"},{"key":"10.3233\/IDA-230214_ref26","doi-asserted-by":"crossref","unstructured":"K. He, X. Zhang, S. Ren and J. Sun, Deep residual learning for image recognition, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp.\u00a0770\u2013778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"10.3233\/IDA-230214_ref27","doi-asserted-by":"crossref","unstructured":"R. Girshick, Fast r-cnn, in: Proceedings of the IEEE International Conference on Computer Vision, 2015, pp.\u00a01440\u20131448.","DOI":"10.1109\/ICCV.2015.169"},{"key":"10.3233\/IDA-230214_ref32","doi-asserted-by":"crossref","unstructured":"B.A. Plummer, L. Wang, C.M. Cervantes, J.C. Caicedo, J. Hockenmaier and S. Lazebnik, Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models, in: Proceedings of the IEEE International Conference on Computer Vision, 2015, pp.\u00a02641\u20132649.","DOI":"10.1109\/ICCV.2015.303"},{"key":"10.3233\/IDA-230214_ref33","doi-asserted-by":"crossref","unstructured":"T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll\u00e1r and C.L. Zitnick, Microsoft coco: Common objects in context, in: European Conference on Computer Vision, Springer, 2014, pp.\u00a0740\u2013755.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"10.3233\/IDA-230214_ref34","doi-asserted-by":"crossref","unstructured":"A. Karpathy and L. Fei-Fei, Deep visual-semantic alignments for generating image descriptions, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp.\u00a03128\u20133137.","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"10.3233\/IDA-230214_ref35","unstructured":"R. Collobert, K. Kavukcuoglu and C. Farabet, Torch7: A Matlab-like Environment for Machine Learning, in: BigLearn NIPS Workshop, 2011."},{"key":"10.3233\/IDA-230214_ref38","doi-asserted-by":"crossref","unstructured":"L. Wang, Y. Li and S. Lazebnik, Learning deep structure-preserving image-text embeddings, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp.\u00a05005\u20135013.","DOI":"10.1109\/CVPR.2016.541"},{"key":"10.3233\/IDA-230214_ref39","doi-asserted-by":"crossref","unstructured":"Y. Wu, S. Wang, G. Song and Q. Huang, Learning fragment self-attention embeddings for image-text matching, in: Proceedings of the 27th ACM International Conference on Multimedia, 2019, pp.\u00a02088\u20132096.","DOI":"10.1145\/3343031.3350940"},{"issue":"1","key":"10.3233\/IDA-230214_ref40","doi-asserted-by":"publisher","first-page":"388","DOI":"10.1109\/TCSVT.2021.3060713","article-title":"Region reinforcement network with topic constraint for image-text matching","volume":"32","author":"Wu","year":"2022","journal-title":"IEEE Trans. Cir. and Sys. for Video Technol."},{"key":"10.3233\/IDA-230214_ref41","doi-asserted-by":"publisher","DOI":"10.1109\/ICME51207.2021.9428380"},{"issue":"1","key":"10.3233\/IDA-230214_ref42","first-page":"3221","article-title":"Accelerating t-SNE using tree-based algorithms","volume":"15","author":"Van Der Maaten","year":"2014","journal-title":"The Journal of Machine Learning Research"},{"key":"10.3233\/IDA-230214_ref44","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475634"}],"container-title":["Intelligent Data Analysis"],"original-title":[],"link":[{"URL":"https:\/\/content.iospress.com\/download?id=10.3233\/IDA-230214","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:20:29Z","timestamp":1777454429000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/full\/10.3233\/IDA-230214"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,28]]},"references-count":34,"journal-issue":{"issue":"3"},"URL":"https:\/\/doi.org\/10.3233\/ida-230214","relation":{},"ISSN":["1088-467X","1571-4128"],"issn-type":[{"value":"1088-467X","type":"print"},{"value":"1571-4128","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5,28]]}}}