{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T18:29:20Z","timestamp":1776882560567,"version":"3.51.2"},"reference-count":47,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2025,1,30]],"date-time":"2025-01-30T00:00:00Z","timestamp":1738195200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,1,30]],"date-time":"2025-01-30T00:00:00Z","timestamp":1738195200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100005230","name":"Natural Science Foundation of Chongqing Municipality","doi-asserted-by":"publisher","award":["CSTB2022NSCQ-MSX1417"],"award-info":[{"award-number":["CSTB2022NSCQ-MSX1417"]}],"id":[{"id":"10.13039\/501100005230","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Science and Technology Research Program of Chongqing Municipal Education Commission","award":["KJZD-K202200513"],"award-info":[{"award-number":["KJZD-K202200513"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Data Sci. Eng."],"published-print":{"date-parts":[[2025,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Although the fact that current methods have some effects, unsupervised cross-modal hashing methods still face several common challenges. First of all, the text features that have been collected from text data are not comprehensive enough to provide sufficient guidance for building textual modal similarity matrices. Secondly, the fusion of similarity matrices from different modalities lacks adaptability, leading to a less accurate final similarity matrix. This work suggests Enhanced Similarity Attention Fusion Hashing (ESAFH) as a remedy for these problems. Firstly, we construct a text encoder to enrich text features, an adjacency matrix is built to represent the association relationship between pairs of samples. Additionally, it is thought that features can be extracted from the sample and its semantic neighbor samples to enhance text features. Furthermore, we enhance the original similarity matrix by incorporating related information. This step aims to improve the accuracy of similarity estimation by considering the enriched text features obtained in the previous step. Finally, we introduce an enhanced attention fusion mechanism. This mechanism adaptively fuses the similarity matrices from different modalities, creating a unified inter-modal similarity matrix. This fused matrix guides the learning of hash functions by preserving the most relevant information from each modality. Through comprehensive experiments on the three popular datasets, the suggested ESAFH method is thoroughly assessed. The findings show that on these datasets, ESAFH performs satisfactorily in cross-modal retrieval tasks. In conclusion, by boosting text features, improving the similarity matrix, and utilizing an attention fusion mechanism, ESAFH solves the shortcomings of current methods.<\/jats:p>","DOI":"10.1007\/s41019-024-00274-7","type":"journal-article","created":{"date-parts":[[2025,1,30]],"date-time":"2025-01-30T10:56:20Z","timestamp":1738234580000},"page":"258-276","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Enhanced-Similarity Attention Fusion for Unsupervised Cross-Modal Hashing Retrieval"],"prefix":"10.1007","volume":"10","author":[{"given":"Mingyong","family":"Li","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4094-0078","authenticated-orcid":false,"given":"Mingyuan","family":"Ge","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,1,30]]},"reference":[{"key":"274_CR1","doi-asserted-by":"crossref","unstructured":"Jiang Q-Y, Li W-J (2017) Deep cross-modal hashing. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3232\u20133240","DOI":"10.1109\/CVPR.2017.348"},{"key":"274_CR2","doi-asserted-by":"crossref","unstructured":"Yang E, Deng C, Liu W, Liu X, Tao D, Gao X (2017) Pairwise relationship guided deep hashing for cross-modal retrieval. In: Proceedings of the AAAI conference on artificial intelligence, vol 31","DOI":"10.1609\/aaai.v31i1.10719"},{"key":"274_CR3","doi-asserted-by":"crossref","unstructured":"Liu H, Ji R, Wu Y, Huang F, Zhang B (2017) Cross-modality binary code learning via fusion similarity hashing. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 7380\u20137388","DOI":"10.1109\/CVPR.2017.672"},{"issue":"10","key":"274_CR4","doi-asserted-by":"publisher","first-page":"2703","DOI":"10.1109\/TCSVT.2017.2723302","volume":"28","author":"D Wang","year":"2017","unstructured":"Wang D, Wang Q, Gao X (2017) Robust and flexible discrete hashing for cross-modal similarity search. IEEE Trans Circuits Syst Video Technol 28(10):2703\u20132715","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"274_CR5","first-page":"5","volume":"1","author":"G Wu","year":"2018","unstructured":"Wu G, Lin Z, Han J, Liu L, Ding G, Zhang B, Shen J (2018) Unsupervised deep hashing via binary latent factor models for large-scale cross-modal retrieval. IJCAI 1:5","journal-title":"IJCAI"},{"issue":"1","key":"274_CR6","doi-asserted-by":"publisher","first-page":"174","DOI":"10.1109\/TMM.2019.2922128","volume":"22","author":"J Zhang","year":"2019","unstructured":"Zhang J, Peng Y (2019) Multi-pathway generative adversarial hashing for unsupervised cross-modal retrieval. IEEE Trans Multimed 22(1):174\u2013187","journal-title":"IEEE Trans Multimed"},{"key":"274_CR7","doi-asserted-by":"publisher","first-page":"107479","DOI":"10.1016\/j.patcog.2020.107479","volume":"107","author":"D Wang","year":"2020","unstructured":"Wang D, Wang Q, He L, Gao X, Tian Y (2020) Joint and individual matrix factorization hashing for large-scale cross-modal retrieval. Pattern Recognit 107:107479","journal-title":"Pattern Recognit"},{"key":"274_CR8","doi-asserted-by":"crossref","unstructured":"Su S, Zhong Z, Zhang C (2019) Deep joint-semantics reconstructing hashing for large-scale unsupervised cross-modal retrieval. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 3027\u20133035","DOI":"10.1109\/ICCV.2019.00312"},{"key":"274_CR9","doi-asserted-by":"crossref","unstructured":"Yu J, Zhou H, Zhan Y, Tao D (2021) Deep graph-neighbor coherence preserving network for unsupervised cross-modal hashing. In: Proceedings of the AAAI conference on artificial intelligence, vol 35, pp 4626\u20134634","DOI":"10.1609\/aaai.v35i5.16592"},{"issue":"2","key":"274_CR10","doi-asserted-by":"publisher","first-page":"563","DOI":"10.1007\/s11280-020-00859-y","volume":"24","author":"P-F Zhang","year":"2021","unstructured":"Zhang P-F, Luo Y, Huang Z, Xu X-S, Song J (2021) High-order nonlocal hashing for unsupervised cross-modal retrieval. World Wide Web 24(2):563\u2013583","journal-title":"World Wide Web"},{"key":"274_CR11","doi-asserted-by":"crossref","unstructured":"Shen X, Zhang H, Li L, Liu L (2021) Attention-guided semantic hashing for unsupervised cross-modal retrieval. In: 2021 IEEE international conference on multimedia and expo (ICME). IEEE, pp 1\u20136","DOI":"10.1109\/ICME51207.2021.9428330"},{"key":"274_CR12","doi-asserted-by":"crossref","unstructured":"Yang D, Wu D, Zhang W, Zhang H, Li B, Wang W (2020) Deep semantic-alignment hashing for unsupervised cross-modal retrieval. In: Proceedings of the 2020 international conference on multimedia Retrieval, pp 44\u201352","DOI":"10.1145\/3372278.3390673"},{"issue":"10","key":"274_CR13","doi-asserted-by":"publisher","first-page":"7255","DOI":"10.1109\/TCSVT.2022.3172716","volume":"32","author":"Y Shi","year":"2022","unstructured":"Shi Y, Zhao Y, Liu X, Zheng F, Ou W, You X, Peng Q (2022) Deep adaptively-enhanced hashing with discriminative similarity guidance for unsupervised cross-modal retrieval. IEEE Trans Circuits Syst Video Technol 32(10):7255\u20137268","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"274_CR14","doi-asserted-by":"crossref","unstructured":"Weng W, Wu J, Yang L, Liu L, Hu B (2019) Label-based deep semantic hashing for cross-modal retrieval. In: International conference on neural information processing. Springer, pp 24\u201336","DOI":"10.1007\/978-3-030-36718-3_3"},{"key":"274_CR15","doi-asserted-by":"crossref","unstructured":"Li C, Deng C, Li N, Liu W, Gao X, Tao D (2018) Self-supervised adversarial hashing networks for cross-modal retrieval. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4242\u20134251","DOI":"10.1109\/CVPR.2018.00446"},{"key":"274_CR16","unstructured":"Defferrard M, Bresson X, Vandergheynst P (2016) Convolutional neural networks on graphs with fast localized spectral filtering. Advances in neural information processing systems 29"},{"key":"274_CR17","unstructured":"Wu B, Yang Q, Zheng W-S, Wang Y, Wang J (2015) Quantized correlation hashing for fast cross-modal search. In: Twenty-fourth international joint conference on artificial intelligence"},{"key":"274_CR18","doi-asserted-by":"crossref","unstructured":"Cao Y, Long M, Wang J, Zhu H (2016) Correlation autoencoder hashing for supervised cross-modal search. In: Proceedings of the 2016 ACM on international conference on multimedia retrieval, pp 197\u2013204","DOI":"10.1145\/2911996.2912000"},{"key":"274_CR19","doi-asserted-by":"crossref","unstructured":"Ye Z, Peng Y (2018) Multi-scale correlation for sequential cross-modal hashing learning. In: Proceedings of the 26th ACM international conference on multimedia, pp 852\u2013860","DOI":"10.1145\/3240508.3240560"},{"issue":"7","key":"274_CR20","doi-asserted-by":"publisher","first-page":"3490","DOI":"10.1109\/TIP.2019.2897944","volume":"28","author":"Q-Y Jiang","year":"2019","unstructured":"Jiang Q-Y, Li W-J (2019) Discrete latent factor model for cross-modal hashing. IEEE Trans Image Process 28(7):3490\u20133501","journal-title":"IEEE Trans Image Process"},{"issue":"10","key":"274_CR21","doi-asserted-by":"publisher","first-page":"3351","DOI":"10.1109\/TKDE.2020.2970050","volume":"33","author":"HT Shen","year":"2020","unstructured":"Shen HT, Liu L, Yang Y, Xu X, Huang Z, Shen F, Hong R (2020) Exploiting subspace relation in semantic labels for cross-modal hashing. IEEE Trans Knowl Data Eng 33(10):3351\u20133365","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"274_CR22","doi-asserted-by":"publisher","first-page":"109503","DOI":"10.1016\/j.knosys.2022.109503","volume":"253","author":"J Li","year":"2022","unstructured":"Li J, Yu E, Ma J, Chang X, Zhang H, Sun J (2022) Discrete fusion adversarial hashing for cross-modal retrieval. Knowl-Based Syst 253:109503","journal-title":"Knowl-Based Syst"},{"key":"274_CR23","unstructured":"Kumar S, Udupa R (2011) Learning hash functions for cross-view similarity search. In: Twenty-second international joint conference on artificial intelligence"},{"key":"274_CR24","doi-asserted-by":"crossref","unstructured":"Song J, Yang Y, Yang Y, Huang Z, Shen HT (2013) Inter-media hashing for large-scale retrieval from heterogeneous data sources. In: Proceedings of the 2013 ACM SIGMOD international conference on management of data, pp 785\u2013796","DOI":"10.1145\/2463676.2465274"},{"key":"274_CR25","doi-asserted-by":"crossref","unstructured":"Ding G, Guo Y, Zhou J (2014) Collective matrix factorization hashing for multimodal data. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 2075\u20132082","DOI":"10.1109\/CVPR.2014.267"},{"key":"274_CR26","doi-asserted-by":"crossref","unstructured":"Zhou J, Ding G, Guo Y (2014) Latent semantic sparse hashing for cross-modal similarity search. In: Proceedings of the 37th international ACM SIGIR conference on research and development in information retrieval, pp 415\u2013424","DOI":"10.1145\/2600428.2609610"},{"key":"274_CR27","doi-asserted-by":"crossref","unstructured":"Liu S, Qian S, Guan Y, Zhan J, Ying L (2020) Joint-modal distribution-based similarity hashing for large-scale unsupervised deep cross-modal retrieval. In: Proceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval, pp 1379\u20131388","DOI":"10.1145\/3397271.3401086"},{"key":"274_CR28","doi-asserted-by":"publisher","first-page":"466","DOI":"10.1109\/TMM.2021.3053766","volume":"24","author":"P-F Zhang","year":"2021","unstructured":"Zhang P-F, Li Y, Huang Z, Xu X-S (2021) Aggregation-based graph convolutional hashing for unsupervised cross-modal retrieval. IEEE Trans Multimed 24:466\u2013479","journal-title":"IEEE Trans Multimed"},{"key":"274_CR29","doi-asserted-by":"crossref","unstructured":"Mikriukov G, Ravanbakhsh M, Demir B (2022) Unsupervised contrastive hashing for cross-modal retrieval in remote sensing. In: ICASSP 2022-2022 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, pp 4463\u20134467","DOI":"10.1109\/ICASSP43922.2022.9746251"},{"issue":"1","key":"274_CR30","doi-asserted-by":"publisher","first-page":"2","DOI":"10.1007\/s13735-023-00268-7","volume":"12","author":"L Mingyong","year":"2023","unstructured":"Mingyong L, Yewen L, Mingyuan G, Longfei M (2023) Clip-based fusion-modal reconstructing hashing for large-scale unsupervised cross-modal retrieval. Int J Multimed Inform Retr 12(1):2","journal-title":"Int J Multimed Inform Retr"},{"issue":"3","key":"274_CR31","first-page":"3877","volume":"45","author":"P Hu","year":"2022","unstructured":"Hu P, Zhu H, Lin J, Peng D, Zhao Y-P, Peng X (2022) Unsupervised contrastive cross-modal hashing. IEEE Trans Pattern Anal Mach Intell 45(3):3877\u20133889","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"274_CR32","doi-asserted-by":"crossref","unstructured":"Yao H-L, Zhan Y-W, Chen Z-D, Luo X, Xu X-S (2021) Teach: attention-aware deep cross-modal hashing. In: Proceedings of the 2021 international conference on multimedia retrieval, pp 376\u2013384","DOI":"10.1145\/3460426.3463625"},{"key":"274_CR33","unstructured":"Sutskever I, Vinyals O, Le QV (2014) Sequence to sequence learning with neural networks. Advances in neural information processing systems 27"},{"key":"274_CR34","doi-asserted-by":"crossref","unstructured":"Yang Z, Yang D, Dyer C, He X, Smola A, Hovy E (2016) Hierarchical attention networks for document classification. In: Proceedings of the 2016 conference of the North American chapter of the association for computational linguistics: human language technologies, pp 1480\u20131489","DOI":"10.18653\/v1\/N16-1174"},{"key":"274_CR35","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I (2017) Attention is all you need. Advances in neural information processing systems 30"},{"key":"274_CR36","doi-asserted-by":"crossref","unstructured":"Zhang X, Lai H, Feng J (2018) Attention-aware deep adversarial hashing for cross-modal retrieval. In: Proceedings of the European conference on computer vision (ECCV), pp 591\u2013606","DOI":"10.1007\/978-3-030-01267-0_36"},{"issue":"11","key":"274_CR37","doi-asserted-by":"publisher","first-page":"5585","DOI":"10.1109\/TIP.2018.2852503","volume":"27","author":"Y Peng","year":"2018","unstructured":"Peng Y, Qi J, Yuan Y (2018) Modality-specific cross-modal similarity measurement with recurrent attention network. IEEE Trans Image Process 27(11):5585\u20135599","journal-title":"IEEE Trans Image Process"},{"issue":"7","key":"274_CR38","doi-asserted-by":"publisher","first-page":"5397","DOI":"10.1007\/s00521-021-06696-y","volume":"34","author":"J Wu","year":"2022","unstructured":"Wu J, Weng W, Fu J, Liu L, Hu B (2022) Deep semantic hashing with dual attention for cross-modal retrieval. Neural Comput Appl 34(7):5397\u20135416","journal-title":"Neural Comput Appl"},{"key":"274_CR39","doi-asserted-by":"publisher","first-page":"107927","DOI":"10.1016\/j.knosys.2021.107927","volume":"239","author":"X Zou","year":"2022","unstructured":"Zou X, Wu S, Zhang N, Bakker EM (2022) Multi-label modality enhanced attention based self-supervised deep cross-modal hashing. Knowl-Based Syst 239:107927","journal-title":"Knowl-Based Syst"},{"key":"274_CR40","unstructured":"Simonyan K, Zisserman A (2014) Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556"},{"key":"274_CR41","doi-asserted-by":"crossref","unstructured":"Ko Y (2012) A study of term weighting schemes using class information for text classification. In: Proceedings of the 35th international ACM SIGIR conference on research and development in information retrieval, pp 1029\u20131030","DOI":"10.1145\/2348283.2348453"},{"key":"274_CR42","doi-asserted-by":"crossref","unstructured":"Huiskes MJ, Lew MS (2008) The mir flickr retrieval evaluation. In: Proceedings of the 1st ACM international conference on multimedia information retrieval, pp 39\u201343","DOI":"10.1145\/1460096.1460104"},{"key":"274_CR43","doi-asserted-by":"crossref","unstructured":"Chua T-S, Tang J, Hong R, Li H, Luo Z, Zheng Y (2009) Nus-wide: a real-world web image database from national university of Singapore. In: Proceedings of the ACM international conference on image and video retrieval, pp 1\u20139","DOI":"10.1145\/1646396.1646452"},{"key":"274_CR44","doi-asserted-by":"crossref","unstructured":"Lin T-Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Doll\u00e1r P, Zitnick CL (2014) Microsoft coco: common objects in context. In: European conference on computer vision. Springer, pp 740\u2013755","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"274_CR45","unstructured":"Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems 25"},{"key":"274_CR46","unstructured":"Witten IH, Frank E, Hall MA, Pal CJ, Data M (2005) Practical machine learning tools and techniques. In: Data mining, vol 2, pp 403\u2013413 . Elsevier Amsterdam, The Netherlands"},{"issue":"7","key":"274_CR47","doi-asserted-by":"publisher","first-page":"3439","DOI":"10.3390\/s23073439","volume":"23","author":"Y Li","year":"2023","unstructured":"Li Y, Ge M, Li M, Li T, Xiang S (2023) Clip-based adaptive graph attention network for large-scale unsupervised multi-modal hashing retrieval. Sensors 23(7):3439","journal-title":"Sensors"}],"container-title":["Data Science and Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s41019-024-00274-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s41019-024-00274-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s41019-024-00274-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,6]],"date-time":"2025-06-06T08:57:49Z","timestamp":1749200269000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s41019-024-00274-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,30]]},"references-count":47,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,6]]}},"alternative-id":["274"],"URL":"https:\/\/doi.org\/10.1007\/s41019-024-00274-7","relation":{},"ISSN":["2364-1185","2364-1541"],"issn-type":[{"value":"2364-1185","type":"print"},{"value":"2364-1541","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,30]]},"assertion":[{"value":"30 January 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 November 2024","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 November 2024","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 January 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}