{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,23]],"date-time":"2025-10-23T01:33:54Z","timestamp":1761183234028,"version":"build-2065373602"},"reference-count":39,"publisher":"Oxford University Press (OUP)","issue":"10","license":[{"start":{"date-parts":[[2025,4,27]],"date-time":"2025-04-27T00:00:00Z","timestamp":1745712000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/pages\/standard-publication-reuse-rights"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62262006"],"award-info":[{"award-number":["62262006"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100017691","name":"Guangxi Key Research and Development Program","doi-asserted-by":"publisher","award":["AB23026048"],"award-info":[{"award-number":["AB23026048"]}],"id":[{"id":"10.13039\/501100017691","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Development Foundation of the 7th Research Institute of China Electronics Technology Group Corporation","award":["CETC7-F2023WX0089"],"award-info":[{"award-number":["CETC7-F2023WX0089"]}]},{"name":"Guilin Science and Technology Development Program","award":["20230110-1","20220115-1"],"award-info":[{"award-number":["20230110-1","20220115-1"]}]},{"name":"Innovation Project of GUET Graduate Education","award":["2024YCXB08","2024YCXS037","2023YCXS040"],"award-info":[{"award-number":["2024YCXB08","2024YCXS037","2023YCXS040"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,10,22]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Although existing image-text retrieval methods show strong retrieval performance, they generally fail to explicitly model the subtle semantic difference within categories. To address this problem, we propose a new image-text retrieval method, called CLIP-based Semantic Refinement Method For Image-Text Retrieval (CLIP2SRITR). The proposed method consists of a modal alignment part and a semantic matching part, where the former is to learn better alignment between two modalities and the latter is to enhance the aggregation of image-text features intra-class. The pipeline helps model the distinction between image-text pairs with semantic information differences. In addition, CLIP2SRITR leverages additive angular margin for image-text retrieval loss function to maximize classification boundaries, achieving the goal of increasing intra-class similarity and inter-class variability. Through extensive experiments on three image-text benchmarks, e.g. Wikipedia, Pascal-Sentence, and NUS-WIDE, we show that CLIP2SRITR outperforms state-of-the-art methods, verifying the effectiveness of our method.<\/jats:p>","DOI":"10.1093\/comjnl\/bxaf044","type":"journal-article","created":{"date-parts":[[2025,4,15]],"date-time":"2025-04-15T07:32:10Z","timestamp":1744702330000},"page":"1377-1385","source":"Crossref","is-referenced-by-count":0,"title":["CLIP-based semantic refinement method for image-text retrieval"],"prefix":"10.1093","volume":"68","author":[{"given":"Ruidong","family":"Chen","sequence":"first","affiliation":[{"name":"Guangxi Key Laboratory of Image and Graphic Intelligent Processing, Guilin University of Electronic Technology , Guilin 541004 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuiping","family":"Guo","sequence":"additional","affiliation":[{"name":"The 7th Research Institute of China Electronics Technology Group Corporation , Guangzhou 510310 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Baohua","family":"Qiang","sequence":"additional","affiliation":[{"name":"Guangxi Key Laboratory of Image and Graphic Intelligent Processing, Guilin University of Electronic Technology , Guilin 541004 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xianyi","family":"Yang","sequence":"additional","affiliation":[{"name":"Guangxi Key Laboratory of Image and Graphic Intelligent Processing, Guilin University of Electronic Technology , Guilin 541004 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pingping","family":"Sun","sequence":"additional","affiliation":[{"name":"Guangxi Key Laboratory of Image and Graphic Intelligent Processing, Guilin University of Electronic Technology , Guilin 541004 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shihao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Guangxi Key Laboratory of Image and Graphic Intelligent Processing, Guilin University of Electronic Technology , Guilin 541004 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuan","family":"Xie","sequence":"additional","affiliation":[{"name":"Guangxi Key Laboratory of Image and Graphic Intelligent Processing, Guilin University of Electronic Technology , Guilin 541004 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2025,4,27]]},"reference":[{"key":"2025102215362916700_ref1","doi-asserted-by":"publisher","first-page":"557","DOI":"10.1093\/comjnl\/bxad001","article-title":"Toward a novel restful big data-based urban traffic incident data web service for connected vehicles","volume":"67","author":"Hireche","year":"2024","journal-title":"Comput. J."},{"key":"2025102215362916700_ref2","doi-asserted-by":"publisher","first-page":"3031","DOI":"10.1093\/comjnl\/bxae067","article-title":"Image retrieval based on auto-encoder and clustering with centroid update","volume":"67","author":"Nalini Sujantha Bel","year":"2024","journal-title":"Comput. J."},{"key":"2025102215362916700_ref3","doi-asserted-by":"publisher","first-page":"2799","DOI":"10.1093\/comjnl\/bxae045","article-title":"Coverless image steganography using content-based image patch retrieval","volume":"67","author":"Taheri","year":"2024","journal-title":"Comput. J."},{"key":"2025102215362916700_ref4","doi-asserted-by":"publisher","first-page":"2191","DOI":"10.1093\/comjnl\/bxac070","article-title":"Feature extraction based deep indexing by deep fuzzy clustering for image retrieval using Jaro Winkler distance","volume":"66","author":"Kumar","year":"2023","journal-title":"Comput. J."},{"key":"2025102215362916700_ref5","doi-asserted-by":"publisher","first-page":"516","DOI":"10.1093\/comjnl\/bxaa073","article-title":"Mesh-based semantic indexing approach to enhance biomedical information retrieval","volume":"65","author":"Kammoun","year":"2022","journal-title":"Comput J"},{"key":"2025102215362916700_ref6","doi-asserted-by":"publisher","first-page":"1979","DOI":"10.1093\/comjnl\/bxad117","article-title":"Privacy-preserving image retrieval with multi-modal query","volume":"67","author":"Zhou","year":"2024","journal-title":"Comput. J."},{"key":"2025102215362916700_ref7","first-page":"2830","article-title":"Multilateral semantic relations modeling for image text retrieval","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, Vancouver, BC, Canada, 17-24 June","author":"Wang","year":"2023"},{"key":"2025102215362916700_ref8","doi-asserted-by":"publisher","first-page":"3622","DOI":"10.1109\/TIP.2023.3286710","article-title":"Efficient token-guided image-text retrieval with consistent multimodal contrastive training","volume":"32","author":"Liu","year":"2023","journal-title":"IEEE Trans Image Process"},{"key":"2025102215362916700_ref9","doi-asserted-by":"publisher","first-page":"2639","DOI":"10.1162\/0899766042321814","article-title":"Canonical correlation analysis: an overview with application to learning methods","volume":"16","author":"Hardoon","year":"2004","journal-title":"Neural Comput"},{"author":"Wang","key":"2025102215362916700_ref10","article-title":"Large-scale approximate kernel canonical correlation analysis"},{"key":"2025102215362916700_ref11","doi-asserted-by":"publisher","first-page":"1856","DOI":"10.1093\/comjnl\/bxac047","article-title":"Urdu named entity recognition system using deep learning approaches","volume":"66","author":"Haq","year":"2023","journal-title":"Comput. J."},{"key":"2025102215362916700_ref12","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1145\/2647868.2654902","article-title":"Cross-modal retrieval with correspondence autoencoder","volume-title":"Proceedings of the 22nd ACM international conference on multimedia, Orlando, FL, United States, 3-7 November","author":"Feng","year":"2014"},{"key":"2025102215362916700_ref13","first-page":"649","article-title":"Effective multi-modal retrieval based on stacked auto-encoders","volume-title":"Proceedings of the VLDB endowment, Hangzhou, China, 1-5 September","author":"Wang","year":"2014"},{"key":"2025102215362916700_ref14","doi-asserted-by":"crossref","first-page":"965","DOI":"10.1109\/TCSVT.2013.2276704","article-title":"Learning cross-media joint representation with sparse and semisupervised regularization","volume":"24","author":"Zhai","year":"2013","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"2025102215362916700_ref15","doi-asserted-by":"publisher","first-page":"405","DOI":"10.1109\/TMM.2017.2742704","article-title":"CCL: cross-modal correlation learning with multigrained fusion by hierarchical network","volume":"20","author":"Peng","year":"2017","journal-title":"IEEE Trans Multimedia"},{"key":"2025102215362916700_ref16","first-page":"10386","article-title":"Deep supervised cross-modal retrieval","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, Long Beach, CA, United States, 16-20 June","author":"Zhen","year":"2019"},{"key":"2025102215362916700_ref17","doi-asserted-by":"crossref","first-page":"154","DOI":"10.1145\/3123266.3123326","article-title":"Adversarial cross-modal retrieval","volume-title":"Proceedings of the 25th ACM international conference on multimedia, mountain view, CA, United States, 23-27 October","author":"Wang","year":"2017"},{"key":"2025102215362916700_ref18","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2020.107335","article-title":"Modality-specific and shared generative adversarial network for cross-modal retrieval","volume":"104","author":"Wu","year":"2020","journal-title":"Pattern Recognit"},{"key":"2025102215362916700_ref19","doi-asserted-by":"publisher","first-page":"1346","DOI":"10.1093\/comjnl\/bxad063","article-title":"Spatial-aware multi-directional autoencoder for pre-training","volume":"67","author":"Yang","year":"2024","journal-title":"Comput J"},{"key":"2025102215362916700_ref20","first-page":"8748","article-title":"Learning transferable visual models from natural language supervision","volume-title":"Proceedings of the 38th international conference on machine learning, virtual, online, 18-24 July","author":"Radford","year":"2021"},{"key":"2025102215362916700_ref21","article-title":"A comprehensive empirical study of vision-language pre-trained model for supervised cross-modal retrieval","author":"Zeng","year":"2022","journal-title":"CoRR"},{"key":"2025102215362916700_ref22","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3637442","article-title":"Incomplete cross-modal retrieval with deep correlation transfer","volume":"20","author":"Shi","year":"2024","journal-title":"ACM Trans Multimedia Comput Commun Appl"},{"key":"2025102215362916700_ref23","doi-asserted-by":"publisher","first-page":"2465","DOI":"10.1109\/TCSVT.2022.3220297","article-title":"Image-text retrieval with cross-modal semantic importance consistency","volume":"33","author":"Liu","year":"2023","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"2025102215362916700_ref24","doi-asserted-by":"publisher","first-page":"1838","DOI":"10.1109\/TNNLS.2020.2997020","article-title":"Deep semantic multimodal hashing network for scalable image-text and video-text retrievals","volume":"34","author":"Jin","year":"2023","journal-title":"IEEE Trans Neural Netw Learn Syst"},{"key":"2025102215362916700_ref25","doi-asserted-by":"publisher","first-page":"508","DOI":"10.1093\/comjnl\/bxac192","article-title":"A deep learning model for energy-aware task scheduling algorithm based on learning automata for fog computing","volume":"67","author":"Ebrahim Pourian","year":"2024","journal-title":"Comput J"},{"key":"2025102215362916700_ref26","first-page":"1247","article-title":"Deep canonical correlation analysis","volume-title":"Proceedings of the 30th international conference on machine learning, Atlanta, GA, United States, 16-21 June","author":"Andrew","year":"2013"},{"key":"2025102215362916700_ref27","first-page":"3441","article-title":"Deep correlation for matching images and text","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition, Boston, MA, United States, 7-12 June","author":"Yan","year":"2015"},{"key":"2025102215362916700_ref28","first-page":"1083","article-title":"On deep multi-view representation learning","volume-title":"Proceedings of the 32nd international conference on machine learning, Lile, France, 6-11 July","author":"Wang","year":"2015"},{"key":"2025102215362916700_ref29","first-page":"770","article-title":"Deep residual learning for image recognition","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition, Las Vegas, NV, United States, 26 June-1 July","author":"He","year":"2016"},{"key":"2025102215362916700_ref30","first-page":"5998","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv Neural Inf Process Syst"},{"key":"2025102215362916700_ref31","article-title":"Layer normalization","author":"Ba","year":"2016","journal-title":"CoRR"},{"key":"2025102215362916700_ref32","doi-asserted-by":"publisher","first-page":"5962","DOI":"10.1109\/TPAMI.2021.3087709","article-title":"ArcFace: additive angular margin loss for deep face recognition","volume":"44","author":"Deng","year":"2022","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"2025102215362916700_ref33","article-title":"SiT: self-supervised vision transformer","author":"Atito","year":"2021","journal-title":"CoRR"},{"key":"2025102215362916700_ref34","first-page":"1","article-title":"Adam: a method for stochastic optimization","volume-title":"Proceedings of the 3rd international conference on learning representations, San Diego, CA, United States, 7-9 may","author":"Kingma","year":"2015"},{"key":"2025102215362916700_ref35","doi-asserted-by":"publisher","first-page":"521","DOI":"10.1109\/TPAMI.2013.142","article-title":"On the role of correlation and abstraction in cross-modal multimedia retrieval","volume":"36","author":"Pereira","year":"2014","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"2025102215362916700_ref36","first-page":"139","article-title":"Collecting image annotations using amazon\u2019s mechanical turk","volume-title":"Proceedings of the NAACL HLT 2010 workshop on creating speech and language data with Amazon\u2019s mechanical Turk, Los Angeles, CA, United States, 6 June","author":"Rashtchian","year":"2010"},{"key":"2025102215362916700_ref37","first-page":"1","article-title":"NUS-WIDE: a real-world web image database from national university of Singapore","volume-title":"Proceedings of the ACM international conference on image and video retrieval, Santorini Island, Greece, 8-10 July","author":"Chua","year":"2009"},{"key":"2025102215362916700_ref38","first-page":"2787","article-title":"Cross-modal implicit relation reasoning and aligning for text-to-image person retrieval","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, Vancouver, Canada, 17-24 June","author":"Jiang","year":"2023"},{"key":"2025102215362916700_ref39","first-page":"4527","article-title":"Category alignment adversarial learning for cross-modal retrieval","volume":"35","author":"He","year":"2023","journal-title":"IEEE Trans Knowl Data Eng"}],"container-title":["The Computer Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/68\/10\/1377\/63015640\/bxaf044.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/68\/10\/1377\/63015640\/bxaf044.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,22]],"date-time":"2025-10-22T19:36:36Z","timestamp":1761161796000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/comjnl\/article\/68\/10\/1377\/8120571"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,27]]},"references-count":39,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2025,4,27]]},"published-print":{"date-parts":[[2025,10,22]]}},"URL":"https:\/\/doi.org\/10.1093\/comjnl\/bxaf044","relation":{},"ISSN":["0010-4620","1460-2067"],"issn-type":[{"type":"print","value":"0010-4620"},{"type":"electronic","value":"1460-2067"}],"subject":[],"published-other":{"date-parts":[[2025,10]]},"published":{"date-parts":[[2025,4,27]]}}}