{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,30]],"date-time":"2025-12-30T01:15:12Z","timestamp":1767057312929,"version":"3.48.0"},"reference-count":58,"publisher":"World Scientific Pub Co Pte Ltd","issue":"02","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62371144"],"award-info":[{"award-number":["62371144"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62461004"],"award-info":[{"award-number":["62461004"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62172458"],"award-info":[{"award-number":["62172458"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Patt. Recogn. Artif. Intell."],"published-print":{"date-parts":[[2026,2]]},"abstract":"<jats:p>Composed image retrieval (CIR) aims to retrieve target images by combining a reference image with a modification text. Traditional CIR methods often struggle with feature-level multimodal fusion, leading to deviations from the original embedding space. To address this, we propose a Dual-Branch Multi-Scale Network (DMN) that integrates a combining branch and a complete text branch. To enhance the use of captions generated by advanced image captioning models for CIR, the DMN leverages an attribute-driven disentanglement layer to separate features into distinct latent factors and employs a dual-path multimodal fusion module for effective feature integration. Additionally, a multi-scale matching module incorporating both global and local matching strategies is introduced to enhance fine-grained feature discrimination. Experimental results on the FashionIQ, Shoes, and CIRR datasets demonstrate that our DMN model consistently outperforms state-of-the-art methods, achieving improvements of up to 1.43% in mean recall metrics.<\/jats:p>","DOI":"10.1142\/s0218001425540217","type":"journal-article","created":{"date-parts":[[2025,11,4]],"date-time":"2025-11-04T04:21:32Z","timestamp":1762230092000},"source":"Crossref","is-referenced-by-count":0,"title":["Unlocking the Potential of Auxiliary Captions via Dual-Branch Multi-Scale Network for Composed Image Retrieval"],"prefix":"10.1142","volume":"40","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-9918-5783","authenticated-orcid":false,"given":"Jinhong","family":"Xu","sequence":"first","affiliation":[{"name":"School of Computer, Electronics and Information, Guangxi University, Nanning 530004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3489-1887","authenticated-orcid":false,"given":"Lina","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Computer, Electronics and Information, Guangxi University, Nanning 530004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-1921-1651","authenticated-orcid":false,"given":"Xichun","family":"Li","sequence":"additional","affiliation":[{"name":"School of Mathematics and Computer Science, Guangxi Normal University for Nationalities, Chongzuo 532200, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0674-9915","authenticated-orcid":false,"given":"Thomas","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering, Guangxi University, Nanning 530004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6887-130X","authenticated-orcid":false,"given":"Yuan Yan","family":"Tang","sequence":"additional","affiliation":[{"name":"Faculty of Science and Technology, University of Macau, Macau 999078, China"},{"name":"Faculty of Science and Technology, UOW College Hong Kong, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9336-3155","authenticated-orcid":false,"given":"Patrick Shen-Pei","family":"Wang","sequence":"additional","affiliation":[{"name":"Department of Computer and Information Science, Khoury College of Computer Sciences, Northeastern University, MA 02115, Boston, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2025,12,10]]},"reference":[{"key":"S0218001425540217BIB001","unstructured":"R. Anil\n                      et\u00a0al.\n                      , Gemini: A family of highly capable multimodal models, arXiv:2312.11805."},{"key":"S0218001425540217BIB002","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00543"},{"key":"S0218001425540217BIB003","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00045"},{"key":"S0218001425540217BIB004","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00307"},{"key":"S0218001425540217BIB005","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58577-8_7"},{"issue":"6","key":"S0218001425540217BIB006","first-page":"1","volume":"20","author":"Chen Y.","year":"2024","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"S0218001425540217BIB007","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2024.106211"},{"key":"S0218001425540217BIB008","unstructured":"G. Delmas, R. S. de Rezende, G. Csurka and D. Larlus, Artemis: Attention-based retrieval with text-explicit matching and implicit similarity, arXiv:2203.08101."},{"key":"S0218001425540217BIB009","unstructured":"E. Dodds, J. Culpepper, S. Herdade, Y. Zhang and K. Boakye, Modality-agnostic attention fusion for visual search with text feedback, arXiv:2007.00145."},{"key":"S0218001425540217BIB010","unstructured":"F. Faghri, D. J. Fleet, J. R. Kiros and S. Fidler, VSE++: Improving visual-semantic embeddings with hard negatives, arXiv:1707.05612."},{"key":"S0218001425540217BIB011","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2025.129642"},{"key":"S0218001425540217BIB012","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-24797-2_4"},{"key":"S0218001425540217BIB013","first-page":"676","volume":"31","author":"Guo X.","year":"2018","journal-title":"Adv. Neural Information Processing Systems"},{"key":"S0218001425540217BIB014","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00262"},{"key":"S0218001425540217BIB015","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.127202"},{"key":"S0218001425540217BIB016","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"S0218001425540217BIB017","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2024.106200"},{"key":"S0218001425540217BIB018","first-page":"4904","volume-title":"Int. Conf. Machine Learning","author":"Jia C.","year":"2021"},{"key":"S0218001425540217BIB019","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01225-0_13"},{"key":"S0218001425540217BIB020","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00086"},{"key":"S0218001425540217BIB021","first-page":"19730","volume-title":"Int. Conf. Machine Learning","author":"Li J.","year":"2023"},{"key":"S0218001425540217BIB022","first-page":"12888","volume-title":"Int. Conf. Machine Learning","author":"Li J.","year":"2022"},{"issue":"6","key":"S0218001425540217BIB023","first-page":"1","volume":"20","author":"Li S.","year":"2024","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"S0218001425540217BIB024","unstructured":"L. H. Li, M. Yatskar, D. Yin, C.J. Hsieh and K.W. Chang, Visualbert: A simple and performant baseline for vision and language, arXiv:1908.03557."},{"key":"S0218001425540217BIB025","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00213"},{"key":"S0218001425540217BIB026","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00565"},{"key":"S0218001425540217BIB027","first-page":"13","volume":"32","author":"Lu J.","year":"2019","journal-title":"Adv. Neural Information Processing Systems"},{"key":"S0218001425540217BIB028","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.07.028"},{"key":"S0218001425540217BIB029","unstructured":"R. Mokady, A. Hertz and A. H. Bermano, Clipcap: CLIP prefix for image captioning, arXiv:2111.09734."},{"key":"S0218001425540217BIB030","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3208742"},{"key":"S0218001425540217BIB031","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2022.05.008"},{"key":"S0218001425540217BIB032","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462829"},{"key":"S0218001425540217BIB033","first-page":"8748","volume-title":"Int. Conf. Machine Learning","author":"Radford A.","year":"2021"},{"key":"S0218001425540217BIB034","unstructured":"A. Suhr, S. Zhou, A. Zhang, I. Zhang, H. Bai and Y. Artzi, A corpus for reasoning about natural language grounded in photographs, arXiv:1811.00491."},{"key":"S0218001425540217BIB035","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i7.32768"},{"key":"S0218001425540217BIB036","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20047-2_21"},{"key":"S0218001425540217BIB037","doi-asserted-by":"publisher","DOI":"10.1145\/2812802"},{"key":"S0218001425540217BIB038","unstructured":"A. van den Oord, Y. Li and O. Vinyals, Representation learning with contrastive predictive coding, arXiv:1807.03748."},{"key":"S0218001425540217BIB039","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00660"},{"key":"S0218001425540217BIB040","doi-asserted-by":"publisher","DOI":"10.1016\/j.vrih.2023.06.003"},{"key":"S0218001425540217BIB041","doi-asserted-by":"publisher","DOI":"10.1007\/s00371-022-02753-2"},{"key":"S0218001425540217BIB042","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3088863"},{"key":"S0218001425540217BIB043","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462967"},{"key":"S0218001425540217BIB044","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3346434"},{"key":"S0218001425540217BIB045","doi-asserted-by":"publisher","DOI":"10.1145\/3626772.3657727"},{"key":"S0218001425540217BIB046","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611817"},{"key":"S0218001425540217BIB047","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01115"},{"key":"S0218001425540217BIB048","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3235495"},{"key":"S0218001425540217BIB049","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3299791"},{"key":"S0218001425540217BIB050","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i7.28479"},{"key":"S0218001425540217BIB051","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413518"},{"key":"S0218001425540217BIB052","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3204213"},{"key":"S0218001425540217BIB053","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3273466"},{"key":"S0218001425540217BIB054","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548126"},{"key":"S0218001425540217BIB055","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3532047"},{"key":"S0218001425540217BIB056","doi-asserted-by":"publisher","DOI":"10.1145\/3596445"},{"key":"S0218001425540217BIB057","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3521646"},{"key":"S0218001425540217BIB058","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2025.113303"}],"container-title":["International Journal of Pattern Recognition and Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0218001425540217","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,30]],"date-time":"2025-12-30T01:03:39Z","timestamp":1767056619000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/10.1142\/S0218001425540217"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,10]]},"references-count":58,"journal-issue":{"issue":"02","published-print":{"date-parts":[[2026,2]]}},"alternative-id":["10.1142\/S0218001425540217"],"URL":"https:\/\/doi.org\/10.1142\/s0218001425540217","relation":{},"ISSN":["0218-0014","1793-6381"],"issn-type":[{"type":"print","value":"0218-0014"},{"type":"electronic","value":"1793-6381"}],"subject":[],"published":{"date-parts":[[2025,12,10]]},"article-number":"2554021"}}