{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T07:56:45Z","timestamp":1774943805340,"version":"3.50.1"},"reference-count":61,"publisher":"World Scientific Pub Co Pte Ltd","issue":"06","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62371144"],"award-info":[{"award-number":["62371144"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62461004"],"award-info":[{"award-number":["62461004"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62172458"],"award-info":[{"award-number":["62172458"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Patt. Recogn. Artif. Intell."],"published-print":{"date-parts":[[2026,5]]},"abstract":"<jats:p>Image captioning aims to generate natural and accurate textual descriptions of given images. Although significant progress has been made in image captioning models in recent years, most existing approaches heavily rely on high-quality image-text paired datasets that require expensive human annotation, thus limiting model scalability. Current unsupervised image captioning methods primarily focus on leveraging zero-shot learning capabilities of large pre-trained models (e.g. CLIP, GPT-2), yet still face persistent challenges including modality gaps, inefficient inference, and excessive noise incorporation, which constrain model accuracy and generalization capabilities. To address these limitations, we propose SGDR-Cap ( Similarity- Guided Denoising Reconstruction for Captioning), a novel unsupervised image captioning method that bridges the vision-language modality gap through a similarity-guided denoising reconstruction module. Our method leverages similarity information to guide the reconstruction of authentic text features during caption generation while simultaneously forcing the model to learn how to extract crucial image-relevant features and filter out unnecessary noise information. This enhances both coarse- and fine-grained cross-modal alignment. Furthermore, our approach jointly optimizes denoising reconstruction loss and language modeling loss, ensuring accuracy and fluency, and promoting greater diversity. Extensive evaluations on the MSCOCO and Flickr30K benchmarks demonstrate that our method achieves state-of-the-art results across all major metrics, with the most notable gain on the CIDEr score, improving from 101.1 to 104.4.<\/jats:p>","DOI":"10.1142\/s0218001426590056","type":"journal-article","created":{"date-parts":[[2026,1,14]],"date-time":"2026-01-14T04:01:51Z","timestamp":1768363311000},"source":"Crossref","is-referenced-by-count":0,"title":["Similarity-Guided Denoising Reconstruction for Unsupervised Image Captioning"],"prefix":"10.1142","volume":"40","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-7751-6200","authenticated-orcid":false,"given":"Dongnan","family":"Yang","sequence":"first","affiliation":[{"name":"School of Computer, Electronics and Information, Guangxi University Nanning 530004, Guangxi, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3489-1887","authenticated-orcid":false,"given":"Lina","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Computer, Electronics and Information, Guangxi University Nanning 530004, Guangxi, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0674-9915","authenticated-orcid":false,"given":"Thomas","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering, Guangxi University, Nanning 530004, Guangxi, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-1921-1651","authenticated-orcid":false,"given":"Xichun","family":"Li","sequence":"additional","affiliation":[{"name":"School of Mathematics and Computer Science, Guangxi Normal University for Nationalities, Chongzuo 532200, Guangxi, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6887-130X","authenticated-orcid":false,"given":"Yuan Yan","family":"Tang","sequence":"additional","affiliation":[{"name":"Faculty of Science and Technology, UOW College Hong Kong, Hong Kong, P. R. China"},{"name":"Faculty of Science and Technology, University of Macau, Macau 999078, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9336-3155","authenticated-orcid":false,"given":"Patrick","family":"Shen-Pei Wang","sequence":"additional","affiliation":[{"name":"Department of Computer and Information Science, Khoury College of Computer Sciences Northeastern University, Boston MA 02115, Guangxi, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2026,2,21]]},"reference":[{"key":"S0218001426590056BIB001","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00130"},{"key":"S0218001426590056BIB002","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"S0218001426590056BIB003","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46454-1_24"},{"key":"S0218001426590056BIB004","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00512"},{"key":"S0218001426590056BIB005","first-page":"1877","volume":"33","author":"Brown T.","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"S0218001426590056BIB006","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2018\/84"},{"key":"S0218001426590056BIB007","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.667"},{"key":"S0218001426590056BIB008","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01059"},{"key":"S0218001426590056BIB009","volume-title":"Proc. 37th Int. Conf. Neural Information Processing Systems, NIPS \u201923","author":"Dai W.","year":"2024"},{"key":"S0218001426590056BIB010","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/W14-3348"},{"key":"S0218001426590056BIB011","unstructured":"J. Devlin, BERT: Pre-training of deep bidirectional transformers for language understanding, preprint (2018), arXiv:1810.04805."},{"key":"S0218001426590056BIB012","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475439"},{"key":"S0218001426590056BIB013","first-page":"2672","volume-title":"Proc. IEEE\/CVF Int. Conf. Computer Vision","author":"Gu S.","year":"2023"},{"key":"S0218001426590056BIB014","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58520-4_25"},{"key":"S0218001426590056BIB015","volume-title":"Advances in Neural Information Processing Systems","volume":"32","author":"Herdade S.","year":"2019"},{"key":"S0218001426590056BIB016","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00473"},{"key":"S0218001426590056BIB017","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"S0218001426590056BIB018","first-page":"19730","volume-title":"Int. Conf. Machine Learning","author":"Li J.","year":"2023"},{"key":"S0218001426590056BIB019","unstructured":"W. Li, L. Zhu, L. Wen and Y. Yang, Decap: Decoding clip latents for zero-shot captioning via text-only training, preprint (2023), arXiv:2303.03032."},{"key":"S0218001426590056BIB020","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2025.130622"},{"key":"S0218001426590056BIB021","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58577-8_8"},{"key":"S0218001426590056BIB022","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"S0218001426590056BIB023","doi-asserted-by":"publisher","DOI":"10.3115\/1218955.1219032"},{"key":"S0218001426590056BIB024","first-page":"3864","volume":"38","author":"Liu Z.","year":"2024","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"S0218001426590056BIB025","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.345"},{"key":"S0218001426590056BIB026","first-page":"2286","volume":"35","author":"Luo Y.","year":"2021","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"S0218001426590056BIB027","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.07.028"},{"key":"S0218001426590056BIB028","first-page":"4089","volume":"38","author":"Ma F.","year":"2024","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"S0218001426590056BIB029","doi-asserted-by":"publisher","DOI":"10.23919\/ELECO47770.2019.8990630"},{"key":"S0218001426590056BIB030","unstructured":"R. Mokady, A. Hertz and A. H. Bermano, Clipcap: Clip prefix for image captioning, preprint (2021), arXiv:2111.09734."},{"key":"S0218001426590056BIB031","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475604"},{"key":"S0218001426590056BIB032","doi-asserted-by":"crossref","unstructured":"D. Nukrai, R. Mokady and A. Globerson, Text-only training for image captioning using noise-injected clip, preprint (2022), arXiv:2211.00575.","DOI":"10.18653\/v1\/2022.findings-emnlp.299"},{"key":"S0218001426590056BIB033","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01098"},{"key":"S0218001426590056BIB034","first-page":"311","volume-title":"Proc. 40th Annu. Meeting Association for Computational Linguistics","author":"Papineni K.","year":"2002"},{"key":"S0218001426590056BIB035","volume-title":"Advances in Neural Information Processing Systems","volume":"32","author":"Paszke A.","year":"2019"},{"key":"S0218001426590056BIB036","doi-asserted-by":"publisher","DOI":"10.18637\/jss.v109.i03"},{"key":"S0218001426590056BIB037","unstructured":"A. Radford and K. Narasimhan, Improving Language Understanding by Generative Pre-Training, OpenAI Blog, 2018. Available: https:\/\/api.semanticscholar.org\/."},{"issue":"8","key":"S0218001426590056BIB038","first-page":"9","volume":"1","author":"Radford A.","year":"2019","journal-title":"OpenAI Blog"},{"key":"S0218001426590056BIB039","first-page":"8748","volume-title":"Int. Conf. Machine Learning","author":"Radford A.","year":"2021"},{"issue":"140","key":"S0218001426590056BIB040","first-page":"1","volume":"21","author":"Raffel C.","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"S0218001426590056BIB041","unstructured":"A. Ramesh, P. Dhariwal, A. Nichol, C. Chu and M. Chen, Hierarchical text-conditional image generation with clip latents,\n                      arXiv\n                      1\n                      (2) (2022) 3, arXiv:2204.06125."},{"key":"S0218001426590056BIB042","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"S0218001426590056BIB043","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.131"},{"key":"S0218001426590056BIB044","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"S0218001426590056BIB045","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-023-17849-7"},{"key":"S0218001426590056BIB046","unstructured":"Y. Su, T. Lan, Y. Liu, F. Liu, D. Yogatama, Y. Wang, L. Kong and N. Collier, Language models can see: Plugging visual controls in text generation, preprint (2022), arXiv:2205.02655."},{"key":"S0218001426590056BIB047","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01739"},{"key":"S0218001426590056BIB048","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-022-13443-5"},{"key":"S0218001426590056BIB049","volume-title":"Advances in Neural Information Processing Systems","volume":"30","author":"Vaswani A.","year":"2017"},{"key":"S0218001426590056BIB050","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"S0218001426590056BIB051","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"S0218001426590056BIB052","unstructured":"T. Wang, F. Li, L. Zhu, J. Li, Z. Zhang and H. T. Shen, Cross-modal retrieval: A systematic review of methods and future directions, preprint (2023), arXiv:2308.14263."},{"key":"S0218001426590056BIB053","first-page":"2585","volume":"36","author":"Wang Y.","year":"2022","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"S0218001426590056BIB054","unstructured":"Y. Wang, H. Luo, J. Xu, Y. Sun and F. Wang, Text data-centric image captioning with interactive prompts, preprint (2024), arXiv:2403.19193."},{"key":"S0218001426590056BIB055","unstructured":"J. Wang, Y. Zhang, M. Yan, J. Zhang and J. Sang, Zero-shot image captioning by anchor-augmented vision-language space alignment, preprint (2022), arXiv:2211.07275."},{"key":"S0218001426590056BIB056","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2020.3036860"},{"key":"S0218001426590056BIB057","first-page":"2048","volume-title":"Proc. 32nd Int. Conf. Machine Learning","volume":"37","author":"Xu K.","year":"2015"},{"key":"S0218001426590056BIB058","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.503"},{"key":"S0218001426590056BIB059","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00166"},{"key":"S0218001426590056BIB060","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2947482"},{"key":"S0218001426590056BIB061","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01075"}],"container-title":["International Journal of Pattern Recognition and Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0218001426590056","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T06:22:11Z","timestamp":1774938131000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/10.1142\/S0218001426590056"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,21]]},"references-count":61,"journal-issue":{"issue":"06","published-print":{"date-parts":[[2026,5]]}},"alternative-id":["10.1142\/S0218001426590056"],"URL":"https:\/\/doi.org\/10.1142\/s0218001426590056","relation":{},"ISSN":["0218-0014","1793-6381"],"issn-type":[{"value":"0218-0014","type":"print"},{"value":"1793-6381","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,21]]},"article-number":"2659005"}}