{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T19:54:11Z","timestamp":1783972451704,"version":"3.55.0"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2023,8,24]],"date-time":"2023-08-24T00:00:00Z","timestamp":1692835200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62172376"],"award-info":[{"award-number":["62172376"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62072418"],"award-info":[{"award-number":["62072418"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["202042008"],"award-info":[{"award-number":["202042008"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Key Research and Development Program of China","award":["2021YFF0704000"],"award-info":[{"award-number":["2021YFF0704000"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,1,31]]},"abstract":"<jats:p>Image-text retrieval in remote sensing aims to provide flexible information for data analysis and application. In recent years, state-of-the-art methods are dedicated to \u201cscale decoupling\u201d and \u201csemantic decoupling\u201d strategies to further enhance the capability of representation. However, these previous approaches focus on either the disentangling scale or semantics but ignore merging these two ideas in a union model, which extremely limits the performance of cross-modal retrieval models. To address these issues, we propose a novel Scale-Semantic Joint Decoupling Network (SSJDN) for remote sensing image-text retrieval. Specifically, we design the Bidirectional Scale Decoupling (BSD) module, which exploits Salience Extraction Map (SEM) and Salience Suppression Map (SSM) units to adaptively extract potential features and suppress cumbersome features at other scales in a bidirectional pattern to yield different scale clues. Besides, we design the Label-supervised Semantic Decoupling (LSD) module by leveraging the category semantic labels as prior knowledge to supervise images and texts probing significant semantic-related information. Finally, we design a Semantic-guided Triple Loss (STL), which adaptively generates a constant to adjust the loss function to improve the probability of matching the same semantic image and text and shorten the convergence time of the retrieval model. Our proposed SSJDN outperforms state-of-the-art approaches in numerical experiments conducted on four benchmark remote sensing datasets.<\/jats:p>","DOI":"10.1145\/3603628","type":"journal-article","created":{"date-parts":[[2023,6,7]],"date-time":"2023-06-07T11:40:51Z","timestamp":1686138051000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":16,"title":["Scale-Semantic Joint Decoupling Network for Image-Text Retrieval in Remote Sensing"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5948-0032","authenticated-orcid":false,"given":"Chengyu","family":"Zheng","sequence":"first","affiliation":[{"name":"College of Information Science and Engineering, Ocean University of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4209-7387","authenticated-orcid":false,"given":"Ning","family":"Song","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Ocean University of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-9186-8864","authenticated-orcid":false,"given":"Ruoyu","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Ocean University of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4087-3677","authenticated-orcid":false,"given":"Lei","family":"Huang","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Ocean University of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2830-8301","authenticated-orcid":false,"given":"Zhiqiang","family":"Wei","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Ocean University of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4952-7666","authenticated-orcid":false,"given":"Jie","family":"Nie","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Ocean University of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,8,24]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.3390\/rs12030405"},{"key":"e_1_3_1_3_2","article-title":"LSCIDMR: Large-scale satellite cloud image database for meteorological research","author":"Bai Cong","year":"2021","unstructured":"Cong Bai, Minjing Zhang, Jinglin Zhang, Jianwei Zheng, and Shengyong Chen. 2021. LSCIDMR: Large-scale satellite cloud image database for meteorological research. IEEE Transactions on Cybernetics (2021).","journal-title":"IEEE Transactions on Cybernetics"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTARS.2022.3202246"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"e_1_3_1_6_2","first-page":"522","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Chen Tianshui","year":"2019","unstructured":"Tianshui Chen, Muxin Xu, Xiaolu Hui, Hefeng Wu, and Liang Lin. 2019. Learning semantic-specific graph representation for multi-label image recognition. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 522\u2013531."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTARS.2021.3070872"},{"key":"e_1_3_1_8_2","article-title":"Vse++: Improving visual-semantic embeddings with hard negatives","author":"Faghri Fartash","year":"2017","unstructured":"Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2017. Vse++: Improving visual-semantic embeddings with hard negatives. arXiv preprint arXiv:1707.05612 (2017).","journal-title":"arXiv preprint arXiv:1707.05612"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTARS.2020.3013818"},{"key":"e_1_3_1_11_2","first-page":"1","volume-title":"Proceedings of the 2020 Mediterranean and Middle-East Geoscience and Remote Sensing Symposium (M2GARSS)","author":"Hoxha Genc","year":"2020","unstructured":"Genc Hoxha, Farid Melgani, and Jacopo Slaghenauffi. 2020. A new CNN-RNN framework for remote sensing image captioning. In Proceedings of the 2020 Mediterranean and Middle-East Geoscience and Remote Sensing Symposium (M2GARSS). IEEE, 1\u20134."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00645"},{"key":"e_1_3_1_13_2","first-page":"3294","volume-title":"Advances in Neural Information Processing Systems","author":"Kiros Ryan","year":"2015","unstructured":"Ryan Kiros, Yukun Zhu, Russ R. Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Skip-thought vectors. In Advances in Neural Information Processing Systems. 3294\u20133302."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01225-0_13"},{"issue":"6","key":"e_1_3_1_15_2","first-page":"5246","article-title":"Truncation cross entropy loss for remote sensing image captioning","volume":"59","author":"Li Xuelong","year":"2020","unstructured":"Xuelong Li, Xueting Zhang, Wei Huang, and Qi Wang. 2020. Truncation cross entropy loss for remote sensing image captioning. IEEE Transactions on Geoscience and Remote Sensing 59, 6 (2020), 5246\u20135257.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_16_2","first-page":"4324","volume-title":"Proceedings of the -2019 IEEE International Geoscience and Remote Sensing Symposium (IGARSS 2019)","author":"Liu Chao","year":"2019","unstructured":"Chao Liu, Jingjing Ma, Xu Tang, Xiangrong Zhang, and Licheng Jiao. 2019. Adversarial hash-code learning for remote sensing image retrieval. In Proceedings of the -2019 IEEE International Geoscience and Remote Sensing Symposium (IGARSS 2019). IEEE, 4324\u20134327."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTARS.2020.2981372"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2020.2984703"},{"issue":"3","key":"e_1_3_1_19_2","first-page":"1985","article-title":"Sound active attention framework for remote sensing image captioning","volume":"58","author":"Lu Xiaoqiang","year":"2019","unstructured":"Xiaoqiang Lu, Binqiang Wang, and Xiangtao Zheng. 2019. Sound active attention framework for remote sensing image captioning. IEEE Transactions on Geoscience and Remote Sensing 58, 3 (2019), 1985\u20132000.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2017.2776321"},{"key":"e_1_3_1_21_2","first-page":"1","article-title":"Fusion-based correlation learning model for cross-modal remote sensing image retrieval","volume":"19","author":"Lv Yafei","year":"2021","unstructured":"Yafei Lv, Wei Xiong, Xiaohan Zhang, and Yaqi Cui. 2021. Fusion-based correlation learning model for cross-modal remote sensing image retrieval. IEEE Geoscience and Remote Sensing Letters 19 (2021), 1\u20135.","journal-title":"IEEE Geoscience and Remote Sensing Letters"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1080\/01431161.2017.1399472"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/LGRS.2018.2845549"},{"key":"e_1_3_1_24_2","first-page":"1","volume-title":"Proceedings of the 2016 International Conference on Computer, Information and Telecommunication Systems (CITS)","author":"Qu Bo","year":"2016","unstructured":"Bo Qu, Xuelong Li, Dacheng Tao, and Xiaoqiang Lu. 2016. Deep semantic understanding of high resolution remote sensing image. In Proceedings of the 2016 International Conference on Computer, Information and Telecommunication Systems (CITS). IEEE, 1\u20135."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24574-4_28"},{"issue":"6","key":"e_1_3_1_26_2","doi-asserted-by":"crossref","first-page":"3623","DOI":"10.1109\/TGRS.2017.2677464","article-title":"Can a machine generate humanlike language descriptions for a remote sensing image?","volume":"55","author":"Shi Zhenwei","year":"2017","unstructured":"Zhenwei Shi and Zhengxia Zou. 2017. Can a machine generate humanlike language descriptions for a remote sensing image? IEEE Transactions on Geoscience and Remote Sensing 55, 6 (2017), 3623\u20133634.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.dsp.2020.102765"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"issue":"8","key":"e_1_3_1_29_2","doi-asserted-by":"crossref","first-page":"1274","DOI":"10.1109\/LGRS.2019.2893772","article-title":"Semantic descriptions of high-resolution remote sensing images","volume":"16","author":"Wang Binqiang","year":"2019","unstructured":"Binqiang Wang, Xiaoqiang Lu, Xiangtao Zheng, and Xuelong Li. 2019. Semantic descriptions of high-resolution remote sensing images. IEEE Geoscience and Remote Sensing Letters 16, 8 (2019), 1274\u20131278.","journal-title":"IEEE Geoscience and Remote Sensing Letters"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350875"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00586"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTARS.2020.3021390"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/1869790.1869829"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTARS.2022.3226325"},{"issue":"11","key":"e_1_3_1_36_2","doi-asserted-by":"crossref","first-page":"364","DOI":"10.3390\/ijgi6110364","article-title":"Multiple feature hashing learning for large-scale remote sensing image retrieval","volume":"6","author":"Ye Dongjie","year":"2017","unstructured":"Dongjie Ye, Yansheng Li, Chao Tao, Xunwei Xie, and Xiang Wang. 2017. Multiple feature hashing learning for large-scale remote sensing image retrieval. ISPRS International Journal of Geo-Information 6, 11 (2017), 364.","journal-title":"ISPRS International Journal of Geo-Information"},{"key":"e_1_3_1_37_2","article-title":"Text-image matching for cross-modal remote sensing image retrieval via graph neural network","author":"Yu Hongfeng","year":"2022","unstructured":"Hongfeng Yu, Fanglong Yao, Wanxuan Lu, Nayu Liu, Peiguang Li, Hongjian You, and Xian Sun. 2022. Text-image matching for cross-modal remote sensing image retrieval via graph neural network. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (2022).","journal-title":"IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing"},{"key":"e_1_3_1_38_2","article-title":"Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval","author":"Yuan Zhiqiang","year":"2022","unstructured":"Zhiqiang Yuan, Wenkai Zhang, Kun Fu, Xuan Li, Chubo Deng, Hongqi Wang, and Xian Sun. 2022. Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval. IEEE Transactions on Geoscience and Remote Sensing (2022).","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_39_2","first-page":"1","article-title":"A lightweight multi-scale crossmodal text-image retrieval method in remote sensing","volume":"60","author":"Yuan Zhiqiang","year":"2021","unstructured":"Zhiqiang Yuan, Wenkai Zhang, Xuee Rong, Xuan Li, Jialiang Chen, Hongqi Wang, Kun Fu, and Xian Sun. 2021. A lightweight multi-scale crossmodal text-image retrieval method in remote sensing. IEEE Transactions on Geoscience and Remote Sensing 60 (2021), 1\u201319.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jag.2022.103071"},{"key":"e_1_3_1_41_2","first-page":"1","article-title":"Remote sensing cross-modal text-image retrieval based on global and local information","volume":"60","author":"Yuan Zhiqiang","year":"2022","unstructured":"Zhiqiang Yuan, Wenkai Zhang, Changyuan Tian, Xuee Rong, Zhengyuan Zhang, Hongqi Wang, Kun Fu, and Xian Sun. 2022. Remote sensing cross-modal text-image retrieval based on global and local information. IEEE Transactions on Geoscience and Remote Sensing 60 (2022), 1\u201316.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"issue":"4","key":"e_1_3_1_42_2","doi-asserted-by":"crossref","first-page":"2175","DOI":"10.1109\/TGRS.2014.2357078","article-title":"Saliency-guided unsupervised feature learning for scene classification","volume":"53","author":"Zhang Fan","year":"2014","unstructured":"Fan Zhang, Bo Du, and Liangpei Zhang. 2014. Saliency-guided unsupervised feature learning for scene classification. IEEE Transactions on Geoscience and Remote Sensing 53, 4 (2014), 2175\u20132184.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_43_2","first-page":"15661","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Kun","year":"2022","unstructured":"Kun Zhang, Zhendong Mao, Quan Wang, and Yongdong Zhang. 2022. Negative-aware attention framework for image-text matching. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 15661\u201315670."},{"key":"e_1_3_1_44_2","doi-asserted-by":"crossref","first-page":"4798","DOI":"10.1109\/IGARSS.2017.8128075","volume-title":"Proceedings of the 2017 IEEE International Geoscience and Remote Sensing Symposium (IGARSS)","author":"Zhang Xiangrong","year":"2017","unstructured":"Xiangrong Zhang, Xiang Li, Jinliang An, Li Gao, Biao Hou, and Chen Li. 2017. Natural language description of remote sensing images based on deep learning. In Proceedings of the 2017 IEEE International Geoscience and Remote Sensing Symposium (IGARSS). IEEE, 4798\u20134801."},{"key":"e_1_3_1_45_2","first-page":"10039","volume-title":"-Proceedings of the 2019 IEEE International Geoscience and Remote Sensing Symposium (","author":"Zhang Xueting","year":"2019","unstructured":"Xueting Zhang, Qi Wang, Shangdong Chen, and Xuelong Li. 2019. Multi-scale cropping mechanism for remote sensing image captioning. In -Proceedings of the 2019 IEEE International Geoscience and Remote Sensing Symposium (IGARSS 2019). IEEE, 10039\u201310042."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2912290"},{"issue":"10","key":"e_1_3_1_47_2","doi-asserted-by":"crossref","first-page":"4514","DOI":"10.1109\/TNNLS.2020.3018790","article-title":"Inductive structure consistent hashing via flexible semantic calibration","volume":"32","author":"Zhang Zheng","year":"2020","unstructured":"Zheng Zhang, Luyao Liu, Yadan Luo, Zi Huang, Fumin Shen, Heng Tao Shen, and Guangming Lu. 2020. Inductive structure consistent hashing via flexible semantic calibration. IEEE Transactions on Neural Networks and Learning Systems 32, 10 (2020), 4514\u20134528.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_1_48_2","article-title":"Modality-invariant asymmetric networks for cross-modal hashing","author":"Zhang Zheng","year":"2022","unstructured":"Zheng Zhang, Haoyang Luo, Lei Zhu, Guangming Lu, and Heng Tao Shen. 2022. Modality-invariant asymmetric networks for cross-modal hashing. IEEE Transactions on Knowledge and Data Engineering (2022).","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_3_1_49_2","first-page":"1080652","volume-title":"Proceedings of the 10th International Conference on Digital Image Processing (ICDIP 2018)","volume":"10806","author":"Zou Chang","year":"2018","unstructured":"Chang Zou, Showhong Wan, Peiquan Jin, and Xingyue Li. 2018. A novel rotation invariance hashing network for fast remote sensing image retrieval. In Proceedings of the 10th International Conference on Digital Image Processing (ICDIP 2018), Vol. 10806. International Society for Optics and Photonics, 1080652."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3603628","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3603628","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:21Z","timestamp":1750178241000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3603628"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,24]]},"references-count":48,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,1,31]]}},"alternative-id":["10.1145\/3603628"],"URL":"https:\/\/doi.org\/10.1145\/3603628","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,24]]},"assertion":[{"value":"2022-12-12","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-05-22","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-08-24","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}