{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,20]],"date-time":"2026-04-20T14:55:21Z","timestamp":1776696921111,"version":"3.51.2"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"5","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>\n                    Remote sensing image segmentation poses significant challenges in generalizing to unseen categories during the evaluation phase. Existing open-vocabulary segmentation methods, primarily designed for natural images, struggle to cope with the spatial complexity, scale variation, and high-resolution characteristics of remote sensing imagery. Specifically, scale variations during inference can degrade performance, as the model tends to overfit to fixed-scale patterns encountered during training. This also affects the model\u2019s ability to recognize unseen or novel class objects appearing in varying sizes or resolutions during testing. These limitations increase the need for developing open-vocabulary segmentation methods addressing the challenges of geospatial images. In this work, we introduce\n                    <jats:italic toggle=\"yes\">AerOSeg<\/jats:italic>\n                    ++, an open-vocabulary segmentation method in remote sensing, focusing on scale-invariant feature learning. We first compute robust image-text correlation features using rotated input images and domain-specific prompts. These are refined via spatial and class refinement blocks, guided by SAM features to enhance spatial consistency. To upscale the refined correlation features, we propose a multi-scale decoder framework that fuses fine-grained texture features with SAM-derived features. By leveraging texture information across multiple receptive fields, AerOSeg++ effectively captures scale-consistent patterns, facilitating accurate segmentation of objects across varying spatial resolutions. Additionally, our training pipeline incorporates ScaleDrop, a computationally efficient parameter-free feature rescaling module ensuring scale-invariant feature representation learning. Our proposed model has shown significant performance gains compared to the state-of-the-art open-vocabulary methods when evaluated on three benchmark datasets for remote sensing\u2014iSAID, DLRSD, and OpenEarthMap. These results highlight the effectiveness of our scale-invariant design and texture-guided multi-scale feature upsampling in handling the challenges of open-vocabulary segmentation in remote sensing imagery.\n                  <\/jats:p>","DOI":"10.1145\/3787522","type":"journal-article","created":{"date-parts":[[2026,1,28]],"date-time":"2026-01-28T13:16:32Z","timestamp":1769606192000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["AerOSeg++: Scale-Aware and Texture-Guided Open-Vocabulary Segmentation with SAM Features for Remote Sensing Images"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6021-5407","authenticated-orcid":false,"given":"Saikat","family":"Dutta","sequence":"first","affiliation":[{"name":"IITB-Monash Research Academy, Mumbai, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-2964-4754","authenticated-orcid":false,"given":"Akhil","family":"Vasim","sequence":"additional","affiliation":[{"name":"Indian Institute of Technology Bombay, Mumbai, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8659-8773","authenticated-orcid":false,"given":"Hamid","family":"Rezatofighi","sequence":"additional","affiliation":[{"name":"Monash University, Melbourne, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8371-8138","authenticated-orcid":false,"given":"Biplab","family":"Banerjee","sequence":"additional","affiliation":[{"name":"Indian Institute of Technology Bombay, Mumbai, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,20]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.3390\/rs15071860"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jag.2022.102690"},{"key":"e_1_3_1_4_2","article-title":"Zero-shot semantic segmentation","volume":"32","author":"Bucher Maxime","year":"2019","unstructured":"Maxime Bucher, Tuan-Hung Vu, Matthieu Cord, and Patrick P\u00e9rez. 2019. Zero-shot semantic segmentation. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 32.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_5_2","unstructured":"Qinglong Cao Yuntian Chen Chao Ma and Xiaokang Yang. 2024. Open-vocabulary remote sensing image semantic segmentation. arXiv:2409.07683. Retrieved from https:\/\/arxiv.org\/abs\/2409.07683"},{"key":"e_1_3_1_6_2","first-page":"1","article-title":"Open-vocabulary high-resolution remote sensing image semantic segmentation","author":"Cao Qinglong","year":"2025","unstructured":"Qinglong Cao, Yuntian Chen, Chao Ma, and Xiaokang Yang. 2025. Open-vocabulary high-resolution remote sensing image semantic segmentation. IEEE Transactions on Geoscience and Remote Sensing 63 (2025), 1\u201314.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_7_2","unstructured":"Yice Cao Chenchen Liu Zhenhua Wu Wenxin Yao Liu Xiong Jie Chen and Zhixiang Huang. 2024. Remote sensing image segmentation using vision mamba and multi-scale multi-frequency feature fusion. arXiv:2410.05624. Retrieved from https:\/\/arxiv.org\/abs\/2410.05624"},{"issue":"2","key":"e_1_3_1_8_2","doi-asserted-by":"crossref","first-page":"1144","DOI":"10.1109\/TGRS.2017.2760909","article-title":"Multilabel remote sensing image retrieval using a semisupervised graph-theoretic method","volume":"56","author":"Chaudhuri Bindita","year":"2017","unstructured":"Bindita Chaudhuri, Beg\u00fcm Demir, Subhasis Chaudhuri, and Lorenzo Bruzzone. 2017. Multilabel remote sensing image retrieval using a semisupervised graph-theoretic method. IEEE Transactions on Geoscience and Remote Sensing 56, 2 (2017), 1144\u20131158.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00135"},{"key":"e_1_3_1_11_2","first-page":"4113","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Cho Seokju","year":"2024","unstructured":"Seokju Cho, Heeseong Shin, Sunghwan Hong, Anurag Arnab, Paul Hongsuck Seo, and Seungryong Kim. 2024. CAT-Seg: Cost aggregation for open-vocabulary semantic segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 4113\u20134123."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01129"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/tgrs.2020.2994150"},{"key":"e_1_3_1_14_2","first-page":"2254","volume-title":"Proceedings of the Computer Vision and Pattern Recognition Conference","author":"Dutta Saikat","year":"2025","unstructured":"Saikat Dutta, Akhil Vasim, Siddhant Gole, Hamid Rezatofighi, and Biplab Banerjee. 2025. AerOSeg: Harnessing SAM for open-vocabulary segmentation in remote sensing images. In Proceedings of the Computer Vision and Pattern Recognition Conference, 2254\u20132264."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20059-5_31"},{"key":"e_1_3_1_16_2","first-page":"4904","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Jia Chao","year":"2021","unstructured":"Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 4904\u20134916."},{"key":"e_1_3_1_17_2","doi-asserted-by":"crossref","unstructured":"Michael Kampffmeyer Arnt-B\u00f8rre Salberg and Robert Jenssen. 2016. Semantic segmentation of small objects and modeling of uncertainty in urban remote sensing images using deep convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops 1\u20139.","DOI":"10.1109\/CVPRW.2016.90"},{"key":"e_1_3_1_18_2","first-page":"5156","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Katharopoulos Angelos","year":"2020","unstructured":"Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Fran\u00e7ois Fleuret. 2020. Transformers are RNNs: Fast autoregressive transformers with linear attention. In Proceedings of the International Conference on Machine Learning. PMLR, 5156\u20135165."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"e_1_3_1_20_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Li Boyi","year":"2022","unstructured":"Boyi Li, Kilian Q. Weinberger, Serge Belongie, Vladlen Koltun, and Rene Ranftl. 2022. Language-driven semantic segmentation. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_21_2","unstructured":"Kaiyu Li Ruixun Liu Xiangyong Cao Xueru Bai Feng Zhou Deyu Meng and Zhi Wang. 2024. SegEarth-OV: Towards training-free open-vocabulary segmentation for remote sensing images. arXiv:2410.01768. Retrieved from https:\/\/arxiv.org\/abs\/2410.01768"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jag.2023.103497"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00335"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.3390\/rs13163054"},{"key":"e_1_3_1_27_2","unstructured":"Joshua Mitton and Roderick Murray-Smith. 2021. Rotation equivariant deforestation segmentation and driver classification. arXiv:2110.13097. Retrieved from https:\/\/arxiv.org\/abs\/2110.13097"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2025.3531930"},{"key":"e_1_3_1_29_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01196"},{"key":"e_1_3_1_31_2","unstructured":"Nikhila Ravi Valentin Gabeur Yuan-Ting Hu Ronghang Hu Chaitanya Ryali Tengyu Ma Haitham Khedr Roman R\u00e4dle Chloe Rolland Laura Gustafson et al. 2024. SAM 2: Segment anything in images and videos. arXiv:2408.00714. Retrieved from https:\/\/arxiv.org\/abs\/2408.00714"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.12"},{"key":"e_1_3_1_33_2","first-page":"28412","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Shan Xiangheng","year":"2024","unstructured":"Xiangheng Shan, Dongyue Wu, Guilin Zhu, Yuanjie Shao, Nong Sang, and Changxin Gao. 2024. Open-vocabulary semantic segmentation with image embedding balancing. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 28412\u201328421."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.3390\/rs10060964"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_1_36_2","article-title":"SAMRS: Scaling-up remote sensing segmentation dataset with segment anything model","volume":"36","author":"Wang Di","year":"2024","unstructured":"Di Wang, Jing Zhang, Bo Du, Minqiang Xu, Lin Liu, Dacheng Tao, and Liangpei Zhang. 2024. SAMRS: Scaling-up remote sensing segmentation dataset with segment anything model. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.3390\/rs14091956"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.isprsjprs.2022.06.008"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00401"},{"key":"e_1_3_1_40_2","unstructured":"Junde Wu Wei Ji Yuanpei Liu Huazhu Fu Min Xu Yanwu Xu and Yueming Jin. 2023. Medical SAM adapter: Adapting segment anything model for medical image segmentation. arXiv:2304.12620. Retrieved from https:\/\/arxiv.org\/abs\/2304.12620"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00418"},{"key":"e_1_3_1_42_2","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV)","author":"Xia Junshi","year":"2023","unstructured":"Junshi Xia, Naoto Yokoya, Bruno Adriano, and Clifford Broni-Bediako. 2023. OpenEarthMap: A benchmark dataset for global high-resolution land cover mapping. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV)."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00845"},{"key":"e_1_3_1_44_2","first-page":"3426","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Xie Bin","year":"2024","unstructured":"Bin Xie, Jiale Cao, Jin Xie, Fahad Shahbaz Khan, and Yanwei Pang. 2024. Sed: A simple encoder-decoder for open-vocabulary semantic segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 3426\u20133436."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00289"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00288"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.3390\/rs13010071"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.3390\/rs13183585"},{"key":"e_1_3_1_49_2","first-page":"1","article-title":"Scale-aware detailed matching for few-shot aerial image semantic segmentation","volume":"60","author":"Yao Xiwen","year":"2021","unstructured":"Xiwen Yao, Qinglong Cao, Xiaoxu Feng, Gong Cheng, and Junwei Han. 2021. Scale-aware detailed matching for few-shot aerial image semantic segmentation. IEEE Transactions on Geoscience and Remote Sensing 60 (2021), 1\u201311.","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_50_2","doi-asserted-by":"crossref","first-page":"9436","DOI":"10.1609\/aaai.v39i9.33022","article-title":"Towards open-vocabulary remote sensing image semantic segmentation","volume":"39","author":"Ye Chengyang","year":"2025","unstructured":"Chengyang Ye, Yunzhi Zhuge, and Pingping Zhang. 2025. Towards open-vocabulary remote sensing image semantic segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, 9436\u20139444.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_1_51_2","article-title":"Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip","volume":"36","author":"Yu Qihang","year":"2024","unstructured":"Qihang Yu, Ju He, Xueqing Deng, Xiaohui Shen, and Liang-Chieh Chen. 2024. Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2020.114417"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2022.3144894"},{"issue":"4","key":"e_1_3_1_54_2","doi-asserted-by":"crossref","first-page":"590","DOI":"10.3390\/rs17040590","article-title":"RSAM-Seg: A SAM-based model with prior knowledge integration for remote sensing image semantic segmentation","volume":"17","author":"Zhang Jie","year":"2025","unstructured":"Jie Zhang, Yunxin Li, Xubing Yang, Rui Jiang, and Li Zhang. 2025. RSAM-Seg: A SAM-based model with prior knowledge integration for remote sensing image semantic segmentation. Remote Sensing 17, 4 (2025), 590. Retrieved from https:\/\/www.proquest.com\/scholarly-journals\/rsam-seg-sam-based-model-with-prior-knowledge\/docview\/3171210456\/se-2","journal-title":"Remote Sensing"},{"key":"e_1_3_1_55_2","first-page":"3270","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Yi","year":"2024","unstructured":"Yi Zhang, Meng-Hao Guo, Miao Wang, and Shi-Min Hu. 2024. Exploring regional clues in CLIP for zero-shot semantic segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 3270\u20133280."},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2021.3085889"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01075"},{"key":"e_1_3_1_58_2","article-title":"Segment everything everywhere all at once","volume":"36","author":"Zou Xueyan","year":"2024","unstructured":"Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Wang, Lijuan Wang, Jianfeng Gao, and Yong Jae Lee. 2024. Segment everything everywhere all at once. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3787522","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,20]],"date-time":"2026-04-20T14:11:30Z","timestamp":1776694290000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3787522"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,20]]},"references-count":57,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3787522"],"URL":"https:\/\/doi.org\/10.1145\/3787522","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,20]]},"assertion":[{"value":"2025-07-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-03","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}