{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T11:05:57Z","timestamp":1775646357292,"version":"3.50.1"},"reference-count":41,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T00:00:00Z","timestamp":1775606400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100003453","name":"Natural Science Foundation of Guangdong Province","doi-asserted-by":"publisher","award":["2024A1515510030"],"award-info":[{"award-number":["2024A1515510030"]}],"id":[{"id":"10.13039\/501100003453","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Guangzhou Municipal Education Bureau\u2019s University Scientific Research Project","award":["2024312014"],"award-info":[{"award-number":["2024312014"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IJGI"],"abstract":"<jats:p>Remote sensing object detection is fundamental to Earth observation, yet remains challenging when relying on a single sensing modality. While optical imagery provides rich spatial and textural details, it is highly sensitive to illumination and adverse weather; conversely, Synthetic Aperture Radar (SAR) offers robust all-weather acquisition but suffers from speckle noise and limited semantic interpretability. To address these limitations, we leverage the potential of foundation models for optical\u2013SAR object detection via a novel gated\u2013guided fusion approach. By integrating transferable and generalizable representations from foundation models into the detection pipeline, we enhance semantic expressiveness and cross-environment robustness. Specifically, a gated\u2013guided fusion mechanism is designed to selectively merge cross-modal features with foundational priors, enabling the network to prioritize informative cues while suppressing unreliable signals in complex scenes. Furthermore, we propose a dual-stream architecture incorporating attention mechanisms and State Space Models (SSMs) to simultaneously capture local and long-range dependencies. Extensive experiments on the large-scale M4-SAR dataset demonstrate that our method achieves state-of-the-art performance, significantly improving detection accuracy and robustness under challenging sensing conditions.<\/jats:p>","DOI":"10.3390\/ijgi15040160","type":"journal-article","created":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T09:49:44Z","timestamp":1775641784000},"page":"160","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Harnessing Foundation Models for Optical\u2013SAR Object Detection via Gated\u2013Guided Fusion"],"prefix":"10.3390","volume":"15","author":[{"given":"Qianyin","family":"Jiang","sequence":"first","affiliation":[{"name":"School of Artificial Intelligence, Guangzhou Maritime University, Guangzhou 510725, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianshang","family":"Liao","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, Guangzhou Maritime University, Guangzhou 510725, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qiuyu","family":"Lin","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, Guangzhou Maritime University, Guangzhou 510725, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junkang","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Science, Nanjing University of Posts and Telecommunications, Nanjing 210023, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,4,8]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_2","unstructured":"Simonyan, K., and Zisserman, A. (2015, January 7\u20139). Very deep convolutional networks for large-scale image recognition. Proceedings of the 3rd International Conference on Learning Representations (ICLR 2015), San Diego, CA, USA."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"2350011","DOI":"10.1142\/S0218001423500118","article-title":"A Novel Copy\u2013Move Forgery Detection Algorithm via Gradient-Hash Matching and Simplified Cluster-Based Filtering","volume":"37","author":"Yang","year":"2023","journal-title":"Int. J. Pattern Recognit. Artif. Intell."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Andrew, O., Apan, A., Paudyal, D.R., and Perera, K. (2023). Convolutional Neural Network-Based Deep Learning Approach for Automatic Flood Mapping Using NovaSAR-1 and Sentinel-1 Data. ISPRS Int. J. Geo. Inf., 12.","DOI":"10.3390\/ijgi12050194"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"200","DOI":"10.1080\/2150704X.2024.2442111","article-title":"Piecewise Self-Adaption Weighted attention for the detection of concentrated distributions of ships in SAR images","volume":"16","author":"Guo","year":"2025","journal-title":"Remote Sens. Lett."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Wan, S., Yeh, M.L., and Ma, H.L. (2021). An Innovative Intelligent System with Integrated CNN and SVM: Considering Various Crops through Hyperspectral Image Data. ISPRS Int. J. Geo. Inf., 10.","DOI":"10.3390\/ijgi10040242"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"512","DOI":"10.1080\/2150704X.2023.2215892","article-title":"TWC-AWT-Net: A transformer-based method for detecting ships in noisy SAR images","volume":"14","author":"Yu","year":"2023","journal-title":"Remote Sens. Lett."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Aleissaee, A.A., Kumar, A., Anwer, R.M., Khan, S., Cholakkal, H., Xia, G.S., and Khan, F.S. (2023). Transformers in Remote Sensing: A Survey. Remote Sens., 15.","DOI":"10.3390\/rs15071860"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zhang, J., Zhao, H., and Li, J. (2021). TRS: Transformers for Remote Sensing Scene Classification. Remote Sens., 13.","DOI":"10.3390\/rs13204143"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Wang, J., Li, H., Li, Y., and Qin, Z. (2025). A Lightweight CNN-Transformer Implemented via Structural Re-Parameterization and Hybrid Attention for Remote Sensing Image Super-Resolution. Isprs Int. J. Geo. Inf., 14.","DOI":"10.3390\/ijgi14010008"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Ding, K., Wang, Y., Wang, C., and Ma, J. (2024). A New Subject-Sensitive Hashing Algorithm Based on Multi-PatchDrop and Swin-Unet for the Integrity Authentication of HRRS Image. ISPRS Int. J. Geo. Inf., 13.","DOI":"10.3390\/ijgi13090336"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"13995","DOI":"10.1109\/JSTARS.2024.3435739","article-title":"A CNN-Transformer Combined Remote Sensing Imagery Spatiotemporal Fusion Model","volume":"17","author":"Jiang","year":"2024","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_13","unstructured":"Gu, A., and Dao, T. (2024). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv."},{"key":"ref_14","first-page":"103031","article-title":"VMamba: Visual State Space Model","volume":"Volume 37","author":"Liu","year":"2024","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_15","first-page":"628","article-title":"SpecSpatMamba: An efficient hyperspectral image classification method integrating spectral-spatial dual-path and state space model","volume":"28","author":"Liao","year":"2025","journal-title":"Egypt. J. Remote Sens. Space Sci."},{"key":"ref_16","first-page":"8002605","article-title":"RSMamba: Remote Sensing Image Classification with State Space Model","volume":"21","author":"Chen","year":"2024","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_17","first-page":"5626114","article-title":"Diffusion-Noise-Based Augmentation for Long-Tailed Remote Sensing Image Classification","volume":"63","author":"Wang","year":"2025","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Hao, X., Liu, L., Yang, R., Yin, L., Zhang, L., and Li, X. (2023). A Review of Data Augmentation Methods of Remote Sensing Image Target Recognition. Remote Sens., 15.","DOI":"10.3390\/rs15030827"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"6427","DOI":"10.1109\/TPAMI.2025.3557581","article-title":"HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model","volume":"47","author":"Wang","year":"2025","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_20","first-page":"5642123","article-title":"RS5M and GeoRSCLIP: A Large-Scale Vision- Language Dataset and a Large Vision-Language Model for Remote Sensing","volume":"62","author":"Zhang","year":"2024","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1235","DOI":"10.1038\/s42256-025-01078-8","article-title":"A semantic-enhanced multi-modal remote sensing foundation model for Earth observation","volume":"7","author":"Wu","year":"2025","journal-title":"Nat. Mach. Intell."},{"key":"ref_22","unstructured":"Wang, C., Lu, W., Li, X., Yang, J., and Luo, L. (2025). M4-SAR: A Multi-Resolution, Multi-Polarization, Multi-Scene, Multi-Source Dataset and Benchmark for Optical-SAR Fusion Object Detection. arXiv."},{"key":"ref_23","first-page":"5610112","article-title":"A New Learning Paradigm for Foundation Model-Based Remote-Sensing Change Detection","volume":"62","author":"Li","year":"2024","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_24","first-page":"5611711","article-title":"Adapting Segment Anything Model for Change Detection in VHR Remote Sensing Images","volume":"62","author":"Ding","year":"2024","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"121601","DOI":"10.1109\/ACCESS.2025.3587922","article-title":"RFHP-CD: A Prompt-Driven Fine-Tuning Framework of Remote Sensing Foundation Model for Building and Cropland Change Detection","volume":"13","author":"Wang","year":"2025","journal-title":"IEEE Access"},{"key":"ref_26","first-page":"5505005","article-title":"Incremental Classification of Cross-Scene Hyperspectral Images Based on Dual Constraints and Knowledge Transfer","volume":"22","author":"Wang","year":"2025","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Cao, Y., Bin, J., Hamari, J., Blasch, E., and Liu, Z. (2023, January 18\u201322). Multimodal Object Detection by Channel Switching and Spatial Attention. Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vancouver, BC, Canada.","DOI":"10.1109\/CVPRW59228.2023.00046"},{"key":"ref_28","unstructured":"Zhang, J., Cao, M., Xie, W., Lei, J., Li, D., Huang, W., Li, Y., and Yang, X. (2024, January 9\u201315). E2E-MFD: Towards end-to-end synchronous multimodal fusion detection. Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS\u201924), Red Hook, NY, USA."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Liu, B., Ren, B., Hou, B., and Gu, Y. (2023, January 16\u201321). Multi-Source Fusion Network for Remote Sensing Image Segmentation with Hierarchical Transformer. Proceedings of the IGARSS 2023\u20142023 IEEE International Geoscience and Remote Sensing Symposium, Pasadena, CA, USA.","DOI":"10.1109\/IGARSS52108.2023.10282984"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"7677","DOI":"10.1109\/JSTARS.2022.3203508","article-title":"Cloud Removal Based on SAR-Optical Remote Sensing Data Fusion via a Two-Flow Network","volume":"15","author":"Mao","year":"2022","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"7449","DOI":"10.1109\/TIV.2024.3398429","article-title":"Misaligned Visible-Thermal Object Detection: A Drone-Based Benchmark and Baseline","volume":"9","author":"Song","year":"2024","journal-title":"IEEE Trans. Intell. Veh."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wei, T., Chen, H., Wang, J., and Liu, W. (2024, January 20\u201322). MDFNet: Multimodal Feature Decomposition and Fusion Network for Multimodal Remote Sensing Image Semantic Segmentation. Proceedings of the 2024 IEEE International Conference on Signal, Information and Data Processing (ICSIDP), Zhuhai, China.","DOI":"10.1109\/ICSIDP62679.2024.10868650"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"108786","DOI":"10.1016\/j.patcog.2022.108786","article-title":"Cross-Modality Attentive Feature Fusion for Object Detection in Multispectral Remote Sensing Imagery","volume":"130","author":"Fang","year":"2022","journal-title":"Pattern Recognit."},{"key":"ref_34","unstructured":"Khanam, R., and Hussain, M. (2024). YOLOv11: An Overview of the Key Architectural Enhancements. arXiv."},{"key":"ref_35","unstructured":"Lin, Z., Nikishin, E., He, X., and Courville, A. (2025, January 24\u201328). Forgetting Transformer: Softmax Attention with a Forget Gate. Proceedings of the Thirteenth International Conference on Learning Representations, Singapore."},{"key":"ref_36","unstructured":"Chen, K., Wang, J., Pang, J., Cao, Y., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., and Xu, J. (2019). MMDetection: Open MMLab Detection Toolbox and Benchmark. arXiv."},{"key":"ref_37","unstructured":"Li, X., Wang, W., Wu, L., Chen, S., Hu, X., Li, J., Tang, J., and Yang, J. (2020, January 6\u201312). Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection. Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS\u201920), Red Hook, NY, USA."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft COCO: Common Objects in Context. Proceedings of the European Conference on Computer Vision (ECCV), Z\u00fcrich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_39","unstructured":"He, X., Tang, C., Zou, X., and Zhang, W. (November, January 29). Multispectral Object Detection via Cross-Modal Conflict-Aware Learning. Proceedings of the 31st ACM International Conference on Multimedia, New York, NY, USA."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"109913","DOI":"10.1016\/j.patcog.2023.109913","article-title":"ICAFusion: Iterative cross-attention guided feature fusion for multispectral object detection","volume":"145","author":"Shen","year":"2024","journal-title":"Pattern Recognit."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"11198","DOI":"10.1109\/TCSVT.2024.3418965","article-title":"MMI-Det: Exploring Multi-Modal Integration for Visible and Infrared Object Detection","volume":"34","author":"Zeng","year":"2024","journal-title":"IEEE Trans. Circuits Syst. Video Technol."}],"container-title":["ISPRS International Journal of Geo-Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2220-9964\/15\/4\/160\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T10:30:35Z","timestamp":1775644235000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2220-9964\/15\/4\/160"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,8]]},"references-count":41,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2026,4]]}},"alternative-id":["ijgi15040160"],"URL":"https:\/\/doi.org\/10.3390\/ijgi15040160","relation":{},"ISSN":["2220-9964"],"issn-type":[{"value":"2220-9964","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,8]]}}}