{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T19:47:26Z","timestamp":1784836046641,"version":"3.55.0"},"reference-count":51,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2024,2,25]],"date-time":"2024-02-25T00:00:00Z","timestamp":1708819200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Natural Science Basic Research Program of Shaanxi","award":["2024JC-YBMS-467"],"award-info":[{"award-number":["2024JC-YBMS-467"]}]},{"name":"Natural Science Basic Research Program of Shaanxi","award":["D023030002"],"award-info":[{"award-number":["D023030002"]}]},{"name":"Aeronautical Science Foundation of China","award":["2024JC-YBMS-467"],"award-info":[{"award-number":["2024JC-YBMS-467"]}]},{"name":"Aeronautical Science Foundation of China","award":["D023030002"],"award-info":[{"award-number":["D023030002"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Current CNN-based methods for infrared and visible image fusion are limited by the low discrimination of extracted structural features, the adoption of uniform loss functions, and the lack of inter-modal feature interaction, which make it difficult to obtain optimal fusion results. To alleviate the above problems, a framework for multimodal feature learning fusion using a cross-attention Transformer is proposed. To extract rich structural features at different scales, residual U-Nets with mixed receptive fields are adopted to capture salient object information at various granularities. Then, a hybrid attention fusion strategy is employed to integrate the complementing information from the input images. Finally, adaptive loss functions are designed to achieve optimal fusion results for different modal features. The fusion framework proposed in this study is thoroughly evaluated using the TNO, FLIR, and LLVIP datasets, encompassing diverse scenes and varying illumination conditions. In the comparative experiments, HATF achieved competitive results on three datasets, with EN, SD, MI, and SSIM metrics reaching the best performance on the TNO dataset, surpassing the second-best method by 2.3%, 18.8%, 4.2%, and 2.2%, respectively. These results validate the effectiveness of the proposed method in terms of both robustness and image fusion quality compared to several popular methods.<\/jats:p>","DOI":"10.3390\/rs16050803","type":"journal-article","created":{"date-parts":[[2024,2,26]],"date-time":"2024-02-26T10:40:17Z","timestamp":1708944017000},"page":"803","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["HATF: Multi-Modal Feature Learning for Infrared and Visible Image Fusion via Hybrid Attention Transformer"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2751-6096","authenticated-orcid":false,"given":"Xiangzeng","family":"Liu","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ziyao","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haojie","family":"Gao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiang","family":"Li","sequence":"additional","affiliation":[{"name":"NavInfo Co., Ltd., Beijing 100094, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lei","family":"Wang","sequence":"additional","affiliation":[{"name":"NavInfo Co., Ltd., Beijing 100094, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2872-388X","authenticated-orcid":false,"given":"Qiguang","family":"Miao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Xidian University, Xi\u2019an 710071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,2,25]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"153","DOI":"10.1016\/j.inffus.2018.02.004","article-title":"Infrared and visible image fusion methods and applications: A survey","volume":"45","author":"Ma","year":"2019","journal-title":"Inf. Fusion"},{"key":"ref_2","unstructured":"Tian, Z., Shen, C., Chen, H., and He, T. (November, January 27). Fcos: Fully convolutional one-stage object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"025041","DOI":"10.1088\/2631-8695\/ab9b36","article-title":"A monocular wide-field vision system for geolocation with uncertainties in urban scenes","volume":"2","author":"Arroyo","year":"2020","journal-title":"Eng. Res. Express"},{"key":"ref_4","first-page":"198","article-title":"Feature level image fusion of optical imagery and Synthetic Aperture Radar (SAR) for invasive alien plant species detection and mapping","volume":"10","author":"Rajah","year":"2018","journal-title":"Remote Sens. Appl. Soc. Environ."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"110","DOI":"10.1016\/j.inffus.2020.04.006","article-title":"Pan-GAN: An unsupervised pan-sharpening method for remote sensing image fusion","volume":"62","author":"Ma","year":"2020","journal-title":"Inf. Fusion"},{"key":"ref_6","first-page":"1","article-title":"A Dual-Domain Super-Resolution Image Fusion Method with SIRV and GALCA Model for PolSAR and Panchromatic Images","volume":"60","author":"Liu","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_7","first-page":"1","article-title":"Unaligned hyperspectral image fusion via registration and interpolation modeling","volume":"60","author":"Ying","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_8","unstructured":"Kumar, K.S., Kavitha, G., Subramanian, R., and Ramesh, G. (2011). MATLAB-A Ubiquitous Tool for the Practical Engineer, IntechOpen."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"375","DOI":"10.1016\/j.compeleceng.2016.09.019","article-title":"Image fusion based on object region detection and non-subsampled contourlet transform","volume":"62","author":"Meng","year":"2017","journal-title":"Comput. Electr. Eng."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"174","DOI":"10.1016\/j.infrared.2016.02.005","article-title":"Infrared and visible image fusion scheme based on NSCT and low-level visual features","volume":"76","author":"Li","year":"2016","journal-title":"Infrared Phys. Technol."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Toet, A., and Hogervorst, M.A. (2016, January 26\u201329). Multiscale image fusion through guided filtering. Proceedings of the Target and Background Signatures II. SPIE, Edinburgh, UK.","DOI":"10.1117\/12.2239945"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"21869","DOI":"10.1007\/s11042-017-4583-3","article-title":"An image fusion framework using novel dictionary based sparse representation","volume":"76","author":"Aishwarya","year":"2017","journal-title":"Multimed. Tools Appl."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"516","DOI":"10.1016\/j.ins.2017.09.010","article-title":"A novel multi-modality image fusion method based on image decomposition and sparse representation","volume":"432","author":"Zhu","year":"2018","journal-title":"Inf. Sci."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Li, H., and Wu, X.J. (2022). Infrared and visible image fusion using latent low-rank representation. arXiv.","DOI":"10.23919\/CISS51089.2021.9652254"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"4733","DOI":"10.1109\/TIP.2020.2975984","article-title":"MDLatLRR: A novel decomposition method for infrared and visible image fusion","volume":"29","author":"Li","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"91462","DOI":"10.1109\/ACCESS.2021.3090436","article-title":"Improving the performance of infrared and visible image fusion based on latent low-rank representation nested with rolling guided image filtering","volume":"9","author":"Gao","year":"2021","journal-title":"IEEE Access"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"3314","DOI":"10.1109\/TMM.2021.3096088","article-title":"Infrared and visible image fusion based on deep decomposition network and saliency analysis","volume":"24","author":"Jian","year":"2021","journal-title":"IEEE Trans. Multimed."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1016\/j.infrared.2017.02.005","article-title":"Infrared and visible image fusion based on visual saliency map and weighted least square optimization","volume":"82","author":"Ma","year":"2017","journal-title":"Infrared Phys. Technol."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Xu, H., Ma, J., Le, Z., Jiang, J., and Guo, X. (2020, January 7\u201312). Fusiondn: A unified densely connected network for image fusion. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6936"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1016\/j.inffus.2018.09.004","article-title":"FusionGAN: A generative adversarial network for infrared and visible image fusion","volume":"48","author":"Ma","year":"2019","journal-title":"Inf. Fusion"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"79","DOI":"10.1016\/j.inffus.2022.03.007","article-title":"PIAFusion: A progressive infrared and visible image fusion network based on illumination aware","volume":"83","author":"Tang","year":"2022","journal-title":"Inf. Fusion"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1850018","DOI":"10.1142\/S0219691318500182","article-title":"Infrared and visible image fusion with convolutional neural networks","volume":"16","author":"Liu","year":"2018","journal-title":"Int. J. Wavelets Multiresolut. Inf. Process."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"2614","DOI":"10.1109\/TIP.2018.2887342","article-title":"DenseFuse: A fusion approach to infrared and visible images","volume":"28","author":"Li","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"9645","DOI":"10.1109\/TIM.2020.3005230","article-title":"NestFuse: An infrared and visible image fusion architecture based on nest connection and spatial\/channel attention models","volume":"69","author":"Li","year":"2020","journal-title":"IEEE Trans. Instrum. Meas."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1016\/j.inffus.2021.02.023","article-title":"RFN-Nest: An end-to-end residual fusion network for infrared and visible images","volume":"73","author":"Li","year":"2021","journal-title":"Inf. Fusion"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"71","DOI":"10.1016\/j.neucom.2023.01.033","article-title":"THFuse: An infrared and visible image fusion network using transformer and hybrid feature extractor","volume":"527","author":"Chen","year":"2023","journal-title":"Neurocomputing"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"3159","DOI":"10.1109\/TCSVT.2023.3234340","article-title":"DATFuse: Infrared and visible image fusion via dual attention transformer","volume":"33","author":"Tang","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"770","DOI":"10.1109\/TCSVT.2023.3289170","article-title":"Cross-Modal Transformers for Infrared and Visible Image Fusion","volume":"34","author":"Park","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the 18th International Conference\u2014Medical Image Computing and Computer-Assisted Intervention\u2013MICCAI 2015, Munich, Germany. Part III 18.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_30","first-page":"9856669","article-title":"Fault Recognition Method Based on Attention Mechanism and the 3D-UNet","volume":"2022","author":"Yu","year":"2022","journal-title":"Comput. Intell. Neurosci."},{"key":"ref_31","unstructured":"Soni, A., Koner, R., and Villuri, V.G.K. (2019, January 12\u201314). M-unet: Modified u-net segmentation framework with satellite imagery. Proceedings of the Global AI Congress 2019, Kolkata, India."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"5008854","DOI":"10.1155\/2022\/5008854","article-title":"Automatic building extraction on satellite images using Unet and ResNet50","volume":"2022","author":"Alsabhan","year":"2022","journal-title":"Comput. Intell. Neurosci."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Cao, H., Wang, Y., Chen, J., Jiang, D., Zhang, X., Tian, Q., and Wang, M. (2022, January 23\u201328). Swin-unet: Unet-like Pure Transformer for Medical Image Segmentation. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-031-25066-8_9"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Lou, A., Guan, S., and Loew, M. (2021, January 15\u201319). DC-UNet: Rethinking the U-Net architecture with dual channel efficient CNN for medical image segmentation. Proceedings of the Medical Imaging 2021: Image Processing SPIE, Online.","DOI":"10.1117\/12.2582338"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Tran, S.T., Cheng, C.H., Nguyen, T.T., Le, M.H., and Liu, D.G. (2021). Tmd-unet: Triple-unet with multi-scale input features and dense skip connection for medical image segmentation. Healthcare, 9.","DOI":"10.3390\/healthcare9010054"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1856","DOI":"10.1109\/TMI.2019.2959609","article-title":"Unet++: Redesigning skip connections to exploit multiscale features in image segmentation","volume":"39","author":"Zhou","year":"2019","journal-title":"IEEE Trans. Med. Imaging"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"107404","DOI":"10.1016\/j.patcog.2020.107404","article-title":"U2-Net: Going deeper with nested U-structure for salient object detection","volume":"106","author":"Qin","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_38","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Liu, N., Zhang, N., Wan, K., Shao, L., and Han, J. (2021, January 11\u201317). Visual saliency transformer. Proceedings of the IEEE\/CVF international Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00468"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 11\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Zhai, Y., and Shah, M. (2006, January 23\u201327). Visual attention detection in video sequences using spatiotemporal cues. Proceedings of the 14th ACM International Conference on Multimedia, Santa Barbara, CA, USA.","DOI":"10.1145\/1180639.1180824"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","article-title":"Image quality assessment: From error visibility to structural similarity","volume":"13","author":"Wang","year":"2004","journal-title":"IEEE Trans. Image Process."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the 13th European Conference\u2014Computer Vision\u2013ECCV 2014, Zurich, Switzerland. Part V 13.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"249","DOI":"10.1016\/j.dib.2017.09.038","article-title":"The TNO multiband image data collection","volume":"15","author":"Toet","year":"2017","journal-title":"Data Brief"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Jia, X., Zhu, C., Li, M., Tang, W., and Zhou, W. (2021, January 11\u201317). LLVIP: A visible-infrared paired dataset for low-light vision. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00389"},{"key":"ref_46","first-page":"57","article-title":"Infrared and visible images fusion method based on discrete wavelet transform","volume":"28","author":"Zhan","year":"2017","journal-title":"J. Comput."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Sruthy, S., Parameswaran, L., and Sasi, A.P. (2013, January 22\u201323). Image fusion technique using DT-CWT. Proceedings of the 2013 International Mutli-Conference on Automation, Computing, Communication, Control and Compressed Sensing (iMac4s), Kottayam, India.","DOI":"10.1109\/iMac4s.2013.6526400"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1016\/j.inffus.2006.02.001","article-title":"Remote sensing image fusion using the curvelet transform","volume":"8","author":"Nencini","year":"2007","journal-title":"Inf. Fusion"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1016\/j.inffus.2019.07.011","article-title":"IFCNN: A general image fusion framework based on convolutional neural network","volume":"54","author":"Zhang","year":"2020","journal-title":"Inf. Fusion"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"1200","DOI":"10.1109\/JAS.2022.105686","article-title":"SwinFusion: Cross-domain long-range learning for general image fusion via swin transformer","volume":"9","author":"Ma","year":"2022","journal-title":"IEEE\/CAA J. Autom. Sin."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Liu, X., Gao, H., Miao, Q., Xi, Y., Ai, Y., and Gao, D. (2022). MFST: Multi-Modal Feature Self-Adaptive Transformer for Infrared and Visible Image Fusion. Remote Sens., 14.","DOI":"10.3390\/rs14133233"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/16\/5\/803\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:04:38Z","timestamp":1760105078000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/16\/5\/803"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,25]]},"references-count":51,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2024,3]]}},"alternative-id":["rs16050803"],"URL":"https:\/\/doi.org\/10.3390\/rs16050803","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,2,25]]}}}