{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T09:19:35Z","timestamp":1784107175333,"version":"3.55.0"},"reference-count":54,"publisher":"MDPI AG","issue":"13","license":[{"start":{"date-parts":[[2024,6,21]],"date-time":"2024-06-21T00:00:00Z","timestamp":1718928000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62376214"],"award-info":[{"award-number":["62376214"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["92270117"],"award-info":[{"award-number":["92270117"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2023-JC-YB-533"],"award-info":[{"award-number":["2023-JC-YB-533"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Natural Science Basic Research Program of Shaanxi","award":["62376214"],"award-info":[{"award-number":["62376214"]}]},{"name":"Natural Science Basic Research Program of Shaanxi","award":["92270117"],"award-info":[{"award-number":["92270117"]}]},{"name":"Natural Science Basic Research Program of Shaanxi","award":["2023-JC-YB-533"],"award-info":[{"award-number":["2023-JC-YB-533"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The fusion of multi-modal medical images has great significance for comprehensive diagnosis and treatment. However, the large differences between the various modalities of medical images make multi-modal medical image fusion a great challenge. This paper proposes a novel multi-scale fusion network based on multi-dimensional dynamic convolution and residual hybrid transformer, which has better capability for feature extraction and context modeling and improves the fusion performance. Specifically, the proposed network exploits multi-dimensional dynamic convolution that introduces four attention mechanisms corresponding to four different dimensions of the convolutional kernel to extract more detailed information. Meanwhile, a residual hybrid transformer is designed, which activates more pixels to participate in the fusion process by channel attention, window attention, and overlapping cross attention, thereby strengthening the long-range dependence between different modes and enhancing the connection of global context information. A loss function, including perceptual loss and structural similarity loss, is designed, where the former enhances the visual reality and perceptual details of the fused image, and the latter enables the model to learn structural textures. The whole network adopts a multi-scale architecture and uses an unsupervised end-to-end method to realize multi-modal image fusion. Finally, our method is tested qualitatively and quantitatively on mainstream datasets. The fusion results indicate that our method achieves high scores in most quantitative indicators and satisfactory performance in visual qualitative analysis.<\/jats:p>","DOI":"10.3390\/s24134056","type":"journal-article","created":{"date-parts":[[2024,6,21]],"date-time":"2024-06-21T12:27:33Z","timestamp":1718972853000},"page":"4056","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["MDC-RHT: Multi-Modal Medical Image Fusion via Multi-Dimensional Dynamic Convolution and Residual Hybrid Transformer"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3318-8137","authenticated-orcid":false,"given":"Wenqing","family":"Wang","sequence":"first","affiliation":[{"name":"School of Automation and Information Engineering, Xi\u2019an University of Technology, Xi\u2019an 710048, China"},{"name":"Shaanxi Key Laboratory of Complex System Control and Intelligent Information Processing, Xi\u2019an University of Technology, Xi\u2019an 710048, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ji","family":"He","sequence":"additional","affiliation":[{"name":"School of Automation and Information Engineering, Xi\u2019an University of Technology, Xi\u2019an 710048, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6618-1380","authenticated-orcid":false,"given":"Han","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Automation and Information Engineering, Xi\u2019an University of Technology, Xi\u2019an 710048, China"},{"name":"Shaanxi Key Laboratory of Complex System Control and Intelligent Information Processing, Xi\u2019an University of Technology, Xi\u2019an 710048, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5463-2183","authenticated-orcid":false,"given":"Wei","family":"Yuan","sequence":"additional","affiliation":[{"name":"School of Automation and Information Engineering, Xi\u2019an University of Technology, Xi\u2019an 710048, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,6,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Xu, Y., Wang, H., Yin, X., and Tao, L. (2020, January 23\u201325). MRI and PET\/SPECT Image Fusion Based on Adaptive Weighted Guided Image Filtering. Proceedings of the 2020 IEEE 5th International Conference on Signal and Image Processing (ICSIP), Nanjing, China.","DOI":"10.1109\/ICSIP49896.2020.9339463"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"166413","DOI":"10.1016\/j.ijleo.2021.166413","article-title":"Optimum weighted multimodal medical image fusion using particle swarm optimization","volume":"231","author":"Shehanaz","year":"2021","journal-title":"Optik"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"5855","DOI":"10.1109\/TIP.2017.2745202","article-title":"Anatomical-functional image fusion by information of interest in local Laplacian filtering domain","volume":"26","author":"Du","year":"2017","journal-title":"IEEE Trans. Image Process."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"252","DOI":"10.1016\/j.compeleceng.2018.03.037","article-title":"Medical images fusion by using weighted least squares filter and sparse representation","volume":"67","author":"Jiang","year":"2018","journal-title":"Comput. Electr. Eng."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"280","DOI":"10.1016\/j.proeng.2010.11.045","article-title":"Multimodal medical image fusion based on IHS and PCA","volume":"7","author":"He","year":"2010","journal-title":"Procedia Eng."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"326","DOI":"10.1016\/j.neucom.2016.02.047","article-title":"Union Laplacian pyramid with multiple features for medical image fusion","volume":"194","author":"Du","year":"2016","journal-title":"Neurocomputing"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"280","DOI":"10.1002\/ima.20295","article-title":"Contrast enhancement dynamic histogram equalization for medical image processing application","volume":"21","author":"Ismail","year":"2011","journal-title":"Int. J. Imaging Syst. Technol."},{"key":"ref_8","first-page":"160","article-title":"Multimodal medical image fusion based on deep learning neural network for clinical treatment analysis","volume":"11","author":"Rajalingam","year":"2018","journal-title":"Int. J. ChemTech Res."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1882","DOI":"10.1109\/LSP.2016.2618776","article-title":"Image fusion with convolutional sparse representation","volume":"23","author":"Liu","year":"2016","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K.Q. (2017, January 21\u201326). Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_11","first-page":"21","article-title":"Medical image fusion method by deep learning","volume":"2","author":"Li","year":"2021","journal-title":"Int. J. Cogn. Comput. Eng."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"675","DOI":"10.2174\/1573405613666170428154156","article-title":"A review of denoising medical images using machine learning approaches","volume":"14","author":"Kaur","year":"2018","journal-title":"Curr. Med. Imaging"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Maqsood, S., and Javed, U. (2020). Multi-modal medical image fusion based on two-scale image decomposition and sparse representation. Biomed. Signal Process. Control, 57.","DOI":"10.1016\/j.bspc.2019.101810"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Wang, Z., Cui, Z., and Zhu, Y. (2020). Multi-modal medical image fusion by Laplacian pyramid and adaptive sparse representation. Comput. Biol. Med., 123.","DOI":"10.1016\/j.compbiomed.2020.103823"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"532","DOI":"10.1109\/TCOM.1983.1095851","article-title":"The Laplacian Pyramid as a Compact Image Code","volume":"31","author":"Burt","year":"1983","journal-title":"IEEE Trans. Commun."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"837","DOI":"10.1109\/34.531803","article-title":"Texture features for browsing and retrieval of image data","volume":"18","author":"Manjunath","year":"1996","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_17","first-page":"21","article-title":"Image fusion using Wavelet Transform: A Review","volume":"14","author":"Sahu","year":"2014","journal-title":"Glob. J. Comput. Sci. Technol."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"268","DOI":"10.1016\/j.neucom.2015.01.050","article-title":"High quality multi-spectral and panchromatic image fusion technologies based on curvelet transform","volume":"159","author":"Dong","year":"2015","journal-title":"Neurocomputing"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1598","DOI":"10.4236\/cs.2016.78139","article-title":"Multimodal medical image fusion in non-subsampled contourlet transform domain","volume":"7","author":"Gomathi","year":"2016","journal-title":"Circuits Syst."},{"key":"ref_20","first-page":"101245","article-title":"Enhanced JAYA optimization based medical image fusion in adaptive non subsampled shearlet transform domain","volume":"35","author":"Shilpa","year":"2022","journal-title":"Eng. Sci. Technol. Int. J."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"173","DOI":"10.4218\/etrij.17.0116.0568","article-title":"Multimodal medical image fusion based on Sugeno\u2019s intuitionistic fuzzy sets","volume":"39","author":"Tirupal","year":"2017","journal-title":"Etri J."},{"key":"ref_22","first-page":"1319","article-title":"Multimodal medical image fusion towards future research: A review","volume":"35","author":"Khan","year":"2023","journal-title":"J. King Saud Univ.-Comput. Inf. Sci."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1016\/j.inffus.2019.07.011","article-title":"IFCNN: A general image fusion framework based on convolutional neural network","volume":"54","author":"Zhang","year":"2020","journal-title":"Inf. Fusion"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"80","DOI":"10.1016\/j.inffus.2022.11.010","article-title":"MUFusion: A general unsupervised image fusion network based on memory unit","volume":"92","author":"Cheng","year":"2023","journal-title":"Inf. Fusion"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Liu, R., Liu, Y., Wang, H., Hu, K., and Du, S. (2024, January 14\u201319). A Novel Medical Image Fusion Framework Integrating Multi-scale Encoder-Decoder with Discrete Wavelet Decomposition. Proceedings of the ICASSP 2024\u20142024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Seoul, Republic of Korea.","DOI":"10.1109\/ICASSP48485.2024.10446618"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"177","DOI":"10.1016\/j.inffus.2021.06.001","article-title":"EMFusion: An unsupervised enhanced medical image fusion network","volume":"76","author":"Xu","year":"2021","journal-title":"Inf. Fusion"},{"key":"ref_27","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., Fu, Y., Feng, J., Xiang, T., and Torr, P.H. (2021, January 20\u201325). Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00681"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 11\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Virtual, Online, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"5016412","DOI":"10.1109\/TIM.2022.3216413","article-title":"SwinFuse: A residual swin transformer fusion network for infrared and visible images","volume":"71","author":"Wang","year":"2022","journal-title":"IEEE Trans. Instrum. Meas."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Li, W., Zhang, Y., Wang, G., Huang, Y., and Li, R. (2023). DFENet: A dual-branch feature enhanced network integrating transformers and convolutional feature learning for multimodal medical image fusion. Biomed. Signal Process. Control, 80.","DOI":"10.1016\/j.bspc.2022.104402"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"5134","DOI":"10.1109\/TIP.2022.3193288","article-title":"MATR: Multimodal medical image fusion via multiscale adaptive transformer","volume":"31","author":"Tang","year":"2022","journal-title":"IEEE Trans. Image Process."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Chen, X., Wang, X., Zhou, J., Qiao, Y., and Dong, C. (2023, January 18\u201322). Activating more pixels in image super-resolution transformer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02142"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Dong, X., Bao, J., Chen, D., Zhang, W., Yu, N., Yuan, L., Chen, D., and Guo, B. (2022, January 19\u201324). Cswin transformer: A general vision transformer backbone with cross-shaped windows. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01181"},{"key":"ref_35","unstructured":"Wu, S., Wu, T., Tan, H., and Guo, G. (March, January 22). Pale transformer: A general vision transformer backbone with pale-shaped attention. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual."},{"key":"ref_36","unstructured":"Li, C., Zhou, A., and Yao, A. (2022, January 25\u201329). Omni-Dimensional Dynamic Convolution. Proceedings of the International Conference on Learning Representations, Virtual."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","article-title":"Image quality assessment: From error visibility to structural similarity","volume":"13","author":"Wang","year":"2004","journal-title":"IEEE Trans. Image Process."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"4980","DOI":"10.1109\/TIP.2020.2977573","article-title":"DDcGAN: A dual-discriminator conditional generative adversarial network for multi-resolution image fusion","volume":"29","author":"Ma","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"502","DOI":"10.1109\/TPAMI.2020.3012548","article-title":"U2Fusion: A unified unsupervised image fusion network","volume":"44","author":"Xu","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Li, W., Peng, X., Fu, J., Wang, G., Huang, Y., and Chao, F. (2022). A multiscale double-branch residual attention network for anatomical\u2013functional medical image fusion. Comput. Biol. Med., 141.","DOI":"10.1016\/j.compbiomed.2021.105005"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"2959","DOI":"10.1109\/26.477498","article-title":"Image quality measures and their performance","volume":"43","author":"Eskicioglu","year":"1995","journal-title":"IEEE Trans. Commun."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"287","DOI":"10.1016\/j.physd.2004.11.001","article-title":"A nonlinear correlation measure for multivariable data set","volume":"200","author":"Wang","year":"2005","journal-title":"Phys. D Nonlinear Phenom."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"1890","DOI":"10.1016\/j.aeue.2015.09.004","article-title":"A new image quality metric for image fusion: The sum of the correlations of differences","volume":"69","author":"Aslantas","year":"2015","journal-title":"Aeu-Int. J. Electron. Commun."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"3345","DOI":"10.1109\/TIP.2015.2442920","article-title":"Perceptual quality assessment for multi-exposure image fusion","volume":"24","author":"Ma","year":"2015","journal-title":"IEEE Trans. Image Process."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"313","DOI":"10.1049\/el:20020212","article-title":"Information measure for performance of image fusion","volume":"38","author":"Qu","year":"2002","journal-title":"Electron. Lett."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1016\/j.inffus.2011.08.002","article-title":"A new image fusion performance metric based on visual information fidelity","volume":"14","author":"Han","year":"2013","journal-title":"Inf. Fusion"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"626","DOI":"10.1049\/el:20060693","article-title":"Image fusion metric based on mutual information and Tsallis entropy","volume":"42","author":"Cvejic","year":"2006","journal-title":"Electron. Lett."},{"key":"ref_48","unstructured":"Piella, G., and Heijmans, H. (2003, January 14\u201317). A new quality metric for image fusion. Proceedings of the Proceedings 2003 International Conference on Image Processing (Cat. No.03CH37429), Barcelona, Spain."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"308","DOI":"10.1049\/el:20000267","article-title":"Objective image fusion performance measure","volume":"36","author":"Xydeas","year":"2000","journal-title":"Electron. Lett."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"162","DOI":"10.1111\/j.1365-313X.2004.02281.x","article-title":"High-throughput protein localization in Arabidopsis using Agrobacterium-mediated transient expression of GFP-ORF fusions","volume":"41","author":"Koroleva","year":"2005","journal-title":"Plant J."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"110159","DOI":"10.1016\/j.jneumeth.2024.110159","article-title":"A novel approach of brain-computer interfacing (BCI) and Grad-CAM based explainable artificial intelligence: Use case scenario for smart healthcare","volume":"408","author":"Lamba","year":"2024","journal-title":"J. Neurosci. Methods"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Wani, N.A., Kumar, R., and Bedi, J. (2024). DeepXplainer: An interpretable deep learning based approach for lung cancer detection using explainable artificial intelligence. Comput. Methods Programs Biomed., 243.","DOI":"10.1016\/j.cmpb.2023.107879"},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"102550","DOI":"10.1016\/j.eclinm.2024.102550","article-title":"Predicting skin cancer risk from facial images with an explainable artificial intelligence (XAI) based approach: A proof-of-concept study","volume":"71","author":"Liu","year":"2024","journal-title":"eClinicalMedicine"},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"102135","DOI":"10.1016\/j.tele.2024.102135","article-title":"EXplainable Artificial Intelligence (XAI) for facilitating recognition of algorithmic bias: An experiment from imposed users\u2019 perspectives","volume":"91","author":"Chuan","year":"2024","journal-title":"Telemat. Inform."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/13\/4056\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:02:37Z","timestamp":1760108557000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/13\/4056"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,21]]},"references-count":54,"journal-issue":{"issue":"13","published-online":{"date-parts":[[2024,7]]}},"alternative-id":["s24134056"],"URL":"https:\/\/doi.org\/10.3390\/s24134056","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,21]]}}}