{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T04:46:22Z","timestamp":1787028382133,"version":"build-2736575974"},"reference-count":50,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2025,12,2]],"date-time":"2025-12-02T00:00:00Z","timestamp":1764633600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2024YFB4303201"],"award-info":[{"award-number":["2024YFB4303201"]}]},{"name":"Ganzhou City Key Research and Development Program","award":["GZ2024ZDZ007"],"award-info":[{"award-number":["GZ2024ZDZ007"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IJGI"],"abstract":"<jats:p>Pedestrian detection under low illumination and complex environments remains a significant challenge for vision-based systems, particularly in safety-critical applications such as urban rail transit. To address the limitations of single-modality detection in adverse conditions, this paper proposes IVIFusion, a lightweight yet robust pedestrian detection framework that fuses infrared and visible images at the feature level. The method integrates a dual-branch Transformer-based backbone for modality-specific feature extraction and introduces a Cross-Modality Attention Fusion Module (CMAFM) to adaptively enhance cross-modal representations while suppressing noise. Furthermore, a dedicated small-object detection layer is incorporated to improve the recall of distant and occluded pedestrians. Extensive experiments conducted on the public LLVIP dataset and the custom HGPD dataset demonstrate the superior performance of IVIFusion, achieving mAP0.5 scores of 98.6% and 97.2%, respectively. The results validate the effectiveness of the proposed architecture in handling challenging lighting conditions while maintaining real-time efficiency and low computational cost.<\/jats:p>","DOI":"10.3390\/ijgi14120477","type":"journal-article","created":{"date-parts":[[2025,12,4]],"date-time":"2025-12-04T16:07:38Z","timestamp":1764864458000},"page":"477","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Infrared\u2013Visible Fusion via Cross-Modality Attention and Small-Object Enhancement for Pedestrian Detection"],"prefix":"10.3390","volume":"14","author":[{"given":"Jie","family":"Yang","sequence":"first","affiliation":[{"name":"School of Electrical Engineering, Shanghai Dianji University, Shanghai 201306, China"},{"name":"Department of Electrical Engineering and Automation, Jiangxi University of Science and Technology, Ganzhou 341000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanxuan","family":"Jiang","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering and Automation, Jiangxi University of Science and Technology, Ganzhou 341000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dengyin","family":"Jiang","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering, Shanghai Dianji University, Shanghai 201306, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7150-4914","authenticated-orcid":false,"given":"Zhichao","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering, Shanghai Dianji University, Shanghai 201306, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,12,2]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Chen, Z., Yang, J., Li, F., Feng, Z., Chen, L., Jia, L., and Li, P. (2025). Foreign Object Detection Method for Railway Catenary Based on a Scarce Image Generation Model and Lightweight Perception Architecture. IEEE Transactions on Circuits and Systems for Video Technology, IEEE.","DOI":"10.1109\/TCSVT.2025.3567319"},{"key":"ref_2","unstructured":"Zeng, S., Chang, X., Xie, M., Liu, X., Bai, Y., Pan, Z., Xu, M., and Wei, X. (2025). FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"37175","DOI":"10.1109\/JIOT.2025.3582636","article-title":"RailVoxelDet: A Lightweight 3-D Object Detection Method for Railway Transportation Driven by Onboard LiDAR Data","volume":"12","author":"Chen","year":"2025","journal-title":"IEEE Internet Things J."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"14100","DOI":"10.1109\/TTE.2025.3609347","article-title":"Generalized Koopman Neural Operator for Data-Driven Modeling of Electric Railway Pantograph\u2013Catenary Systems","volume":"11","author":"Wang","year":"2025","journal-title":"IEEE Trans. Transp. Electrif."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Yang, H., Liu, Z., and Cui, H. (2025). An Electrified Railway Catenary Component Anomaly Detection Frame Based on Invariant Normal Region Prototype with Segment Anything Model. IEEE Transactions on Transportation Electrification, IEEE.","DOI":"10.1109\/TTE.2025.3628607"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1987","DOI":"10.1109\/TCSVT.2024.3486347","article-title":"Unifying Motion and Appearance Cues for Visual Tracking via Shared Queries","volume":"35","author":"Xue","year":"2025","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Xue, C., Zhong, B., Liang, Q., Zheng, Y., Li, N., Xue, Y., and Song, S. (2025, January 10\u201313). Similarity-guided layer-adaptive vision transformer for UAV tracking. Proceedings of the Computer Vision and Pattern Recognition Conference, Gold Coast, Australia.","DOI":"10.1109\/CVPR52734.2025.00631"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"9716","DOI":"10.1109\/TIFS.2025.3608672","article-title":"A Semantically Guided and Focused Network for Occluded Person Re-identification","volume":"20","author":"Lin","year":"2025","journal-title":"IEEE Trans. Inf. Forensics Secur."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1038\/s41597-024-02918-9","article-title":"RailFOD23: A dataset for foreign object detection on railroad transmission lines","volume":"11","author":"Chen","year":"2024","journal-title":"Sci. Data"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s40864-024-00231-7","article-title":"Urban Rail Transit in China: Progress Report and Analysis (2015\u20132023)","volume":"11","author":"Lu","year":"2025","journal-title":"Urban Rail Transit"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"You, K., Gu, Y., Lin, Y., and Wang, Y. (2025). A novel physical constraint-guided quadratic neural networks for interpretable bearing fault diagnosis under zero-fault sample. Nondestruct. Test. Eval., 1\u201331.","DOI":"10.1080\/10589759.2025.2534429"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"3234","DOI":"10.1109\/TITS.2020.2993926","article-title":"Deep neural network based vehicle and pedestrian detection for autonomous driving: A survey","volume":"22","author":"Chen","year":"2021","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"477","DOI":"10.1016\/j.inffus.2022.10.034","article-title":"DIVFusion: Darkness-free infrared and visible image fusion","volume":"91","author":"Tang","year":"2023","journal-title":"Inf. Fusion"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"16116","DOI":"10.1109\/TITS.2025.3581391","article-title":"Target-Distractor Aware UAV Tracking via Global Agent","volume":"26","author":"Xue","year":"2025","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Xue, Y., Jin, G., Zhong, B., Shen, T., Tan, L., Xue, C., and Zheng, Y. (2025). FMTrack: Frequency-aware Interaction and Multi-Expert Fusion for RGB-T Tracking. IEEE Trans. Circuits Syst. Video Technol., 1\u201313.","DOI":"10.1109\/TCSVT.2025.3601598"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Wei, R., Zhu, D., Zhan, W., and Hao, Z. (2019, January 12\u201314). Infrared and visible image fusion based on RPCA and NSST. Proceedings of the 2019 IEEE International Conference on Power, Intelligent Computing and Systems (ICPICS), Shenyang, China.","DOI":"10.1109\/ICPICS47731.2019.8942573"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1016\/j.inffus.2021.04.005","article-title":"Attribute filter based infrared and visible image fusion","volume":"75","author":"Mo","year":"2021","journal-title":"Inf. Fusion"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"105069","DOI":"10.1016\/j.autcon.2023.105069","article-title":"Efficient railway track region segmentation algorithm based on lightweight neural network and cross-fusion decoder","volume":"155","author":"Chen","year":"2023","journal-title":"Autom. Constr."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"105432","DOI":"10.1016\/j.tust.2023.105432","article-title":"Towards automated 3D evaluation of water leakage on a tunnel face via improved GAN and self-attention DL model","volume":"142","author":"Wu","year":"2023","journal-title":"Tunn. Undergr. Space Technol."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Chen, L., Chen, Z., Yan, L., Cheng, Y., Guan, F., and Li, P. (2025, January 16\u201322). Optimal distributed training with co-adaptive data parallelism in heterogeneous environments. Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, Montreal, QC, Canada.","DOI":"10.24963\/ijcai.2025\/4"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"322","DOI":"10.23919\/cje.2023.00.065","article-title":"Ipfa-net: Important points feature aggregating net for point cloud classification and segmentation","volume":"34","author":"Wang","year":"2025","journal-title":"Chin. J. Electron."},{"key":"ref_22","first-page":"509","article-title":"Multispectral Pedestrian Detection using Deep Fusion Convolutional Neural Networks","volume":"587","author":"Wagner","year":"2016","journal-title":"ESANN"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"161","DOI":"10.1016\/j.patcog.2018.08.005","article-title":"Illumination-aware faster R-CNN for robust multispectral pedestrian detection","volume":"85","author":"Li","year":"2019","journal-title":"Pattern Recognit."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"138117","DOI":"10.1109\/ACCESS.2020.3012558","article-title":"An improved faster R-CNN pedestrian detection algorithm based on feature fusion and context analysis","volume":"8","author":"Zhai","year":"2020","journal-title":"IEEE Access"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhao, Z., Bai, H., Zhang, J., Zhang, Y., Xu, S., Lin, Z., Timofte, R., and Van Gool, L. (2023, January 17\u201324). Cddfuse: Correlation-driven dual-branch feature decomposition for multi-modality image fusion. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00572"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Sun, J., Yin, M., Wang, Z., Xie, T., and Bei, S. (2024). Multispectral object detection based on multilevel feature fusion and dual feature modulation. Electronics, 13.","DOI":"10.3390\/electronics13020443"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhao, F., Lou, W., Feng, H., Ding, N., and Li, C. (2024). MFMG-Net: Multispectral Feature Mutual Guidance Network for Visible\u2013Infrared Object Detection. Drones, 8.","DOI":"10.3390\/drones8030112"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Liu, S., He, H., Zhang, Z., and Zhou, Y. (2024). LI-YOLO: An Object Detection Algorithm for UAV Aerial Images in Low-Illumination Scenes. Drones, 8.","DOI":"10.3390\/drones8110653"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"103324","DOI":"10.1016\/j.ecoinf.2025.103324","article-title":"Mamba-based super-resolution and semi-supervised YOLOv10 for freshwater mussel detection using acoustic video camera: A case study at Lake Izunuma, Japan","volume":"90","author":"Zhao","year":"2025","journal-title":"Ecol. Inform."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"13232","DOI":"10.1109\/TNNLS.2023.3266452","article-title":"LRAF-Net: Long-Range Attention Fusion Network for Visible\u2013Infrared Object Detection","volume":"35","author":"Fu","year":"2024","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_31","first-page":"1","article-title":"Global\u2013Local Feature Fusion Network for Visible\u2013Infrared Vehicle Detection","volume":"21","author":"Kang","year":"2024","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3505244","article-title":"Transformers in vision: A survey","volume":"54","author":"Khan","year":"2022","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"3159","DOI":"10.1109\/TCSVT.2023.3234340","article-title":"DATFuse: Infrared and visible image fusion via dual attention transformer","volume":"33","author":"Tang","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TIM.2022.3218574","article-title":"CGTF: Convolution-guided transformer for infrared and visible image fusion","volume":"71","author":"Li","year":"2022","journal-title":"IEEE Trans. Instrum. Meas."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"1614","DOI":"10.1109\/LSP.2022.3180672","article-title":"Learn to search a lightweight architecture for target-aware infrared and visible image fusion","volume":"29","author":"Liu","year":"2022","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_36","unstructured":"Jocher, G., Chaurasia, A., Qiu, T., and Stoken, A. (2024, June 22). YOLOv5: You Only Look Once Version 5. Available online: https:\/\/github.com\/ultralytics\/yolov5."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"111655","DOI":"10.1016\/j.measurement.2022.111655","article-title":"Fast vehicle detection algorithm in traffic scene based on improved SSD","volume":"201","author":"Chen","year":"2022","journal-title":"Measurement"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Jia, X., Zhu, C., Li, M., Tang, W., and Zhou, W. (2021, January 11\u201317). LLVIP: A visible-infrared paired dataset for low-light vision. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Virtual.","DOI":"10.1109\/ICCVW54120.2021.00389"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1186\/s13634-023-01002-5","article-title":"Decision-level fusion detection method of visible and infrared images under low light conditions","volume":"2023","author":"Hu","year":"2023","journal-title":"EURASIP J. Adv. Signal Process."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"14164","DOI":"10.1109\/TNNLS.2023.3274926","article-title":"Image Enhancement Guided Object Detection in Visually Degraded Scenes","volume":"35","author":"Liu","year":"2024","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_43","unstructured":"Zhang, H., Li, F., Liu, S., Zhang, L., Su, H., Zhu, J., Ni, L., and Shum, H.Y. (2023, January 1\u20135). DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection. Proceedings of the The Eleventh International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"9243","DOI":"10.1007\/s11042-022-13644-y","article-title":"Object detection using YOLO: Challenges, architectural successors, datasets and applications","volume":"82","author":"Diwan","year":"2023","journal-title":"Multimed. Tools Appl."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Medeiros, H.R., Pena, F.A.G., Aminbeidokhti, M., Dubail, T., Granger, E., and Pedersoli, M. (2024, January 4\u20138). HalluciDet: Hallucinating RGB Modality for Person Detection Through Privileged Information. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV57701.2024.00147"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Chen, Y.T., Shi, J., Ye, Z., Mertz, C., Ramanan, D., and Kong, S. (2022, January 23\u201327). Multimodal object detection via probabilistic ensembling. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-20077-9_9"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Zhao, W., Xie, S., Zhao, F., He, Y., and Lu, H. (2023, January 17\u201324). Metafusion: Infrared and visible image fusion via meta-feature embedding from object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01341"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Cao, Y., Bin, J., Hamari, J., Blasch, E., and Liu, Z. (2023, January 17\u201324). Multimodal object detection by channel switching and spatial attention. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPRW59228.2023.00046"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Zhao, Z., Bai, H., Zhu, Y., Zhang, J., Xu, S., Zhang, Y., Zhang, K., Meng, D., Timofte, R., and Van Gool, L. (2023, January 2\u20136). DDFM: Denoising diffusion model for multi-modality image fusion. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.00742"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"4776","DOI":"10.1109\/TMM.2023.3326296","article-title":"Camf: An interpretable infrared and visible image fusion network based on class activation mapping","volume":"26","author":"Tang","year":"2023","journal-title":"IEEE Trans. Multimed."}],"container-title":["ISPRS International Journal of Geo-Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2220-9964\/14\/12\/477\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,5]],"date-time":"2025-12-05T05:14:51Z","timestamp":1764911691000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2220-9964\/14\/12\/477"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,2]]},"references-count":50,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["ijgi14120477"],"URL":"https:\/\/doi.org\/10.3390\/ijgi14120477","relation":{},"ISSN":["2220-9964"],"issn-type":[{"value":"2220-9964","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,2]]}}}