{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T16:52:41Z","timestamp":1783788761709,"version":"3.55.0"},"reference-count":53,"publisher":"MDPI AG","issue":"17","license":[{"start":{"date-parts":[[2019,8,28]],"date-time":"2019-08-28T00:00:00Z","timestamp":1566950400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Reliable vision in challenging illumination conditions is one of the crucial requirements of future autonomous automotive systems. In the last decade, thermal cameras have become more easily accessible to a larger number of researchers. This has resulted in numerous studies which confirmed the benefits of the thermal cameras in limited visibility conditions. In this paper, we propose a learning-based method for visible and thermal image fusion that focuses on generating fused images with high visual similarity to regular truecolor (red-green-blue or RGB) images, while introducing new informative details in pedestrian regions. The goal is to create natural, intuitive images that would be more informative than a regular RGB camera to a human driver in challenging visibility conditions. The main novelty of this paper is the idea to rely on two types of objective functions for optimization: a similarity metric between the RGB input and the fused output to achieve natural image appearance; and an auxiliary pedestrian detection error to help defining relevant features of the human appearance and blending them into the output. We train a convolutional neural network using image samples from variable conditions (day and night) so that the network learns the appearance of humans in the different modalities and creates more robust results applicable in realistic situations. Our experiments show that the visibility of pedestrians is noticeably improved especially in dark regions and at night. Compared to existing methods we can better learn context and define fusion rules that focus on the pedestrian appearance, while that is not guaranteed with methods that focus on low-level image quality metrics.<\/jats:p>","DOI":"10.3390\/s19173727","type":"journal-article","created":{"date-parts":[[2019,8,28]],"date-time":"2019-08-28T11:23:18Z","timestamp":1566991398000},"page":"3727","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":72,"title":["Deep Visible and Thermal Image Fusion for Enhanced Pedestrian Visibility"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8487-9598","authenticated-orcid":false,"given":"Ivana","family":"Shopovska","sequence":"first","affiliation":[{"name":"TELIN-IPI, Ghent University - imec, St-Pietersnieuwstraat 41, B-9000 Gent, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8790-1116","authenticated-orcid":false,"given":"Ljubomir","family":"Jovanov","sequence":"additional","affiliation":[{"name":"TELIN-IPI, Ghent University - imec, St-Pietersnieuwstraat 41, B-9000 Gent, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wilfried","family":"Philips","sequence":"additional","affiliation":[{"name":"TELIN-IPI, Ghent University - imec, St-Pietersnieuwstraat 41, B-9000 Gent, Belgium"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2019,8,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Ziebinski, A., Cupek, R., Erdogan, H., and Waechter, S. (2016). A Survey of ADAS Technologies for the Future Perspective of Sensor Fusion. Lecture Notes in Computer Science, Proceedings of the ICCCI 2016, Halkidiki, Greece, 28\u201330 September 2016, Springer.","DOI":"10.1007\/978-3-319-45246-3_13"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"153","DOI":"10.1016\/j.inffus.2018.02.004","article-title":"Infrared and visible image fusion methods and applications: A survey","volume":"45","author":"Ma","year":"2019","journal-title":"Inf. Fusion"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"20676","DOI":"10.1109\/ACCESS.2019.2897320","article-title":"Poisson Reconstruction-Based Fusion of Infrared and Visible Images via Saliency Detection","volume":"7","author":"Li","year":"2019","journal-title":"IEEE Access"},{"key":"ref_4","unstructured":"Commission, E. (2016). Advanced Driver Assistance Systems, European Commission, Directorate General for Transport. Technical Report."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1850018","DOI":"10.1142\/S0219691318500182","article-title":"Infrared and visible image fusion with convolutional neural networks","volume":"16","author":"Liu","year":"2018","journal-title":"Int. J. Wavelets Multiresolution Inf. Process."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1016\/j.inffus.2015.11.003","article-title":"Perceptual fusion of infrared and visible images through a hybrid multi-scale decomposition with Gaussian and bilateral filters","volume":"30","author":"Zhou","year":"2016","journal-title":"Inf. Fusion"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2864","DOI":"10.1109\/TIP.2013.2244222","article-title":"Image fusion with guided filtering","volume":"22","author":"Li","year":"2013","journal-title":"IEEE Trans. Image Process."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Li, H., and Wu, X.J. (2018). Infrared and Visible Image Fusion with ResNet and zero-phase component analysis. arXiv.","DOI":"10.1016\/j.infrared.2019.103039"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Li, H. (2018, January 20\u201324). CODE: Infrared and Visible Image Fusion using a Deep Learning Framework. Proceedings of the International Conference on Pattern Recognition 2018, Beijing, China.","DOI":"10.1109\/ICPR.2018.8546006"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.ijleo.2018.06.123","article-title":"Structure-aware image fusion","volume":"172","author":"Li","year":"2018","journal-title":"Optik"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1193","DOI":"10.1007\/s11760-013-0556-9","article-title":"Image fusion based on pixel significance using cross bilateral filter","volume":"9","author":"Kumar","year":"2015","journal-title":"Signal Image Video Process."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Bavirisetti, D.P., Xiao, G., and Liu, G. (2017, January 10\u201313). Multi-Sensor Image Fusion Based on Fourth Order Partial Differential Equations. Proceedings of the 2017 20th International Conference on Information Fusion (Fusion), Xi\u2019an, China.","DOI":"10.23919\/ICIF.2017.8009719"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"52","DOI":"10.1016\/j.infrared.2016.01.009","article-title":"Two-scale image fusion of visible and infrared images using saliency detection","volume":"76","author":"Bavirisetti","year":"2016","journal-title":"Infrared Phys. Technol."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"479","DOI":"10.14429\/dsj.61.705","article-title":"Image fusion technique using multi-resolution singular value decomposition","volume":"61","author":"Naidu","year":"2011","journal-title":"Def. Sci. J."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Hafner, D., and Weickert, J. (2016). Variational Image Fusion with Optimal Local Contrast, Wiley Online Library. Computer Graphics Forum.","DOI":"10.1111\/cgf.12690"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"100","DOI":"10.1016\/j.inffus.2016.02.001","article-title":"Infrared and visible image fusion via gradient transfer and total variation minimization","volume":"31","author":"Ma","year":"2016","journal-title":"Inf. Fusion"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1016\/j.neucom.2016.03.009","article-title":"Infrared and visible image fusion using total variation model","volume":"202","author":"Ma","year":"2016","journal-title":"Neurocomputing"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"182","DOI":"10.1016\/j.neucom.2016.11.051","article-title":"A novel infrared and visible image fusion algorithm based on shift-invariant dual-tree complex shearlet transform and sparse representation","volume":"226","author":"Yin","year":"2017","journal-title":"Neurocomputing"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"201","DOI":"10.1016\/j.infrared.2017.01.012","article-title":"Fusion of visible and infrared images using global entropy and gradient constrained regularization","volume":"81","author":"Zhao","year":"2017","journal-title":"Infrared Phys. Technol."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1016\/j.infrared.2017.02.005","article-title":"Infrared and visible image fusion based on visual saliency map and weighted least square optimization","volume":"82","author":"Ma","year":"2017","journal-title":"Infrared Phys. Technol."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"94","DOI":"10.1016\/j.infrared.2017.04.018","article-title":"Infrared and visible image fusion method based on saliency detection in sparse domain","volume":"83","author":"Liu","year":"2017","journal-title":"Infrared Phys. Technol."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Alldieck, T., Bahnsen, C., and Moeslund, T. (2016). Context-aware fusion of RGB and thermal imagery for traffic monitoring. Sensors, 16.","DOI":"10.3390\/s16111947"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1203","DOI":"10.1016\/j.ijleo.2015.02.092","article-title":"Multi-level image fusion and enhancement for target detection","volume":"126","author":"He","year":"2015","journal-title":"Optik"},{"key":"ref_24","unstructured":"Choi, E.J., and Park, D.J. (December, January 30). Human Detection Using Image Fusion of Thermal and Visible Image with New Joint Bilateral Filter. Proceedings of the IEEE 5th International Conference on Computer Sciences and Convergence Information Technology, Seoul, Korea."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Thomanek, J., Ritter, M., Lietz, H., and Wanielik, G. (2011, January 6\u20138). Comparing Visual Data Fusion tEchniques Using Fir and Visible Light Sensors to Improve Pedestrian Detection. Proceedings of the IEEE 2011 International Conference on Digital Image Computing: Techniques and Applications, Noosa, Australia.","DOI":"10.1109\/DICTA.2011.27"},{"key":"ref_26","unstructured":"Thomanek, J., and Wanielik, G. (2014, January 7\u201310). A New Pixel-Based Fusion Framework to Enhance Object Detection in Automotive Applications. Proceedings of the IEEE 17th International Conference on Information Fusion (FUSION), Salamanca, Spain."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"194","DOI":"10.1109\/TPAMI.2011.146","article-title":"Image signature: Highlighting sparse salient regions","volume":"34","author":"Hou","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Hwang, S., Park, J., Kim, N., Choi, Y., and So Kweon, I. (2015, January 7\u201312). Multispectral Pedestrian Detection: Benchmark Dataset And Baseline. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298706"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Gonz\u00e1lez, A., Fang, Z., Socarras, Y., Serrat, J., V\u00e1zquez, D., Xu, J., and L\u00f3pez, A.M. (2016). Pedestrian detection at day\/night time with visible and FIR cameras: A comparison. Sensors, 16.","DOI":"10.3390\/s16060820"},{"key":"ref_30","unstructured":"Wagner, J., Fischer, V., Herman, M., and Behnke, S. (2016, January 19). Multispectral Pedestrian Detection Using Deep Fusion Convolutional Neural Networks. Proceedings of the 24th European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN), Bruges, Belgium."},{"key":"ref_31","unstructured":"Choi, H., Kim, S., Park, K., and Sohn, K. (2016, January 4\u20138). Multi-Spectral Pedestrian Detection Based on Accumulated Object Proposal with Fully Convolutional Networks. Proceedings of the IEEE 23rd International Conference on Pattern Recognition (ICPR), Cancun, Mexico."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Liu, J., Zhang, S., Wang, S., and Metaxas, D.N. (2016). Multispectral Deep Neural Networks for Pedestrian Detection. arXiv.","DOI":"10.5244\/C.30.73"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Konig, D., Adam, M., Jarvers, C., Layher, G., Neumann, H., and Teutsch, M. (2017, January 21\u201326). Fully Convolutional Region Proposal Networks for Multispectral Person Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.36"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Xu, D., Ouyang, W., Ricci, E., Wang, X., and Sebe, N. (2017). Learning Cross-Modal Deep Representations for Robust Pedestrian Detection. arXiv.","DOI":"10.1109\/CVPR.2017.451"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Washington, DC, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_36","unstructured":"Benenson, R., Omran, M., Hosang, J., and Schiele, B. (2014). Ten Years of Pedestrian Detection, What Have We Learned?, Springer. European Conference on Computer Vision."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1016\/j.neucom.2018.01.092","article-title":"Computer vision and deep learning techniques for pedestrian detection and tracking: A survey","volume":"300","author":"Brunetti","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"1791","DOI":"10.1109\/TCYB.2018.2813971","article-title":"Deep attention-based spatially recursive networks for fine-grained visual recognition","volume":"49","author":"Wu","year":"2018","journal-title":"IEEE Trans. Cybern."},{"key":"ref_39","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015, January 7\u201312). Faster R-CNN Towards Real-Time Object Detection with Region Proposal Networks. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_40","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014, January 8\u201313). Generative Adversarial Nets. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1016\/j.inffus.2018.09.004","article-title":"FusionGAN: A generative adversarial network for infrared and visible image fusion","volume":"48","author":"Ma","year":"2019","journal-title":"Inf. Fusion"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Lin, K.Y., and Wang, G. (2018, January 18\u201322). Hallucinated-IQA: No-Reference Image Quality Assessment via Adversarial Learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00083"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015). Fast R-CNN. arXiv.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_45","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_46","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014). Microsoft coco: Common objects in context. Lecture Notes in Computer Science, Proceedings of the European Conference on Computer Vision, Switzerland, 6\u201312 September 2014, Springer.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"478","DOI":"10.1016\/j.infrared.2017.07.010","article-title":"A survey of infrared and visual image fusion methods","volume":"85","author":"Jin","year":"2017","journal-title":"Infrared Phys. Technol."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"313","DOI":"10.1049\/el:20020212","article-title":"Information measure for performance of image fusion","volume":"38","author":"Qu","year":"2002","journal-title":"Electron. Lett."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"308","DOI":"10.1049\/el:20000267","article-title":"Objective image fusion performance measure","volume":"36","author":"Xydeas","year":"2000","journal-title":"Electron. Lett."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Petrovic, V.V., Cootes, T., and Pavlovic, R. (2007, January 9\u201312). Dynamic Image Fusion Performance Evaluation. Proceedings of the IEEE 10th International Conference on Information Fusion, Quebec, QC, Canada.","DOI":"10.1109\/ICIF.2007.4408120"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, Faster, Stronger. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Vedaldi, A., and Lenc, K. (2015, January 23\u201326). MatConvNet\u2014Convolutional Neural Networks for MATLAB. In Proceeding of the ACM International Conference on Multimedia Retrieval, Shanghai, China.","DOI":"10.1145\/2733373.2807412"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/17\/3727\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:14:46Z","timestamp":1760188486000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/17\/3727"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,8,28]]},"references-count":53,"journal-issue":{"issue":"17","published-online":{"date-parts":[[2019,9]]}},"alternative-id":["s19173727"],"URL":"https:\/\/doi.org\/10.3390\/s19173727","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,8,28]]}}}