{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,14]],"date-time":"2026-05-14T05:10:21Z","timestamp":1778735421258,"version":"3.51.4"},"reference-count":56,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T00:00:00Z","timestamp":1777334400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>In autonomous driving and intelligent transportation systems, the degradation of image quality under low-light conditions severely impacts the reliability of subsequent object detection. Existing methods predominantly employ data-driven deep learning models for image enhancement, often lacking physical interpretability and struggling to maintain robustness in complex lighting-varying traffic scenarios. To address this, this paper proposes a Physically Guided Transformer\u2013CNN Hybrid Network (Physically Guided Transformer\u2013CNN Hybrid Network, PGT-Net) for end-to-end joint optimization of low-light enhancement and object detection. PGT-Net innovatively integrates the atmospheric scattering physical model with deep learning architecture: first, a learnable physical guidance branch estimates the scene\u2019s atmospheric illumination map and transmittance map, providing explicit physical priors for the network; second, a dual-branch enhancement backbone is designed, where the local CNN branch (based on an improved UNet) restores fine textures, while the Global Transformer Branch (based on Swin Transformer) models long-range dependencies to correct global uneven illumination, with features adaptively combined via a Physical Fusion Module to ensure enhancement results align with physical laws while retaining rich visual features; finally, the enhanced images are directly fed into a lightweight detection head (e.g., YOLOv7) for joint training and optimization. Comprehensive experiments on public datasets (ExDark, BDD100K-night, etc.) demonstrate that PGT-Net significantly outperforms mainstream methods (e.g., RetinexNet, KinD, Zero-DCE) in both low-light image enhancement quality (PSNR\/SSIM) and object detection accuracy (mAP), while maintaining high inference efficiency. This research offers an interpretable, high-performance solution for visual perception tasks under adverse lighting conditions, holding strong theoretical significance and practical value.<\/jats:p>","DOI":"10.3390\/jimaging12050191","type":"journal-article","created":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T11:33:37Z","timestamp":1777376017000},"page":"191","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["PGT-Net: A Physics-Guided Transformer\u2013CNN Hybrid Network for Low-Light Image Enhancement and Object Detection in Traffic Scenes"],"prefix":"10.3390","volume":"12","author":[{"given":"Bin","family":"Chen","sequence":"first","affiliation":[{"name":"School of Computer and Artificial Intelligence, Zhengzhou University, Zhengzhou 450002, China"},{"name":"XJ Electric Co., Ltd., Xuchang 461000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jian","family":"Qiao","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Alternate Electrical Power System with Renewable Energy Sources, North China Electric Power University (Baoding), Baoding 071003, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Baowei","family":"Li","sequence":"additional","affiliation":[{"name":"XJ Electric Co., Ltd., Xuchang 461000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shipeng","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Cyber Science and Engineering, Zhengzhou University, Zhengzhou 450002, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8876-3763","authenticated-orcid":false,"given":"Wei","family":"She","sequence":"additional","affiliation":[{"name":"School of Cyber Science and Engineering, Zhengzhou University, Zhengzhou 450002, China"},{"name":"State Key Laboratory of Target Vulnerability Assessment, Luoyang 471023, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,4,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"58443","DOI":"10.1109\/ACCESS.2020.2983149","article-title":"A Survey of Autonomous Driving: Common Practices and Emerging Technologies","volume":"8","author":"Yurtsever","year":"2020","journal-title":"IEEE Access"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1561\/0600000079","article-title":"Computer vision for autonomous vehicles: Problems, datasets and state of the art","volume":"12","author":"Janai","year":"2020","journal-title":"Found. Trends\u00ae  Comput. Graph. Vis."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Wang, R., Zhang, Q., Fu, C.W., Shen, X., Zheng, W.S., and Jia, J. (2019, January 15\u201320). Underexposed photo enhancement using deep illumination estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00701"},{"key":"ref_4","unstructured":"World Health Organization (2018). Global Status Report on Road Safety 2018, World Health Organization."},{"key":"ref_5","unstructured":"Almasri, F., and Debeir, O. (2020). Visibility enhancement for drivers based on image defogging and semantic segmentation in foggy weather conditions. Proceedings of the 2020 IEEE International Conference on Image Processing (ICIP), Virtual, 25\u201328 October 2020, IEEE."},{"key":"ref_6","unstructured":"Wei, C., Wang, W., Yang, W., and Liu, J. (2018, January 3\u20136). Deep retinex decomposition for low-light enhancement. Proceedings of the British Machine Vision Conference, Newcastle, UK."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Zhang, J., and Guo, X. (2019). Kindling the darkness: A practical low-light image enhancer. Proceedings of the 27th ACM International Conference on Multimedia, Nice, France, 21\u201325 October 2019, Association for Computing Machinery.","DOI":"10.1145\/3343031.3350926"},{"key":"ref_8","unstructured":"Heckbert, P.S. (1994). Contrast limited adaptive histogram equalization. Graphics Gems IV, Academic Press Professional."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1364\/JOSA.61.000001","article-title":"Lightness and retinex theory","volume":"61","author":"Land","year":"1971","journal-title":"Josa"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"982","DOI":"10.1109\/TIP.2016.2639450","article-title":"LIME: Low-light image enhancement via illumination map estimation","volume":"26","author":"Guo","year":"2016","journal-title":"IEEE Trans. Image Process."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"391","DOI":"10.1007\/s11760-025-03914-1","article-title":"Deep decomposer and refiner for low-light image enhancement","volume":"19","author":"Vaish","year":"2025","journal-title":"Signal Image Video Process."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"650","DOI":"10.1016\/j.patcog.2016.06.008","article-title":"LLNet: A deep autoencoder approach to natural low-light image enhancement","volume":"61","author":"Lore","year":"2017","journal-title":"Pattern Recognit."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"2340","DOI":"10.1109\/TIP.2021.3051462","article-title":"EnlightenGAN: Deep light enhancement without paired supervision","volume":"30","author":"Jiang","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Guo, C., Li, C., Guo, J., Loy, C.C., Hou, J., Kwong, S., and Cong, R. (2020, January 13\u201319). Zero-reference deep curve estimation for low-light image enhancement. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00185"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Wang, Z., Cun, X., Bao, J., Zhou, W., Liu, J., and Li, H. (2022). Uformer: A general U-shaped transformer for image restoration. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18\u201324 June 2022, IEEE.","DOI":"10.1109\/CVPR52688.2022.01716"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zamir, S.W., Arora, A., Khan, S., Hayat, M., Khan, F.S., and Yang, M.H. (2022). Restormer: Efficient transformer for high-resolution image restoration. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18\u201324 June 2022, IEEE.","DOI":"10.1109\/CVPR52688.2022.00564"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Liu, W., Ren, G., Yu, R., Guo, S., Zhu, J., and Zhang, L. Image-adaptive YOLO for object detection in adverse weather conditions. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 22 February\u20131 March 2022, AAAI Press.","DOI":"10.1609\/aaai.v36i2.20072"},{"key":"ref_18","first-page":"1723","article-title":"Attention-based context aggregation network for monocular depth estimation","volume":"129","author":"Chen","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yu, C., Mai, Y., Yang, C., Zheng, J., Liu, Y., and Yu, C. (2024). IA-YOLO: A Vatica Segmentation Model Based on an Inverted Attention Block for Drone Cameras. Agriculture, 14.","DOI":"10.20944\/preprints202409.1095.v1"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"492","DOI":"10.1109\/TIP.2018.2867951","article-title":"Benchmarking single-image dehazing and beyond","volume":"28","author":"Li","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Yang, W., Tan, R.T., Feng, J., Liu, J., Guo, Z., and Yan, S. (2017). Deep joint rain detection and removal from a single image. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21\u201326 July 2017, IEEE.","DOI":"10.1109\/CVPR.2017.183"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"5187","DOI":"10.1109\/TIP.2016.2598681","article-title":"Dehazenet: An end-to-end system for single image haze removal","volume":"25","author":"Cai","year":"2016","journal-title":"IEEE Trans. Image Process."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Liu, R., Ma, L., Zhang, J., Fan, X., and Luo, Z. (2021). Retinex-inspired unrolling with cooperative prior architecture search for low-light image enhancement. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20\u201325 June 2021, IEEE.","DOI":"10.1109\/CVPR46437.2021.01042"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Rosa, E., Vaccaro, M., Placidi, E., D\u2019Andrea, M.L., Liporace, F., Natali, G.L., Secinaro, A., and Napolitano, A. (2025). Quantum Neural Networks in Magnetic Resonance Imaging: Advancing Diagnostic Precision Through Emerging Computational Paradigms. Computers, 14.","DOI":"10.3390\/computers14120529"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1023\/A:1016328200723","article-title":"Vision and the atmosphere","volume":"48","author":"Narasimhan","year":"2002","journal-title":"Int. J. Comput. Vis."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhou, Z., Siddiquee, M.M.R., Tajbakhsh, N., and Liang, J. (2018). Unet++: A nested u-net architecture for medical image segmentation. Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support, Springer.","DOI":"10.1007\/978-3-030-00889-5_1"},{"key":"ref_28","unstructured":"Oktay, O., Schlemper, J., Folgoc, L.L., Lee, M., Heinrich, M., Misawa, K., Mori, K., McDonagh, S., Hammerla, N.Y., and Kainz, B. (2018). Attention u-net: Learning where to look for the pancreas. arXiv."},{"key":"ref_29","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4\u20139 December 2017, Curran Associates Inc."},{"key":"ref_30","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10\u201317 October 2021, IEEE.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27\u201330 June 2016, IEEE.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016). SSD: Single shot multibox detector. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wang, C.Y., Bochkovskiy, A., and Liao, H.Y.M. (2022). YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv.","DOI":"10.1109\/CVPR52729.2023.00721"},{"key":"ref_36","unstructured":"Cortes, C., Lawrence, N.D., Lee, D.D., Sugiyama, M., and Garnett, R. (2015). Faster R-CNN: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems 28, Curran Associates, Inc."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020). End-to-end object detection with transformers. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_39","unstructured":"Zheng, Z., Wang, P., Liu, W., Li, J., Ye, R., and Ren, D. (2020, January 7\u201312). Distance-IoU loss: Faster and better learning for bounding box regression. Proceedings of the AAAI Conference on Artificial Intelligence 2020, New York, NY, USA."},{"key":"ref_40","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV) 2018, Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Wang, C.Y., Liao, H.Y.M., Wu, Y.H., Chen, P.Y., Hsieh, J.W., and Yeh, I.H. (2020, January 14\u201319). CSPNet: A new backbone that can enhance learning capability of CNN. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops 2020, Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00203"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2017, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201323). Path aggregation network for instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Yu, F., Chen, H., Wang, X., Xian, W., Chen, Y., Liu, F., Madhavan, V., and Darrell, T. (2020, January 13\u201319). BDD100K: A diverse driving dataset for heterogeneous multitask learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2020, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00271"},{"key":"ref_47","unstructured":"Loshchilov, I., and Hutter, F. (2017). Decoupled weight decay regularization. arXiv."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"30","DOI":"10.1016\/j.cviu.2018.10.010","article-title":"Getting to know low-light images with the exclusively dark dataset","volume":"178","author":"Loh","year":"2019","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Chen, C., Chen, Q., Xu, J., and Koltun, V. (2018, January 18\u201323). Learning to see in the dark. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00347"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Zhang, R., Isola, P., Efros, A.A., Shechtman, E., and Wang, O. (2018, January 18\u201323). The unreasonable effectiveness of deep features as a perceptual metric. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00068"},{"key":"ref_51","unstructured":"Wang, C.-Y., Yeh, I.-H., and Liao, H.-Y.M. (2024). YOLOv21: A Scalable and Efficient Architecture for Real-Time Object Detection. arXiv."},{"key":"ref_52","unstructured":"Jocher, G. (2020). YOLOv5 by Ultralytics, Zenodo. Computer software."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"1680","DOI":"10.3390\/make5040083","article-title":"A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS","volume":"5","author":"Terven","year":"2023","journal-title":"Mach. Learn. Knowl. Extr."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. (2017, January 22\u201329). Grad-cam: Visual explanations from deep networks via gradient-based localization. Proceedings of the IEEE International Conference on Computer Vision 2017, Venice, Italy.","DOI":"10.1109\/ICCV.2017.74"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Huang, X., and Belongie, S. (2017, January 22\u201329). Arbitrary style transfer in real-time with adaptive instance normalization. Proceedings of the IEEE International Conference on Computer Vision 2017, Venice, Italy.","DOI":"10.1109\/ICCV.2017.167"},{"key":"ref_56","unstructured":"Hinton, G., Vinyals, O., and Dean, J. (2015). Distilling the knowledge in a neural network. arXiv."}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/5\/191\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,14]],"date-time":"2026-05-14T04:23:23Z","timestamp":1778732603000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/5\/191"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,28]]},"references-count":56,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2026,5]]}},"alternative-id":["jimaging12050191"],"URL":"https:\/\/doi.org\/10.3390\/jimaging12050191","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,28]]}}}