{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T16:40:19Z","timestamp":1783615219194,"version":"3.55.0"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"10","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62176027"],"award-info":[{"award-number":["62176027"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Chongqing Talent","award":["cstc2024ycjh-bgzxm0082"],"award-info":[{"award-number":["cstc2024ycjh-bgzxm0082"]}]},{"name":"Central University Operating Expenses","award":["2024CDJGF-044"],"award-info":[{"award-number":["2024CDJGF-044"]}]},{"name":"Key Research Capability Enhancement Project of Qiannan Normal University for Nationalities","award":["2024zdzk12"],"award-info":[{"award-number":["2024zdzk12"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,10,31]]},"abstract":"<jats:p>In this article, we propose an end-to-end Transformer-based and structure-aware dual-stream network for low-light image enhancement. First, we divide the dual-stream network into a main stream and a structure stream. The main stream is used to recover an enhanced image and supply structural information to the structure stream, whereas the structure stream is designed to extract structural features from the main stream to provide rich structural information for the enhancement process. Second, we devise a structure-gated Transformer to balance the extraction of global and local features through parallel multihead self-attention with convolution operations following a multilayer perceptron, thus extracting sufficient structural information from the main stream in the encoder part of the dual-stream network. Finally, we develop a cross-attention-based feature fusion module that divides different window sizes into distinct feature fusion stages to achieve multistage, multiscale and multitype feature fusion in the decoder part of the dual-stream network. The experimental results show that our method not only works effectively but also offers favorable computational and storage overheads.<\/jats:p>","DOI":"10.1145\/3758097","type":"journal-article","created":{"date-parts":[[2025,8,5]],"date-time":"2025-08-05T15:22:40Z","timestamp":1754407360000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Transformer-Based and Structure-Aware Dual-Stream Network for Low-Light Image Enhancement"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1874-3641","authenticated-orcid":false,"given":"Mingliang","family":"Zhou","sequence":"first","affiliation":[{"name":"School of Computer Science, Chongqing University, Chongqing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0153-7948","authenticated-orcid":false,"given":"Shuqi","family":"Han","sequence":"additional","affiliation":[{"name":"School of Computer Science, Chongqing University, Chongqing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3639-916X","authenticated-orcid":false,"given":"Jun","family":"Luo","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Mechanical Transmissions, Chongqing University, Chongqing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4762-5994","authenticated-orcid":false,"given":"Xu","family":"Zhuang","sequence":"additional","affiliation":[{"name":"Guangdong OPPO Mobile Telecommunications Corp., Ltd., Dongguan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-6146-7794","authenticated-orcid":false,"given":"Qin","family":"Mao","sequence":"additional","affiliation":[{"name":"The School of Computer and Information Technology, Qiannan Normal University for Nationalities, Duyun, China, and The Key Laboratory of Complex Systems and Intelligent Optimization of Guizhou Province, Dunyun, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4525-1204","authenticated-orcid":false,"given":"Zhengguo","family":"Li","sequence":"additional","affiliation":[{"name":"Infocomm Research, Agency for Science, Technology and Research, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,10,14]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"12504","volume-title":"IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Cai Yuanhao","year":"2023","unstructured":"Yuanhao Cai, Hao Bian, Jing Lin, Haoqian Wang, Radu Timofte, and Yulun Zhang. 2023. Retinexformer: One-stage Retinex-based transformer for low-light image enhancement. In IEEE\/CVF International Conference on Computer Vision (ICCV), 12504\u201312513."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00328"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00347"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00041"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2019.00151"},{"key":"e_1_3_1_7_2","first-page":"12299","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Chen Hanting","year":"2021","unstructured":"Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. 2021. Pre-trained image processing transformer. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12299\u201312310."},{"key":"e_1_3_1_8_2","volume-title":"International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2021. An image is worth 16\u2009\u00d7\u200916 words: Transformers for image recognition at scale. In International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=YicbFdNTTy"},{"key":"e_1_3_1_9_2","first-page":"1787","article-title":"A semantic-aware detail adaptive network for image enhancement","author":"Fan Linlin","year":"2024","unstructured":"Linlin Fan, Xuekai Wei, Mingliang Zhou, Jielu Yan, Huayan Pu, Jun Luo, and Zhengguo Li. 2024. A semantic-aware detail adaptive network for image enhancement. IEEE Transactions on Circuits and Systems for Video Technology 35 (2024), 1787\u20131800.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00185"},{"issue":"3","key":"e_1_3_1_11_2","first-page":"1","article-title":"Image defogging based on regional gradient constrained prior","volume":"20","author":"Guo Qiang","year":"2023","unstructured":"Qiang Guo, Zhi Zhang, Mingliang Zhou, Hong Yue, Huayan Pu, and Jun Luo. 2023. Image defogging based on regional gradient constrained prior. ACM Transactions on Multimedia Computing, Communications and Applications 20, 3 (2023), 1\u201317.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TBC.2022.3187816"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2016.2639450"},{"key":"e_1_3_1_14_2","unstructured":"Zhilin Huang Chujun Qin Ruixin Liu Zhenyu Weng and Yuesheng Zhu. 2021. Structure-aware image inpainting with two parallel streams. arXiv:2111.03414. Retrieved from https:\/\/arxiv.org\/abs\/2111.03414"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3051462"},{"key":"e_1_3_1_16_2","doi-asserted-by":"crossref","first-page":"11296","DOI":"10.1609\/aaai.v34i07.6790","article-title":"Unpaired image enhancement featuring reinforcement-learning-controlled image editing software","author":"Kosugi Satoshi","year":"2020","unstructured":"Satoshi Kosugi and Toshihiko Yamasaki. 2020. Unpaired image enhancement featuring reinforcement-learning-controlled image editing software. In AAAI Conference on Artificial Intelligence, 11296\u201311303.","journal-title":"AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2013.2284059"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3063604"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2018.01.010"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2810539"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3689642"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW54120.2021.00210"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3545609"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01042"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2018.10.010"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2016.06.008"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413925"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2015.2442920"},{"key":"e_1_3_1_30_2","first-page":"5637","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Ma Long","year":"2022","unstructured":"Long Ma, Tengyu Ma, Risheng Liu, Xin Fan, and Zhongxuan Luo. 2022. Toward fast, flexible, and robust low-light image enhancement. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 5637\u20135646."},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01284"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48506.2021.9561051"},{"key":"e_1_3_1_33_2","first-page":"234","volume-title":"International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI \u201915)","author":"Ronneberger O.","year":"2015","unstructured":"O. Ronneberger, P. Fischer, and T. Brox. 2015. U-Net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI \u201915). Springer, 234\u2013241."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3256763"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-017-4783-x"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00701"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2013.2261309"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i3.25364"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01716"},{"key":"e_1_3_1_40_2","article-title":"Deep Retinex decomposition for low-light enhancement","author":"Wei Chen","year":"2018","unstructured":"Chen Wei, Wenjing Wang, Wenhan Yang, and Jiaying Liu. 2018. Deep Retinex decomposition for low-light enhancement. In British Machine Vision Conference.","journal-title":"British Machine Vision Conference"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3497746"},{"key":"e_1_3_1_42_2","first-page":"823","volume-title":"IEEE Transactions on Circuits and Systems for Video Technology","author":"Wu Wanyu","year":"2024","unstructured":"Wanyu Wu, Wei Wang, Zheng Wang, Kui Jiang, and Zhengguo Li. 2024. For overall nighttime visibility: Integrate irregular glow removal with glow-aware enhancement. IEEE Transactions on Circuits and Systems for Video Technology 35 (2024), 823\u2013837."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00235"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01719"},{"key":"e_1_3_1_45_2","first-page":"28522","volume-title":"Advances in Neural Information Processing Systems","author":"Xu Yufei","year":"2021","unstructured":"Yufei Xu, Qiming Zhang, Jing Zhang, and Dacheng Tao. 2021. ViTAE: Vision transformer advanced by exploring intrinsic inductive bias. In Advances in Neural Information Processing Systems. M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (Eds.), Vol. 34. Curran Associates, Inc., 28522\u201328535. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2021\/file\/efb76cff97aaf057654ef2f38cd77d73-Paper.pdf"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2022.3167331"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00313"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3062184"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3050850"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612145"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2017.356"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3321512"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00564"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58595-2_30"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3026740"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350926"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3590965"},{"key":"e_1_3_1_58_2","first-page":"1","article-title":"Blind image quality assessment: Exploring content fidelity perceptibility via quality adversarial learning","author":"Zhou Mingliang","year":"2025","unstructured":"Mingliang Zhou, Wenhao Shen, Xuekai Wei, Jun Luo, Fan Jia, Xu Zhuang, and Weijia Jia. 2025. Blind image quality assessment: Exploring content fidelity perceptibility via quality adversarial learning. International Journal of Computer Vision 133 (2025), 1\u201317.","journal-title":"International Journal of Computer Vision"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3268867"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3758097","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,14]],"date-time":"2025-10-14T21:25:19Z","timestamp":1760477119000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3758097"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,14]]},"references-count":58,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2025,10,31]]}},"alternative-id":["10.1145\/3758097"],"URL":"https:\/\/doi.org\/10.1145\/3758097","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,14]]},"assertion":[{"value":"2024-08-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-29","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}