{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T14:16:54Z","timestamp":1781533014477,"version":"3.54.5"},"reference-count":40,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T00:00:00Z","timestamp":1781222400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100006407","name":"Natural Science Foundation of Henan","doi-asserted-by":"crossref","award":["252300421063"],"award-info":[{"award-number":["252300421063"]}],"id":[{"id":"10.13039\/501100006407","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100009101","name":"Henan Province Basic Research Special Program for Key Research Projects of Higher Education Institutions","doi-asserted-by":"publisher","award":["24ZX005"],"award-info":[{"award-number":["24ZX005"]}],"id":[{"id":"10.13039\/501100009101","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62076223"],"award-info":[{"award-number":["62076223"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"award":["62076223"],"award-info":[{"award-number":["62076223"]}],"id":[{"id":"https:\/\/ror.org\/01h0zpd94","id-type":"ROR","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Text detection models based on DBNet have demonstrated strong performance in natural scene text detection. However, these models still suffer from the following three issues. Firstly, the amplifying factor hyperparameter in the differentiable binarization (DB) makes it difficult for the text detection model to achieve optimal performance. Secondly, the integration of low-level and high-level features within the backbone\u2019s feature pyramid lacks specific optimization strategies. Thirdly, the deconvolution operation in the prediction head may damage text contours. To tackle the aforementioned issues, this paper presents a text detection model termed ASM-DBNet, which mainly consists of three innovations. For the first issue, an adaptive differentiable binarization (ADB) scheme is proposed. It can independently predict amplifying factor for feature points at different spatial locations and replace the original amplifying factor hyperparameter, thereby improving the overall optimization performance of the model. For the second issue, a spatial-channel self-attention (SCA) module is proposed to optimize the fusion of high-level and low-level features. On the one hand, spatial self-attention is used to enhance the spatial localization ability of high-level features; on the other hand, channel self-attention based on a grouped transformer is used to optimize the fusion results of high-level and low-level features. For the third issue, a multi-scale context-enhanced dynamic upsampling (MC-DyUpS) module is proposed to replace the deconvolution operation in the prediction head. It enhances contextual perception in the region of interpolation points through multi-scale context feature extraction, and then accurately predicts coordinate offsets of interpolation points. The position correction based on these offsets effectively suppresses the spatial deviation caused by deconvolution. Ablation studies demonstrate the effectiveness of the SCA module, MC-DyUpS module, ADB scheme, and their arbitrary combinations. Comprehensive quantitative evaluations demonstrate that ASM-DBNet achieves competitive F1-scores of 84.1%, 84.2%, and 85.7% on the ICDAR 2015, Total-Text, and MSRA-TD500 datasets, respectively, with improvements of 1.8%, 1.4%, and 2.9% over the baseline model.<\/jats:p>","DOI":"10.3390\/info17060585","type":"journal-article","created":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T09:49:51Z","timestamp":1781257791000},"page":"585","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["ASM-DBNet: Introducing Adaptive Differentiable Binarization, Spatial-Channel Self-Attention and Multi-Scale Context-Enhanced Dynamic Upsampling for Natural Scene Text Detection"],"prefix":"10.3390","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4328-6411","authenticated-orcid":false,"given":"Xiaoliang","family":"Qian","sequence":"first","affiliation":[{"name":"College of Electrical and Information Engineering, Zhengzhou University of Light Industry, Zhengzhou 450002, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pengfei","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Electrical and Information Engineering, Zhengzhou University of Light Industry, Zhengzhou 450002, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Li","family":"Zeng","sequence":"additional","affiliation":[{"name":"College of Electrical and Information Engineering, Zhengzhou University of Light Industry, Zhengzhou 450002, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mengyang","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Electrical and Information Engineering, Zhengzhou University of Light Industry, Zhengzhou 450002, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wandian","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Electrical and Information Engineering, Zhengzhou University of Light Industry, Zhengzhou 450002, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinchao","family":"Guo","sequence":"additional","affiliation":[{"name":"College of Electrical and Information Engineering, Zhengzhou University of Light Industry, Zhengzhou 450002, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanfang","family":"Mao","sequence":"additional","affiliation":[{"name":"State Grid Nantong Power Supply Company, Nantong 226000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,6,12]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"136","DOI":"10.1007\/s10462-024-10779-2","article-title":"A review of deep learning methods for digitisation of complex documents and engineering diagrams","volume":"57","author":"Jamieson","year":"2024","journal-title":"Artif. Intell. Rev."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"126702","DOI":"10.1016\/j.neucom.2023.126702","article-title":"A survey of text detection and recognition algorithms based on deep learning technology","volume":"556","author":"Wang","year":"2023","journal-title":"Neurocomputing"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Wang, W., Xie, E., Li, X., Hou, W., Lu, T., Yu, G., and Shao, S. (2019). Shape Robust Text Detection with Progressive Scale Expansion Network. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15\u201320 June 2019, IEEE.","DOI":"10.1109\/CVPR.2019.00956"},{"key":"ref_4","first-page":"11474","article-title":"Real-Time Scene Text Detection with Differentiable Binarization","volume":"34","author":"Liao","year":"2020","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"919","DOI":"10.1109\/TPAMI.2022.3155612","article-title":"Real-Time Scene Text Detection with Differentiable Binarization and Adaptive Scale Fusion","volume":"45","author":"Liao","year":"2023","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"e12212","DOI":"10.1049\/tje2.12212","article-title":"Text detection method based on HDBNet in natural scenes","volume":"2023","author":"Wang","year":"2023","journal-title":"J. Eng."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Li, N., Wang, Z., Huang, Y., Tian, J., Li, X., and Xiao, Z. (2024). A Multi-Scale Natural Scene Text Detection Method Based on Attention Feature Extraction and Cascade Feature Fusion. Sensors, 24.","DOI":"10.3390\/s24123758"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Li, Y., Ibrayim, M., and Hamdulla, A. (2021). CSFF-Net: Scene Text Detection Based on Cross-Scale Feature Fusion. Information, 12.","DOI":"10.3390\/info12120524"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Huang, B., Bai, A., Wu, Y., Yang, C., and Sun, H. (2024). DB-EAC and LSTR: DBnet based seal text detection and Lightweight Seal Text Recognition. PLoS ONE, 19.","DOI":"10.1371\/journal.pone.0301862"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"2678","DOI":"10.1109\/TIP.2023.3272826","article-title":"Tripartite Feature Enhanced Pyramid Network for Dense Prediction","volume":"32","author":"Liu","year":"2023","journal-title":"IEEE Trans. Image Process."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Noh, H., Hong, S., and Han, B. (2015). Learning Deconvolution Network for Semantic Segmentation. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7\u201313 December 2015, IEEE.","DOI":"10.1109\/ICCV.2015.178"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Wang, J., Chen, K., Xu, R., Liu, Z., Loy, C.C., and Lin, D. (2019). CARAFE: Content-Aware ReAssembly of FEatures. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October\u20132 November 2019, IEEE.","DOI":"10.1109\/ICCV.2019.00310"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"14691","DOI":"10.1109\/TNNLS.2025.3538806","article-title":"S3INet: Semantic-Information Space Sharing Interaction Network for Arbitrary Shape Text Detection","volume":"36","author":"Wang","year":"2025","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"19770","DOI":"10.1109\/TITS.2024.3479884","article-title":"TTDNet: An End-to-End Traffic Text Detection Framework for Open Driving Environments","volume":"25","author":"Wang","year":"2024","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"15997","DOI":"10.1109\/JIOT.2025.3530682","article-title":"A Text Detection Method Based on Multiscale Selective Fusion Feature Pyramid and Multisemantic Spatial Network for Visual IoT","volume":"12","author":"Wang","year":"2025","journal-title":"IEEE Internet Things J."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"7480","DOI":"10.1109\/TCSVT.2023.3274673","article-title":"Detect Arbitrary-Shaped Text via Adaptive Thresholding and Localization Quality Estimation","volume":"33","author":"Cheng","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_17","first-page":"7069","article-title":"Explicit Relational Reasoning Network for Scene Text Detection","volume":"39","author":"Su","year":"2025","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_18","first-page":"5998","article-title":"Attention is All you Need","volume":"Volume 30","author":"Vaswani","year":"2017","journal-title":"Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4\u20139 December 2017"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Chollet, F. (2017). Xception: Deep Learning with Depthwise Separable Convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21\u201326 July 2017, IEEE.","DOI":"10.1109\/CVPR.2017.195"},{"key":"ref_20","unstructured":"Hendrycks, D., and Gimpel, K. (2016). Gaussian error linear units (GELUs). arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Karatzas, D., Gomez-Bigorda, L., Nicolaou, A., Ghosh, S., Bagdanov, A., Iwamura, M., Matas, J., Neumann, L., Chandrasekhar, V.R., and Lu, S. (2015). ICDAR 2015 competition on Robust Reading. Proceedings of the International Conference on Document Analysis and Recognition (ICDAR), Tunis, Tunisia, 23\u201326 August 2015, IEEE.","DOI":"10.1109\/ICDAR.2015.7333942"},{"key":"ref_22","first-page":"935","article-title":"Total-Text: A Comprehensive Dataset for Scene Text Detection and Recognition","volume":"Volume 1","author":"Chan","year":"2017","journal-title":"Proceedings of the International Conference on Document Analysis and Recognition (ICDAR), Kyoto, Japan, 9\u201315 November 2017"},{"key":"ref_23","unstructured":"Yao, C., Bai, X., Liu, W., Ma, Y., and Tu, Z. (2012). Detecting texts of arbitrary orientations in natural images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Providence, RI, USA, 16\u201321 June 2012, IEEE."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"15745","DOI":"10.1109\/TNNLS.2023.3289327","article-title":"Zoom Text Detector","volume":"35","author":"Yang","year":"2024","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"287","DOI":"10.1109\/TMM.2024.3521797","article-title":"Focus Entirety and Perceive Environment for Arbitrary-Shaped Text Detection","volume":"27","author":"Han","year":"2025","journal-title":"IEEE Trans. Multimed."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"9429","DOI":"10.1109\/TMM.2025.3613181","article-title":"Masked Text Pre-Training for Scene Text Detection","volume":"27","author":"Xie","year":"2025","journal-title":"IEEE Trans. Multimed."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"4737","DOI":"10.1109\/TIP.2014.2353813","article-title":"A Unified Framework for Multioriented Text Detection and Recognition","volume":"23","author":"Yao","year":"2014","journal-title":"IEEE Trans. Image Process."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27\u201330 June 2016, IEEE.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Gupta, A., Vedaldi, A., and Zisserman, A. (2016). Synthetic Data for Text Localisation in Natural Images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27\u201330 June 2016, IEEE.","DOI":"10.1109\/CVPR.2016.254"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Long, S., Ruan, J., Zhang, W., He, X., Wu, W., and Yao, C. (2018). TextSnake: A Flexible Representation for Detecting Text of Arbitrary Shapes. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8\u201314 September 2018, Springer.","DOI":"10.1007\/978-3-030-01216-8_2"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wang, W., Xie, E., Song, X., Zang, Y., Wang, W., Lu, T., Yu, G., and Shen, C. (2019). Efficient and Accurate Arbitrary-Shaped Text Detection with Pixel Aggregation Network. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October\u20132 November 2019, IEEE.","DOI":"10.1109\/ICCV.2019.00853"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"5349","DOI":"10.1109\/TPAMI.2021.3095916","article-title":"PAN++: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text","volume":"44","author":"Wang","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TIP.2022.3201467","article-title":"Fuzzy Semantics for Arbitrary-Shaped Scene Text Detection","volume":"32","author":"Wang","year":"2023","journal-title":"IEEE Trans. Image Process."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"8731","DOI":"10.1109\/TNNLS.2022.3152596","article-title":"Kernel Proposal Network for Arbitrary Shape Text Detection","volume":"34","author":"Zhang","year":"2023","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1648","DOI":"10.1109\/JAS.2024.125022","article-title":"LTDNet: A Lightweight Text Detector for Real-Time Arbitrary-Shape Traffic Text Detection","volume":"12","author":"Wang","year":"2025","journal-title":"IEEE\/CAA J. Autom. Sin."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Xue, C., Lu, S., and Zhang, W. (2019). MSR: Multi-Scale Shape Regression for Scene Text Detection. Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), Macao, China, 10\u201316 August 2019, AAAI Press.","DOI":"10.24963\/ijcai.2019\/139"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"106954","DOI":"10.1016\/j.patcog.2019.06.020","article-title":"SegLink++: Detecting Dense and Arbitrary-shaped Scene Text by Instance-aware Component Grouping","volume":"96","author":"Tang","year":"2019","journal-title":"Pattern Recognit."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Dai, P., Zhang, S., Zhang, H., and Cao, X. (2021). Progressive Contour Regression for Arbitrary-Shape Scene Text Detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20\u201325 June 2021, IEEE.","DOI":"10.1109\/CVPR46437.2021.00731"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"14200","DOI":"10.1109\/TITS.2023.3305686","article-title":"HFENet: Hybrid Feature Enhancement Network for Detecting Texts in Scenes and Traffic Panels","volume":"24","author":"Liang","year":"2023","journal-title":"IEEE Trans. Intell. Transp. Syst."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/17\/6\/585\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,14]],"date-time":"2026-06-14T04:21:17Z","timestamp":1781410877000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/17\/6\/585"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,12]]},"references-count":40,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2026,6]]}},"alternative-id":["info17060585"],"URL":"https:\/\/doi.org\/10.3390\/info17060585","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,12]]}}}