{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,4]],"date-time":"2026-08-04T04:28:58Z","timestamp":1785817738818,"version":"3.56.0"},"reference-count":33,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2023,6,28]],"date-time":"2023-06-28T00:00:00Z","timestamp":1687910400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62166043"],"award-info":[{"award-number":["62166043"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U2003207"],"award-info":[{"award-number":["U2003207"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Text recognition is an important research topic in computer vision. Scene text, which refers to the text in real scenes, sometimes needs to meet the requirement of attracting attention, and there is the situation such as deformation. At the same time, the image acquisition process is affected by factors such as occlusion, noise, and obstruction, making scene text recognition tasks more challenging. In this paper, we improve the CRNN model for text recognition, which has relatively low accuracy, poor performance in recognizing irregular text, and only considers obtaining text sequence information from a single aspect, resulting in incomplete information acquisition. Firstly, to address the problems of low text recognition accuracy and poor recognition of irregular text, we add label smoothing to ensure the model\u2019s generalization ability. Then, we introduce the smoothing loss function from speech recognition into the field of text recognition, and add a language model to increase information acquisition channels, ultimately achieving the goal of improving text recognition accuracy. This method was experimentally verified on six public datasets and compared with other advanced methods. The experimental results show that this method performs well in most benchmark tests, and the improved model outperforms the original model in recognition performance.<\/jats:p>","DOI":"10.3390\/info14070369","type":"journal-article","created":{"date-parts":[[2023,6,29]],"date-time":"2023-06-29T01:15:47Z","timestamp":1688001347000},"page":"369","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["Scene Text Recognition Based on Improved CRNN"],"prefix":"10.3390","volume":"14","author":[{"given":"Wenhua","family":"Yu","sequence":"first","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, Urumqi 830017, China"},{"name":"Xinjiang Key Laboratory of Signal Detection and Processing, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8766-0647","authenticated-orcid":false,"given":"Mayire","family":"Ibrayim","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, Urumqi 830017, China"},{"name":"Xinjiang Key Laboratory of Signal Detection and Processing, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2321-308X","authenticated-orcid":false,"given":"Askar","family":"Hamdulla","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, Urumqi 830017, China"},{"name":"Xinjiang Key Laboratory of Multilingual Information Technology, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,6,28]]},"reference":[{"key":"ref_1","first-page":"1330","article-title":"A deep learning approach for natural scene text detection and recognition","volume":"26","author":"Liu","year":"2021","journal-title":"Chin. J. Graph."},{"key":"ref_2","unstructured":"Shi, B., Bai, X., and Yao, C. (2015). An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Graves, A., Fern\u00e1ndez, S., Gomez, F., and Schmidhuber, J. (2006, January 25\u201329). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. Proceedings of the 23rd International Conference on Machine Learning, Pittsburgh, PA, USA.","DOI":"10.1145\/1143844.1143891"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Liu, Y., Wang, Y., and Shi, H. (2023). A Convolutional Recurrent Neural-Network-Based Machine Learning for Scene Text Recognition Application. Symmetry, 15.","DOI":"10.3390\/sym15040849"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"861","DOI":"10.1007\/s00138-018-0942-y","article-title":"Scene text recognition using residual convolutional recurrent neural network","volume":"29","author":"Lei","year":"2018","journal-title":"Mach. Vis. Appl."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1109\/72.279181","article-title":"Learning long-term dependencies with gradient descent is difficult","volume":"5","author":"Bengio","year":"1994","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Graves, A., Mohamed, A., and Hinton, G. (2013, January 26\u201331). Speech recognition with deep recurrent neural networks. Proceedings of the 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Shi, B., Wang, X., Lyu, P., Yao, C., and Bai, X. (2016, January 27\u201330). Robust scene text recognition with automatic rectification. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.452"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Lee, C.-Y., and Osindero, S. (2016, January 27\u201330). Recursive recurrent nets with attention modeling for OCR in the wild. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.245"},{"key":"ref_10","first-page":"7","article-title":"Star-net: A spatial attention residue network for scene text recognition","volume":"2","author":"Liu","year":"2016","journal-title":"BMVC"},{"key":"ref_11","first-page":"334","article-title":"Gated recurrent convolution neural network for OCR","volume":"30","author":"Wang","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Borisyuk, F., Gordo, A., and Sivakumar, V. (2018, January 19\u201323). Rosetta: Large scale system for text detection and recognition in images. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, London, UK.","DOI":"10.1145\/3219819.3219861"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Baek, J., Kim, G., Lee, J., Park, S., Han, D., Yun, S., Oh, S.J., and Lee, H. (November, January 27). What is wrong with scene text recognition model comparisons? Dataset and model analysis. Proceedings of the 2019 IEEE\/CVF international Conference on Computer Vision, Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00481"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Qiao, Z., Zhou, Y., Yang, D., Zhou, Y., and Wang, W. (2020, January 13\u201319). Seed: Semantics enhanced encoder-decoder framework for scene text recognition. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01354"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Atienza, R. (2021, January 5\u201310). Vision transformer for fast and efficient scene text recognition. Proceedings of the Document Analysis and Recognition\u2014ICDAR 2021: 16th International Conference, Lausanne, Switzerland. Proceedings, Part I 16.","DOI":"10.1007\/978-3-030-86549-8_21"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Baek, J., Matsui, Y., and Aizawa, K. (2021, January 20\u201325). What if we only use real datasets for scene text recognition? Toward scene text recognition with fewer labels. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00313"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zhang, M., Ma, M., and Wang, P. (2021, January 21\u201324). Scene text recognition with cascade attention network. Proceedings of the 2021 International Conference on Multimedia Retrieval, New York, NY, USA.","DOI":"10.1145\/3460426.3463639"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Bhunia, A.K., Sain, A., Chowdhury, P.N., and Song, Y.-Z. (2021, January 10\u201317). Text is text, no matter what: Unifying text recognition using knowledge distillation. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00102"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Liu, C., Yang, C., and Yin, X.C. (2022, January 18\u201324). Open-Set Text Recognition via Character-Context Decoupling. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00448"},{"key":"ref_20","first-page":"1","article-title":"A cervical cell classification method based on migration learning and label smoothing strategy","volume":"28","author":"Liu","year":"2022","journal-title":"Mod. Comput."},{"key":"ref_21","first-page":"422","article-title":"When does label smoothing help?","volume":"32","author":"Kornblith","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_22","unstructured":"Zhao, L. (2021). Research on User Behavior Recognition Based on CNN and LSTM. [Master\u2019s Thesis, Nanjing University of Information Engineering]."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Kim, S., Seltzer, M.L., Li, J., and Zhao, R. (2017). Improved training for online end-to-end speech recognition systems. arXiv.","DOI":"10.21437\/Interspeech.2018-2517"},{"key":"ref_24","unstructured":"Qin, C. (2020). Research on End-to-End Speech Recognition Technology. [Ph.D. Thesis, Strategic Support Force Information Engineering University]."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Fang, S., Xie, H., Wang, Y., Mao, Z., and Zhang, Y. (2021, January 20\u201325). Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00702"},{"key":"ref_26","first-page":"627","article-title":"A multi-scale deformable convolution network model for text recognition","volume":"Volume 12083","author":"Cheng","year":"2022","journal-title":"Proceedings of the Thirteenth International Conference on Graphics and Image Processing (ICGIP 2021)"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Gupta, A., Vedaldi, A., and Zisserman, A. (2016, January 27\u201330). Synthetic data for text localisation in natural images. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.254"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Mishra, A., Alahari, K., and Jawahar, C.V. (2012, January 16\u201321). Top-down and bottom-up cues for scene text recognition. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6247990"},{"key":"ref_29","unstructured":"Wang, K., Babenko, B., and Belongie, S. (2011, January 6\u201313). End-to-end scene text recognition. Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Karatzas, D., Shafait, F., Uchida, S., Iwamura, M., Bigorda, L.G., Mestre, S.R., Mas, J., Mota, D.F., Almazan, J.A., and de las Heras, L.P. (2013, January 25\u201328). ICDAR 2013 robust reading competition. Proceedings of the 2013 12th International Conference on Document Analysis and Recognition, Washington, DC, USA.","DOI":"10.1109\/ICDAR.2013.221"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Phan, T.Q., Shivakumara, P., Tian, S., and Tan, C.L. (2013, January 1\u20138). Recognizing text with perspective distortion in natural scenes. Proceedings of the 2013 IEEE International Conference on Computer Vision, Sydney, NSW, Australia.","DOI":"10.1109\/ICCV.2013.76"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"8027","DOI":"10.1016\/j.eswa.2014.07.008","article-title":"A robust arbitrary text detection system for natural scene images","volume":"41","author":"Risnumawan","year":"2014","journal-title":"Expert Syst. Appl."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Cheng, Z., Xu, Y., Bai, F., Niu, Y., Pu, S., and Zhou, S. (2018, January 18\u201323). Aon: Towards arbitrarily-oriented text recognition. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00584"}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/14\/7\/369\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:03:09Z","timestamp":1760126589000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/14\/7\/369"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,28]]},"references-count":33,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2023,7]]}},"alternative-id":["info14070369"],"URL":"https:\/\/doi.org\/10.3390\/info14070369","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,28]]}}}