{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T13:01:56Z","timestamp":1784120516417,"version":"3.55.0"},"reference-count":51,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2023,5,5]],"date-time":"2023-05-05T00:00:00Z","timestamp":1683244800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China Joint Fund Project","award":["U1603262"],"award-info":[{"award-number":["U1603262"]}]},{"name":"National Natural Science Foundation of China Joint Fund Project","award":["U1911401"],"award-info":[{"award-number":["U1911401"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Scene text recognition (STR) has been a hot research field in computer vision, aiming to recognize text in natural scenes using computers. Currently, attention-based encoder\u2013decoder frameworks struggle to precisely align feature regions with the target object when dealing with complex and low-quality images, a phenomenon known as attention drift. Additionally, with the rise of Transformer, the increasing size of parameters results in higher computational costs. In order to solve the above problems, based on the latest research results of Vision Transformer (ViT), we utilize an additional position-enhancement branch to alleviate attention drift and dynamically fused position information with visual information to achieve better recognition accuracy. The experimental results demonstrate that our model achieves a 3% higher average recognition accuracy on the test set compared to the baseline. Meanwhile, our model maintains the advantage of a small number of parameters and fast inference speed, achieving a good balance between accuracy, speed, and computational load.<\/jats:p>","DOI":"10.3390\/s23094490","type":"journal-article","created":{"date-parts":[[2023,5,5]],"date-time":"2023-05-05T02:56:51Z","timestamp":1683255411000},"page":"4490","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Lightweight Scene Text Recognition Based on Transformer"],"prefix":"10.3390","volume":"23","author":[{"given":"Xin","family":"Luan","sequence":"first","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"},{"name":"Xinjiang Laboratory of Multi-Language Information Technology, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"},{"name":"Xinjiang Multilingual Information Technology Research Center, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinwei","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"},{"name":"Xinjiang Laboratory of Multi-Language Information Technology, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"},{"name":"Xinjiang Multilingual Information Technology Research Center, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Miaomiao","family":"Xu","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"},{"name":"Xinjiang Laboratory of Multi-Language Information Technology, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"},{"name":"Xinjiang Multilingual Information Technology Research Center, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wushouer","family":"Silamu","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"},{"name":"Xinjiang Laboratory of Multi-Language Information Technology, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"},{"name":"Xinjiang Multilingual Information Technology Research Center, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanbing","family":"Li","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"},{"name":"Xinjiang Laboratory of Multi-Language Information Technology, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"},{"name":"Xinjiang Multilingual Information Technology Research Center, Xinjiang University, No. 777 Huarui Street, Urumqi 830017, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,5,5]]},"reference":[{"key":"ref_1","unstructured":"Mandavia, K., Badelia, P., Ghosh, S., and Chaudhuri, A. (2017). Optical Character Recognition Systems for Different Languages with Soft Computing, Springer."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1109\/34.824820","article-title":"Twenty years of document image analysis in PAMI","volume":"22","author":"Nagy","year":"2000","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_3","first-page":"21","article-title":"Automatic number plate recognition system (anpr): A survey","volume":"69","author":"Patel","year":"2013","journal-title":"Int. J. Comput. Appl."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Laroca., R., Cardoso., E.V., Lucio., D.R., Estevam., V., and Menotti., D. (2022, January 6\u20138). On the Cross-dataset Generalization in License Plate Recognition. Proceedings of the 17th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, Vienna, Austria.","DOI":"10.5220\/0010846800003124"},{"key":"ref_5","unstructured":"Hwang, W., Kim, S., Seo, M., Yim, J., Park, S., Park, S., Lee, J., Lee, B., and Lee, H. (2019, January 8\u201314). Post-OCR parsing: Building simple and robust parser via BIO tagging. Proceedings of the Workshop on Document Intelligence at NeurIPS 2019, Vancouver, BC, Canada."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Limonova, E., Bezmaternykh, P., Nikolaev, D., and Arlazarov, V. (2016, January 18\u201320). Slant rectification in Russian passport OCR system using fast Hough transform. Proceedings of the Ninth International Conference on Machine Vision (ICMV 2016), Nice, France.","DOI":"10.1117\/12.2268725"},{"key":"ref_7","unstructured":"Yao, C., Bai, X., Sang, N., Zhou, X., Zhou, S., and Cao, Z. (2016). Scene text detection via holistic, multi-channel prediction. arXiv."},{"key":"ref_8","unstructured":"Liu, J., Liu, X., Sheng, J., Liang, D., Li, X., and Liu, Q. (2019). Pyramid mask text detector. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s11263-015-0823-z","article-title":"Reading text in the wild with convolutional neural networks","volume":"116","author":"Jaderberg","year":"2016","journal-title":"Int. J. Comput. Vis."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Borisyuk, F., Gordo, A., and Sivakumar, V. (2018, January 19\u201323). Rosetta: Large scale system for text detection and recognition in images. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, London, UK.","DOI":"10.1145\/3219819.3219861"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Luong, T., Pham, H., and Manning, C.D. (2015, January 17\u201321). Effective Approaches to Attention-based Neural Machine Translation. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Lisbon, Portugal.","DOI":"10.18653\/v1\/D15-1166"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Lee, C.Y., and Osindero, S. (2016, January 27\u201330). Recursive recurrent nets with attention modeling for ocr in the wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.245"},{"key":"ref_13","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the NIPS\u201917: Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Cheng, Z., Bai, F., Xu, Y., Zheng, G., Pu, S., and Zhou, S. (2017, January 22\u201329). Focusing attention: Towards accurate text recognition in natural images. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.543"},{"key":"ref_15","unstructured":"Zheng, T., Chen, Z., Fang, S., Xie, H., and Jiang, Y.G. (2021). Cdistnet: Perceiving multi-domain character distance for robust text recognition. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Yue, X., Kuang, Z., Lin, C., Sun, H., and Zhang, W. (2020, January 23\u201328). Robustscanner: Dynamically enhancing positional clues for robust text recognition. Proceedings of the Computer Vision\u2014ECCV 2020: 16th European Conference, Glasgow, UK.","DOI":"10.1007\/978-3-030-58529-7_9"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Atienza, R. (2021, January 5\u201310). Vision transformer for fast and efficient scene text recognition. Proceedings of the Document Analysis and Recognition\u2014ICDAR 2021: 16th International Conference, Lausanne, Switzerland.","DOI":"10.1007\/978-3-030-86549-8_21"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Neumann, L., and Matas, J. (2012, January 16\u201321). Real-time scene text localization and recognition. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248097"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yao, C., Bai, X., Shi, B., and Liu, W. (2014, January 23\u201328). Strokelets: A learned multi-scale representation for scene text recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.515"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Yang, X., He, D., Zhou, Z., Kifer, D., and Giles, C.L. (2017, January 19\u201325). Learning to read irregular text with attention mechanisms. Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, Melbourne, Australia.","DOI":"10.24963\/ijcai.2017\/458"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Lee, C.Y., Bhardwaj, A., Di, W., Jagadeesh, V., and Piramuthu, R. (2014, January 23\u201328). Region-based discriminative feature pooling for scene text recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.516"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Cho, K., van Merri\u00ebnboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014, January 25\u201329). Learning Phrase Representations using RNN Encoder\u2013Decoder for Statistical Machine Translation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Bai, F., Cheng, Z., Niu, Y., Pu, S., and Zhou, S. (2018, January 18\u201322). Edit probability for scene text recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00163"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Wan, Z., Zhang, J., Zhang, L., Luo, J., and Yao, C. (2020, January 13\u201319). On vocabulary reliance in scene text recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01144"},{"key":"ref_25","unstructured":"Liao, M., Zhang, J., Wan, Z., Xie, F., Liang, J., Lyu, P., Yao, C., and Bai, X. (February, January 27). Scene text recognition from two-dimensional perspective. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wan, Z., He, M., Chen, H., Bai, X., and Yao, C. (2020, January 7\u201312). Textscanner: Reading characters in order for robust scene text recognition. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6891"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Liu, W., Chen, C., Wong, K.Y.K., Su, Z., and Han, J. (2016, January 19\u201322). Star-net: A spatial attention residue network for scene text recognition. Proceedings of the BMVC, York, UK.","DOI":"10.5244\/C.30.43"},{"key":"ref_28","unstructured":"Wan, Z., Xie, F., Liu, Y., Bai, X., and Yao, C. (2019). 2D-CTC for scene text recognition. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"2298","DOI":"10.1109\/TPAMI.2016.2646371","article-title":"An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition","volume":"39","author":"Shi","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_30","unstructured":"Gao, Y., Chen, Y., Wang, J., and Lu, H. (2017). Reading scene text with attention convolutional sequence modeling. arXiv."},{"key":"ref_31","unstructured":"Li, H., Wang, P., Shen, C., and Zhang, G. (February, January 27). Show, attend and read: A simple and strong baseline for irregular text recognition. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_32","unstructured":"Wang, T., Zhu, Y., Jin, L., Luo, C., Chen, X., Wu, Y., Wang, Q., and Cai, M. (2020, January 7\u201312). Decoupled attention network for text recognition. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA."},{"key":"ref_33","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3\u20137). An Image is Worth 16 \u00d7 16 Words: Transformers for Image Recognition at Scale. Proceedings of the International Conference on Learning Representations, Virtual."},{"key":"ref_34","unstructured":"Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and J\u00e9gou, H. (2021, January 18\u201324). Training data-efficient image transformers & distillation through attention. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_35","unstructured":"Jaderberg, M., Simonyan, K., Vedaldi, A., and Zisserman, A. (2014). Synthetic data and artificial neural networks for natural scene text recognition. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Gupta, A., Vedaldi, A., and Zisserman, A. (2016, January 27\u201330). Synthetic data for text localisation in natural images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.254"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Bautista, D., and Atienza, R. (2022, January 23\u201327). Scene text recognition with permuted autoregressive sequence models. Proceedings of the Computer Vision\u2014ECCV 2022: 17th European Conference, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19815-1_11"},{"key":"ref_38","unstructured":"Yin, F., Liu, C.L., and Lu, S. (2017, January 9\u201315). A new artwork dataset for experiments in artwork retrieval and recognition. Proceedings of the 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), Kyoto, Japan."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_40","unstructured":"Liu, Y., Jin, L., Zhang, S., and Zhang, Z. (February, January 27). LSVT: Large scale vocabulary training for scene text recognition. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_41","unstructured":"Liu, Y., Chen, H., Shen, C., He, T., Jin, L., Wang, L., and Zhang, Z. (2019, January 15\u201320). Multi-lingual scene text dataset and benchmarks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA."},{"key":"ref_42","unstructured":"Yin, F., Zhang, Y., Liu, C.L., and Lu, S. (2018, January 24\u201327). ICDAR2017 competition on reading Chinese text in the wild (RCTW-17). Proceedings of the 2018 13th IAPR International Workshop on Document Analysis Systems (DAS), Vienna, Austria."},{"key":"ref_43","unstructured":"Krylov, I., Nosov, S., and Sovrasov, V. (2021, January 17\u201319). Open images v5 text annotation and yet another mask text spotter. Proceedings of the Asian Conference on Machine Learning, Virtually."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Mishra, A., Shekhar, R., and Jawahar, C. (2012, January 7\u201313). Scene text recognition using higher order language priors. Proceedings of the European Conference on Computer Vision, Florence, Italy.","DOI":"10.5244\/C.26.127"},{"key":"ref_45","first-page":"1468","article-title":"ReCTS: A Large-Scale Reusable Corpus for Text Spotting","volume":"21","author":"Shen","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_46","unstructured":"Schneider, J., Puigcerver, J., Ciss\u00e9, M., Ginev, D., Gruenstein, A., Gutherie, D., Jayaraman, D., Kassis, T., Kazemzadeh, F., and Llados, J. (2019). U-BER: A Large-Scale Dataset for Street Scene Text Reading. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Mishra, A., Alahari, K., and Jawahar, C.V. (2012, January 18\u201320). Top-Down and Bottom-Up Cues for Scene Text Recognition. Proceedings of the 2012 International Conference on Frontiers in Handwriting Recognition, Bari, Italy.","DOI":"10.1109\/CVPR.2012.6247990"},{"key":"ref_48","unstructured":"Wang, K., Babenko, B., and Belongie, S. (2011, January 18\u201321). End-to-End Scene Text Recognition. Proceedings of the 2011 International Conference on Document Analysis and Recognition, Beijing, China."},{"key":"ref_49","unstructured":"Yin, X.Y., Wang, X., Zhang, P., Wen, L., Bai, X., Lei, Z., and Li, S.Z. (2016, January 4\u20138). Robust Text Detection in Natural Images with Edge-Enhanced Maximally Stable Extremal Regions. Proceedings of the 23rd International Conference on Pattern Recognition (ICPR), Cancun, Mexico."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Karatzas, D., Gomez-Bigorda, L., Nicolaou, A., Ghosh, S., Bagdanov, A., Iwamura, M., Matas, J., Neumann, L., Chandrasekhar, V., and Lu, S. (2015, January 23\u201326). ICDAR 2015 Competition on Robust Reading. Proceedings of the 2015 13th International Conference on Document Analysis and Recognition (ICDAR), Tunis, Tunisia.","DOI":"10.1109\/ICDAR.2015.7333942"},{"key":"ref_51","first-page":"1834","article-title":"Recognizing text in perspective view of scene images based on unsupervised feature learning","volume":"46","author":"Cheung","year":"2013","journal-title":"Pattern Recognit."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/9\/4490\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:29:30Z","timestamp":1760124570000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/9\/4490"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,5]]},"references-count":51,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2023,5]]}},"alternative-id":["s23094490"],"URL":"https:\/\/doi.org\/10.3390\/s23094490","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,5]]}}}