{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T12:05:30Z","timestamp":1776168330378,"version":"3.50.1"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>Accurate Chinese text recognition (CTR) is vital for applications such as document digitization, but remains challenging due to high inter-class visual similarity, complex hierarchical structures, and diverse visual degradations in real-world scenes. To address these challenges, we propose a robust CTR framework that synergizes structure-aware visual discrimination with cross-modal reasoning. The hybrid encoder integrates multi-scale attention modulation and global self-attention, guided by hierarchical structural supervision, to capture fine-grained structure-aware visual representations essential for distinguishing visually similar characters. Complementing this, an iterative vision-language decoder, trained with a stochastic masking strategy, learns to reconstruct text from partially observed visual and contextual cues, enabling complementary vision-language reasoning that effectively resolves visually ambiguous or degraded characters. Extensive experiments on public benchmarks demonstrate that our method achieves state-of-the-art performance, validating its effectiveness in addressing the challenges of Chinese text recognition.<\/jats:p>","DOI":"10.1145\/3793249","type":"journal-article","created":{"date-parts":[[2026,1,23]],"date-time":"2026-01-23T21:08:51Z","timestamp":1769202531000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Improving Chinese Text Recognition with Multi-Granularity Features and Vision-Language Reasoning"],"prefix":"10.1145","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0441-6655","authenticated-orcid":false,"given":"Tao","family":"Li","sequence":"first","affiliation":[{"name":"School of Information Science and Technology, University of Science and Technology of China","place":["Hefei, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1859-900X","authenticated-orcid":false,"given":"Zengfu","family":"Wang","sequence":"additional","affiliation":[{"name":"Chinese Academy of Sciences Institute of Intelligent Machines","place":["Hefei, China"]},{"name":"School of Information Science and Technology, University of Science and Technology of China","place":["Hefei, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,14]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19815-1_11"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00520"},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","first-page":"1571","DOI":"10.1109\/ICDAR.2019.00252","volume-title":"Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR)","author":"Chng Chee Kheng","year":"2019","unstructured":"Chee Kheng Chng, Yuliang Liu, Yipeng Sun, Chun Chet Ng, Canjie Luo, Zihan Ni, ChuanMing Fang, Shuaitao Zhang, Junyu Han, Errui Ding, et\u00a0al. 2019. Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art. In Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR). IEEE, 1571\u20131576."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.195"},{"key":"e_1_3_1_6_2","first-page":"322","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Da Cheng","year":"2022","unstructured":"Cheng Da, Peng Wang, and Cong Yao. 2022. Levenshtein ocr. In Proceedings of the European Conference on Computer Vision. Springer, 322\u2013338."},{"key":"e_1_3_1_7_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et\u00a0al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_8_2","first-page":"2798","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Du Yongkun","year":"2025","unstructured":"Yongkun Du, Zhineng Chen, Caiyan Jia, Xieping Gao, and Yu-Gang Jiang. 2025. Out of length text recognition with sub-string matching. In Proceedings of the AAAI Conference on Artificial Intelligence. 2798\u20132806."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","unstructured":"Yongkun Du Zhineng Chen Caiyan Jia Xiaoting Yin Chenxia Li Yuning Du and Yu-Gang Jiang. 2025. Context perception parallel decoder for scene text recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 47 6 (2025) 4668\u20134683. DOI:10.1109\/TPAMI.2025.3545453","DOI":"10.1109\/TPAMI.2025.3545453"},{"key":"e_1_3_1_10_2","first-page":"884","volume-title":"Proceedings of the 31st International Joint Conference on Artificial Intelligence, IJCAI-22","author":"Du Yongkun","year":"2022","unstructured":"Yongkun Du, Zhineng Chen, Caiyan Jia, Xiaoting Yin, Tianlun Zheng, Chenxia Li, Yuning Du, and Yu-Gang Jiang. 2022. SVTR: Scene text recognition with a single visual model. In Proceedings of the 31st International Joint Conference on Artificial Intelligence, IJCAI-22. 884\u2013890."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","unstructured":"Yongkun Du Zhineng Chen Yuchen Su Caiyan Jia and Yu-Gang Jiang. 2025. Instruction-guided scene text recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 47 4 (2025) 2723\u20132738. DOI:10.1109\/TPAMI.2025.3525526","DOI":"10.1109\/TPAMI.2025.3525526"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00702"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143891"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01186"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_16_2","first-page":"7","volume-title":"Proceedings of the 2018 24th International Conference on Pattern Recognition (ICPR)","author":"He Mengchao","year":"2018","unstructured":"Mengchao He, Yuliang Liu, Zhibo Yang, Sheng Zhang, Canjie Luo, Feiyu Gao, Qi Zheng, Yongpan Wang, Xin Zhang, and Lianwen Jin. 2018. ICPR2018 contest on robust reading for multi-type web images. In Proceedings of the 2018 24th International Conference on Pattern Recognition (ICPR). IEEE, 7\u201312."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.245"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW50498.2020.00281"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018610"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i11.26538"},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","unstructured":"Yanyu Li Geng Yuan Yang Wen Ju Hu Georgios Evangelidis Sergey Tulyakov Yanzhi Wang and Jian Ren. 2022. Efficientformer: Vision transformers at mobilenet speed. Advances in Neural Information Processing Systems (NeurIPS 2022) 35 (2022) 12934\u201312949.","DOI":"10.52202\/068431-0940"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00553"},{"key":"e_1_3_1_23_2","first-page":"274","volume-title":"Proceedings of the Document Analysis and Recognition\u2013ICDAR 2021: 16th International Conference, Lausanne, Switzerland, September 5\u201310, 2021, Part III 16","author":"Liu Brian","year":"2021","unstructured":"Brian Liu, Weicong Sun, Wenjing Kang, and Xianchao Xu. 2021. Searching from the prediction of visual and language model for handwritten Chinese text recognition. In Proceedings of the Document Analysis and Recognition\u2013ICDAR 2021: 16th International Conference, Lausanne, Switzerland, September 5\u201310, 2021, Part III 16. Springer, 274\u2013288."},{"key":"e_1_3_1_24_2","doi-asserted-by":"crossref","unstructured":"Canjie Luo Lianwen Jin and Zenghui Sun. 2019. Moran: A multi-object rectified attention network for scene text recognition. Pattern Recognition 90 C (2019) 109\u2013118.","DOI":"10.1016\/j.patcog.2019.01.020"},{"key":"e_1_3_1_25_2","unstructured":"Pengyuan Lyu Chengquan Zhang Shanshan Liu Meina Qiao Yangliu Xu Liang Wu Kun Yao Junyu Han Errui Ding and Jingdong Wang. 2023. MaskOCR: Text Recognition with Masked Encoder-Decoder Pretraining. arXiv:2206.00311. Retrieved from https:\/\/arxiv.org\/abs\/2206.00311"},{"key":"e_1_3_1_26_2","doi-asserted-by":"crossref","unstructured":"Tianlong Ma Xiangcheng Du Xingjiao Wu Zhao Zhou Yingbin Zheng and Cheng Jin. 2023. Reading scene text with aggregated temporal convolutional encoder. ACM Transactions on Asian and Low-Resource Language Information Processing 22 11 (2023) 1\u201316.","DOI":"10.1145\/3625822"},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","unstructured":"Dezhi Peng Lianwen Jin Weihong Ma Canyu Xie Hesuo Zhang Shenggao Zhu and Jing Li. 2022. Recognition of handwritten Chinese text by segmentation: A segment-annotation-free approach. IEEE Transactions on Multimedia 25 (2022) 2368\u20132381.","DOI":"10.1109\/TMM.2022.3146771"},{"key":"e_1_3_1_28_2","first-page":"13528","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Qiao Zhi","year":"2020","unstructured":"Zhi Qiao, Yu Zhou, Dongbao Yang, Yucan Zhou, and Weiping Wang. 2020. Seed: Semantics enhanced encoder-decoder framework for scene text recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 13528\u201313537."},{"key":"e_1_3_1_29_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et\u00a0al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","unstructured":"Baoguang Shi Xiang Bai and Cong Yao. 2016. An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 39 11 (2016) 2298\u20132304.","DOI":"10.1109\/TPAMI.2016.2646371"},{"key":"e_1_3_1_31_2","doi-asserted-by":"crossref","unstructured":"Baoguang Shi Mingkun Yang Xinggang Wang Pengyuan Lyu Cong Yao and Xiang Bai. 2018. Aster: An attentional scene text recognizer with flexible rectification. IEEE Transactions on Pattern Analysis and Machine Intelligence 41 9 (2018) 2035\u20132048.","DOI":"10.1109\/TPAMI.2018.2848939"},{"key":"e_1_3_1_32_2","first-page":"1429","volume-title":"Proceedings of the 2017 14th iapr International Conference on Document Analysis and Recognition (ICDAR)","author":"Shi Baoguang","year":"2017","unstructured":"Baoguang Shi, Cong Yao, Minghui Liao, Mingkun Yang, Pei Xu, Linyan Cui, Serge Belongie, Shijian Lu, and Xiang Bai. 2017. Icdar2017 competition on reading chinese text in the wild (rctw-17). In Proceedings of the 2017 14th iapr International Conference on Document Analysis and Recognition (ICDAR). IEEE, 1429\u20131434."},{"key":"e_1_3_1_33_2","first-page":"1557","volume-title":"Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR)","author":"Sun Yipeng","year":"2019","unstructured":"Yipeng Sun, Zihan Ni, Chee-Kheng Chng, Yuliang Liu, Canjie Luo, Chun Chet Ng, Junyu Han, Errui Ding, Jingtuo Liu, Dimosthenis Karatzas, et\u00a0al. 2019. ICDAR 2019 competition on large-scale street view text with partial labeling-RRC-LSVT. In Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR). IEEE, 1557\u20131562."},{"key":"e_1_3_1_34_2","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems (NeurIPS 2017) 30 (2017) 6000\u20136010."},{"key":"e_1_3_1_35_2","first-page":"12120","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Wan Zhaoyi","year":"2020","unstructured":"Zhaoyi Wan, Minghang He, Haoran Chen, Xiang Bai, and Cong Yao. 2020. Textscanner: Reading characters in order for robust scene text recognition. In Proceedings of the AAAI Conference on Artificial Intelligence. 12120\u201312127."},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","unstructured":"Kai Wang Mingliang Zhou Qing Lin Guanglin Niu and Xiaowei Zhang. 2025. Geometry-guided point generation for 3D object detection. IEEE Signal Processing Letters 32 (2025) 136\u2013140.","DOI":"10.1109\/LSP.2024.3503359"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01393"},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","unstructured":"Shilian Wu Yongrui Li and Zengfu Wang. 2023. Chinese text recognition enhanced by glyph and character semantic information. International Journal on Document Analysis and Recognition 27 1 (2023) 1\u201312.","DOI":"10.1007\/s10032-023-00444-9"},{"key":"e_1_3_1_39_2","first-page":"792","volume-title":"Proceedings of the 2023 IEEE International Conference on Multimedia and Expo (ICME)","author":"Wu Shilian","year":"2023","unstructured":"Shilian Wu, Yongrui Li, and Zengfu Wang. 2023. Improving CTC-based handwritten Chinese text recognition with cross-modality knowledge distillation and feature aggregation. In Proceedings of the 2023 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 792\u2013797."},{"key":"e_1_3_1_40_2","doi-asserted-by":"crossref","unstructured":"Yi-Chao Wu Fei Yin and Cheng-Lin Liu. 2017. Improving handwritten Chinese text recognition using neural network language models and convolutional neural network shape models. Pattern Recognition 65 C (2017) 251\u2013264.","DOI":"10.1016\/j.patcog.2016.12.026"},{"key":"e_1_3_1_41_2","doi-asserted-by":"crossref","unstructured":"Mingkun Yang Biao Yang Minghui Liao Yingying Zhu and Xiang Bai. 2024. Class-aware mask-guided feature refinement for scene text recognition. Pattern Recognition 149 Article 110244 (2024).","DOI":"10.1016\/j.patcog.2023.110244"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01213"},{"key":"e_1_3_1_43_2","unstructured":"Haiyang Yu Jingye Chen Bin Li Jianqi Ma Mengnan Guan Xixi Xu Xiaocong Wang Shaobo Qu and Xiangyang Xue. 2022. Benchmarking Chinese Text Recognition: Datasets Baselines and an Empirical Study. arXiv:2112.15093. Retrieved from https:\/\/arxiv.org\/abs\/2112.15093"},{"key":"e_1_3_1_44_2","first-page":"11943","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Yu Haiyang","year":"2023","unstructured":"Haiyang Yu, Xiaocong Wang, Bin Li, and Xiangyang Xue. 2023. Chinese text recognition with a pre-trained CLIP-Like model through image-IDS aligning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 11943\u201311952."},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","unstructured":"Tai-Ling Yuan Zhe Zhu Kun Xu Cheng-Jun Li Tai-Jiang Mu and Shi-Min Hu. 2019. A large chinese text dataset in the wild. Journal of Computer Science and Technology 34 3 (2019) 509\u2013521.","DOI":"10.1007\/s11390-019-1923-y"},{"key":"e_1_3_1_46_2","doi-asserted-by":"crossref","unstructured":"Hesuo Zhang Lingyu Liang and Lianwen Jin. 2020. SCUT-HCCDoc: A new benchmark dataset of handwritten chinese text in unconstrained camera-captured documents. Pattern Recognition 108 Article 107559 (2020).","DOI":"10.1016\/j.patcog.2020.107559"},{"key":"e_1_3_1_47_2","doi-asserted-by":"crossref","first-page":"1577","DOI":"10.1109\/ICDAR.2019.00253","volume-title":"Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR)","author":"Zhang Rui","year":"2019","unstructured":"Rui Zhang, Yongsheng Zhou, Qianyi Jiang, Qi Song, Nan Li, Kai Zhou, Lei Wang, Dong Wang, Minghui Liao, Mingkun Yang, et\u00a0al. 2019. Icdar 2019 robust reading challenge on reading chinese text on signboard. In Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR). IEEE, 1577\u20131581."},{"key":"e_1_3_1_48_2","doi-asserted-by":"crossref","unstructured":"Xiaowei Zhang Xiao Dou Xinpeng Zhao Guocong Li and Zekang Wang. 2024. Instance-aware diversity feature generation for unsupervised person re-identification. Displays 83 (2024) 102717.","DOI":"10.1016\/j.displa.2024.102717"},{"key":"e_1_3_1_49_2","doi-asserted-by":"crossref","unstructured":"Xiaowei Zhang Jianwei Ma Hong Liu Hai-Miao Hu and Peng Yang. 2022. Dual attentional Siamese network for visual tracking. Displays 74 (2022) 102205.","DOI":"10.1016\/j.displa.2022.102205"},{"key":"e_1_3_1_50_2","first-page":"7441","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Zhang Ziyin","year":"2024","unstructured":"Ziyin Zhang, Ning Lu, Minghui Liao, Yongshuai Huang, Cheng Li, Min Wang, and Wei Peng. 2024. Self-distillation regularized connectionist temporal classification loss for text recognition: A simple yet effective approach. In Proceedings of the AAAI Conference on Artificial Intelligence. 7441\u20137449."},{"key":"e_1_3_1_51_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Zhao Shuai","year":"2024","unstructured":"Shuai Zhao, Xiaohan Wang, Linchao Zhu, and Yi Yang. 2024. Test-time adaptation with CLIP reward for zero-shot generalization in vision-language models. In Proceedings of the 12th International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=kIP0duasBb"},{"key":"e_1_3_1_52_2","doi-asserted-by":"crossref","unstructured":"Tianlun Zheng Zhineng Chen Shancheng Fang Hongtao Xie and Yu-Gang Jiang. 2024. Cdistnet: Perceiving multi-domain character distance for robust text recognition. International Journal of Computer Vision 132 2 (2024) 300\u2013318.","DOI":"10.1007\/s11263-023-01880-0"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547827"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3793249","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T11:24:50Z","timestamp":1776165890000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3793249"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,14]]},"references-count":52,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3793249"],"URL":"https:\/\/doi.org\/10.1145\/3793249","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,14]]},"assertion":[{"value":"2024-05-04","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-17","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}