{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T15:44:16Z","timestamp":1778082256824,"version":"3.51.4"},"reference-count":85,"publisher":"Association for Computing Machinery (ACM)","issue":"6","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62376266 and 62406318"],"award-info":[{"award-number":["62376266 and 62406318"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key Laboratory of Ethnic Language Intelligent Analysis"},{"name":"Security Governance of MOE, Minzu University of China, Beijing, China"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,6,30]]},"abstract":"<jats:p>Existing scene text spotters are designed to locate and transcribe texts from images. However, it is challenging for a spotter to achieve precise detection and recognition of scene texts simultaneously. Inspired by the glimpse-focus spotting pipeline of human beings and impressive performances of Pre-trained Language Models (PLMs) on visual tasks, we ask: (1) \u201cCan machines spot texts without precise detection just like human beings?\u201d, and if yes, (2) \u201cIs text block another alternative for scene text spotting other than word or character?\u201d To this end, our proposed scene text spotter leverages advanced PLMs to enhance performance without fine-grained detection. Specifically, we first use a simple detector for block-level text detection to obtain rough positional information. Then, we fine-tune a PLM using a large-scale OCR dataset to achieve accurate recognition. Benefiting from the comprehensive language knowledge gained during the pre-training phase, the PLM-based recognition module effectively handles complex scenarios, including multi-line, reversed, occluded, and incomplete-detection texts. Taking advantage of the fine-tuned language model on scene recognition benchmarks and the paradigm of text block detection, extensive experiments demonstrate the superior performance of our scene text spotter across multiple public benchmarks. Additionally, we attempt to spot texts directly from an entire scene image to demonstrate the potential of PLMs, even Large Language Models (LLMs).<\/jats:p>","DOI":"10.1145\/3734872","type":"journal-article","created":{"date-parts":[[2025,5,8]],"date-time":"2025-05-08T15:35:27Z","timestamp":1746718527000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2051-8045","authenticated-orcid":false,"given":"Jiahao","family":"Lyu","sequence":"first","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China and School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2016-0901","authenticated-orcid":false,"given":"Jin","family":"Wei","sequence":"additional","affiliation":[{"name":"Lenovo Group Ltd, Lenovo Research, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2696-8549","authenticated-orcid":false,"given":"Gangyan","family":"Zeng","sequence":"additional","affiliation":[{"name":"School of Cyber Science and Engineering, Nanjing University of Science and Technology, Jiangyin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-7075-1869","authenticated-orcid":false,"given":"Zeng","family":"Li","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China and School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9218-3005","authenticated-orcid":false,"given":"Enze","family":"Xie","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China and School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6847-469X","authenticated-orcid":false,"given":"Wei","family":"Wang","sequence":"additional","affiliation":[{"name":"Shanghai Artificial Intelligence Laboratory, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-2307-5002","authenticated-orcid":false,"given":"Can","family":"Ma","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4188-9953","authenticated-orcid":false,"given":"Yu","family":"Zhou","sequence":"additional","affiliation":[{"name":"VCIP &amp; TMCC &amp; DISSec, College of Computer Science, Nankai University, Tianjin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,7,8]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00983"},{"key":"e_1_3_1_3_2","first-page":"457","volume-title":"ECCV","author":"Wang W.","year":"2020","unstructured":"W. Wang, X. Liu, X. Ji, E. Xie, D. Liang, Z. Yang, T. Lu, C. Shen, and P. Luo. 2020. AE Textspotter: Learning visual and linguistic representation for ambiguous text spotting. In ECCV. Springer, 457\u2013473."},{"key":"e_1_3_1_4_2","first-page":"9519","volume-title":"CVPR","author":"Zhang X.","year":"2022","unstructured":"X. Zhang, Y. Su, S. Tripathi, and Z. Tu. 2022. Text spotting transformers. In CVPR, 9519\u20139528."},{"key":"e_1_3_1_5_2","first-page":"5014","volume-title":"ACM MM","author":"Wang W.","year":"2022","unstructured":"W. Wang, Y. Zhou, J. Lv, D. Wu, G. Zhao, N. Jiang, and W. Wang. 2022. TPSNet: Reverse thinking of thin plate splines for arbitrary shape scene text representation. In ACM MM, 5014\u20135025."},{"key":"e_1_3_1_6_2","first-page":"1851","volume-title":"ACM MM","author":"Shu Y.","year":"2023","unstructured":"Y. Shu, W. Wang, Y. Zhou, S. Liu, A. Zhang, D. Yang, and W. Wang. 2023. Perceiving ambiguity and semantics without recognition: An efficient and effective ambiguous scene text detector. In ACM MM, 1851\u20131862."},{"key":"e_1_3_1_7_2","first-page":"10\u2009134","volume-title":"ACM MM","author":"Qiao Q.","year":"2024","unstructured":"Q. Qiao, Y. Xie, J. Gao, T. Wu, S. Huang, J. Fan, Z. Cao, Z. Wang, and Y. Zhang. 2024. Dntextspotter: Arbitrary-shaped scene text spotting via improved denoising training. In ACM MM, 10\u2009134\u201310\u2009143."},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","first-page":"5919","DOI":"10.1609\/aaai.v39i6.32632","article-title":"Arbitrary reading order scene text spotter with local semantics guidance. In","volume":"39","author":"Lyu J.","year":"2025","unstructured":"J. Lyu, W. Wang, D. Yang, J. Zhong, and Y. Zhou. 2025. Arbitrary reading order scene text spotter with local semantics guidance. In AAAI, Vol. 39, 5919\u20135927.","journal-title":"AAAI"},{"issue":"4","key":"e_1_3_1_9_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3356728","article-title":"AB-LSTM: Attention-based bidirectional LSTM model for scene text detection","volume":"15","author":"Liu Z.","year":"2019","unstructured":"Z. Liu, W. Zhou, and H. Li. 2019. AB-LSTM: Attention-based bidirectional LSTM model for scene text detection. TOMM 15, 4 (2019), 1\u201323.","journal-title":"TOMM"},{"key":"e_1_3_1_10_2","first-page":"1883","article-title":"Accurate scene text detection via scale-aware data augmentation and shape similarity constraint","volume":"24","author":"Dai P.","year":"2021","unstructured":"P. Dai, Y. Li, H. Zhang, J. Li, and X. Cao. 2021. Accurate scene text detection via scale-aware data augmentation and shape similarity constraint. IEEE TMM 24 (2021), 1883\u20131895.","journal-title":"IEEE TMM"},{"key":"e_1_3_1_11_2","first-page":"414","volume-title":"ACM MM","author":"Qin X.","year":"2021","unstructured":"X. Qin, Y. Zhou, Y. Guo, D. Wu, Z. Tian, N. Jiang, H. Wang, and W. Wang. 2021. Mask is all you need: Rethinking mask R-CNN for dense and arbitrary-shaped scene text detection. In ACM MM, 414\u2013423."},{"key":"e_1_3_1_12_2","first-page":"1","volume-title":"ICASSP","author":"Shu Y.","year":"2023","unstructured":"Y. Shu, S. Liu, Y. Zhou, H. Xu, and F. Jiang. 2023. EI \\({}^{2}\\) SR: Learning an enhanced intra-instance semantic relationship for arbitrary-shaped scene text detection. In ICASSP. IEEE, 1\u20135."},{"key":"e_1_3_1_13_2","first-page":"2025","volume-title":"ACM MM","author":"Qin X.","year":"2023","unstructured":"X. Qin, P. Lyu, C. Zhang, Y. Zhou, K. Yao, P. Zhang, H. Lin, and W. Wang. 2023. Towards robust real-time scene text detection: From semantic to instance representation learning. In ACM MM, 2025\u20132034."},{"issue":"3","key":"e_1_3_1_14_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3440087","article-title":"MFECN: Multi-level feature enhanced cumulative network for scene text detection","volume":"17","author":"Liu Z.","year":"2021","unstructured":"Z. Liu, W. Zhou, and H. Li. 2021. MFECN: Multi-level feature enhanced cumulative network for scene text detection. TOMM 17, 3 (2021), 1\u201322.","journal-title":"TOMM"},{"issue":"1","key":"e_1_3_1_15_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3231737","article-title":"Convolutional attention networks for scene text recognition","volume":"15","author":"Xie H.","year":"2019","unstructured":"H. Xie, S. Fang, Z.-J. Zha, Y. Yang, Y. Li, and Y. Zhang. 2019. Convolutional attention networks for scene text recognition. TOMM 15, 1s (2019), 1\u201317.","journal-title":"TOMM"},{"key":"e_1_3_1_16_2","first-page":"13528","volume-title":"CVPR","author":"Qiao Z.","year":"2020","unstructured":"Z. Qiao, Y. Zhou, D. Yang, Y. Zhou, and W. Wang. 2020. SEED: Semantics enhanced encoder-decoder framework for scene text recognition. In CVPR, 13528\u201313537."},{"key":"e_1_3_1_17_2","first-page":"2046","volume-title":"ACM MM","author":"Qiao Z.","year":"2021","unstructured":"Z. Qiao, Y. Zhou, J. Wei, W. Wang, Y. Zhang, N. Jiang, H. Wang, and W. Wang. 2021. PIMNet: A parallel, iterative and mimicking network for scene text recognition. In ACM MM, 2046\u20132055."},{"key":"e_1_3_1_18_2","first-page":"884","volume-title":"IJCAI","author":"Du Y.","year":"2022","unstructured":"Y. Du, Z. Chen, C. Jia, X. Yin, T. Zheng, C. Li, Y. Du, and Y.-G. Jiang. 2022. SVTR: Scene text recognition with a single visual model. In IJCAI, 884\u2013890."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3656476"},{"issue":"4","key":"e_1_3_1_20_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3633517","article-title":"Scene text recognition via dual-path network with shape-driven attention alignment","volume":"20","author":"Hu Y.","year":"2024","unstructured":"Y. Hu, B. Dong, K. Huang, L. Ding, W. Wang, X. Huang, and Q.-F. Wang. 2024. Scene text recognition via dual-path network with shape-driven attention alignment. TOMM 20, 4 (2024), 1\u201320.","journal-title":"TOMM"},{"key":"e_1_3_1_21_2","volume-title":"CVPR","author":"Zhang Y.","year":"2025","unstructured":"Y. Zhang, C. Liu, J. Wei, X. Yang, Y. Zhou, C. Ma, and X. Ji. 2025. Linguistics-aware masked image modeling for self-supervised scene text recognition. In CVPR."},{"key":"e_1_3_1_22_2","first-page":"1369","article-title":"Divide rows and conquer cells: Towards structure recognition for large tables","author":"Shen H.","year":"2023","unstructured":"H. Shen, X. Gao, J. Wei, L. Qiao, Y. Zhou, Q. Li, and Z. Cheng. 2023. Divide rows and conquer cells: Towards structure recognition for large tables. In IJCAI, 1369\u20131377.","journal-title":"IJCAI"},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","unstructured":"J. Wang L. Jin and K. Ding. 2022. Lilt: A simple yet effective language-independent layout Transformer for structured document understanding. arXiv:2202.13669. Retrieved from https:\/\/arxiv.org\/abs\/2202.13669","DOI":"10.18653\/v1\/2022.acl-long.534"},{"key":"e_1_3_1_24_2","doi-asserted-by":"crossref","first-page":"6805","DOI":"10.1609\/aaai.v39i7.32730","article-title":"LDP: Generalizing to multilingual visual information extraction by language decoupled pretraining","volume":"39","author":"Shen H.","year":"2025","unstructured":"H. Shen, G. Li, J. Zhong, and Y. Zhou. 2025. LDP: Generalizing to multilingual visual information extraction by language decoupled pretraining. In AAAI, Vol. 39, 6805\u20136813.","journal-title":"AAAI"},{"key":"e_1_3_1_25_2","first-page":"19462","volume-title":"ICCV","author":"Da C.","year":"2023","unstructured":"C. Da, C. Luo, Q. Zheng, and C. Yao. 2023. Vision grid transformer for document layout analysis. In ICCV, 19462\u201319472."},{"key":"e_1_3_1_26_2","first-page":"1631","volume-title":"ICME","author":"Yang X.","year":"2023","unstructured":"X. Yang, D. Yang, Y. Zhou, Y. Guo, and W. Wang. 2023. Mask-guided stamp erasure for real document image. In ICME. IEEE, 1631\u20131636."},{"key":"e_1_3_1_27_2","first-page":"109337","article-title":"Beyond OCR+ VQA: Towards end-to-end reading and reasoning for robust and accurate TextVQA","volume":"138","author":"Zeng G.","year":"2023","unstructured":"G. Zeng, Y. Zhang, Y. Zhou, X. Yang, N. Jiang, G. Zhao, W. Wang, and X.-C. Yin. 2023. Beyond OCR+ VQA: Towards end-to-end reading and reasoning for robust and accurate TextVQA. PR 138 (2023), 109337.","journal-title":"PR"},{"key":"e_1_3_1_28_2","first-page":"1261","volume-title":"ACM MM","author":"Zeng G.","year":"2023","unstructured":"G. Zeng, Y. Zhang, Y. Zhou, B. Fang, G. Zhao, X. Wei, and W. Wang. 2023. Filling in the blank: Rationale-augmented prompt tuning for textvqa. In ACM MM, 1261\u20131272."},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","first-page":"10275","DOI":"10.1609\/aaai.v39i10.33115","article-title":"Track the answer: Extending textvqa from image to video with spatio-temporal clues. In","volume":"39","author":"Zhang Y.","year":"2025","unstructured":"Y. Zhang, G. Zeng, H. Shen, D. Wu, Y. Zhou, and C. Ma. 2025. Track the answer: Extending textvqa from image to video with spatio-temporal clues. In AAAI, Vol. 39, 10275\u201310283.","journal-title":"AAAI"},{"key":"e_1_3_1_30_2","first-page":"138569","article-title":"Textctrl: Diffusion-based scene text editing with prior guidance control","volume":"37","author":"Zeng W.","year":"2024","unstructured":"W. Zeng, Y. Shu, Z. Li, D. Yang, and Y. Zhou. 2024. Textctrl: Diffusion-based scene text editing with prior guidance control. NeurIPS 37 (2024), 138569\u2013138594.","journal-title":"NeurIPS"},{"key":"e_1_3_1_31_2","first-page":"346","volume-title":"ECAI","author":"Li Z.","year":"2024","unstructured":"Z. Li, Y. Shu, W. Zeng, D. Yang, and Y. Zhou. 2024. First creating backgrounds then rendering texts: A new paradigm for visual text blending. In ECAI. IOS Press, 346\u2013353."},{"key":"e_1_3_1_32_2","first-page":"67","volume-title":"ECCV","author":"Lyu P.","year":"2018","unstructured":"P. Lyu, M. Liao, C. Yao, W. Wu, and X. Bai. 2018. Mask Textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes. In ECCV, 67\u201383."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2937086"},{"key":"e_1_3_1_34_2","first-page":"6","volume-title":"ECCV","author":"Liao M.","year":"2020","unstructured":"M. Liao, G. Pang, J. Huang, T. Hassner, and X. Bai. 2020. Mask Textspotter v3: Segmentation proposal network for robust scene text spotting. In ECCV. Springer, 6\u2013722."},{"key":"e_1_3_1_35_2","doi-asserted-by":"crossref","unstructured":"Y. Liu C. Shen L. Jin T. He P. Chen C. Liu and H. Chen. 2021. ABCNet v2: Adaptive bezier-curve network for real-time end-to-end text spotting. arXiv:2105.03620. Retrieved from https:\/\/arxiv.org\/abs\/2105.03620","DOI":"10.1109\/TPAMI.2021.3107437"},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","unstructured":"R. Liu N. Lu D. Chen C. Li Z. Yuan and W. Peng. 2023. PBFormer: Capturing complex scene text shape with polynomial band Transformer. arXiv:2308.15004. Retrieved from https:\/\/arxiv.org\/abs\/2308.15004","DOI":"10.1145\/3581783.3612059"},{"issue":"2022","key":"e_1_3_1_37_2","first-page":"7123","article-title":"ABINet++: Autonomous, bidirectional and iterative language modeling for scene text spotting","volume":"6","author":"Fang S.","year":"2022","unstructured":"S. Fang, Z. Mao, H. Xie, Y. Wang, C. Yan, and Y. Zhang. 2022. ABINet++: Autonomous, bidirectional and iterative language modeling for scene text spotting. IEEE TPAMI 45, 6 (2022), 7123\u20137141.","journal-title":"IEEE TPAMI"},{"key":"e_1_3_1_38_2","first-page":"4703","volume-title":"ICCV","author":"Qin S.","year":"2019","unstructured":"S. Qin, A. Bissacco, M. Raptis, Y. Fujii, and Y. Xiao. 2019. Towards unconstrained end-to-end text spotting. In ICCV, 4703\u20134713."},{"key":"e_1_3_1_39_2","first-page":"5892","volume-title":"ACM MM","author":"Wei J.","year":"2022","unstructured":"J. Wei, Y. Zhang, Y. Zhou, G. Zeng, Z. Qiao, Y. Guo, H. Wu, H. Wang, and W. Wang. 2022. Textblock: Towards scene text spotting without fine-grained detection. In ACM MM, 5892\u20135902."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i3.16348"},{"key":"e_1_3_1_41_2","first-page":"5973","volume-title":"CVPR","author":"Wang T.","year":"2021","unstructured":"T. Wang, Y. Zhu, L. Jin, D. Peng, Z. Li, M. He, Y. Wang, and C. Luo. 2021. Implicit feature alignment: Learn to convert text recognizer to text spotter. In CVPR, 5973\u20135982."},{"key":"e_1_3_1_42_2","first-page":"19348","volume-title":"CVPR","author":"Ye M.","year":"2023","unstructured":"M. Ye, J. Zhang, S. Zhao, J. Liu, T. Liu, B. Du, and D. Tao. 2023. DeepSolo: Let Transformer decoder with explicit points solo for text spotting. In CVPR, 19348\u201319357."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01786"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.322"},{"key":"e_1_3_1_45_2","volume-title":"ICLR","author":"Zhu X.","year":"2021","unstructured":"X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai. 2021. Deformable DETR: Deformable Transformers for end-to-end object detection. In ICLR."},{"key":"e_1_3_1_46_2","first-page":"9076","volume-title":"ICCV","author":"Feng W.","year":"2019","unstructured":"W. Feng, W. He, F. Yin, X.-Y. Zhang, and C.-L. Liu. 2019. Textdragon: An end-to-end framework for arbitrary shaped text spotting. In ICCV, 9076\u20139085."},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00922"},{"key":"e_1_3_1_48_2","first-page":"504","volume-title":"ECCV","author":"Baek Y.","year":"2020","unstructured":"Y. Baek, S. Shin, J. Baek, S. Park, J. Lee, D. Nam, and H. Lee. 2020. Character region attention for text spotting. In ECCV. Springer, 504\u2013521."},{"issue":"9","key":"e_1_3_1_49_2","first-page":"5349","article-title":"PAN++: Towards efficient and accurate end-to-end spotting of arbitrarily-shaped text","volume":"44","author":"Wang W.","year":"2021","unstructured":"W. Wang, E. Xie, X. Li, X. Liu, D. Liang, Z. Yang, T. Lu, and C. Shen. 2021. PAN++: Towards efficient and accurate end-to-end spotting of arbitrarily-shaped text. IEEE TPAMI 44, 9 (2021), 5349\u20135367.","journal-title":"IEEE TPAMI"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00112"},{"key":"e_1_3_1_51_2","first-page":"903","volume-title":"WACV","author":"Long S.","year":"2024","unstructured":"S. Long, S. Qin, Y. Fujii, A. Bissacco, and M. Raptis. 2024. Hierarchical text spotter for joint text spotting and layout analysis. In WACV, 903\u2013913."},{"key":"e_1_3_1_52_2","first-page":"4593","volume-title":"CVPR","author":"Huang M.","year":"2022","unstructured":"M. Huang, Y. Liu, Z. Peng, C. Liu, D. Lin, S. Zhu, N. Yuan, K. Ding, and L. Jin. 2022. SwinTextSpotter: Scene text spotting via better synergy between text detection and text recognition. In CVPR, 4593\u20134603."},{"key":"e_1_3_1_53_2","first-page":"4272","volume-title":"ACM MM","author":"Peng D.","year":"2022","unstructured":"D. Peng, X. Wang, Y. Liu, J. Zhang, M. Huang, S. Lai, J. Li, S. Zhu, D. Lin, C. Shen, et al. 2022. SPTS: Single-point text spotting. In ACM MM, 4272\u20134281."},{"key":"e_1_3_1_54_2","first-page":"15665","volume-title":"IEEE TPAMI","author":"Liu Y.","year":"2023","unstructured":"Y. Liu, J. Zhang, D. Peng, M. Huang, X. Wang, J. Tang, C. Huang, D. Lin, C. Shen, X. Bai, et al. 2023. SPTS v2: Single-point scene text spotting. IEEE TPAMI 45, 12 (2023), 15665\u201315679."},{"key":"e_1_3_1_55_2","first-page":"5238","volume-title":"ICCV","author":"Li H.","year":"2017","unstructured":"H. Li, P. Wang, and C. Shen. 2017. Towards end-to-end text spotting with convolutional recurrent neural networks. In ICCV, 5238\u20135246."},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2646371"},{"key":"e_1_3_1_57_2","first-page":"5020","volume-title":"CVPR","author":"He T.","year":"2018","unstructured":"T. He, Z. Tian, W. Huang, C. Shen, Y. Qiao, and C. Sun. 2018. An end-to-end Textspotter with explicit alignment and attention. In CVPR, 5020\u20135029."},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00702"},{"key":"e_1_3_1_59_2","first-page":"4154","volume-title":"ACM MM","author":"Tang J.","year":"2022","unstructured":"J. Tang, S. Qiao, B. Cui, Y. Ma, S. Zhang, and D. Kanoulas. 2022. You can even annotate text with voice: Transcription-only-supervised text spotting. In ACM MM, 4154\u20134163."},{"key":"e_1_3_1_60_2","article-title":"Deep contextualized word representations","volume":"5","author":"Peters M. E.","year":"2018","unstructured":"M. E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, Luke Zettlemoyer. 2018. Deep contextualized word representations. In NAACL, vol. 5.","journal-title":"NAACL"},{"key":"e_1_3_1_61_2","first-page":"4171","volume-title":"NAACL","author":"Devlin J.","year":"2019","unstructured":"J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. 2019. BERT: Pre-training of deep bidirectional Transformers for language understanding. In NAACL, 4171\u20134186."},{"key":"e_1_3_1_62_2","unstructured":"A. Radford K. Narasimhan T. Salimans I. Sutskever. 2018. Improving language understanding by generative pre-training. OpenAI Blog. Retrieved from https:\/\/openai.com\/index\/language-unsupervised"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.5555\/3455716.3455856"},{"key":"e_1_3_1_64_2","first-page":"15558","article-title":"K-lite: Learning transferable visual models with external knowledge","volume":"35","author":"Shen S.","year":"2022","unstructured":"S. Shen, C. Li, X. Hu, Y. Xie, J. Yang, P. Zhang, Z. Gan, L. Wang, L. Yuan, C. Liu, et al. 2022. K-lite: Learning transferable visual models with external knowledge. NeurIPS 35 (2022), 15558\u201315573.","journal-title":"NeurIPS"},{"key":"e_1_3_1_65_2","first-page":"18786","volume-title":"ICCV","author":"Ma W.","year":"2023","unstructured":"W. Ma, S. Li, J. Zhang, C. H. Liu, J. Kang, Y. Wang, and G. Huang. 2023. Borrowing knowledge from pre-trained language model: A new data-efficient visual learning paradigm. In ICCV, 18786\u201318797."},{"key":"e_1_3_1_66_2","first-page":"8025","volume-title":"WACV","author":"Fujitake M.","year":"2024","unstructured":"M. Fujitake. 2024. DTrOCR: Decoder-only transformer for optical character recognition. In WACV, 8025\u20138035."},{"key":"e_1_3_1_67_2","unstructured":"W. X. Zhao K. Zhou J. Li T. Tang X. Wang Y. Hou Y. Min B. Zhang J. Zhang Z. Dong et al. 2023. A survey of large language models. arXiv:2303.18223. Retrieved from https:\/\/arxiv.org\/abs\/2303.18223"},{"key":"e_1_3_1_68_2","first-page":"284","volume-title":"ECCV","author":"Xue C.","year":"2022","unstructured":"C. Xue, W. Zhang, Y. Hao, S. Lu, P. H. Torr, and S. Bai. 2022. Language matters: A weakly supervised vision-language pre-training approach for scene text detection and spotting. In ECCV. Springer, 284\u2013302."},{"issue":"34","key":"e_1_3_1_69_2","first-page":"226","article-title":"A density-based algorithm for discovering clusters in large spatial databases with noise","volume":"96","author":"Ester M.","year":"1996","unstructured":"M. Ester, H.-P. Kriegel, J. Sander, X. Xu. 1996. A density-based algorithm for discovering clusters in large spatial databases with noise. In KDD 96, 34 (1996), 226\u2013231.","journal-title":"KDD"},{"key":"e_1_3_1_70_2","article-title":"Unified language model pre-training for natural language understanding and generation","volume":"32","author":"Dong L.","year":"2019","unstructured":"L. Dong, N. Yang, W. Wang, F. Wei, X. Liu, Y. Wang, J. Gao, M. Zhou, and H.-W. Hon. 2019. Unified language model pre-training for natural language understanding and generation. NeurIPS 32 (2019).","journal-title":"NeurIPS"},{"key":"e_1_3_1_71_2","first-page":"1156","volume-title":"ICDAR","author":"Karatzas D.","year":"2015","unstructured":"D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. Ghosh, A. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V. R. Chandrasekhar, S. Lu. et al. 2015. ICDAR 2015 competition on robust reading. In ICDAR. IEEE, 1156\u20131160."},{"key":"e_1_3_1_72_2","first-page":"935","volume-title":"ICDAR","volume":"1","author":"Ch\u2019ng C. K.","year":"2017","unstructured":"C. K. Ch\u2019ng and C. S. Chan. 2017. Total-text: A comprehensive dataset for scene text detection and recognition. In ICDAR, Vol. 1, IEEE, 935\u2013942."},{"key":"e_1_3_1_73_2","unstructured":"L. Yuliang J. Lianwen Z. Shuaitao and Z. Sheng 2017. Detecting curve text in the wild: New dataset and new solution. arXiv:1712.02170. Retrieved from https:\/\/arxiv.org\/abs\/1712.02170"},{"key":"e_1_3_1_74_2","first-page":"2315","volume-title":"CVPR","author":"Gupta A.","year":"2016","unstructured":"A. Gupta, A. Vedaldi, and A. Zisserman. 2016. Synthetic data for text localization in natural images. In CVPR, 2315\u20132324."},{"key":"e_1_3_1_75_2","first-page":"20543","volume-title":"ICCV","author":"Jiang Q.","year":"2023","unstructured":"Q. Jiang, J. Wang, D. Peng, C. Liu, and L. Jin. 2023. Revisiting scene text recognition: A data perspective. In ICCV, 20543\u201320554."},{"key":"e_1_3_1_76_2","first-page":"109","volume-title":"ICDAR","author":"Yim M.","year":"2021","unstructured":"M. Yim, Y. Kim, H.-C. Cho, and S. Park. 2021. Synthtiger: Synthetic text image generator towards better text recognition models. In ICDAR. Springer, 109\u2013124."},{"key":"e_1_3_1_77_2","first-page":"3791","volume-title":"ACM MM","author":"Kuang Z.","year":"2021","unstructured":"Z. Kuang, H. Sun, Z. Li, X. Yue, T. H. Lin, J. Chen, H. Wei, Y. Zhu, T. Gao, W. Zhang, et al. 2021. MMOCR: A comprehensive toolbox for text detection, recognition and understanding. In ACM MM, 3791\u20133794."},{"key":"e_1_3_1_78_2","first-page":"95","volume-title":"ICDAR","author":"Hao J.","year":"2021","unstructured":"J. Hao, Y. Wen, J. Deng, J. Gan, S. Ren, H. Tan, and X. Chen. 2021. Eem: An end-to-end evaluation metric for scene text detection and recognition. In ICDAR. Springer, 95\u2013108."},{"issue":"11","key":"e_1_3_1_79_2","doi-asserted-by":"crossref","first-page":"13094","DOI":"10.1609\/aaai.v37i11.26538","article-title":"TrOCR: Transformer-based optical character recognition with pre-trained models","volume":"37","author":"Li M.","year":"2023","unstructured":"M. Li, T. Lv, J. Chen, L. Cui, Y. Lu, D. Florencio, C. Zhang, Z. Li, and F. Wei. 2023. TrOCR: Transformer-based optical character recognition with pre-trained models. In AAAI 37, 11 (2023), 13094\u201313102.","journal-title":"AAAI"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19815-1_11"},{"key":"e_1_3_1_81_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018610"},{"issue":"07","key":"e_1_3_1_82_2","doi-asserted-by":"crossref","first-page":"12216","DOI":"10.1609\/aaai.v34i07.6903","article-title":"Decoupled attention network for text recognition","volume":"34","author":"Wang T.","year":"2020","unstructured":"T. Wang, Y. Zhu, L. Jin, C. Luo, X. Chen, Y. Wu, Q. Wang, and M. Cai. 2020. Decoupled attention network for text recognition. In AAAI 34, 07 (2020), 12216\u201312224.","journal-title":"AAAI"},{"key":"e_1_3_1_83_2","first-page":"12113","article-title":"Towards accurate scene text recognition with semantic reasoning networks","author":"Yu D.","year":"2020","unstructured":"D. Yu, X. Li, C. Zhang, T. Liu, J. Han, J. Liu, and E. Ding. 2020. Towards accurate scene text recognition with semantic reasoning networks. In CVPR, 12113\u201312122.","journal-title":"CVPR"},{"key":"e_1_3_1_84_2","first-page":"14194","volume-title":"ICCV","author":"Wang Y.","year":"2021","unstructured":"Y. Wang, H. Xie, S. Fang, J. Wang, S. Zhu, and Y. Zhang. 2021. From two to one: A new scene text recognizer with visual language modeling network. In ICCV, 14194\u201314203."},{"key":"e_1_3_1_85_2","first-page":"446","volume-title":"ECCV","author":"Na B.","year":"2022","unstructured":"B. Na, Y. Kim, and S. Park. 2022. Multi-modal text recognition networks: Interactive enhancements between visual and semantic features. In ECCV. Springer, 446\u2013463."},{"key":"e_1_3_1_86_2","unstructured":"R. OpenAI. 2023. GPT-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3734872","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,8]],"date-time":"2025-07-08T17:08:24Z","timestamp":1751994504000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3734872"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,30]]},"references-count":85,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,6,30]]}},"alternative-id":["10.1145\/3734872"],"URL":"https:\/\/doi.org\/10.1145\/3734872","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,30]]},"assertion":[{"value":"2024-10-12","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}