{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T23:02:23Z","timestamp":1776207743656,"version":"3.50.1"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"11","license":[{"start":{"date-parts":[[2023,11,20]],"date-time":"2023-11-20T00:00:00Z","timestamp":1700438400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100003399","name":"Shanghai Municipal Science and Technology Commission","doi-asserted-by":"crossref","award":["12511505303"],"award-info":[{"award-number":["12511505303"]}],"id":[{"id":"10.13039\/501100003399","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Shanghai Archives Research Program","award":["2108"],"award-info":[{"award-number":["2108"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,11,30]]},"abstract":"<jats:p>Reading scene text in the natural image is of fundamental importance in many real-world problems. Text recognition has a profound effect on information processing by enabling automated extraction and interpretation. Recent scene text recognition methods employ the encoder-decoder framework, which constructs the encoder by obtaining the visual representations based on the last layer of the backbone network and then feeding them into a sequence model. In this article, we propose a novel encoder structure that performs the feature extractor and the sequence modeling within a unified framework. The introduced Aggregated Temporal Convolutional Encoder (ATCE) first incorporates the temporal convolutional layers to consider the long-term temporal relationship in the encoder stage. The aggregation of these temporal convolution modules is designed to utilize visual features from different levels, by augmenting the standard architecture with deeper aggregation to better fuse information across modules. We also study the impact of different attention modules in convolutional blocks for learning accurate text representations. We conduct comparisons on several scene text recognition benchmarks for both Chinese and English; the experiments demonstrate the complementary ability with different decoder variants and the effectiveness of our proposed approach.<\/jats:p>","DOI":"10.1145\/3625822","type":"journal-article","created":{"date-parts":[[2023,10,12]],"date-time":"2023-10-12T14:53:42Z","timestamp":1697122422000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Reading Scene Text with Aggregated Temporal Convolutional Encoder"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4079-7897","authenticated-orcid":false,"given":"Tianlong","family":"Ma","sequence":"first","affiliation":[{"name":"School of Computer Science, Fudan University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4268-6114","authenticated-orcid":false,"given":"Xiangcheng","family":"Du","sequence":"additional","affiliation":[{"name":"School of Computer Science, Fudan University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9146-051X","authenticated-orcid":false,"given":"Xingjiao","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Computer Science, Fudan University &amp; Technology Innovation Center of Digital Creation and Applications of Chinese Calligraphy and Painting, Ministry of Culture and Tourism, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7069-6275","authenticated-orcid":false,"given":"Zhao","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Computer Science, Fudan University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5590-9292","authenticated-orcid":false,"given":"Yingbin","family":"Zheng","sequence":"additional","affiliation":[{"name":"Videt Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3063-1957","authenticated-orcid":false,"given":"Cheng","family":"Jin","sequence":"additional","affiliation":[{"name":"School of Computer Science, Fudan University &amp; Technology Innovation Center of Digital Creation and Applications of Chinese Calligraphy and Painting, Ministry of Culture and Tourism, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,11,20]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/2812809"},{"key":"e_1_3_2_3_2","first-page":"4715","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Baek Jeonghun","year":"2019","unstructured":"Jeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park, Dongyoon Han, Sangdoo Yun, Seong Joon Oh, and Hwalsuk Lee. 2019. What is wrong with scene text recognition model comparisons? Dataset and model analysis. In Proceedings of the International Conference on Computer Vision. 4715\u20134723."},{"key":"e_1_3_2_4_2","doi-asserted-by":"crossref","DOI":"10.1145\/3594631","article-title":"Low-resource multilingual neural translation using linguistic feature based relevance mechanisms","author":"Chakrabarty Abhisek","year":"2023","unstructured":"Abhisek Chakrabarty, Raj Dabre, Chenchen Ding, Masao Utiyama, and Eiichiro Sumita. 2023. Low-resource multilingual neural translation using linguistic feature based relevance mechanisms. ACM Transactions on Asian and Low-Resource Language Information Processing.","journal-title":"ACM Transactions on Asian and Low-Resource Language Information Processing."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3511807"},{"key":"e_1_3_2_6_2","first-page":"12026","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Chen Jingye","year":"2021","unstructured":"Jingye Chen, Bin Li, and Xiangyang Xue. 2021. Scene text telescope: Text-focused scene image super-resolution. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 12026\u201312035."},{"key":"e_1_3_2_7_2","article-title":"Benchmarking Chinese text recognition: Datasets, baselines, and an empirical study","author":"Chen Jingye","year":"2021","unstructured":"Jingye Chen, Haiyang Yu, Jianqi Ma, Mengnan Guan, Xixi Xu, Xiaocong Wang, Shaobo Qu, Bin Li, and Xiangyang Xue. 2021. Benchmarking Chinese text recognition: Datasets, baselines, and an empirical study. arXiv preprint arXiv:2112.15093 (2021).","journal-title":"arXiv preprint arXiv:2112.15093"},{"key":"e_1_3_2_8_2","first-page":"5076","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Cheng Zhanzhan","year":"2017","unstructured":"Zhanzhan Cheng, Fan Bai, Yunlu Xu, Gang Zheng, Shiliang Pu, and Shuigeng Zhou. 2017. Focusing attention: Towards accurate text recognition in natural images. In Proceedings of the International Conference on Computer Vision. 5076\u20135084."},{"key":"e_1_3_2_9_2","first-page":"5571","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Cheng Zhanzhan","year":"2018","unstructured":"Zhanzhan Cheng, Yangliu Xu, Fan Bai, Yi Niu, Shiliang Pu, and Shuigeng Zhou. 2018. AON: Towards arbitrarily-oriented text recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5571\u20135579."},{"key":"e_1_3_2_10_2","first-page":"2383","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing","author":"Du Xiangcheng","year":"2020","unstructured":"Xiangcheng Du, Tianlong Ma, Yingbin Zheng, Hao Ye, Xingjiao Wu, and Liang He. 2020. Scene text recognition with temporal convolutional encoder. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing. 2383\u20132387."},{"key":"e_1_3_2_11_2","article-title":"Study on machine translation teaching model based on translation parallel corpus and exploitation for multimedia Asian information processing","author":"Gong Yan","year":"2022","unstructured":"Yan Gong. 2022. Study on machine translation teaching model based on translation parallel corpus and exploitation for multimedia Asian information processing. ACM Transactions on Asian and Low-Resource Language Information Processing. Published November 7, 2022.","journal-title":"ACM Transactions on Asian and Low-Resource Language Information Processing."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143891"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.254"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_15_2","article-title":"Reading scene text in deep convolutional sequences","author":"He Pan","year":"2015","unstructured":"Pan He, Weilin Huang, Yu Qiao, Chen Change Loy, and Xiaoou Tang. 2015. Reading scene text in deep convolutional sequences. arXiv:1506.04395 (2015).","journal-title":"arXiv:1506.04395"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_3_2_17_2","article-title":"Synthetic data and artificial neural networks for natural scene text recognition","author":"Jaderberg Max","year":"2014","unstructured":"Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014. Synthetic data and artificial neural networks for natural scene text recognition. arXiv:1406.2227 (2014).","journal-title":"arXiv:1406.2227"},{"key":"e_1_3_2_18_2","first-page":"2017","volume-title":"Advances in Neural Information Processing Systems","author":"Jaderberg Max","year":"2015","unstructured":"Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu. 2015. Spatial transformer networks. In Advances in Neural Information Processing Systems. 2017\u20132025."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/3162077"},{"key":"e_1_3_2_20_2","first-page":"1156","volume-title":"Proceedings of the International Conference on Document Analysis and Recognition","author":"Karatzas Dimosthenis","year":"2015","unstructured":"Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, Faisal Shafait, Seiichi Uchida, and Ernest Valveny. 2015. ICDAR 2015 competition on Robust Reading. In Proceedings of the International Conference on Document Analysis and Recognition. 1156\u20131160."},{"key":"e_1_3_2_21_2","first-page":"1484","volume-title":"Proceedings of the International Conference on Document Analysis and Recognition","author":"Karatzas Dimosthenis","year":"2013","unstructured":"Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernandez Mota, Jon Almazan Almazan, and Lluis Pere De Las Heras. 2013. ICDAR 2013 Robust Reading competition. In Proceedings of the International Conference on Document Analysis and Recognition. 1484\u20131493."},{"key":"e_1_3_2_22_2","article-title":"Semantic and context understanding for sentiment analysis in Hindi handwritten character recognition using a multiresolution technique","author":"Kumar Ankit","year":"2022","unstructured":"Ankit Kumar, Surbhi Bhatiya, Mohammad R. Khosravi, Arwa Mashat, and Parul Agarwal. 2022. Semantic and context understanding for sentiment analysis in Hindi handwritten character recognition using a multiresolution technique. ACM Transactions on Asian and Low-Resource Language Information Processing. Published October 6, 2022.","journal-title":"ACM Transactions on Asian and Low-Resource Language Information Processing."},{"key":"e_1_3_2_23_2","first-page":"47","volume-title":"Proceedings of the ECCV Workshops","author":"Lea Colin","year":"2016","unstructured":"Colin Lea, Rene Vidal, Austin Reiter, and Gregory D. Hager. 2016. Temporal convolutional networks: A unified approach to action segmentation. In Proceedings of the ECCV Workshops. 47\u201354."},{"key":"e_1_3_2_24_2","first-page":"2231","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Lee Chen-Yu","year":"2016","unstructured":"Chen-Yu Lee and Simon Osindero. 2016. Recursive recurrent nets with attention modeling for OCR in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2231\u20132239."},{"key":"e_1_3_2_25_2","first-page":"8610","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Li Hui","year":"2019","unstructured":"Hui Li, Peng Wang, Chunhua Shen, and Guyu Zhang. 2019. Show, attend and read: A simple and strong baseline for irregular text recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 8610\u20138617."},{"key":"e_1_3_2_26_2","first-page":"8714","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Liao Minghui","year":"2019","unstructured":"Minghui Liao, Jian Zhang, Zhaoyi Wan, Fengming Xie, Jiajun Liang, Pengyuan Lyu, Cong Yao, and Xiang Bai. 2019. Scene text recognition from two-dimensional perspective. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 8714\u20138721."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.106"},{"key":"e_1_3_2_28_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Liu Zichuan","year":"2018","unstructured":"Zichuan Liu, Yixing Li, Fengbo Ren, Wang Ling Goh, and Hao Yu. 2018. SqueezedText: A real-time scene text recognition by binary convolutional encoder-decoder network. In Proceedings of the AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2019.01.020"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2818020"},{"key":"e_1_3_2_32_2","first-page":"4187","volume-title":"Proceedings of the IEEE International Conference on Big Data","author":"Ma Tianlong","year":"2022","unstructured":"Tianlong Ma, Xiangcheng Du, Yanlong Wang, and Xiutao Cui. 2022. Scene text recognition with heuristic local attention. In Proceedings of the IEEE International Conference on Big Data. 4187\u20134194."},{"key":"e_1_3_2_33_2","volume-title":"Proceedings of the British Machine Vision Conference","author":"Mishra Anand","year":"2012","unstructured":"Anand Mishra, Karteek Alahari, and C. V. Jawahar. 2012. Scene text recognition using higher order language priors. In Proceedings of the British Machine Vision Conference."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3450273"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01354"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2646371"},{"key":"e_1_3_2_37_2","first-page":"4168","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Shi Baoguang","year":"2016","unstructured":"Baoguang Shi, Xinggang Wang, Pengyuan Lyu, Cong Yao, and Xiang Bai. 2016. Robust scene text recognition with automatic rectification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4168\u20134176."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2848939"},{"key":"e_1_3_2_39_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_2_40_2","first-page":"5998","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.04.129"},{"key":"e_1_3_2_42_2","first-page":"335","volume-title":"Advances in Neural Information Processing Systems","author":"Wang Jianfeng","year":"2017","unstructured":"Jianfeng Wang and Xiaolin Hu. 2017. Gated recurrent convolution neural network for OCR. In Advances in Neural Information Processing Systems. 335\u2013344."},{"key":"e_1_3_2_43_2","first-page":"1457","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Wang Kai","year":"2011","unstructured":"Kai Wang, Boris Babenko, and Serge Belongie. 2011. End-to-end scene text recognition. In Proceedings of the International Conference on Computer Vision. 1457\u20131464."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6903"},{"key":"e_1_3_2_45_2","first-page":"844","volume-title":"Proceedings of the International Conference on Document Analysis and Recognition","author":"Wojna Zbigniew","year":"2017","unstructured":"Zbigniew Wojna, Alexander N. Gorban, Dar-Shyang Lee, Kevin Murphy, Qian Yu, Yeqing Li, and Julian Ibarz. 2017. Attention-based extraction of structured information from street view imagery. In Proceedings of the International Conference on Document Analysis and Recognition. 844\u2013850."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.04.071"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2021.07.020"},{"key":"e_1_3_2_49_2","first-page":"6538","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Xie Zecheng","year":"2019","unstructured":"Zecheng Xie, Yaoxiong Huang, Yuanzhi Zhu, Lianwen Jin, Yuliang Liu, and Lele Xie. 2019. Aggregation cross-entropy for sequence recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 6538\u20136547."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.07.010"},{"key":"e_1_3_2_51_2","first-page":"9147","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Yang Mingkun","year":"2019","unstructured":"Mingkun Yang, Yushuo Guan, Minghui Liao, Xin He, Kaigui Bian, Song Bai, Cong Yao, and Xiang Bai. 2019. Symmetry-constrained rectification network for scene text recognition. In Proceedings of the International Conference on Computer Vision. 9147\u20139156."},{"key":"e_1_3_2_52_2","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence","author":"Yang Xiao","year":"2017","unstructured":"Xiao Yang, Dafang He, Zihan Zhou, Daniel Kifer, and C. Lee Giles. 2017. Learning to read irregular text with attention mechanisms. In Proceedings of the International Joint Conference on Artificial Intelligence."},{"key":"e_1_3_2_53_2","article-title":"ADADELTA: An adaptive learning rate method","author":"Zeiler Matthew D.","year":"2012","unstructured":"Matthew D. Zeiler. 2012. ADADELTA: An adaptive learning rate method. arXiv preprint arXiv:1212.5701 (2012).","journal-title":"arXiv preprint arXiv:1212.5701"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2020.107559"},{"key":"e_1_3_2_55_2","first-page":"751","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Zhang Hui","year":"2020","unstructured":"Hui Zhang, Quanming Yao, Mingkun Yang, Yongchao Xu, and Xiang Bai. 2020. AutoSTR: Efficient backbone search for scene text recognition. In Proceedings of the European Conference on Computer Vision. 751\u2013767."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3625822","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3625822","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:12Z","timestamp":1750182552000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3625822"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,20]]},"references-count":54,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2023,11,30]]}},"alternative-id":["10.1145\/3625822"],"URL":"https:\/\/doi.org\/10.1145\/3625822","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,20]]},"assertion":[{"value":"2023-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-20","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}