{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,8]],"date-time":"2025-09-08T06:34:55Z","timestamp":1757313295917,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":43,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T00:00:00Z","timestamp":1602460800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Science Foundation","award":["1838193"],"award-info":[{"award-number":["1838193"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,12]]},"DOI":"10.1145\/3394171.3413913","type":"proceedings-article","created":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T12:26:25Z","timestamp":1602505585000},"page":"4051-4059","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["LAL: Linguistically Aware Learning for Scene Text Recognition"],"prefix":"10.1145","author":[{"given":"Yi","family":"Zheng","sequence":"first","affiliation":[{"name":"Boston University, Boston, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenda","family":"Qin","sequence":"additional","affiliation":[{"name":"Boston University, Boston, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Derry","family":"Wijaya","sequence":"additional","affiliation":[{"name":"Boston University, Boston, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Margrit","family":"Betke","sequence":"additional","affiliation":[{"name":"Boston University, Boston, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,10,12]]},"reference":[{"volume-title":"Character Region Awareness for Text Detection","author":"Baek Youngmin","key":"e_1_3_2_2_1_1","unstructured":"Youngmin Baek , Bado Lee , Dongyoon Han , Sangdoo Yun , and Hwalsuk Lee . 2019. Character Region Awareness for Text Detection . In CVPR. Computer Vision Foundation \/ IEEE , 9365--9374. Youngmin Baek, Bado Lee, Dongyoon Han, Sangdoo Yun, and Hwalsuk Lee. 2019. Character Region Awareness for Text Detection. In CVPR. Computer Vision Foundation \/ IEEE, 9365--9374."},{"key":"e_1_3_2_2_2_1","unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural Machine Translation by Jointly Learning to Align and Translate. http:\/\/arxiv.org\/abs\/1409.0473 cite arxiv:1409.0473Comment: Accepted at ICLR 2015 as oral presentation.  Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural Machine Translation by Jointly Learning to Align and Translate. http:\/\/arxiv.org\/abs\/1409.0473 cite arxiv:1409.0473Comment: Accepted at ICLR 2015 as oral presentation."},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.24792"},{"key":"e_1_3_2_2_4_1","volume-title":"Chi-squared and Fisher-Irwin tests of two-by-two tables with small sample recommendations. Statistics in medicine","author":"Campbell Ian","year":"2007","unstructured":"Ian Campbell . 2007. Chi-squared and Fisher-Irwin tests of two-by-two tables with small sample recommendations. Statistics in medicine , Vol. 26 19 ( 2007 ), 3661--75. Ian Campbell. 2007. Chi-squared and Fisher-Irwin tests of two-by-two tables with small sample recommendations. Statistics in medicine, Vol. 26 19 (2007), 3661--75."},{"key":"e_1_3_2_2_5_1","volume-title":"Focusing Attention: Towards Accurate Text Recognition in Natural Images. CoRR","author":"Cheng Zhanzhan","year":"2017","unstructured":"Zhanzhan Cheng , Fan Bai , Yunlu Xu , Gang Zheng , Shiliang Pu , and Shuigeng Zhou . 2017 a. Focusing Attention: Towards Accurate Text Recognition in Natural Images. CoRR , Vol. abs\/ 1709 .02054 (2017). arxiv: 1709.02054 http:\/\/arxiv.org\/abs\/1709.02054 Zhanzhan Cheng, Fan Bai, Yunlu Xu, Gang Zheng, Shiliang Pu, and Shuigeng Zhou. 2017a. Focusing Attention: Towards Accurate Text Recognition in Natural Images. CoRR, Vol. abs\/1709.02054 (2017). arxiv: 1709.02054 http:\/\/arxiv.org\/abs\/1709.02054"},{"key":"e_1_3_2_2_6_1","volume-title":"Arbitrarily-Oriented Text Recognition. CoRR","author":"Cheng Zhanzhan","year":"2017","unstructured":"Zhanzhan Cheng , Xuyang Liu , Fan Bai , Yi Niu , Shiliang Pu , and Shuigeng Zhou . 2017b. Arbitrarily-Oriented Text Recognition. CoRR , Vol. abs\/ 1711 .04226 ( 2017 ). arxiv: 1711.04226 http:\/\/arxiv.org\/abs\/1711.04226 Zhanzhan Cheng, Xuyang Liu, Fan Bai, Yi Niu, Shiliang Pu, and Shuigeng Zhou. 2017b. Arbitrarily-Oriented Text Recognition. CoRR, Vol. abs\/1711.04226 (2017). arxiv: 1711.04226 http:\/\/arxiv.org\/abs\/1711.04226"},{"key":"e_1_3_2_2_7_1","volume-title":"Attention-Based Models for Speech Recognition. CoRR","author":"Chorowski Jan","year":"2015","unstructured":"Jan Chorowski , Dzmitry Bahdanau , Dmitriy Serdyuk , KyungHyun Cho , and Yoshua Bengio . 2015. Attention-Based Models for Speech Recognition. CoRR , Vol. abs\/ 1506 .07503 ( 2015 ). arxiv: 1506.07503 http:\/\/arxiv.org\/abs\/1506.07503 Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, KyungHyun Cho, and Yoshua Bengio. 2015. Attention-Based Models for Speech Recognition. CoRR, Vol. abs\/1506.07503 (2015). arxiv: 1506.07503 http:\/\/arxiv.org\/abs\/1506.07503"},{"key":"e_1_3_2_2_8_1","volume-title":"KyungHyun Cho, and Yoshua Bengio.","author":"Chung Junyoung","year":"2014","unstructured":"Junyoung Chung , cC aglar G\u00fc lcc ehre , KyungHyun Cho, and Yoshua Bengio. 2014 . Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. CoRR , Vol. abs\/ 1412 .3555 (2014). arxiv: 1412.3555 http:\/\/arxiv.org\/abs\/1412.3555 Junyoung Chung, cC aglar G\u00fc lcc ehre, KyungHyun Cho, and Yoshua Bengio. 2014. Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling. CoRR, Vol. abs\/1412.3555 (2014). arxiv: 1412.3555 http:\/\/arxiv.org\/abs\/1412.3555"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1080\/14786440009463897"},{"key":"e_1_3_2_2_10_1","volume-title":"Synthetic Data for Text Localisation in Natural Images. CoRR","author":"Gupta Ankush","year":"2016","unstructured":"Ankush Gupta , Andrea Vedaldi , and Andrew Zisserman . 2016. Synthetic Data for Text Localisation in Natural Images. CoRR , Vol. abs\/ 1604 .06646 ( 2016 ). arxiv: 1604.06646 http:\/\/arxiv.org\/abs\/1604.06646 Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman. 2016. Synthetic Data for Text Localisation in Natural Images. CoRR, Vol. abs\/1604.06646 (2016). arxiv: 1604.06646 http:\/\/arxiv.org\/abs\/1604.06646"},{"key":"e_1_3_2_2_11_1","volume-title":"Deep Residual Learning for Image Recognition. CoRR","author":"He Kaiming","year":"2015","unstructured":"Kaiming He , Xiangyu Zhang , Shaoqing Ren , and Jian Sun . 2015. Deep Residual Learning for Image Recognition. CoRR , Vol. abs\/ 1512 .03385 ( 2015 ). arxiv: 1512.03385 http:\/\/arxiv.org\/abs\/1512.03385 Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Deep Residual Learning for Image Recognition. CoRR, Vol. abs\/1512.03385 (2015). arxiv: 1512.03385 http:\/\/arxiv.org\/abs\/1512.03385"},{"key":"e_1_3_2_2_12_1","volume-title":"Long short-term memory. Neural computation","author":"Hochreiter Sepp","year":"1997","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber . 1997. Long short-term memory. Neural computation , Vol. 9 , 8 ( 1997 ), 1735--1780. Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long short-term memory. Neural computation, Vol. 9, 8 (1997), 1735--1780."},{"key":"e_1_3_2_2_13_1","volume-title":"Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition. CoRR","author":"Jaderberg Max","year":"2014","unstructured":"Max Jaderberg , Karen Simonyan , Andrea Vedaldi , and Andrew Zisserman . 2014. Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition. CoRR , Vol. abs\/ 1406 .2227 ( 2014 ), 10. http:\/\/arxiv.org\/abs\/1406.2227 Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014. Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition. CoRR, Vol. abs\/1406.2227 (2014), 10. http:\/\/arxiv.org\/abs\/1406.2227"},{"key":"e_1_3_2_2_14_1","volume-title":"Spatial Transformer Networks. CoRR","author":"Jaderberg Max","year":"2025","unstructured":"Max Jaderberg , Karen Simonyan , Andrew Zisserman , and Koray Kavukcuoglu . 2015. Spatial Transformer Networks. CoRR , Vol. abs\/ 1506 .0 2025 (2015). arxiv: 1506.02025 http:\/\/arxiv.org\/abs\/1506.02025 Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu. 2015. Spatial Transformer Networks. CoRR, Vol. abs\/1506.02025 (2015). arxiv: 1506.02025 http:\/\/arxiv.org\/abs\/1506.02025"},{"key":"e_1_3_2_2_15_1","volume-title":"ICDAR 2015 competition on Robust Reading. In ICDAR. IEEE Computer Society, 1156--1160","author":"Karatzas Dimosthenis","year":"2015","unstructured":"Dimosthenis Karatzas , Lluis Gomez-Bigorda , Anguelos Nicolaou , Suman K. Ghosh , Andrew D. Bagdanov , Masakazu Iwamura , Jiri Matas , Lukas Neumann , Vijay Ramaseshan Chandrasekhar , Shijian Lu , Faisal Shafait , Seiichi Uchida , and Ernest Valveny . 2015 . ICDAR 2015 competition on Robust Reading. In ICDAR. IEEE Computer Society, 1156--1160 . Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman K. Ghosh, Andrew D. Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, Faisal Shafait, Seiichi Uchida, and Ernest Valveny. 2015. ICDAR 2015 competition on Robust Reading. In ICDAR. IEEE Computer Society, 1156--1160."},{"volume-title":"ICDAR 2013 Robust Reading Competition. In ICDAR. IEEE Computer Society, 1484--1493","author":"Karatzas Dimosthenis","key":"e_1_3_2_2_16_1","unstructured":"Dimosthenis Karatzas , Faisal Shafait , Seiichi Uchida , Masakazu Iwamura , Lluis Gomez i Bigorda , Sergi Robles Mestre , Joan Mas , David Fern\u00e1 ndez Mota , Jon Almaz\u00e1 n, and Llu'i s-Pere de las Heras. 2013 . ICDAR 2013 Robust Reading Competition. In ICDAR. IEEE Computer Society, 1484--1493 . Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fern\u00e1 ndez Mota, Jon Almaz\u00e1 n, and Llu'i s-Pere de las Heras. 2013. ICDAR 2013 Robust Reading Competition. In ICDAR. IEEE Computer Society, 1484--1493."},{"key":"e_1_3_2_2_17_1","volume-title":"Recursive Recurrent Nets with Attention Modeling for OCR in the Wild. CoRR","author":"Lee Chen-Yu","year":"2016","unstructured":"Chen-Yu Lee and Simon Osindero . 2016. Recursive Recurrent Nets with Attention Modeling for OCR in the Wild. CoRR , Vol. abs\/ 1603 .03101 ( 2016 ). arxiv: 1603.03101 http:\/\/arxiv.org\/abs\/1603.03101 Chen-Yu Lee and Simon Osindero. 2016. Recursive Recurrent Nets with Attention Modeling for OCR in the Wild. CoRR, Vol. abs\/1603.03101 (2016). arxiv: 1603.03101 http:\/\/arxiv.org\/abs\/1603.03101"},{"key":"e_1_3_2_2_18_1","volume-title":"Attend and Read: A Simple and Strong Baseline for Irregular Text Recognition. CoRR","author":"Li Hui","year":"2018","unstructured":"Hui Li , Peng Wang , Chunhua Shen , and Guyu Zhang . 2018. Show , Attend and Read: A Simple and Strong Baseline for Irregular Text Recognition. CoRR , Vol. abs\/ 1811 .00751 ( 2018 ). arxiv: 1811.00751 http:\/\/arxiv.org\/abs\/1811.00751 Hui Li, Peng Wang, Chunhua Shen, and Guyu Zhang. 2018. Show, Attend and Read: A Simple and Strong Baseline for Irregular Text Recognition. CoRR, Vol. abs\/1811.00751 (2018). arxiv: 1811.00751 http:\/\/arxiv.org\/abs\/1811.00751"},{"volume-title":"TextBoxes: A Fast Text Detector with a Single Deep Neural Network","author":"Liao Minghui","key":"e_1_3_2_2_19_1","unstructured":"Minghui Liao , Baoguang Shi , Xiang Bai , Xinggang Wang , and Wenyu Liu . 2017. TextBoxes: A Fast Text Detector with a Single Deep Neural Network . In AAAI. AAAI Press , 4161--4167. Minghui Liao, Baoguang Shi, Xiang Bai, Xinggang Wang, and Wenyu Liu. 2017. TextBoxes: A Fast Text Detector with a Single Deep Neural Network. In AAAI. AAAI Press, 4161--4167."},{"key":"e_1_3_2_2_20_1","volume-title":"Scene Text Recognition from Two-Dimensional Perspective. CoRR","author":"Liao Minghui","year":"2018","unstructured":"Minghui Liao , Jian Zhang , Zhaoyi Wan , Fengming Xie , Jiajun Liang , Pengyuan Lyu , Cong Yao , and Xiang Bai . 2018. Scene Text Recognition from Two-Dimensional Perspective. CoRR , Vol. abs\/ 1809 .06508 ( 2018 ). arxiv: 1809.06508 http:\/\/arxiv.org\/abs\/1809.06508 Minghui Liao, Jian Zhang, Zhaoyi Wan, Fengming Xie, Jiajun Liang, Pengyuan Lyu, Cong Yao, and Xiang Bai. 2018. Scene Text Recognition from Two-Dimensional Perspective. CoRR, Vol. abs\/1809.06508 (2018). arxiv: 1809.06508 http:\/\/arxiv.org\/abs\/1809.06508"},{"volume-title":"Char-Net: A Character-Aware Neural Network for Distorted Scene Text Recognition","author":"Liu Wei","key":"e_1_3_2_2_21_1","unstructured":"Wei Liu , Chaofeng Chen , and Kwan-Yee K. Wong . 2018a. Char-Net: A Character-Aware Neural Network for Distorted Scene Text Recognition . In AAAI. AAAI Press , 7154--7161. Wei Liu, Chaofeng Chen, and Kwan-Yee K. Wong. 2018a. Char-Net: A Character-Aware Neural Network for Distorted Scene Text Recognition. In AAAI. AAAI Press, 7154--7161."},{"volume-title":"STAR-Net: A SpaTial Attention Residue Network for Scene Text Recognition","author":"Liu Wei","key":"e_1_3_2_2_22_1","unstructured":"Wei Liu , Chaofeng Chen , Kwan-Yee K. Wong , Zhizhong Su , and Junyu Han . 2016. STAR-Net: A SpaTial Attention Residue Network for Scene Text Recognition . In BMVC. BMVA Press . Wei Liu, Chaofeng Chen, Kwan-Yee K. Wong, Zhizhong Su, and Junyu Han. 2016. STAR-Net: A SpaTial Attention Residue Network for Scene Text Recognition. In BMVC. BMVA Press."},{"key":"e_1_3_2_2_23_1","volume-title":"FOTS: Fast Oriented Text Spotting with a Unified Network. CoRR","author":"Liu Xuebo","year":"2018","unstructured":"Xuebo Liu , Ding Liang , Shi Yan , Dagui Chen , Yu Qiao , and Junjie Yan . 2018 b. FOTS: Fast Oriented Text Spotting with a Unified Network. CoRR , Vol. abs\/ 1801 .01671 (2018). arxiv: 1801.01671 http:\/\/arxiv.org\/abs\/1801.01671 Xuebo Liu, Ding Liang, Shi Yan, Dagui Chen, Yu Qiao, and Junjie Yan. 2018b. FOTS: Fast Oriented Text Spotting with a Unified Network. CoRR, Vol. abs\/1801.01671 (2018). arxiv: 1801.01671 http:\/\/arxiv.org\/abs\/1801.01671"},{"key":"e_1_3_2_2_24_1","volume-title":"Proceedings of the Seventh International Conference on Document Analysis and Recognition -","volume":"2","author":"Lucas S. M.","unstructured":"S. M. Lucas , A. Panaretos , L. Sosa , A. Tang , S. Wong , and R. Young . 2003. ICDAR 2003 Robust Reading Competitions . In Proceedings of the Seventh International Conference on Document Analysis and Recognition - Volume 2 (ICDAR '03). IEEE Computer Society, USA, 682. S. M. Lucas, A. Panaretos, L. Sosa, A. Tang, S. Wong, and R. Young. 2003. ICDAR 2003 Robust Reading Competitions. In Proceedings of the Seventh International Conference on Document Analysis and Recognition - Volume 2 (ICDAR '03). IEEE Computer Society, USA, 682."},{"key":"e_1_3_2_2_25_1","volume-title":"2D Attentional Irregular Scene Text Recognizer. CoRR","author":"Lyu Pengyuan","year":"2019","unstructured":"Pengyuan Lyu , Zhicheng Yang , Xinhang Leng , Xiaojun Wu , Ruiyu Li , and Xiaoyong Shen . 2019. 2D Attentional Irregular Scene Text Recognizer. CoRR , Vol. abs\/ 1906 .05708 ( 2019 ). arxiv: 1906.05708 http:\/\/arxiv.org\/abs\/1906.05708 Pengyuan Lyu, Zhicheng Yang, Xinhang Leng, Xiaojun Wu, Ruiyu Li, and Xiaoyong Shen. 2019. 2D Attentional Irregular Scene Text Recognizer. CoRR, Vol. abs\/1906.05708 (2019). arxiv: 1906.05708 http:\/\/arxiv.org\/abs\/1906.05708"},{"key":"e_1_3_2_2_26_1","volume-title":"Luk\u00e1 s Burget, Jan Cernock\u00fd, and Sanjeev Khudanpur.","author":"Mikolov Tomas","year":"2010","unstructured":"Tomas Mikolov , Martin Karafi\u00e1 t , Luk\u00e1 s Burget, Jan Cernock\u00fd, and Sanjeev Khudanpur. 2010 . Recurrent neural network based language model. In INTERSPEECH. ISCA , 1045--1048. Tomas Mikolov, Martin Karafi\u00e1 t, Luk\u00e1 s Burget, Jan Cernock\u00fd, and Sanjeev Khudanpur. 2010. Recurrent neural network based language model. In INTERSPEECH. ISCA, 1045--1048."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"crossref","unstructured":"Anand Mishra Karteek Alahari and C. V. Jawahar. 2012. Scene Text Recognition using Higher Order Language Priors. In BMVC. BMVA Press 1--11.  Anand Mishra Karteek Alahari and C. V. Jawahar. 2012. Scene Text Recognition using Higher Order Language Priors. In BMVC. BMVA Press 1--11.","DOI":"10.5244\/C.26.127"},{"key":"e_1_3_2_2_28_1","volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","year":"1912","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , Alban Desmaison , Andreas K\u00f6pf , Edward Yang , Zach DeVito , Martin Raison , Alykhan Tejani , Sasank Chilamkurthy , Benoit Steiner , Lu Fang , Junjie Bai , and Soumith Chintala . 2019. PyTorch: An Imperative Style , High-Performance Deep Learning Library . arxiv: 1912 .01703 [cs.LG] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K\u00f6pf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. arxiv: 1912.01703 [cs.LG]"},{"volume-title":"Recognizing Text with Perspective Distortion in Natural Scenes","author":"Phan Trung Quy","key":"e_1_3_2_2_29_1","unstructured":"Trung Quy Phan , Palaiahnakote Shivakumara , Shangxuan Tian , and Chew Lim Tan . 2013. Recognizing Text with Perspective Distortion in Natural Scenes . In ICCV. IEEE Computer Society , 569--576. Trung Quy Phan, Palaiahnakote Shivakumara, Shangxuan Tian, and Chew Lim Tan. 2013. Recognizing Text with Perspective Distortion in Natural Scenes. In ICCV. IEEE Computer Society, 569--576."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2014.07.008"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2646371"},{"volume-title":"Robust Scene Text Recognition with Automatic Rectification","author":"Shi Baoguang","key":"e_1_3_2_2_32_1","unstructured":"Baoguang Shi , Xinggang Wang , Pengyuan Lyu , Cong Yao , and Xiang Bai . 2016. Robust Scene Text Recognition with Automatic Rectification . In CVPR. IEEE Computer Society , 4168--4176. Baoguang Shi, Xinggang Wang, Pengyuan Lyu, Cong Yao, and Xiang Bai. 2016. Robust Scene Text Recognition with Automatic Rectification. In CVPR. IEEE Computer Society, 4168--4176."},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2848939"},{"key":"e_1_3_2_2_34_1","volume-title":"Ralf Schl\u00fc ter, and Hermann Ney","author":"Sundermeyer Martin","year":"2012","unstructured":"Martin Sundermeyer , Ralf Schl\u00fc ter, and Hermann Ney . 2012 . LS\u2122 Neural Networks for Language Modeling. In INTERSPEECH. ISCA , 194--197. Martin Sundermeyer, Ralf Schl\u00fc ter, and Hermann Ney. 2012. LS\u2122 Neural Networks for Language Modeling. In INTERSPEECH. ISCA, 194--197."},{"key":"e_1_3_2_2_35_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008."},{"key":"e_1_3_2_2_36_1","volume-title":"Belongie","author":"Wang Kai","year":"2011","unstructured":"Kai Wang , Boris Babenko , and Serge J . Belongie . 2011 . End-to-end scene text recognition. In ICCV. IEEE Computer Society , 1457--1464. Kai Wang, Boris Babenko, and Serge J. Belongie. 2011. End-to-end scene text recognition. In ICCV. IEEE Computer Society, 1457--1464."},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW50498.2020.00278"},{"key":"e_1_3_2_2_38_1","unstructured":"Donghyeon Won Zachary C. Steinert-Threlkeld and Jungseock Joo. 2017. Protest Activity Detection and Perceived Violence Estimation from Social Media Images. In ACM Multimedia. ACM 786--794.  Donghyeon Won Zachary C. Steinert-Threlkeld and Jungseock Joo. 2017. Protest Activity Detection and Perceived Violence Estimation from Social Media Images. In ACM Multimedia. ACM 786--794."},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"crossref","unstructured":"Xiao Yang Dafang He Zihan Zhou Daniel Kifer and C. Lee Giles. 2017. Learning to Read Irregular Text with Attention Mechanisms. In IJCAI. ijcai.org 3280--3286.  Xiao Yang Dafang He Zihan Zhou Daniel Kifer and C. Lee Giles. 2017. Learning to Read Irregular Text with Attention Mechanisms. In IJCAI. ijcai.org 3280--3286.","DOI":"10.24963\/ijcai.2017\/458"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2014.2353813"},{"key":"e_1_3_2_2_41_1","volume-title":"ADADELTA: An Adaptive Learning Rate Method. CoRR","author":"Zeiler Matthew D.","year":"2012","unstructured":"Matthew D. Zeiler . 2012 . ADADELTA: An Adaptive Learning Rate Method. CoRR , Vol. abs\/ 1212 .5701 (2012). Matthew D. Zeiler. 2012. ADADELTA: An Adaptive Learning Rate Method. CoRR, Vol. abs\/1212.5701 (2012)."},{"key":"e_1_3_2_2_42_1","volume-title":"ESIR: End-To-End Scene Text Recognition via Iterative Image Rectification","author":"Zhan Fangneng","year":"2019","unstructured":"Fangneng Zhan and Shijian Lu . 2019 . ESIR: End-To-End Scene Text Recognition via Iterative Image Rectification . In CVPR. Computer Vision Foundation \/ IEEE , 2059--2068. Fangneng Zhan and Shijian Lu. 2019. ESIR: End-To-End Scene Text Recognition via Iterative Image Rectification. In CVPR. Computer Vision Foundation \/ IEEE, 2059--2068."},{"key":"e_1_3_2_2_43_1","volume-title":"Deep Neural Network for Semantic-based Text Recognition in Images. CoRR","author":"Zheng Yi","year":"2020","unstructured":"Yi Zheng , Qitong Wang , and Margrit Betke . 2020. Deep Neural Network for Semantic-based Text Recognition in Images. CoRR , Vol. abs\/ 1908 .01403 ( 2020 ), 9. http:\/\/arxiv.org\/abs\/1908.01403 Yi Zheng, Qitong Wang, and Margrit Betke. 2020. Deep Neural Network for Semantic-based Text Recognition in Images. CoRR, Vol. abs\/1908.01403 (2020), 9. http:\/\/arxiv.org\/abs\/1908.01403"}],"event":{"name":"MM '20: The 28th ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Seattle WA USA","acronym":"MM '20"},"container-title":["Proceedings of the 28th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413913","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3394171.3413913","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:32:06Z","timestamp":1750195926000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413913"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,12]]},"references-count":43,"alternative-id":["10.1145\/3394171.3413913","10.1145\/3394171"],"URL":"https:\/\/doi.org\/10.1145\/3394171.3413913","relation":{},"subject":[],"published":{"date-parts":[[2020,10,12]]},"assertion":[{"value":"2020-10-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}