{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:12:10Z","timestamp":1750219930059,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":53,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,3,17]],"date-time":"2023-03-17T00:00:00Z","timestamp":1679011200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,3,17]]},"DOI":"10.1145\/3594315.3594356","type":"proceedings-article","created":{"date-parts":[[2023,8,3]],"date-time":"2023-08-03T00:14:16Z","timestamp":1691021656000},"page":"452-458","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["A Decoupled Language Model Based on Contrastive Attention Mechanism for Scene Text Recognition"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2734-1310","authenticated-orcid":false,"given":"Junwei","family":"Zhou","sequence":"first","affiliation":[{"name":"Chinese Academy of Sciences, Institute of Information Engineering, China and University of Chinese Academy of Sciences, School of Cyber Security, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6642-8160","authenticated-orcid":false,"given":"Xi","family":"Wang","sequence":"additional","affiliation":[{"name":"Chinese Academy of Sciences, Institute of Information Engineering, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3559-8009","authenticated-orcid":false,"given":"Jiao","family":"Dai","sequence":"additional","affiliation":[{"name":"Chinese Academy of Sciences, Institute of Information Engineering, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1107-3873","authenticated-orcid":false,"given":"Jizhong","family":"Han","sequence":"additional","affiliation":[{"name":"Chinese Academy of Sciences, Institute of Information Engineering, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,8,2]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00959"},{"key":"e_1_3_2_1_2_1","volume-title":"Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473","author":"Bahdanau Dzmitry","year":"2014","unstructured":"Dzmitry Bahdanau , Kyunghyun Cho , and Yoshua Bengio . 2014. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473 ( 2014 ). Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473 (2014)."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00163"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01467"},{"key":"e_1_3_2_1_5_1","volume-title":"Big self-supervised models are strong semi-supervised learners. Advances in neural information processing systems 33","author":"Chen Ting","year":"2020","unstructured":"Ting Chen , Simon Kornblith , Kevin Swersky , Mohammad Norouzi , and Geoffrey\u00a0 E Hinton . 2020. Big self-supervised models are strong semi-supervised learners. Advances in neural information processing systems 33 ( 2020 ), 22243\u201322255. Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey\u00a0E Hinton. 2020. Big self-supervised models are strong semi-supervised learners. Advances in neural information processing systems 33 (2020), 22243\u201322255."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.543"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00584"},{"key":"e_1_3_2_1_8_1","volume-title":"Svtr: Scene text recognition with a single visual model. IJCAI","author":"Du Yongkun","year":"2022","unstructured":"Yongkun Du , Zhineng Chen , Caiyan Jia , Xiaoting Yin , Tianlun Zheng , Chenxia Li , Yuning Du , and Yu-Gang Jiang . 2022 . Svtr: Scene text recognition with a single visual model. IJCAI (2022). Yongkun Du, Zhineng Chen, Caiyan Jia, Xiaoting Yin, Tianlun Zheng, Chenxia Li, Yuning Du, and Yu-Gang Jiang. 2022. Svtr: Scene text recognition with a single visual model. IJCAI (2022)."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00702"},{"key":"e_1_3_2_1_10_1","volume-title":"Attention and Language Ensemble for Scene Text Recognition with Convolutional Sequence Modeling. ACM MM","author":"Fang Shancheng","year":"2018","unstructured":"Shancheng Fang , Hongtao Xie , Zheng-Jun Zha , Nannan Sun , Jianlong Tan , and Yongdong Zhang . 2018. Attention and Language Ensemble for Scene Text Recognition with Convolutional Sequence Modeling. ACM MM ( 2018 ). Shancheng Fang, Hongtao Xie, Zheng-Jun Zha, Nannan Sun, Jianlong Tan, and Yongdong Zhang. 2018. Attention and Language Ensemble for Scene Text Recognition with Convolutional Sequence Modeling. ACM MM (2018)."},{"key":"e_1_3_2_1_11_1","volume-title":"Generative adversarial nets. Advances in neural information processing systems 27","author":"Goodfellow Ian","year":"2014","unstructured":"Ian Goodfellow , Jean Pouget-Abadie , Mehdi Mirza , Bing Xu , David Warde-Farley , Sherjil Ozair , Aaron Courville , and Yoshua Bengio . 2014. Generative adversarial nets. Advances in neural information processing systems 27 ( 2014 ). Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. Advances in neural information processing systems 27 (2014)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143891"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.254"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0823-z"},{"key":"e_1_3_2_1_16_1","volume-title":"Spatial Transformer Networks. arXiv","author":"Jaderberg Max","year":"2015","unstructured":"Max Jaderberg , Karen Simonyan , Andrew Zisserman , and Koray Kavukcuoglu . 2015. Spatial Transformer Networks. arXiv ( 2015 ). Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu. 2015. Spatial Transformer Networks. arXiv (2015)."},{"key":"e_1_3_2_1_17_1","volume-title":"ICDAR 2015 competition on robust reading. In 2015 13th International Conference on Document Analysis and Recognition (ICDAR). IEEE, 1156\u20131160","author":"Karatzas Dimosthenis","year":"2015","unstructured":"Dimosthenis Karatzas , Lluis Gomez-Bigorda , Anguelos Nicolaou , Suman Ghosh , Andrew Bagdanov , Masakazu Iwamura , Jiri Matas , Lukas Neumann , Vijay\u00a0Ramaseshan Chandrasekhar , Shijian Lu , 2015 . ICDAR 2015 competition on robust reading. In 2015 13th International Conference on Document Analysis and Recognition (ICDAR). IEEE, 1156\u20131160 . Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay\u00a0Ramaseshan Chandrasekhar, Shijian Lu, 2015. ICDAR 2015 competition on robust reading. In 2015 13th International Conference on Document Analysis and Recognition (ICDAR). IEEE, 1156\u20131160."},{"key":"e_1_3_2_1_18_1","volume-title":"ICDAR 2013 robust reading competition. In 2013 12th International Conference on Document Analysis and Recognition (ICDAR). IEEE, 1484\u20131493","author":"Karatzas Dimosthenis","year":"2013","unstructured":"Dimosthenis Karatzas , Faisal Shafait , Seiichi Uchida , Masakazu Iwamura , Lluis\u00a0Gomez i Bigorda , Sergi\u00a0Robles Mestre , Joan Mas , David\u00a0Fernandez Mota , Jon\u00a0Almazan Almazan , and Lluis\u00a0Pere De\u00a0Las\u00a0Heras . 2013 . ICDAR 2013 robust reading competition. In 2013 12th International Conference on Document Analysis and Recognition (ICDAR). IEEE, 1484\u20131493 . Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis\u00a0Gomez i Bigorda, Sergi\u00a0Robles Mestre, Joan Mas, David\u00a0Fernandez Mota, Jon\u00a0Almazan Almazan, and Lluis\u00a0Pere De\u00a0Las\u00a0Heras. 2013. ICDAR 2013 robust reading competition. In 2013 12th International Conference on Document Analysis and Recognition (ICDAR). IEEE, 1484\u20131493."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018610"},{"key":"e_1_3_2_1_20_1","volume-title":"Char-Net: A Character-Aware Neural Network for Distorted Scene Text Recognition. AAAI Conference on Artificial Intelligence (AAAI)","author":"Liu Wei","year":"2018","unstructured":"Wei Liu , Chaofeng Chen , and Kwanyee\u00a0 K Wong . 2018 . Char-Net: A Character-Aware Neural Network for Distorted Scene Text Recognition. AAAI Conference on Artificial Intelligence (AAAI) (2018), 7154\u20137161. Wei Liu, Chaofeng Chen, and Kwanyee\u00a0K Wong. 2018. Char-Net: A Character-Aware Neural Network for Distorted Scene Text Recognition. AAAI Conference on Artificial Intelligence (AAAI) (2018), 7154\u20137161."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Wei Liu Chaofeng Chen Kwan-Yee\u00a0K Wong Zhizhong Su and Junyu Han. 2016. Star-net: a spatial attention residue network for scene text recognition.. In BMVC Vol.\u00a02. 7.  Wei Liu Chaofeng Chen Kwan-Yee\u00a0K Wong Zhizhong Su and Junyu Han. 2016. Star-net: a spatial attention residue network for scene text recognition.. In BMVC Vol.\u00a02. 7.","DOI":"10.5244\/C.30.43"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00983"},{"key":"e_1_3_2_1_23_1","volume-title":"Hierarchical question-image co-attention for visual question answering. Advances in neural information processing systems 29","author":"Lu Jiasen","year":"2016","unstructured":"Jiasen Lu , Jianwei Yang , Dhruv Batra , and Devi Parikh . 2016. Hierarchical question-image co-attention for visual question answering. Advances in neural information processing systems 29 ( 2016 ). Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016. Hierarchical question-image co-attention for visual question answering. Advances in neural information processing systems 29 (2016)."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10032-004-0134-3"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2019.01.020"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-020-01411-1"},{"key":"e_1_3_2_1_27_1","volume-title":"Scene text recognition using higher order language priors. BMVC","author":"Mishra Anand","year":"2012","unstructured":"Anand Mishra , Karteek Alahari , and CV Jawahar . 2012. Scene text recognition using higher order language priors. BMVC ( 2012 ), 1\u201311. Anand Mishra, Karteek Alahari, and CV Jawahar. 2012. Scene text recognition using higher order language priors. BMVC (2012), 1\u201311."},{"key":"e_1_3_2_1_28_1","volume-title":"Recurrent models of visual attention. Advances in neural information processing systems 27","author":"Mnih Volodymyr","year":"2014","unstructured":"Volodymyr Mnih , Nicolas Heess , Alex Graves , 2014. Recurrent models of visual attention. Advances in neural information processing systems 27 ( 2014 ). Volodymyr Mnih, Nicolas Heess, Alex Graves, 2014. Recurrent models of visual attention. Advances in neural information processing systems 27 (2014)."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2012.6248097"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01354"},{"key":"e_1_3_2_1_31_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV). 569\u2013576","author":"Quy\u00a0Phan Trung","year":"2013","unstructured":"Trung Quy\u00a0Phan , Palaiahnakote Shivakumara , Shangxuan Tian , and Chew Lim\u00a0Tan . 2013 . Recognizing text with perspective distortion in natural scenes . In Proceedings of the IEEE International Conference on Computer Vision (ICCV). 569\u2013576 . Trung Quy\u00a0Phan, Palaiahnakote Shivakumara, Shangxuan Tian, and Chew Lim\u00a0Tan. 2013. Recognizing text with perspective distortion in natural scenes. In Proceedings of the IEEE International Conference on Computer Vision (ICCV). 569\u2013576."},{"key":"e_1_3_2_1_32_1","volume-title":"International conference on machine learning. PMLR, 8748\u20138763","author":"Radford Alec","year":"2021","unstructured":"Alec Radford , Jong\u00a0Wook Kim , Chris Hallacy , Aditya Ramesh , Gabriel Goh , Sandhini Agarwal , Girish Sastry , Amanda Askell , Pamela Mishkin , Jack Clark , 2021 . Learning transferable visual models from natural language supervision . In International conference on machine learning. PMLR, 8748\u20138763 . Alec Radford, Jong\u00a0Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PMLR, 8748\u20138763."},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2014.07.008"},{"key":"e_1_3_2_1_34_1","volume-title":"An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition","author":"Shi Baoguang","year":"2016","unstructured":"Baoguang Shi , Xiang Bai , and Cong Yao . 2016. An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition . IEEE transactions on pattern analysis and machine intelligence 39, 11 ( 2016 ), 2298\u20132304. Baoguang Shi, Xiang Bai, and Cong Yao. 2016. An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE transactions on pattern analysis and machine intelligence 39, 11 (2016), 2298\u20132304."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2646371"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.452"},{"key":"e_1_3_2_1_37_1","volume-title":"Robust Scene Text Recognition with Automatic Rectification. arXiv","author":"Shi Baoguang","year":"2016","unstructured":"Baoguang Shi , Xinggang Wang , Pengyuan Lyu , Cong Yao , and Xiang Bai . 2016. Robust Scene Text Recognition with Automatic Rectification. arXiv ( 2016 ). Baoguang Shi, Xinggang Wang, Pengyuan Lyu, Cong Yao, and Xiang Bai. 2016. Robust Scene Text Recognition with Automatic Rectification. arXiv (2016)."},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2848939"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00918"},{"key":"e_1_3_2_1_40_1","volume-title":"Visual-semantic transformer for scene text recognition. AAAI","author":"Tang Xin","year":"2022","unstructured":"Xin Tang , Yongquan Lai , Ying Liu , Yuanyuan Fu , and Rui Fang . 2022. Visual-semantic transformer for scene text recognition. AAAI ( 2022 ). Xin Tang, Yongquan Lai, Ying Liu, Yuanyuan Fu, and Rui Fang. 2022. Visual-semantic transformer for scene text recognition. AAAI (2022)."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00436"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00436"},{"key":"e_1_3_2_1_43_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan\u00a0 N Gomez , \u0141ukasz Kaiser , and Illia Polosukhin . 2017. Attention is all you need. Advances in neural information processing systems 30 ( 2017 ). Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan\u00a0N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_3_2_1_44_1","volume-title":"International Conference on Computer Vision (ICCV)","author":"Wang Kai","year":"2011","unstructured":"Kai Wang , Boris Babenko , and Serge Belongie . 2011 . End-to-end scene text recognition . International Conference on Computer Vision (ICCV) (2011), 1457\u20131464. Kai Wang, Boris Babenko, and Serge Belongie. 2011. End-to-end scene text recognition. International Conference on Computer Vision (ICCV) (2011), 1457\u20131464."},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-15549-9_43"},{"key":"e_1_3_2_1_46_1","volume-title":"A simple and robust convolutional-attention network for irregular text recognition. arXiv preprint arXiv:1904.01375 6, 2","author":"Wang Peng","year":"2019","unstructured":"Peng Wang , Lu Yang , Hui Li , Yuyan Deng , Chunhua Shen , and Yanning Zhang . 2019. A simple and robust convolutional-attention network for irregular text recognition. arXiv preprint arXiv:1904.01375 6, 2 ( 2019 ), 1. Peng Wang, Lu Yang, Hui Li, Yuyan Deng, Chunhua Shen, and Yanning Zhang. 2019. A simple and robust convolutional-attention network for irregular text recognition. arXiv preprint arXiv:1904.01375 6, 2 (2019), 1."},{"key":"e_1_3_2_1_47_1","volume-title":"Scene Text Recognition via Gated Cascade Attention. ICME","author":"Wang Siwei","year":"2019","unstructured":"Siwei Wang , Yongtao Wang , Xiaoran Qin , Qijie Zhao , and Zhi Tang . 2019. Scene Text Recognition via Gated Cascade Attention. ICME ( 2019 ), 1018\u20131023. Siwei Wang, Yongtao Wang, Xiaoran Qin, Qijie Zhao, and Zhi Tang. 2019. Scene Text Recognition via Gated Cascade Attention. ICME (2019), 1018\u20131023."},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01393"},{"key":"e_1_3_2_1_49_1","volume-title":"International conference on machine learning. PMLR","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu , Jimmy Ba , Ryan Kiros , Kyunghyun Cho , Aaron Courville , Ruslan Salakhudinov , Rich Zemel , and Yoshua Bengio . 2015 . Show, attend and tell: Neural image caption generation with visual attention . In International conference on machine learning. PMLR , 2048\u20132057. Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In International conference on machine learning. PMLR, 2048\u20132057."},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.515"},{"key":"e_1_3_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01213"},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58529-7_9"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00216"}],"event":{"name":"ICCAI 2023: 2023 9th International Conference on Computing and Artificial Intelligence","acronym":"ICCAI 2023","location":"Tianjin China"},"container-title":["Proceedings of the 2023 9th International Conference on Computing and Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3594315.3594356","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3594315.3594356","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:48:40Z","timestamp":1750182520000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3594315.3594356"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,17]]},"references-count":53,"alternative-id":["10.1145\/3594315.3594356","10.1145\/3594315"],"URL":"https:\/\/doi.org\/10.1145\/3594315.3594356","relation":{},"subject":[],"published":{"date-parts":[[2023,3,17]]},"assertion":[{"value":"2023-08-02","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}