{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T15:14:19Z","timestamp":1784301259802,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":44,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"the Science and Technology Foundation of Guangzhou Huangpu Development District","award":["2020GH17"],"award-info":[{"award-number":["2020GH17"]}]},{"name":"NSFC","award":["61936003"],"award-info":[{"award-number":["61936003"]}]},{"name":"GD-NSF","award":["2017A030312006,2021A1515011870"],"award-info":[{"award-number":["2017A030312006,2021A1515011870"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3547942","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:42:46Z","timestamp":1665416566000},"page":"4272-4281","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":63,"title":["SPTS: Single-Point Text Spotting"],"prefix":"10.1145","author":[{"given":"Dezhi","family":"Peng","sequence":"first","affiliation":[{"name":"South China University of Technology, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinyu","family":"Wang","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuliang","family":"Liu","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiaxin","family":"Zhang","sequence":"additional","affiliation":[{"name":"South China University of Technology, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mingxin","family":"Huang","sequence":"additional","affiliation":[{"name":"South China University of Technology, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Songxuan","family":"Lai","sequence":"additional","affiliation":[{"name":"Huawei Cloud Computing Technologies, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jing","family":"Li","sequence":"additional","affiliation":[{"name":"Huawei Cloud Computing Technologies, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shenggao","family":"Zhu","sequence":"additional","affiliation":[{"name":"Huawei Cloud Computing Technologies, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dahua","family":"Lin","sequence":"additional","affiliation":[{"name":"Chinese University of Hong Kong, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chunhua","family":"Shen","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiang","family":"Bai","sequence":"additional","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lianwen","family":"Jin","sequence":"additional","affiliation":[{"name":"South China University of Technology, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00959"},{"key":"e_1_3_2_2_2_1","volume-title":"Proc. AAAI Conf. Artificial Intell. 6674--6681","author":"Bartz Christian","year":"2018","unstructured":"Christian Bartz , Haojin Yang , and Christoph Meinel . 2018 . SEE: Towards semisupervised end-to-end scene text recognition . In Proc. AAAI Conf. Artificial Intell. 6674--6681 . Christian Bartz, Haojin Yang, and Christoph Meinel. 2018. SEE: Towards semisupervised end-to-end scene text recognition. In Proc. AAAI Conf. Artificial Intell. 6674--6681."},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.242"},{"key":"e_1_3_2_2_4_1","volume-title":"Proc. Int. Conf. Learn. Represent.","author":"Chen Ting","year":"2022","unstructured":"Ting Chen , Saurabh Saxena , Lala Li , David J Fleet , and Geoffrey Hinton . 2022 . Pix2Seq: A language modeling framework for object detection . In Proc. Int. Conf. Learn. Represent. Ting Chen, Saurabh Saxena, Lala Li, David J Fleet, and Geoffrey Hinton. 2022. Pix2Seq: A language modeling framework for object detection. In Proc. Int. Conf. Learn. Represent."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2017.157"},{"key":"e_1_3_2_2_6_1","volume-title":"ICDAR2019 robust reading challenge on arbitrary-shaped text-RRC-ArT. In Proc. Int. Conf. Doc. Anal. and Recognit. 1571--1576","author":"Chng Chee Kheng","year":"2019","unstructured":"Chee Kheng Chng , Yuliang Liu , Yipeng Sun , Chun Chet Ng , Canjie Luo , Zihan Ni , ChuanMing Fang , Shuaitao Zhang , Junyu Han , Errui Ding , 2019 . ICDAR2019 robust reading challenge on arbitrary-shaped text-RRC-ArT. In Proc. Int. Conf. Doc. Anal. and Recognit. 1571--1576 . Chee Kheng Chng, Yuliang Liu, Yipeng Sun, Chun Chet Ng, Canjie Luo, Zihan Ni, ChuanMing Fang, Shuaitao Zhang, Junyu Han, Errui Ding, et al. 2019. ICDAR2019 robust reading challenge on arbitrary-shaped text-RRC-ArT. In Proc. Int. Conf. Doc. Anal. and Recognit. 1571--1576."},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-009-0275-4"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00917"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.254"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00527"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00815"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.529"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0823-z"},{"key":"e_1_3_2_2_14_1","volume-title":"ICDAR 2015 competition on robust reading. In Proc. Int. Conf. Doc. Anal. and Recognit. IEEE, 1156--1160","author":"Karatzas Dimosthenis","year":"2015","unstructured":"Dimosthenis Karatzas , Lluis Gomez-Bigorda , Anguelos Nicolaou , Suman Ghosh , Andrew Bagdanov , Masakazu Iwamura , Jiri Matas , Lukas Neumann , Vijay Ramaseshan Chandrasekhar , Shijian Lu , 2015 . ICDAR 2015 competition on robust reading. In Proc. Int. Conf. Doc. Anal. and Recognit. IEEE, 1156--1160 . Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, et al. 2015. ICDAR 2015 competition on robust reading. In Proc. Int. Conf. Doc. Anal. and Recognit. IEEE, 1156--1160."},{"key":"e_1_3_2_2_15_1","volume-title":"ICDAR 2011 robust reading competition-challenge 1: reading text in born-digital images (web and email). In Proc. Int. Conf. Doc. Anal. and Recognit. 1485--1490","author":"Karatzas Dimosthenis","year":"2011","unstructured":"Dimosthenis Karatzas , S Robles Mestre , Joan Mas , Farshad Nourbakhsh , and P Pratim Roy . 2011 . ICDAR 2011 robust reading competition-challenge 1: reading text in born-digital images (web and email). In Proc. Int. Conf. Doc. Anal. and Recognit. 1485--1490 . Dimosthenis Karatzas, S Robles Mestre, Joan Mas, Farshad Nourbakhsh, and P Pratim Roy. 2011. ICDAR 2011 robust reading competition-challenge 1: reading text in born-digital images (web and email). In Proc. Int. Conf. Doc. Anal. and Recognit. 1485--1490."},{"key":"e_1_3_2_2_16_1","volume-title":"ICDAR 2013 robust reading competition. In Proc. Int. Conf. Doc. Anal. and Recognit. IEEE, 1484--1493","author":"Karatzas Dimosthenis","year":"2013","unstructured":"Dimosthenis Karatzas , Faisal Shafait , Seiichi Uchida , Masakazu Iwamura , Lluis Gomez i Bigorda , Sergi Robles Mestre , Joan Mas , David Fernandez Mota , Jon Almazan Almazan , and Lluis Pere De Las Heras . 2013 . ICDAR 2013 robust reading competition. In Proc. Int. Conf. Doc. Anal. and Recognit. IEEE, 1484--1493 . Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernandez Mota, Jon Almazan Almazan, and Lluis Pere De Las Heras. 2013. ICDAR 2013 robust reading competition. In Proc. Int. Conf. Doc. Anal. and Recognit. IEEE, 1484--1493."},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.560"},{"key":"e_1_3_2_2_18_1","volume-title":"Mask TextSpotter: An end-to-end trainable neural network for spotting text with arbitrary shapes","author":"Liao Minghui","year":"2019","unstructured":"Minghui Liao , Pengyuan Lyu , Minghang He , Cong Yao , Wenhao Wu , and Xiang Bai . 2019. Mask TextSpotter: An end-to-end trainable neural network for spotting text with arbitrary shapes . IEEE Trans. Pattern Anal. Mach. Intell . ( 2019 ), 532--548. Minghui Liao, Pengyuan Lyu, Minghang He, Cong Yao, Wenhao Wu, and Xiang Bai. 2019. Mask TextSpotter: An end-to-end trainable neural network for spotting text with arbitrary shapes. IEEE Trans. Pattern Anal. Mach. Intell. (2019), 532--548."},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58621-8_41"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.11196"},{"key":"e_1_3_2_2_21_1","volume-title":"Proc. IEEE Int. Conf. Comp. Vis. 9126--9136","author":"Linjie Xing","unstructured":"Xing Linjie , Tian Zhi , Huang Weilin , and R. Scott Matthew . 2019. Convolutional Character Networks . In Proc. IEEE Int. Conf. Comp. Vis. 9126--9136 . Xing Linjie, Tian Zhi, Huang Weilin, and R. Scott Matthew. 2019. Convolutional Character Networks. In Proc. IEEE Int. Conf. Comp. Vis. 9126--9136."},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00595"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00983"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.368"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2019.02.002"},{"key":"e_1_3_2_2_26_1","volume-title":"ABCNet v2: Adaptive Bezier-Curve Network for Real-time End-to-end Text Spotting","author":"Liu Yuliang","year":"2021","unstructured":"Yuliang Liu , Chunhua Shen , Lianwen Jin , Tong He , Peng Chen , Chongyu Liu , and Hao Chen . 2021. ABCNet v2: Adaptive Bezier-Curve Network for Real-time End-to-end Text Spotting . IEEE Trans. Pattern Anal. Mach. Intell . ( 2021 ), 1--1. https:\/\/doi.org\/10.1109\/TPAMI.2021.3107437 Yuliang Liu, Chunhua Shen, Lianwen Jin, Tong He, Peng Chen, Chongyu Liu, and Hao Chen. 2021. ABCNet v2: Adaptive Bezier-Curve Network for Real-time End-to-end Text Spotting. IEEE Trans. Pattern Anal. Mach. Intell. (2021), 1--1. https:\/\/doi.org\/10.1109\/TPAMI.2021.3107437"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01216-8_2"},{"key":"e_1_3_2_2_28_1","volume-title":"Proc. Int. Conf. Learn. Representations","author":"Loshchilov Ilya","year":"2018","unstructured":"Ilya Loshchilov and Frank Hutter . 2018 . Decoupled weight decay regularization . Proc. Int. Conf. Learn. Representations (2018). Ilya Loshchilov and Frank Hutter. 2018. Decoupled weight decay regularization. Proc. Int. Conf. Learn. Representations (2018)."},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_5"},{"key":"e_1_3_2_2_30_1","volume-title":"ICDAR 2019 Robust Reading Challenge on Multi-lingual Scene Text Detection and Recognition--RRC-MLT-2019. In Proc. Int. Conf. Doc. Anal. and Recognit. 1582--1587","author":"Nayef Nibal","year":"2019","unstructured":"Nibal Nayef , Yash Patel , Michal Busta , Pinaki Nath Chowdhury , Dimosthenis Karatzas , Wafa Khlif , Jiri Matas , Umapada Pal , Jean-Christophe Burie , Cheng-lin Liu, 2019 . ICDAR 2019 Robust Reading Challenge on Multi-lingual Scene Text Detection and Recognition--RRC-MLT-2019. In Proc. Int. Conf. Doc. Anal. and Recognit. 1582--1587 . Nibal Nayef, Yash Patel, Michal Busta, Pinaki Nath Chowdhury, Dimosthenis Karatzas, Wafa Khlif, Jiri Matas, Umapada Pal, Jean-Christophe Burie, Cheng-lin Liu, et al. 2019. ICDAR 2019 Robust Reading Challenge on Multi-lingual Scene Text Detection and Recognition--RRC-MLT-2019. In Proc. Int. Conf. Doc. Anal. and Recognit. 1582--1587."},{"key":"e_1_3_2_2_31_1","volume-title":"ICDAR 2017 robust reading challenge on multi-lingual scene text detection and script identification-RRC-MLT. In Proc. Int. Conf. Doc. Anal. and Recognit.","volume":"1","author":"Nayef Nibal","year":"2017","unstructured":"Nibal Nayef , Fei Yin , Imen Bizid , Hyunsoo Choi , Yuan Feng , Dimosthenis Karatzas , Zhenbo Luo , Umapada Pal , Christophe Rigaud , Joseph Chazalon , 2017 . ICDAR 2017 robust reading challenge on multi-lingual scene text detection and script identification-RRC-MLT. In Proc. Int. Conf. Doc. Anal. and Recognit. , Vol. 1 . IEEE, 1454--1459. Nibal Nayef, Fei Yin, Imen Bizid, Hyunsoo Choi, Yuan Feng, Dimosthenis Karatzas, Zhenbo Luo, Umapada Pal, Christophe Rigaud, Joseph Chazalon, et al. 2017. ICDAR 2017 robust reading challenge on multi-lingual scene text detection and script identification-RRC-MLT. In Proc. Int. Conf. Doc. Anal. and Recognit., Vol. 1. IEEE, 1454--1459."},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i3.16348"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00480"},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.166"},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_4"},{"key":"e_1_3_2_2_36_1","volume-title":"Proc. Advances in Neural Inf. Process. Syst.","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , ? ukasz Kaiser, and Illia Polosukhin . 2017 . Attention is All you Need . In Proc. Advances in Neural Inf. Process. Syst. , Vol. 30 . Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ? ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Proc. Advances in Neural Inf. Process. Syst., Vol. 30."},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413819"},{"key":"e_1_3_2_2_38_1","volume-title":"Proc. AAAI Conf. Artificial Intell. 12160--12167","author":"Lu Pu","year":"2020","unstructured":"HaoWang, Pu Lu , Hui Zhang , Mingkun Yang , Xiang Bai , Yongchao Xu , Mengchao He , Yongpan Wang , and Wenyu Liu . 2020 . All You Need Is Boundary: Toward Arbitrary-Shaped Text Spotting . In Proc. AAAI Conf. Artificial Intell. 12160--12167 . HaoWang, Pu Lu, Hui Zhang, Mingkun Yang, Xiang Bai, Yongchao Xu, Mengchao He, Yongpan Wang, and Wenyu Liu. 2020. All You Need Is Boundary: Toward Arbitrary-Shaped Text Spotting. In Proc. AAAI Conf. Artificial Intell. 12160--12167."},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i4.16383"},{"key":"e_1_3_2_2_40_1","volume-title":"PAN: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text","author":"Wang Wenhai","year":"2021","unstructured":"Wenhai Wang , Enze Xie , Xiang Li , Xuebo Liu , Ding Liang , Yang Zhibo , Tong Lu , and Chunhua Shen . 2021 . PAN: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text . IEEE Transactions on Pattern Analysis and Machine Intelligence ( 2021), 1--1. https:\/\/doi.org\/10.1109\/TPAMI.2021.3077555 Wenhai Wang, Enze Xie, Xiang Li, Xuebo Liu, Ding Liang, Yang Zhibo, Tong Lu, and Chunhua Shen. 2021. PAN: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text. IEEE Transactions on Pattern Analysis and Machine Intelligence (2021), 1--1. https:\/\/doi.org\/10.1109\/TPAMI.2021.3077555"},{"key":"e_1_3_2_2_41_1","volume-title":"Proc. Int. Conf. Mach. Learn. 10524--10533","author":"Xiong Ruibin","year":"2020","unstructured":"Ruibin Xiong , Yunchang Yang , Di He , Kai Zheng , Shuxin Zheng , Chen Xing , Huishuai Zhang , Yanyan Lan , Liwei Wang , and Tieyan Liu . 2020 . On layer normalization in the Transformer architecture . In Proc. Int. Conf. Mach. Learn. 10524--10533 . Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. 2020. On layer normalization in the Transformer architecture. In Proc. Int. Conf. Mach. Learn. 10524--10533."},{"key":"e_1_3_2_2_42_1","volume-title":"Proc. IEEE Conf. Comp. Vis. Patt. Recogn. IEEE, 1083--1090","author":"Yao Cong","year":"2012","unstructured":"Cong Yao , Xiang Bai , Wenyu Liu , Yi Ma , and Zhuowen Tu . 2012 . Detecting texts of arbitrary orientations in natural images . In Proc. IEEE Conf. Comp. Vis. Patt. Recogn. IEEE, 1083--1090 . Cong Yao, Xiang Bai, Wenyu Liu, Yi Ma, and Zhuowen Tu. 2012. Detecting texts of arbitrary orientations in natural images. In Proc. IEEE Conf. Comp. Vis. Patt. Recogn. IEEE, 1083--1090."},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.283"},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00314"}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3547942","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3547942","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:31Z","timestamp":1750186831000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3547942"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":44,"alternative-id":["10.1145\/3503161.3547942","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3547942","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}