{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T06:59:48Z","timestamp":1777705188421,"version":"3.51.4"},"reference-count":16,"publisher":"SAGE Publications","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IFS"],"published-print":{"date-parts":[[2024,4,18]]},"abstract":"<jats:p>\u00a0This research focuses on Scene Text Recognition (STR), a crucial component in various applications of artificial intelligence such as image retrieval, office automation, and intelligent traffic systems. Recent studies have shown that semantic-aware approaches significantly improve the performance of STR tasks, with context-aware STR methods becoming mainstream. Among these, the fusion of visual and language models has shown remarkable effectiveness. We propose a novel method (PABINet) that incorporates three key components: a Visual-Language Decoder, a Language Model, and a Fusion Model. First, during training, the Visual-Language Decoder masks the original labels in the Transformer decoder using permutation masks, with each mask being unique. This enhances word memorization and learning through contextual semantic information, resulting in robust semantic knowledge. During the inference stage, the Visual-Language Decoder employs autonomous Autoregressive model (AR) inference to generate results. Subsequently, the Language Model scrutinizes and corrects the output of the Visual-Language Encoder using a cloze mask approach, achieving context-aware, autonomous, bidirectional inference. Finally, the Fusion Model concatenates and refines the outputs of both models through iterative layers.Experimental results demonstrate that our PABINet performs exceptionally well when handling various quality images. When trained with synthetic data, PABINet achieves a new STR benchmark (average accuracy of 92.41%), and when trained with real data, it establishes new state-of-the-art results (average accuracy of 96.28%).<\/jats:p>","DOI":"10.3233\/jifs-237135","type":"journal-article","created":{"date-parts":[[2024,2,23]],"date-time":"2024-02-23T11:31:19Z","timestamp":1708687879000},"page":"8605-8616","source":"Crossref","is-referenced-by-count":0,"title":["Scene text recognition with context-aware autonomous bidirectional iterative models"],"prefix":"10.1177","volume":"46","author":[{"given":"Xiaoqing","family":"Zhao","sequence":"first","affiliation":[{"name":"College of Computer Science and Technology, Xinjiang University, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Miaomiao","family":"Xu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Xinjiang University, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yanbing","family":"Li","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Xinjiang University, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hao","family":"Huang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Xinjiang University, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wushour","family":"Silamu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Xinjiang University, Urumqi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","reference":[{"key":"10.3233\/JIFS-237135_ref3","doi-asserted-by":"crossref","first-page":"161","DOI":"10.1007\/s11263-020-01369-0","article-title":"Scene text detection and recognition: The deep learning era","volume":"129","author":"Shangbang Long","year":"2021","journal-title":"International Journal of Computer Vision"},{"key":"10.3233\/JIFS-237135_ref4","doi-asserted-by":"crossref","first-page":"19","DOI":"10.1007\/s11704-015-4488-0","article-title":"Scene text detection and recognition: Recent advances and future trends","volume":"10","author":"Yingying Zhu","year":"2016","journal-title":"Frontiers of Computer Science"},{"issue":"18","key":"10.3233\/JIFS-237135_ref14","doi-asserted-by":"crossref","first-page":"8027","DOI":"10.1016\/j.eswa.2014.07.008","article-title":"A robust arbitrary text detection system for natural scene images,\u2013","volume":"41","author":"Anhar Risnumawan","year":"2014","journal-title":"Expert Systems with Applications"},{"key":"10.3233\/JIFS-237135_ref16","doi-asserted-by":"crossref","first-page":"5585","DOI":"10.1109\/TIP.2022.3197981","article-title":"Petr: Rethinking the capability of transformer-based language model in scene text recognition","volume":"31","author":"Yuxin Wang","year":"2022","journal-title":"IEEE Transactions on Image Processing"},{"key":"10.3233\/JIFS-237135_ref19","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N. Gomez , \u0141ukasz Kaiser , Illia Polosukhin Attention is all you need, Advances in neural information processing systems 30 (2017)."},{"key":"10.3233\/JIFS-237135_ref21","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s11263-015-0823-z","article-title":"Reading text in the wild with convolutional neural networks","volume":"116","author":"Max Jaderberg","year":"2016","journal-title":"International journal of computer vision"},{"issue":"11","key":"10.3233\/JIFS-237135_ref22","first-page":"2298","article-title":"An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition","volume":"39","author":"Baoguang Shi","year":"2016","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"10.3233\/JIFS-237135_ref24","doi-asserted-by":"crossref","first-page":"397","DOI":"10.1016\/j.patcog.2016.10.016","article-title":"Accurate recognition of words in scenes without character segmentation using recurrent neural network","volume":"63","author":"Bolan Su","year":"2017","journal-title":"Pattern Recognition"},{"issue":"07","key":"10.3233\/JIFS-237135_ref25","doi-asserted-by":"crossref","first-page":"11005","DOI":"10.1609\/aaai.v34i07.6735","article-title":"Gtc: Guided training of CTC towards efficient and accurate scene text recognition, in","volume":"34","author":"Wenyang Hu","year":"2020","journal-title":"Proceedings of the AAAI conference on artificial intelligence"},{"issue":"01","key":"10.3233\/JIFS-237135_ref26","doi-asserted-by":"crossref","first-page":"8714","DOI":"10.1609\/aaai.v33i01.33018714","article-title":"Scene text recognition from two-dimensional perspective, in","volume":"33","author":"Minghui Liao","year":"2019","journal-title":"Proceedings of the AAAI conference on artificial intelligence"},{"issue":"9","key":"10.3233\/JIFS-237135_ref30","first-page":"2035","article-title":"Aster: An attentional scene text recognizer with flexible rectification","volume":"41","author":"Baoguang Shi","year":"2018","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"10.3233\/JIFS-237135_ref42","first-page":"1429","article-title":"Icdarcompetition on reading Chinese text in the wild (rctw-17), in","volume":"1","author":"Baoguang Shi","year":"2017","journal-title":"2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR)"},{"key":"10.3233\/JIFS-237135_ref43","first-page":"5","article-title":"Uber-text: A large-scale dataset for optical character recognition from street-level imagery, in","volume":"2017","author":"Ying Zhang","year":"2017","journal-title":"SUNw: Scene Understanding Workshop-CVPR"},{"issue":"3","key":"10.3233\/JIFS-237135_ref48","first-page":"18","volume":"2","author":"Ivan Krasin","year":"2017","journal-title":"Openimages: A public dataset for large-scale multi-label and multi-class image classification"},{"issue":"18","key":"10.3233\/JIFS-237135_ref52","doi-asserted-by":"crossref","first-page":"8027","DOI":"10.1016\/j.eswa.2014.07.008","article-title":"A robust arbitrary text detection system for natural scene images","volume":"41","author":"Anhar Risnumawan","year":"2014","journal-title":"Expert Systems with Applications"},{"key":"10.3233\/JIFS-237135_ref57","first-page":"369","article-title":"Super-convergence: Very fast training of neural networks using large learning rates, in","volume":"11006","author":"Leslie Smith","year":"2019","journal-title":"Artificial Intelligence and Machine Learning for Multi-Domain Operations Applications"}],"container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"original-title":[],"link":[{"URL":"https:\/\/content.iospress.com\/download?id=10.3233\/JIFS-237135","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:43:09Z","timestamp":1777455789000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/full\/10.3233\/JIFS-237135"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,18]]},"references-count":16,"journal-issue":{"issue":"4"},"URL":"https:\/\/doi.org\/10.3233\/jifs-237135","relation":{},"ISSN":["1064-1246","1875-8967"],"issn-type":[{"value":"1064-1246","type":"print"},{"value":"1875-8967","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,18]]}}}