{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,17]],"date-time":"2026-01-17T12:30:29Z","timestamp":1768653029776,"version":"3.49.0"},"reference-count":61,"publisher":"Oxford University Press (OUP)","issue":"12","license":[{"start":{"date-parts":[[2024,9,18]],"date-time":"2024-09-18T00:00:00Z","timestamp":1726617600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/pages\/standard-publication-reuse-rights"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2024,12,20]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Recent advancements in scene text recognition have predominantly focused on leveraging textual semantics. However, an over-reliance on linguistic priors can impede a model\u2019s ability to handle irregular text scenes, including non-standard word usage, occlusions, severe distortions, or stretching. The key challenges lie in effectively localizing occlusions, perceiving multi-scale text, and inferring text based on scene context. To address these challenges and enhance visual capabilities, we introduce the Graph Reasoning Model (GRM). The GRM employs a novel feature fusion method to align spatial context information across different scales, beginning with a feature aggregation stage that extracts rich spatial contextual information from various feature maps. Visual reasoning representations are then obtained through graph convolution. We integrate the GRM module with a language model to form a two-stream architecture called GRNet. This architecture combines pure visual predictions with joint visual-linguistic predictions to produce the final recognition results. Additionally, we propose a dynamic iteration refinement for the language model to prevent over-correction of prediction results, ensuring a balanced contribution from both visual and linguistic cues. Extensive experiments demonstrate that GRNet achieves state-of-the-art average recognition accuracy across six mainstream benchmarks. These results highlight the efficacy of our multi-modal approach in scene text recognition, particularly in challenging scenarios where visual reasoning plays a crucial role.<\/jats:p>","DOI":"10.1093\/comjnl\/bxae085","type":"journal-article","created":{"date-parts":[[2024,9,19]],"date-time":"2024-09-19T08:02:48Z","timestamp":1726732968000},"page":"3239-3250","source":"Crossref","is-referenced-by-count":1,"title":["GRNet: a graph reasoning network for enhanced multi-modal learning in scene text recognition"],"prefix":"10.1093","volume":"67","author":[{"given":"Zeguang","family":"Jia","sequence":"first","affiliation":[{"name":"School of Control Science and Engineering , Tiangong University, Tianjin, 300387,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianming","family":"Wang","sequence":"additional","affiliation":[{"name":"Tianjin Key Laboratory Autonomous Intelligence Technology and Systems , Tiangong University, Tianjin, 300387,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rize","family":"Jin","sequence":"additional","affiliation":[{"name":"School of Software , Tiangong University, Tianjin, 300387,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2024,9,18]]},"reference":[{"key":"2025010523430708700_ref1","doi-asserted-by":"crossref","first-page":"7789","DOI":"10.1109\/JIOT.2020.3039359","article-title":"Empowering things with intelligence: a survey of the progress, challenges, and opportunities in artificial intelligence of things","volume":"8","author":"Zhang","year":"2020","journal-title":"IEEE Internet Things J"},{"key":"2025010523430708700_ref2","first-page":"2950","article-title":"Fine-grained image classification and retrieval by combining visual and locally pooled textual features","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Snowmass Village, USA, 2\u20135, March","author":"Mafla","year":"2020"},{"key":"2025010523430708700_ref3","first-page":"1693","article-title":"Ontological supervision for fine grained classification of street view storefronts","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7\u201312, June","author":"Movshovitz-Attias","year":"2015"},{"key":"2025010523430708700_ref4","first-page":"4023","article-title":"Multi-modal reasoning graph for scene-text based fine-grained image classification and retrieval","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Mafla","year":"2021"},{"key":"2025010523430708700_ref5","doi-asserted-by":"publisher","first-page":"297","DOI":"10.1109\/TAI.2021.3116216","article-title":"Character-level street view text spotting based on deep multisegmentation network for smarter autonomous driving","volume":"3","author":"Zhang","year":"2021","journal-title":"IEEE Trans Artif Intell"},{"key":"2025010523430708700_ref6","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s11263-015-0823-z","article-title":"Reading text in the wild with convolutional neural networks","volume":"116","author":"Jaderberg","year":"2016","journal-title":"Int J Comput Vis"},{"key":"2025010523430708700_ref7","first-page":"4715","article-title":"What is wrong with scene text recognition model comparisons? Dataset and model analysis","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea, 27 October\u20132 November","author":"Baek","year":"2019"},{"key":"2025010523430708700_ref8","first-page":"8610","article-title":"Show, attend and read: a simple and strong baseline for irregular text recognition","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence, Hawaii, USA 27, January\u20141, February","author":"Li","year":"2019"},{"key":"2025010523430708700_ref9","doi-asserted-by":"crossref","first-page":"66322","DOI":"10.1109\/ACCESS.2018.2878899","article-title":"Integrating scene text and visual appearance for fine-grained image classification","volume":"6","author":"Bai","year":"2018","journal-title":"IEEE Access"},{"key":"2025010523430708700_ref10","first-page":"844","article-title":"Attention-based extraction of structured information from street view imagery","volume-title":"2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), Kyoto, Japan,10\u201315 November","author":"Wojna","year":"2017"},{"key":"2025010523430708700_ref11","first-page":"4558","article-title":"Scene text retrieval via joint text detection and similarity learning","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition,Long Beach,USA,10\u201325 June","author":"Wang","year":"2021"},{"key":"2025010523430708700_ref12","doi-asserted-by":"crossref","first-page":"1063","DOI":"10.1109\/TMM.2016.2638622","article-title":"Words matter: Scene text for image classification and retrieval","volume":"19","author":"Karaoglu","year":"2016","journal-title":"IEEE Trans Multimed"},{"key":"2025010523430708700_ref13","first-page":"9147","article-title":"Symmetry-constrained rectification network for scene text recognition","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision,Seoul, Korea, 27 October,2 November","author":"Yang","year":"2019"},{"key":"2025010523430708700_ref14","first-page":"3","article-title":"Learning to read irregular text with attention mechanisms","volume-title":"IJCAI,Melbourne, Australia,19\u201325 August","author":"Yang","year":"2017"},{"key":"2025010523430708700_ref15","first-page":"2231","article-title":"Recursive recurrent nets with attention modeling for ocr in the wild","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition,Las Vegas, USA, 27\u201330 June","author":"Lee","year":"2016"},{"key":"2025010523430708700_ref16","article-title":"Spatial transformer networks","volume":"28","author":"Jaderberg","year":"2015","journal-title":"Adv Neural Inf Process Syst"},{"key":"2025010523430708700_ref17","first-page":"2059","article-title":"Esir: end-to-end scene text recognition via iterative image rectification","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 15\u201320 June","author":"Zhan","year":"2019"},{"key":"2025010523430708700_ref18","first-page":"8714","article-title":"Scene text recognition from two-dimensional perspective","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence,Hawaii, USA,27 January, 1 February","author":"Liao","year":"2019"},{"key":"2025010523430708700_ref19","first-page":"4159","article-title":"Multi-oriented text detection with fully convolutional networks","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, 27\u201330 June","author":"Zhang","year":"2016"},{"key":"2025010523430708700_ref20","first-page":"11425","article-title":"On vocabulary reliance in scene text recognition","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA","author":"Wan","year":"2020"},{"key":"2025010523430708700_ref21","first-page":"3304","article-title":"End-to-end text recognition with convolutional neural networks","volume-title":"Proceedings of the 21st International Conference on Pattern Recognition,Tsukuba Japan 10\u201316 November","author":"Wang","year":"2012"},{"key":"2025010523430708700_ref22","doi-asserted-by":"crossref","first-page":"3538","DOI":"10.1109\/CVPR.2012.6248097","article-title":"Real-time scene text localization and recognition","volume-title":"2012 IEEE Conference on Computer Vision and Pattern recognition, Providence, USA, 16\u201321 June","author":"Neumann","year":"2012"},{"key":"2025010523430708700_ref23","first-page":"35","article-title":"Accurate scene text recognition based on recurrent neural network","volume-title":"12th Asian Conference on Computer Vision, Singapore, 1\u20135 November","author":"Su","year":"2015"},{"key":"2025010523430708700_ref24","volume-title":"Star-Net: A Spatial Attention Residue Network for Scene Text Recognition. BMVC, Scotland, UK, 19\u201322 September 7","author":"Liu","year":"2016"},{"key":"2025010523430708700_ref25","first-page":"67","article-title":"Mask textspotter: an end-to-end trainable neural network for spotting text with arbitrary shapes","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany,8\u201314 September","author":"Lyu","year":"2018"},{"key":"2025010523430708700_ref26","article-title":"2D attentional irregular scene text recognizer","author":"Lyu","year":"2019"},{"key":"2025010523430708700_ref27","first-page":"178","article-title":"Scene text recognition with permuted autoregressive sequence models","volume-title":"European Conference on Computer Vision,Tel Aviv, Israel, 23\u201327 October","author":"Bautista","year":"2022"},{"key":"2025010523430708700_ref28","doi-asserted-by":"crossref","first-page":"512","DOI":"10.1007\/978-3-319-10593-2_34","article-title":"Deep features for text spotting","volume-title":"Computer Vision\u2013ECCV 2014: 13th European Conference, Zurich, Switzerland, 6\u201312 September","author":"Jaderberg","year":"2014"},{"key":"2025010523430708700_ref29","article-title":"Deep structured output learning for unconstrained text recognition","author":"Jaderberg","year":"2014"},{"key":"2025010523430708700_ref30","first-page":"5076","article-title":"Focusing attention: towards accurate text recognition in natural images","volume-title":"Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22\u201329 October","author":"Cheng","year":"2017"},{"key":"2025010523430708700_ref31","first-page":"12113","article-title":"Towards accurate scene text recognition with semantic reasoning networks","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA,13\u201319 June","author":"Yu","year":"2020"},{"key":"2025010523430708700_ref32","first-page":"13528","article-title":"Seed: Semantics enhanced encoder\u2013decoder framework for scene text recognition","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 13\u201319 June","author":"Qiao","year":"2020"},{"key":"2025010523430708700_ref33","first-page":"28522","article-title":"Vitae: vision transformer advanced by exploring intrinsic inductive bias","volume":"34","author":"Xu","year":"2021","journal-title":"Adv Neural Inf Process Syst"},{"key":"2025010523430708700_ref34","first-page":"7098","article-title":"Read like humans: autonomous, bidirectional and iterative language modeling for scene text recognition","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 10\u201325 June","author":"Fang","year":"2021"},{"key":"2025010523430708700_ref35","first-page":"433","article-title":"Graph-based global reasoning networks","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 15\u201320 June","author":"Chen","year":"2019"},{"key":"2025010523430708700_ref36","first-page":"4050","article-title":"Region-based discriminative feature pooling for scene text recognition","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, USA, 23\u201328 June","author":"Lee","year":"2014"},{"key":"2025010523430708700_ref37","article-title":"Graph attention networks","author":"Veli\u010dkovi\u0107","year":"2017"},{"key":"2025010523430708700_ref38","first-page":"11005","article-title":"Gtc: guided training of ctc towards efficient and accurate scene text recognition","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence, New York USA, 7\u201312 February","author":"Hu","year":"2020"},{"key":"2025010523430708700_ref39","article-title":"Inductive representation learning on large graphs","volume":"30","author":"Hamilton","year":"2017","journal-title":"Adv Neural Inf Process Syst"},{"key":"2025010523430708700_ref40","first-page":"284","article-title":"Primitive representation learning for scene text recognition","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, 10\u201325 June","author":"Yan","year":"2021"},{"key":"2025010523430708700_ref41","first-page":"888","article-title":"Visual semantics allow for textual reasoning better in scene text recognition","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence, New York, USA, 7\u201312 February","author":"He","year":"2022"},{"key":"2025010523430708700_ref42","doi-asserted-by":"crossref","first-page":"775","DOI":"10.1007\/978-3-030-58452-8_45","article-title":"Semantic flow for fast and accurate scene parsing","volume-title":"Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, 23\u201328 August","author":"Li","year":"2020"},{"key":"2025010523430708700_ref54","first-page":"2315","article-title":"Synthetic data for text localisation in natural images","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, 27\u201330 June","author":"Gupta","year":"2016"},{"key":"2025010523430708700_ref55","first-page":"1","article-title":"Synthetic data and artificial neural networks for natural scene text recognition","volume":"abs\/1406.2227","author":"Jaderberg","year":"2014","journal-title":"ArXiv"},{"key":"2025010523430708700_ref56","first-page":"1484","article-title":"Lcdar robust reading competition","volume-title":"2013 12th International Conference on Document Analysis and Recognition,Washington, USA, 22\u201328 August","author":"Karatzas","year":"2013"},{"key":"2025010523430708700_ref57","doi-asserted-by":"crossref","DOI":"10.5244\/C.26.127","article-title":"Scene text recognition using higher order language priors","volume-title":"BMVC-British Machine Vision Conference, Surrey, UK, 9\u201313 September","author":"Mishra","year":"2012"},{"key":"2025010523430708700_ref58","volume-title":"International Conference on Computer Vision, Barcelona, Spain, 6\u201313 November","author":"Wang","year":"2011"},{"key":"2025010523430708700_ref59","first-page":"1156","article-title":"ICDAR 2015 competition on robust reading","volume-title":"2015 13th International Conference on Document Analysis and Recognition (ICDAR), Tunis, Tunisia, 23\u201326 August","author":"Karatzas","year":"2015"},{"key":"2025010523430708700_ref60","first-page":"569","article-title":"Recognizing text with perspective distortion in natural scenes","volume-title":"Proceedings of the IEEE International Conference on Computer Vision,Sydney, Australia 1\u20138 December","author":"Phan","year":"2013"},{"key":"2025010523430708700_ref61","doi-asserted-by":"publisher","first-page":"8027","DOI":"10.1016\/j.eswa.2014.07.008","article-title":"A robust arbitrary text detection system for natural scene images","volume":"41","author":"Risnumawan","year":"2014","journal-title":"Exp Syst Appl"},{"key":"2025010523430708700_ref43","doi-asserted-by":"crossref","first-page":"2298","DOI":"10.1109\/TPAMI.2016.2646371","article-title":"An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition","volume":"39","author":"Shi","year":"2016","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"2025010523430708700_ref44","doi-asserted-by":"crossref","first-page":"2035","DOI":"10.1109\/TPAMI.2018.2848939","article-title":"Aster: an attentional scene text recognizer with flexible rectification","volume":"41","author":"Shi","year":"2018","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"2025010523430708700_ref45","doi-asserted-by":"crossref","first-page":"751","DOI":"10.1007\/978-3-030-58586-0_44","article-title":"AutoSTR: efficient backbone search for scene text recognition","volume-title":"Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, 23\u201328 August","author":"Zhang","year":"2020"},{"key":"2025010523430708700_ref46","doi-asserted-by":"crossref","first-page":"135","DOI":"10.1007\/978-3-030-58529-7_9","article-title":"Robustscanner: dynamically enhancing positional clues for robust text recognition","volume-title":"Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, 23\u201328 August","author":"Yue","year":"2020"},{"key":"2025010523430708700_ref47","first-page":"12216","article-title":"Decoupled attention network for text recognition","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence, New York, USA, 7\u201312 February","author":"Wang","year":"2020"},{"key":"2025010523430708700_ref48","first-page":"339","article-title":"Multi-granularity prediction for scene text recognition","volume-title":"European Conference on Computer Vision, Tel Aviv, Israel, 23\u201327 October","author":"Wang","year":"2022"},{"key":"2025010523430708700_ref49","first-page":"464","article-title":"SGBANet: semantic GAN and balanced attention network for arbitrarily oriented scene text recognition","volume-title":"European Conference on Computer Vision, Tel Aviv, Israel, 23\u201327 October","author":"Zhong","year":"2022"},{"key":"2025010523430708700_ref50","first-page":"322","article-title":"Levenshtein OCR","volume-title":"European Conference on Computer Vision, Tel Aviv, Israel, 23\u201327 October","author":"Da","year":"2022"},{"key":"2025010523430708700_ref51","first-page":"19541","article-title":"LISTER: neighbor decoding for length-insensitive scene text recognition","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France,2\u20136 October","author":"Cheng","year":"2023"},{"key":"2025010523430708700_ref52","first-page":"15285","article-title":"Self-supervised implicit glyph attention for text recognition","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, Canada, 18\u201322 June","author":"Guan","year":"2023"},{"key":"2025010523430708700_ref53","first-page":"319","article-title":"Vision transformer for fast and efficient scene text recognition","volume-title":"International Conference on Document Analysis and Recognition, Switzerland, 5\u201310 September","author":"Atienza","year":"2021"}],"container-title":["The Computer Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/67\/12\/3239\/59182328\/bxae085.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/67\/12\/3239\/59182328\/bxae085.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,6]],"date-time":"2025-01-06T04:32:26Z","timestamp":1736137946000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/comjnl\/article\/67\/12\/3239\/7760133"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,18]]},"references-count":61,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2024,9,18]]},"published-print":{"date-parts":[[2024,12,20]]}},"URL":"https:\/\/doi.org\/10.1093\/comjnl\/bxae085","relation":{},"ISSN":["0010-4620","1460-2067"],"issn-type":[{"value":"0010-4620","type":"print"},{"value":"1460-2067","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024,12]]},"published":{"date-parts":[[2024,9,18]]}}}