{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T10:25:12Z","timestamp":1779359112690,"version":"3.51.4"},"reference-count":100,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2024,9,9]],"date-time":"2024-09-09T00:00:00Z","timestamp":1725840000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Deanship of Scientific Research (DSR) at King Abdulaziz University, Jeddah","award":["1177-612-2024"],"award-info":[{"award-number":["1177-612-2024"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["BDCC"],"abstract":"<jats:p>Text localization and recognition from natural scene images has gained a lot of attention recently due to its crucial role in various applications, such as autonomous driving and intelligent navigation. However, two significant gaps exist in this area: (1) prior research has primarily focused on recognizing English text, whereas Arabic text has been underrepresented, and (2) most prior research has adopted separate approaches for scene text localization and recognition, as opposed to one integrated framework. To address these gaps, we propose a novel bilingual end-to-end approach that localizes and recognizes both Arabic and English text within a single natural scene image. Specifically, our approach utilizes pre-trained CNN models (ResNet and EfficientNetV2) with kernel representation for localization text and RNN models (LSTM and BiLSTM) with an attention mechanism for text recognition. In addition, the AraElectra Arabic language model was incorporated to enhance Arabic text recognition. Experimental results on the EvArest, ICDAR2017, and ICDAR2019 datasets demonstrated that our model not only achieves superior performance in recognizing horizontally oriented text but also in recognizing multi-oriented and curved Arabic and English text in natural scene images.<\/jats:p>","DOI":"10.3390\/bdcc8090117","type":"journal-article","created":{"date-parts":[[2024,9,9]],"date-time":"2024-09-09T09:21:00Z","timestamp":1725873660000},"page":"117","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["An End-to-End Scene Text Recognition for Bilingual Text"],"prefix":"10.3390","volume":"8","author":[{"given":"Bayan M.","family":"Albalawi","sequence":"first","affiliation":[{"name":"Department of Computer Science, Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah 21589, Saudi Arabia"},{"name":"Department of Computer Science, Faculty of Computers and Information Technology, University of Tabuk, Tabuk 71491, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Amani T.","family":"Jamal","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah 21589, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lama A.","family":"Al Khuzayem","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah 21589, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5883-4479","authenticated-orcid":false,"given":"Olaa A.","family":"Alsaedi","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah 21589, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,9,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Wang, C., Bochkovskiy, A., and Liao, H. (2022). YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv.","DOI":"10.1109\/CVPR52729.2023.00721"},{"key":"ref_2","unstructured":"Wang, W., Xie, E., Song, X., Zang, Y., Wang, W., Lu, T., Yu, G., and Shen, C. (November, January 27). Efficient and accurate arbitrary-shaped text detection with pixel aggregation network. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1016\/j.patcog.2019.01.020","article-title":"Moran: A multi-object rectified attention network for scene text recognition","volume":"90","author":"Luo","year":"2019","journal-title":"Pattern Recognit."},{"key":"ref_4","first-page":"449","article-title":"A bilingual text detection in natural images using heuristic and unsupervised learning","volume":"10","author":"Bayatpour","year":"2022","journal-title":"J. AI Data Min."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Huang, M., Liu, Y., Peng, Z., Liu, C., Lin, D., Zhu, S., Yuan, N., Ding, K., and Jin, L. (2022, January 18\u201324). Swintextspotter: Scene text spotting via better synergy between text detection and text recognition. Proceedings of the IEEE\/CVF Conference on Compute Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00455"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"3239","DOI":"10.1007\/s10462-020-09930-6","article-title":"Deep learning approaches to scene text detection: A comprehensive review","volume":"54","author":"Khan","year":"2021","journal-title":"Artif. Intell. Rev."},{"key":"ref_7","first-page":"178","article-title":"Deep neural networks combined with STN for multi-oriented text detection and recognition","volume":"11","author":"Katper","year":"2020","journal-title":"Int. J. Adv. Comput. Sci. Appl."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Yao, C., Zhang, X., Bai, X., Liu, W., Ma, Y., and Tu, Z. (2013). Rotation-invariant features for multi-oriented text detection in natural images. PLoS ONE, 8.","DOI":"10.1371\/journal.pone.0070173"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Ranjitha, P., and Rajashekar, K. (2020, January 2\u20134). A Review on text detection from multi-oriented text images in different approaches. Proceedings of the 2020 International Conference on Electronics and Sustainable Communication Systems (ICESC), IEEE, Coimbatore, India.","DOI":"10.1109\/ICESC48915.2020.9156002"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"107046","DOI":"10.1109\/ACCESS.2021.3100717","article-title":"Arabic scene text recognition in the deep learning era: Analysis on a novel dataset","volume":"9","author":"Hassan","year":"2021","journal-title":"IEEE Access"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"7266","DOI":"10.1109\/TPAMI.2021.3095916","article-title":"Towards end-to-end text spotting in natural scenes","volume":"44","author":"Wang","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Ahmed, S.B., Razzak, M.I., and Yusof, R. (2020). Cursive Script Text Recognition in Natural Scene Images, Springer.","DOI":"10.1007\/978-981-15-1297-1"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"367","DOI":"10.1016\/j.ipm.2017.08.004","article-title":"Approaches for preserving content integrity of sensitive online Arabic content: A survey and research challenges","volume":"56","author":"Hakak","year":"2019","journal-title":"Inf. Process. Manag."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"102121","DOI":"10.1016\/j.ipm.2019.102121","article-title":"Arabic text classification using deep learning models","volume":"57","author":"Elnagar","year":"2020","journal-title":"Inf. Process. Manag."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"9943","DOI":"10.1007\/s13369-021-06363-3","article-title":"Arabic handwritten recognition using deep learning: A survey","volume":"47","author":"Alrobah","year":"2022","journal-title":"Arab. J. Sci. Eng."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"387","DOI":"10.1016\/j.jesit.2016.07.005","article-title":"Using features of local densities, statistics and HMM toolkit (HTK) for offline Arabic handwriting text recognition","volume":"4","author":"Hicham","year":"2017","journal-title":"J. Electr. Syst. Inf. Technol."},{"key":"ref_17","first-page":"209896493","article-title":"Handwritten Arabic text recognition using principal component analysis and support vector machines","volume":"10","author":"Aloun","year":"2019","journal-title":"Int. J. Adv. Comput. Sci. Appl."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"89882","DOI":"10.1109\/ACCESS.2020.2994248","article-title":"Exploring deep learning approaches to recognize handwritten Arabic texts","volume":"8","author":"Eltay","year":"2020","journal-title":"IEEE Access"},{"key":"ref_19","first-page":"211029354","article-title":"A deep learning approach for handwritten Arabic names recognition","volume":"11","author":"Mustafa","year":"2020","journal-title":"Int. J. Adv. Comput. Sci. Appl."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Eltay, M., Zidouri, A., Ahmad, I., and Elarian, Y. (2022). Generative adversarial network based adaptive data augmentation for handwritten Arabic text recognition. PeerJ Comput. Sci., 8.","DOI":"10.7717\/peerj-cs.861"},{"key":"ref_21","first-page":"5349","article-title":"Pan++: Towards efficient and accurate end-to-end spotting of arbitrarily-shaped text","volume":"44","author":"Wang","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"3011","DOI":"10.1007\/s00521-020-05137-6","article-title":"Automatic recognition of handwritten Arabic characters: A comprehensive review","volume":"33","author":"Balaha","year":"2021","journal-title":"Neural Comput. Appl."},{"key":"ref_23","first-page":"42","article-title":"Text recognition in the wild: A survey","volume":"54","author":"Chen","year":"2021","journal-title":"ACM Comput. Surv."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"433","DOI":"10.1007\/s11831-019-09315-1","article-title":"Review of scene text detection and recognition","volume":"27","author":"Lin","year":"2020","journal-title":"Arch. Comput. Methods Eng."},{"key":"ref_25","unstructured":"Neumann, L., and Matas, J. (2010, January 8\u201312). A method for text localization and recognition in real-world images. Proceedings of the Computer Vision\u2013ACCV 2010: 10th Asian Conference on Computer Vision, Queenstown, New Zealand. Revised Selected Papers, Part III 10."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Epshtein, B., Ofek, E., and Wexler, Y. (2010, January 13\u201318). Detecting text in natural scenes with stroke width transform. Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, IEEE, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5540041"},{"key":"ref_27","first-page":"800","article-title":"A hybrid approach to detect and localize texts in natural scene images","volume":"20","author":"Pan","year":"2010","journal-title":"IEEE Trans. Image Process."},{"key":"ref_28","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015, January 7\u201312). Faster R-CNN: Towards real-time object detection with region proposal networks. Proceedings of the Advances in Neural Information Processing Systems 28 (NIPS 2015), Montreal, QC, Canada."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhou, X., Yao, C., Wen, H., Wang, Y., Zhou, S., He, W., and Liang, J. (2017, January 21\u201326). EAST: An efficient and accurate scene text detector. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.283"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016, January 11\u201314). SSD: Single shot multibox detector. Proceedings of the Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands. Proceedings, Part I 14.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"3676","DOI":"10.1109\/TIP.2018.2825107","article-title":"Textboxes++: A single-shot oriented scene text detector","volume":"27","author":"Liao","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Tian, Z., Huang, W., He, T., He, P., and Qiao, Y. (2016, January 11\u201314). Detecting text in natural image with connectionist text proposal network. Proceedings of the Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands. Proceedings, Part VIII 14.","DOI":"10.1007\/978-3-319-46484-8_4"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhang, C., Liang, B., Huang, Z., En, M., Han, J., Ding, E., and Ding, X. (2019, January 15\u201320). Look more than once: An accurate detector for text of arbitrary shapes. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01080"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"855","DOI":"10.1109\/TPAMI.2008.137","article-title":"A novel connectionist system for unconstrained handwriting recognition","volume":"31","author":"Graves","year":"2008","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Graves, A., Fernandez, S., Gomez, F., and Schmidhuber, J. (2006, January 25\u201329). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. Proceedings of the 23rd International Conference on Machine Learning, Pittsburgh, PA, USA.","DOI":"10.1145\/1143844.1143891"},{"key":"ref_36","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Tounsi, M., Moalla, I., Alimi, A.M., and Lebouregois, F. (2015, January 23\u201326). Arabic characters recognition in natural scenes using sparse coding for feature representations. Proceedings of the 2015 13th International Conference on Document Analysis and Recognition (ICDAR), IEEE, Tunis, Tunisia.","DOI":"10.1109\/ICDAR.2015.7333919"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Tounsi, M., Moalla, I., and Alimi, A.M. (2017, January 3\u20135). ARASTI: A database for arabic scene text recognition. Proceedings of the 2017 1st International Workshop on Arabic Script Analysis and Recognition (ASAR), IEEE, Nancy, France.","DOI":"10.1109\/ASAR.2017.8067776"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Nayef, N., Yin, F., Bizid, I., Choi, H., Feng, Y., Karatzas, D., Luo, Z., Pal, U., Rigaud, C., and Chazalon, J. (2017, January 9\u201315). ICDAR2017 robust reading challenge on multi-lingual scene text detection and script identification-RRC-MLT. Proceedings of the 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), IEEE, Kyoto, Japan.","DOI":"10.1109\/ICDAR.2017.237"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"19801","DOI":"10.1109\/ACCESS.2019.2895876","article-title":"A novel dataset for English-Arabic scene text recognition (EASTR)-42K and its evaluation using invariant feature extraction on detected extremal regions","volume":"7","author":"Ahmed","year":"2019","journal-title":"IEEE Access"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Nayef, N., Patel, Y., Busta, M., Chowdhury, P.N., Karatzas, D., Khlif, W., Matas, J., Pal, U., Burie, J.C., and Liu, C.l. (2019, January 20\u201325). ICDAR2019 robust reading challenge on multi-lingual scene text detection and recognition\u2013RRC-MLT-2019. Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR), IEEE, Sydney, Australia.","DOI":"10.1109\/ICDAR.2019.00254"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"3026","DOI":"10.1109\/TITS.2020.3029451","article-title":"ASAYAR: A dataset for Arabic\u2013Latin scene text localization in highway traffic panels","volume":"23","author":"Akallouch","year":"2020","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_43","first-page":"1634","article-title":"Real-time Arabic scene text detection using fully convolutional neural networks","volume":"11","author":"Moumen","year":"2021","journal-title":"Int. J. Electr. Comput. Eng."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"93937","DOI":"10.1109\/ACCESS.2021.3092821","article-title":"ATTICA: A dataset for Arabic text-based traffic panels detection","volume":"9","author":"Boujemaa","year":"2021","journal-title":"IEEE Access"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1016\/j.patrec.2022.03.016","article-title":"Reduced annotation based on deep active learning for Arabic text detection in natural scene images","volume":"157","author":"Boukthir","year":"2022","journal-title":"Pattern Recognit. Lett."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Gaddour, H., Kanoun, S., and Vincent, N. (June, January 30). A new method for arabic text detection in natural scene image based on the color homogeneity. Proceedings of the Image and Signal Processing: 7th International Conference, ICISP 2016, Trois-Rivi\u00e8res, QC, Canada. Proceedings 7.","DOI":"10.1007\/978-3-319-33618-3_14"},{"key":"ref_47","first-page":"94","article-title":"Active Deep Learning Reduces Annotation Burden in Automatic Cell Segmentation","volume":"Volume 11603","author":"Chowdhury","year":"2021","journal-title":"Proceedings of the Medical Imaging 2021: Digital Pathology"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Yang, L., Zhang, Y., Chen, J., Zhang, S., and Chen, D.Z. (2017, January 11\u201313). Suggestive annotation: A deep active learning framework for biomedical image segmentation. Proceedings of the Medical Image Computing and Computer Assisted Intervention\u2013MICCAI 2017: 20th International Conference, Quebec City, QC, Canada. Proceedings, Part III 20.","DOI":"10.1007\/978-3-319-66179-7_46"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Liao, M., Shi, B., Bai, X., Wang, X., and Liu, W. (2017, January 4\u20139). Textboxes: A fast text detector with a single deep neural network. Proceedings of the AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.11196"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Shi, B., Bai, X., and Belongie, S. (2017, January 21\u201326). Detecting oriented text in natural images by linking segments. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.371"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Wang, W., Xie, E., Li, X., Hou, W., Lu, T., Yu, G., and Shao, S. (2019, January 15\u201320). Shape robust text detection with progressive scale expansion network. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00956"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Baek, Y., Lee, B., Han, D., Yun, S., and Lee, H. (2019, January 15\u201320). Character region awareness for text detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00959"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Dai, P., Zhang, S., Zhang, H., and Cao, X. (2021, January 20\u201325). Progressive contour regression for arbitrary-shape scene text detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00731"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Ye, M., Zhang, J., Zhao, S., Liu, J., Du, B., and Tao, D. (2023, January 7\u201314). Dptext-detr: Towards better scene text detection with dynamic points in transformer. Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA.","DOI":"10.1609\/aaai.v37i3.25430"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Ahmed, S.B., Naz, S., Razzak, M.I., and Yousaf, R. (2017, January 3\u20135). Deep learning based isolated arabic scene character recognition. Proceedings of the 2017 1st International Workshop on Arabic Script Analysis and Recognition (ASAR), IEEE, Nancy, France.","DOI":"10.1109\/ASAR.2017.8067758"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Jain, M., Mathew, M., and Jawahar, C. (2017, January 3\u20135). Unconstrained scene text and video text recognition for rabic script. Proceedings of the 2017 1st International Workshop on Arabic Script Analysis and Recognition (ASAR), IEEE, Nancy, France.","DOI":"10.1109\/ASAR.2017.8067754"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Alsaeedi, A., Al Mutawa, H., Snoussi, S., Natheer, S., Omri, K., and Al Subhi, W. (2018, January 12\u201314). Arabic words recognition using CNN and TNN on a smartphone. Proceedings of the 2018 IEEE 2nd International Workshop on Arabic and Derived Script Analysis and Recognition (ASAR), IEEE, London, UK.","DOI":"10.1109\/ASAR.2018.8480267"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Ahmed, S.B., Naz, S., Razzak, I., and Prasad, M. (2020, January 19\u201324). Unconstrained arabic scene text analysis using concurrent invariant points. Proceedings of the 2020 International Joint Conference on Neural Networks (IJCNN), IEEE, Glasgow, UK.","DOI":"10.1109\/IJCNN48605.2020.9207283"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Bissacco, A., Cummins, M., Netzer, Y., and Neven, H. (2013, January 2\u20138). Photoocr: Reading text in uncontrolled conditions. Proceedings of the IEEE International Conference on Computer Vision, Washington, DC, USA.","DOI":"10.1109\/ICCV.2013.102"},{"key":"ref_60","unstructured":"Liu, W., Chen, C., Wong, K.Y.K., Su, Z., and Han, J. (2016, January 19\u201322). Star-net: A spatial attention residue network for scene text recognition. Proceedings of the BMVC, York, UK."},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"2298","DOI":"10.1109\/TPAMI.2016.2646371","article-title":"An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition","volume":"39","author":"Shi","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_62","unstructured":"Wang, J., and Hu, X. (2017, January 4\u20139). Gated recurrent convolution neural network for ocr. Proceedings of the Advances in Neural Information Processing Systems 30 (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Borisyuk, F., Gordo, A., and Sivakumar, V. (2018, January 19\u201323). Rosetta: Large scale system for text detection and recognition in images. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, London, UK.","DOI":"10.1145\/3219819.3219861"},{"key":"ref_64","unstructured":"Shi, B., Wang, X., Lyu, P., Yao, C., and Bai, X. (July, January 26). Robust scene text recognition with automatic rectification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_65","unstructured":"Lee, C.Y., and Osindero, S. (July, January 26). Recursive recurrent nets with attention modeling for ocr in the wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"Zhan, F., and Lu, S. (2019, January 15\u201320). Esir: End-to-end scene text recognition via iterative image rectification. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00216"},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Hassan, H., Torki, M., and Hussein, M.E. (2021, January 8\u201310). SCAN: Sequence-character aware network for text recognition. Proceedings of the 16th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications (VISIGRAPP 2021), Vienna, Austria.","DOI":"10.5220\/0010321106020609"},{"key":"ref_68","doi-asserted-by":"crossref","unstructured":"Cheng, C., Wang, P., Da, C., Zheng, Q., and Yao, C. (2023, January 4\u20136). LISTER: Neighbor decoding for length-insensitive scene text recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.01790"},{"key":"ref_69","first-page":"8048","article-title":"Abcnet v2: Adaptive bezier-curve network for real-time end-to-end text spotting","volume":"44","author":"Liu","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_70","doi-asserted-by":"crossref","unstructured":"Zhang, X., Su, Y., Tripathi, S., and Tu, Z. (2022, January 18\u201324). Text spotting transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00930"},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"Kittenplon, Y., Lavi, I., Fogel, S., Bar, Y., Manmatha, R., and Perona, P. (2022, January 18\u201324). Towards weakly-supervised text spotting using a multi-task transformer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00456"},{"key":"ref_72","doi-asserted-by":"crossref","unstructured":"Huang, M., Zhang, J., Peng, D., Lu, H., Huang, C., Liu, Y., Bai, X., and Jin, L. (2023, January 2\u20133). Estextspotter: Towards better scene text spotting with explicit synergy in transformer. Proceedings of the IEEE\/CVF International Conference on Computer 1446 Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.01786"},{"key":"ref_73","doi-asserted-by":"crossref","unstructured":"Kil, T., Kim, S., Seo, S., Kim, Y., and Kim, D. (2023, January 17\u201324). Towards unified scene text spotting based on sequence generation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01461"},{"key":"ref_74","doi-asserted-by":"crossref","unstructured":"Ye, M., Zhang, J., Zhao, S., Liu, J., Liu, T., Du, B., and Tao, D. (2023, January 17\u201324). Deepsolo: Let transformer decoder with explicit points solo for text spotting. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01854"},{"key":"ref_75","doi-asserted-by":"crossref","unstructured":"Das, A., Biswas, S., Banerjee, A., Llad\u00f3s, J., Pal, U., and Bhattacharya, S. (2024, January 1\u20136). Harnessing the power of multi-lingual datasets for pre-training: Towards enhancing text spotting performance. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV57701.2024.00077"},{"key":"ref_76","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_77","unstructured":"Tan, M., and Le, Q. (2021, January 18\u201324). EfficientNetV2: Smaller models and faster training. Proceedings of the International Conference on Machine Learning, PMLR, Virtual."},{"key":"ref_78","unstructured":"Tan, M., and Le, Q. (2019, January 9\u201315). EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of the International Conference on Machine Learning, PMLR, Long Beach, CA, USA."},{"key":"ref_79","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.C. (2018, January 18\u201323). MobileNetV2: Inverted residuals and linear bottlenecks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_80","doi-asserted-by":"crossref","unstructured":"Tan, M., Chen, B., Pang, R., Vasudevan, V., Sandler, M., Howard, A., and Le, Q.V. (2019, January 15\u201320). MNASNet: Platform-aware neural architecture search for mobile. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00293"},{"key":"ref_81","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_82","unstructured":"Sifre, L., and Mallat, S. (2014). Rigid-motion scattering for texture classification. arXiv."},{"key":"ref_83","unstructured":"Gupta, S., and Tan, M. (2019). EfficientNet-EdgeTPU: Creating Accelerator-Optimized Neural Networks with AutoML, Google AI Blog."},{"key":"ref_84","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_85","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_86","unstructured":"Le, Q.V., Jaitly, N., and Hinton, G.E. (2015). A simple way to initialize recurrent networks of rectified linear units. arXiv."},{"key":"ref_87","unstructured":"Salehinejad, H., Sankar, S., Barfett, J., Colak, E., and Valaee, S. (2017). Recent advances in recurrent neural networks. arXiv."},{"key":"ref_88","doi-asserted-by":"crossref","first-page":"2673","DOI":"10.1109\/78.650093","article-title":"Bidirectional recurrent neural networks","volume":"45","author":"Schuster","year":"1997","journal-title":"IEEE Trans. Signal Process."},{"key":"ref_89","doi-asserted-by":"crossref","unstructured":"Sun, S., Sun, J., Wang, Z., Zhou, Z., and Cai, W. (2022). Prediction of battery SOH by CNN-BiLSTM network fused with attention mechanism. Energies, 15.","DOI":"10.3390\/en15124428"},{"key":"ref_90","doi-asserted-by":"crossref","unstructured":"Adil, M., Wu, J.Z., Chakrabortty, R.K., Alahmadi, A., Ansari, M.F., and Ryan, M.J. (2021). Attention-based STL-BiLSTM network to forecast tourist arrival. Processes, 9.","DOI":"10.3390\/pr9101759"},{"key":"ref_91","unstructured":"Clark, K., Luong, M.T., Le, Q.V., and Manning, C.D. (2020). Electra: Pre-training text encoders as discriminators rather than generators. arXiv."},{"key":"ref_92","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014, January 8\u201313). Generative adversarial nets. Proceedings of the Advances in Neural Information Processing Systems 27 (NIPS 2014), Montreal, QC, Canada."},{"key":"ref_93","unstructured":"Antoun, W., Baly, F., and Hajj, H. (2020). AraELECTRA: Pre-training text discriminators for Arabic language understanding. arXiv."},{"key":"ref_94","unstructured":"Antoun, W., Baly, F., and Hajj, H. (2020). Arabert: Transformer-based model for rabic language understanding. arXiv."},{"key":"ref_95","doi-asserted-by":"crossref","unstructured":"Karatzas, D., Shafait, F., Uchida, S., Iwamura, M., I Bigorda, L.G., Mestre, S.R., Mas, J., Mota, D.F., Almazan, J.A., and De Las Heras, L.P. (2013, January 25\u201328). ICDAR 2013 robust reading competition. Proceedings of the 2013 12th International Conference on Document Analysis and Recognition, IEEE, Washington, DC, USA.","DOI":"10.1109\/ICDAR.2013.221"},{"key":"ref_96","doi-asserted-by":"crossref","unstructured":"Karatzas, D., Gomez-Bigorda, L., Nicolaou, A., Ghosh, S., Bagdanov, A., Iwamura, M., Matas, J., Neumann, L., Chandrasekhar, V.R., and Lu, S. (2015, January 23\u201326). ICDAR 2015 competition on robust reading. Proceedings of the 2015 13th International Conference on Document Analysis and Recognition (ICDAR), IEEE, Tunis, Tunisia.","DOI":"10.1109\/ICDAR.2015.7333942"},{"key":"ref_97","unstructured":"Veit, A., Matera, T., Neumann, L., Matas, J., and Belongie, S. (2016). Coco-text: Ddtaset and benchmark for text detection and recognition in natural images. arXiv."},{"key":"ref_98","doi-asserted-by":"crossref","unstructured":"Ch\u2019ng, C.K., and Chan, C.S. (2017, January 9\u201315). Total-text: A comprehensive dataset for scene text detection and recognition. Proceedings of the 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), IEEE, Kyoto, Japan.","DOI":"10.1109\/ICDAR.2017.157"},{"key":"ref_99","doi-asserted-by":"crossref","first-page":"98","DOI":"10.1007\/s11263-014-0733-5","article-title":"The Pascal visual object classes challenge: A retrospective","volume":"111","author":"Everingham","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_100","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."}],"container-title":["Big Data and Cognitive Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2504-2289\/8\/9\/117\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:52:14Z","timestamp":1760111534000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2504-2289\/8\/9\/117"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,9]]},"references-count":100,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2024,9]]}},"alternative-id":["bdcc8090117"],"URL":"https:\/\/doi.org\/10.3390\/bdcc8090117","relation":{},"ISSN":["2504-2289"],"issn-type":[{"value":"2504-2289","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,9,9]]}}}