{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T23:24:54Z","timestamp":1783553094321,"version":"3.55.0"},"reference-count":28,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2022,10,3]],"date-time":"2022-10-03T00:00:00Z","timestamp":1664755200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Symbiosis International (Deemed) University"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["JSAN"],"abstract":"<jats:p>Optical Character Recognition has made large strides in the field of recognizing printed and properly formatted text. However, the effort attributed to developing systems that are able to reliably apply OCR to both printed as well as handwritten text simultaneously, such as hand-filled forms, is lackadaisical. As Machine printed\/typed text follows specific formats and fonts while handwritten texts are variable and non-uniform, it is very hard to classify and recognize using traditional OCR only. A pre-processing methodology employing semantic segmentation to identify, segment and crop boxes containing relevant text on a given image in order to improve the results of conventional online-available OCR engines is proposed here. In this paper, the authors have also provided a comparison of popular OCR engines like Microsoft Cognitive Services, Google Cloud Vision and AWS recognitions. We have proposed a pixel-wise classification technique to accurately identify the area of an image containing relevant text, to feed them to a conventional OCR engine in the hopes of improving the quality of the output. The proposed methodology also supports the digitization of mixed typed text documents with amended performance. The experimental study shows that the proposed pipeline architecture provides reliable and quality inputs through complex image preprocessing to Conventional OCR, which results in better accuracy and improved performance.<\/jats:p>","DOI":"10.3390\/jsan11040063","type":"journal-article","created":{"date-parts":[[2022,10,9]],"date-time":"2022-10-09T01:43:11Z","timestamp":1665279791000},"page":"63","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["Enhancing Optical Character Recognition on Images with Mixed Text Using Semantic Segmentation"],"prefix":"10.3390","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4903-1540","authenticated-orcid":false,"given":"Shruti","family":"Patil","sequence":"first","affiliation":[{"name":"Symbiosis Centre for Applied Artificial Intelligence (SCAAI), Symbiosis Institute of Technology, Symbiosis International (Deemed University), Pune 412115, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3752-7220","authenticated-orcid":false,"given":"Vijayakumar","family":"Varadarajan","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, The University of New South Wales, Sydney, NSW 2052, Australia"},{"name":"School of NUOVOS, Ajeenkya D Y Patil University, Pune 412105, India"},{"name":"Swiss School of Business and Management, SSBM Geneva, 1213 Geneva, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3845-2240","authenticated-orcid":false,"given":"Supriya","family":"Mahadevkar","sequence":"additional","affiliation":[{"name":"Ph.D Research Scholar, Symbiosis Institute of Technology, Symbiosis International (Deemed University), Pune 412115, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rohan","family":"Athawade","sequence":"additional","affiliation":[{"name":"Department of Computer Science Engineering, Symbiosis Institute of Technology, Symbiosis International (Deemed University), Pune 412115, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lakhan","family":"Maheshwari","sequence":"additional","affiliation":[{"name":"Department of Computer Science Engineering, Symbiosis Institute of Technology, Symbiosis International (Deemed University), Pune 412115, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shrushti","family":"Kumbhare","sequence":"additional","affiliation":[{"name":"Department of Computer Science Engineering, Symbiosis Institute of Technology, Symbiosis International (Deemed University), Pune 412115, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yash","family":"Garg","sequence":"additional","affiliation":[{"name":"Department of Computer Science Engineering, Symbiosis Institute of Technology, Symbiosis International (Deemed University), Pune 412115, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2540-6942","authenticated-orcid":false,"given":"Deepak","family":"Dharrao","sequence":"additional","affiliation":[{"name":"Department of Computer Science Engineering, Symbiosis Institute of Technology, Symbiosis International (Deemed University), Pune 412115, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7597-0197","authenticated-orcid":false,"given":"Pooja","family":"Kamat","sequence":"additional","affiliation":[{"name":"Department of Computer Science Engineering, Symbiosis Institute of Technology, Symbiosis International (Deemed University), Pune 412115, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2653-3780","authenticated-orcid":false,"given":"Ketan","family":"Kotecha","sequence":"additional","affiliation":[{"name":"Symbiosis Centre for Applied Artificial Intelligence (SCAAI), Symbiosis Institute of Technology, Symbiosis International (Deemed University), Pune 412115, India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,10,3]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Ranjan, A., Behera, V.N.J., and Reza, M. (2021). OCR Using Computer Vision and Machine Learning. Machine Learning Algorithms for Industrial Applications, Springer.","DOI":"10.1007\/978-3-030-50641-4_6"},{"key":"ref_2","unstructured":"(2022, January 05). Available online: http:\/\/www.capturedocs.com\/thread\/handwritten-invoices\/."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Rabby, A.K.M., Islam, M., Hasan, N., Nahar, J., and Rahman, F. (2021). A Deep Learning Solution to Detect Text-Types Using a Convolutional Neural Network. Proceedings of the International Conference on Machine Intelligence and Data Science Applications, Springer.","DOI":"10.1007\/978-981-33-4087-9_58"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"337","DOI":"10.1109\/TPAMI.2004.1262324","article-title":"Machine printed text and handwriting identification in noisy document images","volume":"26","author":"Zheng","year":"2004","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"4412","DOI":"10.35940\/ijitee.J9835.0881019","article-title":"Demystifying User Data Privacy in the World of IOT","volume":"8","author":"Patil","year":"2019","journal-title":"Int. J. Innov. Technol. Explor. Eng."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Bidwe, R.V., Mishra, S., Patil, S., Shaw, K., Vora, D.R., Kotecha, K., and Zope, B. (2022). Deep Learning Approaches for Video Compression: A Bibliometric Analysis. Big Data Cogn. Comput., 6.","DOI":"10.3390\/bdcc6020044"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"72894","DOI":"10.1109\/ACCESS.2021.3072900","article-title":"Efficient Automated Processing of the Unstructured Documents Using Artificial Intelligence: A Systematic Literature Review and Future Directions","volume":"9","author":"Baviskar","year":"2021","journal-title":"IEEE Access"},{"key":"ref_8","first-page":"4798","article-title":"Estimating Remaining Useful Life in Machines Using Artificial Intelligence: A Scoping Review","volume":"2021","author":"Sayyad","year":"2021","journal-title":"Libr. Philos. Pract."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Chaudhuri, A., Mandaviya, K., Badelia, P., and Ghosh, S.K. (2016). Optical Character Recognition Systems. Optical Character Recognition Systems for Different Languages with Soft Computing, Springer.","DOI":"10.1007\/978-3-319-50252-6"},{"key":"ref_10","first-page":"1","article-title":"Text recognition in the wild: A survey","volume":"54","author":"Chen","year":"2021","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"142642","DOI":"10.1109\/ACCESS.2020.3012542","article-title":"Handwritten Optical Character Recognition (OCR): A Comprehensive Systematic Literature Review (SLR)","volume":"8","author":"Memon","year":"2020","journal-title":"IEEE Access"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1007\/s13735-017-0141-z","article-title":"A review of semantic segmentation using deep neural networks","volume":"7","author":"Guo","year":"2017","journal-title":"Int. J. Multimed. Inf. Retr."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"102643","DOI":"10.1016\/j.bspc.2021.102643","article-title":"Dilated MultiResUNet: Dilated multiresidual blocks network based on U-Net for biomedical image segmentation","volume":"68","author":"Yang","year":"2021","journal-title":"Biomed. Signal Process. Control"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"640","DOI":"10.1109\/TPAMI.2016.2572683","article-title":"Fully convolutional networks for semantic segmentation","volume":"39","author":"Shelhamer","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_15","first-page":"420","article-title":"Page Segmentation in OCR System-A Review","volume":"4","author":"Kaur","year":"2013","journal-title":"Int. J. Comput. Sci. Inf. Technol."},{"key":"ref_16","unstructured":"Reisswig, C., Katti, A., Spinaci, M., and H\u00f6hne, J. (2019, January 14). Chargrid-OCR: End-to-end trainable Optical Character Recognition through Semantic Segmentation and Object Detection. Proceedings of the Workshop on Document Intelligence at NeurIPS 2019, Vancouver, BC, Canada."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1007\/978-981-16-3915-9_5","article-title":"Handwriting Recognition Using Deep Learning","volume":"2021","author":"Shubh","year":"2021","journal-title":"Emerg. Trends Data Driven Comput. Commun. Proc."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Boualam, M., Elfakir, Y., Khaissidi, G., and Mrabti, M. (2021). Arabic Handwriting Word Recognition Based on Convolutional Recurrent Neural Network. WITS 2020, Springer.","DOI":"10.1007\/978-981-33-6893-4_79"},{"key":"ref_19","unstructured":"Huo, Q. (2022, February 12). Underline Detection and Removal in a Document Image Usingmultiple Strategies. Available online: https:\/\/www.researchgate.net\/publication\/4090302_Underline_detection_and_removal_in_a_document_image_using_multiple_strategies."},{"key":"ref_20","first-page":"73","article-title":"Skew Correction of Textural Documents","volume":"15","author":"Abuhaiba","year":"2003","journal-title":"J. King Saud Univ.-Comput. Inf. Sci."},{"key":"ref_21","unstructured":"Patrick, J. (1995). Handprinted Forms and Character Database, NIST Special Database 19, National Institute of Standards and Technology."},{"key":"ref_22","unstructured":"(2022, February 20). Google Cloud Vision API Documentation. Available online: https:\/\/cloud.google.com\/vision\/docs\/drag-and-drop."},{"key":"ref_23","unstructured":"Dataturks.com (2022, March 01). Image Text Recognition APIs Showdown. Google Vision vs Microsoft Cognitive Services vs AWS Rekognition. Available online: https:\/\/dataturks.com\/blog\/compare-image-text-recognition-apis.php."},{"key":"ref_24","first-page":"12018","article-title":"Image Segmentation Based on Improved Unet","volume":"Volume 1815","author":"Li","year":"2021","journal-title":"Journal of Physics: Conference Series"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"124945","DOI":"10.1109\/ACCESS.2021.3110787","article-title":"MMU-OCR-21: Towards End-to-End Urdu Text Recognition Using Deep Learning","volume":"9","author":"Nasir","year":"2021","journal-title":"IEEE Access"},{"key":"ref_26","unstructured":"(2022, April 04). U-Net Architecture Image, 2011, LMB, University of Freiburg Department of Computer Science Faculty of Engineering. Available online: https:\/\/lmb.informatik.uni-freiburg.de\/people\/ronneber\/u-net\/."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Hwang, S.-M., and Yeom, H.-G. (2021). An Implementation of a System for Video Translation Using OCR. Software Engineering in IoT, Big Data, Cloud and Mobile Computing, Springer.","DOI":"10.1007\/978-3-030-64773-5_4"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Edupuganti, S.A., Koganti, V.D., Lakshmi, C.S., Kumar, R.N., and Paruchuri, R. (2021, January 7\u20139). Text and Speech Recognition for Visually Impaired People using Google Vision. Proceedings of the 2021 2nd International Conference on Smart Electronics and Communication (ICOSEC), Tiruchirappalli, India.","DOI":"10.1109\/ICOSEC51865.2021.9591829"}],"container-title":["Journal of Sensor and Actuator Networks"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2224-2708\/11\/4\/63\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T00:45:48Z","timestamp":1760143548000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2224-2708\/11\/4\/63"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,3]]},"references-count":28,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2022,12]]}},"alternative-id":["jsan11040063"],"URL":"https:\/\/doi.org\/10.3390\/jsan11040063","relation":{},"ISSN":["2224-2708"],"issn-type":[{"value":"2224-2708","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,10,3]]}}}