{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T11:56:44Z","timestamp":1781265404155,"version":"3.54.1"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T00:00:00Z","timestamp":1781222400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    Optical Character Recognition (OCR) has accelerated the digitization of printed Indic texts, yet recognition remains error-prone due to ligatures, matras, conjuncts, orthographic variability, and long-range grammatical dependencies. Hence, it is imperative to employ post-OCR error-correction techniques that can exploit broader linguistic cues. This article presents a context-aware correction framework that leverages a larger sentence-level context around the erroneous span. The correction model inputs the OCR-generated sentence, along with an auxiliary context sentence, and outputs a corrected sequence. Our correction model is a pre-trained language model finetuned on a small, supervised corpus. Experiments demonstrate substantial reductions in the character error rates (10.20% to 6.73% for Hindi, 6.10% to 1.39% for Gujarati, and 8.19% to 3.29% for Marathi) as well as the word error rates (28.53% to 14.57% for Hindi, 20.85% to 4.16% for Gujarati, and 24.89% to 8.28% for Marathi). These results outperform the seq2seq baselines. Error-type analysis indicates the largest improvements for character confusion and conjunct\/ligature errors. These results demonstrate that supplying a larger context consistently improves post-OCR correction for Indic scripts. The dataset is publicly available via Hugging Face at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/huggingface.co\/datasets\/AbhishekBhandari\/Indic-post-ocr-correction\">https:\/\/huggingface.co\/datasets\/AbhishekBhandari\/Indic-post-ocr-correction<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3815575","type":"journal-article","created":{"date-parts":[[2026,5,11]],"date-time":"2026-05-11T11:30:24Z","timestamp":1778499024000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["A Framework and Dataset for Contextual Post-OCR Correction"],"prefix":"10.1145","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5193-7213","authenticated-orcid":false,"given":"Abhishek","family":"Bhandari","sequence":"first","affiliation":[{"name":"Computer Science and Engineering, Indian Institute of Technology Jodhpur","place":["Jodhpur, India"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7943-0123","authenticated-orcid":false,"given":"Gaurav","family":"Harit","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, Indian Institute of Technology Jodhpur","place":["Jodhpur, India"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,12]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/SKIMA.2016.7916243"},{"issue":"4","key":"e_1_3_1_3_2","doi-asserted-by":"crossref","first-page":"269","DOI":"10.1007\/s100320100066","article-title":"Partitioning and searching dictionary for correction of optically read Devanagari character strings","volume":"4","author":"Bansal Veena","year":"2002","unstructured":"Veena Bansal and RMK Sinha. 2002. Partitioning and searching dictionary for correction of optically read Devanagari character strings. International Journal on Document Analysis and Recognition 4, 4 (2002), 269\u2013280.","journal-title":"International Journal on Document Analysis and Recognition"},{"key":"e_1_3_1_4_2","first-page":"102","volume-title":"Proceedings of the Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portugal, April 14\u201317, 2020, Part II 42","author":"Bazzo Guilherme Torresan","year":"2020","unstructured":"Guilherme Torresan Bazzo, Gustavo Acauan Lorentz, Danny Suarez Vargas, and Viviane P. Moreira. 2020. Assessing the impact of OCR errors in information retrieval. In Proceedings of the Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portugal, April 14\u201317, 2020, Part II 42. Springer, 102\u2013109."},{"issue":"4","key":"e_1_3_1_5_2","first-page":"1","article-title":"Post-ocr text correction for Bulgarian historical documents","volume":"26","author":"Beshirov Angel","year":"2025","unstructured":"Angel Beshirov, Milena Dobreva, Dimitar Dimitrov, Momchil Hardalov, Ivan Koychev, and Preslav Nakov. 2025. Post-ocr text correction for Bulgarian historical documents. International Journal on Digital Libraries 26, 4 (2025), 1\u201311.","journal-title":"International Journal on Digital Libraries"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.1996.546947"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-70645-5_2"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-acl.145"},{"key":"e_1_3_1_9_2","first-page":"655","volume-title":"Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR)","author":"Das Deepayan","year":"2019","unstructured":"Deepayan Das, Jerin Philip, Minesh Mathew, and C. V. Jawahar. 2019. A cost efficient approach to correct OCR errors in large document collections. In Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR). IEEE, 655\u2013662."},{"key":"e_1_3_1_10_2","first-page":"1","volume-title":"Proceedings of the 2025 ACM Symposium on Document Engineering","author":"Ara\u00fajo S\u00e1vio Santos de","year":"2025","unstructured":"S\u00e1vio Santos de Ara\u00fajo, Byron Leite Dantas Bezerra, and Arthur Flor de Sousa Neto. 2025. A proposal of post-OCR spelling correction using monolingual byte-level language models. In Proceedings of the 2025 ACM Symposium on Document Engineering. 1\u20134."},{"key":"e_1_3_1_11_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2019)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2019). Association for Computational Linguistics, Minneapolis, MN, 4171\u20134186."},{"key":"e_1_3_1_12_2","unstructured":"Sumanth Doddapaneni Rahul Aralikatte Gowtham Ramesh Shreyansh Goyal Mitesh M. Khapra Anoop Kunchukuttan and Pratyush Kumar. 2022. Towards leaving no Indic language behind: Building monolingual corpora benchmark and models for Indic languages. arXiv:2212.05409. Retrieved from https:\/\/arxiv.org\/abs\/2212.05409"},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","first-page":"6036","DOI":"10.18653\/v1\/2024.findings-acl.361","volume-title":"Findings of the Association for Computational Linguistics ACL 2024","author":"Guan Shuhao","year":"2024","unstructured":"Shuhao Guan and Derek Greene. 2024. Advancing post-OCR correction: A comparative study of synthetic data. In Findings of the Association for Computational Linguistics ACL 2024. Association for Computational Linguistics, 6036\u20136047."},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","unstructured":"Shuhao Guan Moule Lin Cheng Xu Xinyi Liu Jinman Zhao Jiexin Fan Qi Xu and Derek Greene. 2025. Prep-OCR: A complete pipeline for document image restoration and enhanced OCR accuracy. arXiv:2505.20429. Retrieved from https:\/\/arxiv.org\/abs\/2505.20429","DOI":"10.18653\/v1\/2025.acl-long.749"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_16_2","unstructured":"Jenna Kanerva Cassandra Ledins Siiri K\u00e4pyaho and Filip Ginter. 2025. OCR error post-correction with LLMs in historical documents: No free lunches. arXiv:2502.01205. Retrieved from https:\/\/arxiv.org\/abs\/2502.01205"},{"key":"e_1_3_1_17_2","unstructured":"Harshvivek Kashid and Pushpak Bhattacharyya. 2025. RoundTripOCR: A data generation technique for enhancing post-OCR error correction in low-resource Devanagari languages. arXiv:2412.15248. Retrieved from https:\/\/arxiv.org\/abs\/2412.15248"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","first-page":"345","DOI":"10.18653\/v1\/K18-1034","volume-title":"Proceedings of the 22nd Conference on Computational Natural Language Learning","author":"Krishna Amrith","year":"2018","unstructured":"Amrith Krishna, Bodhisattwa P. Majumder, Rajesh Bhat, and Pawan Goyal. 2018. Upcycle your OCR: Reusing OCRs for post-OCR text correction in Romanised Sanskrit. In Proceedings of the 22nd Conference on Computational Natural Language Learning. 345\u2013355."},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","first-page":"358","DOI":"10.1007\/3-540-70659-3_37","volume-title":"Proceedings of the Structural, Syntactic, and Statistical Pattern Recognition: Joint IAPR International Workshops SSPR 2002 and SPR 2002 Windsor, Ontario, Canada, August 6\u20139, 2002","author":"Lehal G. S.","year":"2002","unstructured":"G. S. Lehal and Chandan Singh. 2002. A complete OCR system for Gurmukhi script. In Proceedings of the Structural, Syntactic, and Statistical Pattern Recognition: Joint IAPR International Workshops SSPR 2002 and SPR 2002 Windsor, Ontario, Canada, August 6\u20139, 2002. Springer, 358\u2013367."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2001.953957"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.acl-industry.17"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00343"},{"key":"e_1_3_1_24_2","doi-asserted-by":"crossref","unstructured":"Ayush Maheshwari Nikhil Singh Amrith Krishna and Ganesh Ramakrishnan. 2022. A benchmark and dataset for post-OCR text correction in Sanskrit. arXiv:2211.07980. Retrieved from https:\/\/arxiv.org\/abs\/2211.07980","DOI":"10.18653\/v1\/2022.findings-emnlp.466"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/1815330.1815394"},{"issue":"1","key":"e_1_3_1_26_2","first-page":"91","article-title":"Spell checker for OCR","volume":"4","author":"Mohapatra Yogomaya","year":"2013","unstructured":"Yogomaya Mohapatra, Ashis Kumar Mishra, and Anil Kumar Mishra. 2013. Spell checker for OCR. International Journal of Computer Science and Information Technologies 4, 1 (2013), 91\u201397.","journal-title":"International Journal of Computer Science and Information Technologies"},{"key":"e_1_3_1_27_2","first-page":"1238","volume-title":"Proceedings of the 9th International Conference on Document Analysis and Recognition (ICDAR 2007)","author":"Namboodiri A.","year":"2007","unstructured":"A. Namboodiri, P. Narayanan, and C. Jawahar. 2007. On using classical poetry structure for Indian language post-processing. In Proceedings of the 9th International Conference on Document Analysis and Recognition (ICDAR 2007). IEEE, 1238\u20131242."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3283340"},{"key":"e_1_3_1_29_2","unstructured":"Aditya Pal and Abhijit Mustafi. 2020. Vartani spellcheck\u2013automatic context-sensitive spelling correction of OCR-generated Hindi text using BERT and Levenshtein distance. arXiv:2012.07652. Retrieved from https:\/\/arxiv.org\/abs\/2012.07652"},{"issue":"6","key":"e_1_3_1_30_2","first-page":"903","article-title":"OCR error correction of an inflectional indian language using morphological parsing","volume":"16","author":"Pal U.","year":"2000","unstructured":"U. Pal, Pulak K. Kundu, and Bidyut Baran Chaudhuri. 2000. OCR error correction of an inflectional indian language using morphological parsing. Journal of Information Science and Engineering 16, 6 (2000), 903\u2013922.","journal-title":"Journal of Information Science and Engineering"},{"key":"e_1_3_1_31_2","first-page":"1","volume-title":"Proceedings of the 2014 International Conference on Green Computing Communication and Electrical Engineering (ICGCCEE)","author":"Patel Dhruv B.","year":"2014","unstructured":"Dhruv B. Patel and Mukesh M. Goswami. 2014. Word level correction in Gujarati document using probabilistic approach. In Proceedings of the 2014 International Conference on Green Computing Communication and Electrical Engineering (ICGCCEE). IEEE, 1\u20135."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.5555\/3455716.3455856"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ITACT.2015.7492681"},{"key":"e_1_3_1_34_2","first-page":"17","volume-title":"Proceedings of the 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR)","author":"Saluja Rohit","year":"2017","unstructured":"Rohit Saluja, Devaraj Adiga, Parag Chaudhuri, Ganesh Ramakrishnan, and Mark Carman. 2017. Error detection and corrections in Indic OCR using LSTMs. In Proceedings of the 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR). IEEE, 17\u201322."},{"key":"e_1_3_1_35_2","first-page":"25","volume-title":"Proceedings of the 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR)","author":"Saluja Rohit","year":"2017","unstructured":"Rohit Saluja, Devaraj Adiga, Ganesh Ramakrishnan, Parag Chaudhuri, and Mark Carman. 2017. A framework for document specific error detection and corrections in indic ocr. In Proceedings of the 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR). IEEE, 25\u201330."},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","first-page":"160","DOI":"10.1109\/ICDAR.2019.00034","volume-title":"Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR)","author":"Saluja Rohit","year":"2019","unstructured":"Rohit Saluja, Mayur Punjabi, Mark Carman, Ganesh Ramakrishnan, and Parag Chaudhuri. 2019. Sub-word embeddings for OCR corrections in highly fusional indic languages. In Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR). IEEE, 160\u2013165."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/0031-3203(87)90075-6"},{"key":"e_1_3_1_38_2","first-page":"5926","volume-title":"Proceedings of the 36th International Conference on Machine Learning (ICML 2019), Proceedings of Machine Learning Research","volume":"97","author":"Song Kaitao","year":"2019","unstructured":"Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019. MASS: Masked sequence to sequence pre-training for language generation. In Proceedings of the 36th International Conference on Machine Learning (ICML 2019), Proceedings of Machine Learning Research. Vol. 97, PMLR, 5926\u20135936."},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","first-page":"284","DOI":"10.18653\/v1\/2021.wnut-1.31","volume-title":"Proceedings of the 7th Workshop on Noisy User-generated Text (W-NUT 2021)","author":"Soper Elizabeth","year":"2021","unstructured":"Elizabeth Soper, Stanley Fujimoto, and Yen-Yun Yu. 2021. BART for post-correction of OCR newspaper text. In Proceedings of the 7th Workshop on Noisy User-generated Text (W-NUT 2021). 284\u2013290."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.3390\/math13152372"},{"key":"e_1_3_1_41_2","first-page":"116","volume-title":"Proceedings of the 3rd Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA) @ LREC-COLING-2024","author":"Thomas Alan","year":"2024","unstructured":"Alan Thomas, Robert Gaizauskas, and Haiping Lu. 2024. Leveraging LLMs for post-OCR correction of historical newspapers. In Proceedings of the 3rd Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA) @ LREC-COLING-2024. ELRA and ICCL, 116\u2013121."},{"key":"e_1_3_1_42_2","first-page":"5998","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017), 5998\u20136008.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_43_2","first-page":"32","volume-title":"Proceedings of the 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR)","author":"Vinitha V. S.","year":"2017","unstructured":"V. S. Vinitha, Minesh Mathew, and C. V. Jawahar. 2017. An empirical study of effectiveness of post-processing in indic scripts. In Proceedings of the 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR). IEEE, 32\u201336."},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00461"},{"key":"e_1_3_1_45_2","first-page":"483","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2021)","author":"Xue Linting","year":"2021","unstructured":"Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2021). Association for Computational Linguistics, 483\u2013498."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3815575","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T11:06:11Z","timestamp":1781262371000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3815575"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,12]]},"references-count":44,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3815575"],"URL":"https:\/\/doi.org\/10.1145\/3815575","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,12]]},"assertion":[{"value":"2024-09-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}