{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,14]],"date-time":"2026-03-14T12:30:03Z","timestamp":1773491403658,"version":"3.50.1"},"reference-count":22,"publisher":"Association for Computing Machinery (ACM)","issue":"3","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2026,3,31]]},"abstract":"<jats:p>\n                    State-of-the-art multilingual Automatic Speech Recognition (ASR) models produce systematic errors when applied to low-resource languages like Rajasthani, for which they lack dedicated training data. This article addresses this challenge by introducing a post-ASR correction framework that leverages the complementary error patterns in the outputs (termed as views) from two distinct models: Whisper-large-v3 and MMS-1B-All. We propose a multi-view,\n                    <jats:xref ref-type=\"fn\">\n                      <jats:sup>1<\/jats:sup>\n                    <\/jats:xref>\n                    character-level sequence-to-sequence (Seq2Seq) model that uses a gated fusion mechanism to dynamically weigh information from the two ASR outputs. On a new benchmark created from the IndicTTS Rajasthani corpus, our gated model achieves a Character Error Rate (CER) of 7.86% and a Word Error Rate (WER) of 29.98%. This outperforms the best single-view baselines (8.01% CER and 30.33% WER), simple multi-view concatenation (8.21% CER and 30.05% WER), as well as Llama-3.2-3B and mBART-50-large, both fine-tuned on Whisper and MMS inputs. It also surpasses powerful Large Language Models (LLMs) like GPT-4o and Gemini 2.5 Pro in a zero-shot setting. This work establishes the first baseline for post-ASR correction in Rajasthani, demonstrating that a compact, specialized model is more effective than general-purpose LLMs for this targeted, low-resource task.\n                  <\/jats:p>","DOI":"10.1145\/3793254","type":"journal-article","created":{"date-parts":[[2026,1,24]],"date-time":"2026-01-24T20:45:16Z","timestamp":1769287516000},"page":"1-13","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Post-ASR Correction for Low-Resource Rajasthani Language"],"prefix":"10.1145","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5193-7213","authenticated-orcid":false,"given":"Abhishek","family":"Bhandari","sequence":"first","affiliation":[{"name":"Computer Science and Engineering, Indian Institute of Technology Jodhpur","place":["Jodhpur, India"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7943-0123","authenticated-orcid":false,"given":"Gaurav","family":"Harit","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, Indian Institute of Technology Jodhpur","place":["Jodhpur, India"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,2,27]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv preprintarxiv:1409.0473 (2014)."},{"key":"e_1_3_2_3_2","unstructured":"Gheorghe Comanici Eric Bieber Mike Schaekermann Ice Pasupat Noveen Sachdeva Inderjit Dhillon Marcel Blistein Ori Ram Dan Zhang Evan Rosen et\u00a0al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning multimodality long context and next generation agentic capabilities. arXiv preprintarxiv:2507.06261 (2025)."},{"key":"e_1_3_2_4_2","unstructured":"Speech Technology Consortium Hema A Murthy and S Umesh. 2023. Indic TTS: A Text-to-Speech Database for Indian Languages. Retrieved from https:\/\/www.iitm.ac.in\/donlab\/indictts\/"},{"key":"e_1_3_2_5_2","unstructured":"Samrat Dutta Shreyansh Jain Ayush Maheshwari Souvik Pal Ganesh Ramakrishnan and Preethi Jyothi. 2022. Error correction in asr using sequence-to-sequence models. arXiv preprintarxiv:2202.01157 (2022)."},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU.1997.659110"},{"key":"e_1_3_2_7_2","unstructured":"Aaron Grattafiori Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman Akhil Mathur Alan Schelten Alex Vaughan et\u00a0al. 2024. The llama 3 herd of models. arXiv preprintarxiv:2407.21783 (2024)."},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","unstructured":"Alex Graves and J\u00fcrgen Schmidhuber. 2005. Framewise phoneme classification with bidirectional LSTM and other neural network architectures. Neural Networks 18 5-6 (2005) 602\u2013610.","DOI":"10.1016\/j.neunet.2005.06.042"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long short-term memory. Neural Computation 9 8 (1997) 1735\u20131780.","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_10_2","unstructured":"Aaron Hurst Adam Lerer Adam P Goucher Adam Perelman Aditya Ramesh Aidan Clark AJ Ostrow Akila Welihinda Alan Hayes Alec Radford et\u00a0al. 2024. Gpt-4o system card. arXiv preprintarxiv:2410.21276 (2024)."},{"key":"e_1_3_2_11_2","unstructured":"Amrith Krishna Bodhisattwa Prasad Majumder Rajesh Shreedhar Bhat and Pawan Goyal. 2018. Upcycle your OCR: Reusing OCRs for post-OCR text correction in romanised sanskrit. arXiv preprintarxiv:1809.02147 (2018)."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2024-368"},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","unstructured":"Yinhan Liu Jiatao Gu Naman Goyal Xian Li Sergey Edunov Marjan Ghazvininejad Mike Lewis and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics 8 (2020) 726\u2013742.","DOI":"10.1162\/tacl_a_00343"},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","unstructured":"Ayush Maheshwari Nikhil Singh Amrith Krishna and Ganesh Ramakrishnan. 2022. A benchmark and dataset for post-OCR text correction in Sanskrit. arXiv preprintarxiv:2211.07980 (2022).","DOI":"10.18653\/v1\/2022.findings-emnlp.466"},{"key":"e_1_3_2_15_2","unstructured":"Aditya Pal and Abhijit Mustafi. 2020. Vartani spellcheck\u2013automatic context-sensitive spelling correction of OCR-generated hindi text using BERT and levenshtein distance. arXiv preprintarxiv:2012.07652 (2020)."},{"key":"e_1_3_2_16_2","unstructured":"Vineel Pratap Andros Tjandra Bowen Shi Paden Tomasello Arun Babu Sayani Kundu Ali Elkahky Zhaoheng Ni Apoorv Vyas Maryam Fazel-Zarandi Alexei Baevski Yossi Adi Xiaohui Zhang Wei-Ning Hsu Alexis Conneau and Michael Auli. 2023. Scaling speech technology to 1 000+ languages. arXiv (2023)."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","unstructured":"Alec Radford Jong Wook Kim Tao Xu Greg Brockman Christine McLeavey and Ilya Sutskever. 2022. Robust Speech Recognition via Large-Scale Weak Supervision. DOI:10.48550\/ARXIV.2212.04356","DOI":"10.48550\/ARXIV.2212.04356"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.478"},{"key":"e_1_3_2_19_2","doi-asserted-by":"crossref","unstructured":"Shruti Rijhwani Daisy Rosenblum Antonios Anastasopoulos and Graham Neubig. 2021. Lexically aware semi-supervised learning for OCR post-correction. Transactions of the Association for Computational Linguistics 9 (2021) 1285\u20131302.","DOI":"10.1162\/tacl_a_00427"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2017.13"},{"key":"e_1_3_2_21_2","unstructured":"Abigail See Peter J Liu and Christopher D Manning. 2017. Get to the point: Summarization with pointer-generator networks. arXiv preprintarxiv:1704.04368 (2017)."},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","unstructured":"Kai Shen Yichong Leng Xu Tan Siliang Tang Yuan Zhang Wenjie Liu and Edward Lin. 2022. Mask the correct tokens: An embarrassingly simple approach for error correction. arXiv preprintarxiv:2211.13252 (2022).","DOI":"10.18653\/v1\/2022.emnlp-main.708"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053213"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3793254","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,14]],"date-time":"2026-03-14T11:20:57Z","timestamp":1773487257000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3793254"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,27]]},"references-count":22,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,3,31]]}},"alternative-id":["10.1145\/3793254"],"URL":"https:\/\/doi.org\/10.1145\/3793254","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,27]]},"assertion":[{"value":"2025-08-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-17","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}