{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T01:03:05Z","timestamp":1760058185001,"version":"build-2065373602"},"reference-count":30,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2025,3,18]],"date-time":"2025-03-18T00:00:00Z","timestamp":1742256000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"\u201cRomanian Hub for Artificial Intelligence\u2014HRIA\u201d, Smart Growth, Digitization and Financial Instruments Program, 2021\u20132027, MySMIS","award":["334906"],"award-info":[{"award-number":["334906"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Nowadays, grammatical error correction (GEC) has a significant role in writing since even native speakers often face challenges with proficient writing. This research is focused on developing a methodology to correct grammatical errors in the Romanian language, a less-resourced language for which there are currently no up-to-date GEC solutions. Our main contributions include an open-source synthetic dataset of 345,403 Romanian sentences, a manually curated dataset of 3054 social media comments, a two-phased GEC approach, and a comparison with several Romanian models, including RoMistral and RoLama3, but also LanguageTool, GPT-4o mini, and GPT-4o. We consider a synthetic dataset to finetune our models, while we rely on two real-life datasets with genuine human mistakes (i.e., CNA and RoComments) to evaluate performance. Building an artificial dataset was necessary because of the scarcity of real-life mistake datasets, whereas introducing RoComments, a new genuine dataset, is argued by the necessity to cover errors amongst native speakers encountered in social media comments. We also introduce a two-phased approach, where we first identify the location of erroneous tokens in the sentence; next, the erroneous tokens are replaced by an encoder\u2013decoder model. Our approach achieved an F0.5 of 0.57 on CNA and 0.64 on RoComments, surpassing by a considerable margin LanguageTool as well as an end-to-end version based on Flan-T5 and mT0 in most setups. While our two-phased method did not outperform GPT-4o, arguably by its smaller size and language exposure, it obtained on-par results with GPT-4o mini and achieved higher performance than all Romanian LLMs.<\/jats:p>","DOI":"10.3390\/info16030242","type":"journal-article","created":{"date-parts":[[2025,3,18]],"date-time":"2025-03-18T07:20:00Z","timestamp":1742282400000},"page":"242","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Show Me All Writing Errors: A Two-Phased Grammatical Error Corrector for Romanian"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-8782-5037","authenticated-orcid":false,"given":"Mihai-Cristian","family":"Tudose","sequence":"first","affiliation":[{"name":"Faculty of Automatic Control and Computers, National University of Science and Technology POLITEHNICA Bucharest, 313 Splaiul Independentei, 060042 Bucharest, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0380-6814","authenticated-orcid":false,"given":"Stefan","family":"Ruseti","sequence":"additional","affiliation":[{"name":"Faculty of Automatic Control and Computers, National University of Science and Technology POLITEHNICA Bucharest, 313 Splaiul Independentei, 060042 Bucharest, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4815-9227","authenticated-orcid":false,"given":"Mihai","family":"Dascalu","sequence":"additional","affiliation":[{"name":"Faculty of Automatic Control and Computers, National University of Science and Technology POLITEHNICA Bucharest, 313 Splaiul Independentei, 060042 Bucharest, Romania"},{"name":"Academy of Romanian Scientists, Str. Ilfov, Nr. 3, 050044 Bucharest, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,3,18]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Jes\u00fas-Ortiz, E., and Calvo-Ferrer, J.R. (2023). His or Her? Errors in Possessive Determiners Made by L2-English Native Spanish Speakers. Languages, 8.","DOI":"10.3390\/languages8040278"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Rei, M., and Yannakoudakis, H. (2017, January 8). Auxiliary Objectives for Neural Error Detection Models. Proceedings of the 12th Workshop on Innovative Use of NLP for Building Educational Applications, Association for Computational Linguistics, Copenhagen, Denmark.","DOI":"10.18653\/v1\/W17-5004"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Ng, H.T., Wu, S.M., Briscoe, T., Hadiwinoto, C., Susanto, R.H., and Bryant, C. (2014, January 26\u201327). The CoNLL-2014 Shared Task on Grammatical Error Correction. Proceedings of the 18th Conference on Computational Natural Language Learning: Shared Task, Baltimore, MD, USA.","DOI":"10.3115\/v1\/W14-1701"},{"key":"ref_4","unstructured":"Yannakoudakis, H., Briscoe, T., and Medlock, B. (2011, January 19\u201324). A New Dataset and Method for Automatically Grading ESOL Texts. Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Portland, OR, USA."},{"key":"ref_5","unstructured":"Naber, D. (2025, January 24). A Rule-Based Style and Grammar Checker. Available online: https:\/\/www.researchgate.net\/publication\/239556866_A_Rule-Based_Style_and_Grammar_Checker."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Cotet, T.M., Ruseti, S., and Dascalu, M. (2020, January 9\u201311). Neural grammatical error correction for romanian. Proceedings of the 2020 IEEE 32nd International Conference on Tools with Artificial Intelligence (ICTAI), Baltimore, MD, USA.","DOI":"10.1109\/ICTAI50040.2020.00101"},{"key":"ref_7","unstructured":"Dahlmeier, D., Ng, H.T., and Wu, S.M. (2013, January 13). Building a large annotated corpus of learner English: The NUS corpus of learner English. Proceedings of the 8th Workshop on Innovative Use of NLP for Building Educational Applications, Atlanta, GA, USA."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Mizumoto, T., and Matsumoto, Y. (2016, January 12\u201317). Discriminative reranking for grammatical error correction with statistical machine translation. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego, CA, USA.","DOI":"10.18653\/v1\/N16-1133"},{"key":"ref_9","first-page":"71","article-title":"Features and functions of the HSK dynamic composition corpus","volume":"4","author":"Zhang","year":"2009","journal-title":"Int. Chin. Lang. Educ."},{"key":"ref_10","unstructured":"Omelianchuk, K., Liubonko, A., Skurzhanskyi, O., Chernodub, A., Korniienko, O., and Samokhin, I. (2024, January 20). Pillars of Grammatical Error Correction: Comprehensive Inspection Of Contemporary Approaches In The Era of Large Language Models. Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2024), Mexico City, Mexico."},{"key":"ref_11","unstructured":"Knight, K., Nenkova, A., and Rambow, O. (2016, January 12\u201317). Grammatical error correction using neural machine translation. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego, CA, USA."},{"key":"ref_12","unstructured":"Zong, C., Xia, F., Li, W., and Navigli, R. (2021, January 1\u20136). A Simple Recipe for Multilingual Grammatical Error Correction. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), Online."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C. (2021, January 6\u201311). mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Online.","DOI":"10.18653\/v1\/2021.naacl-main.41"},{"key":"ref_14","first-page":"1","article-title":"Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer","volume":"21","author":"Raffel","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Lin, N., Fu, Y., and Lin, X. (2023). A New Evaluation Method: Evaluation Data and Metrics for Chinese Grammar Error Correction. arxiv.","DOI":"10.21203\/rs.3.rs-2299197\/v1"},{"key":"ref_16","unstructured":"Webber, B., Cohn, T., He, Y., and Liu, Y. (2020, January 16\u201320). Improving the Efficiency of Grammatical Error Correction with Erroneous Span Detection and Correction. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"146772","DOI":"10.1109\/ACCESS.2019.2940607","article-title":"A two-stage model for Chinese grammatical error correction","volume":"7","author":"Qiu","year":"2019","journal-title":"IEEE Access"},{"key":"ref_18","unstructured":"Keita, M.K., Homan, C., Hamani, S.A., Bremang, A., Zampieri, M., Alfari, H.A., Ibrahim, E.A., and Owusu, D. (2024). Grammatical Error Correction for Low-Resource Languages: The Case of Zarma. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Maity, S., Deroy, A., and Sarkar, S. (2024). How Ready Are Generative Pre-trained Large Language Models for Explaining Bengali Grammatical Errors?. arXiv.","DOI":"10.36227\/techrxiv.171665626.60610675\/v1"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Lytvyn, V., Pukach, P., Vysotska, V., Vovk, M., and Kholodna, N. (2023). Identification and Correction of Grammatical Errors in Ukrainian Texts Based on Machine Learning Technology. Mathematics, 11.","DOI":"10.3390\/math11040904"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Musyafa, A., Gao, Y., Solyman, A., Wu, C., and Khan, S. (2022). Automatic correction of indonesian grammatical errors based on transformer. Appl. Sci., 12.","DOI":"10.3390\/app122010380"},{"key":"ref_22","unstructured":"Barzilay, R., and Kan, M.Y. (August, January 30). Automatic Annotation and Evaluation of Error Types for Grammatical Error Correction. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vancouver, BC, Canada."},{"key":"ref_23","unstructured":"RoTex (2024, May 16). RoTex Corpus Builder. Available online: https:\/\/github.com\/aleris\/ReadME-RoTex-Corpus-Builder."},{"key":"ref_24","unstructured":"Bucur, C. (2024, May 16). Cele mai Frecvente Gre\u0219eli de Gramatic\u0103 din Limba Rom\u00e2n\u0103. Available online: https:\/\/life.ro\/cele-mai-frecvente-greseli-de-gramatica-din-limba-romana\/."},{"key":"ref_25","unstructured":"SpotMedia (2024, May 16). Cum Scriem Corect: Nu Face Sau nu f\u0103? Cum Folosim Nega\u021bia la Imperativ. Available online: https:\/\/spotmedia.ro\/stiri\/educatie\/cum-scriem-corect-nu-face-sau-nu-fa-cum-folosim-negatia-la-imperativ."},{"key":"ref_26","unstructured":"Mititelu, V.B., Irimia, E., Perez, C.A., Ion, R., Simionescu, R., and Popel, M. (2024, May 21). UD Romanian RRT. Available online: https:\/\/universaldependencies.org\/treebanks\/ro_rrt\/index.html."},{"key":"ref_27","unstructured":"Destepti.ro (2024, May 16). Cum Este Corect\u2013Al\u0163ii Sau Al\u0163i?. Available online: https:\/\/destepti.ro\/cum-este-corect-altii-sau-alti\/."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Muennighoff, N., Wang, T., Sutawika, L., Roberts, A., Biderman, S., Le Scao, T., Bari, M.S., Shen, S., Yong, Z.X., and Schoelkopf, H. (2023, January 9\u201314). Crosslingual Generalization through Multitask Finetuning. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Toronto, ON, Canada.","DOI":"10.18653\/v1\/2023.acl-long.891"},{"key":"ref_29","unstructured":"Chung, H.W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., and Brahma, S. (2022). Scaling Instruction-Finetuned Language Models. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Dumitrescu, S., Avram, A.M., and Pyysalo, S. (2020, January 16\u201320). The birth of Romanian BERT. Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2020, Online.","DOI":"10.18653\/v1\/2020.findings-emnlp.387"}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/3\/242\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T16:55:53Z","timestamp":1760028953000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/3\/242"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,18]]},"references-count":30,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2025,3]]}},"alternative-id":["info16030242"],"URL":"https:\/\/doi.org\/10.3390\/info16030242","relation":{},"ISSN":["2078-2489"],"issn-type":[{"type":"electronic","value":"2078-2489"}],"subject":[],"published":{"date-parts":[[2025,3,18]]}}}