{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T15:55:35Z","timestamp":1768319735395,"version":"3.49.0"},"reference-count":51,"publisher":"Cambridge University Press (CUP)","issue":"5","license":[{"start":{"date-parts":[[2015,3,18]],"date-time":"2015-03-18T00:00:00Z","timestamp":1426636800000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2016,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>A spelling error detection and correction application is typically based on three main components: a dictionary (or reference word list), an error model and a language model. While most of the attention in the literature has been directed to the language model, we show how improvements in any of the three components can lead to significant cumulative improvements in the overall performance of the system. We develop our dictionary of 9.2 million fully-inflected Arabic words (types) from a morphological transducer and a large corpus, validated and manually revised. We improve the error model by analyzing error types and creating an edit distance re-ranker. We also improve the language model by analyzing the level of noise in different data sources and selecting an optimal subset to train the system on. Testing and evaluation experiments show that our system significantly outperforms Microsoft Word 2013, OpenOffice Ayaspell 3.4 and Google Docs.<\/jats:p>","DOI":"10.1017\/s1351324915000030","type":"journal-article","created":{"date-parts":[[2015,3,18]],"date-time":"2015-03-18T18:19:37Z","timestamp":1426702777000},"page":"751-773","source":"Crossref","is-referenced-by-count":18,"title":["Arabic spelling error detection and correction"],"prefix":"10.1017","volume":"22","author":[{"given":"MOHAMMED","family":"ATTIA","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"PAVEL","family":"PECINA","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"YOUNES","family":"SAMIH","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"KHALED","family":"SHAALAN","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"JOSEF","family":"VAN GENABITH","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2015,3,18]]},"reference":[{"key":"S1351324915000030_ref040","unstructured":"Shaalan K. , Magdy M. , and Fahmy A. 2013. Analysis and feedback of erroneous arabic verbs. Journal of Natural Language Engineering, Cambridge University, UK. FirstView: 1\u201353."},{"key":"S1351324915000030_ref050","doi-asserted-by":"crossref","unstructured":"Zribi C. B. O. , and Ben Ahmed M. 2003. Efficient automatic correction of misspelled arabic words based on contextual information. Lecture Notes in Computer Science, Springer, 2773: 770\u2013777.","DOI":"10.1007\/978-3-540-45224-9_104"},{"key":"S1351324915000030_ref009","first-page":"31","volume-title":"Proceedings of the Workshop on Computational Approaches to Arabic Script-based Languages, Association for Computational Linguistics","author":"Buckwalter","year":"2004"},{"key":"S1351324915000030_ref034","unstructured":"Och F. J. , and Genzel D. 2013. Automatic spelling correction for machine translation. Patent US 20130144592 A1. June 6, 2013."},{"key":"S1351324915000030_ref041","unstructured":"Shaalan K. , Samih Y. , Attia M. , Pecina P. , and van Genabith J. 2012. Arabic word generation and modelling for spell checking. In Language Resources and Evaluation (LREC), Istanbul, Turkey. pp. 719\u2013725."},{"key":"S1351324915000030_ref046","doi-asserted-by":"crossref","unstructured":"Watson J. 2002. The Phonology and Morphology of Arabic, New York: Oxford University.","DOI":"10.1093\/oso\/9780199257591.001.0001"},{"key":"S1351324915000030_ref012","doi-asserted-by":"publisher","DOI":"10.1007\/BF01889984"},{"key":"S1351324915000030_ref032","unstructured":"Moussa M. , Fakhr M. W. , and Darwish K. 2012. Statistical denormalization for arabic text. In Proceedings of KONVENS 2012, Vienna, pp. 228\u2013232."},{"key":"S1351324915000030_ref026","doi-asserted-by":"crossref","unstructured":"Kiraz G. A. 2001. Computational Nonlinear Morphology: With Emphasis on Semitic Languages, Cambridge University. Cambridge, United Kingdom.","DOI":"10.1017\/CBO9780511497933"},{"key":"S1351324915000030_ref014","first-page":"45","volume-title":"Proceedings of the Workshop on Semitic Languages in the Seventh International Conference on Language Resources and Evaluation (LREC)","author":"El Kholy","year":"2010"},{"key":"S1351324915000030_ref037","doi-asserted-by":"crossref","unstructured":"Ratcliffe R. R. 1998. The Broken Plural Problem in Arabic and Comparative Semitic: Allomorphy and Analogy in Non-concatenative Morphology, Amsterdam Studies in the Theory and History of Linguistic Science, Series IV, Current issues in linguistic theory, vol. 168. Amsterdam, Philadelphia: J. Benjamins.","DOI":"10.1075\/cilt.168"},{"key":"S1351324915000030_ref030","volume-title":"English Spelling and the Computer","author":"Mitton","year":"1996"},{"key":"S1351324915000030_ref013","doi-asserted-by":"publisher","DOI":"10.1145\/363958.363994"},{"key":"S1351324915000030_ref010","unstructured":"Buckwalter T. 2004b. Buckwalter Arabic Morphological Analyzer (BAMA) Version 2.0. Linguistic Data Consortium (LDC) catalogue number: LDC2004L02."},{"key":"S1351324915000030_ref033","unstructured":"Norvig P. 2009. Natural language corpus data. In T. Segaran and J. Hammerbacher (eds.), Beautiful Data, pp. 219\u2013242. Sebastopol, California: O\u2019Reilly."},{"key":"S1351324915000030_ref048","first-page":"17","article-title":"Integrating dictionary and web N-grams for chinese spell checking","volume":"18","author":"Wu","year":"2013","journal-title":"Computational Linguistics and Chinese Language Processing"},{"key":"S1351324915000030_ref016","first-page":"573","volume-title":"Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics","author":"Habash","year":"2005"},{"key":"#cr-split#-S1351324915000030_ref001.1","unstructured":"11. Alfaifi A., and Atwell E. 2012. Arabic learner corpora"},{"key":"#cr-split#-S1351324915000030_ref001.2","unstructured":"12. (ALC): a taxonomy of coding errors. In Proceedings of the 8th International Computing Conference in Arabic (ICCA 2012), Cairo, Egypt."},{"key":"S1351324915000030_ref047","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324907004676"},{"key":"S1351324915000030_ref035","first-page":"73","article-title":"Error-tolerant finite-state recognition with applications to morphological analysis and spelling correction","volume":"22","author":"Oflazer","year":"1996","journal-title":"Computational Linguistics"},{"key":"S1351324915000030_ref021","first-page":"913","volume-title":"IJCNLP","author":"Hassan","year":"2008"},{"key":"S1351324915000030_ref003","first-page":"48","volume-title":"The Challenge of Arabic for NLP\/MT Conference","author":"Attia","year":"2006"},{"key":"S1351324915000030_ref044","doi-asserted-by":"crossref","unstructured":"Ukkonen E. 1983. On approximate string matching. In Foundations of Computation Theory, vol. 158, pp. 487\u2013495. Lecture Notes in Computer Science, Berlin: Springer.","DOI":"10.1007\/3-540-12689-9_129"},{"key":"S1351324915000030_ref002","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2012.2197612"},{"key":"S1351324915000030_ref043","first-page":"88","volume-title":"Proceedings of the 4th Workshop on Very Large Corpora","author":"Tong","year":"1996"},{"key":"S1351324915000030_ref022","doi-asserted-by":"publisher","DOI":"10.1016\/j.system.2007.09.007"},{"key":"S1351324915000030_ref039","first-page":"240","volume-title":"Proceedings of the 4th Conference on Language Engineering, Egyptian Society of Language Engineering (ELSE)","author":"Shaalan","year":"2003"},{"key":"S1351324915000030_ref006","volume-title":"Finite State Morphology","author":"Beesley","year":"2003"},{"key":"S1351324915000030_ref029","first-page":"408","volume-title":"Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing","author":"Magdy","year":"2006"},{"key":"S1351324915000030_ref020","volume-title":"Data Mining, Southeast Asia Edition: Concepts and Techniques","author":"Han","year":"2006"},{"key":"S1351324915000030_ref036","unstructured":"Parker R. , Graff D. , Chen K. , Kong J. , and Maeda K. 2011. Arabic Gigaword Fifth Edition. LDC Catalog No.: LDC2011T11."},{"key":"S1351324915000030_ref031","first-page":"3","article-title":"ACM SIGKDD explorations newsletter","volume":"7","author":"Mooney","year":"2005","journal-title":"Natural Language Processing and Text Mining"},{"key":"S1351324915000030_ref023","first-page":"57","volume-title":"Proceedings of the 25th Conference of the Spanish Society for Natural Language Processing (SEPLN)","author":"Hulden","year":"2009"},{"key":"S1351324915000030_ref024","first-page":"29","volume-title":"Proceedings of the 12th Conference of the European Chapter of the Association for Computational Linguistics","author":"Hulden","year":"2009"},{"key":"S1351324915000030_ref025","unstructured":"Kernigan M. , Church K. , and Gale W. 1990. A spelling correction program based on a noisy channel model. AT & T Laboratories, 600 Mountain Ave., Murray Hill, NJ, pp. 205\u2013210."},{"key":"S1351324915000030_ref027","doi-asserted-by":"publisher","DOI":"10.1145\/146370.146380"},{"key":"S1351324915000030_ref049","first-page":"26","volume-title":"The 9th Edition of the Language Resources and Evaluation Conference (LREC)","author":"Zaghouani","year":"2014"},{"key":"S1351324915000030_ref008","first-page":"467","article-title":"Class-based n-gram models of natural language","volume":"18","author":"Brown","year":"1992","journal-title":"Computational Linguistics"},{"key":"S1351324915000030_ref028","first-page":"707","article-title":"Binary codes capable of correcting deletions, insertions, and reversals","volume":"10","author":"Levenshtein","year":"1966","journal-title":"Soviet Physics Doklady"},{"key":"S1351324915000030_ref042","unstructured":"Stolcke A. , Zheng J. , Wang W. , and Abrash V. 2011. SRILM at sixteen: update and outlook. In Proceedings of the IEEE Automatic Speech Recognition and Understanding Workshop, Waikoloa, Hawaii."},{"key":"S1351324915000030_ref007","doi-asserted-by":"crossref","unstructured":"Brill E. , and Moore R. C. 2000. An improved error model for noisy channel spelling correction. In Proceedings of the 38th Annual Meeting of the Association for Computational Linguistics, Hong Kong, pp. 286\u2013293.","DOI":"10.3115\/1075218.1075255"},{"key":"S1351324915000030_ref019","first-page":"368","volume-title":"Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics","author":"Han","year":"2011"},{"key":"S1351324915000030_ref004","unstructured":"Attia M. , Pecina P. , Tounsi L. , Toral A. , and van Genabith J. 2011. An Open-source finite state morphological transducer for modern standard arabic. In International Workshop on Finite State Methods and Natural Language Processing (FSMNLP), Blois, France, pp. 125\u2013133."},{"key":"S1351324915000030_ref011","doi-asserted-by":"publisher","DOI":"10.1007\/s10032-007-0054-0"},{"key":"S1351324915000030_ref018","first-page":"53","volume-title":"Proceedings of the 4th Workshop on Treebanks and Linguistic Theories (TLT)","author":"Haji\u010d","year":"2005"},{"key":"S1351324915000030_ref017","doi-asserted-by":"publisher","DOI":"10.1142\/S0219427907001706"},{"key":"S1351324915000030_ref015","first-page":"358","volume-title":"Proceedings of the 23rd International Conference on Computational Linguistics","author":"Gao","year":"2010"},{"key":"S1351324915000030_ref005","doi-asserted-by":"publisher","DOI":"10.3115\/1621753.1621763"},{"key":"S1351324915000030_ref045","doi-asserted-by":"crossref","unstructured":"van Delden S. , Bracewell D. B. , and Gomez F. 2004. Supervised and unsupervised automatic spelling correction algorithms. In Proceedings of the 2004 IEEE International Conference on Web Services, pp. 530\u2013535.","DOI":"10.1109\/IRI.2004.1431515"},{"key":"S1351324915000030_ref038","doi-asserted-by":"publisher","DOI":"10.3115\/1557690.1557721"}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324915000030","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,7]],"date-time":"2024-06-07T22:32:23Z","timestamp":1717799543000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324915000030\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,3,18]]},"references-count":51,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2016,9]]}},"alternative-id":["S1351324915000030"],"URL":"https:\/\/doi.org\/10.1017\/s1351324915000030","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2015,3,18]]}}}