{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T17:17:39Z","timestamp":1784567859254,"version":"3.55.0"},"reference-count":39,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2021,12,13]],"date-time":"2021-12-13T00:00:00Z","timestamp":1639353600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Deanship of Scientific Research at King Saud University","award":["RG-1441-332"],"award-info":[{"award-number":["RG-1441-332"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2022,5,31]]},"abstract":"<jats:p>Diacritic restoration (also known as diacritization or vowelization) is the process of inserting the correct diacritical markings into a text. Modern Arabic is typically written without diacritics, e.g., newspapers. This lack of diacritical markings often causes ambiguity, and though natives are adept at resolving, there are times they may fail. Diacritic restoration is a classical problem in computer science. Still, as most of the works tackle the full (heavy) diacritization of text, we, however, are interested in diacritizing the text using a fewer number of diacritics. Studies have shown that a fully diacritized text is visually displeasing and slows down the reading. This article proposes a system to diacritize homographs using the least number of diacritics, thus the name \u201clight.\u201d There is a large class of words that fall under the homograph category, and we will be dealing with the class of words that share the spelling but not the meaning. With fewer diacritics, we do not expect any effect on reading speed, while eye strain is reduced. The system contains morphological analyzer and context similarities. The morphological analyzer is used to generate all word candidates for diacritics. Then, through a statistical approach and context similarities, we resolve the homographs. Experimentally, the system shows very promising results, and our best accuracy is 85.6%.<\/jats:p>","DOI":"10.1145\/3486675","type":"journal-article","created":{"date-parts":[[2021,12,13]],"date-time":"2021-12-13T14:36:25Z","timestamp":1639406185000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Light Diacritic Restoration to Disambiguate Homographs in Modern Arabic Texts"],"prefix":"10.1145","volume":"21","author":[{"given":"Aqil M.","family":"Azmi","sequence":"first","affiliation":[{"name":"King Saud University, Riyadh, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rehab M.","family":"Alnefaie","sequence":"additional","affiliation":[{"name":"King Saud University, Saudi Arabia and Prince Sattam bin Abdulaziz University, Al Kharj, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hatim A.","family":"Aboalsamh","sequence":"additional","affiliation":[{"name":"King Saud University, Riyadh, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,12,13]]},"reference":[{"issue":"4","key":"e_1_3_2_2_2","first-page":"964","article-title":"Homonymy in English and Arabic: A contrastive study","volume":"18","author":"Abdul-Ameer Ahmed Mohammed Ali","year":"2010","unstructured":"Ahmed Mohammed Ali Abdul-Ameer and Areej As\u2018ad Ja\u2018far Altaie. 2010. Homonymy in English and Arabic: A contrastive study. J. Univ. Babylon 18, 4 (2010), 964\u2013984.","journal-title":"J. Univ. Babylon"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10586-017-0918-0"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007906222227"},{"key":"e_1_3_2_5_2","volume-title":"Dirasat Fi Fiqh Al-lughah (in Arabic)","author":"Al-Salih Subhi","year":"1981","unstructured":"Subhi Al-Salih. 1981. Dirasat Fi Fiqh Al-lughah (in Arabic). Dar al-\u2018Ilm lil-Malayin, Beirut, Lebanon."},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.3009217"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2017.10.106"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSP-SPE.2013.6642556"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/AEECT.2017.8257765"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1007\/s13369-019-04024-0"},{"key":"e_1_3_2_11_2","volume-title":"A large-scale computational processor of the Arabic morphology and applications","author":"Attia M.","year":"2000","unstructured":"M. Attia. 2000. A large-scale computational processor of the Arabic morphology and applications. Master\u2019s Thesis. Faculty of Engineering, Cairo University, Giza, Egypt."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-019-09692-w"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10209-017-0522-3"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324913000284"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.chb.2014.07.003"},{"key":"e_1_3_2_16_2","first-page":"295","volume-title":"Proceedings of the 3rd International WordNet Conference","author":"Black William","year":"2006","unstructured":"William Black, Sabri Elkateb, Horacio Rodriguez, Musa Alkhalifa, Piek Vossen, Adam Pease, and Christiane Fellbaum. 2006. Introducing the Arabic wordnet project. In Proceedings of the 3rd International WordNet Conference. 295\u2013300."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jksuci.2016.05.002"},{"key":"e_1_3_2_18_2","volume-title":"AraMorph","author":"Brihaye P.","unstructured":"P. Brihaye. [n. d.]. AraMorph. Retrieved from http:\/\/www.nongnu.org\/aramorph\/english\/index.html."},{"key":"e_1_3_2_19_2","volume-title":"Buckwalter Arabic Morphological Analyzer Version 2.0. Linguistic Data Consortium (LDC)","author":"Buckwalter Tim","year":"2004","unstructured":"Tim Buckwalter. 2004. Buckwalter Arabic Morphological Analyzer Version 2.0. Linguistic Data Consortium (LDC), University of Pennsylvania. Technical Report."},{"key":"e_1_3_2_20_2","first-page":"1349","volume-title":"Proceedings of the 11th International Conference on Language Resources and Evaluation (LREC\u201918)","author":"Gorman Kyle","year":"2018","unstructured":"Kyle Gorman, Gleb Mazovetskiy, and Vitaly Nikolaev. 2018. Improving homograph disambiguation with supervised machine learning. In Proceedings of the 11th International Conference on Language Resources and Evaluation (LREC\u201918). 1349\u20131352."},{"issue":"1","key":"e_1_3_2_21_2","doi-asserted-by":"crossref","first-page":"27","DOI":"10.21248\/jlcl.32.2017.213","article-title":"A survey and comparative study of Arabic diacritization tools","volume":"32","author":"Hamed Osama","year":"2017","unstructured":"Osama Hamed and Torsten Zesch. 2017. A survey and comparative study of Arabic diacritization tools. J. Lang. Technol. Comput. Ling. 32, 1 (2017), 27\u201347.","journal-title":"J. Lang. Technol. Comput. Ling."},{"key":"e_1_3_2_22_2","volume-title":"Aspects of word and sentence processing during reading Arabic: Evidence from eye movements","author":"Hermena Ehab","year":"2016","unstructured":"Ehab Hermena. 2016. Aspects of word and sentence processing during reading Arabic: Evidence from eye movements. Ph.D. Dissertation. University of Southampton."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2018.2865098"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2019.2933721"},{"key":"e_1_3_2_25_2","first-page":"328","volume-title":"Proceedings of the 7th International Conference on Software Paradigm Trends (ICSOFT\u201912)","author":"Jani F.","year":"2012","unstructured":"F. Jani and A. H. Pilevar. 2012. Word sense disambiguation of Persian homographs. In Proceedings of the 7th International Conference on Software Paradigm Trends (ICSOFT\u201912). 328\u2013331."},{"key":"e_1_3_2_26_2","article-title":"Stemming arabic text","author":"Khoja Shereen","year":"1999","unstructured":"Shereen Khoja and Roger Garside. 1999. Stemming arabic text. Computing Department, Lancaster University, Lancaster, UK.","journal-title":"Computing Department, Lancaster University, Lancaster, UK"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4020-6046-5_12"},{"key":"e_1_3_2_28_2","first-page":"466","volume-title":"Proceedings of the NEMLAR Conference on Arabic Language Resources and Tools","volume":"27","author":"Maamouri Mohamed","year":"2004","unstructured":"Mohamed Maamouri, Ann Bies, Tim Buckwalter, and Wigdan Mekki. 2004. The Penn Arabic treebank: Building a large-scale annotated Arabic corpus. In Proceedings of the NEMLAR Conference on Arabic Language Resources and Tools, Vol. 27. 466\u2013467."},{"key":"e_1_3_2_29_2","volume-title":"LDC Standard Arabic Morphological Analyzer (SAMA) Version 3.1","author":"Maamouri Mohamed","year":"2010","unstructured":"Mohamed Maamouri, David Graff, Basma Bouziri, Sondos Krouna, Ann Bies, and Seth Kulick. 2010. LDC Standard Arabic Morphological Analyzer (SAMA) Version 3.1. Retrieved from https:\/\/catalog.ldc.upenn.edu\/LDC2010L01."},{"key":"e_1_3_2_30_2","first-page":"730","volume-title":"Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP\u201919)","author":"Masmoudi Abir","year":"2019","unstructured":"Abir Masmoudi, Mariem Ellouze Khemekhem, et\u00a0al. 2019. Automatic diacritization of Tunisian dialect text using recurrent neural network. In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP\u201919). 730\u2013739."},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3297278"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.5555\/1621787.1621802"},{"key":"e_1_3_2_33_2","first-page":"1094","volume-title":"Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201914)","volume":"14","author":"Pasha Arfath","year":"2014","unstructured":"Arfath Pasha, Mohamed Al-Badrashiny, Mona T. Diab, Ahmed El Kholy, Ramy Eskander, Nizar Habash, Manoj Pooleery, Owen Rambow, and Ryan Roth. 2014. MADAMIRA: A fast, comprehensive tool for morphological analysis and disambiguation of Arabic. In Proceedings of the International Conference on Language Resources and Evaluation (LREC\u201914), Vol. 14. 1094\u20131101."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.5555\/2817174.2817183"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-444-70113-8.50064-3"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","first-page":"441","DOI":"10.1007\/978-0-387-44641-7_46","volume-title":"Intelligent Information Processing III","author":"Shaalan Khaled","year":"2007","unstructured":"Khaled Shaalan, Azza Abdel Monem, and Ahmed Rafea. 2007. Arabic morphological generation from interlingua: A rule-based approach. In Intelligent Information Processing III, Zhongzhi Shi, K. Shimohara, and D. Feng (Eds.). Springer U.S., 441\u2013451."},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1142\/9789813229396_0003"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICINIS.2013.69"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCAIE.2016.7575067"},{"key":"e_1_3_2_40_2","unstructured":"T. Zerrouki. 2014. Arabic Corpora Resources Tashkila Collection from the Arabic Al-Shamela Library."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3486675","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3486675","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:18:46Z","timestamp":1750191526000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3486675"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,13]]},"references-count":39,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,5,31]]}},"alternative-id":["10.1145\/3486675"],"URL":"https:\/\/doi.org\/10.1145\/3486675","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12,13]]},"assertion":[{"value":"2020-05-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-12-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}