{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T22:13:12Z","timestamp":1780611192956,"version":"3.54.1"},"reference-count":24,"publisher":"Association for Computing Machinery (ACM)","issue":"8","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2025,8,31]]},"abstract":"<jats:p>Word Sense Disambiguation (WSD) poses a significant challenge in Natural Language Processing (NLP), particularly for languages with complex morphology and semantics like Urdu. In this article, we present Multi-layered FrAmeworK for WSD (MAKS), a comprehensive framework designed to address the nuances of Urdu WSD. MAKS comprises four integral layers: corpus creation, pre-processing, transfer layer, and classifier layer. The Corpus Creation stage involves the compilation of a diverse and representative corpus of Urdu text data. Pre-processing entails standardization and normalization of raw text through tokenization and other techniques. The Transfer Layer utilizes advanced methods such as XLM tokenization and XLM-RoBERTa for feature extraction and data preparation. Finally, the Classifier Layer employs an ensemble approach with Support Vector Machines (SVM) and Random Forests (RF) to classify tokenized data and enhance WSD accuracy. Through a systematic integration of these layers, MAKS achieves robustness and context-awareness in resolving word senses within Urdu text. Experimental evaluations demonstrate the accuracy and effectiveness of MAKS in addressing the unique challenges of Urdu WSD. Overall, MAKS offers a versatile and powerful framework for advancing Urdu language processing and understanding.<\/jats:p>","DOI":"10.1145\/3748319","type":"journal-article","created":{"date-parts":[[2025,7,15]],"date-time":"2025-07-15T11:27:16Z","timestamp":1752578836000},"page":"1-12","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Breaking Barriers in URDU WSD: The Transfer Learning Enriched MAKS Framework"],"prefix":"10.1145","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4952-8693","authenticated-orcid":false,"given":"Sarfraz","family":"Bibi","sequence":"first","affiliation":[{"name":"UIIT, Arid Agriculture University","place":["Rawalpindi, Pakistan"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6883-3584","authenticated-orcid":false,"given":"Sohail","family":"Asghar","sequence":"additional","affiliation":[{"name":"Computer Science, COMSATS University Islamabad","place":["Islamabad, Pakistan"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3999-6581","authenticated-orcid":false,"given":"Muhammad","family":"Zubair","sequence":"additional","affiliation":[{"name":"Riphah International University","place":["Islamabad, Pakistan"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,8,21]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10586-017-0918-0"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.4236\/jdaip.2020.84020"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2019.10.118"},{"key":"e_1_3_2_5_2","first-page":"748","volume-title":"Proceedings of the 28th International Conference on Neural Information Processing, ICONIP 2021, Part V 28","author":"Chawla Avi","year":"2021","unstructured":"Avi Chawla, Nidhi Mulay, Vikas Bishnoi, Gaurav Dhama, and Anil Kumar Singh. 2021. A comparative study of transformers on word sense disambiguation. In Proceedings of the 28th International Conference on Neural Information Processing, ICONIP 2021, Part V 28. Springer, 748\u2013756."},{"key":"e_1_3_2_6_2","first-page":"1","volume-title":"Proceedings of the 2022 IEEE International Conference on Blockchain and Distributed Systems Security (ICBDS)","author":"Choure Aditya A.","year":"2022","unstructured":"Aditya A. Choure, Rahul B. Adhao, and Vinod K. Pachghare. 2022. NER in Hindi language using transformer model: XLM-roberta. In Proceedings of the 2022 IEEE International Conference on Blockchain and Distributed Systems Security (ICBDS). IEEE, 1\u20135."},{"key":"e_1_3_2_7_2","volume-title":"Proceedings of the 2017 International Electrical Engineering Congress (iEECON)","author":"Dar Kamran Shaukat","year":"2017","unstructured":"Kamran Shaukat Dar, Ahmad Bin Shafat, and Muhammad Umair Hassan. 2017. An efficient stop word elimination algorithm for Urdu language. In Proceedings of the 2017 International Electrical Engineering Congress (iEECON). IEEE."},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-56485-8","volume-title":"Random Forests","author":"Genuer Robin","year":"2020","unstructured":"Robin Genuer, Jean-Michel Poggi, Robin Genuer, and Jean-Michel Poggi. 2020. Random Forests. Springer."},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","unstructured":"Abdul Ghafoor Ali Shariq Imran Sher Muhammad Daudpota Zenun Kastrati Rakhi Batra and Mudasir Ahmad Wani. 2021. The impact of translating resource-rich datasets to low-resource languages through multi-lingual text processing. IEEE Access 9 (2021) 124478\u2013124490.","DOI":"10.1109\/ACCESS.2021.3110285"},{"key":"e_1_3_2_10_2","first-page":"2938","volume-title":"Proceedings of the 9th International Conference on Language Resources and Evaluation, LREC 2014","author":"Jawaid B.","year":"2014","unstructured":"B. Jawaid, A. Kamran, and O. Bojar. 2014. A tagged corpus and a tagger for Urdu. In Proceedings of the 9th International Conference on Language Resources and Evaluation, LREC 2014. 2938\u20132943."},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"131","DOI":"10.1007\/978-981-13-8759-3_5","article-title":"Random forest-based sarcastic tweet classification using multiple feature collection","author":"Kumar Rajeev","year":"2020","unstructured":"Rajeev Kumar and Jasandeep Kaur. 2020. Random forest-based sarcastic tweet classification using multiple feature collection. Multimedia Big Data Computing for IoT Applications: Concepts, Paradigms and Solutions (2020), 131\u2013160.","journal-title":"Multimedia Big Data Computing for IoT Applications: Concepts, Paradigms and Solutions"},{"key":"e_1_3_2_12_2","doi-asserted-by":"crossref","unstructured":"R. Steinberger and B. Pouliquen. 2007. Cross-lingual named entity recognition. Lingvistic\u00e6 Investigationes 30 1 (2021) 135\u2013162.","DOI":"10.1075\/li.30.1.09ste"},{"key":"e_1_3_2_13_2","article-title":"Supervised Word Sense Disambiguation for Urdu Using Bayesian Classification","author":"Naseer A.","year":"2009","unstructured":"A. Naseer and S. Hussain. 2009. Supervised Word Sense Disambiguation for Urdu Using Bayesian Classification. Center for Research in Urdu Language Processing, Lahore, Pakistan .","journal-title":"Center for Research in Urdu Language Processing, Lahore, Pakistan"},{"issue":"1","key":"e_1_3_2_14_2","first-page":"45","article-title":"A comparative study of syntax in Punjabi and Urdu","volume":"1","author":"Naz Bushra","year":"2022","unstructured":"Bushra Naz and Aamir Mahmood. 2022. A comparative study of syntax in Punjabi and Urdu. Cosmic Journal of Linguistics 1, 1 (2022), 45\u201358.","journal-title":"Cosmic Journal of Linguistics"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007692713085"},{"key":"e_1_3_2_16_2","first-page":"243","volume-title":"Proceedings of the International Conference on Applications of Natural Language to Information Systems","author":"Pan Ronghao","year":"2023","unstructured":"Ronghao Pan, Jos\u00e9 Antonio Garc\u00eda-D\u00edaz, and Rafael Valencia-Garc\u00eda. 2023. Evaluation of transformer-based models for punctuation and capitalization restoration in Spanish and Portuguese. In Proceedings of the International Conference on Applications of Natural Language to Information Systems. Springer, 243\u2013256."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-12-815739-8.00006-7"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1120"},{"key":"e_1_3_2_19_2","first-page":"730","volume-title":"Proceedings of the 2022 International Conference on Applied Artificial Intelligence and Computing (ICAAIC)","author":"Rathod Naitik","year":"2022","unstructured":"Naitik Rathod, Nishit Mistry, Dhruv Talati, Manan Parikh, Aniket Kore, and Pratik Kanani. 2022. Marathi social media opinion mining using XLM-R. In Proceedings of the 2022 International Conference on Applied Artificial Intelligence and Computing (ICAAIC). IEEE, 730\u2013736."},{"key":"e_1_3_2_20_2","first-page":"19","volume-title":"Developing Resources and Techniques for Urdu Word Sense Disambiguation","author":"Saeed A.","year":"2019","unstructured":"A. Saeed. 2019. Data and methods for Urdu lexical sample word sense disambiguation. In Developing Resources and Techniques for Urdu Word Sense Disambiguation. N. U. Din, M. Raza, and A. Hassan (Eds.), Springer International Publishing, 19\u201340."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-018-9438-7"},{"key":"e_1_3_2_22_2","first-page":"461","volume-title":"Proceedings of the 2nd International Conference on Intelligent Technologies and Applications, INTAP 2019, Revised Selected Papers","volume":"1198","author":"Sarim Muhammad","year":"2020","unstructured":"Muhammad Sarim. 2020. Urdu natural language processing issues and challenges: A review study. In Proceedings of the 2nd International Conference on Intelligent Technologies and Applications, INTAP 2019, Revised Selected Papers, Vol. 1198. Springer Nature, 461."},{"issue":"1","key":"e_1_3_2_23_2","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1007\/s41133-020-00032-0","article-title":"A comparative analysis of logistic regression, random forest and KNN models for the text classification","volume":"5","author":"Shah Kanish","year":"2020","unstructured":"Kanish Shah, Henil Patel, Devanshi Sanghvi, and Manan Shah. 2020. A comparative analysis of logistic regression, random forest and KNN models for the text classification. Augmented Human Research 5, 1 (2020), 12.","journal-title":"Augmented Human Research"},{"issue":"5103","key":"e_1_3_2_24_2","article-title":"Developing an Urdu lemmatizer using a dictionary-based lookup approach","volume":"13","author":"Shaukat Saima","year":"2023","unstructured":"Saima Shaukat, Muhammad Asad, and Asmara Akram. 2023. Developing an Urdu lemmatizer using a dictionary-based lookup approach. Applied Sciences 13, 5103 (2023), 1\u201313.","journal-title":"Applied Sciences"},{"key":"e_1_3_2_25_2","volume-title":"Proceedings of the Conference on Language and Technology","author":"Urooj S.","year":"2014","unstructured":"S. Urooj, S. Shams, S. Hussain, and F. Adeeba. 2014. Sense tagged CLE Urdu digest corpus. In Proceedings of the Conference on Language and Technology."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3748319","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T13:28:01Z","timestamp":1755782881000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3748319"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,21]]},"references-count":24,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2025,8,31]]}},"alternative-id":["10.1145\/3748319"],"URL":"https:\/\/doi.org\/10.1145\/3748319","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,21]]},"assertion":[{"value":"2024-05-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-31","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}