{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T04:14:16Z","timestamp":1779336856676,"version":"3.51.4"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"10","license":[{"start":{"date-parts":[[2023,10,13]],"date-time":"2023-10-13T00:00:00Z","timestamp":1697155200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,10,31]]},"abstract":"<jats:p>Text pre-processing is a crucial step in Natural Language Processing (NLP) applications, particularly for handling informal and noisy content on social media. Word-level tokenization plays a vital role in text pre-processing by removing stop words, filtering irrelevant characters, and retaining relevant tokens. These tokens are essential for constructing meaningful n-grams within advanced NLP frameworks used for data modeling. However, tokenization in low-resource languages like Urdu presents challenges due to language complexity and limited resources. Conventional space-based methods and direct application of language-specific tools often result in erroneous tokens in Urdu Language Processing (ULP). This hinders language models from effectively learning language-specific and domain-specific tokens, leading to sub-optimal results for downstream tasks such as aspect mining, topic modeling, and Named Entity Recognition (NER). To address this issue for Urdu, we have proposed a data pre-processing technique that detects outliers using the Inter-Quartile-Range (IQR) method and proposed normalization algorithms for creating useful lexicons in conjunction with existing technologies. We have collected approximately 50 million Urdu tweets using the Twitter API and conducted the performance analysis of existing language-specific tokenizers (Urduhack and Space-based tokenizer). Dataset variants were created based on the language-specific tokenizers, and we performed statistical analysis tests and visualization techniques to compare tokenization results before and after applying the proposed outlier detection and normalization method. Our findings highlighted the noticeable improvement in token size distributions, handling of informal language tokens, and misspelled and lengthy tokens. The Urduhack tokenizer combined with the proposed outlier detection and normalization yielded tokens with the best-fitted distribution in ULP. Its effectiveness has been evaluated through the task of topic modeling using Non-negative Matrix Factorization (NMF) and Latent Dirichlet allocation (LDA). The results demonstrated new and distinct topics using unigram features while achieving highly coherent topics when utilizing bigram features. For the traditional space-based method, the results consistently demonstrated improved coherence and precision scores. However, the NMF topic modeling with bigram features outperformed LDA topic modeling with bigram features.<\/jats:p>","DOI":"10.1145\/3622939","type":"journal-article","created":{"date-parts":[[2023,9,8]],"date-time":"2023-09-08T12:09:47Z","timestamp":1694174987000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Assessing Urdu Language Processing Tools via Statistical and Outlier Detection Methods on Urdu Tweets"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9404-3903","authenticated-orcid":false,"family":"Zoya","sequence":"first","affiliation":[{"name":"School of Electrical Engineering and Computer Science (SEECS), National University of Sciences and Technology (NUST), Pakistan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5801-1568","authenticated-orcid":false,"given":"Seemab","family":"Latif","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering and Computer Science (SEECS), National University of Sciences and Technology (NUST), Pakistan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5304-5948","authenticated-orcid":false,"given":"Rabia","family":"Latif","sequence":"additional","affiliation":[{"name":"Prince Sultan University, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7770-5899","authenticated-orcid":false,"given":"Hammad","family":"Majeed","sequence":"additional","affiliation":[{"name":"FAST National University of Computer and Emerging Sciences, Pakistan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0392-6506","authenticated-orcid":false,"given":"Nor Shahida Mohd","family":"Jamail","sequence":"additional","affiliation":[{"name":"Prince Sultan University, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,10,13]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Syed Zain Abbas Dr Rahman Abdul Basit Mughal Syed Mujtaba Haider et\u00a0al. 2022. Urdu news article recommendation model using natural language processing techniques. arxiv:2206.11862. Retrieved from https:\/\/arxiv.org\/abs\/2206.11862"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1080\/17517575.2020.1755455"},{"key":"e_1_3_2_4_2","unstructured":"Ikram ALi. 2020. Urduhack: A Python Library for Urdu Language Processing. Retrieved from https:\/\/docs.urduhack.com\/en\/stable\/#urduhack."},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3112500"},{"key":"e_1_3_2_6_2","volume-title":"Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics Companion Volume Proceedings of the Demo and Poster Sessions","author":"Ananiadou Sophia","year":"2007","unstructured":"Sophia Ananiadou (Ed.). 2007. Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics Companion Volume Proceedings of the Demo and Poster Sessions. Association for Computational Linguistics."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177729437"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-020-05321-8"},{"key":"e_1_3_2_9_2","unstructured":"Gerlof Bouma. 2009. Normalized (pointwise) mutual information in collocation extraction. In Proceedings of the Conference of the German Society for Computational Linguistics and Language Technology Universit\u00e4t Potsdam Potsdam Tubingen 31\u201340. https:\/\/api.semanticscholar.org\/CorpusID:2762657"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.2307\/2684359"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.3390\/electronics12122662"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-016-9482-x"},{"key":"e_1_3_2_13_2","unstructured":"Ozan Dogan. 2021. Fitter: Fitter - Fits Distribution to Data in. One Line of Code. https:\/\/github.com\/fittercommunity\/fitter"},{"key":"e_1_3_2_14_2","unstructured":"Jack Dorsey. 2022. Twitter by the Numbers: Stats Demographics & Fun Facts. Retrieved from https:\/\/www.omnicoreagency.com\/twitter-statistics\/#::text=Twitter"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177731944"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/INMIC48123.2019.9022736"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3310205"},{"key":"e_1_3_2_18_2","series-title":"Working Notes of FIRE 2021\u2014Forum for Information Retrieval Evaluation","first-page":"799","volume":"3159","author":"Kalra Sakshi","year":"2021","unstructured":"Sakshi Kalra, Yash Bansal, and Yashvardhan Sharma. 2021. Detection of abusive records by analyzing the tweets in Urdu language exploring transformer based models. In Working Notes of FIRE 2021\u2014Forum for Information Retrieval Evaluation(CEUR Workshop Proceedings, Vol. 3159), Parth Mehta, Thomas Mandl, Prasenjit Majumder, and Mandar Mitra (Eds.). CEUR-WS.org, 799\u2013805."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1177\/01655515221137270"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3093078"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eij.2020.04.003"},{"key":"e_1_3_2_22_2","series-title":"Working Notes of FIRE 2020\u2014Forum for Information Retrieval Evaluation (FIRE-WN\u201920)","first-page":"452","volume":"2826","author":"Khiljia Abdullah Faiz Ur Rahman","year":"2020","unstructured":"Abdullah Faiz Ur Rahman Khiljia, Sahinur Rahman Laskara, Partha Pakraya, and Sivaji Bandyopadhyaya. 2020. Urdu fake news detection using generalized autoregressors. In Working Notes of FIRE 2020\u2014Forum for Information Retrieval Evaluation (FIRE-WN\u201920)(CEUR Workshop Proceedings, Vol. 2826), Parth Mehta, Thomas Mandl, Prasenjit Majumder, and Mandar Mitra (Eds.). CEUR-WS.org, 452\u2013457."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.2307\/2280779"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.7717\/peerj-cs.713"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.2307\/2529310"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177730491"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.3390\/app11156769"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.18637\/jss.v008.i18"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.5555\/2390524.2390572"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12559-017-9481-5"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3104308"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-demos.14"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICoDT252288.2021.9441524"},{"key":"e_1_3_2_34_2","first-page":"79","volume-title":"Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing: Student Research Workshop","author":"Rani Sadaf","year":"2020","unstructured":"Sadaf Rani and Muhammad Waqas Anwar. 2020. Resource creation and evaluation of aspect based sentiment analysis in Urdu. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing: Student Research Workshop. Association for Computational Linguistics, 79\u201384."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDIM.2018.8847044"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/IALP.2012.11"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICoDT252288.2021.9441508"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/2684822.2685324"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/BESC.2018.8697243"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.2307\/2333709"},{"key":"e_1_3_2_41_2","first-page":"37","volume-title":"Mexican International Conference on Artificial Intelligence","author":"Haq Ehsan ul","year":"2010","unstructured":"Ehsan ul Haq, Sahar Rauf, Sarmad Hussain, and Kashif Javed. 2010. Corpus of aspect-based sentiment for Urdu political data. In Mexican International Conference on Artificial Intelligence. Springer, 37\u201340."},{"key":"e_1_3_2_42_2","volume-title":"Natural Language Processing with Python and SpaCy: A Practical Introduction","author":"Vasiliev Yuli","year":"2020","unstructured":"Yuli Vasiliev. 2020. Natural Language Processing with Python and SpaCy: A Practical Introduction. No Starch Press, San Francisco, CA."},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1051\/itmconf\/20182300037"},{"key":"e_1_3_2_44_2","first-page":"29","article-title":"Chinese word segmentation as character tagging","volume":"8","author":"Xue Nianwen","year":"2003","unstructured":"Nianwen Xue. 2003. Chinese word segmentation as character tagging. Int. J. Comput. Ling. Chin. Lang. Process. 8 (Feb. 2003), 29\u201348.","journal-title":"Int. J. Comput. Ling. Chin. Lang. Process."},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3112620"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3622939","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3622939","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:57:28Z","timestamp":1750298248000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3622939"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,13]]},"references-count":44,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2023,10,31]]}},"alternative-id":["10.1145\/3622939"],"URL":"https:\/\/doi.org\/10.1145\/3622939","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,13]]},"assertion":[{"value":"2023-05-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-08-24","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}