{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T13:42:09Z","timestamp":1783950129644,"version":"3.55.0"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,3,24]],"date-time":"2023-03-24T00:00:00Z","timestamp":1679616000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,4,30]]},"abstract":"<jats:p>Studies on natural language processing are mainly conducted in English, with very few exploring languages that are under-resourced, including the Dravidian languages. We present a novel work in detecting offensive language using a corpus collected from YouTube containing comments in Tamil. The study specifically aims to compare two machine learning approaches\u2014namely, supervised and unsupervised\u2014to detect offensive patterns in textual communications. In the first setup, offensive language detection models were developed using traditional machine learning algorithms such as Random Forest, Logistic Regression, Support Vector Machine, and AdaBoost, and assessed based on human labeling. Conversely, we used<jats:italic>K<\/jats:italic>-means (<jats:italic>K<\/jats:italic>= 2) to cluster the unlabeled data before training the same set of machine learning algorithms to detect offensive communications. Performance scores indicate unsupervised clustering to be more effective than human labeling with ensemble classifiers achieving an impressive accuracy of 99.70% and 99.87% respectively for balanced and imbalanced datasets, hence showing that the unsupervised approach can be used effectively to detect offensive language in low-resourced languages.<\/jats:p>","DOI":"10.1145\/3575860","type":"journal-article","created":{"date-parts":[[2022,12,16]],"date-time":"2022-12-16T06:43:52Z","timestamp":1671173032000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Tamil Offensive Language Detection: Supervised versus Unsupervised Learning Approaches"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6859-4488","authenticated-orcid":false,"given":"Vimala","family":"Balakrishnan","sequence":"first","affiliation":[{"name":"Universiti Malaya, Malaysia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1875-820X","authenticated-orcid":false,"given":"Vithyatheri","family":"Govindan","sequence":"additional","affiliation":[{"name":"Universiti Malaya, Malaysia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3456-3472","authenticated-orcid":false,"given":"Kumanan N.","family":"Govaichelvan","sequence":"additional","affiliation":[{"name":"Universiti Malaya, Malaysia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,3,24]]},"reference":[{"key":"e_1_3_2_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.osnem.2020.100096"},{"issue":"6","key":"e_1_3_2_3_1","doi-asserted-by":"crossref","first-page":"e138820","DOI":"10.24425\/bpasts.2021.138820","article-title":"Deep learning-based Tamil Parts of Speech (POS) tagger","volume":"69","author":"Anbukkarasi S.","year":"2021","unstructured":"S. Anbukkarasi and S. Varadhaganapathy. 2021. Deep learning-based Tamil Parts of Speech (POS) tagger. Bulletin of the Polish Academy of Sciences: Technical Sciences 69, 6 (2021), e138820\u2013e138820.","journal-title":"Bulletin of the Polish Academy of Sciences: Technical Sciences"},{"key":"e_1_3_2_4_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cosrev.2020.100311"},{"key":"e_1_3_2_5_1","article-title":"IIITG-ADBU@ HASOC-Dravidian-CodeMix-FIRE2020: Offensive content detection in code-mixed Dravidian text","author":"Baruah A.","year":"2021","unstructured":"A. Baruah, K. A. Das, F. A. Barbhuiya, and K. Dey. 2021. IIITG-ADBU@ HASOC-Dravidian-CodeMix-FIRE2020: Offensive content detection in code-mixed Dravidian text. arXiv preprint arXiv:2107.14336 (2021).","journal-title":"arXiv preprint"},{"key":"e_1_3_2_6_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S19-2007"},{"key":"e_1_3_2_7_1","first-page":"1","volume-title":"Proceedings of the 2017 14th IEEE India Council International Conference (INDICON\u201917)","author":"Bharti S. K.","year":"2017","unstructured":"S. K. Bharti, R. Naidu, and K. S. Babu. 2017. Hyperbolic feature-based sarcasm detection in tweets: A machine learning approach. In Proceedings of the 2017 14th IEEE India Council International Conference (INDICON\u201917). IEEE, Los Alamitos, CA, 1\u20136."},{"key":"e_1_3_2_8_1","doi-asserted-by":"publisher","DOI":"10.4236\/jdaip.2020.84020"},{"key":"e_1_3_2_9_1","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1007\/978-3-030-63475-9_3","volume-title":"An Anatomy of Chinese Offensive Words","author":"Carson L.","year":"2021","unstructured":"L. Carson and N. Jiang. 2021. Collecting and categorizing offensive words in Chinese. In An Anatomy of Chinese Offensive Words. Palgrave Macmillan, Cham, Switzerland, 53\u201365."},{"key":"e_1_3_2_10_1","volume-title":"Proceedings of the 9th Global Wordnet Conference","author":"Chakravarthi B. R.","year":"2018","unstructured":"B. R. Chakravarthi, M. Arcan, and J. P. McCrae. 2018. Improving wordnets for under-resourced languages using machine translation. In Proceedings of the 9th Global Wordnet Conference. Singapore, 77--86."},{"key":"e_1_3_2_11_1","article-title":"DravidianCodeMix: Sentiment analysis and offensive language identification dataset for Dravidian languages in code-mixed text","author":"Chakravarthi B. R.","year":"2021","unstructured":"B. R. Chakravarthi, R. Priyadharshini, V. Muralidaran, N. Jose, S. Suryawanshi, E. Sherly, and J. P. McCrae. 2021a. DravidianCodeMix: Sentiment analysis and offensive language identification dataset for Dravidian languages in code-mixed text. arXiv preprint arXiv:2106.09460 (2021).","journal-title":"arXiv preprint"},{"key":"e_1_3_2_12_1","article-title":"Dataset for identification of homophobia and transophobia in multilingual YouTube comments","author":"Chakravarthi B. R.","year":"2021","unstructured":"B. R. Chakravarthi, R. Priyadharshini, R. Ponnusamy, P. K. Kumaresan, K. Sampath, D. Thenmozhi, S. Thangasamy, R. Nallathambi, and J. P. McCrae. 2021b. Dataset for identification of homophobia and transophobia in multilingual YouTube comments. arXiv preprint arXiv:2109.00227 (2021).","journal-title":"arXiv preprint"},{"key":"e_1_3_2_13_1","article-title":"Corpus creation for sentiment analysis in code-mixed Tamil-English text","author":"Chakravarthi B. R.","year":"2020","unstructured":"B. R. Chakravarthi, V. Muralidaran, R. Priyadharshini, and J. P. McCrae. 2020a. Corpus creation for sentiment analysis in code-mixed Tamil-English text. arXiv preprint arXiv:2006.00206 (2020).","journal-title":"arXiv preprint"},{"key":"e_1_3_2_14_1","first-page":"57","volume-title":"Proceedings of the 7th Workshop on NLP for Similar Languages, Varieties, and Dialects","author":"Chakravarthi B. R.","year":"2020","unstructured":"B. R. Chakravarthi, N. Rajasekaran, M. Arcan, K. McGuinness, N. E. O'Connor, and J. P. McCrae. 2020b. Bilingual lexicon induction across orthographically-distinct under-resourced Dravidian languages. In Proceedings of the 7th Workshop on NLP for Similar Languages, Varieties, and Dialects. 57\u201369."},{"key":"e_1_3_2_15_1","first-page":"133","article-title":"Use of the word \u2018fuck\u2019 in pedagogy and higher learning","volume":"8","author":"Cusack C. M.","year":"2014","unstructured":"C. M. Cusack. 2014. Use of the word \u2018fuck\u2019 in pedagogy and higher learning. Journal of Law & Social Deviance 8 (2014), 133.","journal-title":"Journal of Law & Social Deviance"},{"key":"e_1_3_2_16_1","first-page":"721","volume-title":"Proceedings of the Future of Information and Communication Conference","author":"Das S.","year":"2020","unstructured":"S. Das, D. Venugopal, and S. Shiva. 2020. A holistic approach for detecting DDoS attacks by using ensemble unsupervised machine learning. In Proceedings of the Future of Information and Communication Conference. 721\u2013738."},{"key":"e_1_3_2_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-28553-1_10"},{"key":"e_1_3_2_18_1","unstructured":"D. M. Eberhard G. F. Simons and C. D. Fennig. 2019. Ethnologue: Languages of the World . SIL International. Available at https:\/\/www.ethnologue.com."},{"key":"e_1_3_2_19_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2019.102121"},{"key":"e_1_3_2_20_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2021.106662"},{"key":"e_1_3_2_21_1","doi-asserted-by":"publisher","DOI":"10.1177\/0038022920963328"},{"key":"e_1_3_2_22_1","first-page":"76","volume-title":"Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection","author":"Haddad B.","year":"2020","unstructured":"B. Haddad, Z. Orabe, A. Al-Abood, and N. Ghneim. 2020. Arabic offensive language detection with attention-based deep neural networks. In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection. 76\u201381."},{"key":"e_1_3_2_23_1","article-title":"Benchmarking multi-task learning for sentiment analysis and offensive language identification in under-resourced Dravidian languages","author":"Hande A.","year":"2021","unstructured":"A. Hande, S. U. Hegde, R. Priyadharshini, R. Ponnusamy, P. K. Kumaresan, S. Thavareesan, and B. R. Chakravarthi. 2021. Benchmarking multi-task learning for sentiment analysis and offensive language identification in under-resourced Dravidian languages. arXiv preprint arXiv:2108.03867 (2021).","journal-title":"arXiv preprint"},{"key":"e_1_3_2_24_1","doi-asserted-by":"publisher","DOI":"10.3390\/a13040083"},{"key":"e_1_3_2_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2018.08.169"},{"key":"e_1_3_2_26_1","doi-asserted-by":"publisher","DOI":"10.1177\/136248060200600406"},{"key":"e_1_3_2_27_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.12.063"},{"key":"e_1_3_2_28_1","first-page":"1","volume-title":"Proceedings of the 2017 10th International Conference on Contemporary Computing (IC3\u201917)","author":"Jain T.","year":"2017","unstructured":"T. Jain, N. Agrawal, G. Goyal, and N. Aggrawal. 2017. Sarcasm detection of tweets: A comparative study. In Proceedings of the 2017 10th International Conference on Contemporary Computing (IC3\u201917). IEEE, Los Alamitos, CA, 1\u20136."},{"key":"e_1_3_2_29_1","first-page":"141","volume-title":"Proceedings of the 2020 International Conference on Biomedical Innovations and Applications (BIA\u201920)","author":"Kalcheva N.","year":"2020","unstructured":"N. Kalcheva, M. Karova, and I. Penev. 2020. Comparison of the accuracy of SVM kernel functions in text classification. In Proceedings of the 2020 International Conference on Biomedical Innovations and Applications (BIA\u201920). IEEE, Los Alamitos, CA, 141\u2013145."},{"key":"e_1_3_2_30_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2019.07.086"},{"issue":"2","key":"e_1_3_2_31_1","doi-asserted-by":"crossref","first-page":"278","DOI":"10.1080\/01419870.2021.1964558","article-title":"Dirty food: Racism and casteism in India","volume":"45","author":"Kikon D.","year":"2022","unstructured":"D. Kikon. 2022. Dirty food: Racism and casteism in India. Ethnic and Racial Studies 45, 2 (2022), 278\u2013297.","journal-title":"Ethnic and Racial Studies"},{"key":"e_1_3_2_32_1","author":"Kocon J.","year":"2021","unstructured":"J. Kocon, A. Figas, M. Gruza, D. Puchalska, T. Kajdanowicz, and P. Kazienko. 2021. Offensive, aggressive, and hat speech analysis: From data-centric to human-centered approach. Information Processing and Management 58 (2021), 102643. https:\/\/doi.org\/10.1016\/j.ipm.2021.102643","journal-title":"Information Processing and Management"},{"key":"e_1_3_2_33_1","first-page":"1","volume-title":"Proceedings of the 1st Workshop on Trolling, Aggression, and Cyberbullying (TRAC-2018)","author":"Kumar R.","year":"2018","unstructured":"R. Kumar, A. K. Ojha, S. Malmasi, and M. Zampieri. 2018. Benchmarking aggression identification in social media. In Proceedings of the 1st Workshop on Trolling, Aggression, and Cyberbullying (TRAC-2018). 1\u201311."},{"key":"e_1_3_2_34_1","first-page":"1","volume-title":"Intelligent Systems, Technologies, and Applications","author":"Kumar S. S.","year":"2020","unstructured":"S. S. Kumar, M. A. Kumar, K. P. Soman, and P. Poornachandran. 2020. Dynamic mode-based feature with random mapping for sentiment analysis. In Intelligent Systems, Technologies, and Applications. Springer, Singapore, 1\u201315."},{"key":"e_1_3_2_35_1","doi-asserted-by":"publisher","DOI":"10.21744\/lingcure.v5nS1.1477"},{"key":"e_1_3_2_36_1","first-page":"138","volume-title":"Global English Slang","author":"Lambert J.","year":"2014","unstructured":"J. Lambert. 2014. Indian English slang. In Global English Slang. Routledge, 138\u2013146."},{"key":"e_1_3_2_37_1","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1145\/3441501.3441517","volume-title":"Proceedings of the Forum for Information Retrieval Evaluation (FIRE\u201920)","author":"Mandl T.","year":"2020","unstructured":"T. Mandl, S. Modha, M. A. Kumar, and B. R. Chakravarthi. 2020. Overview of the HASOC Track at FIRE 2020: Hate speech and offensive language identification in Tamil, Malayalam, Hindi, English and German. In Proceedings of the Forum for Information Retrieval Evaluation (FIRE\u201920). 29\u201332."},{"key":"e_1_3_2_38_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2021.08.050"},{"key":"e_1_3_2_39_1","unstructured":"C. Newton. 2019. The trauma floor. The Verge . Retrieved November 10 2021 from https:\/\/www.theverge.com\/2019\/2\/25\/18229714\/cognizant-facebook-content-moderator-interviews-trauma-working-conditions-arizona."},{"key":"e_1_3_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2872427.2883062"},{"key":"e_1_3_2_41_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2016.01.012"},{"key":"e_1_3_2_42_1","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-12-813314-9.00011-6"},{"key":"e_1_3_2_43_1","volume-title":"Proceedings of the 20th International Conference on Computational Linguistics and Intelligent Text Processing (CICLing\u201919)","author":"Sane K. R.","year":"2019","unstructured":"K. R. Sane, S. Kolla, S. R. Sane, V. K. Srirangam, and R. Mamidi. 2019. Corpus and baseline system for hate speech detection in Telugu-English code-mixed tweets. In Proceedings of the 20th International Conference on Computational Linguistics and Intelligent Text Processing (CICLing\u201919)."},{"key":"e_1_3_2_44_1","doi-asserted-by":"publisher","DOI":"10.1177\/1470785320921779"},{"key":"e_1_3_2_45_1","doi-asserted-by":"crossref","unstructured":"A. Schmidt and M. Wiegand. 2017. A survey on hate speech detection using natural language processing. In Proceedings of the 5th International Workshop on Natural Language Processing for Social Media Association for Computational Linguistics Valencia 1--10. https:\/\/www.aclweb.org\/anthology\/W17-1101.","DOI":"10.18653\/v1\/W17-1101"},{"key":"e_1_3_2_46_1","article-title":"NLP-CUET@ DravidianLangTech-EACL2021: Offensive language detection from multilingual code-mixed text using Transformers","author":"Sharif O.","year":"2021","unstructured":"O. Sharif, E. Hossain, and M. M. Hoque. 2021. NLP-CUET@ DravidianLangTech-EACL2021: Offensive language detection from multilingual code-mixed text using Transformers. arXiv preprint arXiv:2103.00455 (2021).","journal-title":"arXiv preprint"},{"issue":"22","key":"e_1_3_2_47_1","first-page":"433","article-title":"A comprehensive study on sarcasm detection techniques in sentiment analysis","volume":"118","author":"Sindhu C.","year":"2018","unstructured":"C. Sindhu, G. Vadivu, and M. V. Rao. 2018. A comprehensive study on sarcasm detection techniques in sentiment analysis. International Journal of Pure and Applied Mathematics 118, 22 (2018), 433\u2013442.","journal-title":"International Journal of Pure and Applied Mathematics"},{"key":"e_1_3_2_48_1","doi-asserted-by":"publisher","DOI":"10.4324\/9781315722580"},{"key":"e_1_3_2_49_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ahj.2015.03.009"},{"key":"e_1_3_2_50_1","doi-asserted-by":"crossref","unstructured":"B. Vidgen and L. Derczynski. 2020. Directions in abusive language training data: Garbage in garbage out. arXiv:2004.01670.","DOI":"10.1371\/journal.pone.0243300"},{"key":"e_1_3_2_51_1","article-title":"Offensive language detection: A comparative analysis","author":"Vyshnav M. T.","year":"2020","unstructured":"M. T. Vyshnav, S. Kumar, and K. P. Soman. 2020. Offensive language detection: A comparative analysis. arXiv preprint arXiv:2001.03131 (2020).","journal-title":"arXiv preprint"},{"key":"e_1_3_2_52_1","volume-title":"Proceedings of GermEval 2018, 14th Conference on Natural Language Processing","author":"Wiegand M.","year":"2018","unstructured":"M. Wiegand, M. Siegel, and J. Ruppenhofer. 2018. Overview of the GermEval 2018 shared task on the identification of offensive language. In Proceedings of GermEval 2018, 14th Conference on Natural Language Processing. 1--10."},{"key":"e_1_3_2_53_1","doi-asserted-by":"publisher","DOI":"10.5555\/2382029.2382139"},{"key":"e_1_3_2_54_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2021.02.056"},{"key":"e_1_3_2_55_1","doi-asserted-by":"publisher","DOI":"10.3390\/sym13010110"},{"key":"e_1_3_2_56_1","article-title":"SemEval-2019 Task 6: Identifying and categorizing offensive language in social media (OffensEval)","author":"Zampieri M.","year":"2019","unstructured":"M. Zampieri, S. Malmasi, P. Nakov, S. Rosenthal, N. Farra, and R. Kumar. 2019. SemEval-2019 Task 6: Identifying and categorizing offensive language in social media (OffensEval). arXiv preprint arXiv:1903.08983 (2019).","journal-title":"arXiv preprint"},{"key":"e_1_3_2_57_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1155\/2022\/9879986","article-title":"Sentiment analysis of international and foreign Chinese-language texts with multilevel features","volume":"2022","author":"Zhu M.","year":"2022","unstructured":"M. Zhu. 2022. Sentiment analysis of international and foreign Chinese-language texts with multilevel features. Discrete Dynamics in Nature and Society 2022 (2022), 1\u201312.","journal-title":"Discrete Dynamics in Nature and Society"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3575860","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3575860","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:12Z","timestamp":1750178772000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3575860"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,24]]},"references-count":56,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,4,30]]}},"alternative-id":["10.1145\/3575860"],"URL":"https:\/\/doi.org\/10.1145\/3575860","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,24]]},"assertion":[{"value":"2022-02-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-11-27","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-03-24","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}