{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,19]],"date-time":"2026-04-19T01:22:52Z","timestamp":1776561772314,"version":"3.51.2"},"reference-count":57,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2021,5,12]],"date-time":"2021-05-12T00:00:00Z","timestamp":1620777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Multilingual characteristics, lack of annotated data, and imbalanced sample distribution are the three main challenges for toxic comment analysis in a multilingual setting. This paper proposes a multilingual toxic text classifier which adopts a novel fusion strategy that combines different loss functions and multiple pre-training models. Specifically, the proposed learning pipeline starts with a series of pre-processing steps, including translation, word segmentation, purification, text digitization, and vectorization, to convert word tokens to a vectorized form suitable for the downstream tasks. Two models, multilingual bidirectional encoder representation from transformers (MBERT) and XLM-RoBERTa (XLM-R), are employed for pre-training through Masking Language Modeling (MLM) and Translation Language Modeling (TLM), which incorporate semantic and contextual information into the models. We train six base models and fuse them to obtain three fusion models using the F1 scores as the weights. The models are evaluated on the Jigsaw Multilingual Toxic Comment dataset. Experimental results show that the best fusion model outperforms the two state-of-the-art models, MBERT and XLM-R, in F1 score by 5.05% and 0.76%, respectively, verifying the effectiveness and robustness of the proposed fusion strategy.<\/jats:p>","DOI":"10.3390\/info12050205","type":"journal-article","created":{"date-parts":[[2021,5,12]],"date-time":"2021-05-12T10:59:12Z","timestamp":1620817152000},"page":"205","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["A Study of Multilingual Toxic Text Detection Approaches under Imbalanced Sample Distribution"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2588-7010","authenticated-orcid":false,"given":"Guizhe","family":"Song","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Dalian University of Technology, Dalian 116024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Degen","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Dalian University of Technology, Dalian 116024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3327-8108","authenticated-orcid":false,"given":"Zhifeng","family":"Xiao","sequence":"additional","affiliation":[{"name":"School of Engineering, Penn State Erie, The Behrend College, Erie, PA 16563, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,5,12]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"van Aken, B., Risch, J., Krestel, R., and L\u00f6ser, A. (2018). Challenges for toxic comment classification: An in-depth error analysis. arXiv.","DOI":"10.18653\/v1\/W18-5105"},{"key":"ref_2","unstructured":"Bashar, M.A., and Nayak, R. (2020). QutNocturnal@ HASOC\u201919: CNN for hate speech and offensive content identification in Hindi language. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Moon, J., Cho, W.I., and Lee, J. (2020). BEEP! Korean Corpus of Online News Comments for Toxic Speech Detection. arXiv.","DOI":"10.18653\/v1\/2020.socialnlp-1.4"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Zueva, N., Kabirova, M., and Kalaidin, P. (2020). Reducing Unintended Identity Bias in Russian Hate Speech Detection. arXiv.","DOI":"10.18653\/v1\/2020.alw-1.8"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"114120","DOI":"10.1016\/j.eswa.2020.114120","article-title":"Comparing pre-trained language models for Spanish hate speech detection","volume":"166","year":"2021","journal-title":"Expert Syst. Appl."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Waseem, Z., and Hovy, D. (2016, January 7\u201312). Hateful symbols or hateful people? Predictive features for hate speech detection on twitter. Proceedings of the NAACL Student Research Workshop, Berlin, Germany.","DOI":"10.18653\/v1\/N16-2013"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Davidson, T., Warmsley, D., Macy, M., and Weber, I. (2017, January 15\u201318). Automated hate speech detection and the problem of offensive language. Proceedings of the International AAAI Conference on Web and Social Media, Montr\u00e9al, QC, Canada.","DOI":"10.1609\/icwsm.v11i1.14955"},{"key":"ref_8","unstructured":"Sharma, S., Agrawal, S., and Shrivastava, M. (2018). Degree based classification of harmful speech using twitter data. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Salminen, J., Almerekhi, H., Kamel, A.M., Jung, S.G., and Jansen, B.J. (2019, January 10\u201314). Online hate ratings vary by extremes: A statistical analysis. Proceedings of the 2019 Conference on Human Information Interaction and Retrieval, Glasgow, UK.","DOI":"10.1145\/3295750.3298954"},{"key":"ref_10","unstructured":"Kajla, H., Hooda, J., and Saini, G. (2020, January 13\u201315). Classification of Online Toxic Comments Using Machine Learning Algorithms. Proceedings of the 2020 4th International Conference on Intelligent Computing and Control Systems (ICICCS), Madurai, India."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Greevy, E., and Smeaton, A.F. (2004, January 25\u201329). Classifying racist texts using a support vector machine. Proceedings of the 27th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Sheffield, UK.","DOI":"10.1145\/1008992.1009074"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Alfina, I., Mulia, R., Fanany, M.I., and Ekanata, Y. (2017, January 28\u201329). Hate speech detection in the Indonesian language: A dataset and preliminary study. Proceedings of the 2017 International Conference on Advanced Computer Science and Information Systems (ICACSIS), Jakarta, Indonesia.","DOI":"10.1109\/ICACSIS.2017.8355039"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Kwok, I., and Wang, Y. (2013, January 14\u201318). Locate the hate: Detecting tweets against blacks. Proceedings of the AAAI Conference on Artificial Intelligence, Bellevue, WA, USA.","DOI":"10.1609\/aaai.v27i1.8539"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"060011","DOI":"10.1063\/1.5082126","article-title":"Classification of online toxic comments using the logistic regression and neural networks models","volume":"Volume 2048","author":"Saif","year":"2018","journal-title":"AIP Conference Proceedings"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Georgakopoulos, S.V., Tasoulis, S.K., Vrahatis, A.G., and Plagianakos, V.P. (2018, January 9\u201312). Convolutional neural networks for toxic comment classification. Proceedings of the 10th Hellenic Conference on Artificial Intelligence, Patras, Greece.","DOI":"10.1145\/3200947.3208069"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Jubaer, A., Sayem, A., and Rahman, M.A. (2019, January 22\u201323). Bangla toxic comment classification (machine learning and deep learning approach). Proceedings of the 2019 8th International Conference System Modeling and Advancement in Research Trends (SMART), Moradabad, India.","DOI":"10.1109\/SMART46866.2019.9117286"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Dubey, K., Nair, R., Khan, M.U., and Shaikh, S. (2020, January 11\u201312). Toxic Comment Detection using LSTM. Proceedings of the 2020 Third International Conference on Advances in Electronics, Computers and Communications (ICAECC), Bengaluru, India.","DOI":"10.1109\/ICAECC50550.2020.9339521"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Mahajan, A., Shah, D., and Jafar, G. (EasyChair Preprint, 2020). Explainable AI Approach towards Toxic Comment Classification, EasyChair Preprint.","DOI":"10.1007\/978-981-33-4367-2_81"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"106443","DOI":"10.1016\/j.knosys.2020.106443","article-title":"A machine learning-based investigation utilizing the in-text features for the identification of dominant emotion in an email","volume":"208","author":"Halim","year":"2020","journal-title":"Knowl. Based Syst."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"324","DOI":"10.1016\/j.ijar.2019.07.010","article-title":"Three-way decisions based feature fusion for Chinese irony detection","volume":"113","author":"Jia","year":"2019","journal-title":"Int. J. Approx. Reason."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Tzogka, C., Passalis, N., Iosifidis, A., Gabbouj, M., and Tefas, A. (2019, January 13\u201316). Less Is More: Deep Learning Using Subjective Annotations for Sentiment Analysis from Social Media. Proceedings of the 2019 IEEE 29th International Workshop on Machine Learning for Signal Processing (MLSP), Pittsburgh, PA, USA.","DOI":"10.1109\/MLSP.2019.8918792"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Ranasinghe, T., and Zampieri, M. (2021). MUDES: Multilingual Detection of Offensive Spans. arXiv.","DOI":"10.18653\/v1\/2021.naacl-demos.17"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Ranasinghe, T., and Hettiarachchi, H. (2020). BRUMS at SemEval-2020 task 12: Transformer based multilingual offensive language identification in social media. arXiv.","DOI":"10.18653\/v1\/2020.semeval-1.251"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"684","DOI":"10.1016\/j.ipm.2016.12.008","article-title":"Multilingual emotion classification using supervised learning: Comparative experiments","volume":"53","author":"Becker","year":"2017","journal-title":"Inf. Process. Manag."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Ousidhoum, N., Lin, Z., Zhang, H., Song, Y., and Yeung, D.Y. (2019). Multilingual and multi-aspect hate speech analysis. arXiv.","DOI":"10.18653\/v1\/D19-1474"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3377323","article-title":"A multilingual evaluation for online hate speech detection","volume":"20","author":"Corazza","year":"2020","journal-title":"ACM Trans. Internet Technol."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"102360","DOI":"10.1016\/j.ipm.2020.102360","article-title":"Misogyny detection in twitter: A multilingual and cross-domain study","volume":"57","author":"Pamungkas","year":"2020","journal-title":"Inf. Process. Manag."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1007\/s10590-017-9202-6","article-title":"Cross-lingual sentiment transfer with limited resources","volume":"32","author":"Rasooli","year":"2018","journal-title":"Mach. Transl."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Dong, X., and De Melo, G. (2018, January 2\u20137). Cross-lingual propagation for deep sentiment analysis. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12071"},{"key":"ref_30","unstructured":"Can, E.F., Ezen-Can, A., and Can, F. (2018). Multilingual sentiment analysis: An RNN-based framework for limited data. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Li, X., Li, Z., Sheng, J., and Slamu, W. (2020). Low-Resource Text Classification via Cross-Lingual Language Model Fine-Tuning. China National Conference on Chinese Computational Linguistics, Springer.","DOI":"10.1007\/978-3-030-63031-7_17"},{"key":"ref_32","unstructured":"Roy, S.G., Narayan, U., Raha, T., Abid, Z., and Varma, V. (2021). Leveraging Multilingual Transformers for Hate Speech Detection. arXiv."},{"key":"ref_33","unstructured":"Mohammad, F. (2018). Is preprocessing of text really worth your time for online comment classification?. arXiv."},{"key":"ref_34","unstructured":"Kalouli, A.L., Kaiser, K., Hautli, A., Kaiser, G.A., and Butt, M. (2018, January 7\u201312). A multilingual approach to question classification. Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wang, Z., Lee, S., Li, S., and Zhou, G. (2015, January 26\u201331). Emotion detection in code-switching texts via bilingual and sentimental information. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), Beijing, China.","DOI":"10.3115\/v1\/P15-2125"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Ibrahim, M., Torki, M., and El-Makky, N. (2018, January 17\u201320). Imbalanced toxic comments classification using data augmentation and deep learning. Proceedings of the 2018 17th IEEE International Conference on Machine Learning and Applications (ICMLA), Orlando, FL, USA.","DOI":"10.1109\/ICMLA.2018.00141"},{"key":"ref_37","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017). Attention is all you need. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzm\u00e1n, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V. (2019). Unsupervised cross-lingual representation learning at scale. arXiv.","DOI":"10.18653\/v1\/2020.acl-main.747"},{"key":"ref_39","unstructured":"Huang, X., Xing, L., Dernoncourt, F., and Paul, M.J. (2020). Multilingual Twitter corpus and baselines for evaluating demographic bias in hate speech recognition. arXiv."},{"key":"ref_40","unstructured":"Aluru, S.S., Mathew, B., Saha, P., and Mukherjee, A. (2020). Deep learning models for multilingual hate speech detection. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Mikolov, T., Karafi\u00e1t, M., Burget, L., \u010cernock\u1ef3, J., and Khudanpur, S. (2010, January 26\u201330). Recurrent neural network based language model. Proceedings of the Eleventh Annual Conference of the International Speech Communication Association, Makuhari, Japan.","DOI":"10.21437\/Interspeech.2010-343"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Ghosh, S., Kumar, S., Lepcha, S., and Jain, S.S. (2021). Toxic Text Classification. Data Science and Security, Springer.","DOI":"10.1007\/978-981-15-5309-7_27"},{"key":"ref_43","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Mozafari, M., Farahbakhsh, R., and Crespi, N. (2019). A BERT-based transfer learning approach for hate speech detection in online social media. International Conference on Complex Networks and Their Applications, Springer.","DOI":"10.1007\/978-3-030-36687-2_77"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Joulin, A., Grave, E., Bojanowski, P., and Mikolov, T. (2016). Bag of tricks for efficient text classification. arXiv.","DOI":"10.18653\/v1\/E17-2068"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Kim, Y., Jernite, Y., Sontag, D., and Rush, A. (2016, January 12\u201317). Character-aware neural language models. Proceedings of the AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA.","DOI":"10.1609\/aaai.v30i1.10362"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Kalchbrenner, N., Grefenstette, E., and Blunsom, P. (2014). A convolutional neural network for modelling sentences. arXiv.","DOI":"10.3115\/v1\/P14-1062"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"102544","DOI":"10.1016\/j.ipm.2021.102544","article-title":"A joint learning approach with knowledge injection for zero-shot cross-lingual hate speech detection","volume":"58","author":"Pamungkas","year":"2021","journal-title":"Inf. Process. Manag."},{"key":"ref_49","unstructured":"Conneau, A., Lample, G., Ranzato, M., Denoyer, L., and J\u00e9gou, H. (2017). Word translation without parallel data. arXiv."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Bassignana, E., Basile, V., and Patti, V. (2018, January 10\u201312). Hurtlex: A multilingual lexicon of words to hurt. Proceedings of the 5th Italian Conference on Computational Linguistics, CLiC-it 2018. CEUR-WS, Torino, Italy.","DOI":"10.4000\/books.aaccademia.3085"},{"key":"ref_51","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach. arXiv."},{"key":"ref_52","unstructured":"Lample, G., and Conneau, A. (2019). Cross-lingual language model pretraining. arXiv."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"223","DOI":"10.1002\/poi3.85","article-title":"Cyber hate speech on twitter: An application of machine classification and statistical modeling for policy and decision making","volume":"7","author":"Burnap","year":"2015","journal-title":"Policy Internet"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Gao, L., and Huang, R. (2017). Detecting online hate speech using context aware models. arXiv.","DOI":"10.26615\/978-954-452-049-6_036"},{"key":"ref_55","unstructured":"Zimmerman, S., Kruschwitz, U., and Fox, C. (2018, January 7\u201312). Improving hate speech detection with deep learning ensembles. Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan."},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Zhang, L., Wu, L., Li, S., Wang, Z., and Zhou, G. (2018). Cross-lingual emotion classification with auxiliary and attention neural networks. CCF International Conference on Natural Language Processing and Chinese Computing, Springer.","DOI":"10.1007\/978-3-319-99495-6_36"},{"key":"ref_57","unstructured":"Yin, W., Kann, K., Yu, M., and Sch\u00fctze, H. (2017). Comparative study of CNN and RNN for natural language processing. arXiv."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/12\/5\/205\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:59:36Z","timestamp":1760162376000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/12\/5\/205"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,12]]},"references-count":57,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2021,5]]}},"alternative-id":["info12050205"],"URL":"https:\/\/doi.org\/10.3390\/info12050205","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,5,12]]}}}