{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T23:32:00Z","timestamp":1776209520295,"version":"3.50.1"},"reference-count":40,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,12,1]],"date-time":"2021-12-01T00:00:00Z","timestamp":1638316800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,12,22]],"date-time":"2021-12-22T00:00:00Z","timestamp":1640131200000},"content-version":"vor","delay-in-days":21,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Big Data"],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Social media have become a very viable medium for communication, collaboration, exchange of information, knowledge, and ideas. However, due to anonymity preservation, the incidents of hate speech and cyberbullying have been diversified across the globe. This intimidating problem has recently sought the attention of researchers and scholars worldwide and studies have been undertaken to formulate solution strategies for automatic detection of cyberaggression and hate speech, varying from machine learning models with vast features to more complex deep neural network models and different SN platforms. However, the existing research is directed towards mature languages and highlights a huge gap in newly embraced resource poor languages. One such language that has been recently adopted worldwide and more specifically by south Asian countries for communication on social media is Roman Urdu i-e Urdu language written using Roman scripting. To address this research gap, we have performed extensive preprocessing on Roman Urdu microtext. This typically involves formation of Roman Urdu slang- phrase dictionary and mapping slangs after tokenization. We have also eliminated cyberbullying domain specific stop words for dimensionality reduction of corpus. The unstructured data were further processed to handle encoded text formats and metadata\/non-linguistic features. Furthermore, we performed extensive experiments by implementing RNN-LSTM, RNN-BiLSTM and CNN models varying epochs executions, model layers and tuning hyperparameters to analyze and uncover cyberbullying textual patterns in Roman Urdu. The efficiency and performance of models were evaluated using different metrics to present the comparative analysis. Results highlight that RNN-LSTM and RNN-BiLSTM performed best and achieved validation accuracy of 85.5 and 85% whereas F1 score was 0.7 and 0.67 respectively over aggression class.<\/jats:p>","DOI":"10.1186\/s40537-021-00550-7","type":"journal-article","created":{"date-parts":[[2021,12,22]],"date-time":"2021-12-22T14:02:43Z","timestamp":1640181763000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":78,"title":["Cyberbullying detection: advanced preprocessing techniques &amp; deep learning architecture for Roman Urdu data"],"prefix":"10.1186","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3816-3644","authenticated-orcid":false,"given":"Amirita","family":"Dewani","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2638-4252","authenticated-orcid":false,"given":"Mohsin Ali","family":"Memon","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0887-8083","authenticated-orcid":false,"given":"Sania","family":"Bhatti","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,12,22]]},"reference":[{"issue":"1","key":"550_CR1","doi-asserted-by":"publisher","first-page":"45","DOI":"10.3390\/ijerph17010045","volume":"17","author":"K Hellfeldt","year":"2020","unstructured":"Hellfeldt K, L\u00f3pez-Romero L, Andershed H. Cyberbullying and psychological well-being in young adolescence: the potential protective mediation effects of social support from family, friends, and teachers. Int J Environ Res Public Health. 2020;17(1):45.","journal-title":"Int J Environ Res Public Health"},{"key":"550_CR2","unstructured":"Dadvar M. Experts and machines united against cyberbullying [PhD thesis]. University of Twente. 2014."},{"issue":"1","key":"550_CR3","first-page":"103","volume":"11","author":"H Magsi","year":"2017","unstructured":"Magsi H, Agha N, Magsi I. Understanding cyber bullying in Pakistani context: causes and effects on young female university students in Sindh province. New Horiz. 2017;11(1):103.","journal-title":"New Horiz"},{"issue":"2","key":"550_CR4","first-page":"503","volume":"6","author":"SF Qureshi","year":"2020","unstructured":"Qureshi SF, Abbasi M, Shahzad M. Cyber harassment and women of Pakistan: analysis of female victimization. J Bus Soc Rev Emerg Econ. 2020;6(2):503\u201310.","journal-title":"J Bus Soc Rev Emerg Econ"},{"key":"550_CR5","unstructured":"S. Irfan Ahmed, Cyber bullying doubles during pandemic. https:\/\/www.thenews.com.pk\/tns\/detail\/671918-cyber-bullying-doubles-during-pandemic. Accessed 24 Aug 2020."},{"key":"550_CR6","doi-asserted-by":"publisher","first-page":"189823","DOI":"10.1109\/ACCESS.2020.3031393","volume":"8","author":"M Shahroz","year":"2020","unstructured":"Shahroz M, Mushtaq MF, Mehmood A, Ullah S, Choi GS. RUTUT: roman Urdu to Urdu translator based on character substitution rules and unicode mapping. IEEE Access. 2020;8:189823\u201341.","journal-title":"IEEE Access"},{"key":"550_CR7","doi-asserted-by":"publisher","first-page":"192740","DOI":"10.1109\/ACCESS.2020.3030885","volume":"8","author":"F Mehmood","year":"2020","unstructured":"Mehmood F, Ghani MU, Ibrahim MA, Shahzadi R, Mahmood W, Asim MN. A precisely xtreme-multi channel hybrid approach for roman urdu sentiment analysis. IEEE Access. 2020;8:192740\u201359.","journal-title":"IEEE Access"},{"issue":"21","key":"550_CR8","doi-asserted-by":"publisher","first-page":"2664","DOI":"10.3390\/electronics10212664","volume":"10","author":"M Alotaibi","year":"2021","unstructured":"Alotaibi M, Alotaibi B, Razaque A. A multichannel deep learning framework for cyberbullying detection on social media. Electronics. 2021;10(21):2664.","journal-title":"Electronics"},{"key":"550_CR9","doi-asserted-by":"crossref","unstructured":"Dinakar K, Reichart R, Lieberman H. Modeling the detection of textual cyberbullying. In: 5th international AAAI conference on weblogs and social media. 2011.","DOI":"10.1609\/icwsm.v5i3.14209"},{"key":"550_CR10","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-020-00701-5","author":"C Iwendi","year":"2020","unstructured":"Iwendi C, Srivastava G, Khan S, Maddikunta PKR. Cyberbullying detection solutions based on deep learning architectures. Multimed Syst. 2020. https:\/\/doi.org\/10.1007\/s00530-020-00701-5.","journal-title":"Multimed Syst"},{"issue":"1","key":"550_CR11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s13673-019-0205-6","volume":"10","author":"J Salminen","year":"2020","unstructured":"Salminen J, Hopf M, Chowdhury SA, Jung S, Almerekhi H, Jansen BJ. Developing an online hate classifier for multiple social media platforms. Hum Cent Comput Inf Sci. 2020;10(1):1\u201334.","journal-title":"Hum Cent Comput Inf Sci"},{"issue":"7","key":"550_CR12","doi-asserted-by":"publisher","first-page":"779","DOI":"10.3390\/electronics10070779","volume":"10","author":"D Dess\u00ec","year":"2021","unstructured":"Dess\u00ec D, Recupero DR, Sack H. An assessment of deep learning models and word embeddings for toxicity detection within online textual comments. Electronics. 2021;10(7):779.","journal-title":"Electronics"},{"key":"550_CR13","doi-asserted-by":"crossref","unstructured":"S. A. \u00d6zel, E. Sara\u00e7, S. Akdemir, and H. Aksu, Detection of cyberbullying on social media messages in Turkish, In: 2017 International Conference on Computer Science and Engineering (UBMK), 2017, pp. 366\u2013370.","DOI":"10.1109\/UBMK.2017.8093411"},{"key":"550_CR14","unstructured":"E. C. Ates, E. Bostanci, and M. S. Guzel, Comparative Performance of Machine Learning Algorithms in Cyberbullying Detection: Using Turkish Language Preprocessing Techniques, arXiv Prepr. arXiv2101.12718, 2021."},{"issue":"10","key":"550_CR15","doi-asserted-by":"publisher","first-page":"e0203794","DOI":"10.1371\/journal.pone.0203794","volume":"13","author":"C Van Hee","year":"2018","unstructured":"Van Hee C, et al. Automatic detection of cyberbullying in social media text. PLoS ONE. 2018;13(10):e0203794.","journal-title":"PLoS ONE"},{"issue":"6","key":"550_CR16","doi-asserted-by":"publisher","first-page":"275","DOI":"10.25046\/aj020634","volume":"2","author":"B Haidar","year":"2017","unstructured":"Haidar B, Chamoun M, Serhrouchni A. A multilingual system for cyberbullying detection: Arabic content detection using machine learning. Adv Sci Technol Eng Syst J. 2017;2(6):275\u201384.","journal-title":"Adv Sci Technol Eng Syst J"},{"key":"550_CR17","unstructured":"G\u00f3mez-Adorno H, Bel-Enguix G, Sierra G, S\u00e1nchez O, Quezada D. A machine learning approach for detecting aggressive tweets in Spanish, In: IberEval@ SEPLN. 2018. pp. 102\u2013107."},{"key":"550_CR18","unstructured":"X. Bai, F. Merenda, C. Zaghi, T. Caselli, and M. Nissim, RuG at GermEval: Detecting Offensive Speech in German Social Media, in 14th Conference on Natural Language Processing KONVENS 2018, 2018, p. 63."},{"key":"550_CR19","unstructured":"B. Birkeneder, J. Mitrovic, J. Niemeier, L. Teubert, and S. Handschuh, upInf\u2014Offensive Language Detection in German Tweets, In: Proceedings of the GermEval 2018 Workshop, 2018, pp. 71\u201378."},{"key":"550_CR20","unstructured":"J. M. Schneider, R. Roller, P. Bourgonje, S. Hegele, and G. Rehm, Towards the Automatic Classification of Offensive Language and Related Phenomena in German Tweets, In: 14th Conference on Natural Language Processing KONVENS 2018, 2018, p. 95."},{"key":"550_CR21","unstructured":"H. Margono, X. Yi, and G. K. Raikundalia, Mining Indonesian cyber bullying patterns in social networks, In: Proceedings of the Thirty-Seventh Australasian Computer Science Conference-Volume 147, 2014, pp. 115\u2013124."},{"key":"550_CR22","doi-asserted-by":"crossref","unstructured":"H. Nurrahmi and D. Nurjanah, Indonesian Twitter Cyberbullying Detection using Text Classification and User Credibility, In: 2018 International Conference on Information and Communications Technology (ICOIACT), 2018, pp. 543\u2013548.","DOI":"10.1109\/ICOIACT.2018.8350758"},{"key":"550_CR23","doi-asserted-by":"crossref","unstructured":"A. Bohra, D. Vijay, V. Singh, S. S. Akhtar, and M. Shrivastava, A Dataset of Hindi-English Code-Mixed Social Media Text for Hate Speech Detection, In: Proceedings of the Second Workshop on Computational Modeling of People\u2019s Opinions, Personality, and Emotions in Social Media, 2018, pp. 36\u201341.","DOI":"10.18653\/v1\/W18-1105"},{"key":"550_CR24","unstructured":"A. Roy, P. Kapil, K. Basak, and A. Ekbal, An ensemble approach for aggression identification in english and hindi text, In: Proceedings of the First Workshop on Trolling, Aggression and Cyberbullying (TRAC-2018), 2018, pp. 66\u201373."},{"key":"550_CR25","unstructured":"Association for Computational Linguistics. https:\/\/www.aclweb.org\/portal\/content\/deadline-extension-first-task-automatic-cyberbullying-detection-polish-language. Accessed 09 May 2019."},{"key":"550_CR26","doi-asserted-by":"publisher","DOI":"10.1109\/ICECE.2018.8636797","author":"R Ghosh","year":"2021","unstructured":"Ghosh R, Nowal S, Manju G. Social media cyberbullying detection using machine learning in bengali language. Int J Eng Res Technol. 2021. https:\/\/doi.org\/10.1109\/ICECE.2018.8636797.","journal-title":"Int J Eng Res Technol"},{"issue":"16","key":"550_CR27","doi-asserted-by":"publisher","first-page":"834","DOI":"10.31838\/jcr.07.16.109","volume":"7","author":"KR Talpur","year":"2020","unstructured":"Talpur KR, Yuhaniz SS, Sjarif NNBA, Ali B. Cyberbullying detection in Roman Urdu language using lexicon based approach. J Crit Rev. 2020;7(16):834\u201348. https:\/\/doi.org\/10.31838\/jcr.07.16.109.","journal-title":"J Crit Rev"},{"key":"550_CR28","unstructured":"J. Brownlee, Imbalanced Classification, December 23, 2019. https:\/\/machinelearningmastery.com. Accessed 10 May 2021."},{"issue":"1","key":"550_CR29","doi-asserted-by":"publisher","first-page":"1","DOI":"10.26735\/GBTV9013","volume":"4","author":"M Arif","year":"2021","unstructured":"Arif M. A systematic review of machine learning algorithms in cyberbullying detection: future directions and challenges. J Inf Secur Cybercrimes Res. 2021;4(1):1\u201326.","journal-title":"J Inf Secur Cybercrimes Res"},{"key":"550_CR30","doi-asserted-by":"crossref","unstructured":"A. Dewani, M. Ali Memon, and S. Bhatti, Development of Computational Linguistic Resources for Automated Detection of Textual Cyberbullying Threats in Roman Urdu Language, 3C TIC. Cuad. Desarro. Apl. a las TIC, 101\u2013121., p. 17, 2021.","DOI":"10.17993\/3ctic.2021.102.101-121"},{"key":"550_CR31","doi-asserted-by":"publisher","first-page":"110212","DOI":"10.1016\/j.chaos.2020.110212","volume":"140","author":"F Shahid","year":"2020","unstructured":"Shahid F, Zameer A, Muneeb M. Predictions for COVID-19 with deep learning models of LSTM, GRU and Bi-LSTM. Chaos Solitons Fractals.\u00a02020;140:110212.","journal-title":"GRU and Bi-LSTM. Chaos Solitons Fractals."},{"key":"550_CR32","doi-asserted-by":"crossref","unstructured":"M. Cliche, \u201cBB_twtr at SemEval-2017 task 4: Twitter sentiment analysis with CNNs and LSTMs,\u201d arXiv Prepr. arXiv1704.06125, 2017.","DOI":"10.18653\/v1\/S17-2094"},{"key":"550_CR33","unstructured":"W. Zaremba, I. Sutskever, and O. Vinyals, Recurrent neural network regularization, arXiv Prepr. arXiv1409.2329, 2014."},{"key":"550_CR34","doi-asserted-by":"publisher","first-page":"51522","DOI":"10.1109\/ACCESS.2019.2909919","volume":"7","author":"G Xu","year":"2019","unstructured":"Xu G, Meng Y, Qiu X, Yu Z, Wu X. Sentiment analysis of comment texts based on BiLSTM. Ieee Access. 2019;7:51522\u201332.","journal-title":"Ieee Access"},{"key":"550_CR35","unstructured":"S. Minaee, E. Azimi, and A. Abdolrashidi, Deep-sentiment: Sentiment analysis using ensemble of cnn and bi-lstm models, arXiv Prepr. arXiv1904.04206, 2019."},{"issue":"1","key":"550_CR36","doi-asserted-by":"publisher","first-page":"104","DOI":"10.1016\/j.ipm.2013.08.006","volume":"50","author":"AK Uysal","year":"2014","unstructured":"Uysal AK, Gunal S. The impact of preprocessing on text classification. Inf Process Manag. 2014;50(1):104\u201312.","journal-title":"Inf Process Manag"},{"issue":"1","key":"550_CR37","first-page":"7","volume":"5","author":"S Vijayarani","year":"2015","unstructured":"Vijayarani S, Ilamathi MJ, Nithya M. Preprocessing techniques for text mining-an overview. Int J Comput Sci Commun Networks. 2015;5(1):7\u201316.","journal-title":"Int J Comput Sci Commun Networks"},{"key":"550_CR38","unstructured":"Pycontractions 2.0.1. https:\/\/pypi.org\/project\/pycontractions\/. Accessed 17 Nov 2021."},{"key":"550_CR39","unstructured":"API Documentation. https:\/\/www.tensorflow.org\/api_docs. Accessed 21 Oct 2020."},{"key":"550_CR40","unstructured":"NumPy. https:\/\/numpy.org\/. Accessed 18 Nov 2021."}],"container-title":["Journal of Big Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-021-00550-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s40537-021-00550-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s40537-021-00550-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,20]],"date-time":"2023-01-20T00:38:05Z","timestamp":1674175085000},"score":1,"resource":{"primary":{"URL":"https:\/\/journalofbigdata.springeropen.com\/articles\/10.1186\/s40537-021-00550-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12]]},"references-count":40,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["550"],"URL":"https:\/\/doi.org\/10.1186\/s40537-021-00550-7","relation":{},"ISSN":["2196-1115"],"issn-type":[{"value":"2196-1115","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12]]},"assertion":[{"value":"3 September 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 December 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 December 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing or conflicts of interest to report regarding the present study.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"160"}}