{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,21]],"date-time":"2026-01-21T07:41:41Z","timestamp":1768981301134,"version":"3.49.0"},"reference-count":45,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2021,1,21]],"date-time":"2021-01-21T00:00:00Z","timestamp":1611187200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100010693","name":"Telekom Malaysia Berhad","doi-asserted-by":"publisher","award":["MMUE\/180029"],"award-info":[{"award-number":["MMUE\/180029"]}],"id":[{"id":"10.13039\/501100010693","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Given the excessive foul language identified in audio and video files and the detrimental consequences to an individual\u2019s character and behaviour, content censorship is crucial to filter profanities from young viewers with higher exposure to uncensored content. Although manual detection and censorship were implemented, the methods proved tedious. Inevitably, misidentifications involving foul language owing to human weariness and the low performance in human visual systems concerning long screening time occurred. As such, this paper proposed an intelligent system for foul language censorship through a mechanized and strong detection method using advanced deep Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) through Long Short-Term Memory (LSTM) cells. Data on foul language were collected, annotated, augmented, and analysed for the development and evaluation of both CNN and RNN configurations. Hence, the results indicated the feasibility of the suggested systems by reporting a high volume of curse word identifications with only 2.53% to 5.92% of False Negative Rate (FNR). The proposed system outperformed state-of-the-art pre-trained neural networks on the novel foul language dataset and proved to reduce the computational cost with minimal trainable parameters.<\/jats:p>","DOI":"10.3390\/s21030710","type":"journal-article","created":{"date-parts":[[2021,1,21]],"date-time":"2021-01-21T09:49:21Z","timestamp":1611222561000},"page":"710","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Design and Implementation of Fast Spoken Foul Language Recognition with Different End-to-End Deep Neural Network Architectures"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1225-1723","authenticated-orcid":false,"given":"Abdulaziz Saleh","family":"Ba Wazir","sequence":"first","affiliation":[{"name":"Faculty of Engineering, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7613-4596","authenticated-orcid":false,"given":"Hezerul Abdul","family":"Karim","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mohd Haris Lye","family":"Abdullah","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5522-0033","authenticated-orcid":false,"given":"Nouar","family":"AlDahoul","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sarina","family":"Mansor","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mohammad Faizal Ahmad","family":"Fauzi","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3005-4109","authenticated-orcid":false,"given":"John","family":"See","sequence":"additional","affiliation":[{"name":"Faculty of Computing and Informatics, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ahmad Syazwan","family":"Naim","sequence":"additional","affiliation":[{"name":"IPTV Development, Unifi Content, Telekom Malaysia Berhad, Cyberjaya 63100, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,1,21]]},"reference":[{"key":"ref_1","unstructured":"Tuttle, K. (2021, January 11). The Profanity Problem: Swearing in Movies. In Movie Babble 2018. Available online: https:\/\/moviebabble.com\/2018\/02\/20\/the-profanity-problem-swearing-in-movies\/."},{"key":"ref_2","unstructured":"Day, S. (2018). Cursing negatively affects society. The Baker Orange, Baker University Media."},{"key":"ref_3","first-page":"16","article-title":"A review on speech recognition technique","volume":"10","author":"Gaikwad","year":"2010","journal-title":"Int. J. Comput. Appl."},{"key":"ref_4","first-page":"173","article-title":"Deep Speech 2: End-to-end speech recognition in English and Mandarin","volume":"Volume 48","author":"Amodei","year":"2016","journal-title":"Proceedings of the 33rd International Conference on Machine Learning"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Bahdanau, D., Chorowski, J., Serdyuk, D., Brakl, P., and Bengio, Y. (2016, January 20\u201325). End-to-end attention-based large vocabulary speech recognition. Proceedings of the 41st IEEE International Conference on Acoustics, Speech, and Signal Processing, Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472618"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Chiu, C.-C., Sainath, T.N., Wu, Y., Prabhavalkar, R., Nguyen, P., Chen, Z., Kannan, A., Weiss, R.J., Rao, K., and Gonina, E. (2018, January 15\u201320). State-of-the-art speech recognition with sequence-to-sequence models. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8462105"},{"key":"ref_7","unstructured":"Warden, P. (2018). Speech commands: A dataset for limited-vocabulary speech recognition. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"10767","DOI":"10.1109\/ACCESS.2019.2891838","article-title":"Effective combination of DenseNet and BiLSTM for keyword spotting","volume":"7","author":"Zeng","year":"2019","journal-title":"IEEE Access"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Pervaiz, A., Hussain, F., Israr, H., Tahir, M.A., Raja, F.R., Baloch, N.K., Ishmanov, F., and Bin Zikria, Y. (2020). Incorporating noise robustness in speech command recognition by noise augmentation of training data. Sensors, 20.","DOI":"10.3390\/s20082326"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Anvarjon, T., and Kwon, S. (2020). Deep-Net: A lightweight CNN-based speech emotion recognition system using deep frequency features. Sensors, 20.","DOI":"10.3390\/s20185212"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Wazir, A.S.B., Karim, H.A., Abdullah, M.H.L., Mansor, S., AlDahoul, N., Fauzi, M.F.A., and See, J. (2020, January 21\u201324). Spectrogram-based classification of spoken foul language using deep CNN. Proceedings of the 2020 IEEE 22nd International Workshop on Multimedia Signal Processing (MMSP), Tampere, Finland.","DOI":"10.1109\/MMSP48831.2020.9287133"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Pieropan, A., Salvi, G., Pauwels, K., Kjellstr\u00f6m, H., and Salvi, G. (2014, January 14\u201318). Audio-visual classification and detection of human manipulation actions. Proceedings of the 2014 IEEE\/RSJ International Conference on Intelligent Robots and Systems, Chicago, IL, USA.","DOI":"10.1109\/IROS.2014.6942983"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Salamon, J., and Bello, J.P. (2015, January 19\u201324). Unsupervised feature learning for urban sound classification. Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, QLD, Australia.","DOI":"10.1109\/ICASSP.2015.7177954"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Han, K., He, Y., Bagchi, D., Fosler-lussier, E., and Wang, D. (2015, January 6\u201310). Deep Neural Network Based Spectral Feature Mapping for Robust Speech Recognition. Proceedings of the 16th Annual Conference of the International Speech Communication Association (Interspeech), Dresden, Germany.","DOI":"10.21437\/Interspeech.2015-536"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"62","DOI":"10.1016\/j.specom.2018.02.009","article-title":"Speaker recognition from whispered speech: A tutorial survey and an application of time-varying linear prediction","volume":"99","author":"Vestman","year":"2018","journal-title":"Speech Commun."},{"key":"ref_16","first-page":"241","article-title":"Spoken digit recognition in Portuguese using line spectral frequencies","volume":"Volume 7637","author":"Silva","year":"2012","journal-title":"Computer Vision"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Bozkurt, E., Erzin, E., Erdem, C.E., and Erdem, T. (2010, January 23\u201326). Use of line spectral frequencies for emotion recognition from speech. Proceedings of the 2010 20th International Conference on Pattern Recognition, Istanbul, Turkey.","DOI":"10.1109\/ICPR.2010.903"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Singh, R.K., Saha, R., Pal, P.K., and Singh, G. (2018, January 22\u201323). Novel feature extraction algorithm using DWT and temporal statistical techniques for word dependent speaker\u2019s recognition. Proceedings of the 2018 Fourth International Conference on Research in Computational Intelligence and Communication Networks (ICRCICN), Kolkata, India.","DOI":"10.1109\/ICRCICN.2018.8718681"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhang, L., Wei, Y., Wang, S., Pan, D., Liang, S., and Xu, T. (2017, January 3\u20135). Specific two words lexical semantic recognition based on the wavelet transform of narrowband spectrogram. Proceedings of the 2017 First International Conference on Electronics Instrumentation & Information Systems (EIIS), Harbin, China.","DOI":"10.1109\/EIIS.2017.8298735"},{"key":"ref_20","first-page":"5","article-title":"Designing and implementing of intelligent emotional speech recognition with wavelet and neural network","volume":"7","author":"Zahra","year":"2016","journal-title":"Int. J. Adv. Comput. Sci. Appl."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Badshah, A.M., Ahmad, J., Rahim, N., and Baik, S.W. (2017, January 13\u201315). Speech emotion recognition from spectrograms with deep convolutional neural network. Proceedings of the 2017 International Conference on Platform Technology and Service (PlatCon), Busan, Korea.","DOI":"10.1109\/PlatCon.2017.7883728"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Stolar, M., Lech, M., Bolia, R.S., and Skinner, M. (2018, January 17\u201319). Acoustic Characteristics of Emotional Speech Using Spectrogram Image Classification. Proceedings of the 2018 12th International Conference on Signal Processing and Communication Systems (ICSPCS), Cairns, QLD, Australia.","DOI":"10.1109\/ICSPCS.2018.8631752"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Hyder, R., Ghaffarzadegan, S., Feng, Z., Hansen, J.H., and Hasan, T. (2017, January 20\u201324). Acoustic scene classification using a CNN-SuperVector system trained with auditory and spectrogram image features. Proceedings of the Interspeech 2017, Stockholm, Sweden.","DOI":"10.21437\/Interspeech.2017-431"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhang, H., McLoughlin, I., and Song, Y. (2015, January 19\u201324). Robust sound event recognition using convolutional neural networks. Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, QLD, Australia.","DOI":"10.1109\/ICASSP.2015.7178031"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Dahake, P.P., Shaw, K., and Malathi, P. (2016, January 9\u201310). Speaker dependent speech emotion recognition using MFCC and Support Vector Machine. Proceedings of the 2016 International Conference on Automatic Control and Dynamic Optimization Techniques (ICACDOT), Pune, India.","DOI":"10.1109\/ICACDOT.2016.7877753"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Chan, W., and Jaitly, N. (2017, January 5\u20139). Very deep convolutional networks for end-to-end speech recognition. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7953077"},{"key":"ref_27","unstructured":"Wazir, A.S.M.B., and Chuah, J.H. (2019, January 29). Spoken arabic digits recognition using deep learning. Proceedings of the 2019 IEEE International Conference on Automatic Control and Intelligent Systems (I2CACIS), Selangor, Malaysia."},{"key":"ref_28","unstructured":"Wazir, A.S.B., Karim, H.A., Abdullah, M.H.L., and Mansor, S. (2019, January 17\u201319). Acoustic pornography recognition using recurrent neural network. Proceedings of the 2019 IEEE International Conference on Signal and Image Processing Applications (ICSIPA), Kuala Lumpur, Malaysia."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1016\/j.eswa.2018.08.052","article-title":"Assessing the performances of different neural network architectures for the detection of screams and shouts in public transportation","volume":"117","author":"Laffitte","year":"2019","journal-title":"Expert Syst. Appl."},{"key":"ref_30","unstructured":"(2020, December 11). Act 620, Film Censorship; Commissioner of Law Revision with Percetakan Nasional Malaysia Bhd. Malaysia, Available online: http:\/\/www.agc.gov.my\/agcportal\/uploads\/files\/Publications\/LOM\/EN\/Act%20620.pdf."},{"key":"ref_31","unstructured":"Font, F., Roma, G., and Serra, X. (November, January 29). Freesound technical demo. Proceedings of the 21st ACM international conference on Information and knowledge management-CIKM\u201912, Maui, HI, USA."},{"key":"ref_32","first-page":"3586","article-title":"Audio Augmentation for Speech Recognition","volume":"Volume 2","author":"Ko","year":"2015","journal-title":"Proceedings of the 16th Annual Conference of the International Speech Communication Association (Interspeech)"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"316","DOI":"10.1016\/j.procs.2017.08.003","article-title":"Improving speech recognition using data augmentation and acoustic model fusion","volume":"112","author":"Rebai","year":"2017","journal-title":"Procedia Comput. Sci."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H. (2007, January 4\u20139). Greedy layer-wise training of deep networks. Proceedings of the 19th International Conference on Neural Information Processing Systems (NIPS\u201906), Vancouver, BC, Canada.","DOI":"10.7551\/mitpress\/7503.003.0024"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"451","DOI":"10.1109\/5326.897072","article-title":"Neural networks for classification: A survey","volume":"30","author":"Zhang","year":"2000","journal-title":"IEEE Trans. Syst. Man Cybern. Part. C (Appl. Rev.)"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"209","DOI":"10.1016\/S0893-6080(97)00120-2","article-title":"Multilayer neural networks and Bayes decision theory","volume":"11","author":"Funahashi","year":"1998","journal-title":"Neural Netw."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Bishop, C.M. (1995). Neural Networks for Pattern Recognition, Oxford University Press, Inc.","DOI":"10.1093\/oso\/9780198538493.001.0001"},{"key":"ref_38","unstructured":"Gish, H. (2012, January 25\u201330). A probabilistic approach to the understanding and training of neural network classifiers. Proceedings of the International Conference on Acoustics, Speech, and Signal Processing, Kyoto, Japan."},{"key":"ref_39","unstructured":"Lipton, Z.C., Berkowitz, J., and Elkan, C. (2015). A critical review of recurrent neural networks for sequence learning. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"1235","DOI":"10.1162\/neco_a_01199","article-title":"A review of recurrent neural networks: LSTM cells and network architectures","volume":"31","author":"Yu","year":"2019","journal-title":"Neural Comput."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Yao, Y., Zhang, S., Yang, S., and Gui, G. (2020). Learning attention representation with a multi-scale CNN for gear fault diagnosis under different working conditions. Sensors, 20.","DOI":"10.3390\/s20041233"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Li, T., Shi, J., Li, X., Wu, J., and Pan, F. (2019). Image encryption based on pixel-level diffusion with dynamic filtering and DNA-level permutation with 3D latin cubes. Entropy, 21.","DOI":"10.3390\/e21030319"},{"key":"ref_43","first-page":"2709","article-title":"Deep convolutional neural networks for image classification: A comprehensive review","volume":"2733","author":"Rawat","year":"2018","journal-title":"Neural Comput."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"107020","DOI":"10.1016\/j.apacoust.2019.107020","article-title":"Trends in audio signal feature extraction methods","volume":"158","author":"Sharma","year":"2020","journal-title":"Appl. Acoust."},{"key":"ref_45","unstructured":"Tuomas Virtanen, A.D. (2017, January 15\u201318). Transfer learning of weakly labelled audio. Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, New Paltz, NY, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/3\/710\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:13:27Z","timestamp":1760159607000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/3\/710"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,1,21]]},"references-count":45,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2021,2]]}},"alternative-id":["s21030710"],"URL":"https:\/\/doi.org\/10.3390\/s21030710","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,1,21]]}}}