{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T19:00:33Z","timestamp":1782586833325,"version":"3.54.5"},"reference-count":54,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2020,12,25]],"date-time":"2020-12-25T00:00:00Z","timestamp":1608854400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"TM R&amp;D, Malaysia","award":["MMUE\/180029"],"award-info":[{"award-number":["MMUE\/180029"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Video pornography and nudity detection aim to detect and classify people in videos into nude or normal for censorship purposes. Recent literature has demonstrated pornography detection utilising the convolutional neural network (CNN) to extract features directly from the whole frames and support vector machine (SVM) to classify the extracted features into two categories. However, existing methods were not able to detect the small-scale content of pornography and nudity in frames with diverse backgrounds. This limitation has led to a high false-negative rate (FNR) and misclassification of nude frames as normal ones. In order to address this matter, this paper explores the limitation of the existing convolutional-only approaches focusing the visual attention of CNN on the expected nude regions inside the frames to reduce the FNR. The You Only Look Once (YOLO) object detector was transferred to the pornography and nudity detection application to detect persons as regions of interest (ROIs), which were applied to CNN and SVM for nude\/normal classification. Several experiments were conducted to compare the performance of various CNNs and classifiers using our proposed dataset. It was found that ResNet101 with random forest outperformed other models concerning the F1-score of 90.03% and accuracy of 87.75%. Furthermore, an ablation study was performed to demonstrate the impact of adding the YOLO before the CNN. YOLO\u2013CNN was shown to outperform CNN-only in terms of accuracy, which was increased from 85.5% to 89.5%. Additionally, a new benchmark dataset with challenging content, including various human sizes and backgrounds, was proposed.<\/jats:p>","DOI":"10.3390\/sym13010026","type":"journal-article","created":{"date-parts":[[2020,12,25]],"date-time":"2020-12-25T09:30:19Z","timestamp":1608888619000},"page":"26","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":26,"title":["Transfer Detection of YOLO to Focus CNN\u2019s Attention on Nude Regions for Adult Content Detection"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5522-0033","authenticated-orcid":false,"given":"Nouar","family":"AlDahoul","sequence":"first","affiliation":[{"name":"Faculty of Engineering, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hezerul","family":"Abdul Karim","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohd Haris","family":"Lye Abdullah","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5382-6269","authenticated-orcid":false,"given":"Mohammad Faizal","family":"Ahmad Fauzi","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1225-1723","authenticated-orcid":false,"given":"Abdulaziz Saleh","family":"Ba Wazir","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sarina","family":"Mansor","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3005-4109","authenticated-orcid":false,"given":"John","family":"See","sequence":"additional","affiliation":[{"name":"Faculty of Computing and Informatics, Multimedia University, Cyberjaya 63100, Malaysia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2020,12,25]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Nuraisha, S., Pratama, F.I., Budianita, A., and Soeleman, M.A. (2017, January 7\u20138). Implementation of K-NN based on histogram at image recognition for pornography detection. Proceedings of the 2017 International Seminar on Application for Technology of Information and Communication (iSemantic), Semarang, Indonesia.","DOI":"10.1109\/ISEMANTIC.2017.8251834"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Garcia, M.B., Revano, T.F., Habal, B.G.M., Contreras, J.O., and Enriquez, J.B.R. (December, January 29). A Pornographic Image and Video Filtering Application Using Optimized Nudity Recognition and Detection Algorithm. Proceedings of the 2018 IEEE 10th International Conference on Humanoid, Nanotechnology, Information Technology, Communication and Control, Environment and Management (HNICEM), Baguio City, Philippines.","DOI":"10.1109\/HNICEM.2018.8666227"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"661","DOI":"10.1007\/s11042-012-1132-y","article-title":"A survey on visual adult image recognition","volume":"69","author":"Ries","year":"2014","journal-title":"Multimed. Tools Appl."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Santos, C., Dos Santos, E.M., and Souto, E. (2012, January 2\u20135). Nudity detection based on image zoning. Proceedings of the 2012 11th International Conference on Information Science, Signal Processing and their Applications (ISSPA), Montreal, QC, Canada.","DOI":"10.1109\/ISSPA.2012.6310454"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Moreira, D.C., and Fechine, J.M. (2018, January 8\u201313). A Machine Learning-based Forensic Discriminator of Pornographic and Bikini Images. Proceedings of the 2018 International Joint Conference on Neural Networks (IJCNN), Rio de Janeiro, Brazil.","DOI":"10.1109\/IJCNN.2018.8489100"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"ref_7","unstructured":"Moustafa, M. (2015). Applying deep learning to classify pornographic images and videos. Pacific Rim Symposium on Image and Video Technology. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"(2019). Automated Nudity Recognition using Very Deep Residual Learning Network. Int. J. Recent Technol. Eng., 8, 136\u2013141.","DOI":"10.35940\/ijrte.C1024.1083S19"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Nurhadiyatna, A., Cahyadi, S., Damatraseta, F., and Rianto, Y. (2017, January 23\u201326). Adult content classification through deep convolution neural network. Proceedings of the 2017 International Conference on Computer, Control, Informatics and its Applications (IC3INA), Jakarta, Indonesia.","DOI":"10.1109\/IC3INA.2017.8251749"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"432","DOI":"10.1016\/j.neucom.2017.07.012","article-title":"Adult content detection in videos with convolutional and recurrent neural networks","volume":"272","author":"Wehrmann","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1016\/j.neucom.2016.12.017","article-title":"Video pornography detection through deep learning techniques and motion information","volume":"230","author":"Perez","year":"2017","journal-title":"Neurocomputing"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Wang, Y., Jin, X., and Tan, X. (2016, January 25\u201328). Pornographic image recognition by strongly-supervised deep multiple instance learning. Proceedings of the 2016 IEEE International Conference on Image Processing (ICIP), Phoenix, AZ, USA.","DOI":"10.1109\/ICIP.2016.7533195"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"AlDahoul, N., Karim, H.A., Abdullah, M.H.L., Fauzi, M.F.A., Mansour, S., See, J., and Alfrahoul, N. (2019, January 17\u201319). Local Receptive Field-Extreme Learning Machine based Adult Content Detection. Proceedings of the 2019 IEEE International Conference on Signal and Image Processing Applications (ICSIPA), Kuala Lumpur, Malaysia.","DOI":"10.1109\/ICSIPA45851.2019.8977754"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"462","DOI":"10.1214\/aoms\/1177729392","article-title":"Stochastic Estimation of the Maximum of a Regression Function","volume":"23","author":"Kiefer","year":"1952","journal-title":"Ann. Math. Stat."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"223","DOI":"10.1137\/16M1080173","article-title":"Optimization Methods for Large-Scale Machine Learning","volume":"60","author":"Bottou","year":"2018","journal-title":"SIAM Rev."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009). ImageNet: A large-scale hierarchical image database. IEEE Conf. Comput. Vis. Pattern Recognit., 248\u2013255.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014). Microsoft COCO: Common objects in context. In Proceedings of the Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics). arXiv.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1145\/3065386","article-title":"ImageNet classification with deep convolutional neural networks","volume":"60","author":"Krizhevsky","year":"2017","journal-title":"Commun. ACM"},{"key":"ref_19","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016). Rethinking the inception architecture for computer vision. Conf. Proc., 2818\u20132826.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1109\/5254.708428","article-title":"Support vector machines","volume":"13","author":"Hearst","year":"1998","journal-title":"IEEE Intell. Syst."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long Short-Term Memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Da Silva, M.V., and Marana, A.N. (2019, January 19\u201322). Spatiotemporal CNNs for pornography detection in videos. Proceedings of the Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Madrid, Spain.","DOI":"10.1007\/978-3-030-13469-3_64"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning Spatiotemporal Features with 3D Convolutional Networks. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., and Paluri, M. (2018, January 18\u201323). A Closer Look at Spatiotemporal Convolutions for Action Recognition. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00675"},{"key":"ref_28","unstructured":"(2019, April 01). NPDI Pornography Database, 2013. Available online: https:\/\/sites.google.com\/site\/pornographydatabase\/."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"453","DOI":"10.1016\/j.cviu.2012.09.007","article-title":"Pooling in image representation: The visual codeword point of view","volume":"117","author":"Avila","year":"2013","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1016\/j.neucom.2015.09.135","article-title":"Pornographic image detection utilizing deep convolutional neural networks","volume":"210","author":"Nian","year":"2016","journal-title":"Neurocomputing"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Liu, B.-B., Su, J.-Y., Lu, Z.-M., and Li, Z. (2008, January 3\u20135). Pornographic Images Detection Based on CBIR and Skin Analysis. Proceedings of the 2008 Fourth International Conference on Semantics, Knowledge and Grid, Beijing, China.","DOI":"10.1109\/SKG.2008.48"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Doll\u00e1r, P., Belongie, S., and Perona, P. (September, January 31). The fastest pedestrian detector in the west. Proceedings of the British Machine Vision Conference, BMVC 2010-Proceedings, Aberystwyth, UK.","DOI":"10.5244\/C.24.68"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1639561","DOI":"10.1155\/2018\/1639561","article-title":"Real-Time Human Detection for Aerial Captured Video Sequences via Deep Models","volume":"2018","author":"AlDahoul","year":"2018","journal-title":"Comput. Intell. Neurosci."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"831","DOI":"10.1016\/j.procs.2018.07.112","article-title":"YOLO based Human Action Recognition and Localization","volume":"133","author":"Shinde","year":"2018","journal-title":"Procedia Comput. Sci."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Iva\u0161i\u0107-Kos, M., Kri\u0161to, M., and Pobar, M. (2019, January 16\u201317). Human detection in thermal imaging using YOLO. Proceedings of the 2019 5th International Conference on Computer and Technology Applications, Istanbul, Turkey.","DOI":"10.1145\/3323933.3324076"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Simoes, G.S., Wehrmann, J., and Barros, R.C. (2019, January 14\u201319). Attention-based Adversarial Training for Seamless Nudity Censorship. Proceedings of the 2019 International Joint Conference on Neural Networks (IJCNN), Budapest, Hungary.","DOI":"10.1109\/IJCNN.2019.8851849"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Ion, C., and Minea, C. (2019, January 7\u20139). Application of Image Classification for Fine-Grained Nudity Detection. Proceedings of the Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), Lake Tahoe, NV, USA.","DOI":"10.1007\/978-3-030-33720-9_1"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Connie, T., Al-Shabi, M., and Goh, M. (2018). Smart content recognition from images using a mixture of convolutional neural networks. IT Convergence and Security 2017, Springer.","DOI":"10.1007\/978-981-10-6451-7_2"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"166","DOI":"10.1016\/j.neucom.2018.08.080","article-title":"EFUI: An ensemble framework using uncertain inference for pornographic image recognition","volume":"322","author":"Shen","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"214","DOI":"10.1016\/j.knosys.2015.04.014","article-title":"Pornographic images recognition based on spatial pyramid partition and multi-instance ensemble learning","volume":"84","author":"Li","year":"2015","journal-title":"Knowl. Based Syst."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, Faster, Stronger. Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_43","unstructured":"Redmon, J., and Farhadi, A. (2018). YOLOv3: An incremental improvement. arXiv."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"175","DOI":"10.1080\/00031305.1992.10475879","article-title":"An Introduction to Kernel and Nearest-Neighbor Nonparametric Regression","volume":"46","author":"Altman","year":"1992","journal-title":"Am. Stat."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/A:1010933404324","article-title":"Random Forests","volume":"45","author":"Breiman","year":"2001","journal-title":"Mach. Learn."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"489","DOI":"10.1016\/j.neucom.2005.12.126","article-title":"Extreme learning machine: Theory and applications","volume":"70","author":"Huang","year":"2006","journal-title":"Neurocomputing"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"46","DOI":"10.1016\/j.forsciint.2016.09.010","article-title":"Pornography classification: The hidden clues in video space\u2013time","volume":"268","author":"Moreira","year":"2016","journal-title":"Forensic Sci. Int."},{"key":"ref_48","unstructured":"(2020, September 20). SciKit Learn Library for Machine Learning. Available online: https:\/\/scikit-learn.org\/."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Torralba, A. (July, January 26). Learning Deep Features for Discriminative Localization. Proceedings of Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.319"},{"key":"ref_50","unstructured":"Narayanan, B.N., Silva, M.S.D., Hardie, R.C., Kueterman, N.K., and Ali, R. (2019). Understanding Deep Neural Network Predictions for Medical Imaging Applications. arXiv."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Tan, M., Pang, R., and Le, Q.V. (2020, January 13\u201319). EfficientDet: Scalable and Efficient Object Detection. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"28","DOI":"10.33093\/jetap.2020.2.2.5","article-title":"Convolutional Neural Network-based Transfer Learning and Classification of Visual Contents for Film Censorship","volume":"2","author":"AlDahoul","year":"2020","journal-title":"J. Eng. Technol. Appl. Phys."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Mahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., and van der Maaten, L. (2018, January 8\u201314). Exploring the Limits of Weakly Supervised Pretraining. Proceedings of the European Conference on Computer Vision. In Proceedings of the Lecture Notes in Computer Science, Munich, Germany.","DOI":"10.1007\/978-3-030-01216-8_12"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Flaute, D., and Narayanan, B.N. (2020). Video captioning using weakly supervised convolutional neural networks, Proc. SPIE 11511. Appl. Mach. Learn.","DOI":"10.1117\/12.2568016"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/13\/1\/26\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:46:05Z","timestamp":1760179565000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/13\/1\/26"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,12,25]]},"references-count":54,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2021,1]]}},"alternative-id":["sym13010026"],"URL":"https:\/\/doi.org\/10.3390\/sym13010026","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,12,25]]}}}