{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T06:54:03Z","timestamp":1784098443497,"version":"3.55.0"},"reference-count":49,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2024,1,31]],"date-time":"2024-01-31T00:00:00Z","timestamp":1706659200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100004663","name":"National Science and Technology Council","doi-asserted-by":"publisher","award":["NSTC 112-2221-E-027-076-MY2"],"award-info":[{"award-number":["NSTC 112-2221-E-027-076-MY2"]}],"id":[{"id":"10.13039\/501100004663","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Violent attacks have been one of the hot issues in recent years. In the presence of closed-circuit televisions (CCTVs) in smart cities, there is an emerging challenge in apprehending criminals, leading to a need for innovative solutions. In this paper, the propose a model aimed at enhancing real-time emergency response capabilities and swiftly identifying criminals. This initiative aims to foster a safer environment and better manage criminal activity within smart cities. The proposed architecture combines an image-to-image stable diffusion model with violence detection and pose estimation approaches. The diffusion model generates synthetic data while the object detection approach uses YOLO v7 to identify violent objects like baseball bats, knives, and pistols, complemented by MediaPipe for action detection. Further, a long short-term memory (LSTM) network classifies the action attacks involving violent objects. Subsequently, an ensemble consisting of an edge device and the entire proposed model is deployed onto the edge device for real-time data testing using a dash camera. Thus, this study can handle violent attacks and send alerts in emergencies. As a result, our proposed YOLO model achieves a mean average precision (MAP) of 89.5% for violent attack detection, and the LSTM classifier model achieves an accuracy of 88.33% for violent action classification. The results highlight the model\u2019s enhanced capability to accurately detect violent objects, particularly in effectively identifying violence through the implemented artificial intelligence system.<\/jats:p>","DOI":"10.3390\/fi16020050","type":"journal-article","created":{"date-parts":[[2024,2,1]],"date-time":"2024-02-01T09:43:22Z","timestamp":1706780602000},"page":"50","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":19,"title":["Enhancing Smart City Safety and Utilizing AI Expert Systems for Violence Detection"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-9503-5689","authenticated-orcid":false,"given":"Pradeep","family":"Kumar","sequence":"first","affiliation":[{"name":"Department of Electro-Optical Engineering, National Taipei University of Technology, Taipei 10608, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-2741-8584","authenticated-orcid":false,"given":"Guo-Liang","family":"Shih","sequence":"additional","affiliation":[{"name":"Department of Electro-Optical Engineering, National Taipei University of Technology, Taipei 10608, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bo-Lin","family":"Guo","sequence":"additional","affiliation":[{"name":"Department of Electro-Optical Engineering, National Taipei University of Technology, Taipei 10608, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Siva Kumar","family":"Nagi","sequence":"additional","affiliation":[{"name":"Department of Electro-Optical Engineering, National Taipei University of Technology, Taipei 10608, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9584-661X","authenticated-orcid":false,"given":"Yibeltal Chanie","family":"Manie","sequence":"additional","affiliation":[{"name":"Department of Electro-Optical Engineering, National Taipei University of Technology, Taipei 10608, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3031-6407","authenticated-orcid":false,"given":"Cheng-Kai","family":"Yao","sequence":"additional","affiliation":[{"name":"Department of Electro-Optical Engineering, National Taipei University of Technology, Taipei 10608, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Michael Augustine","family":"Arockiyadoss","sequence":"additional","affiliation":[{"name":"Department of Electro-Optical Engineering, National Taipei University of Technology, Taipei 10608, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peng-Chun","family":"Peng","sequence":"additional","affiliation":[{"name":"Department of Electro-Optical Engineering, National Taipei University of Technology, Taipei 10608, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,1,31]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Baba, M., Gui, V., Cernazanu, C., and Pescaru, D. (2019). A Sensor Network Approach for Violence Detection in Smart Cities Using Deep Learning. Sensors, 19.","DOI":"10.3390\/s19071676"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Bai, T., Fu, S., and Yang, Q. (2022). Privacy-Preserving Object Detection with Secure Convolutional Neural Networks for Vehicular Edge Computing. Future Internet, 14.","DOI":"10.3390\/fi14110316"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Ali, S.A., Elsaid, S.A., Ateya, A.A., ElAffendi, M., and El-Latif, A.A.A. (2023). Enabling Technologies for Next-Generation Smart Cities: A Comprehensive Review and Research Directions. Future Internet, 15.","DOI":"10.3390\/fi15120398"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Ullah, F.U.M., Ullah, A., Muhammad, K., Haq, I.U., and Baik, S.W. (2019). Violence Detection Using Spatiotemporal Features with 3D Convolutional Neural Network. Sensors, 19.","DOI":"10.3390\/s19112472"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Aremu, T., Zhiyuan, L., Alameeri, R., Khan, M., and Saddik, A.E. (2022). SSIVD-Net: A novel salient super image classification & detection technique for weaponized violence. arXiv.","DOI":"10.21203\/rs.3.rs-3024402\/v2"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Jebur, S.A., Hussein, K.A., Hoomod, H.K., and Alzubaidi, L. (2023). Novel Deep Feature Fusion Framework for Multi-Scenario Violence Detection. Computers, 12.","DOI":"10.3390\/computers12090175"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Vosta, S., and Yow, K.-C.A. (2022). CNN-RNN Combined Structure for Real-World Violence Detection in Surveillance Cameras. Appl. Sci., 12.","DOI":"10.3390\/app12031021"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Alrashedy, H.H.N., Almansour, A.F., Ibrahim, D.M., and Hammoudeh, M.A.A. (2022). BrainGAN: Brain MRI Image Generation and Classification Framework Using GAN Architectures and CNN Models. Sensors, 22.","DOI":"10.3390\/s22114297"},{"key":"ref_9","unstructured":"Nichol, A.Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., Mcgrew, B., Sutskever, I., and Chen, M. (2022, January 17\u201323). GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. Proceedings of the International Conference on Machine Learning, PMLR, Baltimore, ML, USA."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Kawar, B., Zada, S., Lang, O., Tov, O., Chang, H., Dekel, T., Mosseri, I., and Irani, M. (2023, January 18\u201322). Imagic: Text-based real image editing with diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00582"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Avrahami, O., Lischinski, D., and Fried, O. (2022, January 18\u201324). Blended diffusion for text-driven editing of natural images. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01767"},{"key":"ref_12","unstructured":"Borji, A. (2022). Generated faces in the wild: Quantitative comparison of stable diffusion, mid-journey and dall-e 2. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022, January 18\u201324). High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"35365","DOI":"10.1109\/ACCESS.2018.2836950","article-title":"Machine learning and deep learning methods for cybersecurity","volume":"6","author":"Xin","year":"2018","journal-title":"IEEE Access"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Khan, S.U., Haq, I.U., Rho, S., Baik, S.W., and Lee, M.Y. (2019). Cover the Violence: A Novel Deep-Learning-Based Approach Towards Violence-Detection in Movies. Appl. Sci., 9.","DOI":"10.3390\/app9224963"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Maity, M., Banerjee, S., and Sinha, C.S. (2021, January 8\u201310). Faster R-CNN and YOLO based Vehicle detection: A Survey. Proceedings of the 5th International Conference on Computing Methodologies and Communication (ICCMC), Erode, India.","DOI":"10.1109\/ICCMC51019.2021.9418274"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Liu, K., Tang, H., He, S., Yu, Q., Xiong, Y., and Wang, N. (2021, January 22\u201324). Performance validation of YOLO variants for object detection. Proceedings of the 2021 International Conference on Bioinformatics and Intelligent Computing, Harbin, China.","DOI":"10.1145\/3448748.3448786"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Hussain, M. (2023). YOLO-v1 to YOLO-v8: The Rise of YOLO and Its Complementary Nature toward Digital Manufacturing and Industrial Defect Detection. Machines, 7.","DOI":"10.3390\/machines11070677"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Chen, D., and Ju, Y. (2020, January 4\u20136). SAR ship detection based on improved YOLOv3. Proceedings of the IET International Radar Conference (IET IRC 2020), Online.","DOI":"10.1049\/icp.2021.0710"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Li, Y., Zhao, Z., Luo, Y., and Qiu, Z. (2020). Real-Time Pattern-Recognition of GPR Images with YOLO v3 Implemented by Tensorflow. Sensors, 20.","DOI":"10.3390\/s20226476"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Wahyutama, A.B., and Hwang, M. (2022). YOLO-Based Object Detection for Separate Collection of Recyclables and Capacity Monitoring of Trash Bins. Electronics, 11.","DOI":"10.3390\/electronics11091323"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhou, F., Deng, H., Xu, Q., and Lan, X. (2023). CNTR-YOLO: Improved YOLOv5 Based on ConvNext and Transformer for Aircraft Detection in Remote Sensing Images. Electronics, 12.","DOI":"10.3390\/electronics12122671"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Xiao, Y., Chang, A., Wang, Y., Huang, Y., Yu, J., and Huo, L. (2022, January 20\u201322). Real-time Object Detection for Substation Security Early-warning with Deep Neural Network based on YOLO-V5. Proceedings of the IEEE IAS Global Conference on Emerging Technologies (GlobConET), Arad, Romania.","DOI":"10.1109\/GlobConET53749.2022.9872338"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Fan, L., Rao, H., and Yang, W. (2021). 3D Hand Pose Estimation Based on Five-Layer Ensemble CNN. Sensors, 21.","DOI":"10.3390\/s21020649"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Luvizon, D.C., Picard, D., and Tabia, H. (2018, January 18\u201322). 2D\/3D Pose Estimation and Action Recognition Using Multitask Deep Learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00539"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"7157","DOI":"10.1109\/TPAMI.2022.3222784","article-title":"AlphaPose: Whole-Body Regional Multi-Person Pose Estimation and Tracking in Real-Time","volume":"45","author":"Fang","year":"2023","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Yu, C., Xiao, B., Gao, C., Yuan, L., Zhang, L., Sang, N., and Wang, J. (2021, January 20\u201325). Lite-HRNet: A Lightweight High-Resolution Network. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01030"},{"key":"ref_28","unstructured":"(2023, June 14). Guns-Knives Object Detection Dataset. Available online: https:\/\/www.kaggle.com\/datasets\/iqmansingh\/guns-knives-object-detection."},{"key":"ref_29","unstructured":"(2023, June 14). Baseball Bat Dataset. Available online: https:\/\/images.cv\/dataset\/baseball-bat-image-classification-dataset."},{"key":"ref_30","first-page":"9975700","article-title":"Weapon Detection Using YOLO V3 for Smart Surveillance System","volume":"2021","author":"Pandey","year":"2021","journal-title":"Math. Probl. Eng."},{"key":"ref_31","unstructured":"Song, J., Meng, C., and Ermon, S. (2022, January 30). Denoising Diffusion Implicit Models. Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Gu, S., Chen, D., Bao, J., Wen, F., Zhang, B., Chen, D., Yuan, L., and Guo, B. (2022, January 18\u201324). Vector quantized diffusion model for text-to-image synthesis. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01043"},{"key":"ref_33","first-page":"36479","article-title":"Photorealistic text-to-image diffusion models with deep language understanding","volume":"Volume 35","author":"Saharia","year":"2022","journal-title":"Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Hemmatirad, K., Babaie, M., Afshari, M., Maleki, D., Saiadi, M., and Tizhoosh, H.R. (2022, January 11\u201314). Quality Control of Whole Slide Images using the YOLO Concept. Proceedings of the IEEE 10th International Conference on Healthcare Informatics (ICHI), Rochester, MN, USA.","DOI":"10.1109\/ICHI54592.2022.00049"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"133936","DOI":"10.1109\/ACCESS.2022.3230894","article-title":"Efficient Detection Model of Steel Strip Surface Defects Based on YOLO-V7","volume":"10","author":"Wang","year":"2022","journal-title":"IEEE Access"},{"key":"ref_36","first-page":"677","article-title":"Underwater Target Detection Based on Improved YOLOv7","volume":"3","author":"Kaiyue","year":"2023","journal-title":"J. Mar. Sci. Eng."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Kumar, P., Shih, G.-L., Yao, C.-K., Hayle, S.T., Manie, Y.C., and Peng, P.-C. (2023). Intelligent Vibration Monitoring System for Smart Industry Utilizing Optical Fiber Sensor Combined with Machine Learning. Electronics, 12.","DOI":"10.3390\/electronics12204302"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Chen, K.-Y., Shin, J., Hasan, M.A.M., Liaw, J.-J., Yuichi, O., and Tomioka, Y. (2022). Fitness Movement Types and Completeness Detection Using a Transfer-Learning-Based Deep Neural Network. Sensors, 22.","DOI":"10.3390\/s22155700"},{"key":"ref_39","unstructured":"(2023, November 14). MediaPipe: Pose Landmark Detection Guide. Available online: https:\/\/developers.google.com\/mediapipe."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Zeng, Y., Ye, W., Stutheit-Zhao, E.Y., Han, M., Bratman, S.V., Pugh, T.J., and He, H.H. (2023). MEDIPIPE: An automated and comprehensive pipeline for cfMeDIP-seq data quality control and analysis. Bioinformatics, 39.","DOI":"10.1093\/bioinformatics\/btad423"},{"key":"ref_41","unstructured":"Staudemeyer, R.C., and Morris, E.R. (2019). Understanding LSTM\u2014A Tutorial into Long Short-Term Memory Recurrent Neural Networks. arXiv."},{"key":"ref_42","unstructured":"Zhou, C., Sun, C., Liu, Z., and Lau, F.C.M. (2015). A C-LSTM Neural Network for Text Classification. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Ghourabi, A., Mahmood, M.A., and Alzubi, Q.M. (2020). A Hybrid CNN-LSTM Model for SMS Spam Detection in Arabic and English Messages. Future Internet, 12.","DOI":"10.3390\/fi12090156"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"428","DOI":"10.1016\/j.sysarc.2019.01.011","article-title":"A Survey on Optimized Implementation of Deep Learning Models on the NVIDIA Jetson Platform","volume":"97","author":"Mittal","year":"2019","journal-title":"J. Syst. Archit."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Shi, Z. (2021, January 12\u201314). Optimized Yolov3 Deployment on Jetson TX2 with Pruning and Quantization. Proceedings of the 2021 IEEE 3rd International Conference on Frontiers Technology of Information and Computer (ICFTIC), Greenville, SC, USA.","DOI":"10.1109\/ICFTIC54370.2021.9647400"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Chumuang, N., Hiranchan, S., Ketcham, M., Yimyam, W., Pramkeaw, P., and Tangwannawit, S. (2020, January 18\u201320). Developed Credit Card Fraud Detection Alert Systems via the Notification of LINE Application. Proceedings of the 2020 15th International Joint Symposium on Artificial Intelligence and Natural Language Processing (iSAI-NLP), Bangkok, Thailand.","DOI":"10.1109\/iSAI-NLP51646.2020.9376829"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Kumar, P., Li, C.-Y., Guo, B.-L., Manie, Y.C., Yao, C.-K., and Peng, P.-C. (2023, January 9\u201311). Detection of Acrimonious Attacks using Deep Learning Techniques and Edge Computing Devices. Proceedings of the 2023 International Conference on Consumer Electronics\u2014Taiwan (ICCE-Taiwan), Pingtung, Taiwan.","DOI":"10.1109\/ICCE-Taiwan58799.2023.10226915"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"121691","DOI":"10.1016\/j.eswa.2023.121691","article-title":"An automatic fine-grained violence detection system for animation based on modified faster R-CNN","volume":"237","author":"Tang","year":"2024","journal-title":"Expert Syst. Appl."},{"key":"ref_49","unstructured":"Tufail, H., Nazeef, U.H., Muhammad, F., and Muhammad, S. (2021, January 20\u201321). Application of Deep Learning for Weapons Detection in Surveillance Videos. Proceedings of the 2021 International Conference on Digital Futures and Transformative Technologies (ICoDT2), Islamabad, Pakistan."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/16\/2\/50\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T13:52:46Z","timestamp":1760104366000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/16\/2\/50"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,31]]},"references-count":49,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2024,2]]}},"alternative-id":["fi16020050"],"URL":"https:\/\/doi.org\/10.3390\/fi16020050","relation":{},"ISSN":["1999-5903"],"issn-type":[{"value":"1999-5903","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,31]]}}}