{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T14:47:16Z","timestamp":1785336436454,"version":"3.55.0"},"reference-count":143,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2023,5,24]],"date-time":"2023-05-24T00:00:00Z","timestamp":1684886400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Ho Chi Minh City Open University, Vietnam"},{"name":"Institute of Automation, Chinese Academy of Sciences"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Anomaly detection in video surveillance is a highly developed subject that is attracting increased attention from the research community. There is great demand for intelligent systems with the capacity to automatically detect anomalous events in streaming videos. Due to this, a wide variety of approaches have been proposed to build an effective model that would ensure public security. There has been a variety of surveys of anomaly detection, such as of network anomaly detection, financial fraud detection, human behavioral analysis, and many more. Deep learning has been successfully applied to many aspects of computer vision. In particular, the strong growth of generative models means that these are the main techniques used in the proposed methods. This paper aims to provide a comprehensive review of the deep learning-based techniques used in the field of video anomaly detection. Specifically, deep learning-based approaches have been categorized into different methods by their objectives and learning metrics. Additionally, preprocessing and feature engineering techniques are discussed thoroughly for the vision-based domain. This paper also describes the benchmark databases used in training and detecting abnormal human behavior. Finally, the common challenges in video surveillance are discussed, to offer some possible solutions and directions for future research.<\/jats:p>","DOI":"10.3390\/s23115024","type":"journal-article","created":{"date-parts":[[2023,5,25]],"date-time":"2023-05-25T02:30:06Z","timestamp":1684981806000},"page":"5024","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":160,"title":["Deep Learning-Based Anomaly Detection in Video Surveillance: A Survey"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2404-4214","authenticated-orcid":false,"given":"Huu-Thanh","family":"Duong","sequence":"first","affiliation":[{"name":"Faculty of Information Technology, Ho Chi Minh City Open University, 97 Vo Van Tan, District 3, Ho Chi Minh City 700000, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2289-8128","authenticated-orcid":false,"given":"Viet-Tuan","family":"Le","sequence":"additional","affiliation":[{"name":"Faculty of Information Technology, Ho Chi Minh City Open University, 97 Vo Van Tan, District 3, Ho Chi Minh City 700000, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3464-3894","authenticated-orcid":false,"given":"Vinh Truong","family":"Hoang","sequence":"additional","affiliation":[{"name":"Faculty of Information Technology, Ho Chi Minh City Open University, 97 Vo Van Tan, District 3, Ho Chi Minh City 700000, Vietnam"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,5,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1541880.1541882","article-title":"Anomaly detection: A survey","volume":"41","author":"Chandola","year":"2009","journal-title":"ACM Comput. Surv."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Sultani, W., Chen, C., and Shah, M. (2018, January 18\u201323). Real-world anomaly detection in surveillance videos. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00678"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1992","DOI":"10.1109\/TIP.2017.2670780","article-title":"Deep-cascade: Cascading 3d deep neural networks for fast anomaly detection and localization in crowded scenes","volume":"26","author":"Sabokrou","year":"2017","journal-title":"IEEE Trans. Image Process."},{"key":"ref_4","unstructured":"Medel, J.R., and Savakis, A. (2016). Anomaly detection in video using predictive convolutional long short-term memory networks. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"277","DOI":"10.1016\/j.patcog.2018.01.025","article-title":"A novel random forests based class incremental learning method for activity recognition","volume":"78","author":"Hu","year":"2018","journal-title":"Pattern Recognit."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"6955","DOI":"10.1007\/s11042-017-4614-0","article-title":"Action recognition based on hierarchical dynamic Bayesian network","volume":"77","author":"Xiao","year":"2018","journal-title":"Multimed. Tools Appl."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"6369","DOI":"10.1109\/JSEN.2018.2845749","article-title":"Activity recognition for incomplete spinal cord injury subjects using hidden Markov models","volume":"18","author":"Sok","year":"2018","journal-title":"IEEE Sens. J."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1007\/s10044-016-0570-y","article-title":"The joint use of sequence features combination and modified weighted SVM for improving daily activity recognition","volume":"21","author":"Abidine","year":"2018","journal-title":"Pattern Anal. Appl."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"258","DOI":"10.1049\/iet-cvi.2015.0271","article-title":"Video anomaly detection using deep incremental slow feature analysis network","volume":"10","author":"Hu","year":"2016","journal-title":"IET Comput. Vis."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"504","DOI":"10.1016\/j.patcog.2017.07.013","article-title":"Human action recognition in RGB-D videos using motion sequence information and deep learning","volume":"72","author":"Ijjina","year":"2017","journal-title":"Pattern Recognit."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"9718","DOI":"10.1109\/JSEN.2018.2866806","article-title":"Multi-resident activity recognition in a smart home using RGB activity image and DCNN","volume":"18","author":"Tan","year":"2018","journal-title":"IEEE Sens. J."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1109\/MCI.2018.2840738","article-title":"Recent trends in deep learning based natural language processing","volume":"13","author":"Young","year":"2018","journal-title":"IEEE Comput. Intell. Mag."},{"key":"ref_13","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20136). Imagenet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"3941","DOI":"10.1007\/s00521-016-2294-8","article-title":"Deep learning in vision-based static hand gesture recognition","volume":"28","author":"Oyedotun","year":"2017","journal-title":"Neural Comput. Appl."},{"key":"ref_15","unstructured":"Weingaertner, T., Hassfeld, S., and Dillmann, R. (1997, January 16). Human motion analysis: A review. Proceedings of the IEEE Nonrigid and Articulated Motion Workshop, San Juan, PR, USA."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"428","DOI":"10.1006\/cviu.1998.0744","article-title":"Human motion analysis: A review","volume":"73","author":"Aggarwal","year":"1999","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"192","DOI":"10.1049\/ip-vis:20041147","article-title":"Intelligent distributed surveillance systems: A review","volume":"152","author":"Valera","year":"2005","journal-title":"IEEE Proc. Vision Image Signal Process."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Ko, T. (2008, January 15\u201317). A survey on behavior analysis in video surveillance for homeland security applications. Proceedings of the 2008 37th IEEE Applied Imagery Pattern Recognition Workshop, Washington, DC, USA.","DOI":"10.1109\/AIPR.2008.4906450"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"865","DOI":"10.1109\/TSMCC.2011.2178594","article-title":"Video-based abnormal human behavior recognition\u2014A review","volume":"42","author":"Popoola","year":"2012","journal-title":"IEEE Trans. Syst. Man Cybern. Part C Appl. Rev."},{"key":"ref_20","first-page":"782783783","article-title":"A first stage comparative survey on vision-based human activity recognition","volume":"24","author":"Tsitsoulis","year":"2013","journal-title":"Int. J. Tools"},{"key":"ref_21","first-page":"13","article-title":"Intelligent video surveillance systems for public spaces\u2014A survey","volume":"8","author":"Frejlichowski","year":"2014","journal-title":"J. Theor. Appl. Comput. Sci."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"367","DOI":"10.1109\/TCSVT.2014.2358029","article-title":"Crowded scene analysis: A survey","volume":"25","author":"Li","year":"2014","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"97","DOI":"10.1016\/j.eswa.2016.06.011","article-title":"A survey on using domain and contextual knowledge for human activity recognition in video streams","volume":"63","author":"Onofri","year":"2016","journal-title":"Expert Syst. Appl."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1016\/j.imavis.2017.01.010","article-title":"Going deeper into action recognition: A survey","volume":"60","author":"Herath","year":"2017","journal-title":"Image Vis. Comput."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"480","DOI":"10.1016\/j.eswa.2017.09.029","article-title":"Abnormal behavior recognition for intelligent video surveillance systems: A review","volume":"91","author":"Mabrouk","year":"2018","journal-title":"Expert Syst. Appl."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"118","DOI":"10.1016\/j.cviu.2018.04.007","article-title":"RGB-D-based human motion recognition with deep learning: A survey","volume":"171","author":"Wang","year":"2018","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3158645","article-title":"Activity recognition with evolving data streams: A review","volume":"51","author":"Abdallah","year":"2018","journal-title":"ACM Comput. Surv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Kiran, B.R., Thomas, D.M., and Parakkal, R. (2018). An overview of deep learning based methods for unsupervised and semi-supervised anomaly detection in videos. J. Imaging, 4.","DOI":"10.3390\/jimaging4020036"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhang, H.B., Zhang, Y.X., Zhong, B., Lei, Q., Yang, L., Du, J.X., and Chen, D.S. (2019). A comprehensive survey of vision-based human action recognition methods. Sensors, 19.","DOI":"10.3390\/s19051005"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"271","DOI":"10.3233\/IDT-170035","article-title":"Survey and analysis of human activity recognition in surveillance videos","volume":"13","author":"Raval","year":"2019","journal-title":"Intell. Decis. Technol."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1016\/j.engappai.2018.08.014","article-title":"A review of state-of-the-art techniques for abnormal human activity recognition","volume":"77","author":"Dhiman","year":"2019","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Ramachandra, B., Jones, M.J., and Vatsavai, R.R. (2020). A Survey of Single-Scene Video Anomaly Detection. arXiv.","DOI":"10.1109\/TPAMI.2020.3040591"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"107561","DOI":"10.1016\/j.patcog.2020.107561","article-title":"Sensor-based and vision-based human activity recognition: A comprehensive survey","volume":"108","author":"Dang","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"19","DOI":"10.1016\/j.jnca.2015.11.016","article-title":"A survey of network anomaly detection techniques","volume":"60","author":"Ahmed","year":"2016","journal-title":"J. Netw. Comput. Appl."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1016\/j.jnca.2018.12.006","article-title":"A holistic review of network anomaly detection systems: A comprehensive survey","volume":"128","author":"Moustafa","year":"2019","journal-title":"J. Netw. Comput. Appl."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"447","DOI":"10.1007\/s11235-018-0475-8","article-title":"A comprehensive survey on network anomaly detection","volume":"70","author":"Fernandes","year":"2019","journal-title":"Telecommun. Syst."},{"key":"ref_37","unstructured":"CASIA (2020, November 20). CASIA Action Database. Available online: http:\/\/www.cbsr.ia.ac.cn\/english\/Action\/%20Databases\/%20EN.asp."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"555","DOI":"10.1109\/TPAMI.2007.70825","article-title":"Robust real-time unusual event detection using multiple fixed-location monitors","volume":"30","author":"Adam","year":"2008","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Mehran, R., Oyama, A., and Shah, M. (2009, January 20\u201325). Abnormal crowd behavior detection using social force model. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206641"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Zaharescu, A., and Wildes, R. (2010, January 5\u201311). Anomalous behaviour detection using spatiotemporal oriented energies, subset inclusion histogram comparison and event-driven processing. Proceedings of the European Conference on Computer Vision, Heraklion Crete, Greece.","DOI":"10.1007\/978-3-642-15549-9_41"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Lu, C., Shi, J., and Jia, J. (2013, January 1\u20138). Abnormal event detection at 150 FPS in MATLAB. Proceedings of the 2013 IEEE international Conference on Computer Vision, Sydney, Australia.","DOI":"10.1109\/ICCV.2013.338"},{"key":"ref_42","unstructured":"Statistical Visual Computing Lab (2020, November 20). UCSD Anomaly Data Set. Available online: http:\/\/www.svcl.ucsd.edu\/projects\/anomaly\/."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Liu, W., Luo, W., Lian, D., and Gao, S. (2018, January 18\u201323). Future frame prediction for anomaly detection\u2014A new baseline. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00684"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Ramachandra, B., and Jones, M. (2020, January 1\u20135). Street Scene: A new dataset and evaluation protocol for video anomaly detection. Proceedings of the 2020 IEEE Winter Conference on Applications of Computer Vision, Snowmass, CO, USA.","DOI":"10.1109\/WACV45572.2020.9093457"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"635","DOI":"10.1016\/j.patcog.2017.09.040","article-title":"A deep convolutional neural network for video sequence background subtraction","volume":"76","author":"Babaee","year":"2018","journal-title":"Pattern Recognit."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1016\/j.ins.2016.04.049","article-title":"Statistical feature bag based background subtraction for local change detection","volume":"366","author":"Subudhi","year":"2016","journal-title":"Inf. Sci."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"406","DOI":"10.1109\/TMC.2015.2418775","article-title":"Real-time and robust compressive background subtraction for embedded camera networks","volume":"15","author":"Shen","year":"2015","journal-title":"IEEE Trans. Mob. Comput."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"2105","DOI":"10.1109\/TCSVT.2017.2711659","article-title":"WeSamBE: A weight-sample-based method for background subtraction","volume":"28","author":"Jiang","year":"2017","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"23023","DOI":"10.1007\/s11042-017-5460-9","article-title":"End-to-end video background subtraction with 3d convolutional neural networks","volume":"77","author":"Sakkos","year":"2018","journal-title":"Multimed. Tools Appl."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Minematsu, T., Shimada, A., Uchiyama, H., and Taniguchi, R.I. (2018). Analytics of deep neural network-based background subtraction. J. Imaging, 4.","DOI":"10.3390\/jimaging4060078"},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"9692","DOI":"10.1109\/TIE.2018.2881943","article-title":"Activity recognition using temporal optical flow convolutional features and multilayer LSTM","volume":"66","author":"Ullah","year":"2018","journal-title":"IEEE Trans. Ind. Electron."},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"16387","DOI":"10.1007\/s00521-018-3951-x","article-title":"Human activity recognition via optical flow: Decomposing activities into basic actions","volume":"32","author":"Ladjailia","year":"2020","journal-title":"Neural Comput. Appl."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"107140","DOI":"10.1016\/j.patcog.2019.107140","article-title":"Human activity recognition from UAV-captured video sequences","volume":"100","author":"Mliki","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"30599","DOI":"10.1007\/s11042-018-6425-3","article-title":"Depth based enlarged temporal dimension of 3D deep convolutional network for activity recognition","volume":"78","author":"Singh","year":"2019","journal-title":"Multimed. Tools Appl."},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"1015","DOI":"10.1016\/j.patcog.2016.07.024","article-title":"Mining intricate temporal rules for recognizing complex activities of daily living under uncertainty","volume":"60","author":"Liu","year":"2016","journal-title":"Pattern Recognit."},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"013106","DOI":"10.1117\/1.OE.57.1.013106","article-title":"Moving target segmentation using Markov random field-based evaluation metric in infrared videos","volume":"57","author":"Sun","year":"2018","journal-title":"Opt. Eng."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Lai, Y., Han, Y., and Wang, Y. (2021, January 7\u201310). Anomaly detection with prototype-guided discriminative latent embeddings. Proceedings of the 2021 IEEE International Conference on Data Mining (ICDM), Auckland, New Zealand.","DOI":"10.1109\/ICDM51629.2021.00040"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Doshi, K., and Yilmaz, Y. (2020, January 14\u201319). Any-shot sequential anomaly detection in surveillance videos. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00475"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Cai, R., Zhang, H., Liu, W., Gao, S., and Hao, Z. (2021, January 2\u20139). Appearance-motion memory consistency network for video anomaly detection. Proceedings of the AAAI Conference on Artificial Intelligence, Online.","DOI":"10.1609\/aaai.v35i2.16177"},{"key":"ref_60","doi-asserted-by":"crossref","first-page":"3543","DOI":"10.1007\/s11042-018-6034-1","article-title":"Human activity recognition in egocentric video using HOG, GiST and color features","volume":"79","author":"Kumar","year":"2020","journal-title":"Multimed. Tools Appl."},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"3198","DOI":"10.1109\/JSEN.2016.2519679","article-title":"A triaxial accelerometer-based human activity recognition via EEMD-based features and game-theory-based feature selection","volume":"16","author":"Wang","year":"2016","journal-title":"IEEE Sens. J."},{"key":"ref_62","unstructured":"Roy, P.K., and Om, H. (2018). Advances in Soft Computing and Machine Learning in Image Processing, Springer."},{"key":"ref_63","doi-asserted-by":"crossref","first-page":"284","DOI":"10.1016\/j.compeleceng.2016.06.004","article-title":"Human action recognition using fusion of features for unconstrained video sequences","volume":"70","author":"Patel","year":"2018","journal-title":"Comput. Electr. Eng."},{"key":"ref_64","first-page":"22","article-title":"Video based human activity detection, recognition and classification of actions using SVM","volume":"6","author":"Jagadeesh","year":"2018","journal-title":"Trans. Mach. Learn. Artif. Intell."},{"key":"ref_65","doi-asserted-by":"crossref","first-page":"60","DOI":"10.18201\/ijisae.2019151257","article-title":"Human Activity Recognition on Real Time and Offline Dataset","volume":"7","author":"Kale","year":"2019","journal-title":"Int. J. Intell. Syst. Appl. Eng."},{"key":"ref_66","unstructured":"Bux, A., Angelov, P., and Habib, Z. (2017). Advances in Computational Intelligence Systems, Springer."},{"key":"ref_67","doi-asserted-by":"crossref","first-page":"573","DOI":"10.1080\/21681163.2017.1298472","article-title":"Activity representation by SURF-based templates","volume":"6","author":"Ahad","year":"2018","journal-title":"Comput. Methods Biomech. Biomed. Eng. Imaging Vis."},{"key":"ref_68","doi-asserted-by":"crossref","first-page":"612","DOI":"10.1016\/j.patcog.2017.12.007","article-title":"Motion analysis: Action detection, recognition and evaluation based on motion capture data","volume":"76","author":"Patrona","year":"2018","journal-title":"Pattern Recognit."},{"key":"ref_69","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1016\/j.patcog.2018.02.011","article-title":"Structured dynamic time warping for continuous hand trajectory gesture recognition","volume":"80","author":"Tang","year":"2018","journal-title":"Pattern Recognit."},{"key":"ref_70","doi-asserted-by":"crossref","first-page":"2567","DOI":"10.1007\/s42835-019-00278-8","article-title":"Vision-based Human Activity recognition system using depth silhouettes: A Smart home system for monitoring the residents","volume":"14","author":"Kim","year":"2019","journal-title":"J. Electr. Eng. Technol."},{"key":"ref_71","doi-asserted-by":"crossref","first-page":"54","DOI":"10.1016\/j.neucom.2015.03.097","article-title":"Recognizing human actions using novel space-time volume binary patterns","volume":"173","author":"Baumann","year":"2016","journal-title":"Neurocomputing"},{"key":"ref_72","doi-asserted-by":"crossref","first-page":"351","DOI":"10.1007\/s00138-014-0652-z","article-title":"Local polynomial space\u2013time descriptors for action classification","volume":"27","author":"Kihl","year":"2016","journal-title":"Mach. Vis. Appl."},{"key":"ref_73","doi-asserted-by":"crossref","first-page":"12645","DOI":"10.1007\/s11042-016-3630-9","article-title":"Sparse coding-based space-time video representation for action recognition","volume":"76","author":"Fu","year":"2017","journal-title":"Multimed. Tools Appl."},{"key":"ref_74","doi-asserted-by":"crossref","first-page":"1709","DOI":"10.1109\/TAES.2018.2799758","article-title":"Deep convolutional autoencoder for radar-based classification of similar aided and unaided human activities","volume":"54","year":"2018","journal-title":"IEEE Trans. Aerosp. Electron. Syst."},{"key":"ref_75","doi-asserted-by":"crossref","first-page":"2123","DOI":"10.1109\/TPAMI.2015.2505295","article-title":"Multimodal multipart learning for action recognition in depth videos","volume":"38","author":"Shahroudy","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_76","doi-asserted-by":"crossref","first-page":"2653","DOI":"10.1109\/TMM.2019.2903455","article-title":"Multi-person pose estimation using bounding box constraint and LSTM","volume":"21","author":"Li","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_77","doi-asserted-by":"crossref","first-page":"402","DOI":"10.1016\/j.patcog.2017.06.006","article-title":"Generation of human depth images with body part labels for complex human pose recognition","volume":"71","author":"Nishi","year":"2017","journal-title":"Pattern Recognit."},{"key":"ref_78","doi-asserted-by":"crossref","first-page":"117","DOI":"10.1016\/j.cviu.2016.10.010","article-title":"Detecting anomalous events in videos by learning deep representations of appearance and motion","volume":"156","author":"Xu","year":"2017","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_79","doi-asserted-by":"crossref","first-page":"443","DOI":"10.1016\/j.patcog.2015.09.005","article-title":"Combining motion and appearance cues for anomaly detection","volume":"51","author":"Zhang","year":"2016","journal-title":"Pattern Recognit."},{"key":"ref_80","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1016\/j.sigpro.2017.08.016","article-title":"Skeleton embedded motion body partition for human action recognition using depth sequences","volume":"143","author":"Ji","year":"2018","journal-title":"Signal Process."},{"key":"ref_81","doi-asserted-by":"crossref","unstructured":"Hasan, M., Choi, J., Neumann, J., Roy-Chowdhury, A.K., and Davis, L.S. (2016, January 27\u201330). Learning temporal regularity in video sequences. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.86"},{"key":"ref_82","doi-asserted-by":"crossref","first-page":"1122","DOI":"10.1049\/el.2016.0440","article-title":"Video anomaly detection and localisation based on the sparsity and reconstruction error of auto-encoder","volume":"52","author":"Sabokrou","year":"2016","journal-title":"Electron. Lett."},{"key":"ref_83","doi-asserted-by":"crossref","first-page":"13173","DOI":"10.1007\/s11042-017-4940-2","article-title":"Dynamic video anomaly detection and localization using sparse denoising autoencoders","volume":"77","author":"Narasimhan","year":"2018","journal-title":"Multimed. Tools Appl."},{"key":"ref_84","doi-asserted-by":"crossref","unstructured":"Zhao, Y., Deng, B., Shen, C., Liu, Y., Lu, H., and Hua, X.S. (2017, January 23\u201327). Spatio-temporal autoencoder for video anomaly detection. Proceedings of the 25th ACM International Conference on Multimedia, Mountain View, CA, USA.","DOI":"10.1145\/3123266.3123451"},{"key":"ref_85","doi-asserted-by":"crossref","first-page":"88","DOI":"10.1016\/j.cviu.2018.02.006","article-title":"Deep-anomaly: Fully convolutional neural network for fast anomaly detection in crowded scenes","volume":"172","author":"Sabokrou","year":"2018","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_86","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1016\/j.patrec.2017.07.016","article-title":"A study of deep convolutional auto-encoders for anomaly detection in videos","volume":"105","author":"Ribeiro","year":"2018","journal-title":"Pattern Recognit. Lett."},{"key":"ref_87","doi-asserted-by":"crossref","unstructured":"Sabzalian, B., Marvi, H., and Ahmadyfard, A. (2019, January 6\u20137). Deep and Sparse features For Anomaly Detection and Localization in video. Proceedings of the 2019 4th International Conference on Pattern Recognition and Image Analysis (IPRIA), Tehran, Iran.","DOI":"10.1109\/PRIA.2019.8786007"},{"key":"ref_88","unstructured":"Landi, F., Snoek, C.G., and Cucchiara, R. (2019). Anomaly Locality in Video Surveillance. arXiv."},{"key":"ref_89","doi-asserted-by":"crossref","first-page":"2537","DOI":"10.1109\/TIFS.2019.2900907","article-title":"Anomalynet: An anomaly detection network for video surveillance","volume":"14","author":"Zhou","year":"2019","journal-title":"IEEE Trans. Inf. Forensics Secur."},{"key":"ref_90","unstructured":"Zhu, Y., and Newsam, S. (2019). Motion-aware feature for improved video anomaly detection. arXiv."},{"key":"ref_91","doi-asserted-by":"crossref","unstructured":"Lin, S., Yang, H., Tang, X., Shi, T., and Chen, L. (2019, January 18\u201321). Social MIL: Interaction-Aware for Crowd Anomaly Detection. Proceedings of the 2019 16th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), Taipei, Taiwan.","DOI":"10.1109\/AVSS.2019.8909882"},{"key":"ref_92","doi-asserted-by":"crossref","unstructured":"Gong, D., Liu, L., Le, V., Saha, B., Mansour, M.R., Venkatesh, S., and van den Hengel, A. (November, January 27). Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly detection. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00179"},{"key":"ref_93","doi-asserted-by":"crossref","first-page":"407","DOI":"10.1016\/j.jvcir.2019.02.035","article-title":"Generalization of feature embeddings transferred from different video anomaly detection domains","volume":"60","author":"Ribeiro","year":"2019","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_94","doi-asserted-by":"crossref","unstructured":"Ionescu, R.T., Khan, F.S., Georgescu, M.I., and Shao, L. (2019, January 15\u201320). Object-centric auto-encoders and dummy anomalies for abnormal event detection in video. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00803"},{"key":"ref_95","doi-asserted-by":"crossref","first-page":"1070","DOI":"10.1109\/TPAMI.2019.2944377","article-title":"Video anomaly detection with sparse coding inspired deep neural networks","volume":"43","author":"Luo","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_96","doi-asserted-by":"crossref","first-page":"394","DOI":"10.1109\/TMM.2019.2929931","article-title":"Video anomaly detection and localization based on an adaptive intra-frame classification network","volume":"22","author":"Xu","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_97","doi-asserted-by":"crossref","first-page":"102920","DOI":"10.1016\/j.cviu.2020.102920","article-title":"Video anomaly detection and localization via Gaussian mixture fully convolutional variational autoencoder","volume":"195","author":"Fan","year":"2020","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_98","unstructured":"Wikipedia (2023, April 04). Autoencoder\u2014Wikipedia, The Free Encyclopedia. Available online: http:\/\/en.wikipedia.org\/w\/index.php?title=Autoencoder&oldid=1141727025."},{"key":"ref_99","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_100","doi-asserted-by":"crossref","unstructured":"Aich, A., Peng, K.C., and Roy-Chowdhury, A.K. (2023, January 2\u20137). Cross-Domain Video Anomaly Detection without Target Domain Adaptation. Proceedings of the 2023 IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV56688.2023.00261"},{"key":"ref_101","doi-asserted-by":"crossref","unstructured":"Yang, Z., Wu, P., Liu, J., and Liu, X. (2022, January 23\u201327). Dynamic Local Aggregation Network with Adaptive Clusterer for Anomaly Detection. Proceedings of the Computer Vision\u2013ECCV 2022: 17th European Conference, Tel Aviv, Israel. Proceedings, Part IV.","DOI":"10.1007\/978-3-031-19772-7_24"},{"key":"ref_102","doi-asserted-by":"crossref","unstructured":"Liu, Y., Liu, J., Zhao, M., Yang, D., Zhu, X., and Song, L. (2022, January 18\u201322). Learning Appearance-Motion Normality for Video Anomaly Detection. Proceedings of the 2022 IEEE International Conference on Multimedia and Expo (ICME), Taipei, Taiwan.","DOI":"10.1109\/ICME52920.2022.9859727"},{"key":"ref_103","doi-asserted-by":"crossref","first-page":"2259","DOI":"10.1109\/TCSVT.2022.3221622","article-title":"Hybrid Attention and Motion Constraint for Anomaly Detection in Crowded Scenes","volume":"33","author":"Zhang","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_104","doi-asserted-by":"crossref","first-page":"1296","DOI":"10.1109\/JAS.2021.1004045","article-title":"A cognitive memory-augmented network for visual anomaly detection","volume":"8","author":"Wang","year":"2021","journal-title":"IEEE\/CAA J. Autom. Sin."},{"key":"ref_105","doi-asserted-by":"crossref","first-page":"103232","DOI":"10.1016\/j.jvcir.2021.103232","article-title":"A comparative study between single and multi-frame anomaly detection and localization in recorded video streams","volume":"79","author":"Bahrami","year":"2021","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_106","doi-asserted-by":"crossref","unstructured":"Liu, Z., Nie, Y., Long, C., Zhang, Q., and Li, G. (2021, January 10\u201317). A hybrid video anomaly detection framework via memory-augmented flow reconstruction and flow-guided frame prediction. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01333"},{"key":"ref_107","doi-asserted-by":"crossref","unstructured":"Tang, W., Feng, Y., and Li, J. (2021, January 28\u201330). An autoencoder with a memory module for video anomaly detection. Proceedings of the 2021 36th Youth Academic Annual Conference of Chinese Association of Automation (YAC), Nanchang, China.","DOI":"10.1109\/YAC53711.2021.9486538"},{"key":"ref_108","doi-asserted-by":"crossref","first-page":"2715","DOI":"10.1007\/s10586-021-03439-5","article-title":"An explainable and efficient deep learning framework for video anomaly detection","volume":"25","author":"Wu","year":"2022","journal-title":"Clust. Comput."},{"key":"ref_109","doi-asserted-by":"crossref","first-page":"103598","DOI":"10.1016\/j.jvcir.2022.103598","article-title":"A3N: Attention-based adversarial autoencoder network for detecting anomalies in video sequence","volume":"87","author":"Aslam","year":"2022","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_110","doi-asserted-by":"crossref","first-page":"79874","DOI":"10.1109\/ACCESS.2022.3195035","article-title":"Architecture for Automatic Recognition of Group Activities Using Local Motions and Context","volume":"10","author":"Sebban","year":"2022","journal-title":"IEEE Access"},{"key":"ref_111","doi-asserted-by":"crossref","unstructured":"Huang, X., Hu, Y., Luo, X., Han, J., Zhang, B., and Cao, X. (IEEE Trans. Circuits Syst. Video Technol., 2022). Boosting Variational Inference with Margin Learning for Few-Shot Scene-Adaptive Anomaly Detection, IEEE Trans. Circuits Syst. Video Technol., early access.","DOI":"10.1109\/TCSVT.2022.3227716"},{"key":"ref_112","doi-asserted-by":"crossref","first-page":"183","DOI":"10.1016\/j.patrec.2022.03.004","article-title":"Context-related video anomaly detection via generative adversarial network","volume":"156","author":"Li","year":"2022","journal-title":"Pattern Recognit. Lett."},{"key":"ref_113","doi-asserted-by":"crossref","unstructured":"Yu, G., Wang, S., Cai, Z., Liu, X., Xu, C., and Wu, C. (2022, January 18\u201324). Deep anomaly discovery from unlabeled videos via normality advantage and self-paced refinement. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01360"},{"key":"ref_114","doi-asserted-by":"crossref","first-page":"103416","DOI":"10.1016\/j.cviu.2022.103416","article-title":"Detecting abnormality with separated foreground and background: Mutual generative adversarial networks for video abnormal event detection","volume":"219","author":"Zhang","year":"2022","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_115","doi-asserted-by":"crossref","first-page":"103739","DOI":"10.1016\/j.jvcir.2022.103739","article-title":"Anomaly detection with dual-stream memory network","volume":"90","author":"Wang","year":"2023","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_116","doi-asserted-by":"crossref","first-page":"3240","DOI":"10.1007\/s10489-022-03613-1","article-title":"Attention-based residual autoencoder for video anomaly detection","volume":"53","author":"Le","year":"2023","journal-title":"Appl. Intell."},{"key":"ref_117","doi-asserted-by":"crossref","unstructured":"Deng, H., Zhang, Z., Zou, S., and Li, X. (2023, January 2\u20137). Bi-Directional Frame Interpolation for Unsupervised Video Anomaly Detection. Proceedings of the 2023 IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV56688.2023.00266"},{"key":"ref_118","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1109\/TNN.2010.2091281","article-title":"Domain adaptation via transfer component analysis","volume":"22","author":"Pan","year":"2010","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_119","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016, January 11\u201314). Ssd: Single shot multibox detector. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_120","unstructured":"MacQueen, J. (July, January 21). Some methods for classification and analysis of multivariate observations. Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, Oakland, CA, USA."},{"key":"ref_121","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014, January 8\u201313). Generative adversarial nets. Proceedings of the Advances in Neural Information Processing Systems, Cambridge, MA, USA."},{"key":"ref_122","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_123","doi-asserted-by":"crossref","first-page":"715","DOI":"10.1162\/089976602317318938","article-title":"Slow feature analysis: Unsupervised learning of invariances","volume":"14","author":"Wiskott","year":"2002","journal-title":"Neural Comput."},{"key":"ref_124","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning spatiotemporal features with 3d convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_125","first-page":"5998","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_126","doi-asserted-by":"crossref","first-page":"107515","DOI":"10.1016\/j.patcog.2020.107515","article-title":"Fast sparse coding networks for anomaly detection in videos","volume":"107","author":"Wu","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_127","doi-asserted-by":"crossref","first-page":"203","DOI":"10.1109\/TMM.2020.2984093","article-title":"Spatial-temporal cascade autoencoder for video anomaly detection in crowded scenes","volume":"23","author":"Li","year":"2020","journal-title":"IEEE Trans. Multimed."},{"key":"ref_128","doi-asserted-by":"crossref","unstructured":"Chang, Y., Tu, Z., Xie, W., and Yuan, J. (2020, January 23\u201328). Clustering driven deep autoencoder for video anomaly detection. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Proceedings, Part XV 16.","DOI":"10.1007\/978-3-030-58555-6_20"},{"key":"ref_129","doi-asserted-by":"crossref","first-page":"108213","DOI":"10.1016\/j.patcog.2021.108213","article-title":"Video anomaly detection with spatio-temporal dissociation","volume":"122","author":"Chang","year":"2022","journal-title":"Pattern Recognit."},{"key":"ref_130","doi-asserted-by":"crossref","unstructured":"Huang, X., Zhao, C., Gao, C., Chen, L., and Wu, Z. (2023). Synthetic Pseudo Anomalies for Unsupervised Video Anomaly Detection: A Simple yet Efficient Framework based on Masked Autoencoder. arXiv.","DOI":"10.1109\/ICASSP49357.2023.10094296"},{"key":"ref_131","doi-asserted-by":"crossref","unstructured":"Sun, X., Chen, J., Shen, X., and Li, H. (2022, January 26\u201327). Transformer with Spatio-Temporal Representation for Video Anomaly Detection. Proceedings of the Structural, Syntactic, and Statistical Pattern Recognition: Joint IAPR International Workshops, S+ SSPR 2022,  Montreal, QC, Canada.","DOI":"10.1007\/978-3-031-23028-8_22"},{"key":"ref_132","doi-asserted-by":"crossref","first-page":"6163475","DOI":"10.1155\/2018\/6163475","article-title":"HuAc: Human activity recognition using crowdsourced WiFi signals and skeleton data","volume":"2018","author":"Guo","year":"2018","journal-title":"Wirel. Commun. Mob. Comput."},{"key":"ref_133","doi-asserted-by":"crossref","unstructured":"Caba Heilbron, F., Escorcia, V., Ghanem, B., and Carlos Niebles, J. (2015, January 7\u201312). Activitynet: A large-scale video benchmark for human activity understanding. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"ref_134","doi-asserted-by":"crossref","first-page":"273","DOI":"10.1016\/j.patcog.2019.04.025","article-title":"Human trajectory prediction in crowded scene using social-affinity Long Short-Term Memory","volume":"93","author":"Pei","year":"2019","journal-title":"Pattern Recognit."},{"key":"ref_135","doi-asserted-by":"crossref","first-page":"20877","DOI":"10.1007\/s11042-019-7392-z","article-title":"Highly refined human action recognition model to handle intraclass variability & interclass similarity","volume":"78","author":"Akila","year":"2019","journal-title":"Multimed. Tools Appl."},{"key":"ref_136","doi-asserted-by":"crossref","first-page":"1061","DOI":"10.1002\/tee.22901","article-title":"Group activity recognition with an interaction force based on low-level features","volume":"14","author":"Wateosot","year":"2019","journal-title":"IEEJ Trans. Electr. Electron. Eng."},{"key":"ref_137","doi-asserted-by":"crossref","first-page":"295","DOI":"10.1016\/j.patcog.2016.08.003","article-title":"Robust human activity recognition from depth video using spatiotemporal multi-fused features","volume":"61","author":"Jalal","year":"2017","journal-title":"Pattern Recognit."},{"key":"ref_138","unstructured":"Carreira, J., Noland, E., Hillier, C., and Zisserman, A. (2019). A short note on the kinetics-700 human action dataset. arXiv."},{"key":"ref_139","doi-asserted-by":"crossref","first-page":"9893","DOI":"10.1109\/ACCESS.2018.2890675","article-title":"InnoHAR: A deep neural network for complex human activity recognition","volume":"7","author":"Xu","year":"2019","journal-title":"IEEE Access"},{"key":"ref_140","doi-asserted-by":"crossref","first-page":"565","DOI":"10.1016\/j.eswa.2018.08.041","article-title":"Seeded transfer learning for regression problems with deep learning","volume":"115","author":"Salaken","year":"2019","journal-title":"Expert Syst. Appl."},{"key":"ref_141","doi-asserted-by":"crossref","unstructured":"Ord\u00f3\u00f1ez, F.J., and Roggen, D. (2016). Deep convolutional and lstm recurrent neural networks for multimodal wearable activity recognition. Sensors, 16.","DOI":"10.3390\/s16010115"},{"key":"ref_142","doi-asserted-by":"crossref","first-page":"983","DOI":"10.1016\/j.cma.2019.01.011","article-title":"NURBS-based postbuckling analysis of functionally graded carbon nanotube-reinforced composite shells","volume":"347","author":"Nguyen","year":"2019","journal-title":"Comput. Methods Appl. Mech. Eng."},{"key":"ref_143","doi-asserted-by":"crossref","unstructured":"Goyal, R., Kahou, S.E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., and Mueller-Freitag, M. (2017, January 22\u201329). The \u201cSomething Something\u201d Video Database for Learning and Evaluating Visual Common Sense. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.622"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/11\/5024\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:40:53Z","timestamp":1760125253000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/11\/5024"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,24]]},"references-count":143,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2023,6]]}},"alternative-id":["s23115024"],"URL":"https:\/\/doi.org\/10.3390\/s23115024","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,24]]}}}