{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T06:57:23Z","timestamp":1781852243929,"version":"3.54.5"},"reference-count":41,"publisher":"MDPI AG","issue":"17","license":[{"start":{"date-parts":[[2020,9,1]],"date-time":"2020-09-01T00:00:00Z","timestamp":1598918400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100010680","name":"H2020 Transport","doi-asserted-by":"publisher","award":["769033"],"award-info":[{"award-number":["769033"]}],"id":[{"id":"10.13039\/100010680","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Autonomous vehicles (AVs) are already operating on the streets of many countries around the globe. Contemporary concerns about AVs do not relate to the implementation of fundamental technologies, as they are already in use, but are rather increasingly centered on the way that such technologies will affect emerging transportation systems, our social environment, and the people living inside it. Many concerns also focus on whether such systems should be fully automated or still be partially controlled by humans. This work aims to address the new reality that is formed in autonomous shuttles mobility infrastructures as a result of the absence of the bus driver and the increased threat from terrorism in European cities. Typically, drivers are trained to handle incidents of passengers\u2019 abnormal behavior, incidents of petty crimes, and other abnormal events, according to standard procedures adopted by the transport operator. Surveillance using camera sensors as well as smart software in the bus will maximize the feeling and the actual level of security. In this paper, an online, end-to-end solution is introduced based on deep learning techniques for the timely, accurate, robust, and automatic detection of various petty crime types. The proposed system can identify abnormal passenger behavior such as vandalism and accidents but can also enhance passenger security via petty crimes detection such as aggression, bag-snatching, and vandalism. The solution achieves excellent results across different use cases and environmental conditions.<\/jats:p>","DOI":"10.3390\/s20174943","type":"journal-article","created":{"date-parts":[[2020,9,1]],"date-time":"2020-09-01T08:53:43Z","timestamp":1598950423000},"page":"4943","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":29,"title":["Real-Time Abnormal Event Detection for Enhanced Security in Autonomous Shuttles Mobility Infrastructures"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6475-5865","authenticated-orcid":false,"given":"Dimitris","family":"Tsiktsiris","sequence":"first","affiliation":[{"name":"Information Technologies Institute, Centre for Research and Technology Hellas, 6th km Charilaou-Thermi, 57001 Thermi, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6650-7758","authenticated-orcid":false,"given":"Nikolaos","family":"Dimitriou","sequence":"additional","affiliation":[{"name":"Information Technologies Institute, Centre for Research and Technology Hellas, 6th km Charilaou-Thermi, 57001 Thermi, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5337-161X","authenticated-orcid":false,"given":"Antonios","family":"Lalas","sequence":"additional","affiliation":[{"name":"Information Technologies Institute, Centre for Research and Technology Hellas, 6th km Charilaou-Thermi, 57001 Thermi, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2180-9752","authenticated-orcid":false,"given":"Minas","family":"Dasygenis","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, University of Western Macedonia, 50100 Kozani, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6381-8326","authenticated-orcid":false,"given":"Konstantinos","family":"Votis","sequence":"additional","affiliation":[{"name":"Information Technologies Institute, Centre for Research and Technology Hellas, 6th km Charilaou-Thermi, 57001 Thermi, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6915-6722","authenticated-orcid":false,"given":"Dimitrios","family":"Tzovaras","sequence":"additional","affiliation":[{"name":"Information Technologies Institute, Centre for Research and Technology Hellas, 6th km Charilaou-Thermi, 57001 Thermi, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2020,9,1]]},"reference":[{"key":"ref_1","unstructured":"Simonyan, K., and Zisserman, A. (2014, January 8\u201313). Two-stream convolutional networks for action recognition in videos. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., and Van Gool, L. (2016). Temporal Segment Networks: Towards Good Practices for Deep Action Recognition. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning spatiotemporal features with 3d convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"118","DOI":"10.1016\/j.cviu.2018.04.007","article-title":"RGB-D-based human motion recognition with deep learning: A survey","volume":"171","author":"Wang","year":"2018","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"201","DOI":"10.3758\/BF03212378","article-title":"Visual perception of biological motion and a model for its analysis","volume":"14","author":"Johansson","year":"1973","journal-title":"Percept. Psychophys."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1109\/MMUL.2012.24","article-title":"Microsoft kinect sensor and its effect","volume":"19","author":"Zhang","year":"2012","journal-title":"IEEE Multimed."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Cao, Z., Simon, T., Wei, S.E., and Sheikh, Y. (2017, January 21\u201326). Realtime multi-person 2d pose estimation using part affinity fields. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.143"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Song, S., Lan, C., Xing, J., Zeng, W., and Liu, J. (2017, January 4\u20139). An end-to-end spatio-temporal attention model for human action recognition from skeleton data. Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.11212"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Du, Y., Fu, Y., and Wang, L. (2015, January 3\u20136). Skeleton based action recognition with convolutional neural network. Proceedings of the 2015 IEEE 3rd IAPR Asian Conference on Pattern Recognition (ACPR), Kuala Lumpur, Malaysia.","DOI":"10.1109\/ACPR.2015.7486569"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Li, C., Zhong, Q., Xie, D., and Pu, S. (2018, January 13\u201319). Co-occurrence feature learning from skeleton data for action recognition and detection with hierarchical aggregation. Proceedings of the 27th International Joint Conference on Artificial Intelligence, Stockholm, Sweden.","DOI":"10.24963\/ijcai.2018\/109"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Ke, Q., Bennamoun, M., An, S., Sohel, F., and Boussaid, F. (2017, January 21\u201326). A new representation of skeleton sequences for 3D action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.486"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., and Lin, D. (2018, January 2\u20137). Spatial temporal graph convolutional networks for skeleton-based action recognition. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Mahadevan, V., Li, W., Bhalodia, V., and Vasconcelos, N. (2010, January 13\u201318). Anomaly detection in crowded scenes. Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539872"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Avil\u00e9s-Cruz, C., Ferreyra-Ram\u00edrez, A., Z\u00fa\u00f1iga-L\u00f3pez, A., and Villegas-Cort\u00e9z, J. (2019). Coarse-fine convolutional deep-learning strategy for human activity recognition. Sensors, 19.","DOI":"10.3390\/s19071556"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Ord\u00f3\u00f1ez, F.J., and Roggen, D. (2016). Deep convolutional and lstm recurrent neural networks for multimodal wearable activity recognition. Sensors, 16.","DOI":"10.3390\/s16010115"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"88","DOI":"10.1016\/j.cviu.2018.02.006","article-title":"Deep-anomaly: Fully convolutional neural network for fast anomaly detection in crowded scenes","volume":"172","author":"Sabokrou","year":"2018","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"548","DOI":"10.1109\/TCYB.2014.2330853","article-title":"Online anomaly detection in crowd scenes via structure analysis","volume":"45","author":"Yuan","year":"2014","journal-title":"IEEE Trans. Cybern."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"466","DOI":"10.1016\/j.neunet.2018.09.002","article-title":"Soft+ hardwired attention: An LSTM framework for human trajectory prediction and abnormal event detection","volume":"108","author":"Fernando","year":"2018","journal-title":"Neural Netw."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Ravanbakhsh, M., Nabi, M., Mousavi, H., Sangineto, E., and Sebe, N. (2018, January 12\u201315). Plug-and-play cnn for crowd motion analysis: An application in abnormal event detection. Proceedings of the 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe, NV, USA.","DOI":"10.1109\/WACV.2018.00188"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wei, H., Jafari, R., and Kehtarnavaz, N. (2019). Fusion of Video and Inertial Sensing for Deep Learning\u2013Based Human Action Recognition. Sensors, 19.","DOI":"10.3390\/s19173680"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Naqvi, R.A., Arsalan, M., Rehman, A., Rehman, A.U., Loh, W.K., and Paul, A. (2020). Deep Learning-Based Drivers Emotion Classification System in Time Series Data for Remote Applications. Remote Sens., 12.","DOI":"10.3390\/rs12030587"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"6701","DOI":"10.1109\/JSEN.2020.2975382","article-title":"Cloud-Based Driver Monitoring System Using a Smartphone","volume":"20","author":"Kashevnik","year":"2020","journal-title":"IEEE Sens. J."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Khan, M.Q., and Lee, S. (2019). Gaze and Eye Tracking: Techniques and Applications in ADAS. Sensors, 19.","DOI":"10.3390\/s19245540"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Giakoumis, D., Drosou, A., Cipresso, P., Tzovaras, D., Hassapis, G., Gaggioli, A., and Riva, G. (2012). Using activity-related behavioural features towards more effective automatic stress detection. PLoS ONE, 7.","DOI":"10.1371\/journal.pone.0043571"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Dimitriou, N., Kioumourtzis, G., Sideris, A., Stavropoulos, G., Taka, E., Zotos, N., Leventakis, G., and Tzovaras, D. (2017, January 11\u201313). An Integrated Framework for the Timely Detection of Petty Crimes. Proceedings of the 2017 IEEE European Intelligence and Security Informatics Conference (EISIC), Athens, Greece.","DOI":"10.1109\/EISIC.2017.13"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Shahroudy, A., Liu, J., Ng, T.T., and Wang, G. (2016, January 27\u201330). Ntu rgb+ d: A large scale dataset for 3d human activity analysis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.115"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Fang, H.S., Xie, S., Tai, Y.W., and Lu, C. (2017, January 22\u201329). RMPE: Regional Multi-person Pose Estimation. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.256"},{"key":"ref_28","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Hasan, M., Choi, J., Neumann, J., Roy-Chowdhury, A.K., and Davis, L.S. (2016, January 27\u201330). Learning temporal regularity in video sequences. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.86"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Farneb\u00e4ck, G. (2003). Two-frame motion estimation based on polynomial expansion. Scandinavian Conference on Image Analysis, Springer.","DOI":"10.1007\/3-540-45103-X_50"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"El Baf, F., Bouwmans, T., and Vachon, B. (2008). Type-2 fuzzy mixture of Gaussians model: Application to background modeling. International Symposium on Visual Computing, Springer.","DOI":"10.1007\/978-3-540-89639-5_74"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zhang, H.B., Zhang, Y.X., Zhong, B., Lei, Q., Yang, L., Du, J.X., and Chen, D.S. (2019). A comprehensive survey of vision-based human action recognition methods. Sensors, 19.","DOI":"10.3390\/s19051005"},{"key":"ref_33","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (2020). Decoupled Spatial-Temporal Attention Network for Skeleton-Based Action Recognition. arXiv."},{"key":"ref_34","unstructured":"Yang, D., Li, M.M., Fu, H., Fan, J., and Leung, H. (2020). Centrality Graph Convolutional Networks for Skeleton-based Action Recognition. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"3459","DOI":"10.1109\/TIP.2018.2818328","article-title":"Spatio-temporal attention-based LSTM networks for 3D action recognition and detection","volume":"27","author":"Song","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Liu, J., Shahroudy, A., Xu, D., and Wang, G. (2016). Spatio-temporal lstm with trust gates for 3d human action recognition. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46487-9_50"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Yang, X., and Tian, Y. (2014, January 23\u201328). Super normal vector for activity recognition using depth sequences. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.108"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Chong, Y.S., and Tay, Y.H. (2017). Abnormal event detection in videos using spatiotemporal autoencoder. International Symposium on Neural Networks, Springer.","DOI":"10.1007\/978-3-319-59081-3_23"},{"key":"ref_39","unstructured":"Wang, T., and Snoussi, H. (2013, January 15\u201317). Histograms of optical flow orientation for abnormal events detection. Proceedings of the 2013 IEEE International Workshop on Performance Evaluation of Tracking and Surveillance (PETS), Clearwater, FL, USA."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"555","DOI":"10.1109\/TPAMI.2007.70825","article-title":"Robust real-time unusual event detection using multiple fixed-location monitors","volume":"30","author":"Adam","year":"2008","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Mehran, R., Oyama, A., and Shah, M. (2009, January 20\u201325). Abnormal crowd behavior detection using social force model. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206641"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/17\/4943\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:05:24Z","timestamp":1760177124000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/17\/4943"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,1]]},"references-count":41,"journal-issue":{"issue":"17","published-online":{"date-parts":[[2020,9]]}},"alternative-id":["s20174943"],"URL":"https:\/\/doi.org\/10.3390\/s20174943","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,9,1]]}}}