{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T03:23:34Z","timestamp":1784258614181,"version":"3.55.0"},"reference-count":37,"publisher":"MDPI AG","issue":"21","license":[{"start":{"date-parts":[[2022,10,31]],"date-time":"2022-10-31T00:00:00Z","timestamp":1667174400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"European Union funds awarded to Blees Sp. z o.o.","award":["POIR.01.01.01-00-0952\/20-00"],"award-info":[{"award-number":["POIR.01.01.01-00-0952\/20-00"]}]},{"name":"European Union funds awarded to Blees Sp. z o.o.","award":["GA 951911"],"award-info":[{"award-number":["GA 951911"]}]},{"name":"European Union funds awarded to Blees Sp. z o.o.","award":["POR FSE CUP B53D21008060008"],"award-info":[{"award-number":["POR FSE CUP B53D21008060008"]}]},{"name":"the EC H2020 project \u201cAI4media","award":["POIR.01.01.01-00-0952\/20-00"],"award-info":[{"award-number":["POIR.01.01.01-00-0952\/20-00"]}]},{"name":"the EC H2020 project \u201cAI4media","award":["GA 951911"],"award-info":[{"award-number":["GA 951911"]}]},{"name":"the EC H2020 project \u201cAI4media","award":["POR FSE CUP B53D21008060008"],"award-info":[{"award-number":["POR FSE CUP B53D21008060008"]}]},{"name":"research project (RAU-6, 2020) and projects for young scientists of the Silesian University of Technology","award":["POIR.01.01.01-00-0952\/20-00"],"award-info":[{"award-number":["POIR.01.01.01-00-0952\/20-00"]}]},{"name":"research project (RAU-6, 2020) and projects for young scientists of the Silesian University of Technology","award":["GA 951911"],"award-info":[{"award-number":["GA 951911"]}]},{"name":"research project (RAU-6, 2020) and projects for young scientists of the Silesian University of Technology","award":["POR FSE CUP B53D21008060008"],"award-info":[{"award-number":["POR FSE CUP B53D21008060008"]}]},{"name":"research project INAROS (INtelligenza ARtificiale per il mOnitoraggio e Supporto agli anziani), Tuscany","award":["POIR.01.01.01-00-0952\/20-00"],"award-info":[{"award-number":["POIR.01.01.01-00-0952\/20-00"]}]},{"name":"research project INAROS (INtelligenza ARtificiale per il mOnitoraggio e Supporto agli anziani), Tuscany","award":["GA 951911"],"award-info":[{"award-number":["GA 951911"]}]},{"name":"research project INAROS (INtelligenza ARtificiale per il mOnitoraggio e Supporto agli anziani), Tuscany","award":["POR FSE CUP B53D21008060008"],"award-info":[{"award-number":["POR FSE CUP B53D21008060008"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The automatic detection of violent actions in public places through video analysis is difficult because the employed Artificial Intelligence-based techniques often suffer from generalization problems. Indeed, these algorithms hinge on large quantities of annotated data and usually experience a drastic drop in performance when used in scenarios never seen during the supervised learning phase. In this paper, we introduce and publicly release the Bus Violence benchmark, the first large-scale collection of video clips for violence detection on public transport, where some actors simulated violent actions inside a moving bus in changing conditions, such as the background or light. Moreover, we conduct a performance analysis of several state-of-the-art video violence detectors pre-trained with general violence detection databases on this newly established use case. The achieved moderate performances reveal the difficulties in generalizing from these popular methods, indicating the need to have this new collection of labeled data, beneficial for specializing them in this new scenario.<\/jats:p>","DOI":"10.3390\/s22218345","type":"journal-article","created":{"date-parts":[[2022,11,2]],"date-time":"2022-11-02T06:49:02Z","timestamp":1667371742000},"page":"8345","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":24,"title":["Bus Violence: An Open Benchmark for Video Violence Detection on Public Transport"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6985-0439","authenticated-orcid":false,"given":"Luca","family":"Ciampi","sequence":"first","affiliation":[{"name":"Institute of Information Science and Technologies, National Research Council, Via G. Moruzzi 1, 56124 Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5491-9096","authenticated-orcid":false,"given":"Pawe\u0142","family":"Foszner","sequence":"additional","affiliation":[{"name":"Department of Computer Graphics, Vision and Digital Systems, Faculty of Automatic Control, Electronics and Computer Science, Silesian University of Technology, Akademicka 2A, 44-100 Gliwice, Poland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3011-2487","authenticated-orcid":false,"given":"Nicola","family":"Messina","sequence":"additional","affiliation":[{"name":"Institute of Information Science and Technologies, National Research Council, Via G. Moruzzi 1, 56124 Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9659-7451","authenticated-orcid":false,"given":"Micha\u0142","family":"Staniszewski","sequence":"additional","affiliation":[{"name":"Department of Computer Graphics, Vision and Digital Systems, Faculty of Automatic Control, Electronics and Computer Science, Silesian University of Technology, Akademicka 2A, 44-100 Gliwice, Poland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3715-149X","authenticated-orcid":false,"given":"Claudio","family":"Gennaro","sequence":"additional","affiliation":[{"name":"Institute of Information Science and Technologies, National Research Council, Via G. Moruzzi 1, 56124 Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6258-5313","authenticated-orcid":false,"given":"Fabrizio","family":"Falchi","sequence":"additional","affiliation":[{"name":"Institute of Information Science and Technologies, National Research Council, Via G. Moruzzi 1, 56124 Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gianluca","family":"Serao","sequence":"additional","affiliation":[{"name":"Department of Information Engineering, University of Pisa, Via Girolamo Caruso, 16, 56122 Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Micha\u0142","family":"Cogiel","sequence":"additional","affiliation":[{"name":"Blees sp. z o.o., Zygmunta Starego 24a\/10, 44-100 Gliwice, Poland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dominik","family":"Golba","sequence":"additional","affiliation":[{"name":"Blees sp. z o.o., Zygmunta Starego 24a\/10, 44-100 Gliwice, Poland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4354-8258","authenticated-orcid":false,"given":"Agnieszka","family":"Szcz\u0119sna","sequence":"additional","affiliation":[{"name":"Department of Computer Graphics, Vision and Digital Systems, Faculty of Automatic Control, Electronics and Computer Science, Silesian University of Technology, Akademicka 2A, 44-100 Gliwice, Poland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0171-4315","authenticated-orcid":false,"given":"Giuseppe","family":"Amato","sequence":"additional","affiliation":[{"name":"Institute of Information Science and Technologies, National Research Council, Via G. Moruzzi 1, 56124 Pisa, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,10,31]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"117125","DOI":"10.1016\/j.eswa.2022.117125","article-title":"An embedded toolset for human activity monitoring in critical environments","volume":"199","author":"Benedetto","year":"2022","journal-title":"Expert Syst. Appl."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Avvenuti, M., Bongiovanni, M., Ciampi, L., Falchi, F., Gennaro, C., and Messina, N. (July, January 30). A Spatio-Temporal Attentive Network for Video-Based Crowd Counting. Proceedings of the 2022 IEEE Symposium on Computers and Communications (ISCC), Rhodes, Greece.","DOI":"10.1109\/ISCC55528.2022.9913019"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Staniszewski, M., Kloszczyk, M., Segen, J., Wereszczy\u0144ski, K., Drabik, A., and Kulbacki, M. (2016). Recent Developments in Tracking Objects in a Video Sequence. Intelligent Information and Database Systems, Springer.","DOI":"10.1007\/978-3-662-49390-8_42"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Staniszewski, M., Foszner, P., Kostorz, K., Michalczuk, A., Wereszczy\u0144ski, K., Cogiel, M., Golba, D., Wojciechowski, K., and Pola\u0144ski, A. (2020). Application of Crowd Simulations in the Evaluation of Tracking Algorithms. Sensors, 20.","DOI":"10.3390\/s20174960"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"302","DOI":"10.1007\/978-3-030-30642-7_27","article-title":"Learning Pedestrian Detection from Virtual Worlds. Image Analysis and Processing-ICIAP 2019-20th International Conference, Trento, Italy, 9\u201313 September 2019, Proceedings, Part I. Springer","volume":"11751","author":"Amato","year":"2019","journal-title":"Lect. Notes Comput. Sci."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Ciampi, L., Messina, N., Falchi, F., Gennaro, C., and Amato, G. (2020). Virtual to Real Adaptation of Pedestrian Detectors. Sensors, 20.","DOI":"10.3390\/s20185250"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Ma, Y., Bai, T., Zhang, W., Li, S., Hu, J., and Lu, M. (2021, January 5\u20138). Multi-Scale Relation Network for Person Re-identification. Proceedings of the 2021 IEEE Symposium on Computers and Communications (ISCC), Athens, Greece.","DOI":"10.1109\/ISCC53001.2021.9631515"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"P\u0119szor, D., Staniszewski, M., and Wojciechowska, M. (2016). Facial Reconstruction on the Basis of Video Surveillance System for the Purpose of Suspect Identification. Intelligent Information and Database Systems, Springer.","DOI":"10.1007\/978-3-662-49390-8_46"},{"key":"ref_9","unstructured":"Foszner, P., Staniszewski, M., Szcz\u0119sna, A., Cogiel, M., Golba, D., Ciampi, L., Messina, N., Gennaro, C., Falchi, F., and Amato, G. (2022). Bus Violence: A large-scale benchmark for video violence detection in public transport. Zenodo."},{"key":"ref_10","unstructured":"Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., and Natsev, P. (2017). The Kinetics Human Action Video Dataset. arXiv."},{"key":"ref_11","unstructured":"Carreira, J., Noland, E., Banki-Horvath, A., Hillier, C., and Zisserman, A. (2018). A Short Note about Kinetics-600. arXiv."},{"key":"ref_12","unstructured":"Smaira, L., Carreira, J., Noland, E., Clancy, E., Wu, A., and Zisserman, A. (2020). A Short Note on the Kinetics-700-2020 Human Action Dataset. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., and Serre, T. (2011, January 6\u201313). HMDB: A large video database for human motion recognition. Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"ref_14","unstructured":"Soomro, K., Zamir, A.R., and Shah, M. (2012). UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Sultani, W., Chen, C., and Shah, M. (2018, January 18\u201323). Real-World Anomaly Detection in Surveillance Videos. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00678"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"533","DOI":"10.22214\/ijraset.2020.30788","article-title":"Violence Detection in Surveillance Video using Computer Vision Techniques","volume":"8","author":"Padamwar","year":"2020","journal-title":"Int. J. Res. Appl. Sci. Eng. Technol."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Hassner, T., Itcher, Y., and Kliper-Gross, O. (2012, January 16\u201321). Violent flows: Real-time detection of violent crowd behavior. Proceedings of the 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, Providence, RI, USA.","DOI":"10.1109\/CVPRW.2012.6239348"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Perez, M., Kot, A.C., and Rocha, A. (2019, January 12\u201317). Detection of Real-world Fights in Surveillance Videos. Proceedings of the 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683676"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"106587","DOI":"10.1016\/j.dib.2020.106587","article-title":"A dataset for automatic violence detection in videos","volume":"33","author":"Bianculli","year":"2020","journal-title":"Data Brief"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"160580","DOI":"10.1109\/ACCESS.2021.3131315","article-title":"Deep Learning for Automatic Violence Detection: Tests on the AIRTLab Dataset","volume":"9","author":"Sernani","year":"2021","journal-title":"IEEE Access"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Akti, S., Tataroglu, G.A., and Ekenel, H.K. (2019, January 6\u20139). Vision-based Fight Detection from Surveillance Cameras. Proceedings of the 2019 IEEE Ninth International Conference on Image Processing Theory, Tools and Applications (IPTA), Istanbul, Turkey.","DOI":"10.1109\/IPTA.2019.8936070"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Cheng, M., Cai, K., and Li, M. (2021, January 10\u201315). RWF-2000: An Open Large Scale Video Database for Violence Detection. Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy.","DOI":"10.1109\/ICPR48806.2021.9412502"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Soliman, M.M., Kamal, M.H., Nashed, M.A.E.M., Mostafa, Y.M., Chawky, B.S., and Khattab, D. (2019, January 8\u201310). Violence Recognition from Videos using Deep Learning Techniques. Proceedings of the 2019 Ninth International Conference on Intelligent Computing and Information Systems (ICICIS), Cairo, Egypt.","DOI":"10.1109\/ICICIS46948.2019.9014714"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., and Paluri, M. (2018, January 18\u201323). A Closer Look at Spatiotemporal Convolutions for Action Recognition. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00675"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning Spatiotemporal Features with 3D Convolutional Networks. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_26","unstructured":"Feichtenhofer, C., Pinz, A., and Wildes, R.P. (2016, January 5\u201310). Spatiotemporal Residual Networks for Video Action Recognition. Proceedings of the Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, Barcelona, Spain."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., and He, K. (November, January 27). SlowFast Networks for Video Recognition. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00630"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., and Hu, H. (2022, January 18\u201324). Video swin transformer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Sudhakaran, S., and Lanz, O. (September, January 29). Learning to detect violent videos using convolutional long short-term memory. Proceedings of the 2017 14th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), Lecce, Italy.","DOI":"10.1109\/AVSS.2017.8078468"},{"key":"ref_31","unstructured":"Shi, X., Chen, Z., Wang, H., Yeung, D., Wong, W., and Woo, W. (2015, January 7\u201312). Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting. Proceedings of the Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, Montreal, QC, Canada."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Hanson, A., PNVR, K., Krishnagopal, S., and Davis, L. (2018, January 8\u201314). Bidirectional Convolutional LSTM for the Detection of Violence in Videos. Proceedings of the Computer Vision\u2014ECCV 2018 Workshops, Munich, Germany.","DOI":"10.1007\/978-3-030-11012-3_24"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Islam, Z., Rukonuzzaman, M., Ahmed, R., Kabir, M.H., and Farazi, M. (2021, January 18\u201322). Efficient Two-Stream Network for Violence Detection Using Separable Convolutional LSTM. Proceedings of the 2021 International Joint Conference on Neural Networks (IJCNN), Shenzhen, China.","DOI":"10.1109\/IJCNN52387.2021.9534280"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Ciampi, L., Santiago, C., Costeira, J., Gennaro, C., and Amato, G. (2021, January 8\u201310). Domain Adaptation for Traffic Density Estimation. Proceedings of the 16th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, Vienna, Austria.","DOI":"10.5220\/0010303401850195"},{"key":"ref_35","unstructured":"Ciampi, L., Santiago, C., Costeira, J.P., Gennaro, C., and Amato, G. (2020, January 4). Unsupervised vehicle counting via multiple camera domain adaptation. Proceedings of the First International Workshop on New Foundations for Human-Centered AI (NeHuAI) co-located with 24th European Conference on Artificial Intelligence (ECAI 2020), Santiago de Compostella, Spain."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Caron, M., Touvron, H., Misra, I., J\u00e9gou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021, January 10\u201317). Emerging properties in self-supervised vision transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"ref_37","unstructured":"Li, C., Yang, J., Zhang, P., Gao, M., Xiao, B., Dai, X., Yuan, L., and Gao, J. (2021, January 4). Efficient Self-supervised Vision Transformers for Representation Learning. Proceedings of the International Conference on Learning Representations, Vienna, Austria."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/21\/8345\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:06:25Z","timestamp":1760144785000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/21\/8345"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,31]]},"references-count":37,"journal-issue":{"issue":"21","published-online":{"date-parts":[[2022,11]]}},"alternative-id":["s22218345"],"URL":"https:\/\/doi.org\/10.3390\/s22218345","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,10,31]]}}}