{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,23]],"date-time":"2026-01-23T07:39:54Z","timestamp":1769153994992,"version":"3.49.0"},"reference-count":38,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2023,12,21]],"date-time":"2023-12-21T00:00:00Z","timestamp":1703116800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Korea Institute of Civil Engineering and Building Technology","award":["20230143001"],"award-info":[{"award-number":["20230143001"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Weakly supervised video anomaly detection is a methodology that assesses anomaly levels in individual frames based on labeled video data. Anomaly scores are computed by evaluating the deviation of distances derived from frames in an unbiased state. Weakly supervised video anomaly detection encounters the formidable challenge of false alarms, stemming from various sources, with a major contributor being the inadequate reflection of frame labels during the learning process. Multiple instance learning has been a pivotal solution to this issue in previous studies, necessitating the identification of discernible features between abnormal and normal segments. Simultaneously, it is imperative to identify shared biases within the feature space and cultivate a representative model. In this study, we introduce a novel multiple instance learning framework anchored on a memory unit, which augments features based on memory and effectively bridges the gap between normal and abnormal instances. This augmentation is facilitated through the integration of an multi-head attention feature augmentation module and loss function with a KL divergence and a Gaussian distribution estimation-based approach. The method identifies distinguishable features and secures the inter-instance distance, thus fortifying the distance metrics between abnormal and normal instances approximated by distribution. The contribution of this research involves proposing a novel framework based on MIL for performing WSVAD and presenting an efficient integration strategy during the augmentation process. Extensive experiments were conducted on benchmark datasets XD-Violence and UCF-Crime to substantiate the effectiveness of the proposed model.<\/jats:p>","DOI":"10.3390\/s24010058","type":"journal-article","created":{"date-parts":[[2023,12,21]],"date-time":"2023-12-21T10:53:52Z","timestamp":1703156032000},"page":"58","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Cognitive Refined Augmentation for Video Anomaly Detection in Weak Supervision"],"prefix":"10.3390","volume":"24","author":[{"given":"Junyeop","family":"Lee","sequence":"first","affiliation":[{"name":"School of Electrical Engineering, Korea University, Seoul 02841, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hyunbon","family":"Koo","sequence":"additional","affiliation":[{"name":"Korea Institute of Civil Engineering and Building Technology, Goyang-si 10223, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Seongjun","family":"Kim","sequence":"additional","affiliation":[{"name":"Korea Institute of Civil Engineering and Building Technology, Goyang-si 10223, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8744-4514","authenticated-orcid":false,"given":"Hanseok","family":"Ko","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering, Korea University, Seoul 02841, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,12,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"969","DOI":"10.1109\/LGRS.2017.2691741","article-title":"Compact HF surface wave radar data generating simulator for ship detection and tracking","volume":"14","author":"Park","year":"2017","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_2","unstructured":"Kang, B.H., Jeon, C.W., and Ko, H.S. (2012). Pre-Processing Method and Apparatus for Wide Dynamic Range Image Processing. (8,135,235), U.S. Patent."},{"key":"ref_3","unstructured":"Byun, S., Choi, D., Ahn, B., and Ko, H. (1999, January 8\u201312). Traffic incident detection using evidential reasoning based data fusion. Proceedings of the World Congress on Intelligent Transport Systems (ITS), Toronto, ON, Canada."},{"key":"ref_4","unstructured":"Seo, J., and Ko, H. (2004, January 17\u201321). Face detection using support vector domain description in color images. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, Montreal, QC, Canada."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Lee, I., Ko, H., and Han, D.K. (2002, January 13\u201317). Multiple vehicle tracking based on regional estimation in nighttime CCD images. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, Orlando, FL, USA.","DOI":"10.1109\/ICASSP.2002.5745462"},{"key":"ref_6","unstructured":"Kim, K., and Ko, H. (30\u20132, January 30). Hierarchical approach for abnormal acoustic event classification in an elevator. Proceedings of the IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), Klagenfurt, Austria."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"423","DOI":"10.1049\/ip-vis:20000408","article-title":"Spectral subtraction based on phonetic dependency and masking effects","volume":"147","author":"Kim","year":"2000","journal-title":"IEEE Proc. Vis. Image Signal Process."},{"key":"ref_8","first-page":"88","article-title":"ConvGRU-CNN: Spatiotemporal Deep Learning for Real-World Anomaly Detection in Video Surveillance System","volume":"8","year":"2023","journal-title":"Int. J. Interact. Multimed. Artif. Intell."},{"key":"ref_9","first-page":"14","article-title":"Design of integrated artificial intelligence techniques for video surveillance on iot enabled wireless multimedia sensor networks","volume":"7","author":"Mansour","year":"2022","journal-title":"Int. J. Interact. Multimed. Artif. Intell."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Cai, R., Zhang, H., Liu, W., Gao, S., and Hao, Z. (2021, January 2\u20139). Appearance-motion memory consistency network for video anomaly detection. Proceedings of the AAAI Conference on Artificial Intelligence, Virtually.","DOI":"10.1609\/aaai.v35i2.16177"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Park, H., Noh, J., and Ham, B. (2020, January 14\u201319). Learning memory-guided normality for anomaly detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01438"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"2395","DOI":"10.1109\/TIP.2019.2948286","article-title":"BMAN: Bidirectional multi-scale aggregation networks for abnormal event detection","volume":"29","author":"Lee","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"2301","DOI":"10.1109\/TNNLS.2021.3083152","article-title":"Robust unsupervised video anomaly detection by multipath frame prediction","volume":"33","author":"Wang","year":"2021","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_14","unstructured":"Gong, D., Liu, L., Le, V., Saha, B., Mansour, M.R., Venkatesh, S., and Hengel, A.V.D. (November, January 27). Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zaheer, M.Z., Mahmood, A., Khan, M.H., Segu, M., Yu, F., and Lee, S.I. (2022, January 18\u201324). Generative cooperative learning for unsupervised video anomaly detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01433"},{"key":"ref_16","unstructured":"Kim, J.H., Kim, D.H., Yi, S., and Lee, T. (2021). Semi-orthogonal embedding for efficient unsupervised anomaly segmentation. arXiv."},{"key":"ref_17","unstructured":"Wang, J., and Cherian, A. (November, January 27). Gods: Generalized one-class discriminative subspaces for anomaly detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Sultani, W., Chen, C., and Shah, M. (2018, January 18\u201323). Real-world anomaly detection in surveillance videos. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00678"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhang, J., Qing, L., and Miao, J. (2019, January 22\u201325). Temporal convolutional network with complementary inner bag loss for weakly supervised anomaly detection. Proceedings of the 2019 IEEE International Conference on Image Processing (ICIP), Taipei, Taiwan.","DOI":"10.1109\/ICIP.2019.8803657"},{"key":"ref_20","unstructured":"Zaheer, M.Z., Lee, J.h., Astrid, M., Mahmood, A., and Lee, S.I. (2021). Cleaning label noise with clusters for minimally supervised anomaly detection. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhong, J.X., Li, N., Kong, W., Liu, S., Li, T.H., and Li, G. (2019, January 15\u201320). Graph convolutional label noise cleaner: Train a plug-and-play action classifier for anomaly detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00133"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Tian, Y., Pang, G., Chen, Y., Singh, R., Verjans, J.W., and Carneiro, G. (2021, January 11\u201317). Weakly-supervised video anomaly detection with robust temporal feature magnitude learning. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00493"},{"key":"ref_23","unstructured":"Li, S., Liu, F., and Jiao, L. (March, January 22). Self-training multi-sequence learning with transformer for weakly supervised video anomaly detection. Proceedings of the AAAI Conference on Artificial Intelligence, Virtually."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Wu, P., Liu, J., Shi, Y., Sun, Y., Shao, F., Wu, Z., and Yang, Z. (2020, January 23\u201328). Not only look, but also listen: Learning multimodal violence detection under weak supervision. Proceedings of the 16th European Conference on Computer Vision\u2014ECCV 2020, Glasgow, UK.","DOI":"10.1007\/978-3-030-58577-8_20"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"104397","DOI":"10.1016\/j.imavis.2022.104397","article-title":"Batch feature standardization network with triplet loss for weakly-supervised video anomaly detection","volume":"120","author":"Yi","year":"2022","journal-title":"Image Vis. Comput."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zaheer, M.Z., Mahmood, A., Astrid, M., and Lee, S.I. (2020, January 23\u201328). Claws: Clustering assisted weakly supervised learning with normalcy suppression for anomalous event detection. Proceedings of the 16th European Conference on Computer Vision\u2014ECCV 2020, Glasgow, UK.","DOI":"10.1007\/978-3-030-58542-6_22"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Majhi, S., Das, S., and Br\u00e9mond, F. (2021, January 16\u201319). DAM: Dissimilarity attention module for weakly-supervised video anomaly detection. Proceedings of the 2021 17th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), Washington, DC, USA.","DOI":"10.1109\/AVSS52988.2021.9663810"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"5575","DOI":"10.1038\/s41467-020-19266-y","article-title":"State-of-the-art augmented NLP transformer models for direct and single-step retrosynthesis","volume":"11","author":"Tetko","year":"2020","journal-title":"Nat. Commun."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Sun, L., Xia, C., Yin, W., Liang, T., Yu, P.S., and He, L. (2020). Mixup-transformer: Dynamic data augmentation for nlp tasks. arXiv.","DOI":"10.18653\/v1\/2020.coling-main.305"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Morris, J.X., Lifland, E., Yoo, J.Y., Grigsby, J., Jin, D., and Qi, Y. (2020). Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp. arXiv.","DOI":"10.18653\/v1\/2020.emnlp-demos.16"},{"key":"ref_31","unstructured":"Amin-Nejad, A., Ive, J., and Velupillai, S. (2020, January 11\u201316). Exploring transformer text generation for medical dataset augmentation. Proceedings of the Twelfth Language Resources and Evaluation Conference, Marseille, France."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Mercat, J., Gilles, T., El Zoghby, N., Sandou, G., Beauvois, D., and Gil, G.P. (August, January 31). Multi-head attention for multi-modal joint vehicle motion forecasting. Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France.","DOI":"10.1109\/ICRA40945.2020.9197340"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Voita, E., Talbot, D., Moiseev, F., Sennrich, R., and Titov, I. (2019). Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. arXiv.","DOI":"10.18653\/v1\/P19-1580"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Cheng, K.W., Chen, Y.T., and Fang, W.H. (2015, January 7\u201312). Video anomaly detection and localization using hierarchical feature representation and Gaussian process regression. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298909"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Zhao, Y., Deng, B., Shen, C., Liu, Y., Lu, H., and Hua, X.S. (2017, January 23\u201327). Spatio-temporal autoencoder for video anomaly detection. Proceedings of the 25th ACM international conference on Multimedia, Mountain View, CA, USA.","DOI":"10.1145\/3123266.3123451"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1296","DOI":"10.1109\/JAS.2021.1004045","article-title":"A cognitive memory-augmented network for visual anomaly detection","volume":"8","author":"Wang","year":"2021","journal-title":"IEEE\/CAA J. Autom. Sin."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Liu, Z., Nie, Y., Long, C., Zhang, Q., and Li, G. (2021, January 10\u201317). A hybrid video anomaly detection framework via memory-augmented flow reconstruction and flow-guided frame prediction. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01333"},{"key":"ref_38","unstructured":"Hasan, M., Choi, J., Neumann, J., Roy-Chowdhury, A.K., and Davis, L.S. (July, January 26). Learning temporal regularity in video sequences. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Las Vegas, NV, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/1\/58\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:39:49Z","timestamp":1760132389000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/1\/58"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,21]]},"references-count":38,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,1]]}},"alternative-id":["s24010058"],"URL":"https:\/\/doi.org\/10.3390\/s24010058","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,21]]}}}