{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T16:28:29Z","timestamp":1783009709395,"version":"3.54.5"},"reference-count":33,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2023,3,11]],"date-time":"2023-03-11T00:00:00Z","timestamp":1678492800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100000780","name":"European Commission","doi-asserted-by":"publisher","award":["883345"],"award-info":[{"award-number":["883345"]}],"id":[{"id":"10.13039\/501100000780","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Fire detection in videos forms a valuable feature in surveillance systems, as its utilization can prevent hazardous situations. The combination of an accurate and fast model is necessary for the effective confrontation of this significant task. In this work, a transformer-based network for the detection of fire in videos is proposed. It is an encoder\u2013decoder architecture that consumes the current frame that is under examination, in order to compute attention scores. These scores denote which parts of the input frame are more relevant for the expected fire detection output. The model is capable of recognizing fire in video frames and specifying its exact location in the image plane in real-time, as can be seen in the experimental results, in the form of segmentation mask. The proposed methodology has been trained and evaluated for two computer vision tasks, the full-frame classification task (fire\/no fire in frames) and the fire localization task. In comparison with the state-of-the-art models, the proposed method achieves outstanding results in both tasks, with 97% accuracy, 20.4 fps processing time, 0.02 false positive rate for fire localization, and 97% for f-score and recall metrics in the full-frame classification task.<\/jats:p>","DOI":"10.3390\/s23063035","type":"journal-article","created":{"date-parts":[[2023,3,13]],"date-time":"2023-03-13T03:28:33Z","timestamp":1678678113000},"page":"3035","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":22,"title":["Transformer-Based Fire Detection in Videos"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9188-0124","authenticated-orcid":false,"given":"Konstantina","family":"Mardani","sequence":"first","affiliation":[{"name":"Information Technologies Institute (ITI), Centre for Research and Technology Hellas (CERTH), 57001 Thessaloniki, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3604-9685","authenticated-orcid":false,"given":"Nicholas","family":"Vretos","sequence":"additional","affiliation":[{"name":"Information Technologies Institute (ITI), Centre for Research and Technology Hellas (CERTH), 57001 Thessaloniki, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3814-6710","authenticated-orcid":false,"given":"Petros","family":"Daras","sequence":"additional","affiliation":[{"name":"Information Technologies Institute (ITI), Centre for Research and Technology Hellas (CERTH), 57001 Thessaloniki, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,11]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"591","DOI":"10.1007\/s10694-020-01064-z","article-title":"Machine vision based fire detection techniques: A survey","volume":"57","author":"Geetha","year":"2021","journal-title":"Fire Technol."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"941","DOI":"10.1007\/s00138-010-0276-x","article-title":"An integrated fire detection and suppression system based on widely available video surveillance","volume":"21","author":"Yuan","year":"2010","journal-title":"Mach. Vis. Appl."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1943","DOI":"10.1007\/s10694-020-00986-y","article-title":"Video Flame and Smoke Based Fire Detection Algorithms: A Literature Review","volume":"56","author":"Gaur","year":"2020","journal-title":"Fire Technol."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1419","DOI":"10.1109\/TSMC.2018.2830099","article-title":"Efficient deep CNN-based fire detection and localization in video surveillance applications","volume":"49","author":"Muhammad","year":"2018","journal-title":"IEEE Trans. Syst. Man Cybern. Syst."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Aslan, S., G\u00fcd\u00fckbay, U., T\u00f6reyin, B.U., and Cetin, A.E. (2019). Deep convolutional generative adversarial networks based flame detection in video. arXiv.","DOI":"10.1007\/978-3-030-63007-2_63"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Yu, N., and Chen, Y. (2019, January 24\u201326). Video flame detection method based on TwoStream convolutional neural network. Proceedings of the 2019 IEEE 8th Joint International Information Technology and Artificial Intelligence Conference (ITAIC), Chongqing, China.","DOI":"10.1109\/ITAIC.2019.8785841"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2862","DOI":"10.3390\/app9142862","article-title":"A video-based fire detection using deep learning models","volume":"9","author":"Kim","year":"2019","journal-title":"Appl. Sci."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Thomson, W., Bhowmik, N., and Breckon, T.P. (2020, January 14\u201317). Efficient and Compact Convolutional Neural Network Architectures for Non-temporal Real-time Fire Detection. Proceedings of the 2020 19th IEEE International Conference on Machine Learning and Applications (ICMLA), Miami, FL, USA.","DOI":"10.1109\/ICMLA51294.2020.00030"},{"key":"ref_9","unstructured":"Samarth, G., Bhowmik, N., and Breckon, T.P. (2019, January 16\u201319). Experimental exploration of compact convolutional neural network architectures for non-temporal real-time fire detection. Proceedings of the 2019 18th IEEE International Conference on Machine Learning and Applications (ICMLA), Boca Raton, FL, USA."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1419","DOI":"10.1007\/s11760-017-1102-y","article-title":"Video fire detection based on Gaussian Mixture Model and multi-color features","volume":"11","author":"Han","year":"2017","journal-title":"Signal Image Video Process."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1016\/j.firesaf.2015.11.015","article-title":"Fast fire flame detection in surveillance video using logistic regression and temporal smoothing","volume":"79","author":"Kong","year":"2016","journal-title":"Fire Saf. J."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3505244","article-title":"Transformers in vision: A survey","volume":"54","author":"Khan","year":"2022","journal-title":"Acm Comput. Surv."},{"key":"ref_13","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_14","first-page":"213","article-title":"End-to-end object detection with transformers","volume":"Volume 12346","author":"Carion","year":"2020","journal-title":"Computer Vision, Proceedings of the European Conference on Computer Vision (ECCV 2020), Glasgow, UK, 23\u201328 August 2020"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Steffens, C.R., Rodrigues, R.N., and Silva da Costa Botelho, S. (2015, January 29\u201331). An Unconstrained Dataset for Non-Stationary Video Based Fire Detection. Proceedings of the 2015 12th Latin American Robotics Symposium and 2015 3rd Brazilian Symposium on Robotics (LARS-SBR), Uberlandia, Brazil.","DOI":"10.1109\/LARS-SBR.2015.10"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2274","DOI":"10.1109\/TPAMI.2012.120","article-title":"SLIC Superpixels Compared to State-of-the-Art Superpixel Methods","volume":"34","author":"Achanta","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_17","unstructured":"\u00c7elik, T., \u00d6zkaramanl\u0131, H., and Demirel, H. (2007, January 3\u20137). Fire and smoke detection without sensors: Image processing based approach. Proceedings of the 2007 15th European Signal Processing Conference, Poznan, Poland."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"881","DOI":"10.4218\/etrij.10.0109.0695","article-title":"Fast and efficient method for fire detection using image processing","volume":"32","author":"Celik","year":"2010","journal-title":"ETRI J."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"899","DOI":"10.3923\/itj.2010.899.908","article-title":"Early fire detection based on flame contours in video","volume":"9","author":"Zhou","year":"2010","journal-title":"Inf. Technol. J."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Chenebert, A., Breckon, T.P., and Gaszczak, A. (2011, January 11\u201314). A non-temporal texture driven approach to real-time fire detection. Proceedings of the 2011 18th IEEE International Conference on Image Processing, Brussels, Belgium.","DOI":"10.1109\/ICIP.2011.6115796"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1545","DOI":"10.1109\/TCSVT.2015.2392531","article-title":"Real-time fire detection for video-surveillance applications using a combination of experts based on color, shape, and motion","volume":"25","author":"Foggia","year":"2015","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1939171","DOI":"10.1155\/2019\/1939171","article-title":"A real-time fire detection method from video with multifeature fusion","volume":"2019","author":"Gong","year":"2019","journal-title":"Comput. Intell. Neurosci."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zhang, Q., Xu, J., Xu, L., and Guo, H. (2016, January 30\u201331). Deep convolutional neural networks for forest fire detection. Proceedings of the 2016 International Forum on Management, Education and Information Technology Application, Guangzhou, China.","DOI":"10.2991\/ifmeita-16.2016.105"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"712","DOI":"10.3390\/s18030712","article-title":"Saliency detection and deep learning-based wildfire identification in UAV imagery","volume":"18","author":"Zhao","year":"2018","journal-title":"Sensors"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Dunnings, A.J., and Breckon, T.P. (2018, January 7\u201310). Experimentally Defined Convolutional Neural Network Architecture Variants for Non-Temporal Real-Time Fire Detection. Proceedings of the 2018 25th IEEE International Conference on Image Processing (ICIP), Athens, Greece.","DOI":"10.1109\/ICIP.2018.8451657"},{"key":"ref_26","unstructured":"Bello, I., Zoph, B., Vaswani, A., Shlens, J., and Le, Q.V. (November, January 27). Attention Augmented Convolutional Networks. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_27","unstructured":"Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, L., Shazeer, N., Ku, A., and Tran, D. (2018, January 10\u201315). Image Transformer. Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_28","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Li, X., Sun, X., Meng, Y., Liang, J., Wu, F., and Li, J. (2019). Dice loss for data-imbalanced NLP tasks. arXiv.","DOI":"10.18653\/v1\/2020.acl-main.45"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_31","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., and Antiga, L. (2019, January 8\u201314). Pytorch: An imperative style, high-performance deep learning library. Proceedings of the 32 Advances in Neural Information Processing Systems, Vancouver, BC, USA."},{"key":"ref_32","unstructured":"Loshchilov, I., and Hutter, F. (2017). Decoupled weight decay regularization. arXiv."},{"key":"ref_33","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/6\/3035\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:52:40Z","timestamp":1760122360000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/6\/3035"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,11]]},"references-count":33,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["s23063035"],"URL":"https:\/\/doi.org\/10.3390\/s23063035","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,11]]}}}