{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T15:28:12Z","timestamp":1784042892427,"version":"3.55.0"},"reference-count":51,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T00:00:00Z","timestamp":1768262400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Timely and accurate detection of forest fires through unmanned aerial vehicle (UAV) remote sensing target detection technology is of paramount importance. However, multiscale targets and complex environmental interference in UAV remote sensing images pose significant challenges during detection tasks. To address these obstacles, this paper presents FF-Mamba-YOLO, a novel framework based on the principles of Mamba and YOLO (You Only Look Once) that leverages innovative modules and architectures to overcome these limitations. Specifically, we introduce MFEBlock and MFFBlock based on state space models (SSMs) in the backbone and neck parts of the network, respectively, enabling the model to effectively capture global dependencies. Second, we construct CFEBlock, a module that performs feature enhancement before SSM processing, improving local feature processing capabilities. Furthermore, we propose MGBlock, which adopts a dynamic gating mechanism, enhancing the model\u2019s adaptive processing capabilities and robustness. Finally, we enhance the structure of Path Aggregation Feature Pyramid Network (PAFPN) to improve feature fusion quality and introduce DySample to enhance image resolution without significantly increasing computational costs. Experimental results on our self-constructed forest fire image dataset demonstrate that the model achieves 67.4% mAP@50, 36.3% mAP@50:95, and 64.8% precision, outperforming previous state-of-the-art methods. These results highlight the potential of FF-Mamba-YOLO in forest fire monitoring.<\/jats:p>","DOI":"10.3390\/jimaging12010043","type":"journal-article","created":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T18:40:09Z","timestamp":1768329609000},"page":"43","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["FF-Mamba-YOLO: An SSM-Based Benchmark for Forest Fire Detection in UAV Remote Sensing Images"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-0391-2235","authenticated-orcid":false,"given":"Binhua","family":"Guo","sequence":"first","affiliation":[{"name":"College of Computer and Control Engineering, Northeast Forestry University, Harbin 150040, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dinghui","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Mechanical and Electrical Engineering, Northeast Forestry University, Harbin 150040, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhou","family":"Shen","sequence":"additional","affiliation":[{"name":"College of Aulin, Northeast Forestry University, Harbin 150040, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tiebin","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Computer and Control Engineering, Northeast Forestry University, Harbin 150040, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,1,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"eabh2646","DOI":"10.1126\/sciadv.abh2646","article-title":"Increasing forest fire emissions despite the decline in global burned area","volume":"7","author":"Zheng","year":"2021","journal-title":"Sci. Adv."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"eabd2713","DOI":"10.1126\/sciadv.abd2713","article-title":"Tracking and classifying Amazon fire events in near real time","volume":"8","author":"Andela","year":"2022","journal-title":"Sci. Adv."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"eaaz7005","DOI":"10.1126\/science.aaz7005","article-title":"Climate-driven risks to the climate mitigation potential of forests","volume":"368","author":"Anderegg","year":"2020","journal-title":"Science"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"e2213815120","DOI":"10.1073\/pnas.2213815120","article-title":"Anthropogenic climate change impacts exacerbate summer forest fires in California","volume":"120","author":"Turco","year":"2023","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Nikoli\u0107, G., Vujovi\u0107, F., Golijanin, J., \u0160iljeg, A., and Valjarevi\u0107, A. (2023). Modelling of Wildfire Susceptibility in Different Climate Zones in Montenegro Using GIS-MCDA. Atmosphere, 14.","DOI":"10.3390\/atmos14060929"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"0058","DOI":"10.1038\/s41559-016-0058","article-title":"Human exposure and sensitivity to globally extreme wildfire events","volume":"1","author":"Bowman","year":"2017","journal-title":"Nat. Ecol. Evol."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Durlevi\u0107, U., Ili\u0107, V., and Valjarevi\u0107, A. (2025). Wildfire Susceptibility Mapping Using Deep Learning and Machine Learning Models Based on Multi-Sensor Satellite Data Fusion: A Case Study of Serbia. Fire, 8.","DOI":"10.3390\/fire8100407"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Wang, J., Wang, Y., Liu, L., Yin, H., Ye, N., and Xu, C. (2023). Weakly supervised forest fire segmentation in UAV imagery based on foreground-aware pooling and context-aware loss. Remote Sens., 15.","DOI":"10.3390\/rs15143606"},{"key":"ref_9","first-page":"4708523","article-title":"RFWNet: A multiscale remote sensing forest wildfire detection network with digital twinning, adaptive spatial aggregation, and dynamic sparse features","volume":"62","author":"Wang","year":"2024","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_10","unstructured":"Bohush, R., and Brouka, N. (2013, January 26\u201328). Smoke and flame detection in video sequences based on static and dynamic features. Proceedings of the 2013 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA), Poznan, Poland."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1007\/978-3-642-04697-1_2","article-title":"On the evaluation of segmentation methods for wildland fire","volume":"Volume 5807","author":"Philips","year":"2009","journal-title":"Advanced Concepts for Intelligent Vision Systems"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Maeda, N., and Tonooka, H. (2022). Early stage forest fire detection from Himawari-8 AHI images using a modified MOD14 algorithm combined with machine learning. Sensors, 23.","DOI":"10.3390\/s23010210"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Nazarova, T., Martin, P., and Giuliani, G. (2020). Monitoring Vegetation Change in the Presence of High Cloud Cover with Sentinel-2 in a Lowland Tropical Forest Region in Brazil. Remote Sens., 12.","DOI":"10.3390\/rs12111829"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zou, S., Zou, Y., Zhang, M., Luo, S., Chen, Z., and Gao, G. (July, January 30). Fraesormer: Learning Adaptive Sparse Transformer for Efficient Food Recognition. Proceedings of the 2025 IEEE International Conference on Multimedia and Expo (ICME), Nantes, France.","DOI":"10.1109\/ICME59968.2025.11209119"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zou, S., Zou, Y., Zhang, M., Luo, S., Gao, G., and Qi, G. (July, January 30). Learning Dual-Domain Multi-Scale Representations for Single Image Deraining. Proceedings of the 2025 IEEE International Conference on Multimedia and Expo (ICME), Nantes, France.","DOI":"10.1109\/ICME59968.2025.11210243"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"42816","DOI":"10.1109\/ACCESS.2024.3378568","article-title":"YOLOv1 to v8: Unveiling each variant\u2014A comprehensive review of YOLO","volume":"12","author":"Hussain","year":"2024","journal-title":"IEEE Access"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/978-3-031-72751-1_1","article-title":"YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information","volume":"Volume 15089","author":"Leonardis","year":"2025","journal-title":"Computer Vision\u2014ECCV 2024"},{"key":"ref_19","unstructured":"Wang, A., Chen, H., Liu, L., Chen, K., Lin, Z., Han, J., and Ding, G. (2024). YOLOv10: Real-time end-to-end object detection. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Jegham, N., Koh, C.Y., Abdelatti, M., and Hendawi, A. (2025). YOLO evolution: A comprehensive benchmark and architectural review of YOLOv12, YOLO11, and their previous versions. arXiv.","DOI":"10.2139\/ssrn.5175639"},{"key":"ref_21","unstructured":"Lei, M., Li, S., Wu, Y., Hu, H., Zhou, Y., Zheng, X., Ding, G., Du, S., Wu, Z., and Gao, Y. (2025). YOLOv13: Real-time object detection with hypergraph-enhanced adaptive visual perception. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhao, Y., Lv, W., Xu, S., Wei, J., Wang, G., Dang, Q., Liu, Y., and Chen, J. (2024, January 16\u201322). DETRs beat YOLOs on real-time object detection. Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01605"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wang, J., Zhu, W., Wang, P., Yu, X., Liu, L., Omar, M., and Hamid, R. (2023, January 17\u201324). Selective structured state-spaces for long-form video understanding. Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00618"},{"key":"ref_24","unstructured":"Liu, Y., Tian, Y., Zhao, Y., Yu, H., Xie, L., Wang, Y., Ye, Q., Jiao, J., and Liu, Y. (2024). VMamba: Visual state space model. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Wang, J., Zhao, W., Liu, C., Yang, H., and Xu, W. (2024, January 15\u201317). Real-time Object Detection Based on Mamba and YOLOv8. Proceedings of the 2024 4th International Conference on Industrial Automation, Robotics and Control Engineering (IARCE), Chengdu, China.","DOI":"10.1109\/IARCE64300.2024.00055"},{"key":"ref_26","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017). Attention is all you need. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin Transformer: Hierarchical vision transformer using shifted windows. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Lin, Z., Yun, B., and Zheng, Y. (2024). LD-YOLO: A lightweight dynamic forest fire and smoke detection model with dysample and spatial context awareness module. Forests, 15.","DOI":"10.3390\/f15091630"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Han, Y., Duan, B., Guan, R., Yang, G., and Zhen, Z. (2024). LUFFD-YOLO: A lightweight model for UAV remote sensing forest fire detection based on attention mechanism and multi-level feature fusion. Remote Sens., 16.","DOI":"10.3390\/rs16122177"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"105320","DOI":"10.1016\/j.dsp.2025.105320","article-title":"YOLO-SAD for fire detection and localization in real-world images","volume":"165","author":"Yang","year":"2025","journal-title":"Digit. Signal Process."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"113622","DOI":"10.1016\/j.asoc.2025.113622","article-title":"A double-convolution-double-attention transformer network for aircraft cargo hold fire detection","volume":"183","author":"Li","year":"2025","journal-title":"Appl. Soft Comput."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Liu, X., Zhang, C., Huang, F., Xia, S., Wang, G., and Zhang, L. (2025). Vision Mamba: A comprehensive survey and taxonomy. IEEE Trans. Neural Netw. Learn. Syst., 1\u201321.","DOI":"10.1109\/TNNLS.2025.3610435"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"181521","DOI":"10.1109\/ACCESS.2024.3504297","article-title":"Wavelet guided visual state space model and patch resampling enhanced U-shaped structure for skin lesion segmentation","volume":"12","author":"Feng","year":"2024","journal-title":"IEEE Access"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"108435","DOI":"10.1016\/j.bspc.2025.108435","article-title":"Multi-scale vision Mamba-UNet: A Mamba-based method for retinal vessel segmentation","volume":"112","author":"Hu","year":"2026","journal-title":"Biomed. Signal Process. Control"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"117638","DOI":"10.1016\/j.measurement.2025.117638","article-title":"EPDD-YOLO: An efficient benchmark for pavement damage detection based on Mamba-YOLO","volume":"253","author":"Luo","year":"2025","journal-title":"Measurement"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Zheng, X., Kuang, Y., Huo, Y., Zhu, W., Zhang, M., and Wang, H. (2025). HTMNet: Hybrid transformer\u2013Mamba network for hyperspectral target detection. Remote Sens., 17.","DOI":"10.3390\/rs17173015"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Wang, Y., Li, Y., Yang, X., Jiang, R., and Zhang, L. (2025). HDAMNet: Hierarchical dilated adaptive Mamba network for accurate cloud detection in satellite imagery. Remote Sens., 17.","DOI":"10.3390\/rs17172992"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"101351","DOI":"10.1016\/j.atech.2025.101351","article-title":"Evaluating maize emergence quality with multi-task YOLO11-Mamba and UAV-RGB remote sensing","volume":"12","author":"Zhao","year":"2025","journal-title":"Smart Agric. Technol."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Wang, Z., Li, C., Xu, H., Zhu, X., and Li, H. (2024). Mamba YOLO: A simple baseline for object detection with state space model. arXiv.","DOI":"10.1609\/aaai.v39i8.32885"},{"key":"ref_40","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/978-3-030-01261-8_1","article-title":"Group normalization","volume":"Volume 11217","author":"Ferrari","year":"2018","journal-title":"Computer Vision\u2014ECCV 2018"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"721","DOI":"10.1007\/978-981-99-3481-2_55","article-title":"Convolutional gated MLP: Combining convolutions and gMLP","volume":"Volume 1053","author":"Borah","year":"2024","journal-title":"Big Data, Machine Learning, and Applications"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhou, X., Lin, M., and Sun, J. (2018, January 18\u201323). ShuffleNet: An extremely efficient convolutional neural network for mobile devices. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00716"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201323). Path aggregation network for instance segmentation. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Liu, W., Lu, H., Fu, H., and Cao, Z. (2023, January 1\u20136). Learning to upsample by learning to sample. Proceedings of the 2023 IEEE\/CVF International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.00554"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Wang, J., Chen, K., Xu, R., Liu, Z., Loy, C.C., and Lin, D. (November, January 27). CARAFE: Content-aware reassembly of features. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00310"},{"key":"ref_49","unstructured":"Ibn Jafar, A., Islam, A.M., Binta Masud, F., Ullah, J.R., and Ahmed, M.R. (2023). FlameVision: A New Dataset for Wildfire Classification and Detection Using Aerial Imagery. Mendeley Data, V4."},{"key":"ref_50","first-page":"5646316","article-title":"Topology-Aware Hierarchical Mamba for Salient Object Detection in Remote Sensing Imagery","volume":"63","author":"Yang","year":"2025","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Yang, H., Zhao, M., Qiu, Y., Mu, M., Li, Y., and Zhang, B. (2025, January 19\u201321). YOLOv10 Fire Detection Method Combined with Mamba Attention Mechanism. Proceedings of the 2025 5th International Symposium on Artificial Intelligence and Intelligent Manufacturing (AIIM), Chengdu, China.","DOI":"10.1109\/AIIM67611.2025.11232942"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/1\/43\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T18:47:03Z","timestamp":1768330023000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/1\/43"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,13]]},"references-count":51,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1]]}},"alternative-id":["jimaging12010043"],"URL":"https:\/\/doi.org\/10.3390\/jimaging12010043","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,13]]}}}