{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T01:32:12Z","timestamp":1781055132657,"version":"3.54.1"},"reference-count":44,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2025,11,23]],"date-time":"2025-11-23T00:00:00Z","timestamp":1763856000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>With the increasing demand for higher precision and real-time performance in industrial surface defect detection, multimodal detection methods integrating RGB images and 3D point clouds have drawn considerable attention. However, current mainstream methods typically employ computationally expensive Transformer-based models for capturing global features, resulting in significant inference delays that hinder their practical deployment for online inspection tasks. Furthermore, existing approaches exhibit limited capability in deep cross-modal interactions, negatively impacting defect detection and segmentation accuracy. In this paper, we propose a novel multimodal anomaly detection framework based on a bidirectional Mamba network to enhance cross-modal feature interaction and fusion. Specifically, we introduce an anomaly-aware parallel feature extraction network, leveraging a hybrid scanning state space model (SSM) to efficiently capture global and long-range dependencies with linear computational complexity. Additionally, we develop a cross-enhanced feature fusion module to facilitate dynamic interaction and adaptive fusion of multimodal features at multiple scales. Extensive experiments conducted on two publicly available benchmark datasets, MVTec 3D-AD and Eyecandies, demonstrate that the proposed method consistently outperforms existing approaches in both defect detection and segmentation tasks.<\/jats:p>","DOI":"10.3390\/info16121018","type":"journal-article","created":{"date-parts":[[2025,11,24]],"date-time":"2025-11-24T09:02:07Z","timestamp":1763974927000},"page":"1018","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["HFMM-Net: A Hybrid Fusion Mamba Network for Efficient Multimodal Industrial Defect Detection"],"prefix":"10.3390","volume":"16","author":[{"given":"Guo","family":"Zhao","sequence":"first","affiliation":[{"name":"School of Electrical and Electronic Engineering, Hubei University of Technology, Wuhan 430068, China"},{"name":"Hubei Key Laboratory for High-Efficiency Utilization of Solar Energy and Operation Control of Energy Storage System, Hubei University of Technology, Wuhan 430068, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-1359-8838","authenticated-orcid":false,"given":"Liang","family":"Tan","sequence":"additional","affiliation":[{"name":"School of Electrical and Electronic Engineering, Hubei University of Technology, Wuhan 430068, China"},{"name":"Hubei Key Laboratory for High-Efficiency Utilization of Solar Energy and Operation Control of Energy Storage System, Hubei University of Technology, Wuhan 430068, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Musong","family":"He","sequence":"additional","affiliation":[{"name":"School of Electrical and Electronic Engineering, Hubei University of Technology, Wuhan 430068, China"},{"name":"Hubei Key Laboratory for High-Efficiency Utilization of Solar Energy and Operation Control of Energy Storage System, Hubei University of Technology, Wuhan 430068, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qi","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Electrical and Electronic Engineering, Hubei University of Technology, Wuhan 430068, China"},{"name":"Hubei Key Laboratory for High-Efficiency Utilization of Solar Energy and Operation Control of Energy Storage System, Hubei University of Technology, Wuhan 430068, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,11,23]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"104","DOI":"10.1007\/s11633-023-1459-z","article-title":"Deep industrial image anomaly detection: A survey","volume":"21","author":"Liu","year":"2024","journal-title":"Mach. Intell. Res."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Bergmann, P., Fauser, M., Sattlegger, D., and Steger, C. (2020, January 14\u201319). Uninformed students: Student-teacher anomaly detection with discriminative latent embeddings. Proceedings of the IEEE\/CVF Conference on Computer vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00424"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Bergmann, P., Jin, X., Sattlegger, D., and Steger, C. (2021). The mvtec 3d-ad dataset for unsupervised 3d anomaly detection and localization. arXiv.","DOI":"10.5220\/0010865000003124"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Bonfiglioli, L., Toschi, M., Silvestri, D., Fioraio, N., and De Gregorio, D. (2022, January 4\u20138). The eyecandies dataset for unsupervised multimodal anomaly detection and localization. Proceedings of the Asian Conference on Computer Vision, Macao, China.","DOI":"10.1007\/978-3-031-26348-4_27"},{"key":"ref_5","unstructured":"Tu, Y., Zhang, B., Liu, L., Li, Y., Zhang, J., Wang, Y., Wang, C., and Zhao, C. (October, January 29). Self-supervised feature adaptation for 3d industrial anomaly detection. Proceedings of the European Conference on Computer Vision, Milan, Italy."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Rudolph, M., Wehrbein, T., Rosenhahn, B., and Wandt, B. (2023, January 3\u20137). Asymmetric student-teacher networks for industrial anomaly detection. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV56688.2023.00262"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Gu, Z., Zhang, J., Liu, L., Chen, X., Peng, J., Gan, Z., Jiang, G., Shu, A., Wang, Y., and Ma, L. (2024, January 20\u201321). Rethinking reverse distillation for multi-modal anomaly detection. Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada.","DOI":"10.1609\/aaai.v38i8.28687"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"105766","DOI":"10.1016\/j.autcon.2024.105766","article-title":"Anomaly detection of cracks in synthetic masonry arch bridge point clouds using fast point feature histograms and PatchCore","volume":"168","author":"Jing","year":"2024","journal-title":"Autom. Constr."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"110761","DOI":"10.1016\/j.patcog.2024.110761","article-title":"Complementary pseudo multimodal feature for point cloud anomaly detection","volume":"156","author":"Cao","year":"2024","journal-title":"Pattern Recognit."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Horwitz, E., and Hoshen, Y. (2023, January 18\u201322). Back to the feature: Classical 3d features are (almost) all you need for 3d anomaly detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPRW59228.2023.00298"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Wang, Y., Peng, J., Zhang, J., Yi, R., Wang, Y., and Wang, C. (2023, January 18\u201322). Multimodal industrial anomaly detection via hybrid fusion. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00776"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Costanzino, A., Ramirez, P.Z., Lisanti, G., and Di Stefano, L. (2024, January 16\u201322). Multimodal industrial anomaly detection by crossmodal feature mapping. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01631"},{"key":"ref_13","unstructured":"Gu, A., and Dao, T. (2023). Mamba: Linear-time sequence modeling with selective state spaces. arXiv."},{"key":"ref_14","first-page":"103031","article-title":"Vmamba: Visual state space model","volume":"37","author":"Liu","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zhang, H., Zhu, Y., Wang, D., Zhang, L., Chen, T., Wang, Z., and Ye, Z. (2024). A survey on visual mamba. Appl. Sci., 14.","DOI":"10.3390\/app14135683"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"112203","DOI":"10.1016\/j.knosys.2024.112203","article-title":"Semi-Mamba-UNet: Pixel-level contrastive and cross-supervised visual Mamba-based UNet for semi-supervised medical image segmentation","volume":"300","author":"Ma","year":"2024","journal-title":"Knowl.-Based Syst."},{"key":"ref_17","unstructured":"Zou, W., Gao, H., Yang, W., and Liu, T. (November, January 28). Wave-mamba: Wavelet state space model for ultra-high-definition low-light image enhancement. Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, VIC, Australia."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"59808","DOI":"10.52202\/079017-1910","article-title":"Coupled mamba: Enhanced multimodal fusion with coupled state space model","volume":"37","author":"Li","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_19","unstructured":"Perera, P., Nallapati, R., and Xiang, B. (November, January 27). Ocgan: One-class novelty detection using gans with constrained latent representations. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seoul, Republic of Korea."},{"key":"ref_20","unstructured":"Kipf, T.N., and Welling, M. (2016). Variational graph auto-encoders. arXiv."},{"key":"ref_21","unstructured":"Bergmann, P., Fauser, M., Sattlegger, D., and Steger, C. (November, January 27). MVTec AD\u2013A comprehensive real-world dataset for unsupervised anomaly detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seoul, Republic of Korea."},{"key":"ref_22","unstructured":"Gong, D., Liu, L., Le, V., Saha, B., Mansour, M.R., Venkatesh, S., and Hengel, A.v.d. (November, January 27). Memorizing normality to detect anomaly: Memory-augmented deep autoencoder for unsupervised anomaly detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"107706","DOI":"10.1016\/j.patcog.2020.107706","article-title":"Reconstruction by inpainting for visual anomaly detection","volume":"112","author":"Zavrtanik","year":"2021","journal-title":"Pattern Recognit."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Roth, K., Pemula, L., Zepeda, J., Sch\u00f6lkopf, B., Brox, T., and Gehler, P. (2022, January 18\u201324). Towards total recall in industrial anomaly detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01392"},{"key":"ref_25","unstructured":"Horwitz, E., and Hoshen, Y. (2022). An empirical investigation of 3d anomaly detection and segmentation. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Caron, M., Touvron, H., Misra, I., J\u00e9gou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021, January 10\u201317). Emerging properties in self-supervised vision transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"ref_27","first-page":"2440001","article-title":"Masked autoencoders for 3d point cloud self-supervised learning","volume":"1","author":"Pang","year":"2023","journal-title":"World Sci. Annu. Rev. Artif. Intell."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zhao, H., Jiang, L., Jia, J., Torr, P.H., and Koltun, V. (2021, January 10\u201317). Point transformer. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01595"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Wang, Z., Wang, Y., An, L., Liu, J., and Liu, H. (2022). Local transformer network on 3d point cloud semantic segmentation. Information, 13.","DOI":"10.3390\/info13040198"},{"key":"ref_30","unstructured":"Liu, X., Wang, J., Leng, B., and Zhang, S. (2025). Tuned Reverse Distillation: Enhancing Multimodal Industrial Anomaly Detection with Crossmodal Tuners. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Zhu, Q., and Wan, Y. (2025). BiDFNet: A Bidirectional Feature Fusion Network for 3D Object Detection Based on Pseudo-LiDAR. Information, 16.","DOI":"10.20944\/preprints202505.1480.v1"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Ruan, J., Li, J., and Xiang, S. (2024). Vm-unet: Vision mamba unet for medical image segmentation. arXiv.","DOI":"10.1145\/3767748"},{"key":"ref_33","unstructured":"Wang, C., Tsepa, O., Ma, J., and Wang, B. (2024). Graph-mamba: Towards long-range graph sequence modeling with selective state spaces. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"183","DOI":"10.1007\/s11760-024-03798-7","article-title":"Multi-scale representation for image deraining with state space model","volume":"19","author":"Li","year":"2025","journal-title":"Signal Image Video Process."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Tang, Y., Li, Y., Zou, H., and Zhang, X. (2024). Interactive Segmentation for Medical Images Using Spatial Modeling Mamba. Information, 15.","DOI":"10.3390\/info15100633"},{"key":"ref_36","unstructured":"Zhang, T., Yuan, H., Qi, L., Zhang, J., Zhou, Q., Ji, S., Yan, S., and Li, X. (March, January 25). Point cloud mamba: Point cloud learning via state space model. Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA."},{"key":"ref_37","first-page":"32653","article-title":"Pointmamba: A simple state space model for point cloud analysis","volume":"37","author":"Liang","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1007\/s44267-024-00072-9","article-title":"Fusionmamba: Dynamic feature enhancement for multimodal image fusion with mamba","volume":"2","author":"Xie","year":"2024","journal-title":"Vis. Intell."},{"key":"ref_39","unstructured":"Zhu, L., Liao, B., Zhang, Q., Wang, X., Liu, W., and Wang, X. (2024). Vision mamba: Efficient visual representation learning with bidirectional state space model. arXiv."},{"key":"ref_40","unstructured":"Gu, A., Goel, K., and R\u00e9, C. (2021). Efficiently modeling long sequences with structured state spaces. arXiv."},{"key":"ref_41","first-page":"572","article-title":"Combining recurrent, convolutional, and continuous-time models with linear state space layers","volume":"34","author":"Gu","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_42","unstructured":"Smith, J.T., Warrington, A., and Linderman, S.W. (2022). Simplified state space layers for sequence modeling. arXiv."},{"key":"ref_43","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). Pointnet: Deep learning on point sets for 3d classification and segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/12\/1018\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,24]],"date-time":"2025-11-24T09:28:05Z","timestamp":1763976485000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/12\/1018"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,23]]},"references-count":44,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["info16121018"],"URL":"https:\/\/doi.org\/10.3390\/info16121018","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,23]]}}}