{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,1]],"date-time":"2026-03-01T01:15:00Z","timestamp":1772327700517,"version":"3.50.1"},"reference-count":38,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2025,7,9]],"date-time":"2025-07-09T00:00:00Z","timestamp":1752019200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62175111"],"award-info":[{"award-number":["62175111"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Image segmentation is an important method in the field of image processing, while infrared (IR) image segmentation is one of the challenges in this field due to the unique characteristics of IR data. Infrared imaging utilizes the infrared radiation emitted by objects to produce images, which can supplement the performance of visible-light images under adverse lighting conditions to some extent. However, the low spatial resolution and limited texture details in IR images hinder the achievement of high-precision segmentation. To address these issues, an attention mechanism based on symmetrical cross-channel interaction\u2014motivated by symmetry principles in computer vision\u2014was integrated into a Mask Region-Based Convolutional Neural Network (Mask R-CNN) framework. A Bottleneck-enhanced Squeeze-and-Attention (BNSA) module was incorporated into the backbone network, and novel loss functions were designed for both the bounding box (Bbox) regression and mask prediction branches to enhance segmentation performance. Furthermore, a dedicated infrared image dataset was constructed to validate the proposed method. The experimental results demonstrate that the optimized model achieves higher segmentation accuracy and better segmentation performance compared to the original network and other mainstream segmentation models on our dataset, demonstrating how symmetrical design principles can effectively improve complex vision tasks.<\/jats:p>","DOI":"10.3390\/sym17071099","type":"journal-article","created":{"date-parts":[[2025,7,10]],"date-time":"2025-07-10T07:38:27Z","timestamp":1752133107000},"page":"1099","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Attention-Based Mask R-CNN Enhancement for Infrared Image Target Segmentation"],"prefix":"10.3390","volume":"17","author":[{"given":"Liang","family":"Wang","sequence":"first","affiliation":[{"name":"Shaanxi Aerospace Technology Application Research Institute Co., Ltd., Xi\u2019an 710100, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3391-5795","authenticated-orcid":false,"given":"Kan","family":"Ren","sequence":"additional","affiliation":[{"name":"Jiangsu Key Laboratory of Spectral Imaging and Intelligent Sense, Nanjing University of Science and Technology, Nanjing 210094, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,7,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"66","DOI":"10.1016\/j.ins.2017.08.050","article-title":"Automated segmentation of exudates, haemorrhages, microaneurysms using single convolutional neural network","volume":"420","author":"Tan","year":"2017","journal-title":"Inf. Sci."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"726","DOI":"10.1109\/JSTARS.2020.2971061","article-title":"DeepNEM: Deep Network Energy-Minimization for Agricultural Field Segmentation","volume":"13","author":"Torre","year":"2020","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Song, Y., Huang, T., Fu, X., Jiang, Y., Xu, J., Zhao, J., Yan, W., and Wang, X. (2023). A Novel Lane Line Detection Algorithm for Driverless Geographic Information Perception Using Mixed-Attention Mechanism ResNet and Row Anchor Classification. ISPRS Int. J. Geo-Inf., 12.","DOI":"10.3390\/ijgi12030132"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1136","DOI":"10.1007\/s12205-023-0391-7","article-title":"Internal Defect Detection of Structures Based on Infrared Thermography and Deep Learning","volume":"27","author":"Deng","year":"2023","journal-title":"KSCE J. Civ. Eng."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"104678","DOI":"10.1016\/j.infrared.2023.104678","article-title":"Near-infrared vascular image segmentation using improved level set method","volume":"131","author":"Li","year":"2023","journal-title":"Infrared Phys. Technol."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_10","unstructured":"Bolya, D., Zhou, C., Xiao, F., and Lee, Y.J. (November, January 27). Yolact: Real-time instance segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Chen, H., Sun, K., Tian, Z., Shen, C., Huang, Y., and Yan, Y. (2020, January 13\u201319). Blendmask: Top-down meets bottom-up for instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00860"},{"key":"ref_12","unstructured":"Tian, Z., Shen, C., Chen, H., and He, T. (November, January 27). Fcos: Fully convolutional one-stage object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_13","unstructured":"Tian, Z., Shen, C., and Chen, H. (2020, January 23\u201328). Conditional convolutions for instance segmentation. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Proceedings, Part I 16."},{"key":"ref_14","unstructured":"Wang, X., Kong, T., Shen, C., Jiang, Y., and Li, L. (2020, January 23\u201328). Solo: Segmenting objects by locations. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Proceedings, Part XVIII 16."},{"key":"ref_15","first-page":"17721","article-title":"Solov2: Dynamic and fast instance segmentation","volume":"33","author":"Wang","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Tian, Z., Shen, C., Wang, X., and Chen, H. (2021, January 20\u201325). Boxinst: High-performance instance segmentation with box annotations. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00540"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Cheng, T., Wang, X., Chen, S., Zhang, W., Zhang, Q., Huang, C., Zhang, Z., and Liu, W. (2022, January 18\u201324). Sparse instance activation for real-time instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00439"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Ha, Q., Watanabe, K., Karasawa, T., Ushiku, Y., and Harada, T. (2017, January 24\u201328). MFNet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8206396"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"2576","DOI":"10.1109\/LRA.2019.2904733","article-title":"RTFNet: RGB-thermal fusion network for semantic segmentation of urban scenes","volume":"4","author":"Sun","year":"2019","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Shivakumar, S.S., Rodrigues, N., Zhou, A., Miller, I.D., Kumar, V., and Taylor, C.J. (2019). PST900: RGB-Thermal Calibration, Dataset and Segmentation Network. arXiv.","DOI":"10.1109\/ICRA40945.2020.9196831"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"103628","DOI":"10.1016\/j.infrared.2020.103628","article-title":"MCNet: Multi-level correction network for thermal image semantic segmentation of nighttime driving scene","volume":"113","author":"Xiong","year":"2021","journal-title":"Infrared Phys. Technol."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"104193","DOI":"10.1016\/j.infrared.2022.104193","article-title":"MPSA: A multi-level pixel spatial attention network for thermal image segmentation based on Deeplabv3+ architecture","volume":"123","author":"Ren","year":"2022","journal-title":"Infrared Phys. Technol."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Hu, J., Zhang, F., Zhang, J., Yao, K., and Xu, C. (2022, January 18\u201320). Semantic segmentation of infrared ships based on scene-aware priors. Proceedings of the SPIE 12557, AOPC 2022: Optical Sensing, Imaging, and Display Technology, Beijing, China.","DOI":"10.1117\/12.2643034"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1318","DOI":"10.1109\/JSEN.2022.3224837","article-title":"An Improved U-Net Model for Infrared Image Segmentation of Wind Turbine Blade","volume":"23","author":"Yu","year":"2023","journal-title":"IEEE Sens. J."},{"key":"ref_25","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the Medical Image Computing and Computer-Assisted Intervention\u2013MICCAI 2015: 18th International Conference, Munich, Germany. Proceedings, Part III 18."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Jiuzhou, W., and Yang, Y. (2021, January 22\u201324). Infrared Airport Scene Segmentation Based on Aggregation Networks. Proceedings of the 2021 IEEE International Conference on Power Electronics, Computer Applications (ICPECA), Shenyang, China.","DOI":"10.1109\/ICPECA51329.2021.9362565"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"012058","DOI":"10.1088\/1742-6596\/1920\/1\/012058","article-title":"Transfer learning and its application research","volume":"1920","author":"Zhou","year":"2021","journal-title":"J. Phys. Conf. Ser."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1484","DOI":"10.1016\/j.visres.2011.04.012","article-title":"Visual attention: The past 25 years","volume":"51","author":"Carrasco","year":"2011","journal-title":"Vis. Res."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Zhong, Z., Lin, Z.Q., Bidart, R., Hu, X., Daya, I.B., Li, Z., Zheng, W.S., Li, J., and Wong, A. (2020, January 13\u201319). Squeeze-and-attention networks for semantic segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01308"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zhang, H., Dana, K., Shi, J., Zhang, Z., Wang, X., Tyagi, A., and Agrawal, A. (2018, January 18\u201323). Context encoding for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00747"},{"key":"ref_33","unstructured":"Gevorgyan, Z. (2022). SIoU loss: More powerful learning for bounding box regression. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1016\/j.neucom.2022.07.042","article-title":"Focal and efficient IOU loss for accurate bounding box regression","volume":"506","author":"Zhang","year":"2022","journal-title":"Neurocomputing"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Zheng, Z., Wang, P., Liu, W., Li, J., Ye, R., and Ren, D. (2020, January 7\u201312). Distance-IoU loss: Faster and better learning for bounding box regression. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6999"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Zhang, H., Chang, H., Ma, B., Wang, N., and Chen, X. (2020, January 23\u201328). Dynamic R-CNN: Towards high quality object detection via dynamic training. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Proceedings, Part XV 16.","DOI":"10.1007\/978-3-030-58555-6_16"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Milletari, F., Navab, N., and Ahmadi, S.A. (2016, January 25\u201328). V-net: Fully convolutional neural networks for volumetric medical image segmentation. Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV), Stanford, CA, USA.","DOI":"10.1109\/3DV.2016.79"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Berman, M., Triki, A.R., and Blaschko, M.B. (2018, January 18\u201323). The lov\u00e1sz-softmax loss: A tractable surrogate for the optimization of the intersection-over-union measure in neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00464"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/7\/1099\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:06:53Z","timestamp":1760033213000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/7\/1099"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,9]]},"references-count":38,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2025,7]]}},"alternative-id":["sym17071099"],"URL":"https:\/\/doi.org\/10.3390\/sym17071099","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,9]]}}}