{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T02:19:52Z","timestamp":1760149192951,"version":"build-2065373602"},"reference-count":44,"publisher":"MDPI AG","issue":"14","license":[{"start":{"date-parts":[[2023,7,16]],"date-time":"2023-07-16T00:00:00Z","timestamp":1689465600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China","award":["61673396","ZR2022MF260"],"award-info":[{"award-number":["61673396","ZR2022MF260"]}]},{"name":"Natural Science Foundation of Shandong Province","award":["61673396","ZR2022MF260"],"award-info":[{"award-number":["61673396","ZR2022MF260"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Instance segmentation is a challenging task in computer vision, as it requires distinguishing objects and predicting dense areas. Currently, segmentation models based on complex designs and large parameters have achieved remarkable accuracy. However, from a practical standpoint, achieving a balance between accuracy and speed is even more desirable. To address this need, this paper presents ESAMask, a real-time segmentation model fused with efficient sparse attention, which adheres to the principles of lightweight design and efficiency. In this work, we propose several key contributions. Firstly, we introduce a dynamic and sparse Related Semantic Perceived Attention mechanism (RSPA) for adaptive perception of different semantic information of various targets during feature extraction. RSPA uses the adjacency matrix to search for regions with high semantic correlation of the same target, which reduces computational cost. Additionally, we design the GSInvSAM structure to reduce redundant calculations of spliced features while enhancing interaction between channels when merging feature layers of different scales. Lastly, we introduce the Mixed Receptive Field Context Perception Module (MRFCPM) in the prototype branch to enable targets of different scales to capture the feature representation of the corresponding area during mask generation. MRFCPM fuses information from three branches of global content awareness, large kernel region awareness, and convolutional channel attention to explicitly model features at different scales. Through extensive experimental evaluation, ESAMask achieves a mask AP of 45.4 at a frame rate of 45.2 FPS on the COCO dataset, surpassing current instance segmentation methods in terms of the accuracy\u2013speed trade-off, as demonstrated by our comprehensive experimental results. In addition, the high-quality segmentation results of our proposed method for objects of various classes and scales can be intuitively observed from the visualized segmentation outputs.<\/jats:p>","DOI":"10.3390\/s23146446","type":"journal-article","created":{"date-parts":[[2023,7,17]],"date-time":"2023-07-17T01:06:36Z","timestamp":1689555996000},"page":"6446","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["ESAMask: Real-Time Instance Segmentation Fused with Efficient Sparse Attention"],"prefix":"10.3390","volume":"23","author":[{"given":"Qian","family":"Zhang","sequence":"first","affiliation":[{"name":"College of Computer Science and Technology, China University of Petroleum (East China), Qingdao 266580, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lu","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, China University of Petroleum (East China), Qingdao 266580, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mingwen","family":"Shao","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, China University of Petroleum (East China), Qingdao 266580, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hong","family":"Liang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, China University of Petroleum (East China), Qingdao 266580, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jie","family":"Ren","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, China University of Petroleum (East China), Qingdao 266580, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,7,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"101819","DOI":"10.1016\/j.aei.2022.101819","article-title":"UAV imagery based potential safety hazard evaluation for high-speed railroad using Real-time instance segmentation","volume":"55","author":"Wu","year":"2023","journal-title":"Adv. Eng. Inform."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"102569","DOI":"10.1016\/j.media.2022.102569","article-title":"Real-time instance segmentation of surgical instruments using attention and multi-scale feature fusion","volume":"81","author":"Ruiz","year":"2022","journal-title":"Med. Image Anal."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201323). Path aggregation network for instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Huang, Z., Huang, L., Gong, Y., Huang, C., and Wang, X. (2019, January 15\u201320). Mask scoring r-cnn. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00657"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Kirillov, A., Wu, Y., He, K., and Girshick, R. (2020, January 13\u201319). Pointrend: Image segmentation as rendering. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00982"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Tang, C., Chen, H., Li, X., Li, J., Zhang, Z., and Hu, X. (2021, January 20\u201325). Look closer to segment better: Boundary patch refinement for instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01371"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Cheng, T., Wang, X., Huang, L., and Liu, W. (2020, January 23\u201328). Boundary-preserving mask r-cnn. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Proceedings, Part XIV 16.","DOI":"10.1007\/978-3-030-58568-6_39"},{"key":"ref_9","unstructured":"Bolya, D., Zhou, C., Xiao, F., and Lee, Y.G. (November, January 27). Yolact: Real-time instance segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Bolya, D., Zhou, C., Xiao, F., and Lee, Y.G. (2019). Yolact++: Better real-time instance segmentation. arXiv.","DOI":"10.1109\/ICCV.2019.00925"},{"key":"ref_11","unstructured":"Fu, C.Y., Shvets, M., and Berg, A.C. (2019). RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Xie, E., Sun, P., Song, X., Wang, W., Liu, X., Liang, D., Shen, C., and Luo, P. (2020, January 13\u201319). Polarmask: Single shot instance segmentation with polar representation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01221"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Chen, H., Sun, K., Tian, Z., Shen, C., Huang, Y., and Yan, Y. (2020, January 13\u201319). Blendmask: Top-down meets bottom-up for instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00860"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"4063","DOI":"10.1007\/s11042-022-13447-1","article-title":"RISAT: Real-time instance segmentation with adversarial training","volume":"82","author":"Pei","year":"2023","journal-title":"Multimed. Tools Appl."},{"key":"ref_15","unstructured":"Jocher, G., Chaurasia, A., and Qiu, J. (2023, March 06). YOLO by Ultralytics (Version8.0.0) [Computer software]. Available online: https:\/\/github.com\/ultralytics\/ultralytics."},{"key":"ref_16","unstructured":"Jocher, G. (2020, October 08). YOLOv5 by Ultralytics (Version 7.0) [Computer Software]. Available online: https:\/\/zenodo.org\/record\/7347926."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zheng, J., Wu, H., Zhang, H., Wang, Z., and Xu, W. (2022). Insulator-defect detection algorithm based on improved YOLOv7. Sensors, 22.","DOI":"10.3390\/s22228801"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Gallo, I., Rehman, A.U., Dehkordi, R.H., Landro, N., Grassa, R.L., and Boschetti, M. (2023). Deep object detection of crop weeds: Performance of YOLOv7 on a real case dataset from UAV images. Remote Sens., 15.","DOI":"10.3390\/rs15020539"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Dewi, C., Chen, A.P.S., and Christanto, H.J. (2023). Deep Learning for Highly Accurate Hand Recognition Based on Yolov7 Model. Big Data Cogn. Comput., 7.","DOI":"10.3390\/bdcc7010053"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 11\u20137). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Ke, L., Danelljan, M., Li, X., Tai, Y., Tang, C.K., and Yu, F. (2022, January 18\u201322). Mask transfiner for high-quality instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52688.2022.00437"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Fang, Y., Yang, S., Wang, X., Li, Y., Fang, C., Shan, Y., Feng, B., and Liu, W. (2021, January 11\u20137). Instances as queries. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00683"},{"key":"ref_23","first-page":"21898","article-title":"Solq: Segmenting objects by learning queries","volume":"34","author":"Dong","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Cheng, B., Misra, I., Schwing, A.G., Kirillov, A., and Girdhar, R. (2022, January 18\u201322). Masked-attention mask transformer for universal image segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52688.2022.00135"},{"key":"ref_25","first-page":"1","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Hassani, A., Walton, S., Li, J., Li, S., and Shi, H. (2023, January 18\u201322). Neighborhood attention transformer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00599"},{"key":"ref_27","unstructured":"Hassani, A., and Shi, H. (2022). Dilated neighborhood attention transformer. arXiv."},{"key":"ref_28","unstructured":"Li, H., Li, J., Wei, H., Liu, Z., Zhan, Z., and Ren, Q. (2022). Slim-neck by GSConv: A better design paradigm of detector architectures for autonomous vehicles. arXiv."},{"key":"ref_29","unstructured":"Yang, L., Zhang, R.Y., Li, L., and Xie, X. (2021, January 18\u201324). Simam: A simple, parameter-free attention module for convolutional neural networks. Proceedings of the International Conference on Machine Learning, Online."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Faster r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zhang, G., Lu, X., Tan, J., Li, J., Zhang, Z., Li, Q., and Hu, X. (2021, January 20\u201325). Refinemask: Towards high-quality instance segmentation with fine-grained features. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00679"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhu, C., Zhang, X., Li, Y., Qiu, L., Han, K., and Han, X. (2022, January 18\u201322). SharpContour: A contour-based boundary refinement approach for efficient and accurate instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52688.2022.00435"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Lee, Y., and Park, J. (2020, January 13\u201319). Centermask: Real-time anchor-free instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01392"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wang, X., Kong, T., Shen, C., Jiang, Y., and Li, L. (2020, January 23\u201328). Solo: Segmenting objects by locations. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Proceedings, Part XVIII 16.","DOI":"10.1007\/978-3-030-58523-5_38"},{"key":"ref_36","first-page":"17721","article-title":"Solov2: Dynamic and fast instance segmentation","volume":"33","author":"Wang","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Sun, P., Zhang, R., Jiang, Y., Kong, T., Xu, C., Zhan, W., Tomizuka, M., Li, L., Yuan, Z., and Wang, C. (2021, January 20\u201325). Sparse r-cnn: End-to-end object detection with learnable proposals. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01422"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Li, F., Zhang, H., Xu, H., Liu, S., Zhang, L., Ni, L.M., and Shum, H.Y. (2023, January 18\u201322). Mask dino: Towards a unified transformer-based framework for object detection and segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00297"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Nguyen, D.K., Ju, J., Booij, O., Oswald, M.R., and Snoek, C.M. (2022, January 18\u201322). Boxer: Box-attention for 2d and 3d transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52688.2022.00473"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Lee, Y., Hwang, J., Lee, S., Bae, Y., and Park, J. (2019, January 16\u201317). An energy and GPU-computation efficient backbone network for real-time object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Long Beach, CA, USA.","DOI":"10.1109\/CVPRW.2019.00103"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the Computer Vision\u2013ECCV 2014: 13th European Conference, Zurich, Switzerland. Proceedings, Part V 13.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Cheng, T., Wang, X., Chen, S., Zhang, W., Zhang, Q., Huang, C., Zhang, Z., and Liu, W. (2022, January 18\u201322). Sparse instance activation for real-time instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52688.2022.00439"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Zhang, T., Wei, S., and Ji, S. (2022, January 18\u201322). E2ec: An end-to-end contour-based method for high-quality high-speed instance segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52688.2022.00440"},{"key":"ref_44","first-page":"1438","article-title":"Close the loop: A unified bottom-up and top-down paradigm for joint image deraining and segmentation","volume":"36","author":"Li","year":"2022","journal-title":"Proc. AAAI Conf. Artif. Intell."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/14\/6446\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:13:03Z","timestamp":1760127183000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/14\/6446"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,16]]},"references-count":44,"journal-issue":{"issue":"14","published-online":{"date-parts":[[2023,7]]}},"alternative-id":["s23146446"],"URL":"https:\/\/doi.org\/10.3390\/s23146446","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2023,7,16]]}}}