{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T01:06:35Z","timestamp":1779325595401,"version":"3.51.4"},"reference-count":26,"publisher":"MDPI AG","issue":"14","license":[{"start":{"date-parts":[[2023,7,14]],"date-time":"2023-07-14T00:00:00Z","timestamp":1689292800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Natural Science Foundation of Fujian Province","award":["2021J011086"],"award-info":[{"award-number":["2021J011086"]}]},{"name":"Natural Science Foundation of Fujian Province","award":["2023J01964"],"award-info":[{"award-number":["2023J01964"]}]},{"name":"Natural Science Foundation of Fujian Province","award":["2023J01965"],"award-info":[{"award-number":["2023J01965"]}]},{"name":"Natural Science Foundation of Fujian Province","award":["2023J01966"],"award-info":[{"award-number":["2023J01966"]}]},{"name":"Natural Science Foundation of Fujian Province","award":["2021JM-020"],"award-info":[{"award-number":["2021JM-020"]}]},{"name":"Natural Science Basic Research Program of Shaanxi","award":["2021J011086"],"award-info":[{"award-number":["2021J011086"]}]},{"name":"Natural Science Basic Research Program of Shaanxi","award":["2023J01964"],"award-info":[{"award-number":["2023J01964"]}]},{"name":"Natural Science Basic Research Program of Shaanxi","award":["2023J01965"],"award-info":[{"award-number":["2023J01965"]}]},{"name":"Natural Science Basic Research Program of Shaanxi","award":["2023J01966"],"award-info":[{"award-number":["2023J01966"]}]},{"name":"Natural Science Basic Research Program of Shaanxi","award":["2021JM-020"],"award-info":[{"award-number":["2021JM-020"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Understanding and analyzing 2D\/3D sensor data is crucial for a wide range of machine learning-based applications, including object detection, scene segmentation, and salient object detection. In this context, interactive object segmentation is a vital task in image editing and medical diagnosis, involving the accurate separation of the target object from its background based on user annotation information. However, existing interactive object segmentation methods struggle to effectively leverage such information to guide object-segmentation models. To address these challenges, this paper proposes an interactive image-segmentation technique for static images based on multi-level semantic fusion. Our method utilizes user-guidance information both inside and outside the target object to segment it from the static image, making it applicable to both 2D and 3D sensor data. The proposed method introduces a cross-stage feature aggregation module, enabling the effective propagation of multi-scale features from previous stages to the current stage. This mechanism prevents the loss of semantic information caused by multiple upsampling and downsampling of the network, allowing the current stage to make better use of semantic information from the previous stage. Additionally, we incorporate a feature channel attention mechanism to address the issue of rough network segmentation edges. This mechanism captures richer feature details from the feature channel level, leading to finer segmentation edges. In the experimental evaluation conducted on the PASCAL Visual Object Classes (VOC) 2012 dataset, our proposed interactive image segmentation method based on multi-level semantic fusion demonstrates an intersection over union (IOU) accuracy approximately 2.1% higher than the currently popular interactive image segmentation method in static images. The comparative analysis highlights the improved performance and effectiveness of our method. Furthermore, our method exhibits potential applications in various fields, including medical imaging and robotics. Its compatibility with other machine learning methods for visual semantic analysis allows for integration into existing workflows. These aspects emphasize the significance of our contributions in advancing interactive image-segmentation techniques and their practical utility in real-world applications.<\/jats:p>","DOI":"10.3390\/s23146394","type":"journal-article","created":{"date-parts":[[2023,7,14]],"date-time":"2023-07-14T08:40:06Z","timestamp":1689324006000},"page":"6394","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["An Interactive Image Segmentation Method Based on Multi-Level Semantic Fusion"],"prefix":"10.3390","volume":"23","author":[{"given":"Ruirui","family":"Zou","sequence":"first","affiliation":[{"name":"School of Physics and Mechanical and Electrical Engineering, Longyan University, Longyan 364012, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qinghui","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Physics and Mechanical and Electrical Engineering, Longyan University, Longyan 364012, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Falin","family":"Wen","sequence":"additional","affiliation":[{"name":"School of Physics and Mechanical and Electrical Engineering, Longyan University, Longyan 364012, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8934-0738","authenticated-orcid":false,"given":"Yang","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Physics and Mechanical and Electrical Engineering, Longyan University, Longyan 364012, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiale","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Software Engineering, Xi\u2019an Jiaotong University, Xi\u2019an 710049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shaoyi","family":"Du","sequence":"additional","affiliation":[{"name":"Institute of Artificial Intelligence and Robotics, Xi\u2019an Jiaotong University, Xi\u2019an 710049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chengzhi","family":"Yuan","sequence":"additional","affiliation":[{"name":"Department of Mechanical, Industrial and Systems Engineering, University of Rhode Island, Kingston, RI 02881, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,7,14]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"110","DOI":"10.1007\/s11390-017-1681-7","article-title":"Intelligent visual media processing: When graphics meets vision","volume":"32","author":"Cheng","year":"2017","journal-title":"J. Comput. Sci. Technol."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1559","DOI":"10.1109\/TPAMI.2018.2840695","article-title":"DeepIGeoS: A deep interactive geodesic framework for medical image segmentation","volume":"41","author":"Wang","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Papadopoulos, D.P., Uijlings, J.R., Keller, F., and Ferrari, V. (2017, January 21\u201326). Extreme clicking for efficient object annotation. Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/ICCV.2017.528"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Maninis, K.K., Caelles, S., Pont-Tuset, J., and Van Gool, L. (2018, January 18\u201322). Deep extreme cut: From extreme points to object segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00071"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Sofiiuk, K., Petrov, I., Barinova, O., and Konushin, A. (2020, January 14\u201319). f-brs: Rethinking backpropagating refinement for interactive segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, DC, USA.","DOI":"10.1109\/CVPR42600.2020.00865"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"261","DOI":"10.1109\/TII.2022.3157319","article-title":"Rethinking click embedding for deep interactive image segmentation","volume":"19","author":"Ding","year":"2022","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_7","unstructured":"Zhang, S., Liew, J.H., Wei, Y., Wei, S., and Zhao, Y. (14, January 14\u201319). Interactive object segmentation with inside-outside guidance. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1515","DOI":"10.1109\/TPAMI.2018.2838670","article-title":"Video object segmentation without temporal information","volume":"41","author":"Maninis","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zheng, C., Zhu, S., Mendieta, M., Yang, T., Chen, C., and Ding, Z. (2021, January 11\u201317). 3d human pose estimation with spatial and temporal transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.01145"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"118705","DOI":"10.1016\/j.eswa.2022.118705","article-title":"Android malware detection based on multi-head squeeze-and-excitation residual network","volume":"212","author":"Zhu","year":"2023","journal-title":"Expert Syst. Appl."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_12","unstructured":"Zhang, H., Goodfellow, I., Metaxas, D., and Odena, A. (2019, January 9\u201315). Self-attention generative adversarial networks. Proceedings of the International Conference on Machine Learning (PMLR), Long Beach, CA, USA."},{"key":"ref_13","unstructured":"Doulamis, A.D., Doulamis, N.D., Ntalianis, K.S., and Kollias, S.D. (April, January 31). Unsupervised semantic object segmentation of stereoscopic video sequences. Proceedings of the IEEE International Conference on Information Intelligence and Systems, Washington, DC, USA."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Tsapatsoulis, N., Avrithis, Y., and Kollias, S. (2000, January 10\u201313). Efficient face detection for multimedia applications. Proceedings of the IEEE International Conference on Image Processing, Vancouver, BC, Canada.","DOI":"10.1109\/ICIP.2000.899289"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Hao, Y., Liu, Y., Wu, Z., Han, L., Chen, Y., Chen, G., and Lai, B. (2021, January 11\u201317). Edgeflow: Achieving practical interactive segmentation with edge-guided flow. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00180"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 21\u201326). Pyramid scene parsing network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. (2018, January 8\u201314). Encoder-decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Chen, Y., Wang, Z., Peng, Y., Zhang, Z., Yu, G., and Sun, J. (2018, January 18\u201323). Cascaded pyramid network for multi-person pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00742"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201322). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_21","unstructured":"Nair, V., and Hinton, G.E. (2010, January 21\u201324). Rectified linear units improve restricted boltzmann machines. Proceedings of the 27th International Conference on Machine Learning, Haifa, Israel."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","article-title":"The pascal visual object classes (voc) challenge","volume":"88","author":"Everingham","year":"2010","journal-title":"Int. J. Comput. Vis."},{"key":"ref_23","unstructured":"Boykov, Y.Y., and Jolly, M.P. (2001, January 7\u201314). Interactive graph cuts for optimal boundary & region segmentation of objects in ND images. Proceedings of the eighth IEEE International Conference on Computer Vision, Vancouver, BC, Canada."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1768","DOI":"10.1109\/TPAMI.2006.233","article-title":"Random walks for image segmentation","volume":"28","author":"Grady","year":"2006","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Bai, X., and Sapiro, G. (2007, January 14\u201321). A geodesic framework for fast interactive image and video segmentation and matting. Proceedings of the IEEE 11th International Conference on Computer Vision, Rio De Janeiro, Brazil.","DOI":"10.1109\/ICCV.2007.4408931"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Liew, J., Wei, Y., Xiong, W., Ong, S.H., and Feng, J. (2017, January 22\u201329). Regional interactive image segmentation networks. Proceedings of the IEEE international Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.297"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/14\/6394\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:11:53Z","timestamp":1760127113000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/14\/6394"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,14]]},"references-count":26,"journal-issue":{"issue":"14","published-online":{"date-parts":[[2023,7]]}},"alternative-id":["s23146394"],"URL":"https:\/\/doi.org\/10.3390\/s23146394","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,14]]}}}