{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T17:48:47Z","timestamp":1777657727569,"version":"3.51.4"},"reference-count":79,"publisher":"Association for Computing Machinery (ACM)","issue":"2","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62571560"],"award-info":[{"award-number":["62571560"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["42201358"],"award-info":[{"award-number":["42201358"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100031931","name":"Shanghai Artificial Intelligence Laboratory","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100031931","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,2,28]]},"abstract":"<jats:p>\n                    Recent image matting studies are developing toward proposing trimap-free or interactive methods to complete the complex image matting task. Although avoiding the extensive labor of trimap interaction, existing methods still suffer from two limitations: (1) For the single image with multiple objects, it is essential to provide extra interaction information to help determine the matting target; (2) For transparent objects, the accurate regression of alpha matte from the RGB image is much more difficult compared with the opaque ones. In this work, we propose an\n                    <jats:bold>Unified Interactive image Matting (UIM)<\/jats:bold>\n                    method, which solves the limitations and achieves satisfying matting results for any scenario. Specifically, UIM leverages multiple types of user interactions to avoid the ambiguity of multiple matting targets, and we compare the pros and cons of different user interaction types in detail. To unify the matting performance for transparent and opaque objects, we decouple image matting into two stages, i.e., foreground segmentation and transparency prediction. Moreover, we propose a foreground consistency learning strategy to facilitate the feature extraction on the mainstream synthetic matting dataset. Experimental results demonstrate that UIM achieves competitive performance on the Composition-1K test set and a synthetic unified dataset. Our code and models will be released at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/Dinghow\/UIM\">https:\/\/github.com\/Dinghow\/UIM<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3785468","type":"journal-article","created":{"date-parts":[[2025,12,29]],"date-time":"2025-12-29T13:46:11Z","timestamp":1767015971000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Exploring the Interactive Guidance for Unified and Effective Image Matting"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8912-2968","authenticated-orcid":false,"given":"Dinghao","family":"Yang","sequence":"first","affiliation":[{"name":"Shanghai Artificial Intelligence Laboratory, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5625-2966","authenticated-orcid":false,"given":"Bin","family":"Wang","sequence":"additional","affiliation":[{"name":"Shanghai Artificial Intelligence Laboratory, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1838-9176","authenticated-orcid":false,"given":"Weijia","family":"Li","sequence":"additional","affiliation":[{"name":"Sun Yat-Sen University, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8208-3705","authenticated-orcid":false,"given":"Yiqi","family":"Lin","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8697-695X","authenticated-orcid":false,"given":"Conghui","family":"He","sequence":"additional","affiliation":[{"name":"Shanghai Artificial Intelligence Laboratory, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,2,10]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Aksoy Yagiz","year":"2017","unstructured":"Yagiz Aksoy, Tun\u00e7 Ozan Aydin, and Marc Pollefeys. 2017. Designing effective Inter-Pixel information flow for natural image matting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_3_2","first-page":"1","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Bai Xue","year":"2007","unstructured":"Xue Bai and Guillermo Sapiro. 2007. A geodesic framework for fast interactive image and video segmentation and matting. In Proceedings of the International Conference on Computer Vision. IEEE, 1\u20138."},{"key":"e_1_3_2_4_2","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Bearman Amy","year":"2016","unstructured":"Amy Bearman, Olga Russakovsky, Vittorio Ferrari, and Li Fei-Fei. 2016. What\u2019s the point: Semantic segmentation with point supervision. In Proceedings of the European Conference on Computer Vision."},{"key":"e_1_3_2_5_2","first-page":"105","volume-title":"Proceedings of the International Conference on Computer Vision","volume":"1","author":"Boykov Yuri Y.","year":"2001","unstructured":"Yuri Y. Boykov and M.-P. Jolly. 2001. Interactive graph cuts for optimal boundary & region segmentation of objects in ND images. In Proceedings of the International Conference on Computer Vision, Vol. 1. IEEE, 105\u2013112."},{"key":"e_1_3_2_6_2","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Cai Shaofan","year":"2019","unstructured":"Shaofan Cai, Xiaoshuai Zhang, Haoqiang Fan, Haibin Huang, Jiangyu Liu, Jiaming Liu, Jiaying Liu, Jue Wang, and Jian Sun. 2019. Disentangled image matting. In Proceedings of the International Conference on Computer Vision."},{"key":"e_1_3_2_7_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Guanying","year":"2018","unstructured":"Guanying Chen, Kai Han, and Kwan-Yee K. Wong. 2018. TOM-Net: Learning transparent object matting from a single image. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3057493"},{"key":"e_1_3_2_9_2","first-page":"2175","volume-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence","volume":"35","author":"Chen Qifeng","year":"2013","unstructured":"Qifeng Chen, Dingzeyu Li, and Chi-Keung Tang. 2013. KNN matting. IEEE Transactions on Pattern Analysis and Machine Intelligence 35 (2013), 2175\u20132188."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00136"},{"key":"e_1_3_2_11_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chuang Yung-Yu","year":"2001","unstructured":"Yung-Yu Chuang, Brian Curless, David Salesin, and Richard Szeliski. 2001. A bayesian approach to digital matting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_12_2","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Dai Jifeng","year":"2015","unstructured":"Jifeng Dai, Kaiming He, and Jian Sun. 2015. BoxSup: Exploiting bounding boxes to supervise convolutional networks for semantic segmentation. In Proceedings of the International Conference on Computer Vision."},{"key":"e_1_3_2_13_2","first-page":"417","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Ding Henghui","year":"2020","unstructured":"Henghui Ding, Scott Cohen, Brian Price, and Xudong Jiang. 2020. Phraseclick: Toward achieving flexible interactive segmentation by phrase and click. In Proceedings of the European Conference on Computer Vision. Springer, 417\u2013435."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3155958"},{"key":"e_1_3_2_15_2","first-page":"3901","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Enomoto Kenji","year":"2024","unstructured":"Kenji Enomoto, T. J. Rhodes, Brian Price, and Gavin Miller. 2024. PolarMatte: Fully computational ground-truth-quality alpha matte extraction for images and video using polarized screen matting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 3901\u20133909."},{"key":"e_1_3_2_16_2","first-page":"303","volume-title":"International Journal of Computer Vision","volume":"88","author":"Everingham Mark","year":"2010","unstructured":"Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John M. Winn, and Andrew Zisserman. 2010. The pascal visual object classes (VOC) challenge. International Journal of Computer Vision 88 (2010), 303\u2013338."},{"key":"e_1_3_2_17_2","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Feng Xiaoxue","year":"2016","unstructured":"Xiaoxue Feng, Xiaohui Liang, and Zili Zhang. 2016. A cluster sampling method for image matting via sparse coding. In Proceedings of the European Conference on Computer Vision."},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","first-page":"575","DOI":"10.1111\/j.1467-8659.2009.01627.x","article-title":"Shared sampling for real-time alpha matting","volume":"29","author":"Simoes Lopes Gastal Eduardo","year":"2010","unstructured":"Eduardo Simoes Lopes Gastal and Manuel M. Oliveira. 2010. Shared sampling for real-time alpha matting. Computer Graphics Forum 29 (2010), 575\u2013584.","journal-title":"Computer Graphics Forum"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2006.233"},{"key":"e_1_3_2_20_2","first-page":"1551","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Hao Yuying","year":"2021","unstructured":"Yuying Hao, Yi Liu, Zewu Wu, Lin Han, Yizhou Chen, Guowei Chen, Lutao Chu, Shiyu Tang, Zhiliang Yu, Zeyu Chen, et al. 2021. EdgeFlow: Achieving practical interactive segmentation with edge-guided flow. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 1551\u20131560."},{"key":"e_1_3_2_21_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"He Kaiming","year":"2011","unstructured":"Kaiming He, Christoph Rhemann, Carsten Rother, Xiaoou Tang, and Jian Sun. 2011. A global sampling method for alpha matting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_23_2","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Hou Qiqi","year":"2019","unstructured":"Qiqi Hou and Feng Liu. 2019. Context-aware image matting for simultaneous foreground and alpha estimation. In Proceedings of the International Conference on Computer Vision."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3234983"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2018.10.009"},{"key":"e_1_3_2_26_2","unstructured":"Zhanghan Ke Kaican Li Yurou Zhou Qiuhua Wu Xiangyu Mao Qiong Yan and Rynson W. H. Lau. 2020. Is a green screen really necessary for real-time human matting? arXiv:2011.11961. Retrieved from https:\/\/arxiv.org\/abs\/ preprint 2011.11961"},{"key":"e_1_3_2_27_2","doi-asserted-by":"crossref","unstructured":"Alexander Kirillov Eric Mintun Nikhila Ravi Hanzi Mao Chloe Rolland Laura Gustafson Tete Xiao Spencer Whitehead Alexander C. Berg Wan-Yen Lo et al. 2023. Segment anything. arXiv:2304.02643. Retrieved from https:\/\/arxiv.org\/abs\/2304.02643","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"e_1_3_2_28_2","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Lempitsky Victor","year":"2009","unstructured":"Victor Lempitsky, Pushmeet Kohli, Carsten Rother, and Toby Sharp. 2009. Image segmentation with a bounding box prior. In Proceedings of the International Conference on Computer Vision."},{"key":"e_1_3_2_29_2","first-page":"228","volume-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence","author":"Levin Anat","year":"2008","unstructured":"Anat Levin, Dani Lischinski, and Yair Weiss. 2008. A closed-form solution to natural image matting. IEEE Transactions on Pattern Analysis and Machine Intelligence 30 (2008), 228\u2013242."},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2024.3380914"},{"key":"e_1_3_2_31_2","first-page":"6668","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Li Jiachen","year":"2024","unstructured":"Jiachen Li, Roberto Henschel, Vidit Goel, Marianna Ohanyan, Shant Navasardyan, and Humphrey Shi. 2024. Video instance matting. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, 6668\u20136677."},{"key":"e_1_3_2_32_2","first-page":"1775","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Jiachen","year":"2024","unstructured":"Jiachen Li, Jitesh Jain, and Humphrey Shi. 2024. Matting anything. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1775\u20131785."},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01541-0"},{"key":"e_1_3_2_34_2","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence","author":"Li Jizhizi","year":"2021","unstructured":"Jizhizi Li, Jing Zhang, and Dacheng Tao. 2021. Deep automatic natural image matting. In Proceedings of the International Joint Conference on Artificial Intelligence."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02150"},{"key":"e_1_3_2_36_2","first-page":"17397","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Weijia","year":"2023","unstructured":"Weijia Li, Yawen Lai, Linning Xu, Yuanbo Xiangli, Jinhua Yu, Conghui He, Gui-Song Xia, and Dahua Lin. 2023. OmniCity: Omnipotent city understanding with multi-level and multi-view images. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 17397\u201317407."},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.isprsjprs.2023.05.010"},{"key":"e_1_3_2_38_2","first-page":"1958","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Li Weijia","year":"2021","unstructured":"Weijia Li, Wenqian Zhao, Huaping Zhong, Conghui He, and Dahua Lin. 2021. Joint semantic-geometric learning for polygonal building segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35, 1958\u20131965."},{"key":"e_1_3_2_39_2","article-title":"Natural image matting via guided contextual attention","author":"Li Yaoyi","year":"2020","unstructured":"Yaoyi Li and Hongtao Lu. 2020. Natural image matting via guided contextual attention. In Proceedings of the AAAI Conference on Artificial Intelligence.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_2_40_2","first-page":"577","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Zhuwen","year":"2018","unstructured":"Zhuwen Li, Qifeng Chen, and Vladlen Koltun. 2018. Interactive image segmentation with latent diversity. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 577\u2013585."},{"key":"e_1_3_2_41_2","first-page":"2746","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Liew JunHao","year":"2017","unstructured":"JunHao Liew, Yunchao Wei, Wei Xiong, Sim-Heng Ong, and Jiashi Feng. 2017. Regional interactive image segmentation networks. In Proceedings of the International Conference on Computer Vision. IEEE Computer Society, 2746\u20132754."},{"key":"e_1_3_2_42_2","first-page":"3159","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lin Di","year":"2016","unstructured":"Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun. 2016. ScribbleSup: Scribble-supervised convolutional networks for semantic segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 3159\u20133167."},{"key":"e_1_3_2_43_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lin Shanchuan","year":"2021","unstructured":"Shanchuan Lin, Andrey Ryabtsev, Soumyadip Sengupta, Brian L. Curless, Steven M. Seitz, and Ira Kemelmacher-Shlizerman. 2021. Real-time high-resolution background matting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3640017"},{"key":"e_1_3_2_46_2","unstructured":"Songtao Liu Di Huang and Yunhong Wang. 2019. Learning spatial fusion for single-shot object detection. arXiv:1911.09516. Retrieved from https:\/\/arxiv.org\/abs\/1911.09516"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3497747"},{"key":"e_1_3_2_48_2","first-page":"2727","article-title":"Prior-induced information alignment for image matting","author":"Liu Yuhao","year":"2021","unstructured":"Yuhao Liu, Jiake Xie, Yu Qiao, Yong Tang, and Xin Yang. 2021. Prior-induced information alignment for image matting. IEEE Transactions on Multimedia 24 (2021), 2727\u20132738.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-023-01797-8"},{"key":"e_1_3_2_51_2","unstructured":"Sabarinath Mahadevan Paul Voigtlaender and Bastian Leibe. 2018. Iteratively trained interactive segmentation. arXiv:1805.04398. Retrieved from https:\/\/arxiv.org\/abs\/1805.04398s"},{"key":"e_1_3_2_52_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Maninis Kevis-Kokitsi","year":"2018","unstructured":"Kevis-Kokitsi Maninis, Sergi Caelles, Jordi Pont-Tuset, and Luc Van Gool. 2018. Deep extreme cut: From extreme points to object segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_53_2","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Papadopoulos Dim P.","year":"2017","unstructured":"Dim P. Papadopoulos, Jasper R. R. Uijlings, Frank Keller, and Vittorio Ferrari. 2017. Extreme clicking for efficient object annotation. In Proceedings of the International Conference on Computer Vision."},{"key":"e_1_3_2_54_2","first-page":"11696","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Park GyuTae","year":"2022","unstructured":"GyuTae Park, SungJoon Son, JaeYoung Yoo, SeHo Kim, and Nojun Kwak. 2022. MatteFormer: Transformer-based image matting via prior-tokens. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 11696\u201311706."},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3540201"},{"key":"e_1_3_2_56_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Qiao Yu","year":"2020","unstructured":"Yu Qiao, Yuhao Liu, Xin Yang, Dongsheng Zhou, Mingliang Xu, Qiang Zhang, and Xiaopeng Wei. 2020. Attention-guided hierarchical structure aggregation for image matting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_57_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Rhemann Christoph","year":"2009","unstructured":"Christoph Rhemann, Carsten Rother, Jue Wang, Margrit Gelautz, Pushmeet Kohli, and Pamela Rott. 2009. A perceptually motivated online benchmark for image matting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"e_1_3_2_59_2","first-page":"309","volume-title":"ACM Transactions on Graphics","author":"Rother Carsten","year":"2004","unstructured":"Carsten Rother, Vladimir Kolmogorov, and Andrew Blake. 2004. \u201cGrabCut\u201d interactive foreground extraction using iterated graph cuts. ACM Transactions on Graphics 23 (2004), 309\u2013314."},{"key":"e_1_3_2_60_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Sengupta Soumyadip","year":"2020","unstructured":"Soumyadip Sengupta, Vivek Jayaram, Brian Curless, Steven M. Seitz, and Ira Kemelmacher-Shlizerman. 2020. Background matting: The world is your green screen. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_61_2","first-page":"8623","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Sofiiuk Konstantin","year":"2020","unstructured":"Konstantin Sofiiuk, Ilia Petrov, Olga Barinova, and Anton Konushin. 2020. F-BRS: Rethinking backpropagating refinement for interactive segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 8623\u20138632."},{"key":"e_1_3_2_62_2","unstructured":"Konstantin Sofiiuk Ilia A. Petrov and Anton Konushin. 2021. Reviving iterative training with mask guidance for interactive segmentation. arXiv:2102.06583. Retrieved from https:\/\/arxiv.org\/abs\/2102.06583"},{"key":"e_1_3_2_63_2","first-page":"1818","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Tang Meng","year":"2018","unstructured":"Meng Tang, Abdelaziz Djelouah, Federico Perazzi, Yuri Boykov, and Christopher Schroers. 2018. Normalized cut loss for weakly-supervised CNN segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1818\u20131827."},{"key":"e_1_3_2_64_2","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Wang Jue","year":"2005","unstructured":"Jue Wang and Michael F. Cohen. 2005. An iterative optimization approach for unified image segmentation and matting. In Proceedings of the International Conference on Computer Vision."},{"key":"e_1_3_2_65_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wei Tianyi","year":"2021","unstructured":"Tianyi Wei, Dongdong Chen, Wenbo Zhou, Jing Liao, Hanqing Zhao, Weiming Zhang, and Nenghai Yu. 2021. Improved image matting via real-time user clicks and uncertainty estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_66_2","first-page":"418","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Xiao Tete","year":"2018","unstructured":"Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. 2018. Unified perceptual parsing for scene understanding. In Proceedings of the European Conference on Computer Vision (ECCV), 418\u2013434."},{"key":"e_1_3_2_67_2","volume-title":"Proceedings of the British Machine Vision Conference","author":"Xu Ning","year":"2017","unstructured":"Ning Xu, Brian Price, Scott Cohen, Jimei Yang, and Thomas Huang. 2017. Deep GrabCut for object selection. In Proceedings of the British Machine Vision Conference."},{"key":"e_1_3_2_68_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Xu Ning","year":"2017","unstructured":"Ning Xu, Brian L. Price, Scott Cohen, and Thomas S. Huang. 2017. Deep image matting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jag.2023.103609"},{"issue":"4","key":"e_1_3_2_70_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3408323","article-title":"Smart scribbles for image matting","volume":"16","author":"Yang Xin","year":"2020","unstructured":"Xin Yang, Yu Qiao, Shaozhe Chen, Shengfeng He, Baocai Yin, Qiang Zhang, Xiaopeng Wei, and Rynson W. H. Lau. 2020. Smart scribbles for image matting. ACM Transactions on Multimedia Computing, Communications and Applications 16, 4, Article 121 (2020) 1\u201321.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2024.105067"},{"key":"e_1_3_2_72_2","first-page":"25585","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ye Zixuan","year":"2024","unstructured":"Zixuan Ye, Wenze Liu, He Guo, Yujia Liang, Chaoyi Hong, Hao Lu, and Zhiguo Cao. 2024. Unifying automatic and interactive matting with pretrained ViTs. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 25585\u201325594."},{"key":"e_1_3_2_73_2","first-page":"103795","article-title":"CRFFNet: A cross-view reprojection based feature fusion network for fine-grained building segmentation using satellite-view and street-view data","volume":"127","author":"Yu Jinhua","year":"2025","unstructured":"Jinhua Yu, Junyan Ye, Yi Lin, and Weijia Li. 2025. CRFFNet: A cross-view reprojection based feature fusion network for fine-grained building segmentation using satellite-view and street-view data. Information Fusion 127 (2025), 103795.","journal-title":"Information Fusion"},{"key":"e_1_3_2_74_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yu Qihang","year":"2021","unstructured":"Qihang Yu, Jianming Zhang, He Zhang, Yilin Wang, Zhe Lin, Ning Xu, Yutong Bai, and Alan L. Yuille. 2021. Mask guided matting via progressive refinement network. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_75_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Shiyin","year":"2020","unstructured":"Shiyin Zhang, Jun Hao Liew, Yunchao Wei, Shikui Wei, and Yao Zhao. 2020. Interactive object segmentation with inside\u2013outside guidance. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_76_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Yunke","year":"2019","unstructured":"Yunke Zhang, Lixue Gong, Lubin Fan, Peiran Ren, Qixing Huang, Hujun Bao, and Weiwei Xu. 2019. A late fusion CNN for digital matting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.660"},{"key":"e_1_3_2_78_2","volume-title":"Proceedings of the International Conference on Computer Vision","author":"Zheng Yuanjie","year":"2009","unstructured":"Yuanjie Zheng and Chandra Kambhamettu. 2009. Learning based digital matting. In Proceedings of the International Conference on Computer Vision."},{"key":"e_1_3_2_79_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhou Bolei","year":"2017","unstructured":"Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. 2017. Scene parsing through ADE20K dataset. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_80_2","doi-asserted-by":"crossref","first-page":"5828","DOI":"10.1109\/TCSVT.2023.3260025","article-title":"Sampling propagation attention with trimap generation network for natural image matting","author":"Zhou Yuhongze","year":"2023","unstructured":"Yuhongze Zhou, Liguang Zhou, Tin Lun Lam, and Yangsheng Xu. 2023. Sampling propagation attention with trimap generation network for natural image matting. IEEE Transactions on Circuits and Systems for Video Technology 33 (2023), 5828\u20135843.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3785468","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,10]],"date-time":"2026-02-10T12:13:46Z","timestamp":1770725626000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3785468"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,10]]},"references-count":79,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,2,28]]}},"alternative-id":["10.1145\/3785468"],"URL":"https:\/\/doi.org\/10.1145\/3785468","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,10]]},"assertion":[{"value":"2024-10-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}