{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T06:36:42Z","timestamp":1768286202403,"version":"3.49.0"},"reference-count":56,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2021,5,17]],"date-time":"2021-05-17T00:00:00Z","timestamp":1621209600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Natural Science Foundation of China","award":["62076099"],"award-info":[{"award-number":["62076099"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Weakly supervised instance segmentation (WSIS) provides a promising way to address instance segmentation in the absence of sufficient labeled data for training. Previous attempts on WSIS usually follow a proposal-based paradigm, critical to which is the proposal scoring strategy. These works mostly rely on certain heuristic strategies for proposal scoring, which largely hampers the sustainable advances concerning WSIS. Towards this end, this paper introduces a novel framework for weakly supervised instance segmentation, called Weakly Supervised R-CNN (WS-RCNN). The basic idea is to deploy a deep network to learn to score proposals, under the special setting of weak supervision. To tackle the key issue of acquiring proposal-level pseudo labels for model training, we propose a so-called Attention-Guided Pseudo Labeling (AGPL) strategy, which leverages the local maximal (peaks) in image-level attention maps and the spatial relationship among peaks and proposals to infer pseudo labels. We also suggest a novel training loss, called Entropic OpenSet Loss, to handle background proposals more effectively so as to further improve the robustness. Comprehensive experiments on two standard benchmarking datasets demonstrate that the proposed WS-RCNN can outperform the state-of-the-art by a large margin, with an improvement of 11.6% on PASCAL VOC 2012 and 10.7% on MS COCO 2014 in terms of mAP50, which indicates that learning-based proposal scoring and the proposed WS-RCNN framework might be a promising way towards WSIS.<\/jats:p>","DOI":"10.3390\/s21103475","type":"journal-article","created":{"date-parts":[[2021,5,17]],"date-time":"2021-05-17T04:25:07Z","timestamp":1621225507000},"page":"3475","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["WS-RCNN: Learning to Score Proposals for Weakly Supervised Instance Segmentation"],"prefix":"10.3390","volume":"21","author":[{"given":"Jia-Rong","family":"Ou","sequence":"first","affiliation":[{"name":"School of Automation Science and Engineering, South China University of Technology, Guangzhou 510641, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shu-Le","family":"Deng","sequence":"additional","affiliation":[{"name":"School of Automation Science and Engineering, South China University of Technology, Guangzhou 510641, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jin-Gang","family":"Yu","sequence":"additional","affiliation":[{"name":"School of Automation Science and Engineering, South China University of Technology, Guangzhou 510641, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,5,17]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Hariharan, B., Arbel\u00e1ez, P., Girshick, R., and Malik, J. (2014, January 6\u201312). Simultaneous detection and segmentation. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10584-0_20"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201323). Path aggregation network for instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Chen, L.C., Hermans, A., Papandreou, G., Schroff, F., Wang, P., and Adam, H. (2018, January 18\u201323). Masklab: Instance segmentation by refining object detection with semantic and direction features. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00422"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Chen, K., Pang, J., Wang, J., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., Shi, J., and Ouyang, W. (2019, January 16\u201320). Hybrid task cascade for instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00511"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Jeong, D., Kim, B.G., and Dong, S.Y. (2020). Deep joint spatiotemporal network (DJSTN) for efficient facial expression recognition. Sensors, 20.","DOI":"10.3390\/s20071936"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zhou, Y., Zhu, Y., Ye, Q., Qiu, Q., and Jiao, J. (2018, January 18\u201323). Weakly supervised instance segmentation using class peak response. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00399"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Ge, W., Guo, S., Huang, W., and Scott, M.R. (2019, January 16\u201320). Label-PEnet: Sequential label propagation and enhancement Networks for Weakly Supervised Instance Segmentation. Proceedings of the IEEE International Conference on Computer Vision, Long Beach, CA, USA.","DOI":"10.1109\/ICCV.2019.00344"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Khoreva, A., Benenson, R., Hosang, J., Hein, M., and Schiele, B. (2017, January 22\u201329). Simple does it: Weakly supervised instance and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Venice, Italy.","DOI":"10.1109\/CVPR.2017.181"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Remez, T., Huang, J., and Brown, M. (2018, January 18\u201323). Learning to segment via cut-and-paste. Proceedings of the European Conference on Computer Vision, Salt Lake City, UT, USA.","DOI":"10.1007\/978-3-030-01234-2_3"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Kuo, W., Angelova, A., Malik, J., and Lin, T.Y. (2019, January 16\u201320). Shapemask: Learning to segment novel objects by refining shape priors. Proceedings of the IEEE International Conference on Computer Vision, Long Beach, CA, USA.","DOI":"10.1109\/ICCV.2019.00930"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zhu, Y., Zhou, Y., Xu, H., Ye, Q., Doermann, D., and Jiao, J. (2019, January 16\u201320). Learning instance activation maps for weakly supervised instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00323"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Ahn, J., Cho, S., and Kwak, S. (2019, January 16\u201320). Weakly supervised learning of instance segmentation with inter-pixel relations. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00231"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Arun, A., Jawahar, C., and Kumar, M.P. (2020, January 23\u201328). Weakly supervised instance segmentation by learning annotation consistent instances. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58604-1_16"},{"key":"ref_16","unstructured":"Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Torralba, A. (2014). Object detectors emerge in deep scene cnns. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1084","DOI":"10.1007\/s11263-017-1059-x","article-title":"Top-Down Neural Attention by Excitation Backprop","volume":"126","author":"Zhang","year":"2018","journal-title":"Int. J. Comput. Vis."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Torralba, A. (2016, January 27\u201330). Learning deep features for discriminative localization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.319"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zeiler, M.D., and Fergus, R. (2014, January 6\u201312). Visualizing and understanding convolutional networks. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10590-1_53"},{"key":"ref_20","unstructured":"Simonyan, K., Vedaldi, A., and Zisserman, A. (2013). Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 13\u201316). Fast r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 6\u201312). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Zurich, Switzerland.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_23","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015, January 7\u201312). Faster r-cnn: Towards real-time object detection with region proposal networks. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Geng, C., Huang, S.j., and Chen, S. (2020). Recent advances in open set recognition: A survey. IEEE Trans. Pattern Anal. Mach. Intell.","DOI":"10.1109\/TPAMI.2020.2981604"},{"key":"ref_25","unstructured":"Dhamija, A.R., G\u00fcnther, M., and Boult, T. (2018, January 3\u20138). Reducing network agnostophobia. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Laradji, I.H., Vazquez, D., and Schmidt, M. (2019). Where are the Masks: Instance Segmentation with Image-level Supervision. arXiv.","DOI":"10.1109\/ICIP40778.2020.9190782"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 13\u201316). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Santiago, Chile.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Kolesnikov, A., and Lampert, C.H. (2016, January 8\u201316). Seed, expand and constrain: Three principles for weakly-supervised image segmentation. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46493-0_42"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Roy, A., and Todorovic, S. (2017, January 21\u201326). Combining bottom-up, top-down, and smoothness cues for weakly supervised image segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.770"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Pinheiro, P.O., and Collobert, R. (2015, January 7\u201312). From image-level to pixel-level labeling with convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298780"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Papandreou, G., Chen, L.C., Murphy, K., and Yuille, A. (2015). Weakly-and semi-supervised learning of a DCNN for semantic image segmentation. arXiv.","DOI":"10.1109\/ICCV.2015.203"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Huang, Z., Wang, X., Wang, J., Liu, W., and Wang, J. (2018, January 18\u201323). Weakly-supervised semantic segmentation network with deep seeded region growing. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00733"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Ahn, J., and Kwak, S. (2018, January 18\u201323). Learning pixel-level semantic affinity with image-level supervision for weakly supervised semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00523"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wei, Y., Xiao, H., Shi, H., Jie, Z., Feng, J., and Huang, T.S. (2018, January 18\u201323). Revisiting dilated convolution: A simple approach for weakly-and semi-supervised semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00759"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Zhang, X., Wei, Y., Feng, J., Yang, Y., and Huang, T.S. (2018, January 18\u201323). Adversarial complementary learning for weakly supervised object localization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00144"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Li, K., Wu, Z., Peng, K.C., Ernst, J., and Fu, Y. (2018, January 18\u201323). Tell me where to look: Guided attention inference network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00960"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Wang, X., You, S., Li, X., and Ma, H. (2018, January 18\u201323). Weakly-supervised semantic segmentation by iteratively mining common object features. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00147"},{"key":"ref_38","unstructured":"Hou, Q., Jiang, P., Wei, Y., and Cheng, M.M. (2018, January 3\u20138). Self-erasing network for integral object attention. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Fan, R., Hou, Q., Cheng, M.M., Yu, G., Martin, R.R., and Hu, S.M. (2018, January 18\u201323). Associating inter-image salient instances for weakly supervised semantic segmentation. Proceedings of the European Conference on Computer Vision, Salt Lake City, UT, USA.","DOI":"10.1007\/978-3-030-01240-3_23"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Fan, R., Cheng, M.M., Hou, Q., Mu, T.J., Wang, J., and Hu, S.M. (2019, January 16\u201320). S4Net: Single stage salient-instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00626"},{"key":"ref_41","unstructured":"Zhang, C., Platt, J.C., and Viola, P.A. (2006, January 4\u20137). Multiple instance boosting for object detection. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1109\/TPAMI.2015.2456908","article-title":"Weakly supervised large scale object localization with multiple instance learning and bag splitting","volume":"38","author":"Ren","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Wang, X., Zhu, Z., Yao, C., and Bai, X. (2015, January 13\u201326). Relaxed multiple-instance SVM with application to object discovery. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.145"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Bilen, H., and Vedaldi, A. (2016, January 27\u201330). Weakly supervised deep detection networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.311"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Diba, A., Sharma, V., Pazandeh, A., Pirsiavash, H., and Van Gool, L. (2017, January 21\u201326). Weakly supervised cascaded convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.545"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Tang, P., Wang, X., Wang, A., Yan, Y., Liu, W., Huang, J., and Yuille, A. (2018, January 8\u201314). Weakly supervised region proposal network and object detection. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01252-6_22"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"176","DOI":"10.1109\/TPAMI.2018.2876304","article-title":"Pcl: Proposal cluster learning for weakly supervised object detection","volume":"42","author":"Tang","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"2395","DOI":"10.1109\/TPAMI.2019.2898858","article-title":"Min-Entropy Latent Model for Weakly Supervised Object Detection","volume":"41","author":"Wan","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"843","DOI":"10.1109\/TIP.2019.2933735","article-title":"Category-Aware spatial constraint for weakly supervised detection","volume":"29","author":"Shen","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Bendale, A., and Boult, T.E. (2016, January 27\u201330). Towards open set deep networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.173"},{"key":"ref_51","first-page":"128","article-title":"Multiscale combinatorial grouping for image segmentation and object proposal generation","volume":"39","author":"Arbelaez","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"819","DOI":"10.1109\/TPAMI.2017.2700300","article-title":"Convolutional oriented boundaries: From image segmentation to high-level tasks","volume":"40","author":"Maninis","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"98","DOI":"10.1007\/s11263-014-0733-5","article-title":"The pascal visual object classes challenge: A retrospective","volume":"111","author":"Everingham","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Hariharan, B., Arbel\u00e1ez, P., Bourdev, L., Maji, S., and Malik, J. (2011, January 6\u201313). Semantic contours from inverse detectors. Proceedings of the IEEE International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126343"},{"key":"ref_56","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/10\/3475\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:02:25Z","timestamp":1760162545000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/10\/3475"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,17]]},"references-count":56,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2021,5]]}},"alternative-id":["s21103475"],"URL":"https:\/\/doi.org\/10.3390\/s21103475","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,5,17]]}}}