{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T12:18:10Z","timestamp":1777551490371,"version":"3.51.4"},"reference-count":37,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,11,27]],"date-time":"2023-11-27T00:00:00Z","timestamp":1701043200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,11,27]],"date-time":"2023-11-27T00:00:00Z","timestamp":1701043200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Auton. Intell. Syst."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Current studies in few-shot semantic segmentation mostly utilize meta-learning frameworks to obtain models that can be generalized to new categories. However, these models trained on base classes with sufficient annotated samples are biased towards these base classes, which results in semantic confusion and ambiguity between base classes and new classes. A strategy is to use an additional base learner to recognize the objects of base classes and then refine the prediction results output by the meta learner. In this way, the interaction between these two learners and the way of combining results from the two learners are important. This paper proposes a new model, namely Distilling Base and Meta (DBAM) network by using self-attention mechanism and contrastive learning to enhance the few-shot segmentation performance. First, the self-attention-based ensemble module (SEM) is proposed to produce a more accurate adjustment factor for improving the fusion of two predictions of the two learners. Second, the prototype feature optimization module (PFOM) is proposed to provide an interaction between the two learners, which enhances the ability to distinguish the base classes from the target class by introducing contrastive learning loss. Extensive experiments have demonstrated that our method improves on the PASCAL-5<jats:sup><jats:italic>i<\/jats:italic><\/jats:sup> under 1-shot and 5-shot settings, respectively.<\/jats:p>","DOI":"10.1007\/s43684-023-00058-2","type":"journal-article","created":{"date-parts":[[2023,11,27]],"date-time":"2023-11-27T14:02:55Z","timestamp":1701093775000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Distilling base-and-meta network with contrastive learning for few-shot semantic segmentation"],"prefix":"10.1007","volume":"3","author":[{"given":"Xinyue","family":"Chen","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yueyi","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yingyue","family":"Xu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4933-0073","authenticated-orcid":false,"given":"Miaojing","family":"Shi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,11,27]]},"reference":[{"key":"58_CR1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.108164","volume":"120","author":"Z. Yang","year":"2021","unstructured":"Z. Yang, M. Shi, C. Xu, V. Ferrari, Y. Avrithis, Training object detectors from few weakly-labeled and many unlabeled images. Pattern Recognit. 120, 108164 (2021)","journal-title":"Pattern Recognit."},{"issue":"12","key":"58_CR2","doi-asserted-by":"publisher","first-page":"8586","DOI":"10.1109\/TCSVT.2022.3193612","volume":"32","author":"M. Zhang","year":"2022","unstructured":"M. Zhang, M. Shi, L. Li, Mfnet: multiclass few-shot segmentation network with pixel-wise metric learning. IEEE Trans. Circuits Syst. Video Technol. 32(12), 8586\u20138598 (2022)","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"58_CR3","doi-asserted-by":"crossref","unstructured":"Y. Du, M. Shi, F. Wei, G. Li, Boosting zero-shot learning via contrastive optimization of attribute representations (2022). arXiv preprint. arXiv:2207.03824","DOI":"10.1109\/TNNLS.2023.3297134"},{"key":"58_CR4","unstructured":"Z. Tian, H. Zhao, M. Shu, Z. Yang, R. Li, J. Jia, Prior guided feature enrichment network for few-shot segmentation. IEEE Trans. Pattern Anal. Mach. Intell. (2020). arXiv preprint. arXiv:2008.01449"},{"key":"58_CR5","first-page":"9197","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"K. Wang","year":"2019","unstructured":"K. Wang, J.H. Liew, Y. Zou, D. Zhou, J. Feng, Panet: few-shot image semantic segmentation with prototype alignment, in Proceedings of the IEEE\/CVF International Conference on Computer Vision (2019), pp. 9197\u20139206"},{"key":"58_CR6","first-page":"5217","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"C. Zhang","year":"2019","unstructured":"C. Zhang, G. Lin, F. Liu, R. Yao, C. Shen, Canet: class-agnostic segmentation networks with iterative refinement and attentive few-shot learning, in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (2019), pp. 5217\u20135226"},{"key":"58_CR7","first-page":"8334","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"G. Li","year":"2021","unstructured":"G. Li, V. Jampani, L. Sevilla-Lara, D. Sun, J. Kim, J. Kim, Adaptive prototype learning and allocation for few-shot segmentation, in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (2021), pp. 8334\u20138343"},{"key":"58_CR8","first-page":"11573","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Y. Liu","year":"2022","unstructured":"Y. Liu, N. Liu, Q. Cao, X. Yao, J. Han, L. Shao, Learning non-target knowledge for few-shot semantic segmentation, in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (2022), pp. 11573\u201311582"},{"key":"58_CR9","first-page":"9587","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"C. Zhang","year":"2019","unstructured":"C. Zhang, G. Lin, F. Liu, J. Guo, Q. Wu, R. Yao, Pyramid graph networks with connection attentions for region-based one-shot semantic segmentation, in Proceedings of the IEEE\/CVF International Conference on Computer Vision (2019), pp. 9587\u20139595"},{"key":"58_CR10","first-page":"142","volume-title":"European Conference on Computer Vision","author":"Y. Liu","year":"2020","unstructured":"Y. Liu, X. Zhang, S. Zhang, X. He, Part-aware prototype network for few-shot semantic segmentation, in European Conference on Computer Vision (Springer, Berlin, 2020), pp. 142\u2013158"},{"key":"58_CR11","first-page":"8057","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"C. Lang","year":"2022","unstructured":"C. Lang, G. Cheng, B. Tu, J. Han, Learning what not to segment: a new perspective on few-shot segmentation, in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (2022), pp. 8057\u20138067"},{"key":"58_CR12","first-page":"3431","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"J. Long","year":"2015","unstructured":"J. Long, E. Shelhamer, T. Darrell, Fully convolutional networks for semantic segmentation, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2015), pp. 3431\u20133440"},{"key":"58_CR13","first-page":"234","volume-title":"International Conference on Medical Image Computing and Computer-Assisted Intervention","author":"O. Ronneberger","year":"2015","unstructured":"O. Ronneberger, P. Fischer, T. Brox, U-net: convolutional networks for biomedical image segmentation, in International Conference on Medical Image Computing and Computer-Assisted Intervention (Springer, Berlin, 2015), pp. 234\u2013241"},{"key":"58_CR14","unstructured":"F. Yu, V. Koltun, Multi-scale context aggregation by dilated convolutions (2015). arXiv preprint. arXiv:1511.07122"},{"key":"58_CR15","first-page":"2881","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"H. Zhao","year":"2017","unstructured":"H. Zhao, J. Shi, X. Qi, X. Wang, J. Jia, Pyramid scene parsing network, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017), pp. 2881\u20132890"},{"issue":"4","key":"58_CR16","doi-asserted-by":"publisher","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","volume":"40","author":"L.-C. Chen","year":"2017","unstructured":"L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, A.L. Yuille, Deeplab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans. Pattern Anal. Mach. Intell. 40(4), 834\u2013848 (2017)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"58_CR17","first-page":"9258","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Y. Lifchitz","year":"2019","unstructured":"Y. Lifchitz, Y. Avrithis, S. Picard, A. Bursuc, Dense classification and implanting for few-shot learning, in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (2019), pp. 9258\u20139267"},{"key":"58_CR18","first-page":"6372","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"P. Tokmakov","year":"2019","unstructured":"P. Tokmakov, Y.-X. Wang, M. Hebert, Learning compositional representations for few-shot recognition, in Proceedings of the IEEE\/CVF International Conference on Computer Vision (2019), pp. 6372\u20136381"},{"issue":"3","key":"58_CR19","doi-asserted-by":"publisher","first-page":"1091","DOI":"10.1109\/TCSVT.2020.2995754","volume":"31","author":"W. Jiang","year":"2020","unstructured":"W. Jiang, K. Huang, J. Geng, X. Deng, Multi-scale metric learning for few-shot learning. IEEE Trans. Circuits Syst. Video Technol. 31(3), 1091\u20131102 (2020)","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"58_CR20","first-page":"3521","volume":"33","author":"Y. Yang","year":"2020","unstructured":"Y. Yang, F. Wei, M. Shi, G. Li, Restoring negative information in few-shot object detection. Adv. Neural Inf. Process. Syst. 33, 3521\u20133532 (2020)","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"58_CR21","first-page":"14084","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Y. Du","year":"2022","unstructured":"Y. Du, F. Wei, Z. Zhang, M. Shi, Y. Gao, G. Li, Learning to prompt for open-vocabulary object detection with vision-language model, in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (2022), pp. 14084\u201314093"},{"key":"58_CR22","unstructured":"O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra et\u00a0al., Matching networks for one shot learning. Adv. Neural Inf. Process. Syst. 29 (2016). arXiv preprint. arXiv:1606.04080"},{"key":"58_CR23","first-page":"1842","volume-title":"International Conference on Machine Learning","author":"A. Santoro","year":"2016","unstructured":"A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, T. Lillicrap, Meta-learning with memory-augmented neural networks, in International Conference on Machine Learning (2016), pp. 1842\u20131850. PMLR"},{"key":"58_CR24","unstructured":"J. Snell, K. Swersky, R. Zemel, Prototypical networks for few-shot learning. Adv. Neural Inf. Process. Syst. 30 (2017). arXiv preprint. arXiv:1703.05175"},{"key":"58_CR25","first-page":"1199","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"F. Sung","year":"2018","unstructured":"F. Sung, Y. Yang, L. Zhang, T. Xiang, P.H. Torr, T.M. Hospedales, Learning to compare: relation network for few-shot learning, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2018), pp. 1199\u20131208"},{"key":"58_CR26","first-page":"1126","volume-title":"International Conference on Machine Learning","author":"C. Finn","year":"2017","unstructured":"C. Finn, P. Abbeel, S. Levine, Model-agnostic meta-learning for fast adaptation of deep networks, in International Conference on Machine Learning (2017), pp. 1126\u20131135. PMLR"},{"key":"58_CR27","unstructured":"A. Nichol, J. Schulman, Reptile: a scalable metalearning algorithm, (2018). arXiv preprint. arXiv:1803.02999"},{"key":"58_CR28","unstructured":"A.A. Rusu, D. Rao, J. Sygnowski, O. Vinyals, R. Pascanu, S. Osindero, R. Hadsell, Meta-learning with latent embedding optimization (2018). arXiv preprint. arXiv:1807.05960"},{"key":"58_CR29","first-page":"10657","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"K. Lee","year":"2019","unstructured":"K. Lee, S. Maji, A. Ravichandran, S. Soatto, Meta-learning with differentiable convex optimization, in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (2019), pp. 10657\u201310665"},{"key":"58_CR30","doi-asserted-by":"crossref","unstructured":"A. Shaban, S. Bansal, Z. Liu, I. Essa, B. Boots, One-shot learning for semantic segmentation (2017). arXiv preprint. arXiv:1709.03410","DOI":"10.5244\/C.31.167"},{"key":"58_CR31","first-page":"1597","volume-title":"International Conference on Machine Learning","author":"T. Chen","year":"2020","unstructured":"T. Chen, S. Kornblith, M. Norouzi, G. Hinton, A simple framework for contrastive learning of visual representations, in International Conference on Machine Learning (2020), pp. 1597\u20131607. PMLR"},{"key":"58_CR32","first-page":"7303","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"W. Wang","year":"2021","unstructured":"W. Wang, T. Zhou, F. Yu, J. Dai, E. Konukoglu, L. Van Gool, Exploring cross-image pixel contrast for semantic segmentation, in Proceedings of the IEEE\/CVF International Conference on Computer Vision (2021), pp. 7303\u20137313"},{"key":"58_CR33","first-page":"20447","volume":"34","author":"M. Agarwal","year":"2021","unstructured":"M. Agarwal, M. Yurochkin, Y. Sun, On sensitivity of meta-learning to support data. Adv. Neural Inf. Process. Syst. 34, 20447\u201320460 (2021)","journal-title":"Adv. Neural Inf. Process. Syst."},{"issue":"2","key":"58_CR34","doi-asserted-by":"publisher","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","volume":"88","author":"M. Everingham","year":"2010","unstructured":"M. Everingham, L. Van Gool, C.K. Williams, J. Winn, A. Zisserman, The Pascal visual object classes (voc) challenge. Int. J. Comput. Vis. 88(2), 303\u2013338 (2010)","journal-title":"Int. J. Comput. Vis."},{"key":"58_CR35","doi-asserted-by":"publisher","first-page":"991","DOI":"10.1109\/ICCV.2011.6126343","volume-title":"2011 International Conference on Computer Vision","author":"B. Hariharan","year":"2011","unstructured":"B. Hariharan, P. Arbel\u00e1ez, L. Bourdev, S. Maji, J. Malik, Semantic contours from inverse detectors, in 2011 International Conference on Computer Vision (IEEE Press, New York, 2011), pp. 991\u2013998"},{"issue":"9","key":"58_CR36","doi-asserted-by":"publisher","first-page":"3855","DOI":"10.1109\/TCYB.2020.2992433","volume":"50","author":"X. Zhang","year":"2020","unstructured":"X. Zhang, Y. Wei, Y. Yang, T.S. Huang, Sg-one: similarity guidance network for one-shot semantic segmentation. IEEE Trans. Cybern. 50(9), 3855\u20133865 (2020)","journal-title":"IEEE Trans. Cybern."},{"key":"58_CR37","first-page":"770","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"K. He","year":"2016","unstructured":"K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2016), pp. 770\u2013778"}],"container-title":["Autonomous Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s43684-023-00058-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s43684-023-00058-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s43684-023-00058-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,27]],"date-time":"2023-11-27T14:05:33Z","timestamp":1701093933000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s43684-023-00058-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,27]]},"references-count":37,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["58"],"URL":"https:\/\/doi.org\/10.1007\/s43684-023-00058-2","relation":{},"ISSN":["2730-616X"],"issn-type":[{"value":"2730-616X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,27]]},"assertion":[{"value":"17 October 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"31 August 2023","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 November 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 November 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Prof. Miaojing Shi is an editorial board member for Autonomous Intelligent Systems and was not involved in the editorial review, or the decision to publish, this article. All authors declare that there are no other competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"11"}}