{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,3]],"date-time":"2026-02-03T21:43:24Z","timestamp":1770155004115,"version":"3.49.0"},"reference-count":64,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2026,2,3]],"date-time":"2026-02-03T00:00:00Z","timestamp":1770076800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,2,3]],"date-time":"2026-02-03T00:00:00Z","timestamp":1770076800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"other","award":["2024-0011 (ZX20240867)"],"award-info":[{"award-number":["2024-0011 (ZX20240867)"]}]},{"name":"other","award":["2024-0011 (ZX20240867)"],"award-info":[{"award-number":["2024-0011 (ZX20240867)"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Intell."],"published-print":{"date-parts":[[2026,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Semi-supervised semantic segmentation, which aims to utilize large volumes of unlabeled images to achieve accurate segmentation results with fewer human annotations, has attracted increasing attention. Prior methods were primarily based on the assumption that the pseudo-labels generated by each branch were sufficiently reliable to supervise the remaining branches. However, there is a concern that the errors in pseudo-labels could accumulate during co-training, potentially resulting in suboptimal performance. To this end, we present a novel framework called heterogeneous dual-branch voting supervision (HDVS), which is designed to enhance the reliability of pseudo-labels and mitigate the issues arising from pseudo-labeling. Specifically, based on the cross-supervision framework, we introduce a voting mechanism that correlates pseudo-labels from heterogeneous branches to produce pseudo-labels with enhanced reliability. Concurrently, we employ a feature communication module to introduce perturbations at the feature level in each branch to maximize prediction diversity across the dual branches while maintaining convergence. Comprehensive evaluations of the proposed HDVS on two benchmark datasets (PASCAL VOC 2012 and Cityscapes) demonstrate its superiority to the state-of-the-art approaches.<\/jats:p>","DOI":"10.1007\/s44267-026-00107-3","type":"journal-article","created":{"date-parts":[[2026,2,3]],"date-time":"2026-02-03T07:38:49Z","timestamp":1770104329000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["HDVS: semi-supervised semantic segmentation via heterogeneous dual-branch voting supervision"],"prefix":"10.1007","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-3469-5231","authenticated-orcid":false,"given":"Yongqi","family":"Shan","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4288-4516","authenticated-orcid":false,"given":"Yunzhi","family":"Zhuge","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6668-9758","authenticated-orcid":false,"given":"Huchuan","family":"Lu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,2,3]]},"reference":[{"key":"107_CR1","first-page":"1","volume-title":"Proceedings of the 9th international conference on learning representations","author":"A. Dosovitskiy","year":"2021","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2021). An image is worth 16x16 words: transformers for image recognition at scale. In Proceedings of the 9th international conference on learning representations (pp. 1\u201321). Retrieved December 27, 2025, from https:\/\/openreview.net\/forum?id=YicbFdNTTy."},{"key":"107_CR2","first-page":"12077","volume-title":"Proceedings of the 35th international conference on neural information processing systems","author":"E. Xie","year":"2021","unstructured":"Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J. M., & Luo, P. (2021). SegFormer: simple and efficient design for semantic segmentation with transformers. In Proceedings of the 35th international conference on neural information processing systems (pp. 12077\u201312090). Red Hook: Curran Associates."},{"key":"107_CR3","doi-asserted-by":"publisher","first-page":"415","DOI":"10.1007\/s41095-022-0274-8","volume":"8","author":"W. Wang","year":"2022","unstructured":"Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., & Shao, L. (2022). PVT v2: improved baselines with pyramid vision transformer. Computational Visual Media, 8, 415\u2013424.","journal-title":"Computational Visual Media"},{"key":"107_CR4","first-page":"4015","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"A. Kirillov","year":"2023","unstructured":"Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al. (2023). Segment anything. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 4015\u20134026). Piscataway: IEEE."},{"key":"107_CR5","first-page":"19529","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"J. Xu","year":"2023","unstructured":"Xu, J., Xiong, Z., & Bhattacharyya, S. P. (2023). PIDNet: a real-time semantic segmentation network inspired by PID controllers. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 19529\u201319539). Piscataway: IEEE."},{"key":"107_CR6","first-page":"3041","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"F. Li","year":"2023","unstructured":"Li, F., Zhang, H., Xu, H., Liu, S., Zhang, L., Ni, L. M., & Shum, H.-Y. (2023). Mask DINO: towards a unified transformer-based framework for object detection and segmentation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 3041\u20133050). Piscataway: IEEE."},{"key":"107_CR7","doi-asserted-by":"publisher","first-page":"11400","DOI":"10.1109\/TPAMI.2025.3600507","volume":"47","author":"H. Ding","year":"2025","unstructured":"Ding, H., Liu, C., He, S., Ying, K., Jiang, X., Loy, C. C., & Jiang, Y.-G. (2025). MeViS: a multi-modal dataset for referring motion expression video segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 47, 11400\u201311416.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"107_CR8","unstructured":"Ding, H., Ying, K., Liu, C., He, S., Jiang, X., Jiang, Y.-G., Torr, P. H., & Bai, S. (2025). MOSEv2: a more challenging dataset for video object segmentation in complex scenes. arXiv preprint. arXiv:2508.05630."},{"key":"107_CR9","first-page":"20224","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"H. Ding","year":"2023","unstructured":"Ding, H., Liu, C., He, S., Jiang, X., Torr, P. H., & Bai, S. (2023). MOSE: a new dataset for video object segmentation in complex scenes. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 20224\u201320234). Piscataway: IEEE."},{"key":"107_CR10","first-page":"2694","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"H. Ding","year":"2023","unstructured":"Ding, H., Liu, C., He, S., Jiang, X., & Loy, C. C. (2023). MeViS: a large-scale benchmark for video segmentation with motion expressions. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 2694\u20132703). Piscataway: IEEE."},{"key":"107_CR11","first-page":"5688","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"N. Souly","year":"2017","unstructured":"Souly, N., Spampinato, C., & Shah, M. (2017). Semi-supervised semantic segmentation using generative adversarial network. In Proceedings of the IEEE international conference on computer vision (pp. 5688\u20135696). Piscataway: IEEE."},{"key":"107_CR12","first-page":"596","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"K. Sohn","year":"2020","unstructured":"Sohn, K., Berthelot, D., Carlini, N., Zhang, Z., Zhang, H., Raffel, C. A., Cubuk, E. D., Kurakin, A., & Li, C.-L. (2020). FixMatch: simplifying semi-supervised learning with consistency and confidence. In H. Larochelle, M. Ranzato, R. Hadsell, M.-F. Balcan, & H.-T. Lin (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 596\u2013608). Red Hook: Curran Associates."},{"key":"107_CR13","first-page":"4248","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Wang","year":"2022","unstructured":"Wang, Y., Wang, H., Shen, Y., Fei, J., Li, W., Jin, G., Wu, L., Zhao, R., & Le, X. (2022). Semi-supervised semantic segmentation using unreliable pseudo-labels. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 4248\u20134257). Piscataway: IEEE."},{"key":"107_CR14","first-page":"2613","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"X. Chen","year":"2021","unstructured":"Chen, X., Yuan, Y., Zeng, G., & Wang, J. (2021). Semi-supervised semantic segmentation with cross pseudo supervision. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 2613\u20132622). Piscataway: IEEE."},{"key":"107_CR15","first-page":"16009","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"Y. Li","year":"2023","unstructured":"Li, Y., Wang, X., Yang, L., Feng, L., Zhang, W., & Gao, Y. (2023). Diverse cotraining makes strong semi-supervised segmentor. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 16009\u201316021). Piscataway: IEEE"},{"key":"107_CR16","first-page":"1","volume-title":"Proceedings of the 8th international conference on learning representations","author":"D. Berthelot","year":"2020","unstructured":"Berthelot, D., Carlini, N., Cubuk, E. D., Kurakin, A., Sohn, K., Zhang, H., & Raffel, C. (2020). ReMixMatch: semi-supervised learning with distribution matching and augmentation anchoring. In Proceedings of the 8th international conference on learning representations (pp. 1\u201313). Retrieved December 27, 2025, from https:\/\/openreview.net\/forum?id=HklkeR4KPB."},{"key":"107_CR17","first-page":"5026","volume-title":"Proceedings of the 33rd international conference on neural information processing systems","author":"D. Berthelot","year":"2019","unstructured":"Berthelot, D., Carlini, N., Goodfellow, I., Papernot, N., Oliver, A., & Raffel, C. A. (2019). MixMatch: a holistic approach to semi-supervised learning. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d\u2019Alch\u00e9 Buc, E. Fox, & R. Garnett (Eds.), Proceedings of the 33rd international conference on neural information processing systems (pp. 5026\u20135036). Red Hook: Curran Associates."},{"key":"107_CR18","first-page":"1195","volume-title":"Proceedings of the 31st international conference on neural information processing systems","author":"A. Tarvainen","year":"2017","unstructured":"Tarvainen, A., & Valpola, H. (2017). Mean teachers are better role models: weight-averaged consistency targets improve semi-supervised deep learning results. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. V. N. Vishwanathan, & R. Garnett (Eds.), Proceedings of the 31st international conference on neural information processing systems (pp. 1195\u20131204). Red Hook: Curran Associates."},{"key":"107_CR19","first-page":"1","volume-title":"Proceedings of the international joint conference on neural networks","author":"E. Arazo","year":"2020","unstructured":"Arazo, E., Ortego, D., Albert, P., O\u2019Connor, N. E., & McGuinness, K. (2020). Pseudo-labeling and confirmation bias in deep semi-supervised learning. In Proceedings of the international joint conference on neural networks (pp. 1\u20138). Piscataway: IEEE."},{"key":"107_CR20","unstructured":"Ouali, Y., Hudelot, C., & Tami, M. (2020). An overview of deep semi-supervised learning. arXiv preprint. arXiv:2006.05278."},{"key":"107_CR21","first-page":"9947","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"J. Fan","year":"2022","unstructured":"Fan, J., Gao, B., Jin, H., & Jiang, L. (2022). UCC: uncertainty guided cross-head co-training for semi-supervised semantic segmentation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 9947\u20139956). Piscataway: IEEE."},{"key":"107_CR22","first-page":"16348","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"S. Li","year":"2023","unstructured":"Li, S., He, Y., Zhang, W., Zhang, W., Tan, X., Han, J., Ding, E., & Wang, J. (2023). CFCG: semi-supervised semantic segmentation via cross-fusion and contour guidance supervision. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 16348\u201316358). Piscataway: IEEE."},{"key":"107_CR23","first-page":"529","volume-title":"Proceedings of the 18th international conference on neural information processing systems","author":"Y. Grandvalet","year":"2004","unstructured":"Grandvalet, Y., & Bengio, Y. (2004). Semi-supervised learning by entropy minimization. In Proceedings of the 18th international conference on neural information processing systems (pp. 529\u2013536). Red Hook: Curran Associates."},{"key":"107_CR24","doi-asserted-by":"publisher","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","volume":"88","author":"M. Everingham","year":"2010","unstructured":"Everingham, M., Van Gool, L., Williams, C. K., Winn, J., & Zisserman, A. (2010). The PASCAL visual object classes (VOC) challenge. International Journal of Computer Vision, 88, 303\u2013338.","journal-title":"International Journal of Computer Vision"},{"key":"107_CR25","first-page":"3213","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"M. Cordts","year":"2016","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., & Schiele, B. (2016). The Cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3213\u20133223). Piscataway: IEEE."},{"issue":"1","key":"107_CR26","doi-asserted-by":"publisher","first-page":"83","DOI":"10.1007\/s41095-016-0073-1","volume":"3","author":"L. Chen","year":"2017","unstructured":"Chen, L., & Yang, M. (2017). Semi-supervised dictionary learning with label propagation for image classification. Computational Visual Media, 3(1), 83\u201394.","journal-title":"Computational Visual Media"},{"key":"107_CR27","first-page":"18408","volume-title":"Proceedings of the 35th international conference on neural information processing systems","author":"B. Zhang","year":"2021","unstructured":"Zhang, B., Wang, Y., Hou, W., Wu, H., Wang, J., Okumura, M., & Shinozaki, T. (2021). Flexmatch: boosting semi-supervised learning with curriculum pseudo labeling. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P. S. Liang, & J. W. Vaughan (Eds.), Proceedings of the 35th international conference on neural information processing systems (pp. 18408\u201318419). Red Hook: Curran Associates."},{"key":"107_CR28","first-page":"1","volume-title":"Proceedings of the 11th international conference on learning representations","author":"Y. Wang","year":"2023","unstructured":"Wang, Y., Chen, H., Heng, Q., Hou, W., Fan, Y., Wu, Z., Wang, J., Savvides, M., Shinozaki, T., Raj, B., et al. (2023). Freematch: self-adaptive thresholding for semi-supervised learning. In Proceedings of the 11th international conference on learning representations (pp. 1\u201320). Retrieved December 27, 2025, from https:\/\/openreview.net\/forum?id=PDrUPTXJI_A."},{"issue":"2","key":"107_CR29","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1007\/s41095-022-0281-9","volume":"9","author":"C.-Y. Sun","year":"2023","unstructured":"Sun, C.-Y., Yang, Y.-Q., Guo, H.-X., Wang, P.-S., Tong, X., Liu, Y., & Shum, H.-Y. (2023). Semi-supervised 3D shape segmentation with multilevel consistency and part substitution. Computational Visual Media, 9(2), 229\u2013247.","journal-title":"Computational Visual Media"},{"key":"107_CR30","first-page":"3239","volume-title":"Proceedings of the 32nd international conference on neural information processing systems","author":"A. Oliver","year":"2018","unstructured":"Oliver, A., Odena, A., Raffel, C. A., Cubuk, E. D., & Goodfellow, I. (2018). Realistic evaluation of deep semi-supervised learning algorithms. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, & R. Garnett (Eds.), Proceedings of the 32nd international conference on neural information processing systems (pp. 3239\u20133250). Red Hook: Curran Associates."},{"key":"107_CR31","unstructured":"Ding, H., Tang, S., He, S., Liu, C., Wu, Z., & Jiang, Y.-G. (2025). Multimodal referring segmentation: a survey. arXiv preprint. arXiv:2508.00265."},{"key":"107_CR32","first-page":"896","volume-title":"Proceedings of the 30th international conference on machine learning workshop on challenges in representation learning","author":"D.-H. Lee","year":"2013","unstructured":"Lee, D.-H. (2013). Pseudo-label: the simple and efficient semi-supervised learning method for deep neural networks. In Proceedings of the 30th international conference on machine learning workshop on challenges in representation learning (pp. 896\u2013901). Retrieved December 27, 2025, from https:\/\/www.kaggle.com\/blobs\/download\/forum-message-attachment-files\/746\/pseudo_label_final.pdf."},{"key":"107_CR33","first-page":"10687","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Q. Xie","year":"2020","unstructured":"Xie, Q., Luong, M.-T., Hovy, E., & Le, Q. V. (2020). Self-training with noisy student improves ImageNet classification. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 10687\u201310698). Piscataway: IEEE."},{"key":"107_CR34","first-page":"3833","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"B. Zoph","year":"2020","unstructured":"Zoph, B., Ghiasi, G., Lin, T.-Y., Cui, Y., Liu, H., Cubuk, E. D., & Le, Q. (2020). Rethinking pre-training and self-training. In H. Larochelle, M. Ranzato, R. Hadsell, M.-F. Balcan, & H.-T. Lin (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 3833\u20133845). Red Hook: Curran Associates."},{"key":"107_CR35","first-page":"11557","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"H. Pham","year":"2021","unstructured":"Pham, H., Dai, Z., Xie, Q., & Le, Q. V. (2021). Meta pseudo labels. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 11557\u201311568). Piscataway: IEEE."},{"key":"107_CR36","first-page":"663","volume-title":"Proceedings of the 32nd international joint conference on artificial intelligence","author":"C. Ding","year":"2023","unstructured":"Ding, C., Zhang, J., Ding, H., Zhao, H., Wang, Z., Xing, T., & Hu, R. (2023). Decoupling with entropy-based equalization for semi-supervised semantic segmentation. In Proceedings of the 32nd international joint conference on artificial intelligence (pp. 663\u2013671). Cham: Springer."},{"key":"107_CR37","first-page":"10759","volume-title":"Proceedings of the 33rd international conference on neural information processing systems","author":"J. Jeong","year":"2019","unstructured":"Jeong, J., Lee, S., Kim, J., & Kwak, N. (2019). Consistency-based semi-supervised learning for object detection. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d\u2019Alch\u00e9 Buc, E. Fox, & R. Garnett (Eds.), Proceedings of the 33rd international conference on neural information processing systems (pp. 10759\u201310768). Red Hook: Curran Associates."},{"key":"107_CR38","first-page":"1","volume-title":"Proceedings of the 9th international conference on learning representations","author":"M. N. Rizve","year":"2021","unstructured":"Rizve, M. N., Duarte, K., Rawat, Y. S., & Shah, M. (2021). In defense of pseudo-labeling: an uncertainty-aware pseudo-label selection framework for semi-supervised learning. In Proceedings of the 9th international conference on learning representations (pp. 1\u201320). Retrieved December 27, 2025, from https:\/\/openreview.net\/forum?id=-ODN6SbiUU."},{"issue":"8","key":"107_CR39","doi-asserted-by":"publisher","first-page":"1979","DOI":"10.1109\/TPAMI.2018.2858821","volume":"41","author":"T. Miyato","year":"2019","unstructured":"Miyato, T., Maeda, S-i., Koyama, M., & Ishii, S. (2019). Virtual adversarial training: a regularization method for supervised and semi-supervised learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(8), 1979\u20131993.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"107_CR40","first-page":"12674","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Ouali","year":"2020","unstructured":"Ouali, Y., Hudelot, C., & Tami, M. (2020). Semi-supervised semantic segmentation with cross-consistency training. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 12674\u201312684). Piscataway: IEEE."},{"key":"107_CR41","first-page":"6256","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"Q. Xie","year":"2020","unstructured":"Xie, Q., Dai, Z., Hovy, E., Luong, T., & Le, Q. (2020). Unsupervised data augmentation for consistency training. In H. Larochelle, M. Ranzato, R. Hadsell, M.-F. Balcan, & H.-T. Lin (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 6256\u20136268). Red Hook: Curran Associates."},{"key":"107_CR42","first-page":"1","volume-title":"Proceedings of the British machine vision conference","author":"G. French","year":"2020","unstructured":"French, G., Aila, T., Laine, S., Mackiewicz, M., & Finlayson, G. (2020). Semi-supervised semantic segmentation needs strong, high-dimensional perturbations. In Proceedings of the British machine vision conference (pp. 1\u201314). Swansea: BMVA Press."},{"key":"107_CR43","unstructured":"DeVries, T., & Taylor, G. W. (2017). Improved regularization of convolutional neural networks with cutout. arXiv preprint. arXiv:1708.04552."},{"key":"107_CR44","first-page":"6023","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"S. Yun","year":"2019","unstructured":"Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., & Yoo, Y. (2019). CutMix: regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 6023\u20136032). Piscataway: IEEE."},{"key":"107_CR45","first-page":"702","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition workshops","author":"E. D. Cubuk","year":"2020","unstructured":"Cubuk, E. D., Zoph, B., Shlens, J., & Le, Q. V. (2020). Randaugment: practical automated data augmentation with a reduced search space. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition workshops (pp. 702\u2013703). Piscataway: IEEE."},{"key":"107_CR46","doi-asserted-by":"publisher","first-page":"86","DOI":"10.18653\/v1\/P16-1009","volume-title":"Proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: long papers)","author":"R. Sennrich","year":"2016","unstructured":"Sennrich, R., Haddow, B., & Birch, A. (2016). Improving neural machine translation models with monolingual data. In Proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: long papers) (pp. 86\u201396). Stroudsburg: ACL."},{"key":"107_CR47","first-page":"8229","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"J. Yuan","year":"2021","unstructured":"Yuan, J., Liu, Y., Shen, C., Wang, Z., & Li, H. (2021). A simple baseline for semi-supervised semantic segmentation with strong data augmentation. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 8229\u20138238). Piscataway: IEEE."},{"key":"107_CR48","first-page":"4258","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Liu","year":"2022","unstructured":"Liu, Y., Tian, Y., Chen, Y., Liu, F., Belagiannis, V., & Carneiro, G. (2022). Perturbed and strict mean teachers for semi-supervised semantic segmentation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 4258\u20134267). Piscataway: IEEE."},{"key":"107_CR49","first-page":"11350","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Z. Zhao","year":"2023","unstructured":"Zhao, Z., Yang, L., Long, S., Pi, J., Zhou, L., & Wang, J. (2023). Augmentation matters: a simple-yet-effective approach to semi-supervised semantic segmentation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 11350\u201311359). Piscataway: IEEE."},{"key":"107_CR50","first-page":"9957","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"D. Kwon","year":"2022","unstructured":"Kwon, D., & Kwak, S. (2022). Semi-supervised semantic segmentation with error localization network. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 9957\u20139967). Piscataway: IEEE."},{"key":"107_CR51","first-page":"7236","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"L. Yang","year":"2023","unstructured":"Yang, L., Qi, L., Feng, L., Zhang, W., & Shi, Y. (2023). Revisiting weak-to-strong consistency in semi-supervised semantic segmentation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 7236\u20137246). Piscataway: IEEE."},{"key":"107_CR52","first-page":"135","volume-title":"Proceedings of the 15th European conference on computer vision","author":"S. Qiao","year":"2018","unstructured":"Qiao, S., Shen, W., Zhang, Z., Wang, B., & Yuille, A. (2018). Deep co-training for semi-supervised image recognition. In V. Ferrari, M. Hebert, C. Sminchisescu, & Y. Weiss (Eds.), Proceedings of the 15th European conference on computer vision (pp. 135\u2013152). Cham: Springer."},{"key":"107_CR53","first-page":"6728","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"Z. Ke","year":"2019","unstructured":"Ke, Z., Wang, D., Yan, Q., Ren, J., & Lau, R. W. (2019). Dual student: breaking the limits of the teacher in semi-supervised learning. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 6728\u20136736). Piscataway: IEEE."},{"key":"107_CR54","first-page":"22106","volume-title":"Proceedings of the 35th international conference on neural information processing systems","author":"H. Hu","year":"2021","unstructured":"Hu, H., Wei, F., Hu, H., Ye, Q., Cui, J., & Wang, L. (2021). Semi-supervised semantic segmentation via adaptive equalization learning. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P. S. Liang, & J. W. Vaughan (Eds.), Proceedings of the 35th international conference on neural information processing systems (pp. 22106\u201322118). Red Hook: Curran Associates."},{"key":"107_CR55","first-page":"4268","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"L. Yang","year":"2022","unstructured":"Yang, L., Zhuo, W., Qi, L., Shi, Y., & Gao, Y. (2022). ST++: make self-training work better for semi-supervised semantic segmentation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 4268\u20134277). Piscataway: IEEE."},{"key":"107_CR56","first-page":"991","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"B. Hariharan","year":"2011","unstructured":"Hariharan, B., Arbel\u00e1ez, P., Bourdev, L., Maji, S., & Malik, J. (2011). Semantic contours from inverse detectors. In Proceedings of the IEEE international conference on computer vision (pp. 991\u2013998). Piscataway: IEEE."},{"key":"107_CR57","first-page":"801","volume-title":"Proceedings of the 15th European conference on computer vision","author":"L.-C. Chen","year":"2018","unstructured":"Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., & Adam, H. (2018). Encoder-decoder with atrous separable convolution for semantic image segmentation. In V. Ferrari, M. Hebert, C. Sminchisescu, & Y. Weiss (Eds.), Proceedings of the 15th European conference on computer vision (pp. 801\u2013818). Cham: Springer."},{"key":"107_CR58","first-page":"770","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"K. He","year":"2016","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770\u2013778). Piscataway: IEEE."},{"key":"107_CR59","first-page":"761","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"A. Shrivastava","year":"2016","unstructured":"Shrivastava, A., Gupta, A., & Girshick, R. (2016). Training region-based object detectors with online hard example mining. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 761\u2013769). Piscataway: IEEE."},{"key":"107_CR60","first-page":"61792","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"R. Sun","year":"2023","unstructured":"Sun, R., Mai, H., Zhang, T., & Wu, F. (2023). DAW: exploring the better weighting function for semi-supervised semantic segmentation. In A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, & S. Levine (Eds.), Proceedings of the 37th international conference on neural information processing systems (pp. 61792\u201361805). Red Hook: Curran Associates."},{"key":"107_CR61","first-page":"40367","volume-title":"Proceedings of the 37th international conference on neural information processing systems","author":"J. Na","year":"2023","unstructured":"Na, J., Ha, J.-W., Chang, H. J., Han, D., & Hwang, W. (2023). Switching temporary teachers for semi-supervised semantic segmentation. In A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, & S. Levine (Eds.), Proceedings of the 37th international conference on neural information processing systems (pp. 40367\u201340380). Red Hook: Curran Associates."},{"key":"107_CR62","first-page":"3627","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"H. Wang","year":"2024","unstructured":"Wang, H., Zhang, Q., Li, Y., & Li, X. (2024). AllSpark: reborn labeled features from unlabeled in transformer for semi-supervised semantic segmentation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 3627\u20133636). Piscataway: IEEE."},{"key":"107_CR63","first-page":"3097","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"B. Sun","year":"2024","unstructured":"Sun, B., Yang, Y., Zhang, L., Cheng, M.-M., & Hou, Q. (2024). CorrMatch: label propagation via correlation matching for semi-supervised semantic segmentation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 3097\u20133107). Piscataway: IEEE."},{"key":"107_CR64","first-page":"2579","volume":"9","author":"L. van der Maaten","year":"2008","unstructured":"van der Maaten, L., & Hinton, G. (2008). Visualizing data using t-SNE. Journal of Machine Learning Research, 9, 2579\u20132605.","journal-title":"Journal of Machine Learning Research"}],"container-title":["Visual Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-026-00107-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44267-026-00107-3","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-026-00107-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,3]],"date-time":"2026-02-03T09:04:21Z","timestamp":1770109461000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44267-026-00107-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,3]]},"references-count":64,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,12]]}},"alternative-id":["107"],"URL":"https:\/\/doi.org\/10.1007\/s44267-026-00107-3","relation":{},"ISSN":["2097-3330","2731-9008"],"issn-type":[{"value":"2097-3330","type":"print"},{"value":"2731-9008","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,3]]},"assertion":[{"value":"28 August 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 January 2026","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 January 2026","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 February 2026","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"4"}}