{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:21:56Z","timestamp":1760145716233,"version":"build-2065373602"},"reference-count":37,"publisher":"MDPI AG","issue":"17","license":[{"start":{"date-parts":[[2024,8,30]],"date-time":"2024-08-30T00:00:00Z","timestamp":1724976000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Key R &amp; D program of Shaanxi Province","award":["No.2024GX-YBXM-054"],"award-info":[{"award-number":["No.2024GX-YBXM-054"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Metric-based meta-learning methods have demonstrated remarkable success in the domain of few-shot image classification. However, their performance is significantly contingent upon the choice of metric and the feature representation for the support classes. Current approaches, which predominantly rely on holistic image features, may inadvertently disregard critical details necessary for novel tasks, a phenomenon known as \u201csupervision collapse\u201d. Moreover, relying solely on visual features to characterize support classes can prove to be insufficient, particularly in scenarios involving limited sample sizes. In this paper, we introduce an innovative framework named Patch Matching Metric-based Semantic Interaction Meta-Learning (PatSiML), designed to overcome these challenges. To counteract supervision collapse, we have developed a patch matching metric strategy based on the Transformer architecture to transform input images into a set of distinct patch embeddings. This approach dynamically creates task-specific embeddings, facilitated by a graph convolutional network, to formulate precise matching metrics between the support classes and the query image patches. To enhance the integration of semantic knowledge, we have also integrated a label-assisted channel semantic interaction strategy. This strategy merges word embeddings with patch-level visual features across the channel dimension, utilizing a sophisticated language model to combine semantic understanding with visual information. Our empirical findings across four diverse datasets reveal that the PatSiML method achieves a classification accuracy improvement of 0.65% to 21.15% over existing methodologies, underscoring its robustness and efficacy.<\/jats:p>","DOI":"10.3390\/s24175620","type":"journal-article","created":{"date-parts":[[2024,8,30]],"date-time":"2024-08-30T07:45:47Z","timestamp":1725003947000},"page":"5620","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Semantic Interaction Meta-Learning Based on Patch Matching Metric"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0626-2760","authenticated-orcid":false,"given":"Baoguo","family":"Wei","sequence":"first","affiliation":[{"name":"School of Electronic Information, Northwestern Polytechnical University, Xi\u2019an 710129, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinyu","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Electronic Information, Northwestern Polytechnical University, Xi\u2019an 710129, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuetong","family":"Su","sequence":"additional","affiliation":[{"name":"School of Electronic Information, Northwestern Polytechnical University, Xi\u2019an 710129, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yue","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Electronic Information, Northwestern Polytechnical University, Xi\u2019an 710129, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lixin","family":"Li","sequence":"additional","affiliation":[{"name":"School of Electronic Information, Northwestern Polytechnical University, Xi\u2019an 710129, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,8,30]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"3458","DOI":"10.1109\/TNNLS.2020.3011526","article-title":"Learning to learn adaptive classifier\u2013predictor for few-shot learning","volume":"32","author":"Lai","year":"2020","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_2","first-page":"981","article-title":"Crosstransformers: Spatially-aware few-shot transfer","volume":"33","author":"Doersch","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"doi-asserted-by":"crossref","unstructured":"Chen, Y., Liu, Z., Xu, H., Darrell, T., and Wang, X. (2021, January 10\u201317). Meta-baseline: Exploring simple meta-learning for few-shot learning. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","key":"ref_3","DOI":"10.1109\/ICCV48922.2021.00893"},{"doi-asserted-by":"crossref","unstructured":"Kang, S., Hwang, D., Eo, M., Kim, T., and Rhee, W. (2023, January 24). Meta-learning with a geometry-adaptive preconditioner. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","key":"ref_4","DOI":"10.1109\/CVPR52729.2023.01543"},{"key":"ref_5","first-page":"5632","article-title":"Deepemd: Differentiable earth mover\u2019s distance for few-shot learning","volume":"45","author":"Zhang","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"unstructured":"Finn, C., Abbeel, P., and Levine, S. (2017, January 6\u201311). Model-agnostic meta-learning for fast adaptation of deep networks. Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia.","key":"ref_6"},{"unstructured":"Snell, J., Swersky, K., and Zemel, R. (2017, January 4\u20139). Prototypical networks for few-shot learning. Proceedings of the Advances in Neural Information Processing Systems 30 (NIPS 2017), Long Beach, CA, USA.","key":"ref_7"},{"doi-asserted-by":"crossref","unstructured":"Sung, F., Yang, Y., Zhang, L., Xiang, T., Torr, P.H., and Hospedales, T.M. (2018, January 18\u201323). Learning to compare: Relation network for few-shot learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","key":"ref_8","DOI":"10.1109\/CVPR.2018.00131"},{"doi-asserted-by":"crossref","unstructured":"Gidaris, S., and Komodakis, N. (2018, January 18\u201323). Dynamic few-shot visual learning without forgetting. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","key":"ref_9","DOI":"10.1109\/CVPR.2018.00459"},{"key":"ref_10","first-page":"3582","article-title":"Rethinking generalization in few-shot classification","volume":"35","author":"Hiller","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"210102","DOI":"10.1007\/s11432-022-3700-8","article-title":"Sparse spatial transformers for few-shot learning","volume":"66","author":"Chen","year":"2023","journal-title":"Sci. China Inf. Sci."},{"doi-asserted-by":"crossref","unstructured":"Li, A., Huang, W., Lan, X., Feng, J., Li, Z., and Wang, L. (2020, January 17\u201324). Boosting few-shot learning with adaptive margin loss. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","key":"ref_12","DOI":"10.1109\/CVPR42600.2020.01259"},{"doi-asserted-by":"crossref","unstructured":"Yang, F., Wang, R., and Chen, X. (2022, January 3\u20138). Sega: Semantic guided attention on visual prototype for few-shot learning. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","key":"ref_13","DOI":"10.1109\/WACV51458.2022.00165"},{"doi-asserted-by":"crossref","unstructured":"Yan, K., Bouraoui, Z., Wang, P., Jameel, S., and Schockaert, S. (2021, January 16\u201319). Aligning visual prototypes with bert embeddings for few-shot learning. Proceedings of the 2021 International Conference on Multimedia Retrieval, Taipei, Taiwan.","key":"ref_14","DOI":"10.1145\/3460426.3463641"},{"unstructured":"Xing, C., Rostamzadeh, N., Oreshkin, B., and Pinheiro, P.O.O. (2019, January 8\u201314). Adaptive cross-modal few-shot learning. Proceedings of the Advances in Neural Information Processing Systems 32 (NeurIPS 2019), Vancouver, BC, Canada.","key":"ref_15"},{"doi-asserted-by":"crossref","unstructured":"Chen, W., Si, C., Zhang, Z., Wang, L., Wang, Z., and Tan, T. (2023, January 17\u201324). Semantic prompt for few-shot image recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","key":"ref_16","DOI":"10.1109\/CVPR52729.2023.10308797"},{"unstructured":"Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., and Kong, T. (2021). ibot: Image bert pre-training with online tokenizer. arXiv.","key":"ref_17"},{"unstructured":"Bao, H., Dong, L., Piao, S., and Wei, F. (2021). Beit: Bert pre-training of image transformers. arXiv.","key":"ref_18"},{"doi-asserted-by":"crossref","unstructured":"Ye, H.-J., Hu, H., Zhan, D.-C., and Sha, F. (2020, January 17\u201324). Few-shot learning via embedding adaptation with set-to-set functions. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","key":"ref_19","DOI":"10.1109\/CVPR42600.2020.00883"},{"unstructured":"Hou, R., Chang, H., Ma, B., Shan, S., and Chen, X. (2019, January 8\u201314). Cross attention network for few-shot classification. Proceedings of the Advances in Neural Information Processing Systems 32 (NeurIPS 2019), Vancouver, BC, Canada.","key":"ref_20"},{"unstructured":"Kipf, T.N., and Welling, M. (2016). Semi-supervised classification with graph convolutional networks. arXiv.","key":"ref_21"},{"doi-asserted-by":"crossref","unstructured":"Li, W., Wang, L., Xu, J., Huo, J., Gao, Y., and Luo, J. (2019, January 17\u201324). Revisiting local descriptor based image-to-class measure for few-shot learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","key":"ref_22","DOI":"10.1109\/CVPR.2019.00743"},{"unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the 38th International Conference on Machine Learnin, Virtual Event.","key":"ref_23"},{"unstructured":"Vinyals, O., Blundell, C., Lillicrap, T., and Wierstra, D. (2016, January 5\u201310). Matching networks for one shot learning. Proceedings of the Advances in Neural Information Processing Systems 29 (NIPS 2016), Barcelona, Spain.","key":"ref_24"},{"unstructured":"Ren, M., Triantafillou, E., Ravi, S., Snell, J., Swersky, K., Tenenbaum, J.B., Larochelle, H., and Zemel, R.S. (2018). Meta-learning for semi-supervised few-shot classification. arXiv.","key":"ref_25"},{"unstructured":"Bertinetto, L., Henriques, J.F., Torr, P.H., and Vedaldi, A. (2018). Meta-learning with differentiable closed-form solvers. arXiv.","key":"ref_26"},{"unstructured":"Oreshkin, B., L\u00f3pez, P.R., and Lacoste, A. (2018, January 3\u20138). Tadam: Task dependent adaptive metric for improved few-shot learning. Proceedings of the Advances in Neural Information Processing Systems 31 (NeurIPS 2018), Montr\u00e9al, QC, Canada.","key":"ref_27"},{"unstructured":"Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and J\u00e9gou, H. (2021, January 18\u201324). Training data-efficient image transformers & distillation through attention. Proceedings of the 38th International Conference on Machine Learning, Virtual Event.","key":"ref_28"},{"doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","key":"ref_29","DOI":"10.1109\/ICCV48922.2021.00986"},{"unstructured":"Loshchilov, I., and Hutter, F. (2017). Decoupled weight decay regularization. arXiv.","key":"ref_30"},{"doi-asserted-by":"crossref","unstructured":"Zhang, X., Meng, D., Gouk, H., and Hospedales, T.M. (2021, January 10\u201317). Shallow bayesian meta learning for real-world few-shot recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","key":"ref_31","DOI":"10.1109\/ICCV48922.2021.00069"},{"doi-asserted-by":"crossref","unstructured":"Yang, F., Wang, R., and Chen, X. (2023, January 3\u20137). Semantic guided latent parts embedding for few-shot learning. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","key":"ref_32","DOI":"10.1109\/WACV56688.2023.00541"},{"unstructured":"Hu, S.X., Moreno, P.G., Xiao, Y., Shen, X., Obozinski, G., Lawrence, N.D., and Damianou, A. (2020). Empirical bayes transductive meta-learning with synthetic gradients. arXiv.","key":"ref_33"},{"doi-asserted-by":"crossref","unstructured":"Afrasiyabi, A., Lalonde, J.-F., and Gagn\u00e9, C. (2020, January 23\u201328). Associative alignment for few-shot image classification. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Proceedings, Part V 16.","key":"ref_34","DOI":"10.1007\/978-3-030-58558-7_2"},{"doi-asserted-by":"crossref","unstructured":"Dong, B., Zhou, P., Yan, S., and Zuo, W. (2022, January 23\u201327). Self-promoted supervision for few-shot transformer. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","key":"ref_35","DOI":"10.1007\/978-3-031-20044-1_19"},{"doi-asserted-by":"crossref","unstructured":"Reimers, N., and Gurevych, I. (2019). Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv.","key":"ref_36","DOI":"10.18653\/v1\/D19-1410"},{"doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C.D. (2014, January 25\u201329). Glove: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","key":"ref_37","DOI":"10.3115\/v1\/D14-1162"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/17\/5620\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:45:25Z","timestamp":1760111125000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/17\/5620"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,30]]},"references-count":37,"journal-issue":{"issue":"17","published-online":{"date-parts":[[2024,9]]}},"alternative-id":["s24175620"],"URL":"https:\/\/doi.org\/10.3390\/s24175620","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2024,8,30]]}}}