{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T13:49:56Z","timestamp":1740145796655,"version":"3.37.3"},"reference-count":40,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,1,10]],"date-time":"2024-01-10T00:00:00Z","timestamp":1704844800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,1,10]],"date-time":"2024-01-10T00:00:00Z","timestamp":1704844800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Image Video Proc."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Multimodal few-shot learning aims to exploit complementary information inherent in multiple modalities for vision tasks in low data scenarios. Most of the current research focuses on a suitable embedding space for the various modalities. While solutions based on embedding provide state-of-the-art results, they reduce the interpretability of the model. Separate visualization approaches enable the models to become more transparent. In this paper, a multimodal few-shot learning framework that is inherently interpretable is presented. This is achieved by using the textual modality in the form of attributes without embedding them. This enables the model to directly explain which attributes caused it to classify an image into a particular class. The model consists of a variational autoencoder to learn the visual latent representation, which is combined with a semantic latent representation that is learnt from a normal autoencoder, which calculates a semantic loss between the latent representation and a binary attribute vector. A decoder reconstructs the original image from concatenated latent vectors. The proposed model outperforms other multimodal methods when all test classes are used, e.g., 50 classes in a 50-way 1-shot setting, and is comparable for lesser number of ways. Since raw text attributes are used, the datasets for evaluation are CUB, SUN and AWA2. The effectiveness of interpretability provided by the model is evaluated by analyzing how well it has learnt to identify the attributes.<\/jats:p>","DOI":"10.1186\/s13640-024-00620-9","type":"journal-article","created":{"date-parts":[[2024,1,10]],"date-time":"2024-01-10T17:02:27Z","timestamp":1704906147000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Multimodal few-shot classification without attribute embedding"],"prefix":"10.1186","volume":"2024","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5819-6638","authenticated-orcid":false,"given":"Jun Qing","family":"Chang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Deepu","family":"Rajan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nicholas","family":"Vun","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,1,10]]},"reference":[{"unstructured":"C. Finn, P. Abbeel, S. Levine, Model-agnostic meta-learning for fast adaptation of deep networks. In: Proceedings of the 34th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 70, pp. 1126\u20131135 (2017)","key":"620_CR1"},{"doi-asserted-by":"crossref","unstructured":"Y. Tian, Y. Wang, D. Krishnan, J.B. Tenenbaum, P. Isola, Rethinking few-shot image classification: a good embedding is all you need? In: Proceedings of European Conference on Computer Vision, pp. 266\u2013282 (2020)","key":"620_CR2","DOI":"10.1007\/978-3-030-58568-6_16"},{"unstructured":"A. Antoniou, H. Edwards, A. Storkey, How to train your MAML. In: International Conference on Learning Representations (2019)","key":"620_CR3"},{"unstructured":"C. Finn, K. Xu, S. Levine, Probabilistic model-agnostic meta-learning. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems. NIPS\u201918, pp. 9537\u20139548 (2018)","key":"620_CR4"},{"doi-asserted-by":"crossref","unstructured":"F. Pahde, M.M. Puscas, J. Wolff, T. Klein, N. Sebe, M. Nabi, Low-shot learning from imaginary 3d model. 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), 978\u2013985 (2019)","key":"620_CR5","DOI":"10.1109\/WACV.2019.00109"},{"unstructured":"E. Triantafillou, T. Zhu, V. Dumoulin, P. Lamblin, U. Evci, K. Xu, R. Goroshin, C. Gelada, K. Swersky, P.-A. Manzagol, H. Larochelle, Meta-dataset: a dataset of datasets for learning to learn from few examples. In: International Conference on Learning Representations (2020)","key":"620_CR6"},{"doi-asserted-by":"crossref","unstructured":"P. Tokmakov, Y.-X. Wang, M. Hebert, Learning compositional representations for few-shot recognition. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV) (2019)","key":"620_CR7","DOI":"10.1109\/ICCV.2019.00647"},{"doi-asserted-by":"crossref","unstructured":"F. Pahde, M. Puscas, T. Klein, M. Nabi, Multimodal prototypical networks for few-shot learning. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 2644\u20132653 (2021)","key":"620_CR8","DOI":"10.1109\/WACV48630.2021.00269"},{"doi-asserted-by":"crossref","unstructured":"Y.-X. Wang, R. Girshick, M. Hebert, B. Hariharan, Low-shot learning from imaginary data. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2018)","key":"620_CR9","DOI":"10.1109\/CVPR.2018.00760"},{"doi-asserted-by":"crossref","unstructured":"G. Montavon, A. Binder, S. Lapuschkin, W. Samek, K.-R. M\u00fcller, in Explainable AI: Interpreting, Explaining and Visualizing Deep Learning, ed by Samek, W., Montavon, G., Vedaldi, A., Hansen, L.K., M\u00fcller, K.-R. Layer-wise relevance propagation: an overview.\u00a0 (2019) pp. 193\u2013209","key":"620_CR10","DOI":"10.1007\/978-3-030-28954-6_10"},{"issue":"2","key":"620_CR11","doi-asserted-by":"publisher","first-page":"336","DOI":"10.1007\/s11263-019-01228-7","volume":"128","author":"RR Selvaraju","year":"2019","unstructured":"R.R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, Grad-cam: visual explanations from deep networks via gradient-based localization. Int. J. Comput. Vis. 128(2), 336\u2013359 (2019)","journal-title":"Int. J. Comput. Vis."},{"doi-asserted-by":"crossref","unstructured":"W. Liu, R. Li, M. Zheng, S. Karanam, Z. Wu, B. Bhanu, R.J. Radke, O. Camps, Towards visually explaining variational autoencoders. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2020)","key":"620_CR12","DOI":"10.1109\/CVPR42600.2020.00867"},{"doi-asserted-by":"crossref","unstructured":"Y. Xian, S. Sharma, B. Schiele, Z. Akata, F-vaegan-d2: a feature generating framework for any-shot learning. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2019)","key":"620_CR13","DOI":"10.1109\/CVPR.2019.01052"},{"issue":"2","key":"620_CR14","doi-asserted-by":"publisher","first-page":"143","DOI":"10.1016\/0010-0285(82)90007-X","volume":"14","author":"I Biederman","year":"1982","unstructured":"I. Biederman, R.J. Mezzanotte, J.C. Rabinowitz, Scene perception: detecting and judging objects undergoing relational violations. Cogn. Psychol. 14(2), 143\u2013177 (1982)","journal-title":"Cogn. Psychol."},{"doi-asserted-by":"crossref","unstructured":"Y. Xian, C. Lampert, B. Schiele, Z. Akata, Zero-shot learning\u2014a comprehensive evaluation of the good, the bad and the ugly. IEEE Transactions on Pattern Analysis & Machine Intelligence (2017)","key":"620_CR15","DOI":"10.1109\/CVPR.2017.328"},{"doi-asserted-by":"crossref","unstructured":"F. Qi, X. Yang, C. Xu, Zero-shot video emotion recognition via multimodal protagonist-aware transformer network. In: Proceedings of the 29th ACM International Conference on Multimedia, pp. 1074\u20131083 (2021)","key":"620_CR16","DOI":"10.1145\/3474085.3475647"},{"doi-asserted-by":"crossref","unstructured":"E. Schonfeld, S. Ebrahimi, S. Sinha, T. Darrell, Z. Akata, Generalized zero- and few-shot learning via aligned variational autoencoders. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2019)","key":"620_CR17","DOI":"10.1109\/CVPR.2019.00844"},{"issue":"1","key":"620_CR18","first-page":"1","volume":"1","author":"Y Wang","year":"2002","unstructured":"Y. Wang, Q. Yao, J.T. Kwok, L.M. Ni, Generalizing from a few examples: a survey on few-shot learning. ACM Comput. Surv. 1(1), 1\u20131134 (2002)","journal-title":"ACM Comput. Surv."},{"unstructured":"S. Benaim, L. Wolf, One-shot unsupervised cross domain translation. Adv. Neural Inf. Process. Syst. 31 (2018)","key":"620_CR19"},{"unstructured":"H. Gao, Z. Shou, A. Zareian, H. Zhang, S.-F. Chang, Low-shot learning via covariance-preserving adversarial augmentation networks. Adv. Neural Inf. Process. Syst. 31 (2018)","key":"620_CR20"},{"doi-asserted-by":"crossref","unstructured":"A. Miller, A. Fisch, J. Dodge, A.-H. Karimi, A. Bordes, J. Weston, Key-value memory networks for directly reading documents. In: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 1400\u20131409 (2016)","key":"620_CR21","DOI":"10.18653\/v1\/D16-1147"},{"unstructured":"M. Andrychowicz, M. Denil, S. G\u00f3mez, M.W. Hoffman, D. Pfau, T. Schaul, B. Shillingford, N. de Freitas, Learning to learn by gradient descent by gradient descent. In Proceedings of the 29th International Conference on Neural Information Processing Systems (NIPS'16)","key":"620_CR22"},{"doi-asserted-by":"crossref","unstructured":"Y. Song, T. Wang, P. Cai, S.K. Mondal, J.P. Sahoo, A comprehensive survey of few-shot learning: evolution, applications, challenges, and opportunities. ACM Comput. Surv. (2023). Just Accepted","key":"620_CR23","DOI":"10.1145\/3582688"},{"doi-asserted-by":"crossref","unstructured":"M.P. Fortin, B. Chaib-draa, Towards contextual learning in few-shot object classification. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 3279\u20133288 (2021)","key":"620_CR24","DOI":"10.1109\/WACV48630.2021.00332"},{"key":"620_CR25","doi-asserted-by":"publisher","first-page":"142","DOI":"10.1016\/j.patrec.2022.06.012","volume":"160","author":"E Schwartz","year":"2022","unstructured":"E. Schwartz, L. Karlinsky, R. Feris, R. Giryes, A. Bronstein, Baby steps towards few-shot learning with multiple semantics. Pattern Recogn. Lett. 160, 142\u2013147 (2022)","journal-title":"Pattern Recogn. Lett."},{"doi-asserted-by":"crossref","unstructured":"F. Pahde, O. Ostapenko, P. Hnichen, T. Klein, M. Nabi, Self-paced adversarial training for multimodal few-shot learning. In: 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 218\u2013226 (2019)","key":"620_CR26","DOI":"10.1109\/WACV.2019.00029"},{"unstructured":"C. Xing, N. Rostamzadeh, B. Oreshkin, P.O. Pinheiro, Adaptive cross-modal few-shot learning. Adv. Neural Inf. Process. Syst. 32 (2019)","key":"620_CR27"},{"issue":"9","key":"620_CR28","doi-asserted-by":"publisher","first-page":"4594","DOI":"10.1109\/TIP.2019.2910052","volume":"28","author":"Z Chen","year":"2019","unstructured":"Z. Chen, Y. Fu, Y. Zhang, Y.-G. Jiang, X. Xue, L. Sigal, Multi-level semantic feature augmentation for one-shot learning. IEEE Trans. Image Process. 28(9), 4594\u20134605 (2019)","journal-title":"IEEE Trans. Image Process."},{"doi-asserted-by":"crossref","unstructured":"J. Andreas, D. Klein, S. Levine, Learning with latent language. In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 2166\u20132179 (2018)","key":"620_CR29","DOI":"10.18653\/v1\/N18-1197"},{"doi-asserted-by":"crossref","unstructured":"J. Mu, P. Liang, N. Goodman, Shaping visual representations with language for few-shot classification. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 4823\u20134830 (2020)","key":"620_CR30","DOI":"10.18653\/v1\/2020.acl-main.436"},{"doi-asserted-by":"crossref","unstructured":"Z. Chen, Y. Luo, S. Wang, R. Qiu, J. Li, Z. Huang, Mitigating generation shifts for generalized zero-shot learning. In: Proceedings of the 29th ACM International Conference on Multimedia, pp. 844\u2013852 (2021)","key":"620_CR31","DOI":"10.1145\/3474085.3475258"},{"doi-asserted-by":"crossref","unstructured":"D. Samuel, Y. Atzmon, G. Chechik, From generalized zero-shot learning to long-tail with class descriptors. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 286\u2013295 (2021)","key":"620_CR32","DOI":"10.1109\/WACV48630.2021.00033"},{"unstructured":"D.P. Kingma, M. Welling, Auto-encoding variational Bayes. In: 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14\u201316, 2014 (2014)","key":"620_CR33"},{"unstructured":"J. Fajtl, V. Argyriou, D. Monekosso, P. Remagnino, Latent bernoulli autoencoder. In: III, H.D., Singh, A. (eds.) Proceedings of the 37th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 119, pp. 2964\u20132974 (2020)","key":"620_CR34"},{"unstructured":"Y. Zhang, S. Huang, X. Peng, D. Yang, Dizygotic conditional variational autoencoder for multi-modal and partial modality absent few-shot learning (2021). arXiv:2106.14467","key":"620_CR35"},{"unstructured":"C. Wah, S. Branson, P. Welinder, P. Perona, S. Belongie, The Caltech-UCSD Birds-200-2011 Dataset. Technical Report CNS-TR-2011-001, California Institute of Technology (2011)","key":"620_CR36"},{"issue":"1\u20132","key":"620_CR37","doi-asserted-by":"publisher","first-page":"59","DOI":"10.1007\/s11263-013-0695-z","volume":"108","author":"G Patterson","year":"2014","unstructured":"G. Patterson, C. Xu, H. Su, J. Hays, The sun attribute database: beyond categories for deeper scene understanding. Int. J. Comput. Vis. 108(1\u20132), 59\u201381 (2014)","journal-title":"Int. J. Comput. Vis."},{"unstructured":"M. Afham, S. Khan, M.H. Khan, M. Naseer, F.S. Khan, Rich semantics improve few-shot learning. 32nd British Machine Vision Conference (2021)","key":"620_CR38"},{"unstructured":"O. Vinyals, C. Blundell, T. Lillicrap, k. kavukcuoglu, D. Wierstra, Matching networks for one shot learning. Adv. Neural Inf. Process. Syst. 29 (2016)","key":"620_CR39"},{"unstructured":"J. Snell, K. Swersky, R. Zemel, Prototypical networks for few-shot learning. Adv. Neural Inf. Process. Syst. 30 (2017)","key":"620_CR40"}],"container-title":["EURASIP Journal on Image and Video Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13640-024-00620-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13640-024-00620-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13640-024-00620-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,1,10]],"date-time":"2024-01-10T17:04:03Z","timestamp":1704906243000},"score":1,"resource":{"primary":{"URL":"https:\/\/jivp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13640-024-00620-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,10]]},"references-count":40,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["620"],"URL":"https:\/\/doi.org\/10.1186\/s13640-024-00620-9","relation":{},"ISSN":["1687-5281"],"issn-type":[{"type":"electronic","value":"1687-5281"}],"subject":[],"published":{"date-parts":[[2024,1,10]]},"assertion":[{"value":"27 February 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"31 December 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 January 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"4"}}