{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T04:44:29Z","timestamp":1781585069647,"version":"3.54.5"},"reference-count":17,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2023,7,10]],"date-time":"2023-07-10T00:00:00Z","timestamp":1688947200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Ministry of Science and Higher Education","award":["075-10-2021-068"],"award-info":[{"award-number":["075-10-2021-068"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Embeddings, i.e., vector representations of objects, such as texts, images, or graphs, play a key role in deep learning methodologies nowadays. Prior research has shown the importance of analyzing the isotropy of textual embeddings for transformer-based text encoders, such as the BERT model. Anisotropic word embeddings do not use the entire space, instead concentrating on a narrow cone in such a pretrained vector space, negatively affecting the performance of applications, such as textual semantic similarity. Transforming a vector space to optimize isotropy has been shown to be beneficial for improving performance in text processing tasks. This paper is the first comprehensive investigation of the distribution of multimodal embeddings using the example of OpenAI\u2019s CLIP pretrained model. We aimed to deepen the understanding of the embedding space of multimodal embeddings, which has previously been unexplored in this respect, and study the impact on various end tasks. Our initial efforts were focused on measuring the alignment of image and text embedding distributions, with an emphasis on their isotropic properties. In addition, we evaluated several gradient-free approaches to enhance these properties, establishing their efficiency in improving the isotropy\/alignment of the embeddings and, in certain cases, the zero-shot classification accuracy. Significantly, our analysis revealed that both CLIP and BERT models yielded embeddings situated within a cone immediately after initialization and preceding training. However, they were mostly isotropic in the local sense. We further extended our investigation to the structure of multilingual CLIP text embeddings, confirming that the observed characteristics were language-independent. By computing the few-shot classification accuracy and point-cloud metrics, we provide evidence of a strong correlation among multilingual embeddings. Embeddings transformation using the methods described in this article makes it easier to visualize embeddings. At the same time, multiple experiments that we conducted showed that, in regard to the transformed embeddings, the downstream tasks performance does not drop substantially (and sometimes is even improved). This means that one could obtain an easily visualizable embedding space, without substantially losing the quality of downstream tasks.<\/jats:p>","DOI":"10.3390\/info14070392","type":"journal-article","created":{"date-parts":[[2023,7,11]],"date-time":"2023-07-11T01:58:14Z","timestamp":1689040694000},"page":"392","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["On Isotropy of Multimodal Embeddings"],"prefix":"10.3390","volume":"14","author":[{"given":"Kirill","family":"Tyshchuk","sequence":"first","affiliation":[{"name":"Center of Artificial Intelligence Technology (CAIT), Skolkovo Institute of Science and Technology (Skoltech), 121205 Moscow, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Polina","family":"Karpikova","sequence":"additional","affiliation":[{"name":"Center of Artificial Intelligence Technology (CAIT), Skolkovo Institute of Science and Technology (Skoltech), 121205 Moscow, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrew","family":"Spiridonov","sequence":"additional","affiliation":[{"name":"Center of Artificial Intelligence Technology (CAIT), Skolkovo Institute of Science and Technology (Skoltech), 121205 Moscow, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anastasiia","family":"Prutianova","sequence":"additional","affiliation":[{"name":"Center of Artificial Intelligence Technology (CAIT), Skolkovo Institute of Science and Technology (Skoltech), 121205 Moscow, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anton","family":"Razzhigaev","sequence":"additional","affiliation":[{"name":"Center of Artificial Intelligence Technology (CAIT), Skolkovo Institute of Science and Technology (Skoltech), 121205 Moscow, Russia"},{"name":"Artificial Intelligence Research Institute (AIRI), 121170 Moscow, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6097-6118","authenticated-orcid":false,"given":"Alexander","family":"Panchenko","sequence":"additional","affiliation":[{"name":"Center of Artificial Intelligence Technology (CAIT), Skolkovo Institute of Science and Technology (Skoltech), 121205 Moscow, Russia"},{"name":"Artificial Intelligence Research Institute (AIRI), 121170 Moscow, Russia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,7,10]]},"reference":[{"key":"ref_1","unstructured":"Gao, J., He, D., Tan, X., Qin, T., Wang, L., and Liu, T. (2019, January 6\u20139). Representation Degeneration Problem in Training Natural Language Generation Models. Proceedings of the 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Fuster Baggetto, A., and Fresno, V. (2022, January 20). Is anisotropy really the cause of BERT embeddings not being semantic?. Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2022, Abu Dhabi, United Arab Emirates.","DOI":"10.18653\/v1\/2022.findings-emnlp.314"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Li, B., Zhou, H., He, J., Wang, M., Yang, Y., and Li, L. (2020, January 16\u201320). On the Sentence Embeddings from Pre-trained Language Models. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online.","DOI":"10.18653\/v1\/2020.emnlp-main.733"},{"key":"ref_4","unstructured":"Su, J., Cao, J., Liu, W., and Ou, Y. (2021). Whitening Sentence Representations for Better Semantics and Faster Retrieval. arXiv."},{"key":"ref_5","unstructured":"Meila, M., and Zhang, T. (2021, January 18\u201324). Learning Transferable Visual Models From Natural Language Supervision. Proceedings of the 38th International Conference on Machine Learning, PMLR, Virtual. Proceedings of Machine Learning Research."},{"key":"ref_6","unstructured":"Wang, L., Huang, J., Huang, K., Hu, Z., Wang, G., and Gu, Q. (2020, January 26\u201330). Improving neural language generation with spectrum control. Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zhou, W., Lin, B., and Ren, X. (2021, January 2\u20139). IsoBN: Fine-Tuning BERT with Isotropic Batch Normalization. Proceedings of the AAAI Conference on Artificial Intelligence, Virtually.","DOI":"10.1609\/aaai.v35i16.17718"},{"key":"ref_8","unstructured":"Cai, X., Huang, J., Bian, Y., and Church, K. (2021, January 3\u20137). Isotropy in the contextual embedding space: Clusters and manifolds. Proceedings of the International Conference on Learning Representations, Virtual Event."},{"key":"ref_9","unstructured":"Burstein, J., Doran, C., and Solorio, T. (2019, January 2\u20137). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA. Association for Computational Linguistics; Long and Short Papers."},{"key":"ref_10","unstructured":"Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J.R. (2020, January 5\u201310). Unsupervised Cross-lingual Representation Learning at Scale. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online. Association for Computational Linguistics."},{"key":"ref_11","unstructured":"Guyon, I., von Luxburg, U., Bengio, S., Wallach, H.M., Fergus, R., Vishwanathan, S.V.N., and Garnett, R. (2017, January 4\u20139). Attention is All you Need. Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, Long Beach, CA, USA."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Ding, Y., Martinkus, K., Pascual, D., Clematide, S., and Wattenhofer, R. (2021). On Isotropy Calibration of Transformer Models. arXiv.","DOI":"10.18653\/v1\/2022.insights-1.1"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C. (2014). Microsoft COCO: Common Objects in Context. arXiv.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_14","unstructured":"Schamoni, S., Hitschler, J., and Riezler, S. (2018, January 17\u201321). A dataset and reranking method for multimodal MT of user-generated image captions. Proceedings of the 13th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track), Boston, MA, USA."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"McInnes, L., and Healy, J. (2018). UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv.","DOI":"10.21105\/joss.00861"},{"key":"ref_16","first-page":"26","article-title":"Cifar-100 (canadian institute for advanced research). 30 [65] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_17","first-page":"8580","article-title":"Neural tangent kernel: Convergence and generalization in neural networks","volume":"31","author":"Jacot","year":"2018","journal-title":"Adv. Neural Inf. Process. Syst."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/14\/7\/392\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:09:41Z","timestamp":1760126981000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/14\/7\/392"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,10]]},"references-count":17,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2023,7]]}},"alternative-id":["info14070392"],"URL":"https:\/\/doi.org\/10.3390\/info14070392","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,10]]}}}