{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,22]],"date-time":"2025-12-22T22:14:12Z","timestamp":1766441652577,"version":"build-2065373602"},"reference-count":57,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2024,5,29]],"date-time":"2024-05-29T00:00:00Z","timestamp":1716940800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001711","name":"SNF Sinergia project","doi-asserted-by":"publisher","award":["CRSII5-193716"],"award-info":[{"award-number":["CRSII5-193716"]}],"id":[{"id":"10.13039\/501100001711","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>We present a new method of self-supervised learning and knowledge distillation based on multi-views and multi-representations (MV\u2013MR). MV\u2013MR is based on the maximization of dependence between learnable embeddings from augmented and non-augmented views, jointly with the maximization of dependence between learnable embeddings from the augmented view and multiple non-learnable representations from the non-augmented view. We show that the proposed method can be used for efficient self-supervised classification and model-agnostic knowledge distillation. Unlike other self-supervised techniques, our approach does not use any contrastive learning, clustering, or stop gradients. MV\u2013MR is a generic framework allowing the incorporation of constraints on the learnable embeddings via the usage of image multi-representations as regularizers. The proposed method is used for knowledge distillation. MV\u2013MR provides state-of-the-art self-supervised performance on the STL10 and CIFAR20 datasets in a linear evaluation setup. We show that a low-complexity ResNet50 model pretrained using proposed knowledge distillation based on the CLIP ViT model achieves state-of-the-art performance on STL10 and CIFAR100 datasets.<\/jats:p>","DOI":"10.3390\/e26060466","type":"journal-article","created":{"date-parts":[[2024,5,30]],"date-time":"2024-05-30T08:15:54Z","timestamp":1717056954000},"page":"466","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["MV\u2013MR: Multi-Views and Multi-Representations for Self-Supervised Learning and Knowledge Distillation"],"prefix":"10.3390","volume":"26","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5301-9141","authenticated-orcid":false,"given":"Vitaliy","family":"Kinakh","sequence":"first","affiliation":[{"name":"Department of Computer Science, University of Geneva, 1227 Carouge, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6575-4290","authenticated-orcid":false,"given":"Mariia","family":"Drozdova","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Geneva, 1227 Carouge, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0416-9674","authenticated-orcid":false,"given":"Slava","family":"Voloshynovskiy","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Geneva, 1227 Carouge, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,5,29]]},"reference":[{"key":"ref_1","unstructured":"Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., and Kong, T. (2021). ibot: Image bert pre-training with online tokenizer. arXiv."},{"key":"ref_2","first-page":"4071","article-title":"A survey of self-supervised and few-shot object detection","volume":"45","author":"Huang","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_3","unstructured":"Zheng, H., Han, J., Wang, H., Yang, L., Zhao, Z., Wang, C., and Chen, D.Z. (October, January 27). Hierarchical self-supervised learning for medical image segmentation based on multi-domain data aggregation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Strasbourg, France."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s10994-022-06219-3","article-title":"BT-Unet: A self-supervised learning framework for biomedical image segmentation using Barlow Twins with U-Net models","volume":"111","author":"Punn","year":"2022","journal-title":"Mach. Learn."},{"key":"ref_5","unstructured":"Oord, A.v.d., Li, Y., and Vinyals, O. (2018). Representation learning with contrastive predictive coding. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Chen, X., and He, K. (2021, January 20\u201325). Exploring simple siamese representation learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01549"},{"key":"ref_7","unstructured":"Zbontar, J., Jing, L., Misra, I., LeCun, Y., and Deny, S. (2021, January 18\u201324). Barlow twins: Self-supervised learning via redundancy reduction. Proceedings of the International Conference on Machine Learning, PMLR, Virtual."},{"key":"ref_8","unstructured":"Bao, H., Dong, L., Piao, S., and Wei, F. (2021). Beit: Bert pre-training of image transformers. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"He, K., Chen, X., Xie, S., Li, Y., Doll\u00e1r, P., and Girshick, R. (2022, January 18\u201324). Masked autoencoders are scalable vision learners. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"ref_10","first-page":"2769","article-title":"Measuring and testing dependence by correlation of distances","volume":"35","author":"Rizzo","year":"2007","journal-title":"Ann. Stat."},{"key":"ref_11","unstructured":"Bardes, A., Ponce, J., and LeCun, Y. (2021). Vicreg: Variance-invariance-covariance regularization for self-supervised learning. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_13","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning, PMLR, Online."},{"key":"ref_14","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3\u20137). An Image is Worth 16\u00d716 Words: Transformers for Image Recognition at Scale. Proceedings of the ICLR, Virtual Event."},{"key":"ref_15","unstructured":"Coates, A., Ng, A., and Lee, H. (2011, January 11\u201313). An analysis of single-layer networks in unsupervised feature learning. Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, JMLR Workshop and Conference Proceedings, Fort Lauderdale, FL, USA."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Van Gansbeke, W., Vandenhende, S., Georgoulis, S., Proesmans, M., and Van Gool, L. (2020, January 23\u201328). Scan: Learning to classify images without labels. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58607-2_16"},{"key":"ref_17","unstructured":"Krizhevsky, A. (2009). Learning Multiple Layers of Features from Tiny Images, Department of Computer Science, University of Toronto. Technical Report."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"2208","DOI":"10.1109\/TPAMI.2018.2855738","article-title":"Scattering networks for hybrid representation learning","volume":"41","author":"Oyallon","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"655","DOI":"10.1109\/TPAMI.1981.4767166","article-title":"Real-time adaptive contrast enhancement","volume":"6","author":"Narendra","year":"1981","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"886","DOI":"10.1109\/CVPR.2005.177","article-title":"Histograms of oriented gradients for human detection","volume":"Volume 1","author":"Dalal","year":"2005","journal-title":"Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905)"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive image features from scale-invariant keypoints","volume":"60","author":"Loew","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Rublee, E., Rabaud, V., Konolige, K., and Bradski, G. (2011, January 6\u201313). ORB: An efficient alternative to SIFT or SURF. Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126544"},{"key":"ref_24","unstructured":"Pietik\u00e4inen, M., and Zhao, G. (2015). Advances in Independent Component Analysis and Learning Machines, Elsevier."},{"key":"ref_25","unstructured":"Platt, J., Koller, D., Singer, Y., and Roweis, S. (2007). Proceedings of the Advances in Neural Information Processing Systems, Curran Associates, Inc."},{"key":"ref_26","unstructured":"Gidaris, S., Singh, P., and Komodakis, N. (2018). Unsupervised representation learning by predicting image rotations. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Noroozi, M., and Favaro, P. (2016, January 11\u201314). Unsupervised learning of visual representations by solving jigsaw puzzles. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46466-4_5"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., and Efros, A.A. (2016, January 27\u201330). Context encoders: Feature learning by inpainting. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.278"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Larsson, G., Maire, M., and Shakhnarovich, G. (2017, January 21\u201326). Colorization as a proxy task for visual understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.96"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Kinakh, V., Taran, O., and Voloshynovskiy, S. (2021, January 11\u201317). ScatSimCLR: Self-supervised contrastive learning with pretext task regularization for small-scale datasets. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Virtual Conference.","DOI":"10.1109\/ICCVW54120.2021.00129"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Yi, J.S.K., Seo, M., Park, J., and Choi, D.G. (2022). Using Self-Supervised Pretext Tasks for Active Learning. arXiv.","DOI":"10.1007\/978-3-031-19809-0_34"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zaiem, S., Parcollet, T., and Essid, S. (2021). Pretext Tasks selection for multitask self-supervised speech representation learning. arXiv.","DOI":"10.21437\/Interspeech.2021-1027"},{"key":"ref_33","unstructured":"Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. (2020, January 13\u201318). A simple framework for contrastive learning of visual representations. Proceedings of the International Conference on Machine Learning, PMLR, Virtual."},{"key":"ref_34","first-page":"9912","article-title":"Unsupervised learning of visual features by contrasting cluster assignments","volume":"33","author":"Caron","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Caron, M., Touvron, H., Misra, I., J\u00e9gou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021, January 11\u201317). Emerging properties in self-supervised vision transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Caron, M., Bojanowski, P., Joulin, A., and Douze, M. (2018, January 8\u201314). Deep clustering for unsupervised learning of visual features. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_9"},{"key":"ref_37","first-page":"21271","article-title":"Bootstrap your own latent\u2014A new approach to self-supervised learning","volume":"33","author":"Grill","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"1789","DOI":"10.1007\/s11263-021-01453-z","article-title":"Knowledge distillation: A survey","volume":"129","author":"Gou","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_39","unstructured":"Hinton, G., Vinyals, O., and Dean, J. (2015). Distilling the knowledge in a neural network. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Mirzadeh, S.I., Farajtabar, M., Li, A., Levine, N., Matsukawa, A., and Ghasemzadeh, H. (2020, January 20\u201327). Improved knowledge distillation via teacher assistant. Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada.","DOI":"10.1609\/aaai.v34i04.5963"},{"key":"ref_41","unstructured":"Zhang, L., Song, J., Gao, A., Chen, J., Bao, C., and Ma, K. (November, January 27). Be your own teacher: Improve the performance of convolutional neural networks via self distillation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_42","first-page":"2654","article-title":"Do deep nets really need to be deep?","volume":"27","author":"Ba","year":"2014","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_43","unstructured":"Romero, A., Ballas, N., Kahou, S.E., Chassang, A., Gatta, C., and Bengio, Y. (2014). Fitnets: Hints for thin deep nets. arXiv."},{"key":"ref_44","unstructured":"Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. (2021). Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Misra, I., and Maaten, L.v.d. (2020, January 14\u201319). Self-supervised learning of pretext-invariant representations. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00674"},{"key":"ref_46","first-page":"6827","article-title":"What makes for good views for contrastive learning?","volume":"33","author":"Tian","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Gidaris, S., Bursuc, A., Puy, G., Komodakis, N., Cord, M., and Perez, P. (2021, January 19\u201325). Obow: Online bag-of-visual-words generation for self-supervised learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00676"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Haeusser, P., Plapp, J., Golkov, V., Aljalbout, E., and Cremers, D. (2018, January 9\u201312). Associative deep clustering: Training a classification network with no labels. Proceedings of the German Conference on Pattern Recognition, Stuttgart, Germany.","DOI":"10.1007\/978-3-030-12939-2_2"},{"key":"ref_49","unstructured":"Ji, X., Henriques, J.F., and Vedaldi, A. (November, January 27). Invariant information clustering for unsupervised image classification and segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Han, S., Park, S., Park, S., Kim, S., and Cha, M. (2020, January 23\u201328). Mitigating embedding and class assignment mismatch in unsupervised image classification. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58586-0_45"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Park, S., Han, S., Kim, S., Kim, D., Park, S., Hong, S., and Cha, M. (2021, January 19\u201325). Improving unsupervised image clustering with robust learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01210"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Chong, S.S. (2022, January 4\u20136). Loss Function Entropy Regularization for Diverse Decision Boundaries. Proceedings of the 2022 7th International Conference on Big Data Analytics (ICBDA), Guangzhou, China.","DOI":"10.1109\/ICBDA55095.2022.9760312"},{"key":"ref_53","unstructured":"Everingham, M., Van Gool, L., Williams, C.K.I., Winn, J., and Zisserman, A. (2024, May 23). The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results. Available online: http:\/\/host.robots.ox.ac.uk\/pascal\/VOC\/voc2007\/."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Chen, D., Mei, J.P., Zhang, H., Wang, C., Feng, Y., and Chen, C. (2022, January 18\u201324). Knowledge distillation with the reused teacher classifier. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01163"},{"key":"ref_55","first-page":"8026","article-title":"Pytorch: An imperative style, high-performance deep learning library","volume":"32","author":"Paszke","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_56","first-page":"1","article-title":"Kymatio: Scattering Transforms in Python","volume":"21","author":"Andreux","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"ref_57","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/26\/6\/466\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:50:08Z","timestamp":1760107808000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/26\/6\/466"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,29]]},"references-count":57,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2024,6]]}},"alternative-id":["e26060466"],"URL":"https:\/\/doi.org\/10.3390\/e26060466","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2024,5,29]]}}}