{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T06:34:27Z","timestamp":1782369267129,"version":"3.54.5"},"reference-count":85,"publisher":"Springer Science and Business Media LLC","issue":"8","license":[{"start":{"date-parts":[[2025,4,27]],"date-time":"2025-04-27T00:00:00Z","timestamp":1745712000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,4,27]],"date-time":"2025-04-27T00:00:00Z","timestamp":1745712000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100000893","name":"Simons Foundation","doi-asserted-by":"publisher","award":["NC-GB-CULM-00002953-02"],"award-info":[{"award-number":["NC-GB-CULM-00002953-02"]}],"id":[{"id":"10.13039\/100000893","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Caltech Chen Institute","award":["Neuroscience Research Grant Award"],"award-info":[{"award-number":["Neuroscience Research Grant Award"]}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["NIH R01 MH123612A"],"award-info":[{"award-number":["NIH R01 MH123612A"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003006","name":"Eidgen\u00f6ssische Technische Hochschule Z\u00fcrich","doi-asserted-by":"publisher","award":["Doc.Mobility Fellowship"],"award-info":[{"award-number":["Doc.Mobility Fellowship"]}],"id":[{"id":"10.13039\/501100003006","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2025,8]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Self-supervised learning (SSL) is a machine learning approach where the data itself provides supervision, eliminating the need for external labels. The model is forced to learn about the data\u2019s inherent structure or context by solving a pretext task. With SSL, models can learn from abundant and cheap unlabeled data, significantly reducing the cost of training models where labels are expensive or inaccessible. In Computer Vision, SSL is widely used as pre-training followed by a downstream task, such as supervised transfer, few-shot learning on smaller labeled data sets, and\/or unsupervised clustering. Unfortunately, it is infeasible to evaluate SSL methods on all possible downstream tasks and objectively measure the quality of the learned representation. Instead, SSL methods are evaluated using in-domain evaluation protocols, such as fine-tuning, linear probing, and k-nearest neighbors (kNN). However, it is not well understood how well these evaluation protocols estimate the representation quality of a pre-trained model for different downstream tasks under different conditions, such as dataset, metric, and model architecture. In this work, we study how classification-based evaluation protocols for SSL correlate and how well they predict downstream performance on different dataset types. Our study includes eleven common image datasets and 26 models that were pre-trained with different SSL methods or have different model backbones. We find that in-domain linear\/kNN probing protocols are, on average, the best general predictors for out-of-domain performance. We further investigate the importance of batch normalization for the various protocols and evaluate how robust correlations are for different kinds of dataset domain shifts. In addition, we challenge assumptions about the relationship between discriminative and generative self-supervised methods, finding that most of their performance differences can be explained by changes to model backbones.<\/jats:p>","DOI":"10.1007\/s11263-025-02402-w","type":"journal-article","created":{"date-parts":[[2025,4,27]],"date-time":"2025-04-27T05:12:15Z","timestamp":1745730735000},"page":"5013-5025","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["A Closer Look at Benchmarking Self-supervised Pre-training with Image Classification"],"prefix":"10.1007","volume":"133","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8016-1637","authenticated-orcid":false,"given":"Markus","family":"Marks","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5447-4549","authenticated-orcid":false,"given":"Manuel","family":"Knott","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Neehar","family":"Kondapaneni","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6623-0966","authenticated-orcid":false,"given":"Elijah","family":"Cole","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9835-5859","authenticated-orcid":false,"given":"Thijs","family":"Defraeye","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8996-5076","authenticated-orcid":false,"given":"Fernando","family":"Perez-Cruz","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pietro","family":"Perona","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,4,27]]},"reference":[{"key":"2402_CR1","unstructured":"Asano, Y. M., Rupprecht, C., & Vedaldi, A. (2020). Self-labelling via simultaneous clustering and representation learning. In International conference on learning representations."},{"issue":"4","key":"2402_CR2","first-page":"340","volume":"2","author":"I Assent","year":"2012","unstructured":"Assent, I. (2012). Clustering high dimensional data. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 2(4), 340\u2013350.","journal-title":"Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery"},{"key":"2402_CR3","unstructured":"Baevski, A., Hsu, W. N., Xu, Q., Babu, A., Gu, J., & Auli, M. (2022). data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language. In International conference on machine learning (ICML)."},{"key":"2402_CR4","unstructured":"Balestriero, R., Ibrahim, M., Sobal, V., Morcos, A., Shekhar, S., Goldstein, T., & Goldblum, M. (2023). A cookbook of self-supervised learning. arXiv preprint arXiv:2304.12210."},{"key":"2402_CR5","unstructured":"Bao, H., Dong, L., & Wei, F. (2021). BEiT: BERT pre-training of image transformers. arXiv preprint arXiv:2106.08254."},{"key":"2402_CR6","unstructured":"Cabannes, V., Kiani, B., Balestriero, R., LeCun, Y., & Bietti, A. (2023). The ssl interplay: Augmentations, inductive bias, and generalization. In International conference on machine learning (pp. 3252\u20133298). PMLR."},{"key":"2402_CR7","doi-asserted-by":"crossref","unstructured":"Caron, M., Bojanowski, P., Joulin, A., & Douze, M. (2018). Deep Clustering for Unsupervised Learning of Visual Features. In European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-030-01264-9_9"},{"key":"2402_CR8","first-page":"9912","volume":"33","author":"M Caron","year":"2020","unstructured":"Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., & Joulin, A. (2020). Unsupervised learning of visual features by contrasting cluster assignments. Advances in Neural Information Processing Systems, 33, 9912\u20139924.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2402_CR9","doi-asserted-by":"crossref","unstructured":"Caron, M., Touvron, H., Misra, I., J\u00e9gou, H., Mairal, J., Bojanowski, P., & Joulin, A. (2021). Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE\/CVF international conference on computer vision (ICCV) (pp. 9650\u20139660).","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"2402_CR10","unstructured":"Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., & Sutskever, I. (2020a). Generative pretraining from pixels. In Proceedings of the 37th international conference on machine learning (Vol.\u00a0119, pp. 1691\u20131703). PMLR."},{"key":"2402_CR11","unstructured":"Chen, R. J., Ding, T., Lu, M. Y., Williamson, D. F. K., Jaume, G., Chen, B., & Mahmood, F. (2023). A general-purpose self-supervised model for computational pathology. arXiv preprint arXiv:2308.15474."},{"key":"2402_CR12","unstructured":"Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020b). A simple framework for contrastive learning of visual representations. In Proceedings of the 37th international conference on machine learning (Vol.\u00a0119, pp. 1597\u20131607). PMLR."},{"key":"2402_CR13","first-page":"22243","volume":"33","author":"T Chen","year":"2020","unstructured":"Chen, T., Kornblith, S., Swersky, K., Norouzi, M., & Hinton, G. E. (2020). Big self-supervised models are strong semi-supervised learners. Advances in Neural Information Processing Systems, 33, 22243\u201322255.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"1","key":"2402_CR14","doi-asserted-by":"publisher","first-page":"208","DOI":"10.1007\/s11263-023-01852-4","volume":"132","author":"X Chen","year":"2024","unstructured":"Chen, X., Ding, M., Wang, X., Xin, Y., Mo, S., Wang, Y., & Wang, J. (2024). Context autoencoder for self-supervised representation learning. International Journal of Computer Vision, 132(1), 208\u2013223.","journal-title":"International Journal of Computer Vision"},{"key":"2402_CR15","unstructured":"Chen, X., Fan, H., Girshick, R., & He, K. (2020). Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297."},{"key":"2402_CR16","first-page":"15745","volume":"2021","author":"X Chen","year":"2020","unstructured":"Chen, X., & He, K. (2020). Exploring simple siamese representation learning. IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, 15745\u201315753.","journal-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"2402_CR17","doi-asserted-by":"crossref","unstructured":"Chen, X., Xie, S., & He, K. (2021). An empirical study of training self-supervised vision transformers. In 2021 IEEE\/CVF international conference on computer vision (ICCV) (pp. 9620\u20139629).","DOI":"10.1109\/ICCV48922.2021.00950"},{"key":"2402_CR18","doi-asserted-by":"crossref","unstructured":"Cole, E., Yang, X., Wilber, K., Mac\u00a0Aodha, O., & Belongie, S. (2022). When does contrastive visual representation learning work? In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR) (pp. 14755\u201314764).","DOI":"10.1109\/CVPR52688.2022.01434"},{"key":"2402_CR19","doi-asserted-by":"crossref","unstructured":"Dalal, N., & Triggs, B. (2005). Histograms of oriented gradients for human detection. In 2005 IEEE computer society conference on computer vision and pattern recognition (Vol.\u00a01, pp. 886\u2013893).","DOI":"10.1109\/CVPR.2005.177"},{"key":"2402_CR20","unstructured":"Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: Human language technologies (pp. 4171\u20134186)."},{"key":"2402_CR21","first-page":"1422","volume":"2015","author":"C Doersch","year":"2015","unstructured":"Doersch, C., Gupta, A., & Efros, A. A. (2015). Unsupervised visual representation learning by context prediction. IEEE International Conference on Computer Vision (ICCV), 2015, 1422\u20131430.","journal-title":"IEEE International Conference on Computer Vision (ICCV)"},{"key":"2402_CR22","doi-asserted-by":"crossref","unstructured":"Dong, X., Bao, J., Zhang, T., Chen, D., Zhang, W., Yuan, L., & Guo, B. (2023). PeCo: Perceptual codebook for BERT pre-training of vision transformers. In Proceedings of the AAAI conference on artificial intelligence (Vol.\u00a037, pp. 552\u2013560).","DOI":"10.1609\/aaai.v37i1.25130"},{"key":"2402_CR23","doi-asserted-by":"crossref","unstructured":"Ericsson, L., Gouk, H., & Hospedales, T. M. (2021). How well do self-supervised models transfer? In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 5414\u20135423).","DOI":"10.1109\/CVPR46437.2021.00537"},{"issue":"2","key":"2402_CR24","doi-asserted-by":"publisher","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","volume":"88","author":"M Everingham","year":"2010","unstructured":"Everingham, M., Van Gool, L., Williams, C., Winn, J., & Zisserman, A. (2010). The pascal visual object classes (VOC) challenge. International Journal of Computer Vision, 88(2), 303\u2013338. https:\/\/doi.org\/10.1007\/s11263-009-0275-4","journal-title":"International Journal of Computer Vision"},{"key":"2402_CR25","first-page":"19358","volume":"2023","author":"Y Fang","year":"2022","unstructured":"Fang, Y., Wang, W., Xie, B., Sun, Q. S., Wu, L. Y., Wang, X., & Cao, Y. (2022). EVA: Exploring the limits of masked visual representation learning at scale. IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, 19358\u201319369.","journal-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"2402_CR26","first-page":"35946","volume":"32","author":"C Feichtenhofer","year":"2022","unstructured":"Feichtenhofer, C., Fan, H., Li, Y., & He, K. (2022). Masked autoencoders as spatiotemporal learners. Advances in Neural Information Processing Systems, 32, 35946\u201335958.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2402_CR27","doi-asserted-by":"crossref","unstructured":"Gansbeke, W. V., Vandenhende, S., Georgoulis, S., Proesmans, M., & Gool, L. V. (2020). SCAN: Learning to classify images without labels. In European conference on computer vision.","DOI":"10.1007\/978-3-030-58607-2_16"},{"key":"2402_CR28","unstructured":"Gidaris, S., Singh, P., & Komodakis, N. (2018). Unsupervised representation learning by predicting image rotations. In International conference on learning representations."},{"key":"2402_CR29","unstructured":"Goldblum, M., Souri, H., Ni, R., Shu, M., Prabhu, V., Somepalli, G., & Goldstein, T. (2024). Battle of the backbones: A large-scale comparison of pretrained models across computer vision tasks. Advances in Neural Information Processing Systems, 36, ,"},{"key":"2402_CR30","unstructured":"Goyal, P., Caron, M., Lefaudeux, B., Xu, M., Wang, P., Pai, V., & Bojanowski, P. (2021). Self-supervised pretraining of visual features in the wild. arXiv preprint arXiv:2103.01988."},{"key":"2402_CR31","unstructured":"Griffin, G., Holub, A., & Perona, P. (2022). Caltech 256. CaltechDATA. [2023-06-05] https:\/\/data.caltech.edu\/records\/20087"},{"key":"2402_CR32","unstructured":"Grill, J. B., Strub, F., Altch\u00e9, F., Tallec, C., Richemond, P. H., Buchatskaya, E., & Valko, M. (2020). Bootstrap your own latent a new approach to self-supervised learning. In Proceedings of the 34th international conference on neural information processing systems."},{"key":"2402_CR33","doi-asserted-by":"crossref","unstructured":"Gwilliam, M., & Shrivastava, A. (2022). Beyond supervised vs. unsupervised: Representative benchmarking and analysis of image representation learning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 9642\u20139652).","DOI":"10.1109\/CVPR52688.2022.00942"},{"key":"2402_CR34","unstructured":"HaoChen, J. Z., Wei, C., Kumar, A., & Ma, T. (2022). Beyond separability: Analyzing the linear transferability of contrastive representations to related subpopulations. arXiv preprint arXiv:2204.02683."},{"key":"2402_CR35","first-page":"15979","volume":"2022","author":"K He","year":"2021","unstructured":"He, K., Chen, X., Xie, S., Li, Y., Doll\u00e1r, P., & Girshick, R. B. (2021). Masked autoencoders are scalable vision learners. IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, 15979\u201315988.","journal-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"2402_CR36","first-page":"9726","volume":"2020","author":"K He","year":"2019","unstructured":"He, K., Fan, H., Wu, Y., Xie, S., & Girshick, R. B. (2019). Momentum contrast for unsupervised visual representation learning. IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, 9726\u20139735.","journal-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)"},{"issue":"1","key":"2402_CR37","doi-asserted-by":"publisher","first-page":"6456","DOI":"10.1038\/s41467-021-26751-5","volume":"12","author":"I Higgins","year":"2021","unstructured":"Higgins, I., Chang, L., Langston, V., Hassabis, D., Summerfield, C., Tsao, D., & Botvinick, M. (2021). Unsupervised deep learning identifies semantic disentanglement in single inferotemporal face patch neurons. Nature Communications, 12(1), 6456.","journal-title":"Nature Communications"},{"key":"2402_CR38","unstructured":"Hou, Z., Sun, F., Chen, Y. K., Xie, Y., & Kung, S. Y. (2022). MILAN: Masked image pretraining on language assisted representation. arXiv preprint arXiv:2208.06049."},{"key":"2402_CR39","unstructured":"Ibrahim, M., Garrido, Q., Morcos, A., & Bouchacourt, D. (2022). The robustness limits of sota vision models to natural variation. arXiv preprint arXiv:2210.13604."},{"key":"2402_CR40","doi-asserted-by":"publisher","first-page":"4037","DOI":"10.1109\/TPAMI.2020.2992393","volume":"43","author":"L Jing","year":"2019","unstructured":"Jing, L., & Tian, Y. (2019). Self-supervised visual feature learning with deep neural networks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43, 4037\u20134058.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2402_CR41","doi-asserted-by":"crossref","unstructured":"Kim, D., Wang, K., Sclaroff, S., & Saenko, K. (2022). A broad study of pre-training for domain generalization and adaptation. In European conference on computer vision (pp. 621\u2013638).","DOI":"10.1007\/978-3-031-19827-4_36"},{"key":"2402_CR42","first-page":"1920","volume":"2019","author":"A Kolesnikov","year":"2019","unstructured":"Kolesnikov, A., Zhai, X., & Beyer, L. (2019). Revisiting self-supervised visual representation learning. IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 1920\u20131929.","journal-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"2402_CR43","unstructured":"Krizhevsky, A., & Hinton, G. (2009). Learning multiple layers of features from tiny images. https:\/\/www.cs.utoronto.ca\/~kriz\/learning-features-2009-TR.pdf"},{"key":"2402_CR44","unstructured":"Kumar, A., Raghunathan, A., Jones, R., Ma, T., & Liang, P. (2022). Fine-tuning can distort pretrained features and underperform out-of-distribution. arXiv preprint arXiv:2202.10054."},{"key":"2402_CR45","unstructured":"Lee, J. H., Yoon, D., Ji, B., Kim, K., & Hwang, S. (2023). Rethinking evaluation protocols of visual representations learned via self-supervised learning. arXiv preprint arXiv:2304.03456."},{"key":"2402_CR46","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2019.103853","volume":"93","author":"X Li","year":"2020","unstructured":"Li, X., Grandvalet, Y., Davoine, F., Cheng, J., Cui, Y., Zhang, H., & Yang, M. H. (2020). Transfer learning in computer vision tasks: Remember where you come from. Image and Vision Computing, 93, 103853. https:\/\/doi.org\/10.1016\/j.imavis.2019.103853","journal-title":"Image and Vision Computing"},{"issue":"1","key":"2402_CR47","first-page":"857","volume":"35","author":"X Liu","year":"2021","unstructured":"Liu, X., Zhang, F., Hou, Z., Mian, L., Wang, Z., Zhang, J., & Tang, J. (2021). Self-supervised learning: Generative or contrastive. IEEE Transactions on Knowledge and Data Engineering, 35(1), 857\u2013876.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"2402_CR48","unstructured":"Liu, Y., Zhang, S., Chen, J., Chen, K., & Lin, D. (2023). Pixmim: Rethinking pixel reconstruction in masked image modeling. arXiv preprint arXiv:2303.02416."},{"key":"2402_CR49","unstructured":"Locatello, F., Bauer, S., Lucic, M., Raetsch, G., Gelly, S., Sch\u00f6lkopf, B., & Bachem, O. (2019). Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations. In Proceedings of the 36th international conference on machine learning (ICML) (pp. 4114\u20134124). PMLR."},{"key":"2402_CR50","unstructured":"Miller, J. P., Taori, R., Raghunathan, A., Sagawa, S., Koh, P. W., Shankar, V., & Schmidt, L. (2021). Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization. In International conference on machine learning (pp. 7721\u20137735). PMLR."},{"key":"2402_CR51","doi-asserted-by":"crossref","unstructured":"Misra, I., & Maaten, L. V. D. (2019). Self-supervised learning of pretext-invariant representations. In IEEE\/CVF conference on computer vision and pattern recognition (CVPR),2020, 6706\u20136716.","DOI":"10.1109\/CVPR42600.2020.00674"},{"key":"2402_CR52","unstructured":"MMSelfSup Contributors (2021). MMSelfSup: OpenMMLab Self-Supervised Learning Toolbox and Benchmark. https:\/\/github.com\/open-mmlab\/mmselfsup"},{"key":"2402_CR53","doi-asserted-by":"crossref","unstructured":"Musgrave, K., Belongie, S., & Lim, S. N. (2020). A metric learning reality check. In European conference on computer vision (ECCV) (pp. 681\u2013699).","DOI":"10.1007\/978-3-030-58595-2_41"},{"key":"2402_CR54","doi-asserted-by":"crossref","unstructured":"Newell, A., & Deng, J. (2020). How useful is self-supervised pretraining for visual tasks? In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR42600.2020.00737"},{"key":"2402_CR55","doi-asserted-by":"crossref","unstructured":"Noroozi, M., & Favaro, P. (2016). Unsupervised learning of visual representations by solving jigsaw puzzles. In European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-319-46466-4_5"},{"key":"2402_CR56","unstructured":"Oord, A. V. D., Li, Y., & Vinyals, O. (2018). Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748."},{"issue":"10","key":"2402_CR57","doi-asserted-by":"publisher","first-page":"805","DOI":"10.1038\/s41592-018-0109-9","volume":"15","author":"C Pandarinath","year":"2018","unstructured":"Pandarinath, C., O\u2019Shea, D. J., Collins, J., Jozefowicz, R., Stavisky, S. D., Kao, J. C., & Sussillo, D. (2018). Inferring single-trial neural population dynamics using sequential auto-encoders. Nature Methods, 15(10), 805\u2013815.","journal-title":"Nature Methods"},{"key":"2402_CR58","first-page":"8026","volume":"32","author":"A Paszke","year":"2019","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., & Chintala, S. (2019). PyTorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems (NeurIPS), 32, 8026\u20138037.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"2402_CR59","first-page":"2536","volume":"2016","author":"D Pathak","year":"2016","unstructured":"Pathak, D., Kr\u00e4henb\u00fchl, P., Donahue, J., Darrell, T., & Efros, A. A. (2016). Context encoders: Feature learning by inpainting. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, 2536\u20132544.","journal-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"2402_CR60","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., & Dubourg, V. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825\u20132830.","journal-title":"Journal of Machine Learning Research"},{"key":"2402_CR61","unstructured":"Peng, Z., Dong, L., Bao, H., Ye, Q., & Wei, F. (2022). BEiT v2: masked image modeling with vector-quantized visual tokenizers. arXiv preprint arXiv:2208.06366."},{"issue":"10","key":"2402_CR62","doi-asserted-by":"publisher","first-page":"1872","DOI":"10.1007\/s11431-020-1647-3","volume":"63","author":"X Qiu","year":"2020","unstructured":"Qiu, X., Sun, T., Xu, Y., Shao, Y., Dai, N., & Huang, X. (2020). Pre-trained models for natural language processing: A survey. SCIENCE CHINA Technological Sciences, 63(10), 1872\u20131897.","journal-title":"SCIENCE CHINA Technological Sciences"},{"key":"2402_CR63","unstructured":"Recht, B., Roelofs, R., Schmidt, L., & Shankar, V. (2019). Do imagenet classifiers generalize to imagenet? In International conference on machine learning (pp. 5389\u20135400). PMLR."},{"key":"2402_CR64","unstructured":"Rusak, E., Schneider, S., Gehler, P. V., Bringmann, O., Brendel, W., & Bethge, M. (2022). ImageNet-D: A new challenging robustness dataset inspired by domain adaptation. ICML 2022 Shift Happens Workshop."},{"issue":"3","key":"2402_CR65","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., & Fei-Fei, L. (2015). ImageNet large scale visual recognition challenge. International Journal of Computer Vision, 115(3), 211\u2013252. https:\/\/doi.org\/10.1007\/s11263-015-0816-y","journal-title":"International Journal of Computer Vision"},{"key":"2402_CR66","unstructured":"Shi, Y., Daunhawer, I., Vogt, J. E., Torr, P., & Sanyal, A. (2022). How robust are pre-trained models to distribution shift? ICML 2022: Workshop on Spurious Correlations, Invariance, and Stability."},{"key":"2402_CR67","unstructured":"Sun, J. J., Marks, M., Ulmer, A., Chakraborty, D., Geuther, B., Hayes, E., & Kennedy, A. (2023). MABe22: A multi-species multi-task benchmark for learned representations of behavior. In International conference on machine learning (ICML)."},{"key":"2402_CR68","doi-asserted-by":"crossref","unstructured":"Van\u00a0Horn, G., Cole, E., Beery, S., Wilber, K., Belongie, S., & Mac\u00a0Aodha, O. (2021). Benchmarking representation learning for natural world image collections. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 12884\u201312893).","DOI":"10.1109\/CVPR46437.2021.01269"},{"key":"2402_CR69","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., & Polosukhin, I. (2017). Attention is All you Need. In Advances in neural information processing systems (NeurIPS) (pp. 6000\u20136010)."},{"key":"2402_CR70","unstructured":"Wah, C., Branson, S., Welinder, P., Perona, P., & Belongie, S. (2011). The Caltech-UCSD Birds-200-2011 Dataset. California Institute of Technology. (CNS-TR-2011-001)"},{"key":"2402_CR71","first-page":"3023","volume":"2021","author":"X Wang","year":"2020","unstructured":"Wang, X., Zhang, R., Shen, C., Kong, T., & Li, L. (2020). Dense contrastive learning for self-supervised visual pre-training. IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, 3023\u20133032.","journal-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)"},{"issue":"4","key":"2402_CR72","doi-asserted-by":"publisher","first-page":"213","DOI":"10.1109\/MGRS.2022.3198244","volume":"10","author":"Y Wang","year":"2022","unstructured":"Wang, Y., Albrecht, C. M., Braham, N., Mou, L., & Zhu, X. X. (2022). Self-supervised learning in remote sensing: A review. IEEE Geoscience and Remote Sensing Magazine, 10(4), 213\u2013247.","journal-title":"IEEE Geoscience and Remote Sensing Magazine"},{"key":"2402_CR73","doi-asserted-by":"publisher","unstructured":"Wei, C., Fan, H., Xie, S., Wu, C. Y., Yuille, A., & Feichtenhofer, C. (2022). Masked feature prediction for self-supervised visual pre-training. In 2022 IEEE\/CVF conference on computer vision and pattern recognition (CVPR) (pp. 14648\u201314658). https:\/\/doi.org\/10.1109\/CVPR52688.2022.01426","DOI":"10.1109\/CVPR52688.2022.01426"},{"key":"2402_CR74","unstructured":"Wightman, R. (2019). PyTorch Image Models. GitHub. https:\/\/github.com\/rwightman\/pytorch-image-models"},{"key":"2402_CR75","first-page":"3733","volume":"2018","author":"Z Wu","year":"2018","unstructured":"Wu, Z., Xiong, Y., Yu, S. X., & Lin, D. (2018). Unsupervised feature learning via non-parametric instance discrimination. IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2018, 3733\u20133742.","journal-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"2402_CR76","doi-asserted-by":"crossref","unstructured":"Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., & Hu, H. (2021). SimMIM: A simple framework for masked image modeling. In 2022 IEEE\/CVF conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR52688.2022.00943"},{"key":"2402_CR77","first-page":"6508","volume":"2020","author":"X Yan","year":"2019","unstructured":"Yan, X., Misra, I., Gupta, A. K., Ghadiyaram, D., & Mahajan, D. K. (2019). ClusterFit: Improving generalization of visual representations. IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, 6508\u20136517.","journal-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"2402_CR78","doi-asserted-by":"crossref","unstructured":"Yang, L., Zhang, S., Qin, L., Li, Y., Wang, Y., Liu, H., & Zhang, Y. (2022). GLUE-X: Evaluating natural language understanding models from an out-of-distribution generalization perspective. arXiv preprint arXiv:2211.08073.","DOI":"10.18653\/v1\/2023.findings-acl.806"},{"key":"2402_CR79","unstructured":"Yosinski, J., Clune, J., Bengio, Y., & Lipson, H. (2014). How transferable are features in deep neural networks? Advances in Neural Information Processing Systems (Vol.\u00a027)."},{"key":"2402_CR80","unstructured":"Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., & Wu, Y. (2021). Vector-quantized image modeling with improved VQGAN. arXiv preprint arXiv:2110.04627."},{"key":"2402_CR81","doi-asserted-by":"crossref","unstructured":"Yu, X., Tang, L., Rao, Y., Huang, T., Zhou, J., & Lu, J. (2022). Point-bert: Pre-training 3d point cloud transformers with masked point modeling. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 19313\u201319322).","DOI":"10.1109\/CVPR52688.2022.01871"},{"key":"2402_CR82","unstructured":"Zbontar, J., Jing, L., Misra, I., LeCun, Y., & Deny, S. (2021). Barlow twins: Self-supervised learning via redundancy reduction. In International conference on machine learning."},{"key":"2402_CR83","first-page":"645","volume":"2017","author":"R Zhang","year":"2016","unstructured":"Zhang, R., Isola, P., & Efros, A. A. (2016). Split-brain autoencoders: Unsupervised learning by cross-channel prediction. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, 645\u2013654.","journal-title":"IEEE Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"2402_CR84","unstructured":"Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., & Kong, T. (2021). iBOT: Image BERT pre-training with online Tokenizer. arXiv preprint arXiv:2111.07832."},{"key":"2402_CR85","unstructured":"Zhou, Y., Chia, M. A., Wagner, S. K., Ayhan, M. S., Williamson, D. J., Struyven, R. R., & Keane, P. A. (2023). A foundation model for generalizable disease detection from retinal images. Nature, 622(7981), 156\u2013163."}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02402-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-025-02402-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02402-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,6]],"date-time":"2025-09-06T12:21:01Z","timestamp":1757161261000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-025-02402-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,27]]},"references-count":85,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2025,8]]}},"alternative-id":["2402"],"URL":"https:\/\/doi.org\/10.1007\/s11263-025-02402-w","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,27]]},"assertion":[{"value":"31 July 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 February 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 April 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}