{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T07:18:06Z","timestamp":1777360686243,"version":"3.51.4"},"reference-count":117,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2023,9,26]],"date-time":"2023-09-26T00:00:00Z","timestamp":1695686400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,9,26]],"date-time":"2023-09-26T00:00:00Z","timestamp":1695686400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Machine Vision and Applications"],"published-print":{"date-parts":[[2023,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Contrastive representation learning has proven to be an effective self-supervised learning method for images and videos. Most successful approaches are based on Noise Contrastive Estimation (NCE) and use different views of an instance as positives that should be contrasted with other instances, called negatives, that are considered as noise. However, several instances in a dataset are drawn from the same distribution and share underlying semantic information. A good data representation should contain relations between the instances, or semantic similarity and dissimilarity, that contrastive learning harms by considering all negatives as noise. To circumvent this issue, we propose a novel formulation of contrastive learning using semantic similarity between instances called Similarity Contrastive Estimation (SCE). Our training objective is a soft contrastive one that brings the positives closer and estimates a continuous distribution to push or pull negative instances based on their learned similarities. We validate empirically our approach on both image and video representation learning. We show that SCE performs competitively with the state of the art on the ImageNet linear evaluation protocol for fewer pretraining epochs and that it generalizes to several downstream image tasks. We also show that SCE reaches state-of-the-art results for pretraining video representation and that the learned representation can generalize to video downstream tasks. Source code is available here: <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/juliendenize\/eztorch\">https:\/\/github.com\/juliendenize\/eztorch<\/jats:ext-link>. \n<\/jats:p>","DOI":"10.1007\/s00138-023-01444-9","type":"journal-article","created":{"date-parts":[[2023,9,26]],"date-time":"2023-09-26T08:04:04Z","timestamp":1695715444000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Similarity contrastive estimation for image and video soft contrastive self-supervised learning"],"prefix":"10.1007","volume":"34","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2807-1826","authenticated-orcid":false,"given":"Julien","family":"Denize","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jaonary","family":"Rabarisoa","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8652-2250","authenticated-orcid":false,"given":"Astrid","family":"Orcesi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Romain","family":"H\u00e9rault","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,9,26]]},"reference":[{"key":"1444_CR1","unstructured":"Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., Joulin, A.: Unsupervised learning of visual features by contrasting cluster assignments. In: Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (2020)"},{"key":"1444_CR2","unstructured":"Grill, J., Strub, F., Altch\u00e9, F., Tallec, C., Richemond, P.H., Buchatskaya, E., Doersch, C., Pires, B.\u00c1., Guo, Z., Azar, M.G., Piot, B., Kavukcuoglu, K., Munos, R., Valko, M.: Bootstrap your own latent\u2014a new approach to self-supervised learning. In: Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (2020)"},{"key":"1444_CR3","doi-asserted-by":"publisher","unstructured":"Caron, M., Touvron, H., Misra, I., J\u00e9gou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers. In: Proceedings of the International Conference on Computer Vision, pp. 6706\u20136716 (2021). https:\/\/doi.org\/10.1109\/ICCV48922.2021.00951","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"1444_CR4","doi-asserted-by":"publisher","unstructured":"Feichtenhofer, C., Fan, H., Xiong, B., Girshick, R.B., He, K.: A large-scale study on unsupervised spatiotemporal representation learning. In: Conference on Computer Vision and Pattern Recognition, pp. 3299\u20133309 (2021). https:\/\/doi.org\/10.1109\/CVPR46437.2021.00331","DOI":"10.1109\/CVPR46437.2021.00331"},{"key":"1444_CR5","doi-asserted-by":"publisher","unstructured":"Duan, H., Zhao, N., Chen, K., Lin, D.: Transrank: self-supervised video representation learning via ranking-based transformation recognition. In: Conference on Computer Vision and Pattern Recognition, pp. 2990\u20133000 (2022). https:\/\/doi.org\/10.1109\/CVPR52688.2022.00301","DOI":"10.1109\/CVPR52688.2022.00301"},{"key":"1444_CR6","unstructured":"Gutmann, M., Hyv\u00e4rinen, A.: Noise-contrastive estimation: a new estimation principle for unnormalized statistical models. In: 13th International Conference on Artificial Intelligence and Statistics, pp. 297\u2013304 (2010)"},{"key":"1444_CR7","doi-asserted-by":"publisher","unstructured":"He, K., Fan, H., Wu, Y., Xie, S., Girshick, R.B.: Momentum contrast for unsupervised visual representation learning. In: Conference on Computer Vision and Pattern Recognition, pp. 9726\u20139735 (2020). https:\/\/doi.org\/10.1109\/CVPR42600.2020.00975","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"1444_CR8","unstructured":"Chen, T., Kornblith, S., Norouzi, M., Hinton, G.E.: A simple framework for contrastive learning of visual representations. In: Proceedings of the 37th International Conference on Machine Learning, pp. 1597\u20131607 (2020)"},{"key":"1444_CR9","unstructured":"Yang, C., Xu, Y., Dai, B., Zhou, B.: Video representation learning with visual tempo consistency. arXiv:2006.15489 (2020)"},{"key":"1444_CR10","doi-asserted-by":"publisher","unstructured":"Han, T., Xie, W., Zisserman, A.: Memory-augmented dense predictive coding for video representation learning. In: Vedaldi, A., Bischof, H., Brox, T., Frahm, J. (eds.) 16th European Conference on Computer Vision, pp. 312\u2013329 (2020). https:\/\/doi.org\/10.1007\/978-3-030-58580-8_19","DOI":"10.1007\/978-3-030-58580-8_19"},{"key":"1444_CR11","unstructured":"Tian, Y., Sun, C., Poole, B., Krishnan, D., Schmid, C., Isola, P.: What makes for good views for contrastive learning? In: Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (2020)"},{"key":"1444_CR12","doi-asserted-by":"publisher","unstructured":"Qian, R., Meng, T., Gong, B., Yang, M., Wang, H., Belongie, S.J., Cui, Y.: Spatiotemporal contrastive video representation learning. In: Conference on Computer Vision and Pattern Recognition, pp. 6964\u20136974 (2021). https:\/\/doi.org\/10.1109\/CVPR46437.2021.00689","DOI":"10.1109\/CVPR46437.2021.00689"},{"key":"1444_CR13","doi-asserted-by":"publisher","unstructured":"Pan, T., Song, Y., Yang, T., Jiang, W., Liu, W.: Videomoco: Contrastive video representation learning with temporally adversarial examples. In: Conference on Computer Vision and Pattern Recognition, pp. 11205\u201311214 (2021). https:\/\/doi.org\/10.1109\/CVPR46437.2021.01105","DOI":"10.1109\/CVPR46437.2021.01105"},{"key":"1444_CR14","doi-asserted-by":"publisher","unstructured":"Recasens, A., Luc, P., Alayrac, J., Wang, L., Strub, F., Tallec, C., Malinowski, M., Patraaucean, V., Altch\u00e9, F., Valko, M., Grill, J., Oord, A., Zisserman, A.: Broaden your views for self-supervised video learning. In: International Conference on Computer Vision, pp. 1235\u20131245 (2021). https:\/\/doi.org\/10.1109\/ICCV48922.2021.00129","DOI":"10.1109\/ICCV48922.2021.00129"},{"key":"1444_CR15","doi-asserted-by":"publisher","unstructured":"Dave, I.R., Gupta, R., Rizve, M.N., Shah, M.: TCLR: temporal contrastive learning for video representation. Comput. Vis. Image Underst. (2022). https:\/\/doi.org\/10.1016\/j.cviu.2022.103406","DOI":"10.1016\/j.cviu.2022.103406"},{"key":"1444_CR16","unstructured":"Oord, A., Li, Y., Vinyals, O.: Representation learning with contrastive predictive coding. arXiv:1807.03748 (2018)"},{"key":"1444_CR17","doi-asserted-by":"publisher","unstructured":"Wu, Z., Xiong, Y., Yu, S.X., Lin, D.: Unsupervised feature learning via non-parametric instance discrimination. In: Conference on Computer Vision and Pattern Recognition, pp. 3733\u20133742 (2018). https:\/\/doi.org\/10.1109\/CVPR.2018.00393","DOI":"10.1109\/CVPR.2018.00393"},{"key":"1444_CR18","unstructured":"Kalantidis, Y., Sariyildiz, M.B., Pion, N., Weinzaepfel, P., Larlus, D.: Hard negative mixing for contrastive learning. In: Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (2020)"},{"key":"1444_CR19","unstructured":"Robinson, J.D., Chuang, C., Sra, S., Jegelka, S.: Contrastive learning with hard negative samples. In: 9th International Conference on Learning Representations (2021)"},{"key":"1444_CR20","doi-asserted-by":"crossref","unstructured":"Wu, M., Mosse, M., Zhuang, C., Yamins, D., Goodman, N.D.: Conditional negative sampling for contrastive learning of visual representations. In: 9th International Conference on Learning Representations (2021)","DOI":"10.1109\/ICCV48922.2021.00999"},{"key":"1444_CR21","doi-asserted-by":"crossref","unstructured":"Hu, Q., Wang, X., Hu, W., Qi, G.: Adco: Adversarial contrast for efficient learning of unsupervised representations from self-trained negative adversaries. In: Conference on Computer Vision and Pattern Recognition, pp. 1074\u20131083 (2021)","DOI":"10.1109\/CVPR46437.2021.00113"},{"key":"1444_CR22","doi-asserted-by":"publisher","unstructured":"Dwibedi, D., Aytar, Y., Tompson, J., Sermanet, P., Zisserman, A.: With a little help from my friends: nearest-neighbor contrastive learning of visual representations. In: 2021 International Conference on Computer Vision, pp. 9568\u20139577 (2021). https:\/\/doi.org\/10.1109\/ICCV48922.2021.00945","DOI":"10.1109\/ICCV48922.2021.00945"},{"key":"1444_CR23","unstructured":"Cai, T.T., Frankle, J., Schwab, D.J., Morcos, A.S.: Are all negatives created equal in contrastive instance discrimination? arXiv:2010.06682 (2020)"},{"key":"1444_CR24","unstructured":"Wei, C., Wang, H., Shen, W., Yuille, A.L.: CO2: consistent contrast for unsupervised visual representation learning. In: 9th International Conference on Learning Representations (2021)"},{"key":"1444_CR25","unstructured":"Chuang, C., Robinson, J., Lin, Y., Torralba, A., Jegelka, S.: Debiased contrastive learning. In: Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (2020)"},{"key":"1444_CR26","doi-asserted-by":"publisher","unstructured":"Toering, M., Gatopoulos, I., Stol, M., Hu, V.T.: Self-supervised video representation learning with cross-stream prototypical contrasting. In: Winter Conference on Applications of Computer Vision, pp. 846\u2013856 (2022). https:\/\/doi.org\/10.1109\/WACV51458.2022.00092","DOI":"10.1109\/WACV51458.2022.00092"},{"key":"1444_CR27","doi-asserted-by":"publisher","unstructured":"Chen, X., He, K.: Exploring simple SIAMESE representation learning. In: Conference on Computer Vision and Pattern Recognition, pp. 15750\u201315758 (2021). https:\/\/doi.org\/10.1109\/CVPR46437.2021.01549","DOI":"10.1109\/CVPR46437.2021.01549"},{"key":"1444_CR28","unstructured":"Zheng, M., You, S., Wang, F., Qian, C., Zhang, C., Wang, X., Xu, C.: RESSL: relational self-supervised learning with weak augmentation. In: Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems, pp. 2543\u20132555 (2021)"},{"key":"1444_CR29","unstructured":"Chen, X., Fan, H., Girshick, R.B., He, K.: Improved baselines with momentum contrastive learning. arXiv:2003.04297 (2020)"},{"issue":"9","key":"1444_CR30","doi-asserted-by":"publisher","first-page":"1734","DOI":"10.1109\/TPAMI.2015.2496141","volume":"38","author":"A Dosovitskiy","year":"2016","unstructured":"Dosovitskiy, A., Fischer, P., Springenberg, J.T., Riedmiller, M.A., Brox, T.: Discriminative unsupervised feature learning with exemplar convolutional neural networks. Trans. Pattern Anal. Mach. Intell. 38(9), 1734\u20131747 (2016). https:\/\/doi.org\/10.1109\/TPAMI.2015.2496141","journal-title":"Trans. Pattern Anal. Mach. Intell."},{"key":"1444_CR31","doi-asserted-by":"publisher","unstructured":"Doersch, C., Gupta, A., Efros, A.A.: Unsupervised visual representation learning by context prediction. In: International Conference on Computer Vision, pp. 1422\u20131430 (2015). https:\/\/doi.org\/10.1109\/ICCV.2015.167","DOI":"10.1109\/ICCV.2015.167"},{"key":"1444_CR32","doi-asserted-by":"publisher","unstructured":"Zhang, R., Isola, P., Efros, A.A.: Colorful image colorization. In: 14th European Conference on Computer Vision, pp. 649\u2013666 (2016). https:\/\/doi.org\/10.1007\/978-3-319-46487-9_40","DOI":"10.1007\/978-3-319-46487-9_40"},{"key":"1444_CR33","doi-asserted-by":"publisher","unstructured":"Noroozi, M., Favaro, P.: Unsupervised learning of visual representations by solving jigsaw puzzles. In: 14th European Conference on Computer Vision, pp. 69\u201384 (2016). https:\/\/doi.org\/10.1007\/978-3-319-46466-4_5","DOI":"10.1007\/978-3-319-46466-4_5"},{"key":"1444_CR34","doi-asserted-by":"publisher","unstructured":"Noroozi, M., Pirsiavash, H., Favaro, P.: Representation learning by learning to count. In: International Conference on Computer Vision, pp. 5899\u20135907 (2017). https:\/\/doi.org\/10.1109\/ICCV.2017.628","DOI":"10.1109\/ICCV.2017.628"},{"key":"1444_CR35","unstructured":"Gidaris, S., Singh, P., Komodakis, N.: Unsupervised representation learning by predicting image rotations. In: 6th International Conference on Learning Representations (2018)"},{"key":"1444_CR36","unstructured":"Hjelm, R.D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Bachman, P., Trischler, A., Bengio, Y.: Learning deep representations by mutual information estimation and maximization. In: 7th International Conference on Learning Representations (2019)"},{"key":"1444_CR37","doi-asserted-by":"publisher","unstructured":"Tian, Y., Krishnan, D., Isola, P.: Contrastive multiview coding. In: 16th European Conference on Computer Vision, pp. 776\u2013794 (2020). https:\/\/doi.org\/10.1007\/978-3-030-58621-8_45","DOI":"10.1007\/978-3-030-58621-8_45"},{"key":"1444_CR38","doi-asserted-by":"publisher","unstructured":"Misra, I., Maaten, L.: Self-supervised learning of pretext-invariant representations. In: Conference on Computer Vision and Pattern Recognition, pp. 6706\u20136716 (2020). https:\/\/doi.org\/10.1109\/CVPR42600.2020.00674","DOI":"10.1109\/CVPR42600.2020.00674"},{"key":"1444_CR39","doi-asserted-by":"publisher","unstructured":"Wang, G., Wang, K., Wang, G., Torr, P.H.S., Lin, L.: Solving inefficiency of self-supervised representation learning. In: International Conference on Computer Vision, pp. 9485\u20139495 (2021). https:\/\/doi.org\/10.1109\/ICCV48922.2021.00937","DOI":"10.1109\/ICCV48922.2021.00937"},{"key":"1444_CR40","unstructured":"Chen, T., Kornblith, S., Swersky, K., Norouzi, M., Hinton, G.E.: Big self-supervised models are strong semi-supervised learners. In: Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (2020)"},{"key":"1444_CR41","doi-asserted-by":"publisher","unstructured":"Chen, X., Xie, S., He, K.: An empirical study of training self-supervised vision transformers. In: International Conference on Computer Vision, pp. 9620\u20139629 (2021). https:\/\/doi.org\/10.1109\/ICCV48922.2021.00950","DOI":"10.1109\/ICCV48922.2021.00950"},{"key":"1444_CR42","doi-asserted-by":"publisher","unstructured":"Yang, M., Li, Y., Huang, Z., Liu, Z., Hu, P., Peng, X.: Partially view-aligned representation learning with noise-robust contrastive loss. In: Conference on Computer Vision and Pattern Recognition, pp. 1134\u20131143 (2021). https:\/\/doi.org\/10.1109\/CVPR46437.2021.00119","DOI":"10.1109\/CVPR46437.2021.00119"},{"issue":"1","key":"1444_CR43","doi-asserted-by":"publisher","first-page":"1055","DOI":"10.1109\/TPAMI.2022.3155499","volume":"45","author":"M Yang","year":"2023","unstructured":"Yang, M., Li, Y., Hu, P., Bai, J., Lv, J., Peng, X.: Robust multi-view clustering with incomplete information. Trans. Pattern Anal. Mach. Intell. 45(1), 1055\u20131069 (2023). https:\/\/doi.org\/10.1109\/TPAMI.2022.3155499","journal-title":"Trans. Pattern Anal. Mach. Intell."},{"key":"1444_CR44","unstructured":"Zbontar, J., Jing, L., Misra, I., LeCun, Y., Deny, S.: Barlow twins: Self-supervised learning via redundancy reduction. In: 38th International Conference on Machine Learning, pp. 12310\u201312320 (2021)"},{"key":"1444_CR45","unstructured":"Bardes, A., Ponce, J., LeCun, Y.: VICReg: Variance-invariance-covariance regularization for self-supervised learning. In: International Conference on Learning Representations (2022)"},{"key":"1444_CR46","unstructured":"Li, J., Zhou, P., Xiong, C., Hoi, S.C.H.: Prototypical contrastive learning of unsupervised representations. In: 9th International Conference on Learning Representations (2021)"},{"key":"1444_CR47","doi-asserted-by":"publisher","unstructured":"Zheng, M., Wang, F., You, S., Qian, C., Zhang, C., Wang, X., Xu, C.: Weakly supervised contrastive learning. In: International Conference on Computer Vision, pp. 10022\u201310031 (2021). https:\/\/doi.org\/10.1109\/ICCV48922.2021.00989","DOI":"10.1109\/ICCV48922.2021.00989"},{"key":"1444_CR48","unstructured":"Wang, T., Isola, P.: Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In: Proceedings of the 37th International Conference on Machine Learning, pp. 9929\u20139939 (2020)"},{"key":"1444_CR49","unstructured":"Chen, T., Li, L.: Intriguing properties of contrastive losses. arXiv:2011.02803 (2020)"},{"key":"1444_CR50","doi-asserted-by":"publisher","unstructured":"Park, W., Kim, D., Lu, Y., Cho, M.: Relational knowledge distillation. In: Conference on Computer Vision and Pattern Recognition, pp. 3967\u20133976 (2019). https:\/\/doi.org\/10.1109\/CVPR.2019.00409","DOI":"10.1109\/CVPR.2019.00409"},{"key":"1444_CR51","unstructured":"Fang, Z., Wang, J., Wang, L., Zhang, L., Yang, Y., Liu, Z.: SEED: self-supervised distillation for visual representation. In: International Conference on Learning Representations (2021)"},{"key":"1444_CR52","doi-asserted-by":"crossref","unstructured":"Koohpayegani, S.A., Tejankar, A., Pirsiavash, H.: Compress: Self-supervised learning by compressing representations. In: Advances in Neural Information Processing Systems (2020)","DOI":"10.1109\/ICCV48922.2021.01016"},{"key":"1444_CR53","doi-asserted-by":"publisher","unstructured":"Devlin, J., Chang, M., Lee, K., Toutanova, K.: BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2\u20137, 2019, vol. 1 (Long and Short Papers), pp. 4171\u20134186 (2019). https:\/\/doi.org\/10.18653\/v1\/n19-1423","DOI":"10.18653\/v1\/n19-1423"},{"key":"1444_CR54","unstructured":"Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al.: Improving language understanding by generative pre-training (2018)"},{"key":"1444_CR55","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems, pp. 5998\u20136008 (2017)"},{"key":"1444_CR56","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: An image is worth $$16\\times 16$$ words: transformers for image recognition at scale. In: International Conference on Learning Representations (2021)"},{"key":"1444_CR57","doi-asserted-by":"publisher","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B.: Swin transformer: hierarchical vision transformer using shifted windows. In: International Conference on Computer Vision, pp. 9992\u201310002 (2021). https:\/\/doi.org\/10.1109\/ICCV48922.2021.00986","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"1444_CR58","unstructured":"Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., Kong, T.: Image BERT pre-training with online tokenizer. In: International Conference on Learning Representations (2022)"},{"key":"1444_CR59","unstructured":"Bao, H., Dong, L., Piao, S., Wei, F.: Beit: BERT pre-training of image transformers. In: International Conference on Learning Representations (2022)"},{"key":"1444_CR60","doi-asserted-by":"publisher","unstructured":"He, K., Chen, X., Xie, S., Li, Y., Doll\u00e1r, P., Girshick, R.B.: Masked autoencoders are scalable vision learners. In: Conference on Computer Vision and Pattern Recognition, pp. 15979\u201315988 (2022). https:\/\/doi.org\/10.1109\/CVPR52688.2022.01553","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"1444_CR61","unstructured":"Jing, L., Tian, Y.: Self-supervised spatiotemporal feature learning via video rotation prediction. arXiv:1811.11387 (2018)"},{"key":"1444_CR62","doi-asserted-by":"publisher","unstructured":"Kim, D., Cho, D., Kweon, I.S.: Self-supervised video representation learning with space-time cubic puzzles. In: 31st Innovative Applications of Artificial Intelligence Conference, pp. 8545\u20138552 (2019). https:\/\/doi.org\/10.1609\/aaai.v33i01.33018545","DOI":"10.1609\/aaai.v33i01.33018545"},{"key":"1444_CR63","doi-asserted-by":"publisher","unstructured":"Wang, J., Jiao, J., Bao, L., He, S., Liu, Y., Liu, W.: Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics. In: Conference on Computer Vision and Pattern Recognition, pp. 4006\u20134015 (2019). https:\/\/doi.org\/10.1109\/CVPR.2019.00413","DOI":"10.1109\/CVPR.2019.00413"},{"key":"1444_CR64","doi-asserted-by":"publisher","unstructured":"Lee, H., Huang, J., Singh, M., Yang, M.: Unsupervised representation learning by sorting sequences. In: International Conference on Computer Vision, pp. 667\u2013676 (2017). https:\/\/doi.org\/10.1109\/ICCV.2017.79","DOI":"10.1109\/ICCV.2017.79"},{"key":"1444_CR65","doi-asserted-by":"publisher","unstructured":"Misra, I., Zitnick, C.L., Hebert, M.: Shuffle and learn: Unsupervised learning using temporal order verification. In: 14th European Conference on Computer Vision, pp. 527\u2013544 (2016). https:\/\/doi.org\/10.1007\/978-3-319-46448-0_32","DOI":"10.1007\/978-3-319-46448-0_32"},{"key":"1444_CR66","doi-asserted-by":"publisher","unstructured":"Xu, D., Xiao, J., Zhao, Z., Shao, J., Xie, D., Zhuang, Y.: Self-supervised spatiotemporal learning via video clip order prediction. In: Conference on Computer Vision and Pattern Recognition, pp. 10334\u201310343 (2019). https:\/\/doi.org\/10.1109\/CVPR.2019.01058","DOI":"10.1109\/CVPR.2019.01058"},{"key":"1444_CR67","doi-asserted-by":"publisher","unstructured":"Jenni, S., Meishvili, G., Favaro, P.: Video representation learning by recognizing temporal transformations. In: 16th European Conference on Computer Vision, pp. 425\u2013442 (2020). https:\/\/doi.org\/10.1007\/978-3-030-58604-1_26","DOI":"10.1007\/978-3-030-58604-1_26"},{"key":"1444_CR68","doi-asserted-by":"publisher","unstructured":"Benaim, S., Ephrat, A., Lang, O., Mosseri, I., Freeman, W.T., Rubinstein, M., Irani, M., Dekel, T.: Speednet: learning the speediness in videos. In: Conference on Computer Vision and Pattern Recognition, pp. 9919\u20139928 (2020). https:\/\/doi.org\/10.1109\/CVPR42600.2020.00994","DOI":"10.1109\/CVPR42600.2020.00994"},{"key":"1444_CR69","doi-asserted-by":"publisher","unstructured":"Yao, Y., Liu, C., Luo, D., Zhou, Y., Ye, Q.: Video playback rate perception for self-supervised spatio-temporal representation learning. In: Conference on Computer Vision and Pattern Recognition, pp. 6547\u20136556 (2020). https:\/\/doi.org\/10.1109\/CVPR42600.2020.00658","DOI":"10.1109\/CVPR42600.2020.00658"},{"key":"1444_CR70","doi-asserted-by":"publisher","unstructured":"Lorre, G., Rabarisoa, J., Orcesi, A., Ainouz, S., Canu, S.: Temporal contrastive pretraining for video action recognition. In: Winter Conference on Applications of Computer Vision, pp. 651\u2013659 (2020). https:\/\/doi.org\/10.1109\/WACV45572.2020.9093278","DOI":"10.1109\/WACV45572.2020.9093278"},{"key":"1444_CR71","doi-asserted-by":"publisher","unstructured":"Qian, R., Li, Y., Liu, H., See, J., Ding, S., Liu, X., Li, D., Lin, W.: Enhancing self-supervised video representation learning via multi-level feature optimization. In: International Conference on Computer Vision, pp. 7970\u20137981 (2021). https:\/\/doi.org\/10.1109\/ICCV48922.2021.00789","DOI":"10.1109\/ICCV48922.2021.00789"},{"key":"1444_CR72","doi-asserted-by":"publisher","unstructured":"Sun, C., Nagrani, A., Tian, Y., Schmid, C.: Composable augmentation encoding for video representation learning. In: International Conference on Computer Vision, pp. 8814\u20138824 (2021). https:\/\/doi.org\/10.1109\/ICCV48922.2021.00871","DOI":"10.1109\/ICCV48922.2021.00871"},{"key":"1444_CR73","doi-asserted-by":"publisher","unstructured":"Piergiovanni, A.J., Angelova, A., Ryoo, M.S.: Evolving losses for unsupervised video representation learning. In: Conference on Computer Vision and Pattern Recognition, pp. 130\u2013139 (2020). https:\/\/doi.org\/10.1109\/CVPR42600.2020.00021","DOI":"10.1109\/CVPR42600.2020.00021"},{"key":"1444_CR74","doi-asserted-by":"publisher","unstructured":"Wang, J., Jiao, J., Liu, Y.: Self-supervised video representation learning by pace prediction. In: 16th European Conference on Computer Vision, pp. 504\u2013521 (2020). https:\/\/doi.org\/10.1007\/978-3-030-58520-4_30","DOI":"10.1007\/978-3-030-58520-4_30"},{"key":"1444_CR75","doi-asserted-by":"crossref","unstructured":"Chen, P., Huang, D., He, D., Long, X., Zeng, R., Wen, S., Tan, M., Gan, C.: Rspnet: relative speed perception for unsupervised video representation learning. In: 33rd Conference on Innovative Applications of Artificial Intelligence, pp. 1045\u20131053 (2021)","DOI":"10.1609\/aaai.v35i2.16189"},{"key":"1444_CR76","doi-asserted-by":"publisher","unstructured":"Huang, D., Wu, W., Hu, W., Liu, X., He, D., Wu, Z., Wu, X., Tan, M., Ding, E.: Ascnet: self-supervised video representation learning with appearance-speed consistency. In: International Conference on Computer Vision, pp. 8076\u20138085 (2021). https:\/\/doi.org\/10.1109\/ICCV48922.2021.00799","DOI":"10.1109\/ICCV48922.2021.00799"},{"key":"1444_CR77","doi-asserted-by":"publisher","unstructured":"Jenni, S., Jin, H.: Time-equivariant contrastive video representation learning. In: International Conference on Computer Vision (2021). https:\/\/doi.org\/10.1109\/ICCV48922.2021.00982","DOI":"10.1109\/ICCV48922.2021.00982"},{"key":"1444_CR78","unstructured":"Sun, C., Baradel, F., Murphy, K., Schmid, C.: Learning video representations using contrastive bidirectional transformer. arXiv:1906.05743 (2019)"},{"key":"1444_CR79","doi-asserted-by":"publisher","unstructured":"Miech, A., Alayrac, J., Smaira, L., Laptev, I., Sivic, J., Zisserman, A.: End-to-end learning of visual representations from uncurated instructional videos. In: Conference on Computer Vision and Pattern Recognition, pp. 9876\u20139886 (2020). https:\/\/doi.org\/10.1109\/CVPR42600.2020.00990","DOI":"10.1109\/CVPR42600.2020.00990"},{"key":"1444_CR80","unstructured":"Alwassel, H., Mahajan, D., Korbar, B., Torresani, L., Ghanem, B., Tran, D.: Self-supervised learning by cross-modal audio-video clustering. In: Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (2020)"},{"key":"1444_CR81","unstructured":"Han, T., Xie, W., Zisserman, A.: Self-supervised co-training for video representation learning. In: Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems (2020)"},{"key":"1444_CR82","doi-asserted-by":"publisher","unstructured":"Hu, K., Shao, J., Liu, Y., Raj, B., Savvides, M., Shen, Z.: Contrast and order representations for video self-supervised learning. In: International Conference on Computer Vision, pp. 7919\u20137929 (2021). https:\/\/doi.org\/10.1109\/ICCV48922.2021.00784","DOI":"10.1109\/ICCV48922.2021.00784"},{"key":"1444_CR83","doi-asserted-by":"publisher","unstructured":"Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lucic, M., Schmid, C.: Vivit: a video vision transformer. In: International Conference on Computer Vision, pp. 6816\u20136826 (2021). https:\/\/doi.org\/10.1109\/ICCV48922.2021.00676","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"1444_CR84","doi-asserted-by":"publisher","unstructured":"Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., Hu, H.: Video swin transformer. In: Conference on Computer Vision and Pattern Recognition, pp. 3192\u20133201 (2022). https:\/\/doi.org\/10.1109\/CVPR52688.2022.00320","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"1444_CR85","unstructured":"Feichtenhofer, C., Fan, H., Li, Y., He, K.: Masked autoencoders as spatiotemporal learners. In: Advances in Neural Information Processing Systems (2022)"},{"key":"1444_CR86","unstructured":"Tong, Z., Song, Y., Wang, J., Wang, L.: VideoMAE: masked autoencoders are data-efficient learners for self-supervised video pre-training. In: Advances in Neural Information Processing Systems (2022)"},{"key":"1444_CR87","doi-asserted-by":"publisher","unstructured":"Wang, R., Chen, D., Wu, Z., Chen, Y., Dai, X., Liu, M., Jiang, Y., Zhou, L., Yuan, L.: BEVT: BERT pretraining of video transformers. In: Conference on Computer Vision and Pattern Recognition, pp. 14713\u201314723 (2022). https:\/\/doi.org\/10.1109\/CVPR52688.2022.01432","DOI":"10.1109\/CVPR52688.2022.01432"},{"key":"1444_CR88","doi-asserted-by":"publisher","unstructured":"Deng, J., Dong, W., Socher, R., Li, L., Li, K., Fei-Fei, L.: Imagenet: a large-scale hierarchical image database. In: Computer Society Conference on Computer Vision and Pattern Recognition, pp. 248\u2013255 (2009). https:\/\/doi.org\/10.1109\/CVPR.2009.5206848","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"1444_CR89","doi-asserted-by":"publisher","unstructured":"He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Conference on Computer Vision and Pattern Recognition, pp. 770\u2013778 (2016). https:\/\/doi.org\/10.1109\/CVPR.2016.90","DOI":"10.1109\/CVPR.2016.90"},{"key":"1444_CR90","unstructured":"Krizhevsky, A., Hinton, G.: Learning multiple layers of features from tiny images. Cs.Toronto.Edu, pp. 1\u201358 (2009)"},{"key":"1444_CR91","unstructured":"Coates, A., Ng, A.Y., Lee, H.: An analysis of single-layer networks in unsupervised feature learning. In: Proceedings of the 14th International Conference on Artificial Intelligence and Statistics, pp. 215\u2013223 (2011)"},{"key":"1444_CR92","unstructured":"Abai, Z., Rajmalwar, N.: Densenet models for tiny imagenet classification. arXiv:1904.10429 (2019)"},{"key":"1444_CR93","doi-asserted-by":"publisher","unstructured":"Tao, C., Wang, H., Zhu, X., Dong, J., Song, S., Huang, G., Dai, J.: Exploring the equivalence of siamese self-supervised learning via A unified gradient framework. In: Conference on Computer Vision and Pattern Recognition (2022). https:\/\/doi.org\/10.1109\/CVPR52688.2022.01403","DOI":"10.1109\/CVPR52688.2022.01403"},{"key":"1444_CR94","doi-asserted-by":"publisher","unstructured":"Lin, T., Maire, M., Belongie, S.J., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., Zitnick, C.L.: Microsoft COCO: common objects in context. In: 13th European Conference on Computer Vision, pp. 740\u2013755 (2014). https:\/\/doi.org\/10.1007\/978-3-319-10602-1_48","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"1444_CR95","doi-asserted-by":"publisher","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., Girshick, R.B.: Mask r-CNN. In: International Conference on Computer Vision, pp. 2980\u20132988 (2017). https:\/\/doi.org\/10.1109\/ICCV.2017.322","DOI":"10.1109\/ICCV.2017.322"},{"key":"1444_CR96","doi-asserted-by":"publisher","unstructured":"Ericsson, L., Gouk, H., Hospedales, T.M.: How well do self-supervised models transfer? In: Conference on Computer Vision and Pattern Recognition, pp. 5414\u20135423 (2021). https:\/\/doi.org\/10.1109\/CVPR46437.2021.00537","DOI":"10.1109\/CVPR46437.2021.00537"},{"key":"1444_CR97","unstructured":"Maji, S., Rahtu, E., Kannala, J., Blaschko, M.B., Vedaldi, A.: Fine-grained visual classification of aircraft. arXiv:1306.5151 (2013)"},{"key":"1444_CR98","doi-asserted-by":"publisher","unstructured":"Fei-Fei, L., Fergus, R., Perona, P.: Learning generative visual models from few training examples: an incremental Bayesian approach tested on 101 object categories. Comput. Vis. Image Underst. (2007) https:\/\/doi.org\/10.1016\/j.cviu.2005.09.012","DOI":"10.1016\/j.cviu.2005.09.012"},{"key":"1444_CR99","doi-asserted-by":"crossref","unstructured":"Krause, J., Stark, M., Deng, J., Fei-Fei, L.: 3d object representations for fine-grained categorization. In: International Conference on Computer Vision Workshops, pp. 554\u2013561 (2013)","DOI":"10.1109\/ICCVW.2013.77"},{"key":"1444_CR100","doi-asserted-by":"publisher","unstructured":"Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., Vedaldi, A.: Describing textures in the wild. In: Conference on Computer Vision and Pattern Recognition, pp. 3606\u20133613 (2014). https:\/\/doi.org\/10.1109\/CVPR.2014.461","DOI":"10.1109\/CVPR.2014.461"},{"key":"1444_CR101","doi-asserted-by":"publisher","unstructured":"Nilsback, M., Zisserman, A.: Automated flower classification over a large number of classes. In: 6th Indian Conference on Computer Vision, Graphics and Image Processing, pp. 722\u2013729 (2008). https:\/\/doi.org\/10.1109\/ICVGIP.2008.47","DOI":"10.1109\/ICVGIP.2008.47"},{"key":"1444_CR102","doi-asserted-by":"publisher","unstructured":"Bossard, L., Guillaumin, M., Gool, L.V.: Food-101\u2014mining discriminative components with random forests. In: 13th European Conference on Computer Vision, pp. 446\u2013461 (2014). https:\/\/doi.org\/10.1007\/978-3-319-10599-4_29","DOI":"10.1007\/978-3-319-10599-4_29"},{"key":"1444_CR103","doi-asserted-by":"publisher","unstructured":"Parkhi, O.M., Vedaldi, A., Zisserman, A., Jawahar, C.V.: Cats and dogs. In: Conference on Computer Vision and Pattern Recognition, pp. 3498\u20133505 (2012). https:\/\/doi.org\/10.1109\/CVPR.2012.6248092","DOI":"10.1109\/CVPR.2012.6248092"},{"key":"1444_CR104","doi-asserted-by":"publisher","unstructured":"Xiao, J., Hays, J., Ehinger, K.A., Oliva, A., Torralba, A.: SUN database: large-scale scene recognition from abbey to zoo. In: Conference on Computer Vision and Pattern Recognition, pp. 3485\u20133492 (2010). https:\/\/doi.org\/10.1109\/CVPR.2010.5539970","DOI":"10.1109\/CVPR.2010.5539970"},{"key":"1444_CR105","doi-asserted-by":"publisher","unstructured":"Everingham, M., Gool, L.V., Williams, C.K.I., Winn, J.M., Zisserman, A.: The pascal visual object classes (VOC) challenge. Int. J. Comput. Vis. (2010) https:\/\/doi.org\/10.1007\/s11263-009-0275-4","DOI":"10.1007\/s11263-009-0275-4"},{"key":"1444_CR106","doi-asserted-by":"publisher","unstructured":"Xie, S., Sun, C., Huang, J., Tu, Z., Murphy, K.: Rethinking spatiotemporal feature learning: speed-accuracy trade-offs in video classification. In: 15th European Conference on Computer Vision, pp. 318\u2013335 (2018). https:\/\/doi.org\/10.1007\/978-3-030-01267-0_19","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"1444_CR107","unstructured":"Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M., Zisserman, A.: The kinetics human action video dataset. arXiv:1705.06950 (2017)"},{"key":"1444_CR108","unstructured":"Soomro, K., Zamir, A.R., Shah, M.: UCF101: a dataset of 101 human actions classes from videos in the wild. arXiv:1212.0402 (2012)"},{"key":"1444_CR109","doi-asserted-by":"publisher","unstructured":"Kuehne, H., Jhuang, H., Garrote, E., Poggio, T.A., Serre, T.: HMDB: a large video database for human motion recognition. In: International Conference on Computer Vision, pp. 2556\u20132563 (2011). https:\/\/doi.org\/10.1109\/ICCV.2011.6126543","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"1444_CR110","doi-asserted-by":"publisher","unstructured":"Hara, K., Kataoka, H., Satoh, Y.: Can spatiotemporal 3d CNNs retrace the history of 2d CNNs and imagenet? In: Conference on Computer Vision and Pattern Recognition, pp. 6546\u20136555 (2018). https:\/\/doi.org\/10.1109\/CVPR.2018.00685","DOI":"10.1109\/CVPR.2018.00685"},{"key":"1444_CR111","doi-asserted-by":"publisher","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., He, K.: Slowfast networks for video recognition. In: International Conference on Computer Vision, pp. 6201\u20136210 (2019). https:\/\/doi.org\/10.1109\/ICCV.2019.00630","DOI":"10.1109\/ICCV.2019.00630"},{"key":"1444_CR112","doi-asserted-by":"publisher","unstructured":"Gu, C., Sun, C., Ross, D.A., Vondrick, C., Pantofaru, C., Li, Y., Vijayanarasimhan, S., Toderici, G., Ricco, S., Sukthankar, R., Schmid, C., Malik, J.: Ava: a video dataset of spatio-temporally localized atomic visual actions. In: Conference on Computer Vision and Pattern Recognition, pp. 6047\u20136056 (2018). https:\/\/doi.org\/10.1109\/CVPR.2018.00633","DOI":"10.1109\/CVPR.2018.00633"},{"key":"1444_CR113","doi-asserted-by":"publisher","unstructured":"Goyal, R., Kahou, S.E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fr\u00fcnd, I., Yianilos, P., Mueller-Freitag, M., Hoppe, F., Thurau, C., Bax, I., Memisevic, R.: The \u201csomething something\u201d video database for learning and evaluating visual common sense. In: International Conference on Computer Vision, pp. 5843\u20135851 (2017). https:\/\/doi.org\/10.1109\/ICCV.2017.622","DOI":"10.1109\/ICCV.2017.622"},{"key":"1444_CR114","doi-asserted-by":"publisher","unstructured":"Park, J., Lee, J., Kim, I., Sohn, K.: Probabilistic representations for video contrastive learning. In: Conference on Computer Vision and Pattern Recognition, pp. 14691\u201314701 (2022). https:\/\/doi.org\/10.1109\/CVPR52688.2022.01430","DOI":"10.1109\/CVPR52688.2022.01430"},{"key":"1444_CR115","doi-asserted-by":"publisher","unstructured":"Yuan, L., Qian, R., Cui, Y., Gong, B., Schroff, F., Yang, M., Adam, H., Liu, T.: Contextualized spatio-temporal contrastive learning with self-supervision. In: Conference on Computer Vision and Pattern Recognition, pp. 13957\u201313966 (2022). https:\/\/doi.org\/10.1109\/CVPR52688.2022.01359","DOI":"10.1109\/CVPR52688.2022.01359"},{"key":"1444_CR116","unstructured":"Zhang, D., Dai, X., Wang, X., Wang, Y.: S3D: single shot multi-span detector via fully 3d convolutional networks. In: British Machine Vision Conference, p. 293 (2018)"},{"key":"1444_CR117","doi-asserted-by":"publisher","unstructured":"Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., Paluri, M.: A closer look at spatiotemporal convolutions for action recognition. In: Conference on Computer Vision and Pattern Recognition, pp. 6450\u20136459 (2018). https:\/\/doi.org\/10.1109\/CVPR.2018.00675","DOI":"10.1109\/CVPR.2018.00675"}],"container-title":["Machine Vision and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00138-023-01444-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00138-023-01444-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00138-023-01444-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,6]],"date-time":"2023-11-06T17:08:09Z","timestamp":1699290489000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00138-023-01444-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,26]]},"references-count":117,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,11]]}},"alternative-id":["1444"],"URL":"https:\/\/doi.org\/10.1007\/s00138-023-01444-9","relation":{},"ISSN":["0932-8092","1432-1769"],"issn-type":[{"value":"0932-8092","type":"print"},{"value":"1432-1769","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,26]]},"assertion":[{"value":"30 March 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 July 2023","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 August 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 September 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant financial or non-financial interests to disclose.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"111"}}