{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,27]],"date-time":"2026-01-27T09:11:13Z","timestamp":1769505073599,"version":"3.49.0"},"reference-count":81,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2025,3,26]],"date-time":"2025-03-26T00:00:00Z","timestamp":1742947200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2025,3,26]],"date-time":"2025-03-26T00:00:00Z","timestamp":1742947200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2025,7]]},"DOI":"10.1007\/s11263-025-02420-8","type":"journal-article","created":{"date-parts":[[2025,3,29]],"date-time":"2025-03-29T14:05:12Z","timestamp":1743257112000},"page":"4923-4943","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Pre-training for Action Recognition with Automatically Generated Fractal Datasets"],"prefix":"10.1007","volume":"133","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-3475-3859","authenticated-orcid":false,"given":"Davyd","family":"Svyezhentsev","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6734-3575","authenticated-orcid":false,"given":"George","family":"Retsinas","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0534-2707","authenticated-orcid":false,"given":"Petros","family":"Maragos","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,3,26]]},"reference":[{"key":"2420_CR1","unstructured":"Abu-El-Haija, S., Kothari, N., Lee, J., Natsev, P., Toderici, G., Varadarajan, B., & Vijayanarasimhan, S. (2016). Youtube-8m: A large-scale video classification benchmark. arXiv:1609.08675"},{"key":"2420_CR2","doi-asserted-by":"crossref","unstructured":"Anderson, C., & Farrell, R. (2022). Improving fractal pre-training. In Proceedings of the IEEE winter conference on applications of computer vision (WACV).","DOI":"10.1109\/WACV51458.2022.00247"},{"issue":"1","key":"2420_CR3","doi-asserted-by":"publisher","first-page":"3","DOI":"10.3390\/fractalfract4010003","volume":"4","author":"J Anguera","year":"2020","unstructured":"Anguera, J., And\u00fajar, A., Jayasinghe, J., Chakravarthy, V. V. S. S. S., Chowdary, P. S. R., Pijoan, J. L., Ali, T., & Cattani, C. (2020). Fractal antennas: An historical perspective. Fractal Fract, 4(1), 3.","journal-title":"Fractal Fract"},{"key":"2420_CR4","unstructured":"Asano, Y., Rupprecht, C., Zisserman, A., & Vedaldi, A. (2021). Pass: An imagenet replacement for self-supervised pretraining without humans. In Proceedings of the international conference on neural information processing systems (NeurIPS) datasets and benchmarks track."},{"key":"2420_CR5","unstructured":"Baradad\u00a0Jurjo, M., Wulff, J., Wang, T., Isola, P., & Torralba, A. (2021). Learning to see by looking at noise. In Proceedings of the international conference on neural information processing systems (NeurIPS)."},{"key":"2420_CR6","unstructured":"Barnsley, M.F. (1993). Fractals Everywhere. Morgan Kaufmann."},{"key":"2420_CR7","doi-asserted-by":"crossref","unstructured":"Birhane, A., & Prabhu, V.U. (2021). Large image datasets: A pyrrhic win for computer vision? In Proceedings of the IEEE winter conference on applications of computer vision (WACV).","DOI":"10.1109\/WACV48630.2021.00158"},{"key":"2420_CR8","unstructured":"Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. In Proceedings of the conference on fairness, accountability and transparency."},{"key":"2420_CR9","unstructured":"Burch, B., & Hart, J.C. (1997). Linear fractal shape interpolation. In Proceedings of the graphics interface conference."},{"key":"2420_CR10","unstructured":"Carreira, J., Noland, E., Hillier, C., & Zisserman, A. (2019). A short note on the kinetics-700 human action dataset. arXiv:1907.06987"},{"key":"2420_CR11","doi-asserted-by":"crossref","unstructured":"Carreira, J., & Zisserman, A. (2017). Quo vadis, action recognition? a new model and the kinetics dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2017.502"},{"key":"2420_CR12","unstructured":"Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). A simple framework for contrastive learning of visual representations. In Proceedings of the international conference on machine learning (ICML)."},{"key":"2420_CR13","unstructured":"Chen, X., Fan, H., Girshick, R., & He, K. (2020). Improved baselines with momentum contrastive learning. arXiv:2003.04297"},{"key":"2420_CR14","doi-asserted-by":"crossref","unstructured":"Cole, E., Yang, X., Wilber, K., Mac\u00a0Aodha, O., & Belongie, S. (2022). When does contrastive visual representation learning work? In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR52688.2022.01434"},{"key":"2420_CR15","doi-asserted-by":"publisher","first-page":"2455","DOI":"10.1016\/j.jbiomech.2015.12.025","volume":"49","author":"F Costabal","year":"2015","unstructured":"Costabal, F., Hurtado, D., & Kuhl, E. (2015). Generating purkinje networks in the human heart. Journal of Biomechanics, 49, 2455\u20132465.","journal-title":"Journal of Biomechanics"},{"key":"2420_CR16","doi-asserted-by":"crossref","unstructured":"Cubuk, E.D., Zoph, B., Shlens, J., & Le, Q. (2020). Randaugment: Practical automated data augmentation with a reduced search space. Proceedings of the international conference on neural information processing systems (NeurIPS).","DOI":"10.1109\/CVPRW50498.2020.00359"},{"issue":"11","key":"2420_CR17","doi-asserted-by":"publisher","first-page":"4261","DOI":"10.1109\/TSP.2005.857010","volume":"53","author":"A Dimakis","year":"2005","unstructured":"Dimakis, A., & Maragos, P. (2005). Phase-modulated resonances modeled as self-similar processes with application to turbulent sounds. IEEE Transactions on Signal Processing, 53(11), 4261\u20134272.","journal-title":"IEEE Transactions on Signal Processing"},{"key":"2420_CR18","doi-asserted-by":"crossref","unstructured":"Ding, S., Li, M., Yang, T., Qian, R., Xu, H., Chen, Q., Wang, J., & Xiong, H. (2022). Motion-aware contrastive video representation learning via foreground-background merging. In Proceedings of the ieee conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR52688.2022.00949"},{"key":"2420_CR19","doi-asserted-by":"crossref","unstructured":"Draves, S. (2005). The electric sheep screen-saver: A case study in aesthetic evolution. In Proceedings of the European conference on applications of evolutionary computing.","DOI":"10.1007\/978-3-540-32003-6_46"},{"key":"2420_CR20","unstructured":"Draves, S., & Reckase, E. (2008). The fractal flame algorithm. (2008)."},{"key":"2420_CR21","doi-asserted-by":"crossref","unstructured":"Fan, L., Chen, K., Krishnan, D., Katabi, D., Isola, P., & Tian, Y. (2024). Scaling laws of synthetic images for model training ... for now. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR52733.2024.00705"},{"key":"2420_CR22","doi-asserted-by":"crossref","unstructured":"Farin, G. (1988). Curves and surfaces for computer aided geometric design: A practical guide. Academic Press Professional Inc.","DOI":"10.1016\/B978-0-12-460515-2.50020-2"},{"key":"2420_CR23","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Fan, H., Xiong, B., Girshick, R., & He, K. (2021). A large-scale study on unsupervised spatiotemporal representation learning. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR46437.2021.00331"},{"key":"2420_CR24","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Pinz, A., & Wildes, R.P. (2017). Temporal residual networks for dynamic scene recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2017.786"},{"key":"2420_CR25","doi-asserted-by":"crossref","unstructured":"Ghadiyaram, D., Tran, D., & Mahajan, D. (2019). Large-scale weakly-supervised pre-training for video action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2019.01232"},{"key":"2420_CR26","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. In Proceedings of the international conference on neural information processing systems (NeurIPS)."},{"key":"2420_CR27","unstructured":"Goyal, P., Caron, M., Lefaudeux, B., Xu, M., Wang, P., Pai, V., Singh, M., Liptchinsky, V., Misra, I., Joulin, A., & Bojanowski, P. (2021). Self-supervised pretraining of visual features in the wild. arXiv:2103.01988"},{"key":"2420_CR28","doi-asserted-by":"crossref","unstructured":"Goyal, R., Ebrahimi\u00a0Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., Hoppe, F., Thurau, C., Bax, I., & Memisevic, R. (2017). The \"something something\" video database for learning and evaluating visual common sense. In Proceedings of the IEEE international conference on computer vision (ICCV).","DOI":"10.1109\/ICCV.2017.622"},{"key":"2420_CR29","unstructured":"Grill, J.B., Strub, F., Altch\u00e9, F., Tallec, C., Richemond, P., Buchatskaya, E., Doersch, C., Avila\u00a0Pires, B., Guo, Z., Gheshlaghi\u00a0Azar, M., Piot, B., kavukcuoglu, k., Munos, R., & Valko, M. (2020). Bootstrap your own latent - a new approach to self-supervised learning. In Proceedings of the international conference on neural information processing systems (NeurIPS)."},{"key":"2420_CR30","doi-asserted-by":"crossref","unstructured":"Hara, K., Kataoka, H., & Satoh, Y. (2018). Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet? In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2018.00685"},{"key":"2420_CR31","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2016.90"},{"key":"2420_CR32","doi-asserted-by":"crossref","unstructured":"Heilbron, F.C., Escorcia, V., Ghanem, B., & Niebles, J.C. (2015). Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"2420_CR33","unstructured":"Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. In Proceedings of the international conference on neural information processing systems (NeurIPS)."},{"key":"2420_CR34","unstructured":"Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., & Fleet, D.J. (2022). Video diffusion models. In Proceedings of the International Conference on Neural Information Processing Systems (NeurIPS)."},{"issue":"5","key":"2420_CR35","doi-asserted-by":"publisher","first-page":"713","DOI":"10.1512\/iumj.1981.30.30055","volume":"30","author":"JE Hutchinson","year":"1981","unstructured":"Hutchinson, J. E. (1981). Fractals and self similarity. Indiana University Mathematics Journal, 30(5), 713\u2013747.","journal-title":"Indiana University Mathematics Journal"},{"key":"2420_CR36","doi-asserted-by":"crossref","unstructured":"Ibrahim, M.S., Muralidharan, S., Deng, Z., Vahdat, A., & Mori, G. (2016). A hierarchical deep temporal model for group activity recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR.2016.217"},{"issue":"10","key":"2420_CR37","doi-asserted-by":"publisher","first-page":"994","DOI":"10.3182\/20090706-3-FR-2004.00165","volume":"42","author":"C Ionescu","year":"2009","unstructured":"Ionescu, C., Oustaloup, A., Levron, F., Melchior, P., Sabatier, J., & De Keyser, R. (2009). A model of the lungs based on fractal geometrical and structural properties. IFAC Proceedings Volumes, 42(10), 994\u2013999.","journal-title":"IFAC Proceedings Volumes"},{"issue":"1","key":"2420_CR38","doi-asserted-by":"publisher","first-page":"18","DOI":"10.1109\/83.128028","volume":"1","author":"AE Jacquin","year":"1992","unstructured":"Jacquin, A. E. (1992). Image coding based on a fractal theory of iterated contractive image transformations. IEEE Transactions on Image Processing, 1(1), 18\u201330.","journal-title":"IEEE Transactions on Image Processing"},{"key":"2420_CR39","doi-asserted-by":"crossref","unstructured":"Kataoka, H., Hara, K., Hayashi, R., Yamagata, E., & Inoue, N. (2022). Spatiotemporal initialization for 3d cnns with generated motion patterns. In Proceedings of the IEEE winter conference on applications of computer vision (WACV).","DOI":"10.1109\/WACV51458.2022.00081"},{"key":"2420_CR40","doi-asserted-by":"crossref","unstructured":"Kataoka, H., Hayamizu, R., Yamada, R., Nakashima, K., Takashima, S., Zhang, X., Martinez-Noriega, E.J., Inoue, N., & Yokota, R. (2022). Replacing labeled real-image datasets with auto-generated contours. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR52688.2022.02055"},{"key":"2420_CR41","doi-asserted-by":"crossref","unstructured":"Kataoka, H., Matsumoto, A., Yamada, R., Satoh, Y., Yamagata, E., & Inoue, N. (2021). Formula-driven supervised learning with recursive tiling patterns. In Proceedings of the IEEE international conference on computer vision (ICCV) workshops.","DOI":"10.1109\/ICCVW54120.2021.00455"},{"key":"2420_CR42","doi-asserted-by":"crossref","unstructured":"Kataoka, H., Okayasu, K., Matsumoto, A., Yamagata, E., Yamada, R., Inoue, N., Nakamura, A., & Satoh, Y. (2020). Pre-training without natural images. In Proceedings of the Asian Conference on Computer Vision (ACCV).","DOI":"10.1007\/978-3-030-69544-6_35"},{"key":"2420_CR43","unstructured":"Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M., & Zisserman, A. (2017). The kinetics human action video dataset. arXiv:1705.06950"},{"issue":"6","key":"2420_CR44","doi-asserted-by":"publisher","first-page":"1098","DOI":"10.1109\/TSA.2005.852982","volume":"13","author":"I Kokkinos","year":"2005","unstructured":"Kokkinos, I., & Maragos, P. (2005). Nonlinear speech analysis using models for chaotic systems. IEEE Transactions on Speech and Audio Processing, 13(6), 1098\u20131109.","journal-title":"IEEE Transactions on Speech and Audio Processing"},{"key":"2420_CR45","doi-asserted-by":"crossref","unstructured":"Kotar, K., Ilharco, G., Schmidt, L., Ehsani, K., & Mottaghi, R. (2021). Contrasting contrastive self-supervised representation learning pipelines. In Proceedings of the IEEE international conference on computer vision (ICCV).","DOI":"10.1109\/ICCV48922.2021.00980"},{"key":"2420_CR46","doi-asserted-by":"crossref","unstructured":"Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., & Serre, T. (2011). Hmdb: A large video database for human motion recognition. In Proceedings of the IEEE international conference on computer vision (ICCV).","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"2420_CR47","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1023\/A:1011109015675","volume":"41","author":"AB Lee","year":"2004","unstructured":"Lee, A. B., Mumford, D., & Huang, J. (2004). Occlusion models for natural images: A statistical study of a scale-invariant dead leaves model. International Journal of Computer Vision (IJCV), 41, 35\u201359.","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2420_CR48","doi-asserted-by":"crossref","unstructured":"Li, Y., Li, Y., & Vasconcelos, N. (2018). Resound: Towards action recognition without representation bias. In Proceedings of the European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-030-01231-1_32"},{"key":"2420_CR49","doi-asserted-by":"crossref","unstructured":"Li, Y., Liu, M., & Rehg, J.M. (2018). In the eye of beholder: Joint learning of gaze and actions in first person video. In Proceedings of the european conference on computer vision (ECCV).","DOI":"10.1007\/978-3-030-01228-1_38"},{"key":"2420_CR50","doi-asserted-by":"crossref","unstructured":"Lin, J., Gan, C., & Han, S. (2019). Tsm: Temporal shift module for efficient video understanding. In Proceedings of the IEEE international conference on computer vision (ICCV).","DOI":"10.1109\/ICCV.2019.00718"},{"key":"2420_CR51","unstructured":"Loshchilov, I., & Hutter, F. (2017). SGDR: Stochastic gradient descent with warm restarts. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"2420_CR52","unstructured":"Loshchilov, I., & Hutter, F. (2019). Decoupled weight decay regularization. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"2420_CR53","doi-asserted-by":"publisher","first-page":"199","DOI":"10.1016\/S0065-2539(08)60549-1","volume":"88","author":"P Maragos","year":"1994","unstructured":"Maragos, P. (1994). Fractal signal analysis using mathematical morphology. Advances in Electronics and Electron Physics, 88, 199\u2013246.","journal-title":"Advances in Electronics and Electron Physics"},{"key":"2420_CR54","doi-asserted-by":"publisher","first-page":"1925","DOI":"10.1121\/1.426738","volume":"105","author":"P Maragos","year":"1999","unstructured":"Maragos, P., & Potamianos, A. (1999). Fractal dimensions of speech sounds: Computation and application to automatic speech recognition. The Journal of the Acoustical Society of America, 105, 1925\u201332.","journal-title":"The Journal of the Acoustical Society of America"},{"key":"2420_CR55","unstructured":"Melnik, A., Ljubljanac, M., Lu, C., Yan, Q., Ren, W., Ritter, H. (2024). Video diffusion models: A survey. arXiv:2405.03150"},{"key":"2420_CR56","first-page":"502","volume":"42","author":"M Monfort","year":"2020","unstructured":"Monfort, M., Andonian, A., Zhou, B., Ramakrishnan, K., Bargal, S. A., Yan, T., Brown, L., Fan, Q., Gutfreund, D., Vondrick, C., & Oliva, A. (2020). Moments in time dataset: One million videos for event understanding. In IEEE Transactions on Pattern Analysis and Machine Intelligence, 42, 502\u2013508.","journal-title":"In IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2420_CR57","doi-asserted-by":"crossref","unstructured":"Nakashima, K., Kataoka, H., Matsumoto, A., Iwata, K., Inoue, N., & Satoh, Y. (2022). Can vision transformers learn without natural images? In Proceedings of the AAAI conference on artificial intelligence.","DOI":"10.1609\/aaai.v36i2.20094"},{"key":"2420_CR58","doi-asserted-by":"crossref","unstructured":"Perlin, K. (1985). An image synthesizer. In Proceedings of the annual conference on computer graphics and interactive techniques (SIGGRAPH).","DOI":"10.1145\/325334.325247"},{"key":"2420_CR59","doi-asserted-by":"crossref","unstructured":"Perlin, K. (2002). Improving noise. In Proceedings of the annual conference on computer graphics and interactive techniques (SIGGRAPH).","DOI":"10.1145\/566570.566636"},{"issue":"12","key":"2420_CR60","doi-asserted-by":"publisher","first-page":"1206","DOI":"10.1016\/j.specom.2009.06.005","volume":"51","author":"V Pitsikalis","year":"2009","unstructured":"Pitsikalis, V., & Maragos, P. (2009). Analysis and classification of speech signals by generalized fractal dimension features. Speech Communication, 51(12), 1206\u20131223.","journal-title":"Speech Communication"},{"key":"2420_CR61","doi-asserted-by":"crossref","unstructured":"Qu, H., Song, L., & Xue, G. (2013). Shaking video synthesis for video stabilization performance assessment. In Proceedings of the visual communications and image processing (VCIP).","DOI":"10.1109\/VCIP.2013.6706422"},{"issue":"23","key":"2420_CR62","doi-asserted-by":"publisher","first-page":"3385","DOI":"10.1016\/S0042-6989(97)00008-4","volume":"37","author":"DL Ruderman","year":"1997","unstructured":"Ruderman, D. L. (1997). Origins of scaling in natural images. Vision Research, 37(23), 3385\u20133398.","journal-title":"Vision Research"},{"issue":"3","key":"2420_CR63","doi-asserted-by":"publisher","first-page":"1573","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., & Fei-Fei, L. (2015). Imagenet large scale visual recognition challenge. International Journal of Computer Vision (IJCV), 115(3), 1573\u20131405.","journal-title":"International Journal of Computer Vision (IJCV)"},{"key":"2420_CR64","unstructured":"Simonyan, K., & Zisserman, A. (2014). Two-stream convolutional networks for action recognition in videos. In Proceedings of the international conference on neural information processing systems (NeurIPS)."},{"key":"2420_CR65","unstructured":"Soomro, K., Zamir, A.R., & Shah, M. (2012). Ucf101: A dataset of 101 human actions classes from videos in the wild. In arXiv:1212.0402"},{"key":"2420_CR66","doi-asserted-by":"crossref","unstructured":"Steed, R., & Caliskan, A. (2021). Image representations learned with unsupervised pre-training contain human-like biases. In Proceedings of the ACM conference on fairness, accountability, and transparency.","DOI":"10.1145\/3442188.3445932"},{"key":"2420_CR67","doi-asserted-by":"crossref","unstructured":"Thoker, F.M., Doughty, H., Bagad, P., & Snoek, C.G.M. (2022). How severe is benchmark-sensitivity in video self-supervised learning? In Proceedings of the European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-031-19830-4_36"},{"key":"2420_CR68","unstructured":"Tian, Y., Fan, L., Isola, P., Chang, H., & Krishnan, D. (2023). Stablerep: Synthetic images from text-to-image models make strong visual representation learners. In Proceedings of the international conference on neural information processing systems (NeurIPS)."},{"key":"2420_CR69","unstructured":"Tong, Z., Song, Y., Wang, J., & Wang, L. (2022). VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training. In Proceedings of the international conference on neural information processing systems (NeurIPS)."},{"key":"2420_CR70","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., & Paluri, M. (2015). Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision (ICCV).","DOI":"10.1109\/ICCV.2015.510"},{"issue":"1","key":"2420_CR71","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1016\/0165-1684(91)90025-E","volume":"22","author":"L Vincent","year":"1991","unstructured":"Vincent, L. (1991). Morphological transformations of binary images with arbitrary structuring elements. Signal Processing, 22(1), 3\u201323.","journal-title":"Signal Processing"},{"key":"2420_CR72","doi-asserted-by":"crossref","unstructured":"Wang, J., Gao, Y., Li, K., Lin, Y., Ma, A.J., Cheng, H., Peng, P., Huang, F., Ji, R., & Sun, X. (2021). Removing the background by adding the background: Towards background robust self-supervised video representation learning. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR46437.2021.01163"},{"key":"2420_CR73","doi-asserted-by":"crossref","unstructured":"Wang, L., Huang, B., Zhao, Z., Tong, Z., He, Y., Wang, Y., Wang, Y., & Qiao, Y. (2023). Videomae v2: Scaling video masked autoencoders with dual masking. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR52729.2023.01398"},{"key":"2420_CR74","doi-asserted-by":"crossref","unstructured":"Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., & Van\u00a0Gool, L. (2016). Temporal segment networks: Towards good practices for deep action recognition. In Proceedings of the European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"2420_CR75","unstructured":"Wilson, B., Hoffman, J., & Morgenstern, J. (2019). Predictive inequity in object detection. arXiv:1902.11097"},{"key":"2420_CR76","doi-asserted-by":"crossref","unstructured":"Xie, S., Sun, C., Huang, J., Tu, Z., & Murphy, K. (2018). Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In Proceedings of the European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"2420_CR77","unstructured":"Yalniz, I.Z., J\u00e9gou, H., Chen, K., Paluri, M., & Mahajan, D. (2019). Billion-scale semi-supervised learning for image classification. In: arXiv:1905.00546"},{"key":"2420_CR78","doi-asserted-by":"crossref","unstructured":"Yamada, R., Kataoka, H., Chiba, N., Domae, Y., & Ogata, T. (2022). Point cloud pre-training with natural 3d structures. In: Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR52688.2022.02060"},{"key":"2420_CR79","doi-asserted-by":"crossref","unstructured":"Zhao, J., Wang, T., Yatskar, M., Ordonez, V., & Chang, K.W. (2017). Men also like shopping: Reducing gender bias amplification using corpus-level constraints. In Proceedings of the conference on empirical methods in natural language processing (EMNLP).","DOI":"10.18653\/v1\/D17-1323"},{"key":"2420_CR80","doi-asserted-by":"crossref","unstructured":"Zhao, K., Shen, L., Zhang, Y., Zhou, C., Wang, T., Zhang, R., Ding, S., Jia, W., & Shen, W. (2022). B\u00e9zierpalm: A free lunch for palmprint recognition. In Proceedings of the European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-031-19778-9_2"},{"key":"2420_CR81","doi-asserted-by":"publisher","first-page":"737","DOI":"10.1109\/TASL.2012.2231073","volume":"21","author":"N Zlatintsi","year":"2013","unstructured":"Zlatintsi, N., & Maragos, P. (2013). Multiscale fractal analysis of musical instrument signals with application to recognition. Audio, Speech, and Language Processing, IEEE Transactions on, 21, 737\u2013748.","journal-title":"Audio, Speech, and Language Processing, IEEE Transactions on"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02420-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-025-02420-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02420-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,7]],"date-time":"2025-06-07T06:05:11Z","timestamp":1749276311000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-025-02420-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,26]]},"references-count":81,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2025,7]]}},"alternative-id":["2420"],"URL":"https:\/\/doi.org\/10.1007\/s11263-025-02420-8","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,26]]},"assertion":[{"value":"28 December 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 March 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 March 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}