{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T13:25:51Z","timestamp":1784640351524,"version":"3.55.0"},"reference-count":80,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,1,10]],"date-time":"2024-01-10T00:00:00Z","timestamp":1704844800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,1,10]],"date-time":"2024-01-10T00:00:00Z","timestamp":1704844800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"National Key Research and Development Program of China Grant","award":["2018AAA0100400"],"award-info":[{"award-number":["2018AAA0100400"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U21B2013"],"award-info":[{"award-number":["U21B2013"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61971277"],"award-info":[{"award-number":["61971277"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Intell."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper focuses on self-supervised video representation learning. Most existing approaches follow the contrastive learning pipeline to construct positive and negative pairs by sampling different clips. However, this formulation tends to bias the static background and has difficulty establishing global temporal structures. The major reason is that the positive pairs, i.e., different clips sampled from the same video, have limited temporal receptive fields, and usually share similar backgrounds but differ in motions. To address these problems, we propose a framework to jointly utilize local clips and global videos to learn from detailed region-level correspondence as well as general long-term temporal relations. Based on a set of designed controllable augmentations, we implement accurate appearance and motion pattern alignment through soft spatio-temporal region contrast. Our formulation avoids the low-level redundancy shortcut with an adversarial mutual information minimization objective to improve the generalization ability. Moreover, we introduce local-global temporal order dependency to further bridge the gap between clip-level and video-level representations for robust temporal modeling. Extensive experiments demonstrate that our framework is superior on three video benchmarks in action recognition and video retrieval, and captures more accurate temporal dynamics.<\/jats:p>","DOI":"10.1007\/s44267-023-00034-7","type":"journal-article","created":{"date-parts":[[2024,1,10]],"date-time":"2024-01-10T10:02:11Z","timestamp":1704880931000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["Controllable augmentations for video representation learning"],"prefix":"10.1007","volume":"2","author":[{"given":"Rui","family":"Qian","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weiyao","family":"Lin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"John","family":"See","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dian","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,1,10]]},"reference":[{"key":"34_CR1","first-page":"4724","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"J. Carreira","year":"2017","unstructured":"Carreira, J., & Zisserman, A. (2017). Quo vadis, action recognition? A new model and the kinetics dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4724\u20134733). Piscataway: IEEE."},{"key":"34_CR2","first-page":"318","volume-title":"Proceedings of the 15th European conference on computer vision","author":"S. Xie","year":"2018","unstructured":"Xie, S., Sun, C., Huang, J., Tu, Z., & Murphy, K. (2018). Rethinking spatiotemporal feature learning: speed-accuracy trade-offs in video classification. In V. Ferrari, M. Hebert, C. Sminchisescu, et al. (Eds.), Proceedings of the 15th European conference on computer vision (pp. 318\u2013335). Cham: Springer."},{"key":"34_CR3","first-page":"6047","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"C. Gu","year":"2018","unstructured":"Gu, C., Sun, C., Ross, D. A., Vondrick, C., Pantofaru, C., Li, Y., et al. (2018). AVA: a video dataset of spatio-temporally localized atomic visual actions. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 6047\u20136056). Piscataway: IEEE."},{"key":"34_CR4","first-page":"961","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"F.C. Heilbron","year":"2015","unstructured":"Heilbron, F.C., Escorcia, V., Ghanem, B., & Niebles, J. C. (2015). Activitynet: a large-scale video benchmark for human activity understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 961\u2013970). Piscataway: IEEE."},{"key":"34_CR5","unstructured":"Liu, Y., Albanie, S., Nagrani, A., & Zisserman, A. (2019). Use what you have: video retrieval using representations from collaborative experts. arXiv preprint. arXiv:1907.13487."},{"key":"34_CR6","first-page":"2630","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"A. Miech","year":"2019","unstructured":"Miech, A., Zhukov, D., Alayrac, J.-B., Tapaswi, M., Laptev, I., & Sivic, J. (2019). Howto100m: learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 2630\u20132640). Piscataway: IEEE."},{"key":"34_CR7","unstructured":"Soomro, K., Zamir, A. R., & Shah, M. (2012). UCF101: a dataset of 101 human actions classes from videos in the wild. arXiv preprint. arXiv:1212.0402."},{"key":"34_CR8","first-page":"5843","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"R. Goyal","year":"2017","unstructured":"Goyal, R., Kahou, S. E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., et al. (2017). The \u201csomething something\u201d video database for learning and evaluating visual common sense. In Proceedings of the IEEE international conference on computer vision (pp. 5843\u20135851). Piscataway: IEEE."},{"key":"34_CR9","first-page":"520","volume-title":"Proceedings of the 15th European conference on computer vision","author":"Y. Li","year":"2018","unstructured":"Li, Y., Li, Y., & Vasconcelos, N. (2018). RESOUND: towards action recognition without representation bias. In V. Ferrari, M. Hebert, C. Sminchisescu, et al. (Eds.), Proceedings of the 15th European conference on computer vision (pp. 520\u2013535). Cham: Springer."},{"key":"34_CR10","first-page":"9919","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"S. Benaim","year":"2020","unstructured":"Benaim, S., Ephrat, A., Lang, O., Mosseri, I., Freeman, W. T., Rubinstein, M., et al. (2020). SpeedNet: learning the speediness in videos. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 9919\u20139928). Piscataway: IEEE."},{"key":"34_CR11","first-page":"527","volume-title":"Proceedings of the 14th European conference on computer vision","author":"I. Misra","year":"2016","unstructured":"Misra, I., Zitnick, C. L., & Hebert, M. (2016). Shuffle and learn: unsupervised learning using temporal order verification. In B. Leibe, J. Matas, N. Sebe, et al. (Eds.), Proceedings of the 14th European conference on computer vision (pp. 527\u2013544). Cham: Springer."},{"key":"34_CR12","first-page":"8545","volume-title":"Proceedings of the 33rd AAAI conference on artificial intelligence","author":"D. Kim","year":"2019","unstructured":"Kim, D., Cho, D., & Kweon, I. S. (2019). Self-supervised video representation learning with space-time cubic puzzles. In Proceedings of the 33rd AAAI conference on artificial intelligence (pp. 8545\u20138552). Palo Alto: AAAI Press."},{"key":"34_CR13","first-page":"425","volume-title":"Proceedings of the 16th European conference on computer vision","author":"J. Simon","year":"2020","unstructured":"Simon, J., Meishvili, G., & Favaro, P. (2020). Video representation learning by recognizing temporal transformations. In A. Vedaldi, H. Bischof, T. Brox, et al. (Eds.), Proceedings of the 16th European conference on computer vision (pp. 425\u2013442). Cham: Springer."},{"key":"34_CR14","first-page":"10334","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"D. Xu","year":"2019","unstructured":"Xu, D., Xiao, J., Zhao, Z., Shao, J., Xie, D., & Zhuang, Y. (2019). Self-supervised spatiotemporal learning via video clip order prediction. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 10334\u201310343). Piscataway: IEEE."},{"issue":"7","key":"34_CR15","first-page":"3791","volume":"44","author":"J. Wang","year":"2022","unstructured":"Wang, J., Jiao, J., Bao, L., He, S., Liu, W., & Liu, Y. (2022). Self-supervised video representation learning by uncovering spatio-temporal statistics. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(7), 3791\u20133806.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"34_CR16","unstructured":"Gordon, D., Ehsani, K., Fox, D., & Farhadi, A. (2020). Watching the world go by: representation learning from unlabeled videos. arXiv preprint. arXiv:2003.07990."},{"key":"34_CR17","first-page":"6964","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"R. Qian","year":"2021","unstructured":"Qian, R., Meng, T., Gong, B., Yang, M.-H., Wang, H., Belongie, S. J., et al. (2021). Spatiotemporal contrastive video representation learning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 6964\u20136974). Piscataway: IEEE."},{"key":"34_CR18","first-page":"504","volume-title":"Proceedings of the 16th European conference on computer vision","author":"J. Wang","year":"2020","unstructured":"Wang, J., Jiao, J., & Liu, Y.-H. (2020). Self-supervised video representation learning by pace prediction. In A. Vedaldi, H. Bischof, T. Brox, et al. (Eds.), Proceedings of the 16th European conference on computer vision (pp. 504\u2013521). Cham: Springer."},{"key":"34_CR19","first-page":"10656","volume-title":"Proceedings of the 35th AAAI conference on artificial intelligence","author":"T. Yao","year":"2021","unstructured":"Yao, T., Zhang, Y., Qiu, Z., Pan, Y., & Mei, T. (2021). SeCo: exploring sequence supervision for unsupervised representation learning. In Proceedings of the 35th AAAI conference on artificial intelligence (pp. 10656\u201310664). Palo Alto: AAAI Press."},{"key":"34_CR20","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"T. Han","year":"2020","unstructured":"Han, T., Xie, W., & Zisserman, A. (2020). Self-supervised co-training for video representation learning. In H. Larochelle, M. Ranzato, R. Hadsell, et al. (Eds.), Proceedings of the 34th international conference on neural information processing systems, Red Hook: Curran Associates."},{"key":"34_CR21","first-page":"3188","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision workshops","author":"H. Kuang","year":"2021","unstructured":"Kuang, H., Zhu, Y., Zhang, Z., Li, X., Tighe, J., Schwertfeger, S., et al. (2021). Video contrastive learning with global context. In Proceedings of the IEEE\/CVF international conference on computer vision workshops (pp. 3188\u20133197). Piscataway: IEEE."},{"key":"34_CR22","first-page":"11205","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"T. Pan","year":"2021","unstructured":"Pan, T., Song, Y., Yang, T., Jiang, W., & Liu, W. (2021). Videomoco: contrastive video representation learning with temporally adversarial examples. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 11205\u201311214). Piscataway: IEEE."},{"key":"34_CR23","first-page":"11804","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"J. Wang","year":"2021","unstructured":"Wang, J., Gao, Y., Li, K., Lin, Y., Ma, A. J., Cheng, H., et al. (2021). Removing the background by adding the background: towards background robust self-supervised video representation learning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 11804\u201311813). Piscataway: IEEE."},{"key":"34_CR24","first-page":"10129","volume-title":"Proceedings of the 35th AAAI conference on artificial intelligence","author":"J. Wang","year":"2021","unstructured":"Wang, J., Gao, Y., Li, K., Hu, J., Jiang, X., Guo, X., et al. (2021). Enhancing unsupervised video representation learning by decoupling the scene and the motion. In Proceedings of the 35th AAAI conference on artificial intelligence (pp. 10129\u201310137). Menlo Park: AAAI Press."},{"key":"34_CR25","first-page":"9726","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"K. He","year":"2020","unstructured":"He, K., Fan, H., Wu, Y., Xie, S., & Girshick, R. B. (2020). Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 9726\u20139735). Piscataway: IEEE."},{"key":"34_CR26","first-page":"1597","volume-title":"Proceedings of the 37th international conference on machine learning","author":"T. Chen","year":"2020","unstructured":"Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. E. (2020). A simple framework for contrastive learning of visual representations. In Proceedings of the 37th international conference on machine learning (pp. 1597\u20131607). Stroudsburg: International Machine Learning Society."},{"key":"34_CR27","unstructured":"van den Oord, A., Li, Y., & Vinyals, O. (2018). Representation learning with contrastive predictive coding. arXiv preprint. arXiv:1807.03748."},{"key":"34_CR28","first-page":"1735","volume-title":"Proceedings of the IEEE computer society conference on computer vision and pattern recognition","author":"R. Hadsell","year":"2006","unstructured":"Hadsell, R., Chopra, S., & LeCun, Y. (2006). Dimensionality reduction by learning an invariant mapping. In Proceedings of the IEEE computer society conference on computer vision and pattern recognition (pp. 1735\u20131742). Piscataway: IEEE."},{"key":"34_CR29","volume-title":"Proceedings of the 13th international conference on artificial intelligence and statistics","author":"M. Gutmann","year":"2010","unstructured":"Gutmann, M., & Hyv\u00e4rinen, A. (2010). Noise-contrastive estimation: a new estimation principle for unnormalized statistical models. In Y. W. Teh & D. M. Titterington (Eds.), Proceedings of the 13th international conference on artificial intelligence and statistics. Retrieved Novermber 3, 2023, from http:\/\/proceedings.mlr.press\/v9\/gutmann10a.html."},{"key":"34_CR30","first-page":"3733","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Z. Wu","year":"2018","unstructured":"Wu, Z., Xiong, Y., Yu, S. X., & Lin, D. (2018). Unsupervised feature learning via non-parametric instance discrimination. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 3733\u20133742). Piscataway: IEEE."},{"key":"34_CR31","first-page":"776","volume-title":"Proceedings of the 16th European conference on computer vision","author":"Y. Tian","year":"2020","unstructured":"Tian, Y., Krishnan, D., & Isola, P. (2020). Contrastive multiview coding. In A. Vedaldi, H. Bischof, T. Brox, et al. (Eds.), Proceedings of the 16th European conference on computer vision (pp. 776\u2013794). Cham: Springer."},{"key":"34_CR32","volume-title":"Proceedings of the 7th international conference on learning representations","author":"R. D. Hjelm","year":"2019","unstructured":"Hjelm, R. D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Bachman, P., Trischler, A., et al. (2019). Learning deep representations by mutual information estimation and maximization. In Proceedings of the 7th international conference on learning representations. Retrieved November 3, 2023, from https:\/\/openreview.net\/forum?id=Bklr3j0cKX."},{"key":"34_CR33","first-page":"16684","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Z. Xie","year":"2021","unstructured":"Xie, Z., Lin, Y., Zhang, Z., Cao, Y., Lin, S., & Hu, H. (2021). Propagate yourself: exploring pixel-level consistency for unsupervised visual representation learning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 16684\u201316693). Piscataway: IEEE."},{"key":"34_CR34","first-page":"3024","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"X. Wang","year":"2021","unstructured":"Wang, X., Zhang, R., Shen, C., Kong, T., & Li, L. (2021). Dense contrastive learning for self-supervised visual pre-training. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 3024\u20133033). Piscataway: IEEE."},{"key":"34_CR35","first-page":"667","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"H.-Y. Lee","year":"2017","unstructured":"Lee, H.-Y., Huang, J.-B., Singh, M., & Yang, M.-H. (2017). Unsupervised representation learning by sorting sequences. In Proceedings of the IEEE international conference on computer vision (pp. 667\u2013676). Piscataway: IEEE."},{"key":"34_CR36","first-page":"402","volume-title":"Proceedings of the 15th European conference on computer vision","author":"C. Vondrick","year":"2018","unstructured":"Vondrick, C., Shrivastava, A., Fathi, A., Guadarrama, S., & Murphy, K. (2018). Tracking emerges by colorizing videos. In V. Ferrari, M. Hebert, C. Sminchisescu, et al. (Eds.), Proceedings of the 15th European conference on computer vision (pp. 402\u2013419). Cham: Springer."},{"key":"34_CR37","first-page":"2566","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"X. Wang","year":"2019","unstructured":"Wang, X., Jabri, A., & Efros, A. A. (2019). Learning correspondence from the cycle-consistency of time. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 2566\u20132576). Piscataway: IEEE."},{"key":"34_CR38","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"A. Jabri","year":"2020","unstructured":"Jabri, A., Owens, A., & Efros, A. A. (2020). Space-time correspondence as a contrastive random walk. In H. Larochelle, M. Ranzato, R. Hadsell, et al. (Eds.), Proceedings of the 34th international conference on neural information processing systems, Red Hook: Curran Associates."},{"key":"34_CR39","first-page":"317","volume-title":"Proceedings of the 33rd international conference on neural information processing systems","author":"X. Li","year":"2019","unstructured":"Li, X., Liu, S., De Mello, S., Wang, X., Kautz, J., & Yang, M.-H. (2019). Joint-task self-supervised learning for temporal correspondence. In H. M. Wallach, H. Larochelle, A. Beygelzimer, et al. (Eds.), Proceedings of the 33rd international conference on neural information processing systems (pp. 317\u2013327). Red Hook: Curran Associates."},{"key":"34_CR40","volume-title":"Proceedings of the 5th international conference on learning representations","author":"R. Villegas","year":"2017","unstructured":"Villegas, R., Yang, J., Hong, S., Lin, X., & Lee, H. (2017). Decomposing motion and content for natural video sequence prediction. In Proceedings of the 5th international conference on learning representations. Retrieved Novermber 3, 2023, from https:\/\/openreview.net\/forum?id=rkEFLFqee."},{"key":"34_CR41","first-page":"7101","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Z. Luo","year":"2017","unstructured":"Luo, Z., Peng, B., Huang, D.-A., Alahi, A., & Li, F.F. (2017). Unsupervised learning of long-term motion dynamics for videos. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7101\u20137110). Piscataway: IEEE."},{"key":"34_CR42","first-page":"1","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"H. Alwassel","year":"2020","unstructured":"Alwassel, H., Mahajan, D., Korbar, B., Torresani, L., Ghanem, B., & Tran, D. (2020). Self-supervised learning by cross-modal audio-video clustering. In H. Larochelle, M. Ranzato, R. Hadsell, et al. (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 1\u201313). Red Hook: Curran Associates."},{"key":"34_CR43","first-page":"130","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"A. J. Piergiovanni","year":"2020","unstructured":"Piergiovanni, A. J., Angelova, A., & Ryoo, M. S. (2020). Evolving losses for unsupervised video representation learning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 130\u2013139). Piscataway: IEEE."},{"key":"34_CR44","doi-asserted-by":"crossref","unstructured":"Liu, Y., Wang, K., Lan, H., & Lin, L. (2021). Temporal contrastive graph for self-supervised video representation learning. arXiv preprint. arXiv:2101.00820.","DOI":"10.1109\/TIP.2022.3147032"},{"key":"34_CR45","first-page":"1483","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision workshops","author":"T. Han","year":"2019","unstructured":"Han, T., Xie, W., & Zisserman, A. (2019). Video representation learning by dense predictive coding. In Proceedings of the IEEE\/CVF international conference on computer vision workshops (pp. 1483\u20131492). Piscataway: IEEE."},{"key":"34_CR46","first-page":"312","volume-title":"Proceedings of the 16th European conference on computer vision","author":"T. Han","year":"2020","unstructured":"Han, T., Xie, W., & Zisserman, A. (2020). Memory-augmented dense predictive coding for video representation learning. In A. Vedaldi, H. Bischof, T. Brox, et al. (Eds.), Proceedings of the 16th European conference on computer vision (pp. 312\u2013329). Cham: Springer."},{"key":"34_CR47","unstructured":"Yang, C., Xu, Y., Dai, B., & Zhou, B. (2020). Video representation learning with visual tempo consistency. arXiv preprint. arXiv:2006.15489."},{"key":"34_CR48","first-page":"1045","volume-title":"Proceedings of the 35th AAAI conference on artificial intelligence","author":"P. Chen","year":"2021","unstructured":"Chen, P., Huang, D., He, D., Long, X., Zeng, R., Wen, S., et al. (2021). RSPNet: relative speed perception for unsupervised video representation learning. In Proceedings of the 35th AAAI conference on artificial intelligence (pp. 1045\u20131053). Palo Alto: AAAI Press."},{"key":"34_CR49","first-page":"2085","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"R. Li","year":"2021","unstructured":"Li, R., Zhang, Y., Qiu, Z., Yao, T., Liu, D., & Mei, T. (2021). Motion-focused contrastive learning of video representations. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 2085\u20132094). Piscataway: IEEE."},{"key":"34_CR50","first-page":"145","volume-title":"Proceedings of the 17th European conference on computer vision","author":"R. Qian","year":"2022","unstructured":"Qian, R., Ding, S., Liu, X., & Lin, D. (2022). Static and dynamic concepts for self-supervised video representation learning. In S. Avidan, G. J. Brostow, M. Ciss\u00e9, et al. (Eds.), Proceedings of the 17th European conference on computer vision (pp. 145\u2013164). Cham: Springer."},{"key":"34_CR51","doi-asserted-by":"publisher","first-page":"5649","DOI":"10.1145\/3503161.3547783","volume-title":"Proceedings of the 30th ACM international conference on multimedia","author":"S. Ding","year":"2022","unstructured":"Ding, S., Qian, R., & Xiong, H. (2022). Dual contrastive learning for spatio-temporal representation. In J. Magalh\u00e3es, A. Del Bimbo, S. Satoh, et al. (Eds.), Proceedings of the 30th ACM international conference on multimedia (pp. 5649\u20135658). New York: ACM."},{"key":"34_CR52","first-page":"20","volume-title":"Proceedings of the 17th European conference on computer vision workshops","author":"Y. Liu","year":"2022","unstructured":"Liu, Y., Chen, J., & Wu, H. (2022). MoQuad: motion-focused quadruple construction for video contrastive learning. In L. Karlinsky, T. Michaeli, & K. Nishino (Eds.), Proceedings of the 17th European conference on computer vision workshops (pp. 20\u201338). Cham: Springer."},{"key":"34_CR53","first-page":"9706","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"S. Ding","year":"2022","unstructured":"Ding, S., Li, M., Yang, T., Qian, R., Xu, H., Chen, Q., et al. (2022). Motion-aware contrastive video representation learning via foreground-background merging. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 9706\u20139716). Piscataway: IEEE."},{"key":"34_CR54","first-page":"7025","volume-title":"Proceedings of the 35th international conference on neural information processing systems","author":"S. Ma","year":"2021","unstructured":"Ma, S., Zeng, Z., McDuff, D., & Song, Y. (2021). Contrastive learning of global and local video representations. In M. Ranzato, A. Beygelzimer, Y. N. Dauphin, et al. (Eds.), Proceedings of the 35th international conference on neural information processing systems (pp. 7025\u20137040). Red Hook: Curran Associates."},{"key":"34_CR55","first-page":"1235","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"A. Recasens","year":"2021","unstructured":"Recasens, A., Luc, P., Alayrac, J.-B., Wang, L., Strub, F., Tallec, C., et al. (2021). Broaden your views for self-supervised video learning. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 1235\u20131245). Piscataway: IEEE."},{"key":"34_CR56","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2022.103406","volume":"219","author":"I. R. Dave","year":"2022","unstructured":"Dave, I. R., Gupta, R., Rizve, M. N., & Shah, M. (2022). TCLR: temporal contrastive learning for video representation. Computer Vision and Image Understanding, 219, 103406.","journal-title":"Computer Vision and Image Understanding"},{"key":"34_CR57","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"N. Behrmann","year":"2021","unstructured":"Behrmann, N., Fayyaz, M., Gall, J., & Noroozi, M. (2021). Long short view feature decomposition via contrastive video representation learning. In Proceedings of the IEEE\/CVF international conference on computer vision, Piscataway: IEEE."},{"issue":"10","key":"34_CR58","doi-asserted-by":"crossref","first-page":"12408","DOI":"10.1109\/TPAMI.2023.3273415","volume":"45","author":"Z. Qing","year":"2023","unstructured":"Qing, Z., Zhang, S., Huang, Z., Xu, Y., Wang, X., Gao, C., et al. (2023). Self-supervised learning from untrimmed videos via hierarchical consistency. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10), 12408\u201312426.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"34_CR59","first-page":"530","volume-title":"Proceedings of the 35th international conference on machine learning","author":"M.I. Belghazi","year":"2018","unstructured":"Belghazi, M.I., Baratin, A., Rajeswar, S., Ozair, S., Bengio, Y., Hjelm, R. D., et al. (2018). Mutual information neural estimation. In J. G. Dy & A. Krause (Eds.), Proceedings of the 35th international conference on machine learning (pp. 530\u2013539). Stroudsburg: International Machine Learning Society."},{"issue":"1","key":"34_CR60","doi-asserted-by":"publisher","first-page":"71","DOI":"10.1016\/0010-0277(93)90058-4","volume":"48","author":"J. L. Elman","year":"1993","unstructured":"Elman, J. L. (1993). Learning and development in neural networks: the importance of starting small. Cognition, 48(1), 71\u201399.","journal-title":"Cognition"},{"key":"34_CR61","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1145\/1553374.1553380","volume-title":"Proceedings of the 26th annual international conference on machine learning","author":"Y. Bengio","year":"2009","unstructured":"Bengio, Y., Louradour, J., Collobert, R., & Weston, J. (2009). Curriculum learning. In A. P. Danyluk, L. Bottou, & M. L. Littman (Eds.), Proceedings of the 26th annual international conference on machine learning (pp. 41\u201348). Stroudsburg: International Machine Learning Society."},{"key":"34_CR62","first-page":"6453","volume-title":"Proceedings of the IEEE international conference on robotics and automation","author":"A. Murali","year":"2018","unstructured":"Murali, A., Pinto, L., Gandhi, D., & Gupta, A. (2018). CASSL: curriculum accelerated self-supervised learning. In Proceedings of the IEEE international conference on robotics and automation (pp. 6453\u20136460). Piscataway: IEEE."},{"key":"34_CR63","first-page":"2556","volume-title":"IEEE international conference on computer vision","author":"H. Kuehne","year":"2011","unstructured":"Kuehne, H., Jhuang, H., Garrote, E., Poggio, T. A., & Serre, T. (2011). HMDB: a large video database for human motion recognition. In D. N. Metaxas, L. Quan, A. Sanfeliu, et al. (Eds.), IEEE international conference on computer vision (pp. 2556\u20132563). Piscataway: IEEE."},{"key":"34_CR64","first-page":"3154","volume-title":"Proceedings of the IEEE international conference on computer vision workshops","author":"K. Hara","year":"2017","unstructured":"Hara, K., Kataoka, H., & Satoh, Y. (2017). Learning spatio-temporal features with 3D residual networks for action recognition. In Proceedings of the IEEE international conference on computer vision workshops (pp. 3154\u20133160). Piscataway: IEEE."},{"key":"34_CR65","first-page":"11701","volume-title":"Proceedings of the 34th AAAI conference on artificial intelligence","author":"D. Luo","year":"2020","unstructured":"Luo, D., Liu, C., Zhou, Y., Yang, D., Ma, C., Ye, Q., et al. (2020). Video cloze procedure for self-supervised spatio-temporal learning. In Proceedings of the 34th AAAI conference on artificial intelligence (pp. 11701\u201311708). Palo Alto: AAAI Press."},{"key":"34_CR66","unstructured":"Sun, C., Baradel, F., Murphy, K., & Schmid, C. (2019). Learning video representations using contrastive bidirectional transformer. arXiv preprint. arXiv:1906.05743."},{"key":"34_CR67","first-page":"7970","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"R. Qian","year":"2021","unstructured":"Qian, R., Li, Y., Liu, H., See, J., Ding, S., Liu, X., et al. (2021). Enhancing self-supervised video representation learning via multi-level feature optimization. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 7970\u20137981). Piscataway: IEEE."},{"key":"34_CR68","first-page":"14691","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"J. Park","year":"2022","unstructured":"Park, J., Lee, J., Kim, I.-J., & Sohn, K. (2022). Probabilistic representations for video contrastive learning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 14691\u201314701). Piscataway: IEEE."},{"key":"34_CR69","first-page":"8076","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"D. Huang","year":"2021","unstructured":"Huang, D., Wu, W., Hu, W., Liu, X., He, D., Wu, Z., et al. (2021). ASCNet: self-supervised video representation learning with appearance-speed consistency. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 8076\u20138085). Piscataway: IEEE."},{"key":"34_CR70","first-page":"9950","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"J. Simon","year":"2021","unstructured":"Simon, J., & Jin, H. (2021). Time-equivariant contrastive video representation learning. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 9950\u20139960). Piscataway: IEEE."},{"key":"34_CR71","first-page":"3299","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"C. Feichtenhofer","year":"2021","unstructured":"Feichtenhofer, C., Fan, H., Xiong, B., Girshick, R. B., & He, K. (2021). A large-scale study on unsupervised spatiotemporal representation learning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 3299\u20133309). Piscataway: IEEE."},{"key":"34_CR72","first-page":"1","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"Y. M. Asano","year":"2020","unstructured":"Asano, Y. M., Patrick, M., Rupprecht, C., & Vedaldi, A. (2020). Labelling unlabelled videos from scratch with multi-modal self-supervision. In H. Larochelle, M. Ranzato, R. Hadsell, et al. (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 1\u201312). Red Hook: Curran Associates."},{"key":"34_CR73","unstructured":"Patrick, M., Asano, Y. M., Kuznetsova, P., Fong, R., Henriques, J. F., Zweig, G., et al. (2020). Multi-modal self-supervision from generalized data transformations. arXiv preprint. arXiv:2003.04298."},{"key":"34_CR74","first-page":"20","volume-title":"Proceedings of the 14th European conference on computer vision","author":"L. Wang","year":"2016","unstructured":"Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., et al. (2016). Temporal segment networks: towards good practices for deep action recognition. In B. Leibe, J. Matas, N. Sebe, et al. (Eds.), Proceedings of the 14th European conference on computer vision (pp. 20\u201336). Cham: Springer."},{"key":"34_CR75","first-page":"851","volume-title":"Proceedings of the 33rd international conference on neural information processing systems","author":"J. Choi","year":"2019","unstructured":"Choi, J., Gao, C., Messou, J. C. E., & Huang, J.-B. (2019). Why can\u2019t I dance in the mall? Learning to mitigate scene bias in action recognition. In H. M. Wallach, H. Larochelle, A. Beygelzimer, et al. (Eds.), Proceedings of the 33rd international conference on neural information processing systems (pp. 851\u2013863). Red Hook: Curran Associates."},{"key":"34_CR76","first-page":"6547","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Y. Yao","year":"2020","unstructured":"Yao, Y., Liu, C., Luo, D., Zhou, Y., & Ye, Q. (2020). Video playback rate perception for self-supervised spatio-temporal representation learning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 6547\u20136556). Piscataway: IEEE."},{"key":"34_CR77","doi-asserted-by":"crossref","unstructured":"Tao, L., Wang, X., & Yamasaki, T. (2020). Self-supervised video representation using pretext-contrastive learning. arXiv preprint. arXiv:2010.15464.","DOI":"10.1145\/3394171.3413694"},{"key":"34_CR78","first-page":"10451","volume-title":"Proceedings of the 34th AAAI conference on artificial intelligence","author":"K. Baek","year":"2020","unstructured":"Baek, K., Lee, M., & Psynet, H. S. (2020). Self-supervised approach to object localization using point symmetric transformation. In Proceedings of the 34th AAAI conference on artificial intelligence (pp. 10451\u201310459). Palo Alto: AAAI Press."},{"key":"34_CR79","first-page":"1779","volume-title":"Proceedings of the 37th international conference on machine learning","author":"P. Cheng","year":"2020","unstructured":"Cheng, P., Hao, W., Dai, S., Liu, J., Gan, Z., & Carin, L. (2020). CLUB: a contrastive log-ratio upper bound of mutual information. In Proceedings of the 37th international conference on machine learning (pp. 1779\u20131788). Stroudsburg: International Machine Learning Society."},{"key":"34_CR80","first-page":"271","volume-title":"Proceedings of the 30th international conference on neural information processing systems","author":"S. Nowozin","year":"2016","unstructured":"Nowozin, S., Cseke, B., & Tomioka, R. (2016). f-GAN: training generative neural samplers using variational divergence minimization. In D. D. Lee, M. Sugiyama, U. von Luxburg, et al. (Eds.), Proceedings of the 30th international conference on neural information processing systems (pp. 271\u2013279). Red Hook: Curran Associates."}],"container-title":["Visual Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-023-00034-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44267-023-00034-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-023-00034-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,7]],"date-time":"2024-11-07T20:34:55Z","timestamp":1731011695000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44267-023-00034-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,10]]},"references-count":80,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["34"],"URL":"https:\/\/doi.org\/10.1007\/s44267-023-00034-7","relation":{},"ISSN":["2731-9008"],"issn-type":[{"value":"2731-9008","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,10]]},"assertion":[{"value":"22 May 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 December 2023","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 December 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 January 2024","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Weiyao Lin is an Associate Editor for Visual Intelligence and was not involved in the editorial review of, or the decision to publish, this article. The authors declare that there are no other competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"1"}}