{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,13]],"date-time":"2026-02-13T23:18:53Z","timestamp":1771024733323,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":68,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Natural Science Foundation of China","award":["61932022"],"award-info":[{"award-number":["61932022"]}]},{"name":"National Natural Science Foundation of China","award":["61720106001"],"award-info":[{"award-number":["61720106001"]}]},{"name":"National Natural Science Foundation of China","award":["61971285"],"award-info":[{"award-number":["61971285"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3547783","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:43:01Z","timestamp":1665416581000},"page":"5649-5658","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":20,"title":["Dual Contrastive Learning for Spatio-temporal Representation"],"prefix":"10.1145","author":[{"given":"Shuangrui","family":"Ding","sequence":"first","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rui","family":"Qian","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongkai","family":"Xiong","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Self-supervised learning by cross-modal audio-video clustering. arXiv preprint arXiv:1911.12667","author":"Alwassel Humam","year":"2019","unstructured":"Humam Alwassel , Dhruv Mahajan , Bruno Korbar , Lorenzo Torresani , Bernard Ghanem , and Du Tran . 2019. Self-supervised learning by cross-modal audio-video clustering. arXiv preprint arXiv:1911.12667 ( 2019 ). Humam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani, Bernard Ghanem, and Du Tran. 2019. Self-supervised learning by cross-modal audio-video clustering. arXiv preprint arXiv:1911.12667 (2019)."},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6615"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV48630.2021.00171"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00994"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i2.16189"},{"key":"e_1_3_2_2_7_1","volume-title":"International conference on machine learning. PMLR, 1597--1607","author":"Chen Ting","year":"2020","unstructured":"Ting Chen , Simon Kornblith , Mohammad Norouzi , and Geoffrey Hinton . 2020 . A simple framework for contrastive learning of visual representations . In International conference on machine learning. PMLR, 1597--1607 . Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning. PMLR, 1597--1607."},{"key":"e_1_3_2_2_8_1","unstructured":"Jinwoo Choi Chen Gao C. E. Joseph Messou and Jia-Bin Huang. 2019. Why Can't I Dance in the Mall? Learning to Mitigate Scene Bias in Action Recognition. In NeurIPS.  Jinwoo Choi Chen Gao C. E. Joseph Messou and Jia-Bin Huang. 2019. Why Can't I Dance in the Mall? Learning to Mitigate Scene Bias in Action Recognition. In NeurIPS."},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00949"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.167"},{"key":"e_1_3_2_2_11_1","volume-title":"With a little help from my friends: Nearest-neighbor contrastive learning of visual representations. arXiv preprint arXiv:2104.14548","author":"Dwibedi Debidatta","year":"2021","unstructured":"Debidatta Dwibedi , Yusuf Aytar , Jonathan Tompson , Pierre Sermanet , and Andrew Zisserman . 2021. With a little help from my friends: Nearest-neighbor contrastive learning of visual representations. arXiv preprint arXiv:2104.14548 ( 2021 ). Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, and Andrew Zisserman. 2021. With a little help from my friends: Nearest-neighbor contrastive learning of visual representations. arXiv preprint arXiv:2104.14548 (2021)."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00331"},{"key":"e_1_3_2_2_13_1","volume-title":"Unsupervised representation learning by predicting image rotations. arXiv preprint arXiv:1803.07728","author":"Gidaris Spyros","year":"2018","unstructured":"Spyros Gidaris , Praveer Singh , and Nikos Komodakis . 2018. Unsupervised representation learning by predicting image rotations. arXiv preprint arXiv:1803.07728 ( 2018 ). Spyros Gidaris, Praveer Singh, and Nikos Komodakis. 2018. Unsupervised representation learning by predicting image rotations. arXiv preprint arXiv:1803.07728 (2018)."},{"key":"e_1_3_2_2_14_1","volume-title":"Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al.","author":"Grill Jean-Bastien","year":"2020","unstructured":"Jean-Bastien Grill , Florian Strub , Florent Altch\u00e9 , Corentin Tallec , Pierre H Richemond , Elena Buchatskaya , Carl Doersch , Bernardo Avila Pires , Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. 2020 . Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733 (2020). Jean-Bastien Grill, Florian Strub, Florent Altch\u00e9, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. 2020. Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733 (2020)."},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2019.00186"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58580-8_19"},{"key":"e_1_3_2_2_17_1","volume-title":"Self-supervised co-training for video representation learning. arXiv preprint arXiv:2010.09709","author":"Han Tengda","year":"2020","unstructured":"Tengda Han , Weidi Xie , and Andrew Zisserman . 2020b. Self-supervised co-training for video representation learning. arXiv preprint arXiv:2010.09709 ( 2020 ). Tengda Han, Weidi Xie, and Andrew Zisserman. 2020b. Self-supervised co-training for video representation learning. arXiv preprint arXiv:2010.09709 (2020)."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-49409-8_2"},{"key":"e_1_3_2_2_21_1","volume-title":"ASCNet: Self-supervised Video Representation Learning with Appearance-Speed Consistency. arXiv preprint arXiv:2106.02342","author":"Huang Deng","year":"2021","unstructured":"Deng Huang , Wenhao Wu , Weiwen Hu , Xu Liu , Dongliang He , Zhihua Wu , Xiangmiao Wu , Mingkui Tan , and Errui Ding . 2021b. ASCNet: Self-supervised Video Representation Learning with Appearance-Speed Consistency. arXiv preprint arXiv:2106.02342 ( 2021 ). Deng Huang, Wenhao Wu, Weiwen Hu, Xu Liu, Dongliang He, Zhihua Wu, Xiangmiao Wu, Mingkui Tan, and Errui Ding. 2021b. ASCNet: Self-supervised Video Representation Learning with Appearance-Speed Consistency. arXiv preprint arXiv:2106.02342 (2021)."},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01367"},{"key":"e_1_3_2_2_23_1","volume-title":"Space-time correspondence as a contrastive random walk. arXiv preprint arXiv:2006.14613","author":"Jabri Allan","year":"2020","unstructured":"Allan Jabri , Andrew Owens , and Alexei A Efros . 2020. Space-time correspondence as a contrastive random walk. arXiv preprint arXiv:2006.14613 ( 2020 ). Allan Jabri, Andrew Owens, and Alexei A Efros. 2020. Space-time correspondence as a contrastive random walk. arXiv preprint arXiv:2006.14613 (2020)."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58604-1_26"},{"key":"e_1_3_2_2_25_1","volume-title":"Self-supervised spatiotemporal feature learning via video rotation prediction. arXiv preprint arXiv:1811.11387","author":"Jing Longlong","year":"2018","unstructured":"Longlong Jing , Xiaodong Yang , Jingen Liu , and Yingli Tian . 2018. Self-supervised spatiotemporal feature learning via video rotation prediction. arXiv preprint arXiv:1811.11387 ( 2018 ). Longlong Jing, Xiaodong Yang, Jingen Liu, and Yingli Tian. 2018. Self-supervised spatiotemporal feature learning via video rotation prediction. arXiv preprint arXiv:1811.11387 (2018)."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018545"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2018.00092"},{"key":"e_1_3_2_2_28_1","volume-title":"Cycle-contrast for self-supervised video representation learning. arXiv preprint arXiv:2010.14810","author":"Kong Quan","year":"2020","unstructured":"Quan Kong , Wenpeng Wei , Ziwei Deng , Tomoaki Yoshinaga , and Tomokazu Murakami . 2020. Cycle-contrast for self-supervised video representation learning. arXiv preprint arXiv:2010.14810 ( 2020 ). Quan Kong, Wenpeng Wei, Ziwei Deng, Tomoaki Yoshinaga, and Tomokazu Murakami. 2020. Cycle-contrast for self-supervised video representation learning. arXiv preprint arXiv:2010.14810 (2020)."},{"key":"e_1_3_2_2_29_1","volume-title":"Mean Shift for Self-Supervised Learning. arXiv preprint arXiv:2105.07269","author":"Koohpayegani Soroush Abbasi","year":"2021","unstructured":"Soroush Abbasi Koohpayegani , Ajinkya Tejankar , and Hamed Pirsiavash . 2021. Mean Shift for Self-Supervised Learning. arXiv preprint arXiv:2105.07269 ( 2021 ). Soroush Abbasi Koohpayegani, Ajinkya Tejankar, and Hamed Pirsiavash. 2021. Mean Shift for Self-Supervised Learning. arXiv preprint arXiv:2105.07269 (2021)."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW54120.2021.00358"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00211"},{"key":"e_1_3_2_2_33_1","volume-title":"Xiaolong Wang, Jan Kautz, and Ming-Hsuan Yang.","author":"Li Xueting","year":"2019","unstructured":"Xueting Li , Sifei Liu , Shalini De Mello , Xiaolong Wang, Jan Kautz, and Ming-Hsuan Yang. 2019 . Joint-task self-supervised learning for temporal correspondence. arXiv preprint arXiv:1909.11895 (2019). Xueting Li, Sifei Liu, Shalini De Mello, Xiaolong Wang, Jan Kautz, and Ming-Hsuan Yang. 2019. Joint-task self-supervised learning for temporal correspondence. arXiv preprint arXiv:1909.11895 (2019)."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01231-1_32"},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6840"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.751"},{"key":"e_1_3_2_2_37_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 6707--6717","author":"Misra Ishan","unstructured":"Ishan Misra and Laurens van der Maaten. 2020. Self-supervised learning of pretext-invariant representations . In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 6707--6717 . Ishan Misra and Laurens van der Maaten. 2020. Self-supervised learning of pretext-invariant representations. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 6707--6717."},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_32"},{"key":"e_1_3_2_2_39_1","volume-title":"Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748","author":"van den Oord Aaron","year":"2018","unstructured":"Aaron van den Oord , Yazhe Li , and Oriol Vinyals . 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748 ( 2018 ). Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748 (2018)."},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01105"},{"key":"e_1_3_2_2_41_1","volume-title":"Enhancing Self-supervised Video Representation Learning via Multi-level Feature Optimization. arXiv preprint arXiv:2108.02183","author":"Ding Shuangrui","year":"2021","unstructured":"Rui Qian, Yuxi Li, Huabin Liu, John See, Shuangrui Ding , Xian Liu , Dian Li , and Weiyao Lin . 2021. Enhancing Self-supervised Video Representation Learning via Multi-level Feature Optimization. arXiv preprint arXiv:2108.02183 ( 2021 ). Rui Qian, Yuxi Li, Huabin Liu, John See, Shuangrui Ding, Xian Liu, Dian Li, and Weiyao Lin. 2021. Enhancing Self-supervised Video Representation Learning via Multi-level Feature Optimization. arXiv preprint arXiv:2108.02183 (2021)."},{"key":"e_1_3_2_2_42_1","volume-title":"Spatiotemporal contrastive video representation learning. arXiv preprint arXiv:2008.03800","author":"Qian Rui","year":"2020","unstructured":"Rui Qian , Tianjian Meng , Boqing Gong , Ming-Hsuan Yang , Huisheng Wang , Serge Belongie , and Yin Cui . 2020. Spatiotemporal contrastive video representation learning. arXiv preprint arXiv:2008.03800 ( 2020 ). Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, and Yin Cui. 2020. Spatiotemporal contrastive video representation learning. arXiv preprint arXiv:2008.03800 (2020)."},{"key":"e_1_3_2_2_43_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision. 1255--1265","author":"Recasens Adria","unstructured":"Adria Recasens , Pauline Luc , Jean-Baptiste Alayrac , Luyu Wang , Florian Strub , Corentin Tallec , Mateusz Malinowski , Viorica Pua trua ucean, Florent Altch\u00e9, Michal Valko, et al. 2021. Broaden your views for self-supervised video learning . In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 1255--1265 . Adria Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang, Florian Strub, Corentin Tallec, Mateusz Malinowski, Viorica Pua trua ucean, Florent Altch\u00e9, Michal Valko, et al. 2021. Broaden your views for self-supervised video learning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 1255--1265."},{"key":"e_1_3_2_2_44_1","volume-title":"Amir Roshan Zamir, and Mubarak Shah","author":"Soomro Khurram","year":"2012","unstructured":"Khurram Soomro , Amir Roshan Zamir, and Mubarak Shah . 2012 . UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012). Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012)."},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"crossref","unstructured":"Li Tao Xueting Wang and Toshihiko Yamasaki. 2020. Self-supervised video representation learning using inter-intra contrastive framework. In ACM MM. 2193--2201.  Li Tao Xueting Wang and Toshihiko Yamasaki. 2020. Self-supervised video representation learning using inter-intra contrastive framework. In ACM MM. 2193--2201.","DOI":"10.1145\/3394171.3413694"},{"key":"e_1_3_2_2_46_1","volume-title":"Proceedings, Part XI 16","author":"Tian Yonglong","year":"2020","unstructured":"Yonglong Tian , Dilip Krishnan , and Phillip Isola . 2020 . Contrastive multiview coding. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020 , Proceedings, Part XI 16 . Springer, 776--794. Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020. Contrastive multiview coding. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XI 16. Springer, 776--794."},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00675"},{"key":"e_1_3_2_2_48_1","volume-title":"Decomposing motion and content for natural video sequence prediction. arXiv preprint arXiv:1706.08033","author":"Villegas Ruben","year":"2017","unstructured":"Ruben Villegas , Jimei Yang , Seunghoon Hong , Xunyu Lin , and Honglak Lee . 2017. Decomposing motion and content for natural video sequence prediction. arXiv preprint arXiv:1706.08033 ( 2017 ). Ruben Villegas, Jimei Yang, Seunghoon Hong, Xunyu Lin, and Honglak Lee. 2017. Decomposing motion and content for natural video sequence prediction. arXiv preprint arXiv:1706.08033 (2017)."},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.18"},{"key":"e_1_3_2_2_50_1","unstructured":"Jinpeng Wang et al. 2021a. Enhancing unsupervised video representation learning by decoupling the scene and the motion. In AAAI21.  Jinpeng Wang et al. 2021a. Enhancing unsupervised video representation learning by decoupling the scene and the motion. In AAAI21."},{"key":"e_1_3_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01163"},{"key":"e_1_3_2_2_52_1","volume-title":"Self-supervised Video Representation Learning by Uncovering Spatio-temporal Statistics. arXiv preprint arXiv:2008.13426","author":"Wang Jiangliu","year":"2020","unstructured":"Jiangliu Wang , Jianbo Jiao , Linchao Bao , Shengfeng He , Wei Liu , and Yun-hui Liu. 2020b. Self-supervised Video Representation Learning by Uncovering Spatio-temporal Statistics. arXiv preprint arXiv:2008.13426 ( 2020 ). Jiangliu Wang, Jianbo Jiao, Linchao Bao, Shengfeng He, Wei Liu, and Yun-hui Liu. 2020b. Self-supervised Video Representation Learning by Uncovering Spatio-temporal Statistics. arXiv preprint arXiv:2008.13426 (2020)."},{"key":"e_1_3_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58520-4_30"},{"key":"e_1_3_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00267"},{"key":"e_1_3_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00736"},{"key":"e_1_3_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00393"},{"key":"e_1_3_2_2_57_1","volume-title":"MoDist: Motion Distillation for Self-supervised Video Representation Learning. arXiv preprint arXiv:2106.09703","author":"Xiao Fanyi","year":"2021","unstructured":"Fanyi Xiao , Joseph Tighe , and Davide Modolo . 2021. MoDist: Motion Distillation for Self-supervised Video Representation Learning. arXiv preprint arXiv:2106.09703 ( 2021 ). Fanyi Xiao, Joseph Tighe, and Davide Modolo. 2021. MoDist: Motion Distillation for Self-supervised Video Representation Learning. arXiv preprint arXiv:2106.09703 (2021)."},{"key":"e_1_3_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"e_1_3_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01058"},{"key":"e_1_3_2_2_60_1","volume-title":"Bag of Instances Aggregation Boosts Self-supervised Distillation. In International Conference on Learning Representations.","author":"Xu Haohang","year":"2021","unstructured":"Haohang Xu , Jiemin Fang , Xiaopeng Zhang , Lingxi Xie , Xinggang Wang , Wenrui Dai , Hongkai Xiong , and Qi Tian . 2021 a. Bag of Instances Aggregation Boosts Self-supervised Distillation. In International Conference on Learning Representations. Haohang Xu, Jiemin Fang, Xiaopeng Zhang, Lingxi Xie, Xinggang Wang, Wenrui Dai, Hongkai Xiong, and Qi Tian. 2021a. Bag of Instances Aggregation Boosts Self-supervised Distillation. In International Conference on Learning Representations."},{"key":"e_1_3_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3082567"},{"key":"e_1_3_2_2_62_1","volume-title":"Seed the views: Hierarchical semantic alignment for contrastive representation learning","author":"Xu Haohang","year":"2022","unstructured":"Haohang Xu , Xiaopeng Zhang , Hao Li , Lingxi Xie , Wenrui Dai , Hongkai Xiong , and Qi Tian . 2022. Seed the views: Hierarchical semantic alignment for contrastive representation learning . IEEE Transactions on Pattern Analysis and Machine Intelligence ( 2022 ). Haohang Xu, Xiaopeng Zhang, Hao Li, Lingxi Xie, Wenrui Dai, Hongkai Xiong, and Qi Tian. 2022. Seed the views: Hierarchical semantic alignment for contrastive representation learning. IEEE Transactions on Pattern Analysis and Machine Intelligence (2022)."},{"key":"e_1_3_2_2_63_1","volume-title":"Seco: Exploring sequence supervision for unsupervised representation learning. arXiv preprint arXiv:2008.00975","author":"Yao Ting","year":"2020","unstructured":"Ting Yao , Yiheng Zhang , Zhaofan Qiu , Yingwei Pan , and Tao Mei . 2020 b. Seco: Exploring sequence supervision for unsupervised representation learning. arXiv preprint arXiv:2008.00975 (2020). Ting Yao, Yiheng Zhang, Zhaofan Qiu, Yingwei Pan, and Tao Mei. 2020b. Seco: Exploring sequence supervision for unsupervised representation learning. arXiv preprint arXiv:2008.00975 (2020)."},{"key":"e_1_3_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00658"},{"key":"e_1_3_2_2_65_1","volume-title":"How Incomplete is Contrastive Learning? An Inter-intra Variant Dual Representation Method for Self-supervised Video Recognition. arXiv preprint arXiv:2107.01194","author":"Zhang Lin","year":"2021","unstructured":"Lin Zhang , Qi She , Zhengyang Shen , and Changhu Wang . 2021. How Incomplete is Contrastive Learning? An Inter-intra Variant Dual Representation Method for Self-supervised Video Recognition. arXiv preprint arXiv:2107.01194 ( 2021 ). Lin Zhang, Qi She, Zhengyang Shen, and Changhu Wang. 2021. How Incomplete is Contrastive Learning? An Inter-intra Variant Dual Representation Method for Self-supervised Video Recognition. arXiv preprint arXiv:2107.01194 (2021)."},{"key":"e_1_3_2_2_66_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV51458.2022.00105"},{"key":"e_1_3_2_2_67_1","volume-title":"Rynson WH Lau, and Stephen Lin","author":"Zhao Nanxuan","year":"2020","unstructured":"Nanxuan Zhao , Zhirong Wu , Rynson WH Lau, and Stephen Lin . 2020 . What makes instance discrimination good for transfer learning? arXiv preprint arXiv:2006.06606 (2020). Nanxuan Zhao, Zhirong Wu, Rynson WH Lau, and Stephen Lin. 2020. What makes instance discrimination good for transfer learning? arXiv preprint arXiv:2006.06606 (2020)."},{"key":"e_1_3_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.319"}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3547783","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3547783","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:41Z","timestamp":1750188641000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3547783"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":68,"alternative-id":["10.1145\/3503161.3547783","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3547783","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}