{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,12]],"date-time":"2026-07-12T15:01:47Z","timestamp":1783868507803,"version":"3.55.0"},"reference-count":223,"publisher":"Association for Computing Machinery (ACM)","issue":"12","license":[{"start":{"date-parts":[[2024,10,1]],"date-time":"2024-10-01T00:00:00Z","timestamp":1727740800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2024,12,31]]},"abstract":"<jats:p>\n            Video analysis tasks such as action recognition have received increasing research interest with growing applications in fields such as smart healthcare, thanks to the introduction of large-scale datasets and deep learning based representations. However, video models trained on existing datasets suffer from significant performance degradation when deployed directly to real-world applications due to domain shifts between the training public video datasets (source video domains) and real-world videos (target video domains). Further, with the high cost of video annotation, it is more practical to use unlabeled videos for training. To tackle performance degradation and address concerns in high video annotation cost uniformly, video unsupervised domain adaptation (VUDA) is introduced to adapt video models from the labeled source domain to the unlabeled target domain by alleviating video domain shift, improving the generalizability and portability of video models. This article surveys recent progress in VUDA with deep learning. We begin with the motivation of VUDA, followed by its definition, and recent progress of methods for both closed-set VUDA and VUDA under different scenarios, and current benchmark datasets for VUDA research. Eventually, future directions are provided to promote further VUDA research. The repository of this survey is provided at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/xuyu0010\/awesome-video-domain-adaptation\">https:\/\/github.com\/xuyu0010\/awesome-video-domain-adaptation<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3679010","type":"journal-article","created":{"date-parts":[[2024,7,22]],"date-time":"2024-07-22T10:56:16Z","timestamp":1721645776000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Video Unsupervised Domain Adaptation with Deep Learning: A Comprehensive Survey"],"prefix":"10.1145","volume":"56","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4292-7379","authenticated-orcid":false,"given":"Yuecong","family":"Xu","sequence":"first","affiliation":[{"name":"Department of Electrical and Computer Engineering, National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7703-3490","authenticated-orcid":false,"given":"Haozhi","family":"Cao","sequence":"additional","affiliation":[{"name":"School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7137-4136","authenticated-orcid":false,"given":"Lihua","family":"Xie","sequence":"additional","affiliation":[{"name":"School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0762-6562","authenticated-orcid":false,"given":"Xiao-li","family":"Li","sequence":"additional","affiliation":[{"name":"Institute for Infocomm Research, A*STAR, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1719-0328","authenticated-orcid":false,"given":"Zhenghua","family":"Chen","sequence":"additional","affiliation":[{"name":"Institute for Infocomm Research, A*STAR, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8075-0439","authenticated-orcid":false,"given":"Jianfei","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Mechanical and Aerospace Engineering, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,10]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"1","volume-title":"Proceedings of the 2020 International Joint Conference on Neural Networks (IJCNN \u201920)","author":"Arazo Eric","year":"2020","unstructured":"Eric Arazo, Diego Ortego, Paul Albert, Noel E. O\u2019Connor, and Kevin McGuinness. 2020. Pseudo-labeling and confirmation bias in deep semi-supervised learning. In Proceedings of the 2020 International Joint Conference on Neural Networks (IJCNN \u201920). IEEE, 1\u20138."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"e_1_3_1_4_2","first-page":"3575","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Ayazoglu Mustafa","year":"2013","unstructured":"Mustafa Ayazoglu, Burak Yilmaz, Mario Sznaier, and Octavia Camps. 2013. Finding causal interactions in video sequences. In Proceedings of the IEEE International Conference on Computer Vision. IEEE, 3575\u20133582."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV51458.2022.00266"},{"key":"e_1_3_1_6_2","first-page":"1","article-title":"Transfer learning for image classification using VGG19: Caltech-101 image data set","volume":"14","author":"Bansal Monika","year":"2021","unstructured":"Monika Bansal, Munish Kumar, Monika Sachdeva, and Ajay Mittal. 2021. Transfer learning for image classification using VGG19: Caltech-101 image data set. Journal of Ambient Intelligence and Humanized Computing 14 (2021), 1\u201312.","journal-title":"Journal of Ambient Intelligence and Humanized Computing"},{"key":"e_1_3_1_7_2","unstructured":"Oscar Beijbom. 2012. Domain adaptations for computer vision applications. arxiv:1211.4860 [cs.CV] (2012)."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-009-5152-4"},{"issue":"1","key":"e_1_3_1_9_2","doi-asserted-by":"crossref","first-page":"101","DOI":"10.1109\/TCSVT.2016.2595331","article-title":"Lidar-based gait analysis and activity recognition in a 4D surveillance system","volume":"28","author":"Benedek Csaba","year":"2016","unstructured":"Csaba Benedek, Bence G\u00e1lai, Bal\u00e1zs Nagy, and Zsolt Jank\u00f3. 2016. Lidar-based gait analysis and activity recognition in a 4D surveillance system. IEEE Transactions on Circuits and Systems for Video Technology 28, 1 (2016), 101\u2013113.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_10_2","volume-title":"Reproducing Kernel Hilbert Spaces in Probability and Statistics","author":"Berlinet Alain","year":"2011","unstructured":"Alain Berlinet and Christine Thomas-Agnan. 2011. Reproducing Kernel Hilbert Spaces in Probability and Statistics. Springer Science & Business Media, New York, NY."},{"key":"e_1_3_1_11_2","volume-title":"Proceedings of the International Conference on Machine Learning (ICML \u201921)","volume":"2","author":"Bertasius Gedas","year":"2021","unstructured":"Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding? In Proceedings of the International Conference on Machine Learning (ICML \u201921), Vol. 2. 1\u201312."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2008.04.005"},{"issue":"5","key":"e_1_3_1_13_2","first-page":"1","article-title":"Gradient local auto-correlation features for depth human action recognition","volume":"3","author":"Bulbul Mohammad Farhad","year":"2021","unstructured":"Mohammad Farhad Bulbul and Hazrat Ali. 2021. Gradient local auto-correlation features for depth human action recognition. SN Applied Sciences 3, 5 (2021), 1\u201313.","journal-title":"SN Applied Sciences"},{"key":"e_1_3_1_14_2","volume-title":"Proceedings of the India-Norway Workshop on Web Concepts and Technologies","volume":"112","author":"Bungum Lars","year":"2011","unstructured":"Lars Bungum and Bj\u00f6rn Gamb\u00e4ck. 2011. A survey of domain adaptation in machine translation: Towards a refinement of domain space. In Proceedings of the India-Norway Workshop on Web Concepts and Technologies, Vol. 112. 1\u20139."},{"key":"e_1_3_1_15_2","first-page":"11457","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Cai Qi","year":"2019","unstructured":"Qi Cai, Yingwei Pan, Chong-Wah Ngo, Xinmei Tian, Lingyu Duan, and Ting Yao. 2019. Exploring object relation in mean teacher for cross-domain detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 11457\u201311466."},{"key":"e_1_3_1_16_2","first-page":"135","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201918)","author":"Cao Zhangjie","year":"2018","unstructured":"Zhangjie Cao, Lijia Ma, Mingsheng Long, and Jianmin Wang. 2018. Partial adversarial domain adaptation. In Proceedings of the European Conference on Computer Vision (ECCV \u201918). 135\u2013150."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_1_18_2","unstructured":"Santiago Castro and Fabian Caba Heilbron. 2022. FitCLIP: Refining large-scale pretrained image-text models for zero-shot video understanding tasks. arxiv:2203.13371[cs.CV] (2022)."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00753"},{"key":"e_1_3_1_20_2","first-page":"1027","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Chen Jin","year":"2021","unstructured":"Jin Chen, Xinxiao Wu, Yao Hu, and Jiebo Luo. 2021. Spatial-temporal causal inference for partial image-to-video adaptation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 1027\u20131035."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2699184"},{"key":"e_1_3_1_22_2","first-page":"6321","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Chen Min-Hung","year":"2019","unstructured":"Min-Hung Chen, Zsolt Kira, Ghassan AlRegib, Jaekwon Yoo, Ruxin Chen, and Jian Zheng. 2019. Temporal attentive alignment for large-scale video domain adaptation. In Proceedings of the IEEE International Conference on Computer Vision. IEEE, 6321\u20136330."},{"key":"e_1_3_1_23_2","first-page":"605","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Chen Min-Hung","year":"2020","unstructured":"Min-Hung Chen, Baopu Li, Yingze Bao, and Ghassan AlRegib. 2020. Action segmentation with mixed temporal domain adaptation. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. IEEE, 605\u2013614."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00947"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV51458.2022.00085"},{"key":"e_1_3_1_26_2","first-page":"5178","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Chen Pengfei","year":"2021","unstructured":"Pengfei Chen, Leida Li, Jinjian Wu, Weisheng Dong, and Guangming Shi. 2021. Unsupervised curriculum domain adaptation for no-reference video quality assessment. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, 5178\u20135187."},{"key":"e_1_3_1_27_2","first-page":"1597","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In Proceedings of the International Conference on Machine Learning. 1597\u20131607."},{"key":"e_1_3_1_28_2","first-page":"352","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201918)","author":"Chen Yunpeng","year":"2018","unstructured":"Yunpeng Chen, Yannis Kalantidis, Jianshu Li, Shuicheng Yan, and Jiashi Feng. 2018. Multi-fiber networks for video recognition. In Proceedings of the European Conference on Computer Vision (ECCV \u201918). 352\u2013367."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00352"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00063"},{"key":"e_1_3_1_31_2","first-page":"3464","volume-title":"Proceedings of the 2022 26th International Conference on Pattern Recognition (ICPR \u201922)","author":"Choi Jinwoo","year":"2022","unstructured":"Jinwoo Choi, Jia-Bin Huang, and Gaurav Sharma. 2022. Self-supervised cross-video temporal learning for unsupervised video domain adaptation. In Proceedings of the 2022 26th International Conference on Pattern Recognition (ICPR \u201922). IEEE, 3464\u20133470."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093511"},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","first-page":"1706","DOI":"10.1109\/WACV45572.2020.9093511","volume-title":"Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV \u201920)","author":"Choi Jinwoo","year":"2020","unstructured":"Jinwoo Choi, Gaurav Sharma, Manmohan Chandraker, and Jia-Bin Huang. 2020. Unsupervised and semi-supervised domain adaptation for action recognition from drones. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV \u201920). IEEE, 1706\u20131715."},{"key":"e_1_3_1_34_2","first-page":"678","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201920)","author":"Choi Jinwoo","year":"2020","unstructured":"Jinwoo Choi, Gaurav Sharma, Samuel Schulter, and Jia-Bin Huang. 2020. Shuffle and Attend: Video domain adaptation. In Proceedings of the European Conference on Computer Vision (ECCV \u201920). 678\u2013695."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.2197\/ipsjjip.28.413"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.350"},{"key":"e_1_3_1_37_2","doi-asserted-by":"crossref","unstructured":"Gabriela Csurka. 2017. Domain adaptation for visual applications: A comprehensive survey. arXiv:1702.05374 (2017).","DOI":"10.1007\/978-3-319-58347-1_1"},{"key":"e_1_3_1_38_2","first-page":"1181","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Costa Victor G. Turrisi da","year":"2022","unstructured":"Victor G. Turrisi da Costa, Giacomo Zara, Paolo Rota, Thiago Oliveira-Santos, Nicu Sebe, Vittorio Murino, and Elisa Ricci. 2022. Dual-head contrastive domain adaptation for video action recognition. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. IEEE, 1181\u20131190."},{"key":"e_1_3_1_39_2","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201918)","author":"Damen Dima","year":"2018","unstructured":"Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. 2018. Scaling egocentric vision: The EPIC-KITCHENS dataset. In Proceedings of the European Conference on Computer Vision (ECCV \u201918). 720\u2013736."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2012.2211477"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2021.10.009"},{"key":"e_1_3_1_42_2","first-page":"1","article-title":"Discriminative unsupervised feature learning with convolutional neural networks","volume":"27","author":"Dosovitskiy Alexey","year":"2014","unstructured":"Alexey Dosovitskiy, Jost Tobias Springenberg, Martin Riedmiller, and Thomas Brox. 2014. Discriminative unsupervised feature learning with convolutional neural networks. Advances in Neural Information Processing Systems 27 (2014), 1\u20139.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_43_2","first-page":"1110","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Du Yong","year":"2015","unstructured":"Yong Du, Wei Wang, and Liang Wang. 2015. Hierarchical recurrent neural network for skeleton based action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1110\u20131118."},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00298"},{"key":"e_1_3_1_45_2","first-page":"768","volume-title":"Proceedings of the 14th European Conference on Computer Vision (ECCV \u201916)","author":"Escorcia Victor","year":"2016","unstructured":"Victor Escorcia, Fabian Caba Heilbron, Juan Carlos Niebles, and Bernard Ghanem. 2016. DAPs: Deep action proposals for action understanding. In Proceedings of the 14th European Conference on Computer Vision (ECCV \u201916). 768\u2013784."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00630"},{"key":"e_1_3_1_47_2","first-page":"1180","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Ganin Yaroslav","year":"2015","unstructured":"Yaroslav Ganin and Victor Lempitsky. 2015. Unsupervised domain adaptation by backpropagation. In Proceedings of the International Conference on Machine Learning. 1180\u20131189."},{"issue":"1","key":"e_1_3_1_48_2","article-title":"Domain-adversarial training of neural networks","volume":"17","author":"Ganin Yaroslav","year":"2016","unstructured":"Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, Fran\u00e7ois Laviolette, Mario Marchand, and Victor Lempitsky. 2016. Domain-adversarial training of neural networks. Journal of Machine Learning Research 17, 1 (2016), 2096\u20132030.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2016.05.094"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00910"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.2985708"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3038372"},{"issue":"3","key":"e_1_3_1_53_2","first-page":"1147","article-title":"Pairwise two-stream ConvNets for cross-domain action recognition with small data","volume":"33","author":"Gao Zan","year":"2020","unstructured":"Zan Gao, Leming Guo, Tongwei Ren, An-An Liu, Zhi-Yong Cheng, and Shengyong Chen. 2020. Pairwise two-stream ConvNets for cross-domain action recognition with small data. IEEE Transactions on Neural Networks and Learning Systems 33, 3 (2020), 1147\u20131161.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_1_54_2","article-title":"A novel multiple-view adversarial learning network for unsupervised domain adaptation action recognition","author":"Gao Zan","year":"2021","unstructured":"Zan Gao, Yibo Zhao, Hua Zhang, Da Chen, An-An Liu, and Shengyong Chen. 2021. A novel multiple-view adversarial learning network for unsupervised domain adaptation action recognition. IEEE Transactions on Cybernetics. Published Online, September 21, 2021.","journal-title":"IEEE Transactions on Cybernetics."},{"key":"e_1_3_1_55_2","first-page":"3290","article-title":"A unified view of label shift estimation","volume":"33","author":"Garg Saurabh","year":"2020","unstructured":"Saurabh Garg, Yifan Wu, Sivaraman Balakrishnan, and Zachary Lipton. 2020. A unified view of label shift estimation. Advances in Neural Information Processing Systems 33 (2020), 3290\u20133300.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_56_2","unstructured":"Chunjiang Ge Rui Huang Mixue Xie Zihang Lai Shiji Song Shuang Li and Gao Huang. 2022. Domain adaptation via prompt learning. arxiv:2202.06687[cs.CV] (2022)."},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_36"},{"key":"e_1_3_1_58_2","first-page":"1","volume-title":"Proceedings of the 6th International Conference on Learning Representations (ICLR \u201918): Conference Track","author":"Gidaris Spyros","year":"2018","unstructured":"Spyros Gidaris, Praveer Singh, and Nikos Komodakis. 2018. Unsupervised representation learning by predicting image rotations. In Proceedings of the 6th International Conference on Learning Representations (ICLR \u201918): Conference Track. 1\u201316."},{"key":"e_1_3_1_59_2","first-page":"505","volume-title":"Proceedings of the 2019 14th IEEE Conference on Industrial Electronics and Applications (ICIEA \u201919)","author":"Gonog Liang","year":"2019","unstructured":"Liang Gonog and Yimin Zhou. 2019. A review: Generative adversarial networks. In Proceedings of the 2019 14th IEEE Conference on Industrial Electronics and Applications (ICIEA \u201919). IEEE, 505\u2013510."},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3422622"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.622"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2017.10.013"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00795"},{"key":"e_1_3_1_64_2","first-page":"484","article-title":"A history of the unity game engine","volume":"483","author":"Haas John K.","year":"2014","unstructured":"John K. Haas. 2014. A history of the unity game engine. Dissertations of the Worcester Polytechnic Institute 483 (2014), 484.","journal-title":"Dissertations of the Worcester Polytechnic Institute"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2017.01.010"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2926463"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548009"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1080\/00207160701396387"},{"key":"e_1_3_1_70_2","first-page":"448","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the International Conference on Machine Learning. 448\u2013456."},{"key":"e_1_3_1_71_2","first-page":"19545","article-title":"Space-time correspondence as a contrastive random walk","volume":"33","author":"Jabri Allan","year":"2020","unstructured":"Allan Jabri, Andrew Owens, and Alexei Efros. 2020. Space-time correspondence as a contrastive random walk. Advances in Neural Information Processing Systems 33 (2020), 19545\u201319560.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_72_2","first-page":"8866","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Jain Samvit","year":"2019","unstructured":"Samvit Jain, Xin Wang, and Joseph E. Gonzalez. 2019. Accel: A corrective fusion network for efficient semantic segmentation on video. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 8866\u20138875."},{"key":"e_1_3_1_73_2","volume-title":"Proceedings of the British Machine Vision Conference (BMVC \u201918)","volume":"2","author":"Jamal Arshad","year":"2018","unstructured":"Arshad Jamal, Vinay P. Namboodiri, Dipti Deodhare, and K. S. Venkatesh. 2018. Deep domain adaptation in action space. In Proceedings of the British Machine Vision Conference (BMVC \u201918), Vol. 2. 5."},{"key":"e_1_3_1_74_2","doi-asserted-by":"crossref","first-page":"2168","DOI":"10.1109\/CVPR.2012.6247924","volume-title":"Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition","author":"Jhuo I.-Hong","year":"2012","unstructured":"I.-Hong Jhuo, Dong Liu, D. T. Lee, and Shih-Fu Chang. 2012. Robust visual domain adaptation with low-rank reconstruction. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 2168\u20132175."},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.59"},{"issue":"9","key":"e_1_3_1_76_2","doi-asserted-by":"crossref","first-page":"3919","DOI":"10.1109\/TNNLS.2020.3016180","article-title":"Effective visual domain adaptation via generative adversarial distribution matching","volume":"32","author":"Kang Qi","year":"2020","unstructured":"Qi Kang, SiYa Yao, MengChu Zhou, Kai Zhang, and Abdullah Abusorrah. 2020. Effective visual domain adaptation via generative adversarial distribution matching. IEEE Transactions on Neural Networks and Learning Systems 32, 9 (2020), 3919\u20133929.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.223"},{"key":"e_1_3_1_78_2","unstructured":"Will Kay Joao Carreira Karen Simonyan Brian Zhang Chloe Hillier Sudheendra Vijayanarasimhan Fabio Viola Tim Green Trevor Back Paul Natsev Mustafa Suleyman and Andrew Zisserman. 2017. The Kinetics human action video dataset. arxiv:1705.06950[cs.CV] (2017)."},{"key":"e_1_3_1_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.486"},{"key":"e_1_3_1_80_2","first-page":"18661","article-title":"Supervised contrastive learning","volume":"33","author":"Khosla Prannay","year":"2020","unstructured":"Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised contrastive learning. Advances in Neural Information Processing Systems 33 (2020), 18661\u201318673.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_81_2","first-page":"13618","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Kim Donghyun","year":"2021","unstructured":"Donghyun Kim, Yi-Hsuan Tsai, Bingbing Zhuang, Xiang Yu, Stan Sclaroff, Kate Saenko, and Manmohan Chandraker. 2021. Learning cross-modal contrastive features for video domain adaptation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, 13618\u201313627."},{"key":"e_1_3_1_82_2","first-page":"5583","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Kim Wonjae","year":"2021","unstructured":"Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. ViLT: Vision-and-Language Transformer without convolution or region supervision. In Proceedings of the International Conference on Machine Learning. 5583\u20135594."},{"key":"e_1_3_1_83_2","doi-asserted-by":"publisher","DOI":"10.1561\/2200000056"},{"key":"e_1_3_1_84_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2945942"},{"key":"e_1_3_1_85_2","doi-asserted-by":"crossref","first-page":"2556","DOI":"10.1109\/ICCV.2011.6126543","volume-title":"Proceedings of the 2011 International Conference on Computer Vision","author":"Kuehne Hildegard","year":"2011","unstructured":"Hildegard Kuehne, Hueihan Jhuang, Est\u00edbaliz Garrote, Tomaso Poggio, and Thomas Serre. 2011. HMDB: A large video database for human motion recognition. In Proceedings of the 2011 International Conference on Computer Vision. IEEE, 2556\u20132563."},{"key":"e_1_3_1_86_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00651"},{"key":"e_1_3_1_87_2","volume-title":"Proceedings of the Workshop on Challenges in Representation Learning (ICML \u201913)","volume":"3","year":"2013","unstructured":"Dong-Hyun Lee. 2013. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Proceedings of the Workshop on Challenges in Representation Learning (ICML \u201913), Vol. 3. 896."},{"key":"e_1_3_1_88_2","first-page":"6816","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Lee Hyogun","year":"2024","unstructured":"Hyogun Lee, Kyungho Bae, Seong Jong Ha, Yumin Ko, Gyeong-Moon Park, and Jinwoo Choi. 2024. GLAD: Global-local view alignment and background debiasing for unsupervised video domain adaptation with large domain gap. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. IEEE, 6816\u20136825."},{"key":"e_1_3_1_89_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00725"},{"key":"e_1_3_1_90_2","first-page":"6205","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Dongxu","year":"2020","unstructured":"Dongxu Li, Xin Yu, Chenchen Xu, Lars Petersson, and Hongdong Li. 2020. Transferring cross-domain knowledge for video sign language recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 6205\u20136214."},{"key":"e_1_3_1_91_2","first-page":"1","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Li Junlong","year":"2022","unstructured":"Junlong Li, Guangyi Chen, Yansong Tang, Jinan Bao, Kun Zhang, Jie Zhou, and Jiwen Lu. 2022. GAIN: On the generalization of instructional action understanding. In Proceedings of the 11th International Conference on Learning Representations. 1\u201322."},{"key":"e_1_3_1_92_2","first-page":"14643","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Kai","year":"2023","unstructured":"Kai Li, Deep Patel, Erik Kruus, and Martin Renqiang Min. 2023. Source-free video domain adaptation with spatial-temporal-historical consistency learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 14643\u201314652."},{"key":"e_1_3_1_93_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2018.03.005"},{"key":"e_1_3_1_94_2","first-page":"1","article-title":"Transferable time-series forecasting under causal conditional shift","volume":"1","author":"Li Zijian","year":"2023","unstructured":"Zijian Li, Ruichu Cai, Tom Z. J. Fu, Zhifeng Hao, and Kun Zhang. 2023. Transferable time-series forecasting under causal conditional shift. IEEE Transactions on Pattern Analysis and Machine Intelligence 1 (2023), 1\u201318.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"5","key":"e_1_3_1_95_2","first-page":"2760","article-title":"TSM: Temporal shift module for efficient and scalable video understanding on edge devices","volume":"44","author":"Lin Ji","year":"2022","unstructured":"Ji Lin, Chuang Gan, Kuan Wang, and Song Han. 2022. TSM: Temporal shift module for efficient and scalable video understanding on edge devices. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 5 (2022), 2760\u20132774.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_96_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00399"},{"key":"e_1_3_1_97_2","first-page":"698","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Lin Wei","year":"2022","unstructured":"Wei Lin, Anna Kukleva, Kunyang Sun, Horst Possegger, Hilde Kuehne, and Horst Bischof. 2022. CycDA: Unsupervised cycle domain adaptation to learn from image to video. In Proceedings of the European Conference on Computer Vision. 698\u2013715."},{"key":"e_1_3_1_98_2","first-page":"3122","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Lipton Zachary","year":"2018","unstructured":"Zachary Lipton, Yu-Xiang Wang, and Alexander Smola. 2018. Detecting and correcting for label shift with black box predictors. In Proceedings of the International Conference on Machine Learning. 3122\u20133130."},{"key":"e_1_3_1_99_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01020"},{"key":"e_1_3_1_100_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2021.107216"},{"key":"e_1_3_1_101_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2018.2823910"},{"key":"e_1_3_1_102_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2957930"},{"key":"e_1_3_1_103_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3147032"},{"key":"e_1_3_1_104_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_1_105_2","first-page":"10534","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lo Shao-Yuan","year":"2023","unstructured":"Shao-Yuan Lo, Poojan Oza, Sumanth Chennupati, Alejandro Galindo, and Vishal M. Patel. 2023. Spatio-temporal pixel-level contrastive learning-based source-free domain adaptation for video semantic segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 10534\u201310543."},{"key":"e_1_3_1_106_2","first-page":"97","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Long Mingsheng","year":"2015","unstructured":"Mingsheng Long, Yue Cao, Jianmin Wang, and Michael Jordan. 2015. Learning transferable features with deep adaptation networks. In Proceedings of the International Conference on Machine Learning. 97\u2013105."},{"key":"e_1_3_1_107_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413897"},{"key":"e_1_3_1_108_2","first-page":"12435","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Melas-Kyriazi Luke","year":"2021","unstructured":"Luke Melas-Kyriazi and Arjun K. Manrai. 2021. PixMatch: Unsupervised domain adaptation via pixelwise consistency training. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 12435\u201312445."},{"key":"e_1_3_1_109_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00272"},{"key":"e_1_3_1_110_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01435"},{"key":"e_1_3_1_111_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2901464"},{"key":"e_1_3_1_112_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01651"},{"key":"e_1_3_1_113_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00020"},{"key":"e_1_3_1_114_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00473"},{"key":"e_1_3_1_115_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-15552-9_29"},{"issue":"6","key":"e_1_3_1_116_2","doi-asserted-by":"crossref","first-page":"2806","DOI":"10.1109\/TPAMI.2020.3045007","article-title":"A review on deep learning techniques for video prediction","volume":"44","author":"Oprea Sergiu","year":"2020","unstructured":"Sergiu Oprea, Pablo Martinez-Gonzalez, Alberto Garcia-Garcia, John Alejandro Castro-Vargas, Sergio Orts-Escolano, Jose Garcia-Rodriguez, and Antonis Argyros. 2020. A review on deep learning techniques for video prediction. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 6 (2020), 2806\u20132826.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_117_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01269"},{"key":"e_1_3_1_118_2","first-page":"11815","volume-title":"Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI \u201920)","author":"Pan Boxiao","year":"2020","unstructured":"Boxiao Pan, Zhangjie Cao, Ehsan Adeli, and Juan Carlos Niebles. 2020. Adversarial cross-domain action recognition with co-attention. In Proceedings of the 34th AAAI Conference on Artificial Intelligence (AAAI \u201920). 11815\u201311822."},{"key":"e_1_3_1_119_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2009.191"},{"key":"e_1_3_1_120_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2014.2347059"},{"key":"e_1_3_1_121_2","unstructured":"Kunyu Peng Di Wen David Schneider Jiaming Zhang Kailun Yang M. Saquib Sarfraz Rainer Stiefelhagen and Alina Roitberg. 2023. FeatFSDA: Towards few-shot domain adaptation for video-based activity recognition. arxiv:2305.08420[cs.CV] (2023)."},{"key":"e_1_3_1_122_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00149"},{"key":"e_1_3_1_123_2","unstructured":"Xingchao Peng Zijun Huang Yizhe Zhu and Kate Saenko. 2019. Federated adversarial domain adaptation. arXiv:1911.02054 (2019)."},{"key":"e_1_3_1_124_2","unstructured":"Xingchao Peng Ben Usman Neela Kaushik Judy Hoffman Dequan Wang and Kate Saenko. 2017. VisDA: The visual domain adaptation challenge. arxiv:1710.06924[cs.CV] (2017)."},{"key":"e_1_3_1_125_2","first-page":"1807","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Planamente Mirco","year":"2022","unstructured":"Mirco Planamente, Chiara Plizzari, Emanuele Alberti, and Barbara Caputo. 2022. Domain generalization through audio-visual relative norm alignment in first person action recognition. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. IEEE, 1807\u20131818."},{"key":"e_1_3_1_126_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2009.11.014"},{"key":"e_1_3_1_127_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00839"},{"key":"e_1_3_1_128_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00078"},{"key":"e_1_3_1_129_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. 8748\u20138763."},{"key":"e_1_3_1_130_2","doi-asserted-by":"crossref","unstructured":"Alan Ramponi and Barbara Plank. 2020. Neural unsupervised domain adaptation in NLP\u2014A survey. In Proceedings of the 28th International Conference on Computational Linguistics. 6838\u20136855.","DOI":"10.18653\/v1\/2020.coling-main.603"},{"key":"e_1_3_1_131_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00138-012-0450-4"},{"key":"e_1_3_1_132_2","doi-asserted-by":"crossref","unstructured":"Ievgen Redko Amaury Habrard and Marc Sebban. 2017. Theoretical analysis of domain adaptation with optimal transport. In Machine Learning and Knowledge Discovery in Databases. Lecture Notes in Computer Science Vol. 10535. Springer 737\u2013753.","DOI":"10.1007\/978-3-319-71246-8_45"},{"key":"e_1_3_1_133_2","volume-title":"Advances in Domain Adaptation Theory","author":"Redko Ievgen","year":"2019","unstructured":"Ievgen Redko, Emilie Morvant, Amaury Habrard, Marc Sebban, and Younes Bennani. 2019. Advances in Domain Adaptation Theory. Elsevier, Cham."},{"key":"e_1_3_1_134_2","doi-asserted-by":"publisher","DOI":"10.1145\/3472291"},{"key":"e_1_3_1_135_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.243"},{"issue":"7","key":"e_1_3_1_136_2","first-page":"1369","article-title":"Generalization error bounds in semi-supervised classification under the cluster assumption.","volume":"8","author":"Rigollet Philippe","year":"2007","unstructured":"Philippe Rigollet. 2007. Generalization error bounds in semi-supervised classification under the cluster assumption. Journal of Machine Learning Research 8, 7 (2007), 1369\u20131392.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_137_2","doi-asserted-by":"crossref","first-page":"1866","DOI":"10.1109\/WACV.2019.00203","volume-title":"Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV \u201919)","author":"Romijnders Rob","year":"2019","unstructured":"Rob Romijnders, Panagiotis Meletis, and Gijs Dubbelman. 2019. A domain agnostic normalization layer for unsupervised adversarial domain adaptation. In Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV \u201919). IEEE, 1866\u20131875."},{"key":"e_1_3_1_138_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.352"},{"key":"e_1_3_1_139_2","first-page":"23386","article-title":"Contrast and Mix: Temporal contrastive video domain adaptation with background mixing","volume":"34","author":"Sahoo Aadarsh","year":"2021","unstructured":"Aadarsh Sahoo, Rutav Shah, Rameswar Panda, Kate Saenko, and Abir Das. 2021. Contrast and Mix: Temporal contrastive video domain adaptation with background mixing. Advances in Neural Information Processing Systems 34 (2021), 23386\u201323400.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_140_2","doi-asserted-by":"publisher","DOI":"10.1145\/3446374"},{"key":"e_1_3_1_141_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2008.2005605"},{"key":"e_1_3_1_142_2","doi-asserted-by":"publisher","DOI":"10.1162\/089976698300017467"},{"key":"e_1_3_1_143_2","unstructured":"Inkyu Shin Kwanyong Park Sanghyun Woo and In So Kweon. 2021. Unsupervised domain adaptation for video semantic segmentation. arxiv:2107.11052[cs.CV] (2021)."},{"key":"e_1_3_1_144_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00317"},{"key":"e_1_3_1_145_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01519"},{"key":"e_1_3_1_146_2","unstructured":"Lucas Smaira Jo\u00e3o Carreira Eric Noland Ellen Clancy Amy Wu and Andrew Zisserman. 2020. A short note on the Kinetics-700-2020 human action dataset. arXiv:2010.10864 (2020)."},{"key":"e_1_3_1_147_2","first-page":"4263","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"31","author":"Song Sijie","year":"2017","unstructured":"Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng, and Jiaying Liu. 2017. An end-to-end spatio-temporal attention model for human action recognition from skeleton data. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 31. 4263\u20134270."},{"key":"e_1_3_1_148_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00966"},{"key":"e_1_3_1_149_2","unstructured":"Khurram Soomro Amir Roshan Zamir and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv:1212.0402 (2012)."},{"key":"e_1_3_1_150_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093390"},{"issue":"5","key":"e_1_3_1_151_2","first-page":"985","article-title":"Covariate shift adaptation by importance weighted cross validation.","volume":"8","author":"Sugiyama Masashi","year":"2007","unstructured":"Masashi Sugiyama, Matthias Krauledat, and Klaus-Robert M\u00fcller. 2007. Covariate shift adaptation by importance weighted cross validation. Journal of Machine Learning Research 8, 5 (2007), 985\u20131005.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_152_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01708"},{"key":"e_1_3_1_153_2","first-page":"764","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Sultani Waqas","year":"2014","unstructured":"Waqas Sultani and Imran Saleemi. 2014. Human action recognition across datasets by foreground-weighted histogram decomposition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 764\u2013771."},{"key":"e_1_3_1_154_2","first-page":"2058","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"30","author":"Sun Baochen","year":"2016","unstructured":"Baochen Sun, Jiashi Feng, and Kate Saenko. 2016. Return of frustratingly easy domain adaptation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 30. 2058\u20132065."},{"key":"e_1_3_1_155_2","first-page":"9229","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Sun Yu","year":"2020","unstructured":"Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei Efros, and Moritz Hardt. 2020. Test-time training with self-supervision for generalization under distribution shifts. In Proceedings of the International Conference on Machine Learning. 9229\u20139248."},{"key":"e_1_3_1_156_2","first-page":"19276","article-title":"Domain adaptation with conditional distribution matching and generalized label shift","volume":"33","author":"Combes Remi Tachet des","year":"2020","unstructured":"Remi Tachet des Combes, Han Zhao, Yu-Xiang Wang, and Geoffrey J. Gordon. 2020. Domain adaptation with conditional distribution matching and generalized label shift. Advances in Neural Information Processing Systems 33 (2020), 19276\u201319289.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_157_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01424-7_27"},{"key":"e_1_3_1_158_2","first-page":"1415","article-title":"Feature extraction by non-parametric mutual information maximization","author":"Torkkola Kari","year":"2003","unstructured":"Kari Torkkola. 2003. Feature extraction by non-parametric mutual information maximization. Journal of Machine Learning Research 3 (March 2003), 1415\u20131438.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_159_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_1_160_2","first-page":"1","volume-title":"Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition","author":"Turaga Pavan","year":"2008","unstructured":"Pavan Turaga, Ashok Veeraraghavan, and Rama Chellappa. 2008. Statistical analysis on Stiefel and Grassmann manifolds with applications in computer vision. In Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1\u20138."},{"key":"e_1_3_1_161_2","first-page":"1","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017), 1\u201311.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_162_2","first-page":"1","article-title":"Generalizing to unseen domains via adversarial data augmentation","volume":"31","author":"Volpi Riccardo","year":"2018","unstructured":"Riccardo Volpi, Hongseok Namkoong, Ozan Sener, John C. Duchi, Vittorio Murino, and Silvio Savarese. 2018. Generalizing to unseen domains via adversarial data augmentation. Advances in Neural Information Processing Systems 31 (2018), 1\u201311.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_163_2","unstructured":"Dequan Wang Evan Shelhamer Shaoteng Liu Bruno Olshausen and Trevor Darrell. 2020. Tent: Fully test-time adaptation by entropy minimization. arxiv:2006.10726[cs.CV] (2020)."},{"key":"e_1_3_1_164_2","doi-asserted-by":"publisher","DOI":"10.1109\/JAS.2017.7510583"},{"key":"e_1_3_1_165_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"e_1_3_1_166_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2018.05.083"},{"key":"e_1_3_1_167_2","unstructured":"Mengmeng Wang Jiazheng Xing and Yong Liu. 2021. ActionCLIP: A new paradigm for video action recognition. arxiv:2109.08472[cs.CV] (2021)."},{"issue":"3","key":"e_1_3_1_168_2","first-page":"3933","article-title":"Weakly-supervised video object grounding via causal intervention","volume":"45","author":"Wang Wei","year":"2022","unstructured":"Wei Wang, Junyu Gao, and Changsheng Xu. 2022. Weakly-supervised video object grounding via causal intervention. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 3 (2022), 3933\u20133948.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_169_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01271"},{"key":"e_1_3_1_170_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00813"},{"key":"e_1_3_1_171_2","doi-asserted-by":"crossref","unstructured":"Xiyu Wang Yuecong Xu Kezhi Mao and Jianfei Yang. 2022. Calibrating class weights with multi-modal information for partial video domain adaptation. In Proceedings of the 30th ACM International Conference on Multimedia (MM \u201922). 3945\u20133954.","DOI":"10.1145\/3503161.3548095"},{"key":"e_1_3_1_172_2","first-page":"8198","volume-title":"Proceedings of the 2021 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP \u201921)","author":"Wang Yatian","year":"2021","unstructured":"Yatian Wang, Xiaolin Song, Yezhen Wang, Pengfei Xu, Runbo Hu, and Hua Chai. 2021. Dual metric discriminator for open set video domain adaptation. In Proceedings of the 2021 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP \u201921). IEEE, 8198\u20138202."},{"key":"e_1_3_1_173_2","unstructured":"Pengfei Wei Lingdong Kong Xinghua Qu Xiang Yin Zhiqiang Xu Jing Jiang and Zejun Ma. 2022. Unsupervised video domain adaptation: A disentanglement perspective. arXiv:2208.07365 (2022)."},{"key":"e_1_3_1_174_2","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-016-0043-6"},{"key":"e_1_3_1_175_2","doi-asserted-by":"publisher","DOI":"10.1145\/3400066"},{"key":"e_1_3_1_176_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.309"},{"key":"e_1_3_1_177_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2021.11.054"},{"key":"e_1_3_1_178_2","first-page":"11340","article-title":"Online adaptation to label distribution shift","volume":"34","author":"Wu Ruihan","year":"2021","unstructured":"Ruihan Wu, Chuan Guo, Yi Su, and Kilian Q. Weinberger. 2021. Online adaptation to label distribution shift. Advances in Neural Information Processing Systems 34 (2021), 11340\u201311351.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_179_2","first-page":"540","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Wu Yuan","year":"2020","unstructured":"Yuan Wu, Diana Inkpen, and Ahmed El-Roby. 2020. Dual mixup regularized learning for adversarial domain adaptation. In Proceedings of the European Conference on Computer Vision. 540\u2013555."},{"key":"e_1_3_1_180_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2018.12.050"},{"key":"e_1_3_1_181_2","first-page":"6256","article-title":"Unsupervised data augmentation for consistency training","volume":"33","author":"Xie Qizhe","year":"2020","unstructured":"Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le. 2020. Unsupervised data augmentation for consistency training. Advances in Neural Information Processing Systems 33 (2020), 6256\u20136268.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_182_2","unstructured":"Saining Xie Chen Sun Jonathan Huang Zhuowen Tu and Kevin Murphy. 2017. Rethinking spatiotemporal feature learning for video understanding. arXiv:1712.04851 (2017)."},{"key":"e_1_3_1_183_2","doi-asserted-by":"crossref","unstructured":"Yun Xing Dayan Guan Jiaxing Huang and Shijian Lu. 2022. Domain adaptive video segmentation via temporal pseudo supervision. arXiv:2207.02372 (2022).","DOI":"10.1007\/978-3-031-20056-4_36"},{"key":"e_1_3_1_184_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.6123"},{"key":"e_1_3_1_185_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2005.845141"},{"key":"e_1_3_1_186_2","first-page":"9332","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Xu Yuecong","year":"2021","unstructured":"Yuecong Xu, Jianfei Yang, Haozhi Cao, Zhenghua Chen, Qi Li, and Kezhi Mao. 2021. Partial video domain adaptation with partial adversarial temporal attentive network. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, 9332\u20139341."},{"key":"e_1_3_1_187_2","unstructured":"Yuecong Xu Jianfei Yang Haozhi Cao Kezhi Mao Jianxiong Yin and Simon See. 2021. Aligning correlation information for domain adaptation in action recognition. arxiv:2107.04932[cs.CV] (2021)."},{"key":"e_1_3_1_188_2","first-page":"70","volume-title":"Proceedings of the International Workshop on Deep Learning for Human Activity Recognition","author":"Xu Yuecong","year":"2021","unstructured":"Yuecong Xu, Jianfei Yang, Haozhi Cao, Kezhi Mao, Jianxiong Yin, and Simon See. 2021. ARID: A new dataset for recognizing action in the dark. In Proceedings of the International Workshop on Deep Learning for Human Activity Recognition. 70\u201384."},{"key":"e_1_3_1_189_2","unstructured":"Yuecong Xu Jianfei Yang Haozhi Cao Keyu Wu Wu Min and Zhenghua Chen. 2022. Learning temporal consistency for source-free video domain adaptation. In Proceedings of the European Conference on Computer Vision."},{"key":"e_1_3_1_190_2","unstructured":"Yuecong Xu Jianfei Yang Haozhi Cao Keyu Wu Min Wu Rui Zhao and Zhenghua Chen. 2021. Multi-source video domain adaptation with temporal attentive moment alignment. arXiv:2109.09964 (2021)."},{"key":"e_1_3_1_191_2","unstructured":"Yuecong Xu Jianfei Yang Haozhi Cao Min Wu Xiaoli Li Lihua Xie and Zhenghua Chen. 2022. Leveraging endo- and exo-temporal regularization for black-box video domain adaptation. arXiv:2208.05187 (2022)."},{"key":"e_1_3_1_192_2","unstructured":"Yuecong Xu Jianfei Yang Yunjiao Zhou Zhenghua Chen Min Wu and Xiaoli Li. 2023. Augmenting and aligning snippets for few-shot video domain adaptation. arxiv:2303.10451[cs.CV] (2023)."},{"key":"e_1_3_1_193_2","unstructured":"Shen Yan Huan Song Nanxiang Li Lincan Zou and Liu Ren. 2020. Improve unsupervised domain adaptation with mixup training. arXiv:2001.00677 (2020)."},{"key":"e_1_3_1_194_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3001522"},{"key":"e_1_3_1_195_2","first-page":"480","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Yang Jinyu","year":"2020","unstructured":"Jinyu Yang, Weizhi An, Sheng Wang, Xinliang Zhu, Chaochao Yan, and Junzhou Huang. 2020. Label-driven reconstruction for domain adaptation in semantic segmentation. In Proceedings of the European Conference on Computer Vision. 480\u2013498."},{"key":"e_1_3_1_196_2","unstructured":"Jianfei Yang Xiangyu Peng Kai Wang Zheng Zhu Jiashi Feng Lihua Xie and Yang You. 2022. Divide to adapt: Mitigating confirmation bias for domain adaptation of black-box predictors. arXiv:2205.14467 (2022)."},{"key":"e_1_3_1_197_2","article-title":"Advancing imbalanced domain adaptation: Cluster-level discrepancy minimization with a comprehensive benchmark","author":"Yang Jianfei","year":"2021","unstructured":"Jianfei Yang, Jiangang Yang, Shizheng Wang, Shuxin Cao, Han Zou, and Lihua Xie. 2021. Advancing imbalanced domain adaptation: Cluster-level discrepancy minimization with a comprehensive benchmark. IEEE Transactions on Cybernetics. Published Online, August 16, 2021.","journal-title":"IEEE Transactions on Cybernetics."},{"key":"e_1_3_1_198_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2020.12.046"},{"key":"e_1_3_1_199_2","first-page":"589","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Yang Jianfei","year":"2020","unstructured":"Jianfei Yang, Han Zou, Yuxun Zhou, Zhaoyang Zeng, and Lihua Xie. 2020. Mind the discriminability: Asymmetric adversarial domain adaptation. In Proceedings of the European Conference on Computer Vision. 589\u2013606."},{"key":"e_1_3_1_200_2","first-page":"14722","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yang Lijin","year":"2022","unstructured":"Lijin Yang, Yifei Huang, Yusuke Sugano, and Yoichi Sato. 2022. Interact before align: Leveraging cross-modal knowledge for domain adaptive action recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 14722\u201314732."},{"key":"e_1_3_1_201_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-01585-4"},{"key":"e_1_3_1_202_2","doi-asserted-by":"publisher","DOI":"10.1145\/3391743"},{"key":"e_1_3_1_203_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3116945"},{"key":"e_1_3_1_204_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.109023"},{"key":"e_1_3_1_205_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00848"},{"key":"e_1_3_1_206_2","first-page":"10307","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zara Giacomo","year":"2023","unstructured":"Giacomo Zara, Alessandro Conti, Subhankar Roy, St\u00e9phane Lathuili\u00e8re, Paolo Rota, and Elisa Ricci. 2023. The unreasonable effectiveness of large language-vision models for source-free video domain adaptation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. IEEE, 10307\u201310317."},{"key":"e_1_3_1_207_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01107"},{"key":"e_1_3_1_208_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612153"},{"key":"e_1_3_1_209_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00156"},{"key":"e_1_3_1_210_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330961"},{"key":"e_1_3_1_211_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Zhang Hongyi","year":"2018","unstructured":"Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. 2018. mixup: Beyond empirical risk minimization. In Proceedings of the International Conference on Learning Representations. 1\u201313. https:\/\/openreview.net\/forum?id=r1Ddp1-Rb"},{"key":"e_1_3_1_212_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00851"},{"issue":"4","key":"e_1_3_1_213_2","doi-asserted-by":"crossref","first-page":"960","DOI":"10.1109\/TCYB.2016.2535122","article-title":"Semi-supervised image-to-video adaptation for video action recognition","volume":"47","author":"Zhang Jianguang","year":"2016","unstructured":"Jianguang Zhang, Yahong Han, Jinhui Tang, Qinghua Hu, and Jianmin Jiang. 2016. Semi-supervised image-to-video adaptation for video action recognition. IEEE Transactions on Cybernetics 47, 4 (2016), 960\u2013973.","journal-title":"IEEE Transactions on Cybernetics"},{"key":"e_1_3_1_214_2","first-page":"3151","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"29","author":"Zhang Kun","year":"2015","unstructured":"Kun Zhang, Mingming Gong, and Bernhard Sch\u00f6lkopf. 2015. Multi-source domain adaptation: A causal view. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 29. 3151\u20133157."},{"key":"e_1_3_1_215_2","first-page":"819","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Zhang Kun","year":"2013","unstructured":"Kun Zhang, Bernhard Sch\u00f6lkopf, Krikamol Muandet, and Zhikun Wang. 2013. Domain adaptation under target and conditional shift. In Proceedings of the International Conference on Machine Learning. 819\u2013827."},{"key":"e_1_3_1_216_2","unstructured":"Xin Zhang Shixiang Shane Gu Yutaka Matsuo and Yusuke Iwasawa. 2021. Domain prompt learning for efficiently adapting CLIP to unseen domains. arxiv:2111.12853[cs.CV] (2021)."},{"issue":"5","key":"e_1_3_1_217_2","doi-asserted-by":"crossref","first-page":"2775","DOI":"10.1109\/TPAMI.2020.3036956","article-title":"Unsupervised multi-class domain adaptation: Theory, algorithms, and practice","volume":"44","author":"Zhang Yabin","year":"2020","unstructured":"Yabin Zhang, Bin Deng, Hui Tang, Lei Zhang, and Kui Jia. 2020. Unsupervised multi-class domain adaptation: Theory, algorithms, and practice. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 5 (2020), 2775\u20132792.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_218_2","first-page":"13791","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Yunhua","year":"2022","unstructured":"Yunhua Zhang, Hazel Doughty, Ling Shao, and Cees G. M. Snoek. 2022. Audio-adaptive activity recognition across video domains. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 13791\u201313800."},{"key":"e_1_3_1_219_2","first-page":"7404","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Zhang Yuchen","year":"2019","unstructured":"Yuchen Zhang, Tianle Liu, Mingsheng Long, and Michael Jordan. 2019. Bridging theory and algorithm for domain adaptation. In Proceedings of the International Conference on Machine Learning. 7404\u20137413."},{"key":"e_1_3_1_220_2","doi-asserted-by":"crossref","first-page":"1166","DOI":"10.1109\/ICPR48806.2021.9413014","volume-title":"Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR \u201921)","author":"Zhong Tao","year":"2021","unstructured":"Tao Zhong, Wonjik Kim, Masayuki Tanaka, and Masatoshi Okutomi. 2021. Human segmentation with dynamic LiDAR data. In Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR \u201921). IEEE, 1166\u20131172."},{"key":"e_1_3_1_221_2","first-page":"803","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201918)","author":"Zhou Bolei","year":"2018","unstructured":"Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba. 2018. Temporal relational reasoning in videos. In Proceedings of the European Conference on Computer Vision (ECCV \u201918). 803\u2013818."},{"key":"e_1_3_1_222_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2020.3004555"},{"key":"e_1_3_1_223_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00365"},{"key":"e_1_3_1_224_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611897"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3679010","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3679010","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:58:14Z","timestamp":1750294694000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3679010"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10]]},"references-count":223,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2024,12,31]]}},"alternative-id":["10.1145\/3679010"],"URL":"https:\/\/doi.org\/10.1145\/3679010","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10]]},"assertion":[{"value":"2022-11-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-07-07","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-01","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}