{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T13:21:59Z","timestamp":1778592119828,"version":"3.51.4"},"reference-count":89,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2025,3,7]],"date-time":"2025-03-07T00:00:00Z","timestamp":1741305600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62222207, 62072245, 62273264, and 62302208"],"award-info":[{"award-number":["62222207, 62072245, 62273264, and 62302208"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100004608","name":"Natural Science Foundation of Jiangsu Province","doi-asserted-by":"crossref","award":["BK20211520"],"award-info":[{"award-number":["BK20211520"]}],"id":[{"id":"10.13039\/501100004608","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"crossref","award":["2023TQ0151, and 2023M731596"],"award-info":[{"award-number":["2023TQ0151, and 2023M731596"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,4,30]]},"abstract":"<jats:p>\n            Semi-supervised action recognition is a challenging yet prospective task due to its low reliance on costly labeled videos. One high-profile solution is to explore frame-level weak\/strong augmentations for learning abundant representations, inspired by the FixMatch framework dominating the semi-supervised image classification task. However, such a solution mainly brings perturbations in terms of texture and scale, leading to the limitation in learning action representations in videos with spatiotemporal redundancy and complexity. Therefore, we revisit the creative trick of weak\/strong augmentations in FixMatch and then propose, to the best of our knowledge, a novel Frame- and Feature-level augmentation FixMatch (dubbed as F\n            <jats:sup>2<\/jats:sup>\n            -FixMatch) framework to learn more abundant action representations for being robust to complex and dynamic video scenarios. Specifically, we design a new Progressive Augmentation mechanism that implements the weak\/strong augmentations first at the frame level, and further implements the perturbation at the feature level, to obtain abundant four types of augmented features in broader perturbation spaces. Moreover, we present an evolved Multihead Pseudo-Labeling scheme to promote the consistency of features across different augmented versions based on the pseudo labels. We conduct extensive experiments on several public datasets to demonstrate that our F\n            <jats:sup>2<\/jats:sup>\n            -FixMatch achieves the performance gain compared with current state-of-the-art methods. The source codes of F\n            <jats:sup>2<\/jats:sup>\n            -FixMatch are publicly available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/zwtu\/F2FixMatch\">https:\/\/github.com\/zwtu\/F2FixMatch<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3655025","type":"journal-article","created":{"date-parts":[[2024,4,11]],"date-time":"2024-04-11T11:07:20Z","timestamp":1712833640000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Leveraging Frame- and Feature-level Progressive Augmentation for Semi-supervised Action Recognition"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-3246-4848","authenticated-orcid":false,"given":"Zhewei","family":"Tu","sequence":"first","affiliation":[{"name":"Nanjing University of Science and Technology, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4902-4663","authenticated-orcid":false,"given":"Xiangbo","family":"Shu","sequence":"additional","affiliation":[{"name":"Nanjing University of Science and Technology, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-4102-6506","authenticated-orcid":false,"given":"Peng","family":"Huang","sequence":"additional","affiliation":[{"name":"Nanjing University of Science and Technology, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0694-9458","authenticated-orcid":false,"given":"Rui","family":"Yan","sequence":"additional","affiliation":[{"name":"Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9944-7464","authenticated-orcid":false,"given":"Zhenxing","family":"Liu","sequence":"additional","affiliation":[{"name":"Wuhan University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-2637-5095","authenticated-orcid":false,"given":"Jiachao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Nanjing Institute of Technology, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,3,7]]},"reference":[{"issue":"11","key":"e_1_3_2_2_2","doi-asserted-by":"crossref","first-page":"6679","DOI":"10.1109\/TCSVT.2023.3267178","article-title":"Actor-aware self-supervised learning for semi-supervised video representation learning","volume":"33","author":"Assefa Maregu","year":"2023","unstructured":"Maregu Assefa, Wei Jiang, Kumie Gedamu, Getinet Yilma, Deepak Adhikari, Melese Ayalew, Abegaz Mohammed, and Aiman Erbad. 2023. Actor-aware self-supervised learning for semi-supervised video representation learning. IEEE Trans. Circ. Syst. Vid. Technol. 33, 11 (2023), 6679\u20136692.","journal-title":"IEEE Trans. Circ. Syst. Vid. Technol."},{"key":"e_1_3_2_3_2","first-page":"1","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Bachman Philip","year":"2014","unstructured":"Philip Bachman, Ouais Alsharif, and Doina Precup. 2014. Learning with pseudo-ensembles. In Advances in Neural Information Processing Systems (NeurIPS). 1\u20139."},{"key":"e_1_3_2_4_2","first-page":"813","volume-title":"Proceedings of the International Conference on Machine Learning (ICML\u201921)","author":"Bertasius Gedas","year":"2021","unstructured":"Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding? In Proceedings of the International Conference on Machine Learning (ICML\u201921). 813\u2013824."},{"key":"e_1_3_2_5_2","article-title":"Remixmatch: Semi-supervised learning with distribution alignment and augmentation anchoring","author":"Berthelot David","year":"2019","unstructured":"David Berthelot, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin, Kihyuk Sohn, Han Zhang, and Colin Raffel. 2019. Remixmatch: Semi-supervised learning with distribution alignment and augmentation anchoring. arXiv:1911.09785. Retrieved from https:\/\/arxiv.org\/abs\/1911.09785","journal-title":"arXiv:1911.09785"},{"key":"e_1_3_2_6_2","first-page":"1","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Berthelot David","year":"2019","unstructured":"David Berthelot, Nicholas Carlini, Ian Goodfellow, Nicolas Papernot, Avital Oliver, and Colin A. Raffel. 2019. Mixmatch: A holistic approach to semi-supervised learning. In Advances in Neural Information Processing Systems (NeurIPS). 1\u201311."},{"key":"e_1_3_2_7_2","first-page":"961","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915)","author":"Heilbron Fabian Caba","year":"2015","unstructured":"Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201915). 961\u2013970."},{"key":"e_1_3_2_8_2","first-page":"6299","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917)","author":"Carreira Joao","year":"2017","unstructured":"Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? A new model and the kinetics dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917). 6299\u20136308."},{"key":"e_1_3_2_9_2","first-page":"1932","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201909)","author":"Chaudhry Rizwan","year":"2009","unstructured":"Rizwan Chaudhry, Avinash Ravichandran, Gregory Hager, and Ren\u00e9 Vidal. 2009. Histograms of oriented optical flow and binet-cauchy kernels on nonlinear dynamical systems for the recognition of human actions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201909). 1932\u20131939."},{"key":"e_1_3_2_10_2","first-page":"702","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920) Workshops","author":"Cubuk Ekin D.","year":"2020","unstructured":"Ekin D. Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V. Le. 2020. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920) Workshops. 702\u2013703."},{"key":"e_1_3_2_11_2","first-page":"2341","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201923)","author":"Dave Ishan Rajendrakumar","year":"2023","unstructured":"Ishan Rajendrakumar Dave, Mamshad Nayeem Rizve, Chen Chen, and Mubarak Shah. 2023. TimeBalance: Temporally-invariant and temporally-distinctive video representations for semi-supervised action recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201923). 2341\u20132352."},{"key":"e_1_3_2_12_2","first-page":"248","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201909)","author":"Deng Jia","year":"2009","unstructured":"Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201909). 248\u2013255."},{"key":"e_1_3_2_13_2","article-title":"Improved regularization of convolutional neural networks with cutout","author":"DeVries Terrance","year":"2017","unstructured":"Terrance DeVries and Graham W. Taylor. 2017. Improved regularization of convolutional neural networks with cutout. arXiv:1708.04552. Retrieved from https:\/\/arxiv.org\/abs\/1708.04552","journal-title":"arXiv:1708.04552"},{"key":"e_1_3_2_14_2","first-page":"203","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Feichtenhofer Christoph","year":"2020","unstructured":"Christoph Feichtenhofer. 2020. X3d: Expanding architectures for efficient video recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920). 203\u2013213."},{"key":"e_1_3_2_15_2","first-page":"6202","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201919)","author":"Feichtenhofer Christoph","year":"2019","unstructured":"Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. 2019. Slowfast networks for video recognition. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201919). 6202\u20136211."},{"key":"e_1_3_2_16_2","first-page":"1933","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916)","author":"Feichtenhofer Christoph","year":"2016","unstructured":"Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman. 2016. Convolutional two-stream network fusion for video action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916). 1933\u20131941."},{"key":"e_1_3_2_17_2","first-page":"242","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201922)","author":"Gowda Shreyank N.","year":"2022","unstructured":"Shreyank N. Gowda, Marcus Rohrbach, Frank Keller, and Laura Sevilla-Lara. 2022. Learn2Augment: Learning to composite videos for data augmentation in action recognition. In Proceedings of the European Conference on Computer Vision (ECCV\u201922). 242\u2013259."},{"key":"e_1_3_2_18_2","article-title":"Accurate, large minibatch sgd: Training imagenet in 1 hour","author":"Goyal Priya","year":"2017","unstructured":"Priya Goyal, Piotr Doll\u00e1r, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017. Accurate, large minibatch sgd: Training imagenet in 1 hour. arXiv:1706.02677. Retrieved from https:\/\/arxiv.org\/abs\/1706.02677","journal-title":"arXiv:1706.02677"},{"key":"e_1_3_2_19_2","first-page":"1","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Grandvalet Yves","year":"2004","unstructured":"Yves Grandvalet and Yoshua Bengio. 2004. Semi-supervised learning by entropy minimization. In Advances in Neural Information Processing Systems (NeurIPS). 1\u20138."},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3358415"},{"key":"e_1_3_2_21_2","first-page":"1823","volume-title":"Proceedings of the ACM International Conference on Multimedia (ACM MM\u201919)","author":"Guo Dan","year":"2019","unstructured":"Dan Guo, Kun Li, Zheng-Jun Zha, and Meng Wang. 2019. Dadnet: Dilated-attention-deformable convnet for crowd counting. In Proceedings of the ACM International Conference on Multimedia (ACM MM\u201919). 1823\u20131832."},{"key":"e_1_3_2_22_2","first-page":"744","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI\u201919)","author":"Guo Dan","year":"2019","unstructured":"Dan Guo, Shuo Wang, Qi Tian, and Meng Wang. 2019. Dense temporal convolution network for sign language translation. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI\u201919). 744\u2013750."},{"key":"e_1_3_2_23_2","first-page":"6845","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201918)","author":"Guo Dan","year":"2018","unstructured":"Dan Guo, Wengang Zhou, Houqiang Li, and Meng Wang. 2018. Hierarchical LSTM for sign language translation. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201918). 6845\u20136852."},{"issue":"10","key":"e_1_3_2_24_2","doi-asserted-by":"crossref","first-page":"7120","DOI":"10.1109\/TCSVT.2022.3169842","article-title":"Attention in attention: Modeling context correlation for efficient video classification","volume":"32","author":"Hao Yanbin","year":"2022","unstructured":"Yanbin Hao, Shuo Wang, Pei Cao, Xinjian Gao, Tong Xu, Jinmeng Wu, and Xiangnan He. 2022. Attention in attention: Modeling context correlation for efficient video classification. IEEE Trans. Circ. Syst. Vid. Technol. 32, 10 (2022), 7120\u20137132.","journal-title":"IEEE Trans. Circ. Syst. Vid. Technol."},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","first-page":"7279","DOI":"10.1109\/TIP.2022.3221292","article-title":"Spatio-temporal collaborative module for efficient action recognition","volume":"31","author":"Hao Yanbin","year":"2022","unstructured":"Yanbin Hao, Shuo Wang, Yi Tan, Xiangnan He, Zhenguang Liu, and Meng Wang. 2022. Spatio-temporal collaborative module for efficient action recognition. IEEE Trans. Image Process. 31 (2022), 7279\u20137291.","journal-title":"IEEE Trans. Image Process."},{"key":"e_1_3_2_26_2","first-page":"6546","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918)","author":"Hara Kensho","year":"2018","unstructured":"Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. 2018. Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet? In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918). 6546\u20136555."},{"key":"e_1_3_2_27_2","first-page":"770","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916)","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916). 770\u2013778."},{"key":"e_1_3_2_28_2","first-page":"448","volume-title":"Proceedings of the International Conference on Machine Learning (ICML\u201915)","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the International Conference on Machine Learning (ICML\u201915). 448\u2013456."},{"key":"e_1_3_2_29_2","first-page":"1110","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV\u201921)","author":"Jing Longlong","year":"2021","unstructured":"Longlong Jing, Toufiq Parag, Zhe Wu, Yingli Tian, and Hongcheng Wang. 2021. Videossl: Semi-supervised learning for video classification. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV\u201921). 1110\u20131119."},{"key":"e_1_3_2_30_2","article-title":"The kinetics human action video dataset","author":"Kay Will","year":"2017","unstructured":"Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et\u00a0al. 2017. The kinetics human action video dataset. arXiv:1705.06950. Retrieved from https:\/\/arxiv.org\/abs\/1705.06950","journal-title":"arXiv:1705.06950"},{"issue":"1","key":"e_1_3_2_31_2","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1109\/TPAMI.2015.2430335","article-title":"Anticipating human activities using object affordances for reactive robotic response","volume":"38","author":"Koppula Hema S.","year":"2015","unstructured":"Hema S. Koppula and Ashutosh Saxena. 2015. Anticipating human activities using object affordances for reactive robotic response. IEEE Trans. Pattern Anal. Mach. Intell. 38, 1 (2015), 14\u201329.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_32_2","first-page":"1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201919) Workshops","author":"Kopuklu Okan","year":"2019","unstructured":"Okan Kopuklu, Neslihan Kose, Ahmet Gunduz, and Gerhard Rigoll. 2019. Resource efficient 3d convolutional neural networks. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201919) Workshops. 1\u201310."},{"key":"e_1_3_2_33_2","first-page":"2556","volume-title":"Proceedings of the IEEE Conference on International Conference on Computer Vision (ICCV\u201911)","author":"Kuehne Hildegard","year":"2011","unstructured":"Hildegard Kuehne, Hueihan Jhuang, Est\u00edbaliz Garrote, Tomaso Poggio, and Thomas Serre. 2011. HMDB: A large video database for human motion recognition. In Proceedings of the IEEE Conference on International Conference on Computer Vision (ICCV\u201911). 2556\u20132563."},{"key":"e_1_3_2_34_2","first-page":"479","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201920)","author":"Kuo Chia-Wen","year":"2020","unstructured":"Chia-Wen Kuo, Chih-Yao Ma, Jia-Bin Huang, and Zsolt Kira. 2020. Featmatch: Feature-based augmentation for semi-supervised learning. In Proceedings of the European Conference on Computer Vision (ECCV\u201920). 479\u2013495."},{"key":"e_1_3_2_35_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201917)","author":"Laine Samuli","year":"2017","unstructured":"Samuli Laine and Timo Aila. 2017. Temporal ensembling for semi-supervised learning. In Proceedings of the International Conference on Learning Representations (ICLR\u201917). 1\u201313."},{"key":"e_1_3_2_36_2","first-page":"1","volume-title":"Proceedings of the International Conference on Machine Learning (ICML\u201913) Workshop","author":"Lee Dong-Hyun","year":"2013","unstructured":"Dong-Hyun Lee et\u00a0al. 2013. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Proceedings of the International Conference on Machine Learning (ICML\u201913) Workshop. 1\u20136."},{"key":"e_1_3_2_37_2","first-page":"457","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201922)","author":"Li Gang","year":"2022","unstructured":"Gang Li, Xiang Li, Yujie Wang, Yichao Wu, Ding Liang, and Shanshan Zhang. 2022. Pseco: Pseudo labeling and consistency training for semi-supervised object detection. In Proceedings of the European Conference on Computer Vision (ECCV\u201922). 457\u2013472."},{"issue":"8","key":"e_1_3_2_38_2","doi-asserted-by":"crossref","first-page":"1644","DOI":"10.1109\/TPAMI.2013.2297321","article-title":"Prediction of human activity by discovering temporal sequence patterns","volume":"36","author":"Li Kang","year":"2014","unstructured":"Kang Li and Yun Fu. 2014. Prediction of human activity by discovering temporal sequence patterns. IEEE Trans. Pattern Anal. Mach. Intell. 36, 8 (2014), 1644\u20131657.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_39_2","first-page":"1902","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201921)","author":"Li Kun","year":"2021","unstructured":"Kun Li, Dan Guo, and Meng Wang. 2021. Proposal-free video grounding with contextual pyramid network. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201921). 1902\u20131910."},{"issue":"10","key":"e_1_3_2_40_2","doi-asserted-by":"crossref","first-page":"202102","DOI":"10.1007\/s11432-022-3783-3","article-title":"ViGT: Proposal-free video grounding with a learnable token in the transformer","volume":"66","author":"Li Kun","year":"2023","unstructured":"Kun Li, Dan Guo, and Meng Wang. 2023. ViGT: Proposal-free video grounding with a learnable token in the transformer. Sci. Chin. Inf. Sci. 66, 10 (2023), 202102.","journal-title":"Sci. Chin. Inf. Sci."},{"key":"e_1_3_2_41_2","first-page":"7083","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201919)","author":"Lin Ji","year":"2019","unstructured":"Ji Lin, Chuang Gan, and Song Han. 2019. Tsm: Temporal shift module for efficient video understanding. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201919). 7083\u20137093."},{"key":"e_1_3_2_42_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201917)","author":"Loshchilov Ilya","year":"2017","unstructured":"Ilya Loshchilov and Frank Hutter. 2017. Sgdr: Stochastic gradient descent with warm restarts. In Proceedings of the International Conference on Learning Representations (ICLR\u201917). 1\u201316."},{"issue":"8","key":"e_1_3_2_43_2","doi-asserted-by":"crossref","first-page":"1979","DOI":"10.1109\/TPAMI.2018.2858821","article-title":"Virtual adversarial training: A regularization method for supervised and semi-supervised learning","volume":"41","author":"Miyato Takeru","year":"2018","unstructured":"Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii. 2018. Virtual adversarial training: A regularization method for supervised and semi-supervised learning. IEEE Trans. Pattern Anal. Mach. Intell. 41, 8 (2018), 1979\u20131993.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"1","key":"e_1_3_2_44_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3485473","article-title":"Fine-grained adversarial semi-supervised learning","volume":"18","author":"Mugnai Daniele","year":"2022","unstructured":"Daniele Mugnai, Federico Pernici, Francesco Turchini, and Alberto Del Bimbo. 2022. Fine-grained adversarial semi-supervised learning. ACM Trans. Multimedia Comput. Commun. Appl. 18, 1s (2022), 1\u201319.","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_2_45_2","first-page":"581","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201914)","author":"Peng Xiaojiang","year":"2014","unstructured":"Xiaojiang Peng, Changqing Zou, Yu Qiao, and Qiang Peng. 2014. Action recognition with stacked fisher vectors. In Proceedings of the European Conference on Computer Vision (ECCV\u201914). 581\u2013595."},{"key":"e_1_3_2_46_2","first-page":"11557","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921)","author":"Pham Hieu","year":"2021","unstructured":"Hieu Pham, Zihang Dai, Qizhe Xie, and Quoc V. Le. 2021. Meta pseudo labels. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921). 11557\u201311568."},{"key":"e_1_3_2_47_2","first-page":"1","article-title":"MultiMatch: Multi-task learning for semi-supervised domain generalization","author":"Qi Lei","year":"2024","unstructured":"Lei Qi, Hongpeng Yang, Yinghuan Shi, and Xin Geng. 2024. MultiMatch: Multi-task learning for semi-supervised domain generalization. ACM Trans. Multimedia Comput. Commun. Appl. (2024), 1\u201320.","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_2_48_2","first-page":"5533","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917)","author":"Qiu Zhaofan","year":"2017","unstructured":"Zhaofan Qiu, Ting Yao, and Tao Mei. 2017. Learning spatio-temporal representation with pseudo-3d residual networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201917). 5533\u20135541."},{"key":"e_1_3_2_49_2","first-page":"1","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Rasmus Antti","year":"2015","unstructured":"Antti Rasmus, Mathias Berglund, Mikko Honkala, Harri Valpola, and Tapani Raiko. 2015. Semi-supervised learning with ladder networks. In Advances in Neural Information Processing Systems (NeurIPS). 1\u20139."},{"key":"e_1_3_2_50_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201921)","author":"Rizve Mamshad Nayeem","year":"2021","unstructured":"Mamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, and Mubarak Shah. 2021. In defense of pseudo-labeling: An uncertainty-aware pseudo-label selection framework for semi-supervised learning. In Proceedings of the International Conference on Learning Representations (ICLR\u201921). 1\u201320."},{"key":"e_1_3_2_51_2","first-page":"2702","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201916)","author":"Rodomagoulakis Isidoros","year":"2016","unstructured":"Isidoros Rodomagoulakis, Nikolaos Kardaris, Vassilis Pitsikalis, E. Mavroudi, Athanasios Katsamanis, Antigoni Tsiami, and Petros Maragos. 2016. Multimodal human action recognition in assistive human-robot interaction. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201916). 2702\u20132706."},{"key":"e_1_3_2_52_2","first-page":"1","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Sajjadi Mehdi","year":"2016","unstructured":"Mehdi Sajjadi, Mehran Javanmardi, and Tolga Tasdizen. 2016. Regularization with stochastic transformations and perturbations for deep semi-supervised learning. In Advances in Neural Information Processing Systems (NeurIPS). 1\u20139."},{"issue":"3","key":"e_1_3_2_53_2","first-page":"1110","article-title":"Hierarchical long short-term concurrent memory for human interaction recognition","volume":"43","author":"Shu Xiangbo","year":"2019","unstructured":"Xiangbo Shu, Jinhui Tang, Guo-Jun Qi, Wei Liu, and Jian Yang. 2019. Hierarchical long short-term concurrent memory for human interaction recognition. IEEE Trans. Pattern Anal. Mach. Intell. 43, 3 (2019), 1110\u20131118.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"8","key":"e_1_3_2_54_2","doi-asserted-by":"crossref","first-page":"5281","DOI":"10.1109\/TCSVT.2022.3142771","article-title":"Expansion-squeeze-excitation fusion network for elderly activity recognition","volume":"32","author":"Shu Xiangbo","year":"2022","unstructured":"Xiangbo Shu, Jiawen Yang, Rui Yan, and Yan Song. 2022. Expansion-squeeze-excitation fusion network for elderly activity recognition. IEEE Trans. Circ. Syst. Vid. Technol. 32, 8 (2022), 5281\u20135292.","journal-title":"IEEE Trans. Circ. Syst. Vid. Technol."},{"issue":"6","key":"e_1_3_2_55_2","first-page":"3300","article-title":"Spatiotemporal co-attention recurrent neural networks for human-skeleton motion prediction","volume":"44","author":"Shu Xiangbo","year":"2021","unstructured":"Xiangbo Shu, Liyan Zhang, Guo-Jun Qi, Wei Liu, and Jinhui Tang. 2021. Spatiotemporal co-attention recurrent neural networks for human-skeleton motion prediction. IEEE Trans. Pattern Anal. Mach. Intell. 44, 6 (2021), 3300\u20133315.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"2","key":"e_1_3_2_56_2","first-page":"663","article-title":"Host\u2013parasite: Graph LSTM-in-LSTM for group activity recognition","volume":"32","author":"Shu Xiangbo","year":"2020","unstructured":"Xiangbo Shu, Liyan Zhang, Yunlian Sun, and Jinhui Tang. 2020. Host\u2013parasite: Graph LSTM-in-LSTM for group activity recognition. IEEE Trans. Neural Netw. Learn. Syst. 32, 2 (2020), 663\u2013674.","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"e_1_3_2_57_2","first-page":"1","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In Advances in Neural Information Processing Systems (NeurIPS). 1\u20139."},{"key":"e_1_3_2_58_2","first-page":"10389","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921)","author":"Singh Ankit","year":"2021","unstructured":"Ankit Singh, Omprakash Chakraborty, Ashutosh Varshney, Rameswar Panda, Rogerio Feris, Kate Saenko, and Abir Das. 2021. Semi-supervised action recognition with temporal contrastive learning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921). 10389\u201310399."},{"key":"e_1_3_2_59_2","first-page":"596","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Sohn Kihyuk","year":"2020","unstructured":"Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A. Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li. 2020. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. In Advances in Neural Information Processing Systems (NeurIPS). 596\u2013608."},{"key":"e_1_3_2_60_2","article-title":"UCF101: A dataset of 101 human actions classes from videos in the wild","author":"Soomro Khurram","year":"2012","unstructured":"Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv:1212.0402. Retrieved from https:\/\/arxiv.org\/abs\/1212.0402","journal-title":"arXiv:1212.0402"},{"issue":"1","key":"e_1_3_2_61_2","first-page":"1929","article-title":"Dropout: A simple way to prevent neural networks from overfitting","volume":"15","author":"Srivastava Nitish","year":"2014","unstructured":"Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. Dropout: A simple way to prevent neural networks from overfitting. J. Mach. Learn. Res. 15, 1 (2014), 1929\u20131958.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_2_62_2","first-page":"1","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Tarvainen Antti","year":"2017","unstructured":"Antti Tarvainen and Harri Valpola. 2017. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In Advances in Neural Information Processing Systems (NeurIPS). 1\u201310."},{"issue":"3","key":"e_1_3_2_63_2","doi-asserted-by":"crossref","first-page":"1305","DOI":"10.1109\/TCSVT.2022.3210271","article-title":"Semi-supervised action recognition from temporal augmentation using curriculum learning","volume":"33","author":"Tong Anyang","year":"2023","unstructured":"Anyang Tong, Chao Tang, and Wenjian Wang. 2023. Semi-supervised action recognition from temporal augmentation using curriculum learning. IEEE Trans. Circ. Syst. Vid. Technol. 33, 3 (2023), 1305\u20131319.","journal-title":"IEEE Trans. Circ. Syst. Vid. Technol."},{"key":"e_1_3_2_64_2","first-page":"1","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Tong Zhan","year":"2022","unstructured":"Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022. VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training. In Advances in Neural Information Processing Systems (NeurIPS). 1\u201325."},{"key":"e_1_3_2_65_2","first-page":"4489","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201915)","author":"Tran Du","year":"2015","unstructured":"Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015. Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201915). 4489\u20134497."},{"key":"e_1_3_2_66_2","first-page":"5552","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Tran Du","year":"2019","unstructured":"Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli. 2019. Video classification with channel-separated convolutional networks. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 5552\u20135561."},{"key":"e_1_3_2_67_2","first-page":"6450","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918)","author":"Tran Du","year":"2018","unstructured":"Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. 2018. A closer look at spatiotemporal convolutions for action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918). 6450\u20136459."},{"key":"e_1_3_2_68_2","doi-asserted-by":"crossref","first-page":"90","DOI":"10.1016\/j.neunet.2021.10.008","article-title":"Interpolation consistency training for semi-supervised learning","volume":"145","author":"Verma Vikas","year":"2022","unstructured":"Vikas Verma, Kenji Kawaguchi, Alex Lamb, Juho Kannala, Yoshua Bengio, and David Lopez-Paz. 2022. Interpolation consistency training for semi-supervised learning. Neural Netw. 145 (2022), 90\u2013106.","journal-title":"Neural Netw."},{"key":"e_1_3_2_69_2","first-page":"1902","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201924)","author":"Wang Fei","year":"2024","unstructured":"Fei Wang, Dan Guo, Kun Li, and Meng Wang. 2024. EulerMormer: Robust eulerian motion magnification via dynamic filtering within transformer. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201924). 1902\u20131910."},{"key":"e_1_3_2_70_2","first-page":"3169","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201911)","author":"Wang Heng","year":"2011","unstructured":"Heng Wang, Alexander Kl\u00e4ser, Cordelia Schmid, and Cheng-Lin Liu. 2011. Action recognition by dense trajectories. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201911). 3169\u20133176."},{"key":"e_1_3_2_71_2","first-page":"3551","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201913)","author":"Wang Heng","year":"2013","unstructured":"Heng Wang and Cordelia Schmid. 2013. Action recognition with improved trajectories. In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201913). 3551\u20133558."},{"key":"e_1_3_2_72_2","article-title":"Towards good practices for very deep two-stream convnets","author":"Wang Limin","year":"2015","unstructured":"Limin Wang, Yuanjun Xiong, Zhe Wang, and Yu Qiao. 2015. Towards good practices for very deep two-stream convnets. arXiv:1507.02159. Retrieved from https:\/\/arxiv.org\/abs\/1507.02159","journal-title":"arXiv:1507.02159"},{"key":"e_1_3_2_73_2","first-page":"20","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201916)","author":"Wang Limin","year":"2016","unstructured":"Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2016. Temporal segment networks: Towards good practices for deep action recognition. In Proceedings of the European Conference on Computer Vision (ECCV\u201916). 20\u201336."},{"key":"e_1_3_2_74_2","doi-asserted-by":"crossref","first-page":"1483","DOI":"10.1145\/3240508.3240671","volume-title":"Proceedings of the ACM International Conference on Multimedia (ACM MM\u201918)","author":"Wang Shuo","year":"2018","unstructured":"Shuo Wang, Dan Guo, Wen gang Zhou, Zheng jun Zha, and Meng Wang. 2018. Connectionist temporal fusion for sign language translation. In Proceedings of the ACM International Conference on Multimedia (ACM MM\u201918). 1483\u20131491."},{"key":"e_1_3_2_75_2","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201923)","author":"Wang Yidong","year":"2023","unstructured":"Yidong Wang, Hao Chen, Qiang Heng, Wenxin Hou, Marios Savvides, Takahiro Shinozaki, Bhiksha Raj, Zhen Wu, and Jindong Wang. 2023. Freematch: Self-adaptive thresholding for semi-supervised learning. In Proceedings of the International Conference on Learning Representations (ICLR\u201923). 1\u201320."},{"key":"e_1_3_2_76_2","doi-asserted-by":"crossref","first-page":"2215","DOI":"10.1109\/TIP.2023.3265261","article-title":"Neighbor-guided consistent and contrastive learning for semi-supervised action recognition","volume":"32","author":"Wu Jianlong","year":"2023","unstructured":"Jianlong Wu, Wei Sun, Tian Gan, Ning Ding, Feijun Jiang, Jialie Shen, and Liqiang Nie. 2023. Neighbor-guided consistent and contrastive learning for semi-supervised action recognition. IEEE Trans. Image Process. 32 (2023), 2215\u20132227.","journal-title":"IEEE Trans. Image Process."},{"key":"e_1_3_2_77_2","first-page":"3252","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201922)","author":"Xiao Junfei","year":"2022","unstructured":"Junfei Xiao, Longlong Jing, Lin Zhang, Ju He, Qi She, Zongwei Zhou, Alan Yuille, and Yingwei Li. 2022. Learning from temporal gradient for semi-supervised action recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201922). 3252\u20133262."},{"key":"e_1_3_2_78_2","first-page":"6256","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Xie Qizhe","year":"2020","unstructured":"Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le. 2020. Unsupervised data augmentation for consistency training. In Advances in Neural Information Processing Systems (NeurIPS). 6256\u20136268."},{"key":"e_1_3_2_79_2","first-page":"10687","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Xie Qizhe","year":"2020","unstructured":"Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V. Le. 2020. Self-training with noisy student improves imagenet classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201920). 10687\u201310698."},{"key":"e_1_3_2_80_2","first-page":"18816","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201923)","author":"Xing Zhen","year":"2023","unstructured":"Zhen Xing, Qi Dai, Han Hu, Jingjing Chen, Zuxuan Wu, and Yu-Gang Jiang. 2023. Svformer: Semi-supervised video transformer for action recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201923). 18816\u201318826."},{"key":"e_1_3_2_81_2","first-page":"7209","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201921)","author":"Xiong Bo","year":"2021","unstructured":"Bo Xiong, Haoqi Fan, Kristen Grauman, and Christoph Feichtenhofer. 2021. Multiview pseudo-labeling for semi-supervised learning from video. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201921). 7209\u20137219."},{"key":"e_1_3_2_82_2","first-page":"3060","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201921)","author":"Xu Mengde","year":"2021","unstructured":"Mengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang, Lijuan Wang, Fangyun Wei, Xiang Bai, and Zicheng Liu. 2021. End-to-end semi-supervised object detection with soft teacher. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201921). 3060\u20133069."},{"key":"e_1_3_2_83_2","first-page":"2959","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201922)","author":"Xu Yinghao","year":"2022","unstructured":"Yinghao Xu, Fangyun Wei, Xiao Sun, Ceyuan Yang, Yujun Shen, Bo Dai, Bolei Zhou, and Stephen Lin. 2022. Cross-model pseudo-labeling for semi-supervised action recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201922). 2959\u20132968."},{"key":"e_1_3_2_84_2","article-title":"Revisiting weak-to-strong consistency in semi-supervised semantic segmentation","author":"Yang Lihe","year":"2022","unstructured":"Lihe Yang, Lei Qi, Litong Feng, Wayne Zhang, and Yinghuan Shi. 2022. Revisiting weak-to-strong consistency in semi-supervised semantic segmentation. arXiv:2208.09910. Retrieved from https:\/\/arxiv.org\/abs\/2208.09910","journal-title":"arXiv:2208.09910"},{"key":"e_1_3_2_85_2","first-page":"1476","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201919)","author":"Zhai Xiaohua","year":"2019","unstructured":"Xiaohua Zhai, Avital Oliver, Alexander Kolesnikov, and Lucas Beyer. 2019. S4l: Self-supervised semi-supervised learning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV\u201919). 1476\u20131485."},{"key":"e_1_3_2_86_2","first-page":"18408","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Zhang Bowen","year":"2021","unstructured":"Bowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu, Jindong Wang, Manabu Okumura, and Takahiro Shinozaki. 2021. Flexmatch: Boosting semi-supervised learning with curriculum pseudo labeling. In Advances in Neural Information Processing Systems (NeurIPS). 18408\u201318419."},{"key":"e_1_3_2_87_2","article-title":"mixup: Beyond empirical risk minimization","author":"Zhang Hongyi","year":"2017","unstructured":"Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. 2017. mixup: Beyond empirical risk minimization. arXiv:1710.09412. Retrieved from https:\/\/arxiv.org\/abs\/1710.09412","journal-title":"arXiv:1710.09412"},{"issue":"3","key":"e_1_3_2_88_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3506711","article-title":"Balanced and accurate pseudo-labels for semi-supervised image classification","volume":"18","author":"Zhao Jian","year":"2022","unstructured":"Jian Zhao, Xianhui Liu, and Weidong Zhao. 2022. Balanced and accurate pseudo-labels for semi-supervised image classification. ACM Trans. Multimedia Comput. Commun. Appl. 18, 3s (2022), 1\u201318.","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_2_89_2","first-page":"803","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201918)","author":"Zhou Bolei","year":"2018","unstructured":"Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba. 2018. Temporal relational reasoning in videos. In Proceedings of the European Conference on Computer Vision (ECCV\u201918). 803\u2013818."},{"key":"e_1_3_2_90_2","first-page":"40","article-title":"Learning representational invariances for data-efficient action recognition","volume":"227","author":"Zou Yuliang","year":"2022","unstructured":"Yuliang Zou, Jinwoo Choi, Qitong Wang, and Jia-Bin Huang. 2022. Learning representational invariances for data-efficient action recognition. Comput. Vis. Image Understand. 227 (2022), 40\u201353.","journal-title":"Comput. Vis. Image Understand."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3655025","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3655025","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:03:51Z","timestamp":1750291431000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3655025"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,7]]},"references-count":89,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,4,30]]}},"alternative-id":["10.1145\/3655025"],"URL":"https:\/\/doi.org\/10.1145\/3655025","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,7]]},"assertion":[{"value":"2024-02-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-03-25","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}