{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,19]],"date-time":"2026-02-19T02:24:22Z","timestamp":1771467862819,"version":"3.50.1"},"reference-count":58,"publisher":"Springer Science and Business Media LLC","issue":"22","license":[{"start":{"date-parts":[[2024,8,10]],"date-time":"2024-08-10T00:00:00Z","timestamp":1723248000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2024,8,10]],"date-time":"2024-08-10T00:00:00Z","timestamp":1723248000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"funder":[{"name":"The Science and Technology Development Fund (FDCT) in Macau","award":["0071\/2022\/A"],"award-info":[{"award-number":["0071\/2022\/A"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62101316"],"award-info":[{"award-number":["62101316"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007219","name":"Shanghai Municipal Natural Science Foundation","doi-asserted-by":"crossref","award":["20ZR1423500"],"award-info":[{"award-number":["20ZR1423500"]}],"id":[{"id":"10.13039\/100007219","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Appl Intell"],"published-print":{"date-parts":[[2024,11]]},"DOI":"10.1007\/s10489-024-05661-1","type":"journal-article","created":{"date-parts":[[2024,8,10]],"date-time":"2024-08-10T08:02:30Z","timestamp":1723276950000},"page":"11177-11195","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Integrating pseudo labeling with contrastive clustering for transformer-based semi-supervised action recognition"],"prefix":"10.1007","volume":"54","author":[{"given":"Nannan","family":"Li","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kan","family":"Huang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qingtian","family":"Wu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yang","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,8,10]]},"reference":[{"key":"5661_CR1","doi-asserted-by":"crossref","unstructured":"Carreira J, Zisserman A (2017) Quo vadis, action recognition? a new model and the kinetics dataset. In: Proc IEEE\/CVF Conf Comput Vis Pattern Recognit pp 6299\u20136308","DOI":"10.1109\/CVPR.2017.502"},{"key":"5661_CR2","unstructured":"Dong-Hyun L (2013) Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In: Proc Int Conf Mach Learn workshop"},{"key":"5661_CR3","unstructured":"Xie Q, Dai Z, Hovy E, Luong T, Le Q (2020) Unsupervised data augmentation for consistency training. In: Proc Int Conf Neural Inf Process Syst"},{"key":"5661_CR4","unstructured":"Sohn K, Berthelot D, Carlini N, Zhang Z, Zhang H, Raffel CA, Cubuk ED, Kurakin A, Li C-L (2020) \u201cFixmatch: Simplifying semi-supervised learning with consistency and confidence. In: Proc Int Conf Neural Inf Process Syst"},{"key":"5661_CR5","unstructured":"Zhen X, Dai Q, Hu H, Chen J, Wu Z, Jiang Y-G (2023) Svformer: Semi-supervised video transformer for action recognition. In: Proc IEEE\/CVF Conf Comput Vis Pattern Recognit"},{"key":"5661_CR6","unstructured":"Soomro K, Zamir AR, Shah M (2012) Ucf101: A dataset of 101 human actions classes from videos In: the wild. In: CRCV-TR-12-01"},{"key":"5661_CR7","doi-asserted-by":"crossref","unstructured":"Xiong B, Fan H, Grauman K, Feichtenhofer C (2021) Multiview pseudo-labeling for semi-supervised learning from video. In: Proc IEEE Int Conf Comput Vis","DOI":"10.1109\/ICCV48922.2021.00712"},{"key":"5661_CR8","doi-asserted-by":"crossref","unstructured":"Xu Y, Wei F, Sun X, Yang C, Shen Y, Dai B, Zhou B, Lin S (2022) Cross-model pseudo-labeling for semi-supervised action recognition. In: Proc IEEE\/CVF Conf Comput Vis Pattern Recognit","DOI":"10.1109\/CVPR52688.2022.00297"},{"key":"5661_CR9","doi-asserted-by":"crossref","unstructured":"Singh A, Chakraborty O, Varshney A, Panda R, Feris R, Saenko K, Das A (2021) Semi-supervised action recognition with temporal contrastive learning. In: Proc IEEE\/CVF Conf Comput Vis Pattern Recognit","DOI":"10.1109\/CVPR46437.2021.01025"},{"key":"5661_CR10","doi-asserted-by":"crossref","unstructured":"Dave I, Gupta R, Rizve MN, Shah M (2022) Tclr: Temporal contrastive learning for video representation. In: Comput Vis Image Und vol 219, pp 103\u2013106","DOI":"10.1016\/j.cviu.2022.103406"},{"key":"5661_CR11","unstructured":"Kanchana R, Naseer M, Khan S, Khan FS, Ryoo MS (2022) Self-supervised video transformer. In: Proc IEEE\/CVF Conf Comput Vis Pattern Recognit pp 2874\u20132884"},{"key":"5661_CR12","doi-asserted-by":"crossref","unstructured":"Takeru\u00a0Miyato MK, Maeda S-i, Ishii S (2018) Virtual adversarial training: a regularization method for supervised and semi-supervised learning. In: IEEE Trans Pattern Anal Mach Intell, vol\u00a048, pp 1979\u20131993","DOI":"10.1109\/TPAMI.2018.2858821"},{"key":"5661_CR13","unstructured":"Tarvainen A, Valpola H (2017) Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In: Proc Int Conf Neural Inf Process Syst"},{"key":"5661_CR14","doi-asserted-by":"crossref","unstructured":"Chen J, Yang M, Ling J (2021) Attention-based label consistency for semi-supervised deep learning based image classification. In: Neurocomput, vol 453 pp 731\u2013741","DOI":"10.1016\/j.neucom.2020.06.133"},{"key":"5661_CR15","doi-asserted-by":"crossref","unstructured":"Li X, Wu Y, Dai S (2023) Semi-supervised medical imaging segmentation with soft pseudo-label fusion. In: Appl Intell ,vol\u00a053, pp 20\u00a0573\u201320\u00a0765","DOI":"10.1007\/s10489-023-04569-6"},{"key":"5661_CR16","doi-asserted-by":"crossref","unstructured":"Wang X, Kihara D, Luo J, jun Qi G (2021) Enaet: A self-trained framework for semi-supervised and supervised learning with ensemble transformations. In: IEEE Trans Image Process vol\u00a030, pp 1639\u20131647","DOI":"10.1109\/TIP.2020.3044220"},{"key":"5661_CR17","unstructured":"Berthelot D, Carlini N, Goodfellow IJ, Papernot N, Oliver A, Raffel C (2019) Mixmatch: A holistic approach to semi-supervised learning. In: Proc Int Conf Neural Inf Process Syst"},{"key":"5661_CR18","unstructured":"Berthelot D, Carlini N, Cubuk ED, Kurakin A, Sohn K, Zhang H, Raffel C (2020) Remixmatch: Semi-supervised learning with distribution matching and augmentation anchoring. In: Proc Int Conf Learn Representations"},{"key":"5661_CR19","unstructured":"Li J, Socher R, Hoi Sch (2020) Dividemix: Learning with noisy labels as semi-supervised learning. In: Proc Int Conf Learn Representations"},{"key":"5661_CR20","doi-asserted-by":"crossref","unstructured":"Tong A, Tang C, Wang W (2022) Semisupervised action recognition from temporal augmentation using curriculum learning. In: IEEE Trans Circuits Syst Video Technol vol\u00a033, pp 1305\u20131319","DOI":"10.1109\/TCSVT.2022.3210271"},{"key":"5661_CR21","doi-asserted-by":"crossref","unstructured":"Tu Z, Shu X, Huang P, Yan R, Liu Z, Zhang J (2024) Leveraging frame- and feature-level progressive augmentation for semi-supervised action recognition. In: ACM Trans Multimedia Comput Commun Appl","DOI":"10.1145\/3655025"},{"key":"5661_CR22","doi-asserted-by":"crossref","unstructured":"Gao G, Liu Z, Zhang G, Li J, Qin A (2023) Danet: Semi-supervised differentiated auxiliaries guided network for video action recognition. In: Neural Netwworks, vol 158, pp 121\u2013131","DOI":"10.1016\/j.neunet.2022.11.009"},{"key":"5661_CR23","doi-asserted-by":"crossref","unstructured":"Wu J, Sun W, Gan T, Ding N, Jiang F, Shen J, Nie L (2023) Neighbor-guided consistent and contrastive learning for semi-supervised action recognition. In: IEEE Trans Image Process vol\u00a032, pp. 2215\u20132227","DOI":"10.1109\/TIP.2023.3265261"},{"key":"5661_CR24","doi-asserted-by":"crossref","unstructured":"Assefa M, Jiang W, Zhan J, Gedamu K, Yilma G, Ayalew M, Adhikari D (2004) Audio-visual contrastive and consistency learning for semi-supervised action recognition. In: IEEE Trans Multimedia vol\u00a026, pp 3491\u20133504","DOI":"10.1109\/TMM.2023.3312856"},{"key":"5661_CR25","doi-asserted-by":"crossref","unstructured":"Shu X, Xu B, Tab LZ, Tang J (2023) Multi-granularity anchorcontrastive representation learning for semi-supervised skeleton-based action recognition. In: IEEE Trans Pattern Anal Mach Intell vol\u00a045, pp 7559\u20137576","DOI":"10.1109\/TPAMI.2022.3222871"},{"key":"5661_CR26","doi-asserted-by":"crossref","unstructured":"Jun X, Li L, Xu D, Long C, Shao J, Zhang S, Pu S, Zhuang Y (2020) Explore video clip order with self-supervised and curriculum learning for video applications. In: IEEE Trans Multimedia vol\u00a023, pp 3454\u20133466","DOI":"10.1109\/TMM.2020.3025661"},{"key":"5661_CR27","doi-asserted-by":"crossref","unstructured":"Jiang Y, Li X, Chen Y, He Y, Xu Q, Yang Z, Cao X, Huang Q (2023) Maxmatch: Semi-supervised learning with worst-case consistency. In: IEEE Trans Pattern Anal Mach Intell vol\u00a045, pp 5970\u20135987","DOI":"10.1109\/TPAMI.2022.3208419"},{"key":"5661_CR28","doi-asserted-by":"crossref","unstructured":"Park JH, Kim JH, Ngo BH, Kwon JE, Cho SI (2023) Adversarial representation teaching with perturbation-agnostic student-teacher structure for semi-supervised learning. In: Appl Intell vol\u00a053, pp 26\u00a0797\u201326\u00a0809","DOI":"10.1007\/s10489-023-04950-5"},{"key":"5661_CR29","doi-asserted-by":"crossref","unstructured":"Chavoshinejad J, Seyedi SA, Tab FA, Salahian N (2023) Self-supervised semi-supervised nonnegative matrix factorization for data clustering. In: Pattern Recognit vol 137, p 109282","DOI":"10.1016\/j.patcog.2022.109282"},{"key":"5661_CR30","doi-asserted-by":"crossref","unstructured":"Zhai X, Oliver A, Kolesnikov A, Beyer L (2019) S4l: Self-supervised semi-supervised learning. In: Proc IEEE Int Conf Comput Vis","DOI":"10.1109\/ICCV.2019.00156"},{"key":"5661_CR31","doi-asserted-by":"crossref","unstructured":"Jing L, Parag T, Wu Z, Tian Y, Wang H (2021) Videossl: Semi-supervised learning for video classification. In: Proc IEEE\/CVF Win Conf Appl Comput Vis","DOI":"10.1109\/WACV48630.2021.00115"},{"key":"5661_CR32","doi-asserted-by":"crossref","unstructured":"Xiao J, Jing L, Zhang L, He J, She Q, Zhou Z, Yuille A, Li Y (2022) Learning from temporal gradient for semi-supervised action recognition. In: Proc IEEE\/CVF Conf Comput Vis Pattern Recognit","DOI":"10.1109\/CVPR52688.2022.00325"},{"key":"5661_CR33","doi-asserted-by":"crossref","unstructured":"Xu B, Shu X, Song Y (2022) X-invariant contrastive augmentation and representation learning for semi-supervised skeleton-based action recognition. In: IEEE Trans Image Process vol\u00a031, pp 3852\u20133867","DOI":"10.1109\/TIP.2022.3175605"},{"key":"5661_CR34","unstructured":"Kaiming H, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proc IEEE\/CVF Conf Comput Vis Pattern Recognit"},{"key":"5661_CR35","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S (2021) An image is worth 16x16 words: Transformers for image recognition at scale. In: Proc Int Conf Learn Representations"},{"key":"5661_CR36","unstructured":"Bertasius G, Wang H, Torresani L (2021) Is space-time attention all you need for video understanding? In: Proc Int Conf Mach Learn"},{"key":"5661_CR37","doi-asserted-by":"crossref","unstructured":"Liu Z, Ning J, Cao Y, Wei Y, Zhang Z (2022) Video swin transformer. In: Proc IEEE\/CVF Conf Comput Vis Pattern Recognit","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"5661_CR38","doi-asserted-by":"crossref","unstructured":"Ahn D, Kim S, Ko BC (2023) Star++: Rethinking spatio-temporal cross attention transformer for video action recognition. In: Appl Intell vol\u00a053, pp 28\u00a0446\u201328\u00a0459","DOI":"10.1007\/s10489-023-04978-7"},{"key":"5661_CR39","doi-asserted-by":"crossref","unstructured":"Liang J, Cao J, Fan Y, Zhang K, Li RRY, Timofte R, Gool LV (2024) Vrt: A video restoration transformer. In: IEEE Trans image Process vol\u00a033, pp 2171\u20132182","DOI":"10.1109\/TIP.2024.3372454"},{"key":"5661_CR40","doi-asserted-by":"crossref","unstructured":"Fan H, Xiong B, Mangalam K, Li Y, Yan Z, Malik J, Feichtenhofer C (2021) Multiscale vision transformers. In: Proc IEEE Int Conf Comput Vis","DOI":"10.1109\/ICCV48922.2021.00675"},{"key":"5661_CR41","doi-asserted-by":"crossref","unstructured":"Schiappa MC, Rawat YS, Shah M (2023) Self-supervised learning for videos: A survey. In: ACM Computing Surveys, vol\u00a055, pp 1\u201337","DOI":"10.1145\/3577925"},{"key":"5661_CR42","unstructured":"Kaiming H, Fan H, Wu Y, Xie S, Girshick R (2020) Momentum contrast for unsupervised visual representation learning. In: Proc IEEE\/CVF Conf Comput Vis Pattern Recognit pp 9729\u20139738"},{"key":"5661_CR43","unstructured":"Ting C, Kornblith S, Norouzi M, Hinton G (2020) Simclr: A simple framework for contrastive learning of visual representations. In: Proc Int Conf Mach Learn pp 1597\u20131607"},{"key":"5661_CR44","unstructured":"Jean-Bastien G, Strub F, Altch\u00e9 F, Tallec C, Richemond P, Buchatskaya E, Doersch C (2020) Bootstrap your own latent-a new approach to self-supervised learning. In: Proc Int Conf Neural Inf Process Syst"},{"key":"5661_CR45","unstructured":"Hangbo B, Dong L, Piao S, Wei F (2022) Beit: Bert pre-training of image transformers. In: Proc Int Conf Learn Representations"},{"key":"5661_CR46","unstructured":"Kaiming H, Chen X, Xie S, Li Y, Doll\u00e1r P, Girshick R (2022) Masked autoencoders are scalable vision learners. In: Proc IEEE\/CVF Conf Comput Vis Pattern Recognit"},{"key":"5661_CR47","unstructured":"Junnan L, Zhou P, Xiong C, Hoi SC (2021) Prototypical contrastive learning of unsupervised representations. In: Proc Int Conf Learn Representations"},{"key":"5661_CR48","doi-asserted-by":"crossref","unstructured":"Kuehne H, Jhuang H, Garrote E, Poggio T, Serre T (2011) Hmdb: A large video database for human motion recognition. In: Proc Int Conf Comput Vis pp 2556\u20132563","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"5661_CR49","unstructured":"P KD, Ba J (2015) Adam: A method for stochastic optimization. In: Proc Int Conf Learn Representations"},{"key":"5661_CR50","doi-asserted-by":"crossref","unstructured":"Rajendrakumar DI, Rizve MN, Chen C, Shah M (2023) Timebalance: Temporally-invariant and temporally-distinctive video representations for semi-supervised action recognition. In: Proc IEEE\/CVF Conf Comput Vis Pattern Recognit pp 2341\u20132352","DOI":"10.1109\/CVPR52729.2023.00232"},{"key":"5661_CR51","doi-asserted-by":"crossref","unstructured":"Yuliang Z, Choi J, Wang Q, Huang J-B (2023) Learning representational invariances for data-efficient action recognition. In: Comput Vis Image Und vol 227, p 103597","DOI":"10.1016\/j.cviu.2022.103597"},{"key":"5661_CR52","doi-asserted-by":"crossref","unstructured":"Assefa M, Jiang W, Alemu KG, Yilma G, Adhikari D, Ayalew M, Seid AM, Erbad A (2023) Actor-aware self-supervised learning for semi-supervised video representation learning. In: IEEE Trans Circuits Syst Video Technol vol\u00a033, pp 6679\u20136692","DOI":"10.1109\/TCSVT.2023.3267178"},{"key":"5661_CR53","doi-asserted-by":"crossref","unstructured":"Gavrilyuk K, Jain M, Karmanov I, Snoek CG (2021) Motion-augmented self-training for video recognition at smaller scale. In: Proc IEEE\/CVF Conf Comput Vis Pattern Recognit pp 10\u00a0429\u201310\u00a0438","DOI":"10.1109\/ICCV48922.2021.01026"},{"key":"5661_CR54","doi-asserted-by":"crossref","unstructured":"Feichtenhofer C, Fan H, Malik J, He K (2019) Slowfast networks for video recognition. In: Proc IEEE Int Conf Comput Vis pp 6202\u20136211","DOI":"10.1109\/ICCV.2019.00630"},{"key":"5661_CR55","doi-asserted-by":"crossref","unstructured":"Zagoruyko S, Komodakis (2016) Wide residual networks. In: Proc Brit Mach Vis Conf","DOI":"10.5244\/C.30.87"},{"key":"5661_CR56","doi-asserted-by":"crossref","unstructured":"Cubuk ED, Zoph B, Shlens J, Le QV (2020) Randaugment: Practical automated data augmentation with a reduced search space. In: Proc IEEE\/CVF Conf Comput Vis Pattern Recognit Workshops","DOI":"10.1109\/CVPRW50498.2020.00359"},{"key":"5661_CR57","doi-asserted-by":"crossref","unstructured":"Li J, Xiong C, Hoi (2021) Comatch: Semi-supervised learning with contrastive graph regularization. In: Proc IEEE Int Conf Comput Vis","DOI":"10.1109\/ICCV48922.2021.00934"},{"key":"5661_CR58","doi-asserted-by":"crossref","unstructured":"Zhou B, Lu J, Liu K, Xu Y, Cheng Z, Niu Y (2023) Hypermatch:noise-tolerant semi-supervised learning via relaxed contrastive constraiint. In: Proc IEEE Int Conf Comput Vis","DOI":"10.1109\/CVPR52729.2023.02300"}],"container-title":["Applied Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-024-05661-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10489-024-05661-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10489-024-05661-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,18]],"date-time":"2024-09-18T15:22:15Z","timestamp":1726672935000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10489-024-05661-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,10]]},"references-count":58,"journal-issue":{"issue":"22","published-print":{"date-parts":[[2024,11]]}},"alternative-id":["5661"],"URL":"https:\/\/doi.org\/10.1007\/s10489-024-05661-1","relation":{},"ISSN":["0924-669X","1573-7497"],"issn-type":[{"value":"0924-669X","type":"print"},{"value":"1573-7497","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,8,10]]},"assertion":[{"value":"30 June 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 August 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no competing interests to declare that are relevant to the content of this article.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interests"}}]}}