{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T16:11:51Z","timestamp":1781799111466,"version":"3.54.5"},"reference-count":181,"publisher":"American Association for the Advancement of Science (AAAS)","content-domain":{"domain":["spj.science.org"],"crossmark-restriction":true},"short-container-title":["Intell Comput"],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:p>Affective video facial analysis (AVFA) has emerged as a key research field for building emotion-aware intelligent systems, yet this field continues to suffer from limited data availability. In recent years, the self-supervised learning (SSL) technique of masked autoencoders (MAEs) has gained momentum, with growing adaptations in audiovisual contexts. While scaling up has proven essential for breakthroughs in general multimodal learning domains, its specific impact on AVFA remains largely unexplored. Another core challenge in this field is capturing both intra- and intermodal correlations through scalable audiovisual representations. To tackle these issues, we propose AVF-MAE++, a family of audiovisual MAE models designed to efficiently investigate the scaling properties in AVFA while enhancing cross-modal correlation modeling. Our framework introduces a dual masking strategy across audio and visual modalities and strengthens modality encoders with a more holistic design to better support scalable pretraining. Additionally, we present an iterative audiovisual correlation learning module that improves correlation learning within the SSL paradigm, addressing the limitations of previous methods. To support smooth adaptation and reduce overfitting risks, we introduce a progressive semantic injection strategy, organizing the model training into 3 structured stages. Extensive experiments conducted on 17 datasets, covering 3 major AVFA tasks, demonstrate that AVF-MAE++ achieves consistent state-of-the-art performance across multiple benchmarks. Comprehensive ablation studies highlight the importance of each proposed component and provide deeper insights into the design choices driving these improvements. Our code and models have been publicly released at Zenodo.<\/jats:p>","DOI":"10.34133\/icomputing.0246","type":"journal-article","created":{"date-parts":[[2025,11,26]],"date-time":"2025-11-26T13:03:09Z","timestamp":1764162189000},"update-policy":"https:\/\/doi.org\/10.34133\/aaas_crossmark_01","source":"Crossref","is-referenced-by-count":2,"title":["Scalable Audiovisual Masked Autoencoders for Efficient Affective Video Facial Analysis"],"prefix":"10.34133","volume":"5","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6244-0269","authenticated-orcid":false,"given":"Xuecheng","family":"Wu","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Junxiao","family":"Xue","sequence":"additional","affiliation":[{"name":"Research Center for Space Computing System, Zhejiang Lab, Hangzhou, Zhejiang, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinyi","family":"Yin","sequence":"additional","affiliation":[{"name":"School of Cyber Science and Engineering, \rZhengzhou University, Zhengzhou, Henan, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yunyun","family":"Shi","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Liangyu","family":"Fu","sequence":"additional","affiliation":[{"name":"School of Software, \rNorthwestern Polytechnical University, Xi\u2019an, Shaanxi, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Danlei","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yifan","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Advanced Technology, \rUniversity of Science and Technology of China, Hefei, Anhui, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jia","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiayu","family":"Nie","sequence":"additional","affiliation":[{"name":"Inspur Group, Jinan, Shandong, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jun","family":"Wang","sequence":"additional","affiliation":[{"name":"Research Center for Space Computing System, Zhejiang Lab, Hangzhou, Zhejiang, China."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"221","published-online":{"date-parts":[[2026,1,21]]},"reference":[{"issue":"12","key":"e_1_3_3_2_2","doi-asserted-by":"crossref","first-page":"1424","DOI":"10.1109\/34.895976","article-title":"Automatic analysis of facial expressions: The state of the art","volume":"22","author":"Pantic M","year":"2000","unstructured":"Pantic M, Rothkrantz LJM. Automatic analysis of facial expressions: The state of the art. IEEE Trans Pattern Anal Mach Intell. 2000;22(12):1424\u20131445.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"9","key":"e_1_3_3_3_2","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1109\/MC.2018.3620963","article-title":"The age of artificial emotional intelligence","volume":"51","author":"Schuller D","year":"2018","unstructured":"Schuller D, Schuller BW. The age of artificial emotional intelligence. Computer. 2018;51(9):38\u201346.","journal-title":"Computer"},{"key":"e_1_3_3_4_2","doi-asserted-by":"crossref","first-page":"1522","DOI":"10.1109\/TIP.2024.3365248","article-title":"Dual-view curricular optimal transport for cross-lingual cross-modal retrieval","volume":"33","author":"Wang Y","year":"2024","unstructured":"Wang Y, Wang S, Luo H, Dong J, Wang F, Han M, Wang X, Wang M. Dual-view curricular optimal transport for cross-lingual cross-modal retrieval. IEEE Trans Image Process. 2024;33:1522\u20131533.","journal-title":"IEEE Trans Image Process"},{"key":"e_1_3_3_5_2","doi-asserted-by":"crossref","unstructured":"Mehrabian A. Communication without words. In: Communication theory. New York (NY): Routledge; 2017. p. 193\u2013200.","DOI":"10.4324\/9781315080918-15"},{"issue":"1","key":"e_1_3_3_6_2","article-title":"Survey on audiovisual emotion recognition: Databases, features, and data fusion strategies","volume":"3","author":"Wu CH","year":"2014","unstructured":"Wu CH, Lin JC, Wei WL. Survey on audiovisual emotion recognition: Databases, features, and data fusion strategies. APSIPA Trans Signal Inf Process. 2014;3(1): Article e12.","journal-title":"APSIPA Trans Signal Inf Process"},{"issue":"5","key":"e_1_3_3_7_2","doi-asserted-by":"crossref","first-page":"3192","DOI":"10.1109\/TCSVT.2023.3312858","article-title":"Transformer-based multimodal emotional perception for dynamic facial expression recognition in the wild","volume":"34","author":"Zhang X","year":"2023","unstructured":"Zhang X, Li M, Lin S, Xu H, Xiao G. Transformer-based multimodal emotional perception for dynamic facial expression recognition in the wild. IEEE Trans Circuits Syst Video Technol. 2023;34(5):3192\u20133203.","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"issue":"10","key":"e_1_3_3_8_2","doi-asserted-by":"crossref","first-page":"3030","DOI":"10.1109\/TCSVT.2017.2719043","article-title":"Learning affective features with a hybrid deep model for audio\u2013visual emotion recognition","volume":"28","author":"Zhang S","year":"2017","unstructured":"Zhang S, Zhang S, Huang T, Gao W, Tian Q. Learning affective features with a hybrid deep model for audio\u2013visual emotion recognition. IEEE Trans Circuits Syst Video Technol. 2017;28(10):3030\u20133043.","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"e_1_3_3_9_2","doi-asserted-by":"crossref","first-page":"69","DOI":"10.1016\/j.inffus.2018.09.008","article-title":"Emotion recognition using deep learning approach from audio\u2013visual emotional big data","volume":"49","author":"Hossain MS","year":"2019","unstructured":"Hossain MS, Muhammad G. Emotion recognition using deep learning approach from audio\u2013visual emotional big data. Inf Fusion. 2019;49:69\u201378.","journal-title":"Inf Fusion"},{"issue":"1","key":"e_1_3_3_10_2","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1109\/TAFFC.2024.3436913","article-title":"SVFAP: Self-supervised video facial affect perceiver","volume":"16","author":"Sun L","year":"2024","unstructured":"Sun L, Lian Z, Wang K, He Y, Xu M, Sun H, Liu B, Tao J. SVFAP: Self-supervised video facial affect perceiver. IEEE Trans Affect Comput. 2024;16(1):405\u2013422.","journal-title":"IEEE Trans Affect Comput"},{"key":"e_1_3_3_11_2","doi-asserted-by":"crossref","unstructured":"Wu X Sun H Xue J Nie J Kong X Zhai R Huang D He L. Towards emotion analysis in short-form videos: A large-scale dataset and baseline. In: Proceedings of the 2025 International Conference on Multimedia Retrieval. Chicago (IL): ACM; 2025. p. 1497\u20131506.","DOI":"10.1145\/3731715.3733453"},{"key":"e_1_3_3_12_2","doi-asserted-by":"crossref","unstructured":"He K Chen X Xie S Li Y Dollar P Girshick R. Masked autoencoders are scalable vision learners. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. New Orleans (LA): IEEE; 2022. p. 16000\u201316009.","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"e_1_3_3_13_2","doi-asserted-by":"crossref","DOI":"10.1016\/j.inffus.2024.102382","article-title":"HiCMAE: Hierarchical contrastive masked autoencoder for self-supervised audio-visual emotion recognition","volume":"108","author":"Sun L","year":"2024","unstructured":"Sun L, Lian Z, Liu B, Tao J. HiCMAE: Hierarchical contrastive masked autoencoder for self-supervised audio-visual emotion recognition. Inf Fusion. 2024;108: Article 102382.","journal-title":"Inf Fusion"},{"key":"e_1_3_3_14_2","doi-asserted-by":"crossref","unstructured":"Sun L Lian Z Liu B Tao J. MAE-DFER: Efficient masked autoencoder for self-supervised dynamic facial expression recognition. In: Proceedings of the 31st ACM International Conference on Multimedia. Ottawa (ON Canada): ACM; 2023. p. 6110\u20136121.","DOI":"10.1145\/3581783.3612365"},{"key":"e_1_3_3_15_2","unstructured":"Bao H Dong L Piao S Wei F. BEiT: BERT pre-training of image transformers. arXiv. 2021. https:\/\/doi.org\/10.48550\/arXiv.2106.08254"},{"key":"e_1_3_3_16_2","doi-asserted-by":"crossref","unstructured":"Wang L Huang B Zhao Z Tong Z He Y Wang Y Wang Y Qiao Y. VideoMAE V2: Scaling video masked autoencoders with dual masking. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Vancouver (BC Canada): IEEE; 2023. p. 14549\u201314560.","DOI":"10.1109\/CVPR52729.2023.01398"},{"key":"e_1_3_3_17_2","unstructured":"Wu X Sun H Xue J Nie J Kong X Zhai R He L. eMotions: A large-scale dataset for emotion recognition in short videos. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2311.17335"},{"issue":"7","key":"e_1_3_3_18_2","doi-asserted-by":"crossref","first-page":"5299","DOI":"10.1109\/TPAMI.2024.3429301","article-title":"Self-supervised multimodal learning: A survey","volume":"47","author":"Zong Y","year":"2024","unstructured":"Zong Y, MacAodha O, Hospedales TM. Self-supervised multimodal learning: A survey. IEEE Trans Pattern Anal Mach Intell. 2024;47(7):5299\u20135318.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"e_1_3_3_19_2","first-page":"10078","article-title":"VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training","volume":"35","author":"Tong Z","year":"2022","unstructured":"Tong Z, Song Y, Wang J, Wang L. VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Adv Neural Inf Proces Syst. 2022;35:10078\u201310093.","journal-title":"Adv Neural Inf Proces Syst"},{"key":"e_1_3_3_20_2","doi-asserted-by":"crossref","unstructured":"Liu Y Dai W Feng C Wang W Yin G Zeng J Shan S. MAFW: A large-scale multi-modal compound affective database for dynamic facial expression recognition in the wild. In: Proceedings of the 30th ACM International Conference on Multimedia. Lisbon (Portugal): ACM; 2022. p. 24\u201332.","DOI":"10.1145\/3503161.3548190"},{"key":"e_1_3_3_21_2","doi-asserted-by":"crossref","unstructured":"Cheng Z Cheng ZQ He JY Sun J Wang K Lin Y Lian Z Peng X Hauptmann A. Emotion-LLaMA: Multimodal emotion recognition and reasoning with instruction tuning. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2406.11161","DOI":"10.52202\/079017-3518"},{"key":"e_1_3_3_22_2","doi-asserted-by":"crossref","unstructured":"Wang Y Wang L Zhou Q Wang Z Li H Hua G Tang W. Multimodal LLM enhanced cross-lingual cross-modal retrieval. In: Proceedings of the 32nd ACM International Conference on Multimedia. New York (NY): ACM; 2024. p. 8296\u20138305.","DOI":"10.1145\/3664647.3680886"},{"issue":"3","key":"e_1_3_3_23_2","first-page":"2531","article-title":"One-shot talking face generation from single-speaker audio visual correlation learning","volume":"36","author":"Wang S","year":"2022","unstructured":"Wang S, Li L, Ding Y, Yu X. One-shot talking face generation from single-speaker audio visual correlation learning. Proc AAAI Conf Artif Intell. 2022;36(3):2531\u20132539.","journal-title":"Proc AAAI Conf Artif Intell"},{"key":"e_1_3_3_24_2","doi-asserted-by":"crossref","unstructured":"Wang Y Wu X Zhang J Jing M Lu K Yu J Su W Gao F Liu Q Sun J et\u00a0al. Building robust video-level deepfake detection via audio visual local-global interactions. In: Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne (VIC Australia): ACM; 2024. p. 11370\u201311376.","DOI":"10.1145\/3664647.3688985"},{"key":"e_1_3_3_25_2","doi-asserted-by":"crossref","unstructured":"Wu X Huang D Sun H Yin X Wang Y Wang H Zhang J Wang F Guo P Xing S et\u00a0al. HOLA: Enhancing audio-visual deepfake detection via hierarchical contextual aggregations and efficient pre-training. arXiv. 2025. https:\/\/doi.org\/10.48550\/arXiv.2507.22781","DOI":"10.1145\/3746027.3761980"},{"key":"e_1_3_3_26_2","unstructured":"Lian Z Sun H Sun L Gu H Wen Z Zhang Z Chen S Xu M Xu K Chen K et\u00a0al. Explainable multimodal emotion reasoning. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2306.15401"},{"key":"e_1_3_3_27_2","doi-asserted-by":"crossref","unstructured":"Wu X Sun H Wang Y Nie J Zhang J Wang Y Xue J He L. AVF-MAE++: Scaling affective video facial masked autoencoders via efficient audio-visual self-supervised learning. In: Proceedings of the Computer Vision and Pattern Recognition Conference. Piscataway (NJ): IEEE Computer Society; 2025. p. 9142\u20139153.","DOI":"10.1109\/CVPR52734.2025.00854"},{"key":"e_1_3_3_28_2","doi-asserted-by":"crossref","unstructured":"Zhang Z Wu X Huang D Yan S Peng C Cao X. HKD4VLM: A progressive hybrid knowledge distillation framework for robust multimodal hallucination and factuality detection in VLMs. arXiv. 2025. https:\/\/doi.org\/10.48550\/arXiv.2506.13038","DOI":"10.1145\/3746027.3762014"},{"key":"e_1_3_3_29_2","doi-asserted-by":"crossref","unstructured":"Chumachenko K Iosifidis A Gabbouj M. MMA-DFER: Multimodal adaptation of unimodal models for dynamic facial expression recognition in-the-wild. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Piscataway (NJ): IEEE; 2024. p. 4673\u20134682.","DOI":"10.1109\/CVPRW63382.2024.00470"},{"key":"e_1_3_3_30_2","unstructured":"Zhao Z Patras I. Prompting visual-language models for dynamic facial expression recognition. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2308.13382"},{"key":"e_1_3_3_31_2","doi-asserted-by":"crossref","unstructured":"Zhang Z Wang L Yang J. Weakly supervised video emotion detection and prediction via cross-modal temporal erasing network. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Vancouver (Canada): IEEE; 2023. p. 18888\u201318897.","DOI":"10.1109\/CVPR52729.2023.01811"},{"key":"e_1_3_3_32_2","doi-asserted-by":"crossref","unstructured":"Goncalves L Leem SG Lin WC Sisman B Busso C. Versatile audio-visual learning for emotion recognition. In: IEEE Transactions on Affective Computing. Piscataway (NJ): IEEE; 2024.","DOI":"10.1109\/TAFFC.2024.3433386"},{"key":"e_1_3_3_33_2","doi-asserted-by":"crossref","unstructured":"Wang H Zheng S Chen Y Cheng L Chen Q. Cam++: A fast and efficient network for speaker verification using context-aware masking. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2303.00332","DOI":"10.21437\/Interspeech.2023-1513"},{"key":"e_1_3_3_34_2","doi-asserted-by":"crossref","unstructured":"Liu Z Ning J Cao Y Wei Y Zhang Z Lin S Hu H. Video swin transformer. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. New Orleans (LA): IEEE; 2022. p. 3202\u20133211.","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"e_1_3_3_35_2","doi-asserted-by":"crossref","unstructured":"Zhang Z Zhao P Park E Yang J. MART: Masked affective representation learning via masked temporal distribution distillation. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Seattle (WA): IEEE; 2024. p. 12830\u201312840.","DOI":"10.1109\/CVPR52733.2024.01219"},{"key":"e_1_3_3_36_2","doi-asserted-by":"crossref","unstructured":"Hsu WN Bolte B Tsai YHH Lakhotia K Salakhutdinov R Mohamed A. HuBERT: Self supervised speech representation learning by masked prediction of hidden units. In: IEEE\/ACM Transactions on Audio Speech and Language Processing. Piscataway (NJ): IEEE; 2021. p. 3451\u20133460.","DOI":"10.1109\/TASLP.2021.3122291"},{"key":"e_1_3_3_37_2","doi-asserted-by":"crossref","first-page":"184","DOI":"10.1016\/j.inffus.2018.06.003","article-title":"Audio-visual emotion fusion (AVEF): A deep efficient weighted approach","volume":"46","author":"Ma Y","year":"2019","unstructured":"Ma Y, Hao Y, Chen M, Chen J, Lu P, Kosir A. Audio-visual emotion fusion (AVEF): A deep efficient weighted approach. Inf Fusion. 2019;46:184\u2013192.","journal-title":"Inf Fusion"},{"issue":"1","key":"e_1_3_3_38_2","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1109\/TAFFC.2023.3274829","article-title":"Efficient multimodal transformer with dual-level feature restoration for robust multimodal sentiment analysis","volume":"15","author":"Sun L","year":"2023","unstructured":"Sun L, Lian Z, Liu B, Tao J. Efficient multimodal transformer with dual-level feature restoration for robust multimodal sentiment analysis. IEEE Trans Affect Comput. 2023;15(1):309\u2013325.","journal-title":"IEEE Trans Affect Comput"},{"key":"e_1_3_3_39_2","doi-asserted-by":"crossref","unstructured":"Tsai YHH Bai S Liang PP Kolter JZ Morency LP Salakhutdinov R. Multimodal transformer for unaligned multimodal language sequences. In: Proceedings of the Conference. Association for Computational Linguistics. Meeting. NIH Public Access. Stroudsburg (PA): ACL; 2019. p. 6558.","DOI":"10.18653\/v1\/P19-1656"},{"key":"e_1_3_3_40_2","unstructured":"Shi B Hsu WN Lakhotia K Mohamed A. Learning audio-visual speech representation by masked multimodal cluster prediction. In: International Conference on Learning Representations. 2022."},{"key":"e_1_3_3_41_2","doi-asserted-by":"crossref","unstructured":"Guo Y Sun S Ma S Zheng K Bao X MA S Zou W Zheng Y. CrossMAE: Cross-modality masked autoencoders for region aware audio-visual pre-training. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Seattle (WA): IEEE; 2024. p. 26721\u201326731.","DOI":"10.1109\/CVPR52733.2024.02523"},{"key":"e_1_3_3_42_2","doi-asserted-by":"crossref","first-page":"1691","DOI":"10.1007\/s10489-023-05191-2","article-title":"Masked co-attention model for audio-visual event localization","volume":"54","author":"Liu H","year":"2024","unstructured":"Liu H, Gu X. Masked co-attention model for audio-visual event localization. Appl Intell. 2024;54:1691\u20131705.","journal-title":"Appl Intell"},{"key":"e_1_3_3_43_2","doi-asserted-by":"crossref","unstructured":"Woo J Ryu H Senocak A Chung JS. Speech guided masked image modeling for visually grounded speech. In: ICASSP 2024-2024 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). Seoul (Republic of Korea): IEEE; 2024. p. 8361\u20138365.","DOI":"10.1109\/ICASSP48485.2024.10447235"},{"key":"e_1_3_3_44_2","unstructured":"Huang PY Sharma V Xu H Ryali C Fan H Li Y Li SW Ghosh G Malik J Feichtenhofer C. MAViL: Masked audio-video learners. In: Advances in Neural Information Processing Systems. Cambridge (MA): MIT Press; 2024."},{"key":"e_1_3_3_45_2","doi-asserted-by":"crossref","unstructured":"Georgescu MI Fonseca E Ionescu RT Lucic M Schmid C Arnab A. Audiovisual masked autoencoders. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision. Vancouver (Canada): IEEE; 2023. p. 16144\u201316154.","DOI":"10.1109\/ICCV51070.2023.01479"},{"key":"e_1_3_3_46_2","unstructured":"Gong Y Rouditchenko A Liu AH Harwath D Karlinsky L Kuehne H Glass J. Contrastive audio-visual masked autoencoder. arXiv. 2022. https:\/\/doi.org\/10.48550\/arXiv.2210.07839"},{"key":"e_1_3_3_47_2","doi-asserted-by":"crossref","unstructured":"Mo S Morgado P. Unveiling the power of audio-visual early fusion transformers with dense interactions through masked modeling. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Seattle (WA): IEEE; 2024. p. 27186\u201327196.","DOI":"10.1109\/CVPR52733.2024.02567"},{"key":"e_1_3_3_48_2","unstructured":"Sadok S. Audiovisual speech representation learning applied to emotion recognition [thesis]. [Rennes (France)]: Institut d\u2019Electronique et des Technologies du numeRique; 2024."},{"key":"e_1_3_3_49_2","doi-asserted-by":"crossref","unstructured":"Xiang P Lin C Wu K Bai O. MultiMAE-DER: Multimodal masked autoencoder for dynamic emotion recognition. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2404.18327.","DOI":"10.1109\/ICPRS62101.2024.10677820"},{"key":"e_1_3_3_50_2","doi-asserted-by":"crossref","unstructured":"Singh M Duval Q Alwala KV Fan H Aggarwal V Adcock A Joulin A Dollar P Feichtenhofer C Girshick R et\u00a0al. The effectiveness of MAE pre-pretraining for billion-scale pretraining. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision. Paris (France): IEEE; 2023. p. 5484\u20135494.","DOI":"10.1109\/ICCV51070.2023.00505"},{"key":"e_1_3_3_51_2","doi-asserted-by":"crossref","unstructured":"Han Q Zhang G Huang J Gao P Wei Z Lu S. Efficient MAE towards large-scale vision transformers. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. Seattle (WA): IEEE; 2024. p. 606\u2013615.","DOI":"10.1109\/WACV57701.2024.00066"},{"key":"e_1_3_3_52_2","unstructured":"Feichtenhofer C Fan H Li Y He K. Masked autoencoders as spatiotemporal learners. In: Advances in Neural Information Processing Systems. Cambridge (MA): MIT Press; 2022. p. 35946\u201335958."},{"key":"e_1_3_3_53_2","doi-asserted-by":"crossref","unstructured":"Xie Z Zhang Z Cao Y Lin Y Wei Y Dai Q Hu H. On data scaling in masked image modeling. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Vancouver (Canada): IEEE; 2023. p. 10365\u201310374.","DOI":"10.1109\/CVPR52729.2023.00999"},{"key":"e_1_3_3_54_2","unstructured":"Dosovitskiy A Beyer L Kolesnikov A Weissenborn D Zhai X Unterthiner T Dehghani M Minderre M Heigold G Gelly S et\u00a0al. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR; 2021."},{"key":"e_1_3_3_55_2","doi-asserted-by":"crossref","first-page":"218","DOI":"10.1109\/TMM.2023.3263288","article-title":"MAR: Masked autoencoders for efficient action recognition","volume":"26","author":"Qing Z","year":"2023","unstructured":"Qing Z, Zhang S, Huang Z, Wang X, Wang Y, Lv Y, Gao C, Sang N. MAR: Masked autoencoders for efficient action recognition. IEEE Trans Multimed. 2023;26:218\u2013233.","journal-title":"IEEE Trans Multimed"},{"key":"e_1_3_3_56_2","unstructured":"Huang PY Xu H Li JB Baevski A Auli M Galuba W Metze F Feichtenhofer C. Masked autoencoders that listen. In: Advances in Neural Information Processing Systems. Cambridge (MA): MIT Press; 2022. p. 28708\u201328720."},{"key":"e_1_3_3_57_2","doi-asserted-by":"crossref","unstructured":"Cheng Y Wang R Pan Z Feng R Zhang Y. Look listen and attend: Co-attention network for self-supervised audio-visual representation learning. In: Proceedings of the 28th ACM International Conference on Multimedia. Seattle (WA): ACM; 2020. p. 3884\u20133892.","DOI":"10.1145\/3394171.3413869"},{"key":"e_1_3_3_58_2","doi-asserted-by":"crossref","unstructured":"Hu Y Li R Chen C Zou H Zhu Q Chng ES. Cross-modal global interaction and local alignment for audio-visual speech recognition. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2305.09212.","DOI":"10.24963\/ijcai.2023\/564"},{"key":"e_1_3_3_59_2","doi-asserted-by":"crossref","unstructured":"Fan Y Kang J Li L Li KC Chen HL Cheng ST Zhang PY Zhou ZY Cai YQ Wang D. CN-CELEB: A challenging Chinese speaker recognition dataset. In: ICASSP 2020-2020 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). Barcelona (Spain): IEEE; 2020. p. 7604\u20137608.","DOI":"10.1109\/ICASSP40776.2020.9054017"},{"key":"e_1_3_3_60_2","doi-asserted-by":"crossref","unstructured":"Lian Z Sun H Sun L Wen Z Zhang S Chen S Gu H Zhoa J Ma Z Chen X et\u00a0al. MER 2024: Semi-supervised learning noise robustness and open-vocabulary multimodal emotion recognition. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2404.17113","DOI":"10.1145\/3689092.3689959"},{"key":"e_1_3_3_61_2","doi-asserted-by":"crossref","unstructured":"Chung JS Nagrani A Zisserman A. VoxCeleb2: Deep speaker recognition. arXiv. 2018. https:\/\/doi.org\/10.48550\/arXiv.1806.05622","DOI":"10.21437\/Interspeech.2018-1929"},{"key":"e_1_3_3_62_2","doi-asserted-by":"crossref","unstructured":"Ephrat A Mosseri I Lang O Dekel T Wilson K Hassidim A Freeman WT Rubinstein M. Looking to listen at the cocktail party: A speaker independent audio-visual model for speech separation. arXiv. 2018. https:\/\/doi.org\/10.48550\/arXiv.1804.03619","DOI":"10.1145\/3197517.3201357"},{"key":"e_1_3_3_63_2","doi-asserted-by":"crossref","unstructured":"Zhu H Wu W Zhu W Jiang L Tang S Zhang L Liu Z Loy CC. CelebV-HQ: A large-scale video facial attributes dataset. In: European Conference on Computer Vision. Berlin (Germany): Springer; 2022. p. 650\u2013667.","DOI":"10.1007\/978-3-031-20071-7_38"},{"key":"e_1_3_3_64_2","doi-asserted-by":"crossref","unstructured":"Chen C Wang D Zheng TF. CN-CVS: A mandarin audio-visual dataset for large vocabulary continuous visual to speech synthesis. In: ICASSP 2023-2023 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). Rhodes Island (Greece): IEEE; 2023. p. 1\u20135.","DOI":"10.1109\/ICASSP49357.2023.10095796"},{"key":"e_1_3_3_65_2","doi-asserted-by":"crossref","unstructured":"Chung JS Zisserman A. Out of time: Automated lip sync in the wild. In: Workshop on Multi-view Lip-reading. ACCV; Berlin (Germany): Springer; 2016.","DOI":"10.1007\/978-3-319-54427-4_19"},{"issue":"1","key":"e_1_3_3_66_2","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1109\/TAFFC.2016.2515617","article-title":"MSP-IMPROV: An acted corpus of dyadic interactions to study emotion perception","volume":"8","author":"Busso C","year":"2016","unstructured":"Busso C, Parthasarathy S, Burmania A, AbdelWahab M, Sadoughi N, Provost EM. MSP-IMPROV: An acted corpus of dyadic interactions to study emotion perception. IEEE Trans Affect Comput. 2016;8(1):67\u201380.","journal-title":"IEEE Trans Affect Comput"},{"key":"e_1_3_3_67_2","doi-asserted-by":"crossref","unstructured":"Jiang X Zong Y Zheng W Tang C Xia W Lu C Liu J. DFEW: A large-scale database for recognizing dynamic facial expressions in the wild. In: Proceedings of the 28th ACM International Conference on Multimedia. Seattle (WA): ACM; 2020. p. 2881\u20132889.","DOI":"10.1145\/3394171.3413620"},{"key":"e_1_3_3_68_2","doi-asserted-by":"crossref","unstructured":"Lian Z Sun H Sun L Chen K Xu M Wang K Xu K He Y Li Y Zho J et\u00a0al. MER 2023: Multi-label learning modality robustness and semisupervised learning. In: Proceedings of the 31st ACM International Conference on Multimedia. Ottawa (ON Canada): ACM; 2023. p. 9610\u20139164.","DOI":"10.1145\/3581783.3612836"},{"key":"e_1_3_3_69_2","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1007\/s10579-008-9076-6","article-title":"IEMOCAP: Interactive emotional dyadic motion capture database","volume":"42","author":"Busso C","year":"2008","unstructured":"Busso C, Bulut M, Lee CC, Kazemzadeh A, Mower E, Kim S, Chang JN, Lee S, Narayanan SS. IEMOCAP: Interactive emotional dyadic motion capture database. Lang Resour Eval. 2008;42:335\u2013359.","journal-title":"Lang Resour Eval"},{"issue":"4","key":"e_1_3_3_70_2","doi-asserted-by":"crossref","first-page":"377","DOI":"10.1109\/TAFFC.2014.2336244","article-title":"CREMA-D: Crowd-sourced emotional multimodal actors dataset","volume":"5","author":"Cao H","year":"2014","unstructured":"Cao H, Cooper DG, Keutmann MK, Gur RC, Nenkova A, Verma R. CREMA-D: Crowd-sourced emotional multimodal actors dataset. IEEE Trans Affect Comput. 2014;5(4):377\u2013390.","journal-title":"IEEE Trans Affect Comput"},{"key":"e_1_3_3_71_2","doi-asserted-by":"crossref","unstructured":"Tran M Kim Y Su CC Kuo CH Soleymani M. SAAML: A framework for semi-supervised affective adaptation via metric learning. In: Proceedings of the 31st ACM International Conference on Multimedia. Ottawa (ON Canada): ACM; 2023. p. 6004\u20136015.","DOI":"10.1145\/3581783.3612286"},{"issue":"5","key":"e_1_3_3_72_2","doi-asserted-by":"crossref","DOI":"10.1371\/journal.pone.0196391","article-title":"The Ryerson audio-visual database of emotional speech and song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English","volume":"13","author":"Livingstone SR","year":"2018","unstructured":"Livingstone SR, Russo FA. The Ryerson audio-visual database of emotional speech and song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English. PLOS ONE. 2018;13(5): Article e0196391.","journal-title":"PLOS ONE"},{"issue":"1","key":"e_1_3_3_73_2","first-page":"76","article-title":"AVCAffe: A large scale audio-visual dataset of cognitive load and affect for remote work","volume":"37","author":"Sarkar P","year":"2023","unstructured":"Sarkar P, Posen A, Etemad A. AVCAffe: A large scale audio-visual dataset of cognitive load and affect for remote work. Proc AAAI Conf Artif Intell. 2023;37(1):76\u201385.","journal-title":"Proc AAAI Conf Artif Intell"},{"issue":"2","key":"e_1_3_3_74_2","doi-asserted-by":"crossref","first-page":"1201","DOI":"10.1109\/TAFFC.2021.3101563","article-title":"Werewolf-XL: A database for identifying spontaneous affect in large competitive group interactions","volume":"14","author":"Zhang K","year":"2021","unstructured":"Zhang K, Wu X, Xie X, Zhang X, Zhang H, Chen X, Sun L. Werewolf-XL: A database for identifying spontaneous affect in large competitive group interactions. IEEE Trans Affect Comput. 2021;14(2):1201\u20131214.","journal-title":"IEEE Trans Affect Comput"},{"issue":"1","key":"e_1_3_3_75_2","doi-asserted-by":"crossref","DOI":"10.1371\/journal.pone.0086041","article-title":"CASME II: An improved spontaneous micro-expression database and the baseline evaluation","volume":"9","author":"Yan WJ","year":"2014","unstructured":"Yan WJ, Li X, Wang SJ, Zhoa G, Liu JY, Chen YH, Fu X. CASME II: An improved spontaneous micro-expression database and the baseline evaluation. PLOS ONE. 2014;9(1): Article e86041.","journal-title":"PLOS ONE"},{"issue":"1","key":"e_1_3_3_76_2","doi-asserted-by":"crossref","first-page":"116","DOI":"10.1109\/TAFFC.2016.2573832","article-title":"SAMM: A spontaneous micro-facial movement dataset","volume":"9","author":"Davison AK","year":"2016","unstructured":"Davison AK, Lansley C, Costen N, Tan K, Yap MH. SAMM: A spontaneous micro-facial movement dataset. IEEE Trans Affect Comput. 2016;9(1):116\u2013129.","journal-title":"IEEE Trans Affect Comput"},{"key":"e_1_3_3_77_2","doi-asserted-by":"crossref","unstructured":"Li X Pfister T Huang X Zhao G Pietikainen M. A spontaneous micro-expression database: Inducement collection and baseline. In: 2013 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG). Piscataway (NJ): IEEE; 2013. p. 1\u20136.","DOI":"10.1109\/FG.2013.6553717"},{"issue":"3","key":"e_1_3_3_78_2","first-page":"2782","article-title":"CAS(ME)3: A third generation facial spontaneous micro-expression database with depth information and high ecological validity","volume":"45","author":"Li J","year":"2022","unstructured":"Li J, Dong Z, Lu S, Wang SJ, Yan WJ, Ma Y. CAS(ME)3: A third generation facial spontaneous micro-expression database with depth information and high ecological validity. IEEE Trans Pattern Anal Mach Intell. 2022;45(3):2782\u20132800.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"9","key":"e_1_3_3_79_2","first-page":"5826","article-title":"Video-based facial micro-expression analysis: A survey of datasets, features and algorithms","volume":"44","author":"Ben X","year":"2021","unstructured":"Ben X, Ren Y, Zhang J, Wang SJ, Kpalma K, Meng W, Liu YJ. Video-based facial micro-expression analysis: A survey of datasets, features and algorithms. IEEE Trans Pattern Anal Mach Intell. 2021;44(9):5826\u20135846.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"e_1_3_3_80_2","doi-asserted-by":"crossref","unstructured":"Nguyen XB Duong CN Li X Gauch S Seo HS Luu K. Micron-BERT: BERT-based facial micro-expression recognition. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Vancouver (BC Canada): IEEE; 2023. p. 1482\u20131492.","DOI":"10.1109\/CVPR52729.2023.00149"},{"key":"e_1_3_3_81_2","doi-asserted-by":"crossref","DOI":"10.1016\/j.patcog.2021.108275","article-title":"Feature refinement: An expression-specific feature learning and fusion method for micro-expression recognition","volume":"122","author":"Zhou L","year":"2022","unstructured":"Zhou L, Mao Q, Huang X, Zhang F, Zhang Z. Feature refinement: An expression-specific feature learning and fusion method for micro-expression recognition. Pattern Recogn. 2022;122: Article 108275.","journal-title":"Pattern Recogn"},{"key":"e_1_3_3_82_2","unstructured":"Chen Y Li J Zhang Y Hu Z Shan S Wang M Hong R. UniLearn: Enhancing dynamic facial expression recognition through unified pre-training and fine-tuning on images and videos. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2409.06154"},{"key":"e_1_3_3_83_2","doi-asserted-by":"crossref","DOI":"10.1016\/j.neucom.2024.128196","article-title":"HTNet for micro-expression recognition","volume":"602","author":"Wang Z","year":"2024","unstructured":"Wang Z, Zhang K, Luo W, Sankaranarayana R. HTNet for micro-expression recognition. Neurocomputing. 2024;602: Article 128196.","journal-title":"Neurocomputing"},{"key":"e_1_3_3_84_2","unstructured":"Tu S Pan Y Huang Y Han X Xing Z Dai Q Luo C Wu Z Jiang YG. StableAvatar: Infinite-length audio-driven avatar video generation. arXiv. 2025. https:\/\/doi.org\/10.48550\/arXiv.2508.08248"},{"key":"e_1_3_3_85_2","unstructured":"Loshchilov I. Decoupled weight decay regularization. arXiv. 2017. https:\/\/doi.org\/10.48550\/arXiv.1711.05101"},{"key":"e_1_3_3_86_2","doi-asserted-by":"crossref","unstructured":"Hoffer E Ben-Nun T Hubara I Giladi N Hoefler T Soudry D. Augment your batch: Improving generalization through instance repetition. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Seattle (WA): IEEE; 2020. p. 8129\u20138138.","DOI":"10.1109\/CVPR42600.2020.00815"},{"key":"e_1_3_3_87_2","unstructured":"Zhang H. mixup: Beyond empirical risk minimization. arXiv. 2017. https:\/\/doi.org\/10.48550\/arXiv.1710.09412"},{"key":"e_1_3_3_88_2","doi-asserted-by":"crossref","unstructured":"Cubuk ED Zoph B Shlens J Le QV. Randaugment: Practical automated data augmentation with a reduced search space. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops. Seattle (WA): IEEE; 2020. p. 702\u2013703.","DOI":"10.1109\/CVPRW50498.2020.00359"},{"key":"e_1_3_3_89_2","doi-asserted-by":"crossref","unstructured":"Szegedy C Vanhoucke V Ioffe S Shlens J Wojna Z. Rethinking the inception architecture for computer vision. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway (NJ): IEEE; 2016. p. 2818\u20132826.","DOI":"10.1109\/CVPR.2016.308"},{"key":"e_1_3_3_90_2","first-page":"12449","article-title":"wav2vec 2.0: A framework for self-supervised learning of speech representations","volume":"33","author":"Baevski A","year":"2020","unstructured":"Baevski A, Zhou Y, Mohammed A, Auli M. wav2vec 2.0: A framework for self-supervised learning of speech representations. Adv Neural Inf Proces Syst. 2020;33:12449\u201312460.","journal-title":"Adv Neural Inf Proces Syst"},{"key":"e_1_3_3_91_2","doi-asserted-by":"crossref","first-page":"3451","DOI":"10.1109\/TASLP.2021.3122291","article-title":"HuBERT: Self supervised speech representation learning by masked prediction of hidden units","volume":"29","author":"Hsu WN","year":"2021","unstructured":"Hsu WN, Bolte B, Tsai YHH, Lakhotia K, Salakhutdinov R, Mohamed A. HuBERT: Self supervised speech representation learning by masked prediction of hidden units. IEEE\/ACM Trans Audio Speech Lang Process. 2021;29:3451\u20133460.","journal-title":"IEEE\/ACM Trans Audio Speech Lang Process"},{"issue":"6","key":"e_1_3_3_92_2","first-page":"1505","article-title":"WavLM: Large-scale self-supervised pre-training for full stack speech processing","volume":"16","author":"Chen S","year":"2022","unstructured":"Chen S, Wang C, Chen Z, Wu Y, Liu S, Chen Z. WavLM: Large-scale self-supervised pre-training for full stack speech processing. IEEE\/ACM Trans Audio Speech Lang Process. 2022;16(6):1505\u20131518.","journal-title":"IEEE\/ACM Trans Audio Speech Lang Process"},{"key":"e_1_3_3_93_2","doi-asserted-by":"crossref","unstructured":"He K Zhang X Ren S Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas (NV): IEEE; 2016. p. 770\u2013778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_3_94_2","doi-asserted-by":"crossref","unstructured":"Chen Y Li J Shan S Wang M Hong R. From static to dynamic: Adapting landmark aware image models for facial expression recognition in videos. In: IEEE Transactions on Affective Computing. Piscataway (NJ): IEEE; 2024.","DOI":"10.1109\/TAFFC.2024.3453443"},{"key":"e_1_3_3_95_2","doi-asserted-by":"crossref","unstructured":"Tran D Bourdev L Fergus R Torresani L Paluri M. Learning spatiotemporal features with 3D convolutional networks. In: Proceedings of the IEEE International Conference on Computer Vision. Boston (MA): IEEE; 2015. p. 4489\u20134497.","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_3_96_2","doi-asserted-by":"crossref","unstructured":"Gowda SN Gao B Clifton DA. FE-Adapter: Adapting image-based emotion classifiers to videos. In: 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG). Istanbul (Turkiye): IEEE; 2024. p. 1\u20136.","DOI":"10.1109\/FG59268.2024.10581905"},{"key":"e_1_3_3_97_2","doi-asserted-by":"crossref","unstructured":"Zhao Z Liu Q. Former-DFER: Dynamic facial expression recognition transformer. In: Proceedings of the 29th ACM International Conference on Multimedia. China: ACM; 2021. p. 1553\u20131561.","DOI":"10.1145\/3474085.3475292"},{"key":"e_1_3_3_98_2","doi-asserted-by":"crossref","unstructured":"Foteinopoulou NM Patras I. EmoCLIP: A vision-language method for zero-shot video facial expression recognition. In: 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG). Istanbul (Turkiye): IEEE; 2024. p. 1\u201310.","DOI":"10.1109\/FG59268.2024.10581982"},{"key":"e_1_3_3_99_2","unstructured":"Tao Z Wang Y Lin J Wang H Mai X Tong X Zhou Z Yan S Zhao Q Han L et\u00a0al. A3lign-DFER: Pioneering comprehensive dynamic affective alignment for dynamic facial expression recognition with CLIP. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2403.04294"},{"key":"e_1_3_3_100_2","doi-asserted-by":"crossref","unstructured":"Yoon S Dey S Lee H Jung K. Attentive modality hopping mechanism for speech emotion recognition. In: ICASSP 2020-2020 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). Barcelona (Spain): IEEE; 2020. p. 3362\u20133366.","DOI":"10.1109\/ICASSP40776.2020.9054229"},{"key":"e_1_3_3_101_2","doi-asserted-by":"crossref","unstructured":"Chen H Huang H Dong J Zheng M Shao D. FineCLIPER: Multi-modal fine-grained clip for dynamic facial expression recognition with adapters. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2407.02157","DOI":"10.1145\/3664647.3680827"},{"key":"e_1_3_3_102_2","doi-asserted-by":"crossref","unstructured":"Yoon S Byun S Jung K. Multimodal speech emotion recognition using audio and text. In: 2018 IEEE Spoken Language Technology Workshop (SLT). Athens (Greece): IEEE; 2018. p. 112\u2013118.","DOI":"10.1109\/SLT.2018.8639583"},{"key":"e_1_3_3_103_2","doi-asserted-by":"crossref","unstructured":"Rajan V Brutti A Cavallaro A. Is cross-attention preferable to self-attention for multimodal emotion recognition? In: ICASSP 2022-2022 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). Singapore: IEEE; 2022. p. 4693\u20134697.","DOI":"10.1109\/ICASSP43922.2022.9746924"},{"key":"e_1_3_3_104_2","doi-asserted-by":"crossref","unstructured":"Goncalves L Busso C. AuxFormer: Robust approach to audiovisual emotion recognition. In: ICASSP 2022-2022 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). Singapore: IEEE; 2022. p. 7357\u20137361.","DOI":"10.1109\/ICASSP43922.2022.9747157"},{"key":"e_1_3_3_105_2","doi-asserted-by":"crossref","unstructured":"Keesing A Koh YS Yogarajan V Witbrock M. Emotion recognition toolkit (ERTK): Standardising tools for emotion recognition research. In: Proceedings of the 31st ACM International Conference on Multimedia. Ottawa (ON Canada): IEEE; 2023. p. 9693\u20139696.","DOI":"10.1145\/3581783.3613459"},{"issue":"2","key":"e_1_3_3_106_2","doi-asserted-by":"crossref","first-page":"190","DOI":"10.1109\/TAFFC.2015.2457417","article-title":"The Geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing","volume":"7","author":"Eyben F","year":"2015","unstructured":"Eyben F, Scherer KR, Schuller BW, Sundberg J, Andre E, Busso C, Devillers LY, Epps J, Laukka P, Narayanan SS, et al. The Geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing. IEEE Trans Affect Comput. 2015;7(2):190\u2013202.","journal-title":"IEEE Trans Affect Comput"},{"key":"e_1_3_3_107_2","doi-asserted-by":"crossref","unstructured":"Ghaleb E Popa M Asteriadis S. Multimodal and temporal perception of audio-visual cues for emotion recognition. In: 2019 8th International Conference on Affective Computing and Intelligent Interaction (ACII). Cambridge (UK): IEEE; 2019. p. 552\u2013558.","DOI":"10.1109\/ACII.2019.8925444"},{"issue":"4","key":"e_1_3_3_108_2","doi-asserted-by":"crossref","first-page":"2156","DOI":"10.1109\/TAFFC.2022.3216993","article-title":"Robust audiovisual emotion recognition: Aligning modalities, capturing temporal information, and handling missing features","volume":"13","author":"Goncalves L","year":"2022","unstructured":"Goncalves L, Busso C. Robust audiovisual emotion recognition: Aligning modalities, capturing temporal information, and handling missing features. IEEE Trans Affect Comput. 2022;13(4):2156\u20132170.","journal-title":"IEEE Trans Affect Comput"},{"issue":"4","key":"e_1_3_3_109_2","doi-asserted-by":"crossref","first-page":"2954","DOI":"10.1109\/TAFFC.2023.3234777","article-title":"Audio-visual emotion recognition with preference learning based on intended and multi-modal perceived labels","volume":"14","author":"Lei Y","year":"2023","unstructured":"Lei Y, Cao H. Audio-visual emotion recognition with preference learning based on intended and multi-modal perceived labels. IEEE Trans Affect Comput. 2023;14(4):2954\u20132969.","journal-title":"IEEE Trans Affect Comput"},{"key":"e_1_3_3_110_2","doi-asserted-by":"crossref","unstructured":"Tran M Soleymani M. A pre-trained audio-visual transformer for emotion recognition. In: ICASSP 2022-2022 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). Singapore: IEEE; 2022. p. 4698\u20134702.","DOI":"10.1109\/ICASSP43922.2022.9747278"},{"key":"e_1_3_3_111_2","doi-asserted-by":"crossref","unstructured":"Zadeh A Chen M Poria S Cambria E. Morency LP. Tensor fusion network for multimodal sentiment analysis. arXiv. 2017. https:\/\/doi.org\/10.48550\/arXiv.1707.07250.","DOI":"10.18653\/v1\/D17-1115"},{"key":"e_1_3_3_112_2","doi-asserted-by":"crossref","unstructured":"Ghaleb E Niehues J Asteriadis S. Multimodal attention-mechanism for temporal emotion recognition. In: 2020 IEEE International Conference on Image Processing (ICIP). Abu Dhabi (United Arab Emirates): IEEE; 2020. p. 251\u2013255.","DOI":"10.1109\/ICIP40778.2020.9191019"},{"key":"e_1_3_3_113_2","doi-asserted-by":"crossref","unstructured":"Goncalves L Busso C. Learning cross-modal audiovisual representations with ladder networks for emotion recognition. In: ICASSP 2023-2023 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). Rhodes Island (Greece): IEEE; 2023. p. 1\u20135.","DOI":"10.1109\/ICASSP49357.2023.10096138"},{"key":"e_1_3_3_114_2","doi-asserted-by":"crossref","unstructured":"Tran D Wang H Torresani L Ray J LeCun Y Paluri M. A closer look at spatiotemporal convolutions for action recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Salt Lake City (UT): IEEE; 2018. p. 6450\u20136459.","DOI":"10.1109\/CVPR.2018.00675"},{"key":"e_1_3_3_115_2","doi-asserted-by":"crossref","unstructured":"Hara K Kataoka H Satoh Y. Can spatiotemporal 3D CNNs retrace the history of 2D CNNs and ImageNet? In: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition. Salt Lake City (UT): IEEE; 2018. p. 6546\u20136555.","DOI":"10.1109\/CVPR.2018.00685"},{"key":"e_1_3_3_116_2","doi-asserted-by":"crossref","first-page":"182","DOI":"10.1016\/j.ins.2022.03.062","article-title":"Clip-aware expressive feature learning for video-based facial expression recognition","volume":"598","author":"Liu Y","year":"2022","unstructured":"Liu Y, Feng C, Yuan X, Zhou L, Wang W, Qin J, Luo Z. Clip-aware expressive feature learning for video-based facial expression recognition. Inf Sci. 2022;598:182\u2013195.","journal-title":"Inf Sci"},{"key":"e_1_3_3_117_2","doi-asserted-by":"crossref","DOI":"10.1016\/j.patcog.2023.109368","article-title":"Expression snippet transformer for robust video-based facial expression recognition","volume":"138","author":"Liu Y","year":"2023","unstructured":"Liu Y, Wang W, Feng C, Zhang H, Chen Z, Zhan Y. Expression snippet transformer for robust video-based facial expression recognition. Pattern Recogn. 2023;138: Article 109368.","journal-title":"Pattern Recogn"},{"key":"e_1_3_3_118_2","unstructured":"Ma F Sun B Li S. Spatio-temporal transformer for dynamic facial expression recognition in the wild. arXiv. 2022. https:\/\/doi.org\/10.48550\/arXiv.2205.04749"},{"key":"e_1_3_3_119_2","unstructured":"Li H Sui M Zhu Z Zhao F. NR-DFERNet: Noise-robust network for dynamic facial expression recognition. arXiv. 2022. https:\/\/doi.org\/10.48550\/arXiv.2206.04975"},{"key":"e_1_3_3_120_2","doi-asserted-by":"crossref","unstructured":"Wang Y Sun Y Song W Gao S Huang H Chen Z Ge W Zhang W. DPCNet: Dual path multi-excitation collaborative network for facial expression representation learning in videos. In: Proceedings of the 30th ACM International Conference on Multimedia. Lisbon (Portugal): ACM; 2022. p. 101\u2013110.","DOI":"10.1145\/3503161.3547865"},{"issue":"1","key":"e_1_3_3_121_2","first-page":"67","article-title":"Intensity-aware loss for dynamic facial expression recognition in the wild","volume":"37","author":"Li H","year":"2023","unstructured":"Li H, Niu H, Zhu Z, Zhao F. Intensity-aware loss for dynamic facial expression recognition in the wild. Proc AAAI Conf Artif Intell. 2023;37(1):67\u201375.","journal-title":"Proc AAAI Conf Artif Intell"},{"key":"e_1_3_3_122_2","doi-asserted-by":"crossref","unstructured":"Wang H Li B Wu S Shen S Liu F Ding S Zhou A. Rethinking the learning paradigm for dynamic facial expression recognition. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Vancouver (BC Canada): IEEE; 2023. p. 17958\u201317968.","DOI":"10.1109\/CVPR52729.2023.01722"},{"key":"e_1_3_3_123_2","doi-asserted-by":"crossref","unstructured":"Li H Niu H Zhu Z Zhao F. CLIPER: A unified vision-language framework for in-the-wild facial expression recognition. In: 2024 IEEE International Conference on Multimedia and Expo (ICME). Niagara Falls (Canada): IEEE; 2024. p. 1\u20136.","DOI":"10.1109\/ICME57554.2024.10687508"},{"key":"e_1_3_3_124_2","doi-asserted-by":"crossref","unstructured":"He K Fan H Wu Y Xie S Girshick R. Momentum contrast for unsupervised visual representation learning. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Piscataway (NJ): IEEE; 2020. p. 9729\u20139738.","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_3_125_2","doi-asserted-by":"crossref","unstructured":"Tseng Y Berry L Chen YT Chiu IH Lin HH Liu M Peng P Shih YJ Wang HY Wu H et\u00a0al. AV-SUPERB: A multi-task evaluation benchmark for audio visual representation models. In: ICASSP 2024-2024 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). Seoul (Korea): IEEE; 2024. p. 6890\u20136894.","DOI":"10.1109\/ICASSP48485.2024.10445941"},{"key":"e_1_3_3_126_2","unstructured":"Mittal H Morgado P Jain U Gupta A. Learning state-aware visual representations from audible interactions. In: Advances in Neural Information Processing Systems. Cambridge (MA): MIT Press; 2022. p. 23765\u201323779."},{"key":"e_1_3_3_127_2","unstructured":"Lee S Yu Y Kim G Breuel T Kautz J Song Y. Parameter efficient multimodal transformers for video representation learning. In: 9th International Conference on Learning Representations ICLR 2021. 2021."},{"key":"e_1_3_3_128_2","doi-asserted-by":"crossref","unstructured":"Dalal N Triggs B. Histograms of oriented gradients for human detection. In: 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905). San Diego (CA): IEEE; 2005. p. 886\u2013893.","DOI":"10.1109\/CVPR.2005.177"},{"key":"e_1_3_3_129_2","unstructured":"Radford A Kim JW Xu T Brockman G McLeavey C Sutskever I. Robust speech recognition via large-scale weak supervision. In: International Conference on Machine Learning. New York (NY): PMLR; 2023. p. 28492\u201328518."},{"key":"e_1_3_3_130_2","doi-asserted-by":"crossref","unstructured":"Hershey S Chaudhuri S Ellis DP Gemmeke JF Jansen A Moore RC Plakal M Platt D Saurous RA Seybold B et\u00a0al. CNN architectures for large-scale audio classification. In: 2017 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). New Orleans (LA): IEEE; 2017. p. 131\u2013135.","DOI":"10.1109\/ICASSP.2017.7952132"},{"key":"e_1_3_3_131_2","doi-asserted-by":"crossref","unstructured":"Ma Z Zheng Z Ye J Li J Gao Z Zhang S Chen X. emotion2vec: Self-supervised pre-training for speech emotion representation. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2312.15185","DOI":"10.18653\/v1\/2024.findings-acl.931"},{"key":"e_1_3_3_132_2","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1007\/s12193-015-0195-2","article-title":"EmoNets: Multimodal deep learning approaches for emotion recognition in video","volume":"10","author":"Kahou SE","year":"2016","unstructured":"Kahou SE, Bouthillier X, Lamblin P, Gulcehre C, Michalski C, Konda K, Jean S, Froumenty P, Dauphin Y, Baoulanger-Lewandowski N, et al. EmoNets: Multimodal deep learning approaches for emotion recognition in video. J Multimodal User Interfaces. 2016;10:99\u2013111.","journal-title":"J Multimodal User Interfaces"},{"key":"e_1_3_3_133_2","doi-asserted-by":"crossref","unstructured":"Hu J Shen L Sun G. Squeeze-and-excitation networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Salt Lake City (UT): IEEE; 2018. p. 7132\u20131741.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_3_3_134_2","doi-asserted-by":"crossref","first-page":"6544","DOI":"10.1109\/TIP.2021.3093397","article-title":"Learning deep global multi-scale and local attention features for facial expression recognition in the wild","volume":"30","author":"Zhao Z","year":"2021","unstructured":"Zhao Z, Liu Q, Wang S. Learning deep global multi-scale and local attention features for facial expression recognition in the wild. IEEE Trans Image Process. 2021;30:6544\u20136556.","journal-title":"IEEE Trans Image Process"},{"key":"e_1_3_3_135_2","unstructured":"Radford A Kim JW Hallacy C Ramesh A Goh G Agarwal S Sastry G Askell A Mishkin P Clark J et\u00a0al. Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning. New York (NY): PMLR; 2021. p. 8748\u20138763."},{"key":"e_1_3_3_136_2","doi-asserted-by":"crossref","DOI":"10.1016\/j.imavis.2024.105171","article-title":"EVA-02: A visual representation for neon genesis","volume":"149","author":"Fang Y","year":"2024","unstructured":"Fang Y, Sun Q, Wang X, Huang T, Wang X, Cao Y. EVA-02: A visual representation for neon genesis. Image Vis Comput. 2024;149: Article 105171.","journal-title":"Image Vis Comput"},{"key":"e_1_3_3_137_2","unstructured":"Oquab M Darcet T Moutakanni T Vo H Szafraneic M Khalidov V Fernandez P Haziza D Massa F El-Nouby A et\u00a0al. DUBIv2: Learning robust visual features without supervision. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2304.07193"},{"key":"e_1_3_3_138_2","unstructured":"Su L Hu C Li G Cao D. MSAF: Multimodal split attention fusion. arXiv. 2020. https:\/\/doi.org\/10.48550\/arXiv.2012.07175"},{"key":"e_1_3_3_139_2","doi-asserted-by":"crossref","unstructured":"Fukui A Park DH Yang D Rohrbach A Darrell T Rohrbach M. Multimodal compact bilinear pooling for visual question answering and visual grounding. In: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. Stroudsburg (PA): ACL; 2016. p. 457\u2013468.","DOI":"10.18653\/v1\/D16-1044"},{"key":"e_1_3_3_140_2","unstructured":"Joze HRV Shaban A Iuzzolino ML Koishida K. MMTM: Multimodal transfer module for CNN fusion. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Piscataway (NJ): IEEE; 2020. p. 13289\u201313299."},{"key":"e_1_3_3_141_2","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1016\/j.patrec.2022.07.012","article-title":"ERANNs: Efficient residual audio neural networks for audio pattern recognition","volume":"161","author":"Verbitskiy S","year":"2022","unstructured":"Verbitskiy S, Berikov V, Vyshegorodtsev V. ERANNs: Efficient residual audio neural networks for audio pattern recognition. Pattern Recogn Lett. 2022;161:38\u201344.","journal-title":"Pattern Recogn Lett"},{"key":"e_1_3_3_142_2","unstructured":"Fu Z Liu F Wang H Qi J Fu X Zhou A Li Z A cross-modal fusion network based on self-attention and residual structure for multimodal emotion recognition. arXiv. 2021. https:\/\/doi.org\/10.48550\/arXiv.2111.02172"},{"key":"e_1_3_3_143_2","doi-asserted-by":"crossref","unstructured":"Chumachenko K Iosifidis A Gabbouj M. Self-attention fusion for audiovisual emotion recognition with incomplete data. In: 2022 26th International Conference on Pattern Recognition (ICPR). Montreal (QC Canada): IEEE; 2022. p. 2822\u20132828.","DOI":"10.1109\/ICPR56361.2022.9956592"},{"key":"e_1_3_3_144_2","doi-asserted-by":"crossref","unstructured":"Zhang B Lv H Guo P Shao Q Chao Y Xie L. WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition. In: ICASSP 2022-2022 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP). Singapore: IEEE; 2022. p. 6182\u20136186.","DOI":"10.1109\/ICASSP43922.2022.9746682"},{"key":"e_1_3_3_145_2","doi-asserted-by":"crossref","unstructured":"Parkhi O Vedaldi A Zisserman A. Deep face recognition. In: BMVC 2015-Proceedings of the British Machine Vision Conference 2015. Swansea (UK): British Machine Vision Association; 2015.","DOI":"10.5244\/C.29.41"},{"issue":"6","key":"e_1_3_3_146_2","doi-asserted-by":"crossref","first-page":"915","DOI":"10.1109\/TPAMI.2007.1110","article-title":"Dynamic texture recognition using local binary patterns with an application to facial expressions","volume":"29","author":"Zhao G","year":"2007","unstructured":"Zhao G, Pietikainen M. Dynamic texture recognition using local binary patterns with an application to facial expressions. IEEE Trans Pattern Anal Mach Intell. 2007;29(6):915\u2013928.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"e_1_3_3_147_2","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1016\/j.image.2017.11.006","article-title":"Less is more: Micro-expression recognition from video using apex frame","volume":"62","author":"Liong ST","year":"2018","unstructured":"Liong ST, See J, Wong K, Phan RCW. Less is more: Micro-expression recognition from video using apex frame. Signal Process Image Commun. 2018;62:82\u201392.","journal-title":"Signal Process Image Commun"},{"key":"e_1_3_3_148_2","doi-asserted-by":"crossref","unstructured":"Zhang H Zhang H. A review of micro-expression recognition based on deep learning. In: 2022 International Joint Conference on Neural Networks (IJCNN). Padua (Italy): IEEE; 2022. p. 1\u20138.","DOI":"10.1109\/IJCNN55064.2022.9892307"},{"issue":"1","key":"e_1_3_3_149_2","article-title":"On the performance of GoogLeNet and AlexNet applied to sketches","volume":"30","author":"Ballester P","year":"2016","unstructured":"Ballester P, Araujo R. On the performance of GoogLeNet and AlexNet applied to sketches. Proc AAAI Conf Artif Intell. 2016;30(1).","journal-title":"Proc AAAI Conf Artif Intell"},{"key":"e_1_3_3_150_2","doi-asserted-by":"crossref","unstructured":"Liong ST Gan YS See J Khor HQ Huang YC. Shallow triple stream three-dimensional CNN (STSTNet) for micro-expression recognition. In: 2019 14th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2019). Lille (France): IEEE; 2019. p. 1\u20135.","DOI":"10.1109\/FG.2019.8756567"},{"key":"e_1_3_3_151_2","doi-asserted-by":"publisher","DOI":"10.3389\/fnins.2019.00095"},{"key":"e_1_3_3_152_2","doi-asserted-by":"crossref","unstructured":"Van Quang N Chun J Tokuyama T. CapsuleNet for micro-expression recognition. In: 2019 14th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2019). Lille (France): IEEE; 2019. p. 1\u20137.","DOI":"10.1109\/FG.2019.8756544"},{"key":"e_1_3_3_153_2","doi-asserted-by":"crossref","first-page":"8590","DOI":"10.1109\/TIP.2020.3018222","article-title":"Revealing the invisible with model and data shrinking for composite-database micro-expression recognition","volume":"29","author":"Xia Z","year":"2020","unstructured":"Xia Z, Peng W, Khor HQ, Feng X, Zhao G. Revealing the invisible with model and data shrinking for composite-database micro-expression recognition. IEEE Trans Image Process. 2020;29:8590\u20138605.","journal-title":"IEEE Trans Image Process"},{"key":"e_1_3_3_154_2","doi-asserted-by":"crossref","unstructured":"Liu Y Du H Zheng L Gedeon T. A neural micro-expression recognizer. In: 2019 14th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2019). Lille (France): IEEE; 2019. p. 1\u20134.","DOI":"10.1109\/FG.2019.8756583"},{"key":"e_1_3_3_155_2","doi-asserted-by":"crossref","first-page":"129","DOI":"10.1016\/j.image.2019.02.005","article-title":"OFF-ApexNet on micro-expression recognition system","volume":"74","author":"Gan YS","year":"2019","unstructured":"Gan YS, Liong ST, Yau WC, Huang YC, Tan LK. OFF-ApexNet on micro-expression recognition system. Signal Process Image Commun. 2019;74:129\u2013139.","journal-title":"Signal Process Image Commun"},{"key":"e_1_3_3_156_2","doi-asserted-by":"crossref","unstructured":"Zhou L Mao Q Xue L. Dual-inception network for cross-database micro-expression recognition. In: 2019 14th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2019). Lille (France): IEEE; 2019. p. 1\u20135.","DOI":"10.1109\/FG.2019.8756579"},{"issue":"4","key":"e_1_3_3_157_2","doi-asserted-by":"crossref","first-page":"1973","DOI":"10.1109\/TAFFC.2022.3213509","article-title":"Short and long range relation based spatiotemporal transformer for micro-expression recognition","volume":"13","author":"Zhang L","year":"2022","unstructured":"Zhang L, Hong X, Arandjelovic O, Zhao G. Short and long range relation based spatiotemporal transformer for micro-expression recognition. IEEE Trans Affect Comput. 2022;13(4):1973\u20131985.","journal-title":"IEEE Trans Affect Comput"},{"key":"e_1_3_3_158_2","doi-asserted-by":"crossref","unstructured":"Desplanques B Thienpondt J Demuynck K. ECAPA-TDNN: Emphasized channel attention propagation and aggregation in tdnn based speaker verification. arXiv. 2020. https:\/\/doi.org\/10.48550\/arXiv.2005.07143","DOI":"10.21437\/Interspeech.2020-2650"},{"key":"e_1_3_3_159_2","first-page":"652","article-title":"Res2Net: A new multi scale backbone architecture","volume":"43","author":"Gao SH","year":"2019","unstructured":"Gao SH, Cheng MM, Zhao K, Zhang XY, Yang MH, Torr P. Res2Net: A new multi scale backbone architecture. IEEE Trans Pattern Anal Mach Intell. 2019;43:652\u2013662.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"e_1_3_3_160_2","doi-asserted-by":"crossref","unstructured":"Hu J Shen L Sun G. Squeeze-and-excitation networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Salt Lake City (UT): IEEE; 2018. p. 7132\u20137141.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_3_3_161_2","doi-asserted-by":"crossref","first-page":"2880","DOI":"10.1109\/TASLP.2020.3030497","article-title":"PANNs: Large-scale pretrained audio neural networks for audio pattern recognition","volume":"28","author":"Kong Q","year":"2020","unstructured":"Kong Q, Cao Y, Iqbal T, Wang Y, Wang W, Plumbley MD. PANNs: Large-scale pretrained audio neural networks for audio pattern recognition. IEEE\/ACM Trans Audio Speech Lang Process. 2020;28:2880\u20132894.","journal-title":"IEEE\/ACM Trans Audio Speech Lang Process"},{"key":"e_1_3_3_162_2","doi-asserted-by":"crossref","unstructured":"Carreira J Zisserman A. Quo vadis action recognition? A new model and the Kinetics dataset. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Honolulu (HI): IEEE; 2017. p. 6299\u20136308.","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_3_163_2","doi-asserted-by":"crossref","unstructured":"Tran D Bourdev L Fergus R Torresani L Paluri M. Learning spatiotemporal features with 3D convolutional networks. In: Proceedings of the IEEE International Conference on Computer Vision. Santiago (Chile): IEEE; 2015. p. 4489\u20134497.","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_3_164_2","unstructured":"Bertasius G Wang H Torresani L. Is space-time attention all you need for video understanding? In: ICML. New York (NY): PLMR; 2021."},{"key":"e_1_3_3_165_2","unstructured":"OpenAI. OpenAI o3. https:\/\/openai.com\/index\/introducing-o3-and-o4-mini\/. San Francisco USA 2025."},{"key":"e_1_3_3_166_2","unstructured":"Team G. Gemini-2.5-pro-preview-03-25. https:\/\/deepmind.google\/technologies\/gemini\/pro\/. California USA 2025."},{"key":"e_1_3_3_167_2","unstructured":"OpenAI. Hello GPT-4o. https:\/\/openai.com\/index\/hello-gpt-4o\/. San Francisco USA 2024."},{"key":"e_1_3_3_168_2","unstructured":"Team Q. QVQ: To see the world with wisdom. https:\/\/qwenlm.github.io\/blog\/qvq72b-preview\/. Hangzhou China 2024."},{"key":"e_1_3_3_169_2","unstructured":"Bai S Chen K Liu X Wang J Ge W Song S Dang K Peng W Wang S Tang J et\u00a0al. Qwen2.5-vl technical report. arXiv. 2025. https:\/\/doi.org\/10.48550\/arXiv.2502.13923"},{"key":"e_1_3_3_170_2","unstructured":"xAI. xAI Grok 4. https:\/\/x.ai\/news\/grok-4. Nevada USA 2025."},{"key":"e_1_3_3_171_2","unstructured":"Abdin M Jacobs SA Awadala H Awadala A Awan AA Bach N Bahree A Bakhtiari A Bao J Behl H et\u00a0al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2404.14219"},{"key":"e_1_3_3_172_2","unstructured":"ByteDance. Doubao-1.5-vision-pro-32k. https:\/\/volcengine.com\/product\/doubao. Beijing China 2025."},{"key":"e_1_3_3_173_2","doi-asserted-by":"crossref","unstructured":"Zhang W Cun X Wang X Zhang Y Shen X Guo Y Shan Y Wang F. SadTalker: Learning realistic 3D motion coefficients for stylized audio-driven single image talking face animation. arXiv. 2022. https:\/\/doi.org\/10.48550\/arXiv.2211.12194","DOI":"10.1109\/CVPR52729.2023.00836"},{"key":"e_1_3_3_174_2","unstructured":"Wei H Yang Z Wang Z. AniPortrait: Audio-driven synthesis of photorealistic portrait animations. arXiv. 2024. https:\/\/doi.org\/10.48550\/arXiv.2403.17694"},{"key":"e_1_3_3_175_2","doi-asserted-by":"crossref","unstructured":"Ji X Hu X Xu Z Zhu J Lin C He Q Zhang J Luo D Chen Y Lin Q et\u00a0al. Sonic: Shifting focus to global audio perception in portrait animation. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville (TN)L IEEE; 2025. p. 193\u2013203.","DOI":"10.1109\/CVPR52734.2025.00027"},{"issue":"3","key":"e_1_3_3_176_2","first-page":"2403","article-title":"EchoMimic: Lifelike audio-driven portrait animations through editable landmark conditions","volume":"39","author":"Chen Z","year":"2025","unstructured":"Chen Z, Cao J, Chen Z, Li Y, Ma C. EchoMimic: Lifelike audio-driven portrait animations through editable landmark conditions. Proc AAAI Conf Artif Intell. 2025;39(3):2403\u20132410.","journal-title":"Proc AAAI Conf Artif Intell"},{"key":"e_1_3_3_177_2","doi-asserted-by":"crossref","unstructured":"Cui J Li H Zhan Y Shang H Cheng K Ma Y Mu S Zhou H Wang J Zhu S. Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville (TN): IEEE; 2025. p. 21086\u201321095.","DOI":"10.1109\/CVPR52734.2025.01964"},{"key":"e_1_3_3_178_2","doi-asserted-by":"crossref","unstructured":"Wang M Wang Q Jiang F Fan Y Zhang Y Qi Y Zhao Z Xu M. FantasyTalking: Realistic talking portrait generation via coherent motion synthesis. arXiv. 2025. https:\/\/doi.org\/10.48550\/arXiv.2504.04842","DOI":"10.1145\/3746027.3755217"},{"key":"e_1_3_3_179_2","unstructured":"Chen Y Liang S Zhou Z Huang Z Ma Y Tang J Lin Q Zhou Y Lu Q. HunyuanVideo-Avatar: High-fidelity audio-driven human animation for multiple characters. arXiv. 2025. https:\/\/doi.org\/10.48550\/arXiv.2505.20156"},{"key":"e_1_3_3_180_2","unstructured":"Kong Z Gao F Zhang Y Kang Z Wei X Cai X Chen G Luo W. Let them talk: Audio-driven multi-person conversational video generation. arXiv. 2025. https:\/\/doi.org\/10.48550\/arXiv.2505.22647"},{"key":"e_1_3_3_181_2","unstructured":"Gan Q Yang R Zhu J Xue S Hoi S. OmniAvatar: Efficient audio-driven avatar video generation with adaptive body animation. arXiv. 2025. https:\/\/doi.org\/10.48550\/arXiv.2506.18866"},{"key":"e_1_3_3_182_2","unstructured":"Bai J Bai S Yang S Wang S Tan S Wang P Lin J Zhou C Zhou J. Qwen-VL: A versatile vision-language model for understanding localization text reading and beyond. arXiv. 2023. https:\/\/doi.org\/10.48550\/arXiv.2308.12966"}],"container-title":["Intelligent Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/spj.science.org\/doi\/pdf\/10.34133\/icomputing.0246","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,21]],"date-time":"2026-01-21T10:15:14Z","timestamp":1768990514000},"score":1,"resource":{"primary":{"URL":"https:\/\/spj.science.org\/doi\/10.34133\/icomputing.0246"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1]]},"references-count":181,"alternative-id":["10.34133\/icomputing.0246"],"URL":"https:\/\/doi.org\/10.34133\/icomputing.0246","relation":{},"ISSN":["2771-5892"],"issn-type":[{"value":"2771-5892","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1]]},"assertion":[{"value":"2025-05-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-27","order":1,"name":"revised","label":"Revised","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-23","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"0246"}}