{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T14:55:33Z","timestamp":1784904933364,"version":"3.55.0"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T00:00:00Z","timestamp":1782259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100022023","name":"Brandenburgische Technische Universit\u00e4t Cottbus-Senftenberg","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100022023","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>\n                    The emergence of Vision-Language Models (VLMs) like Contrastive Language-Image Pretraining (CLIP) provides appealing solutions to various vision problems including Dynamic Facial Expression Recognition (DFER). However, most of the proposed approaches face major challenges, particularly related to inefficient full fine-tuning of the encoders and the complexity of the models. Moreover, some of the proposed methods seem to struggle with suboptimal performance due to (i) poor alignment between textual and visual representations, and (ii) ineffective temporal modeling. To address these challenges, we propose PE-CLIP, a parameter-efficient fine-tuning (PEFT) framework that elegantly adapts CLIP for dynamic facial expression recognition, requiring significantly reduced number of trainable parameters while maintaining high accuracy. At its core, to enhance efficiency and performance, PE-CLIP introduces two specialized adapters namely a Temporal Dynamic Adapter (TDA) and a Shared Adapter (ShA). The TDA is a GRU-based module with a dynamic scaling mechanism, capturing sequential dependencies while adaptively modulating the contribution of each temporal feature to emphasize the most informative ones while mitigating irrelevant variations. The ShA is a lightweight adapter refine representations within both textual and visual encoders, ensuring consistent feature processing while maintaining parameter efficiency. Additionally, we leverage Multi-modal Prompt Learning (MaPLe), which introduces learnable prompts to both visual and action unit-based textual description inputs, further improving the semantic alignment between modalities and enabling the efficient adaptation of CLIP for dynamic tasks. We evaluate our proposed PE-CLIP on two benchmark datasets, namely DFEW, FERV39K, and AFEW, achieving competitive performance compared to state-of-the-art methods while requiring fewer trainable parameters. By striking an optimal balance between parameter efficiency and performance, PE-CLIP sets a new benchmark in resource-efficient DFER. The source code of the proposed PE-CLIP will be publicly available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/Ibtissam-SAADI\/PE-CLIP\">https:\/\/github.com\/Ibtissam-SAADI\/PE-CLIP<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3786789","type":"journal-article","created":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T14:07:04Z","timestamp":1768831624000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["PE-CLIP: A Parameter-Efficient Fine-Tuning of Vision Language Models for Dynamic Facial Expression Recognition"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-4231-4246","authenticated-orcid":false,"given":"Ibtissam","family":"Saadi","sequence":"first","affiliation":[{"name":"Faculty of Graphical Systems, Univ. BTU Cottbus-Senftenberg, Cottbus, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9092-735X","authenticated-orcid":false,"given":"Abdenour","family":"Hadid","sequence":"additional","affiliation":[{"name":"Sorbonne Center for Artificial Intelligence, Sorbonne University Abu Dhabi, Abu Dhabi, United Arab Emirates"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1419-2552","authenticated-orcid":false,"given":"Douglas W.","family":"Cunningham","sequence":"additional","affiliation":[{"name":"Faculty of Graphical Systems, Univ. BTU Cottbus-Senftenberg, Cottbus, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7218-3799","authenticated-orcid":false,"given":"Abdelmalik","family":"Taleb-Ahmed","sequence":"additional","affiliation":[{"name":"Laboratory of IEMN, CNRS, Centrale Lille, UMR 8520, Univ. Polytechnique Hauts-de-France, Valenciennes, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3980-9902","authenticated-orcid":false,"given":"Yassin El","family":"Hillali","sequence":"additional","affiliation":[{"name":"Laboratory of IEMN, CNRS, Centrale Lille, UMR 8520, Univ. Polytechnique Hauts-de-France, Valenciennes, France"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2006.244"},{"key":"e_1_3_1_3_2","unstructured":"Shuai Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge Sibo Song Kai Dang Peng Wang Shijie Wang Jun Tang et al. 2025. Qwen2.5-VL technical report. arXiv:2502.13923. Retrieved from https:\/\/arxiv.org\/abs\/2502.13923"},{"key":"e_1_3_1_4_2","unstructured":"Shaojie Bai J. Zico Kolter and Vladlen Koltun. 2018. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv:1803.01271. Retrieved from https:\/\/arxiv.org\/abs\/1803.01271"},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","first-page":"1181","DOI":"10.1007\/978-3-030-41862-5_119","volume-title":"Proceedings of the New Trends in Computational Vision and Bio-Inspired Computing: Selected Works Presented at the ICCVBIC 2018","author":"Chattopadhyay Joyati","year":"2020","unstructured":"Joyati Chattopadhyay, Souvik Kundu, Arpita Chakraborty, and Jyoti Sekhar Banerjee. 2020. Facial expression recognition for human computer interaction. In Proceedings of the New Trends in Computational Vision and Bio-Inspired Computing: Selected Works Presented at the ICCVBIC 2018, 1181\u20131192."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1212"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2024.3453443"},{"key":"e_1_3_1_8_2","unstructured":"Kyunghyun Cho Bart Van Merri\u00ebnboer Caglar Gulcehre Dzmitry Bahdanau Fethi Bougares Holger Schwenk and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv:1406.1078. Retrieved from https:\/\/arxiv.org\/abs\/1406.1078"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/MMUL.2012.26"},{"key":"e_1_3_1_10_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_1_11_2","volume-title":"What the Face Reveals: Basic and Applied Studies of Spontaneous Expression Using the Facial Action Coding System (FACS)","author":"Ekman Paul","year":"1997","unstructured":"Paul Ekman and Erika L. Rosenberg. 1997. What the Face Reveals: Basic and Applied Studies of Spontaneous Expression Using the Facial Action Coding System (FACS). Oxford University Press."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1207\/s15516709cog1402_1"},{"key":"e_1_3_1_13_2","first-page":"1","volume-title":"Proceedings of the 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG)","author":"Foteinopoulou Niki Maria","year":"2024","unstructured":"Niki Maria Foteinopoulou and Ioannis Patras. 2024. EmoCLIP: A vision-language method for zero-shot video facial expression recognition. In Proceedings of the 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG). IEEE, 1\u201310."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i20.30206"},{"key":"e_1_3_1_15_2","first-page":"1","volume-title":"Proceedings of the 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG)","author":"Gowda Shreyank N.","year":"2024","unstructured":"Shreyank N. Gowda, Boyan Gao, and David A. Clifton. 2024. FE-Adapter: Adapting image-based emotion classifiers to videos. In Proceedings of the 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG). IEEE, 1\u20136."},{"key":"e_1_3_1_16_2","unstructured":"Albert Gu and Tri Dao. 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv:2312.00752. Retrieved from https:\/\/arxiv.org\/abs\/2312.00752"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSP62122.2024.10743836"},{"key":"e_1_3_1_18_2","unstructured":"Ethan Harris Antonia Marcu Matthew Painter Mahesan Niranjan Adam Pr\u00fcgel-Bennett and Jonathon Hare. 2020. FMix: Enhancing mixed sample data augmentation. arXiv:2002.12047. Retrieved from https:\/\/arxiv.org\/abs\/2002.12047"},{"key":"e_1_3_1_19_2","unstructured":"Junxian He Chunting Zhou Xuezhe Ma Taylor Berg-Kirkpatrick and Graham Neubig. 2021. Towards a unified view of parameter-efficient transfer learning. arXiv:2110.04366. Retrieved from https:\/\/arxiv.org\/abs\/2110.04366"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_21_2","first-page":"2790","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Houlsby Neil","year":"2019","unstructured":"Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the International Conference on Machine Learning. PMLR, 2790\u20132799."},{"key":"e_1_3_1_22_2","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations Vol. 1 3."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19827-4_41"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413620"},{"issue":"6","key":"e_1_3_1_25_2","doi-asserted-by":"crossref","first-page":"2063","DOI":"10.1007\/s13042-023-02016-z","article-title":"Transformer embedded spectral-based graph network for facial expression recognition","volume":"15","author":"Jin Xing","year":"2024","unstructured":"Xing Jin, Xulin Song, Xiyin Wu, and Wenzhu Yan. 2024. Transformer embedded spectral-based graph network for facial expression recognition. International Journal of Machine Learning and Cybernetics 15, 6 (2024), 2063\u20132077.","journal-title":"International Journal of Machine Learning and Cybernetics"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-023-16066-6"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01832"},{"key":"e_1_3_1_28_2","first-page":"5681","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lee Bokyeung","year":"2023","unstructured":"Bokyeung Lee, Hyunuk Shin, Bonhwa Ku, and Hanseok Ko. 2023. Frame level emotion guided dynamic facial expression recognition with emotion grouping. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 5681\u20135691."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.2996086"},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1609\/aaai.v37i1.25077","article-title":"Intensity-aware loss for dynamic facial expression recognition in the wild","volume":"37","author":"Li Hanting","year":"2023","unstructured":"Hanting Li, Hongjing Niu, Zhaoqing Zhu, and Feng Zhao. 2023. Intensity-aware loss for dynamic facial expression recognition in the wild. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, 67\u201375.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_1_31_2","first-page":"1","volume-title":"Proceedings of the 2024 IEEE International Conference on Multimedia and Expo (ICME)","author":"Li Hanting","year":"2024","unstructured":"Hanting Li, Hongjing Niu, Zhaoqing Zhu, and Feng Zhao. 2024. CLIPER: A unified vision-language framework for in-the-wild facial expression recognition. In Proceedings of the 2024 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1\u20136."},{"key":"e_1_3_1_32_2","unstructured":"Hanting Li Mingzhe Sui Zhaoqing Zhu et al. 2022. NR-dFERNet: Noise-robust network for dynamic facial expression recognition. arXiv:2206.04975. Retrieved from https:\/\/arxiv.org\/abs\/2206.04975"},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","first-page":"120953","DOI":"10.1016\/j.ins.2024.120953","article-title":"Dual-STI: Dual-path spatial-temporal interaction learning for dynamic facial expression recognition","volume":"678","author":"Li Min","year":"2024","unstructured":"Min Li, Xiaoqin Zhang, Chenxiang Fan, Tangfei Liao, and Guobao Xiao. 2024. Dual-STI: Dual-path spatial-temporal interaction learning for dynamic facial expression recognition. Information Sciences 678 (2024), 120953.","journal-title":"Information Sciences"},{"key":"e_1_3_1_34_2","unstructured":"Kun-Yu Lin Henghui Ding Jiaming Zhou Yu-Ming Tang Yi-Xing Peng Zhilin Zhao Chen Change Loy and Wei-Shi Zheng. 2024. Rethinking clip-based video learners in cross-domain open-vocabulary action recognition. arXiv:2403.01560. Retrieved from https:\/\/arxiv.org\/abs\/2403.01560"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548190"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2022.03.062"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2023.109368"},{"key":"e_1_3_1_38_2","unstructured":"Fuyan Ma Bin Sun and Shutao Li. 2022. Spatio-temporal transformer for dynamic facial expression recognition in the wild. arXiv:2205.04749. Retrieved from https:\/\/arxiv.org\/abs\/2205.04749"},{"key":"e_1_3_1_39_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201923)","author":"Ma Fuyan","year":"2023","unstructured":"Fuyan Ma, Bin Sun, and Shutao Li. 2023. Logo-former: Local-global spatio-temporal transformer for dynamic facial expression recognition. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201923). IEEE, 1\u20135."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.iswa.2024.200339"},{"key":"e_1_3_1_41_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_1_42_2","first-page":"7492","article-title":"Facial expression monitoring via fine-grained vision-language alignment","volume":"22","author":"Ren Weihong","year":"2024","unstructured":"Weihong Ren, Yu Gao, Xi\u2019ai Chen, Zhi Han, Zhiyong Wang, Jiaole Wang, and Honghai Liu. 2024. Facial expression monitoring via fine-grained vision-language alignment. IEEE Transactions on Automation Science and Engineering 22 (2024), 7492\u20137505.","journal-title":"IEEE Transactions on Automation Science and Engineering"},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","first-page":"122784","DOI":"10.1016\/j.eswa.2023.122784","article-title":"Driver\u2019s facial expression recognition: A comprehensive survey","volume":"242","author":"Saadi Ibtissam","year":"2023","unstructured":"Ibtissam Saadi, Douglas W. Cunningham, Taleb-Ahmed Abdelmalik, Abdenour Hadid, and Yassin El Hillali. 2023. Driver\u2019s facial expression recognition: A comprehensive survey. Expert Systems with Applications 242 (2023), 122784.","journal-title":"Expert Systems with Applications"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612365"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3423327.3423672"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611972"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"issue":"11","key":"e_1_3_1_48_2","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Van der Maaten Laurens","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9, 11 (2008), 2579\u20132605.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_49_2","doi-asserted-by":"crossref","first-page":"17958","DOI":"10.1109\/CVPR52729.2023.01722","volume-title":"Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Wang Hanyang","year":"2023","unstructured":"Hanyang Wang, Bo Li, Shuang Wu, Siyuan Shen, Feng Liu, Shouhong Ding, and Aimin Zhou. 2023. Rethinking the learning paradigm for dynamic facial expression recognition. In Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17958\u201317968."},{"key":"e_1_3_1_50_2","first-page":"20922","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Yan","year":"2022","unstructured":"Yan Wang, Yixuan Sun, Yiwen Huang, Zhongying Liu, Shuyong Gao, Wei Zhang, Weifeng Ge, and Wenqiang Zhang. 2022. FERV39k: A large-scale multi-scene dataset for facial expression recognition in videos. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 20922\u201320931."},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547865"},{"key":"e_1_3_1_52_2","doi-asserted-by":"crossref","first-page":"16085","DOI":"10.1609\/aaai.v38i14.29541","article-title":"VMT-Adapter: Parameter-efficient transfer learning for multi-task dense scene understanding","volume":"38","author":"Xin Yi","year":"2024","unstructured":"Yi Xin, Junlong Du, Qiang Wang, Zhiwen Lin, and Ke Yan. 2024. VMT-Adapter: Parameter-efficient transfer learning for multi-task dense scene understanding. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 16085\u201316093.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_1_53_2","doi-asserted-by":"crossref","first-page":"6540","DOI":"10.1609\/aaai.v38i7.28475","article-title":"DGL: Dynamic global-local prompt tuning for text-video retrieval","volume":"38","author":"Yang Xiangpeng","year":"2024","unstructured":"Xiangpeng Yang, Linchao Zhu, Xiaohan Wang, and Yi Yang. 2024. DGL: Dynamic global-local prompt tuning for text-video retrieval. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 6540\u20136548.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_1_54_2","unstructured":"Hongyi Zhang. 2017. mixup: Beyond empirical risk minimization. arXiv:1710.09412. Retrieved from https:\/\/arxiv.org\/abs\/1710.09412"},{"key":"e_1_3_1_55_2","first-page":"1","volume-title":"Proceedings of the 2024 IEEE International Joint Conference on Biometrics (IJCB)","author":"Zhang Junliang","year":"2024","unstructured":"Junliang Zhang, Xu Liu, Yu Liang, Xiaole Xian, Weicheng Xie, Linlin Shen, and Siyang Song. 2024. CLIP-guided bidirectional prompt and semantic supervision for dynamic facial expression recognition. In Proceedings of the 2024 IEEE International Joint Conference on Biometrics (IJCB), 1\u201310."},{"key":"e_1_3_1_56_2","first-page":"20751","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhang Xiang","year":"2023","unstructured":"Xiang Zhang, Taoyue Wang, Xiaotian Li, Huiyuan Yang, and Lijun Yin. 2023. Weakly-supervised text-driven contrastive learning for facial behavior understanding. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 20751\u201320762."},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2014.06.002"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475292"},{"key":"e_1_3_1_59_2","volume-title":"Proceedings of the 34th British Machine Vision Conference 2023 (BMVC \u201923)","author":"Zhao Zengqun","year":"2023","unstructured":"Zengqun Zhao and Ioannis Patras. 2023. Prompting visual-language models for dynamic facial expression recognition. In Proceedings of the 34th British Machine Vision Conference 2023 (BMVC \u201923). BMVA. Retrieved from https:\/\/papers.bmvc2023.org\/0098.pdf"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01631"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-022-01653-1"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01393"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3786789","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:42:48Z","timestamp":1782312168000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3786789"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":61,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3786789"],"URL":"https:\/\/doi.org\/10.1145\/3786789","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-02-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}