{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T15:48:40Z","timestamp":1783439320624,"version":"3.54.6"},"reference-count":47,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2025,8,6]],"date-time":"2025-08-06T00:00:00Z","timestamp":1754438400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Beijing Outstanding Young Scientist Program","award":["JWZQ20240101027"],"award-info":[{"award-number":["JWZQ20240101027"]}]},{"name":"Beijing Outstanding Young Scientist Program","award":["12071313"],"award-info":[{"award-number":["12071313"]}]},{"name":"National Nature Science Foundation of China","award":["JWZQ20240101027"],"award-info":[{"award-number":["JWZQ20240101027"]}]},{"name":"National Nature Science Foundation of China","award":["12071313"],"award-info":[{"award-number":["12071313"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Multimodal sentiment analysis (MSA) benefits from integrating diverse modalities (e.g., text, video, and audio). However, challenges remain in effectively aligning non-text features and mitigating redundant information, which may limit potential performance improvements. To address these challenges, we propose a Hierarchical Text-Guided Refinement Network (HTRN), a novel framework that refines and aligns non-text modalities using hierarchical textual representations. We introduce Shuffle-Insert Fusion (SIF) and the Text-Guided Alignment Layer (TAL) to enhance crossmodal interactions and suppress irrelevant signals. In SIF, empty tokens are inserted at fixed intervals in unimodal feature sequences, disrupting local correlations and promoting more generalized representations with improved feature diversity. The TAL guides the refinement of audio and visual representations by leveraging textual semantics and dynamically adjusting their contributions through learnable gating factors, ensuring that non-text modalities remain semantically coherent while retaining essential crossmodal interactions. Experiments demonstrate that the HTRN achieves state-of-the-art performance with accuracies of 86.3% (Acc-2) on CMU-MOSI, 86.7% (Acc-2) on CMU-MOSEI, and 80.3% (Acc-2) on CH-SIMS, outperforming existing methods by 0.8\u20133.45%. Ablation studies validate the contributions of SIF and the TAL, showing 1.9\u20132.1% performance gains over baselines. By integrating these components, the HTRN establishes a robust multimodal representation learning framework.<\/jats:p>","DOI":"10.3390\/e27080834","type":"journal-article","created":{"date-parts":[[2025,8,6]],"date-time":"2025-08-06T15:09:53Z","timestamp":1754492993000},"page":"834","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Hierarchical Text-Guided Refinement Network for Multimodal Sentiment Analysis"],"prefix":"10.3390","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-0472-9062","authenticated-orcid":false,"given":"Yue","family":"Su","sequence":"first","affiliation":[{"name":"School of Mathematical Sciences, Capital Normal University, Beijing 100048, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-1367-0692","authenticated-orcid":false,"given":"Xuying","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Mathematical Sciences, Capital Normal University, Beijing 100048, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,8,6]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1016\/j.imavis.2017.08.003","article-title":"A survey of multimodal sentiment analysis","volume":"65","author":"Soleymani","year":"2017","journal-title":"Image Vis. Comput."},{"key":"ref_2","unstructured":"Elmadany, A.A., Mubarak, H., and Magdy, W. (2018, January 8). An Arabic speech-act and sentiment Corpus of Tweets. Proceedings of the 3rd Workshop on Open-Source Arabic Corpora and Processing Tools, Miyazaki, Japan."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"75","DOI":"10.1007\/s10462-024-10988-9","article-title":"A review of Chinese sentiment analysis: Subjects, methods, and trends","volume":"58","author":"Wang","year":"2025","journal-title":"Artif. Intell. Rev."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"306","DOI":"10.1016\/j.inffus.2023.02.028","article-title":"Multimodal sentiment analysis based on fusion methods: A survey","volume":"95","author":"Zhu","year":"2023","journal-title":"Inf. Fusion"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"424","DOI":"10.1016\/j.inffus.2022.09.025","article-title":"Multimodal sentiment analysis: A systematic review of history, datasets, multimodal fusion methods, applications, challenges and future directions","volume":"91","author":"Gandhi","year":"2023","journal-title":"Inf. Fusion"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1479","DOI":"10.1007\/s10462-023-10555-8","article-title":"A comprehensive survey on deep learning-based approaches for multimodal sentiment analysis","volume":"56","author":"Ghorbanali","year":"2023","journal-title":"Artif. Intell. Rev."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1109\/MIS.2016.94","article-title":"Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages","volume":"31","author":"Zadeh","year":"2016","journal-title":"IEEE Intell. Syst."},{"key":"ref_8","unstructured":"Zadeh, A.B., Liang, P.P., Poria, S., Cambria, E., and Morency, L.P. (2018, January 15\u201320). Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Melbourne, Australia."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zadeh, A., Chen, M., Poria, S., Cambria, E., and Morency, L.P. (2017). Tensor fusion network for multimodal sentiment analysis. arXiv.","DOI":"10.18653\/v1\/D17-1115"},{"key":"ref_10","unstructured":"Tsai, Y.H.H., Bai, S., Liang, P.P., Kolter, J.Z., Morency, L.P., and Salakhutdinov, R. (August, January 28). Multimodal transformer for unaligned multimodal language sequences. Proceedings of the Conference, Association for Computational Linguistics, Meeting, Florence, Italy. NIH Public Access."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Yang, D., Huang, S., Kuang, H., Du, Y., and Zhang, L. (2022, January 10\u201314). Disentangled representation learning for multimodal emotion recognition. Proceedings of the 30th ACM International Conference on Multimedia, Lisboa, Portugal.","DOI":"10.1145\/3503161.3547754"},{"key":"ref_12","unstructured":"Hazarika, D., Zimmermann, R., and Poria, S. (2020, January 12\u201316). Misa: Modality-invariant and-specific representations for multimodal sentiment analysis. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Yang, J., Yu, Y., Niu, D., Guo, W., and Xu, Y. (2023, January 10\u201312). ConFEDE: Contrastive Feature Decomposition for Multimodal Sentiment Analysis. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Toronto, ON, Canada.","DOI":"10.18653\/v1\/2023.acl-long.421"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"103229","DOI":"10.1016\/j.ipm.2022.103229","article-title":"PS-mixer: A polar-vector and strength-vector mixer model for multimodal sentiment analysis","volume":"60","author":"Lin","year":"2023","journal-title":"Inf. Process. Manag."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Yu, W., Xu, H., Meng, F., Zhu, Y., Ma, Y., Wu, J., Zou, J., and Yang, K. (2020, January 5\u201310). Ch-sims: A chinese multimodal sentiment analysis dataset with fine-grained annotation of modality. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online.","DOI":"10.18653\/v1\/2020.acl-main.343"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"228","DOI":"10.1007\/s00530-024-01421-w","article-title":"Text-centered cross-sample fusion network for multimodal sentiment analysis","volume":"30","author":"Huang","year":"2024","journal-title":"Multimed. Syst."},{"key":"ref_17","first-page":"1","article-title":"PAMoE-MSA: Polarity-aware mixture of experts network for multimodal sentiment analysis","volume":"14","author":"Huang","year":"2025","journal-title":"Int. J. Multimed. Inf. Retr."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Chen, J., Huang, Q., Huang, C., and Huang, X. (2025). Actual Cause Guided Adaptive Gradient Scaling for Balanced Multimodal Sentiment Analysis. ACM Trans. Multimed. Comput. Commun. Appl.","DOI":"10.1145\/3736415"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Rahman, W., Hasan, M.K., Lee, S., Zadeh, A., Mao, C., Morency, L.P., and Hoque, E. (2020, January 5\u201310). Integrating multimodal information in large pretrained transformers. Proceedings of the Conference, Association for Computational Linguistics, Meeting, Seattle, WA, USA. NIH Public Access.","DOI":"10.18653\/v1\/2020.acl-main.214"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"196","DOI":"10.1007\/s40747-025-01806-y","article-title":"H 2 CAN: Heterogeneous hypergraph attention network with counterfactual learning for multimodal sentiment analysis","volume":"11","author":"Huang","year":"2025","journal-title":"Complex Intell. Syst."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"102725","DOI":"10.1016\/j.inffus.2024.102725","article-title":"AtCAF: Attention-based causality-aware fusion network for multimodal sentiment analysis","volume":"114","author":"Huang","year":"2025","journal-title":"Inf. Fusion"},{"key":"ref_22","unstructured":"Song, Q. (2024). Multimodal Sentiment Analysis Based on Learning Genes and Uncertainty Estimation. [Master\u2019s Thesis, Central China Normal University]. (In Chinese)."},{"key":"ref_23","first-page":"1137","article-title":"A neural probabilistic language model","volume":"3","author":"Bengio","year":"2003","journal-title":"J. Mach. Learn. Res."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Collobert, R., and Weston, J. (2008, January 5\u20139). A unified architecture for natural language processing: Deep neural networks with multitask learning. Proceedings of the 25th International Conference on Machine Learning, Helsinki, Finland.","DOI":"10.1145\/1390156.1390177"},{"key":"ref_25","unstructured":"Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C.D. (2014, January 25\u201329). Glove: Global vectors for word representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_27","unstructured":"Kenton, J.D.M.W.C., and Toutanova, L.K. (2019, January 2\u20137). Bert: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the naacL-HLT, Minneapolis, MN, USA."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Benitez-Quiroz, C.F., Wang, Y., and Martinez, A.M. (2017, January 22\u201329). Recognition of action units in the wild with deep nets and a new global-Local loss. Proceedings of the ICCV, Venice, Italy.","DOI":"10.1109\/ICCV.2017.428"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning spatiotemporal features with 3D convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Lao, S., and Kawade, M. (2004, January 13\u201314). Vision-based face understanding technologies and their applications. Proceedings of the Chinese Conference on Biometric Recognition, Guangzhou, China.","DOI":"10.1007\/978-3-540-30548-4_39"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Littlewort, G., Whitehill, J., Wu, T., Fasel, I., Frank, M., Movellan, J., and Bartlett, M. (2011, January 21\u201324). The computer expression recognition toolbox (CERT). Proceedings of the 2011 IEEE International Conference on Automatic Face & Gesture Recognition (FG), Santa Barbara, CA, USA.","DOI":"10.1109\/FG.2011.5771414"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Baltrusaitis, T., Zadeh, A., Lim, Y.C., and Morency, L.-P. (2018, January 15\u201319). Openface 2.0: Facial behavior analysis toolkit. Proceedings of the 13th IEEE International Conference on Automatic Face & Gesture Recognition, Xi\u2019an, China.","DOI":"10.1109\/FG.2018.00019"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1446","DOI":"10.3758\/s13428-017-0996-1","article-title":"Facial expression analysis with AFFDEX and FACET: A validation study","volume":"50","author":"Borer","year":"2018","journal-title":"Behav. Res. Methods"},{"key":"ref_34","unstructured":"Anand, N., and Verma, P. (2015). Convoluted Feelings Convolutional and Recurrent Nets for Detecting Emotion from Audio Data, Stanford University. Technical Report."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Eyben, F., W\u00f6llmer, M., and Schuller, B. (2009, January 10\u201312). OpenEAR\u2014introducing the Munich open-source emotion and affect recognition toolkit. Proceedings of the 2009 3rd International Conference on Affective Computing and Intelligent Interaction and Workshops, Amsterdam, The Netherlands.","DOI":"10.1109\/ACII.2009.5349350"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Eyben, F., W\u00f6llmer, M., and Schuller, B. (2010, January 25\u201329). Opensmile: The munich versatile and fast open-source audio feature extractor. Proceedings of the 18th ACM International Conference on Multimedia, Firenze, Italy.","DOI":"10.1145\/1873951.1874246"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"McFee, B., Raffel, C., Liang, D., Ellis, D., McVicar, M., Battenberg, E., and Nieto, O. (2015, January 6\u201312). librosa: Audio and Music Signal Analysis in Python. Proceedings of the 14th Python in Science Conference, SciPy, Austin, TX, USA.","DOI":"10.25080\/Majora-7b98e3ed-003"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Degottex, G., Kane, J., Drugman, T., Raitio, T., and Scherer, S. (2014, January 4\u20139). COVAREP\u2014A collaborative voice analysis repository for speech technologies. Proceedings of the 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6853739"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Mao, H., Yuan, Z., Xu, H., Yu, W., Liu, Y., and Gao, K. (2022, January 22\u201327). M-SENA: An Integrated Platform for Multimodal Sentiment Analysis. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Dublin, Ireland.","DOI":"10.18653\/v1\/2022.acl-demo.20"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Zhang, H., Wang, Y., Yin, G., Liu, K., Liu, Y., and Yu, T. (2023, January 6\u201310). Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore.","DOI":"10.18653\/v1\/2023.emnlp-main.49"},{"key":"ref_41","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2025, June 08). Attention Is All you Need. Advances in Neural Information Processing Systems 30 (NIPS 2017). Available online: https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2017\/file\/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Yu, W., Xu, H., Yuan, Z., and Wu, J. (2021, January 2\u20139). Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis. Proceedings of the AAAI Conference on Artificial Intelligence, Virtually.","DOI":"10.1609\/aaai.v35i12.17289"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Liu, Z., Shen, Y., Lakshminarasimhan, V.B., Liang, P.P., Zadeh, A.B., and Morency, L.P. (2018, January 15\u201320). Efficient Low-rank Multimodal Fusion With Modality-Specific Factors. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Melbourne, Australia.","DOI":"10.18653\/v1\/P18-1209"},{"key":"ref_44","unstructured":"Tsai, Y.H.H., Liang, P.P., Zadeh, A., Morency, L.P., and Salakhutdinov, R. (2019, January 6\u20139). Learning Factorized Multimodal Representations. Proceedings of the International Conference on Learning Representations, New Orleans, LA, USA."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Lv, F., Chen, X., Huang, Y., Duan, L., and Lin, G. (2021, January 20\u201325). Progressive modality reinforcement for human multimodal emotion recognition from unaligned multimodal sequences. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00258"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"3020","DOI":"10.1007\/s11263-024-02304-3","article-title":"Noise-resistant multimodal transformer for emotion recognition","volume":"133","author":"Liu","year":"2025","journal-title":"Int. J. Comput. Vis."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"108348","DOI":"10.1016\/j.engappai.2024.108348","article-title":"Token-disentangling Mutual Transformer for multimodal emotion recognition","volume":"133","author":"Yin","year":"2024","journal-title":"Eng. Appl. Artif. Intell."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/8\/834\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:24:07Z","timestamp":1760034247000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/8\/834"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,6]]},"references-count":47,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2025,8]]}},"alternative-id":["e27080834"],"URL":"https:\/\/doi.org\/10.3390\/e27080834","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,6]]}}}