{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,5]],"date-time":"2026-02-05T18:05:21Z","timestamp":1770314721549,"version":"3.49.0"},"reference-count":38,"publisher":"Wiley","issue":"3","license":[{"start":{"date-parts":[[2026,1,28]],"date-time":"2026-01-28T00:00:00Z","timestamp":1769558400000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,1,28]],"date-time":"2026-01-28T00:00:00Z","timestamp":1769558400000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Concurrency and Computation"],"published-print":{"date-parts":[[2026,2]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>Multimodal sentiment analysis (MSA) has recently encountered two major challenges: non\u2010textual modalities are often affected by noise, and sentiment intensity differences are difficult to capture. To address these issues, we propose a Sentiment Intensity Contrastive Text\u2010Enhanced Fusion Network (SICTEF Net), which achieves deep collaboration among text, audio, and visual modalities through three key mechanisms. First, a grouped\u2010channel\u2010attention based Feature Enhancement Module (EMA) is designed to mitigate modality\u2010specific noise and emphasize emotion\u2010sensitive cues by combining spatial\u2013channel interaction mapping with dual\u2010branch attention fusion. Second, a text\u2010centered cross\u2010modal fusion mechanism is introduced, where bidirectional multi\u2010head self\u2010attention and a residual\u2010enhanced encoder jointly enable complementary mappings between text and non\u2010text modalities, thereby producing intermediate representations that preserve semantic primacy while incorporating fine\u2010grained complementary information. Third, a sentiment\u2010intensity weighted contrastive learning strategy dynamically assigns weights to positive and negative sample pairs according to their sentiment intensity differences, allowing the model to more precisely distinguish samples with varying degrees of similarity in the embedding space. Experimental evaluation on the CMU\u2010MOSI and CMU\u2010MOSEI datasets demonstrates that SICTEF Net consistently outperforms state\u2010of\u2010the\u2010art baselines in binary accuracy, F1 score, seven\u2010class accuracy, mean absolute error (MAE), and Pearson correlation. Comprehensive ablation studies further confirm the complementary benefits of EMA, the text\u2010enhanced Transformer, and sentiment\u2010intensity contrastive learning. These results indicate that combining text\u2010driven deep interaction, non\u2010text modality enhancement via channel attention, and contrastive learning can improve the accuracy and robustness of multimodal sentiment analysis.<\/jats:p>","DOI":"10.1002\/cpe.70593","type":"journal-article","created":{"date-parts":[[2026,1,28]],"date-time":"2026-01-28T14:11:42Z","timestamp":1769609502000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Sentiment Intensity Contrastive Text\u2010Enhanced Fusion Network"],"prefix":"10.1002","volume":"38","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-8517-0251","authenticated-orcid":false,"given":"Heng","family":"Jiang","sequence":"first","affiliation":[{"name":"Physics and Electronics Henan University  Kaifeng Henan China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lianke","family":"Shi","sequence":"additional","affiliation":[{"name":"Physics and Electronics Henan University  Kaifeng Henan China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Deyu","family":"Kong","sequence":"additional","affiliation":[{"name":"Physics and Electronics Henan University  Kaifeng Henan China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiahao","family":"Hua","sequence":"additional","affiliation":[{"name":"Physics and Electronics Henan University  Kaifeng Henan China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Honggui","family":"Shang","sequence":"additional","affiliation":[{"name":"Physics and Electronics Henan University  Kaifeng Henan China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lijia","family":"Chen","sequence":"additional","affiliation":[{"name":"Physics and Electronics Henan University  Kaifeng Henan China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,1,28]]},"reference":[{"key":"e_1_2_10_2_1","doi-asserted-by":"crossref","DOI":"10.1016\/j.patcog.2022.108837","article-title":"Multimodal Channel\u2010Wise Attention Transformer Inspired by Multisensory Integration Mechanisms of the Brain","volume":"130","author":"Shi Q.","year":"2022","journal-title":"Pattern Recognition"},{"key":"e_1_2_10_3_1","doi-asserted-by":"crossref","first-page":"296","DOI":"10.1016\/j.patcog.2019.06.013","article-title":"Graph\u2010Based Multimodal Fusion With Metric Learning for Multimodal Classification","volume":"95","author":"Angelou M.","year":"2019","journal-title":"Pattern Recognition"},{"key":"e_1_2_10_4_1","first-page":"107795","article-title":"Vismin: Visual Minimal\u2010Change Understanding","volume":"37","author":"Awal R.","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_10_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01363"},{"key":"e_1_2_10_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2023.102218"},{"key":"e_1_2_10_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547754"},{"key":"e_1_2_10_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3275156"},{"key":"e_1_2_10_9_1","doi-asserted-by":"crossref","first-page":"1033","DOI":"10.1109\/ICDM.2017.134","volume-title":"2017 IEEE International Conference on Data Mining (ICDM)","author":"Poria S.","year":"2017"},{"key":"e_1_2_10_10_1","doi-asserted-by":"crossref","first-page":"2399","DOI":"10.1145\/3132847.3133142","volume-title":"Proceedings of the 2017 ACM on Conference on Information and Knowledge Management","author":"Xu N.","year":"2017"},{"key":"e_1_2_10_11_1","doi-asserted-by":"crossref","first-page":"163","DOI":"10.1145\/3136755.3136801","volume-title":"Proceedings of the 19th ACM International Conference on Multimodal Interaction","author":"Chen M.","year":"2017"},{"key":"e_1_2_10_12_1","first-page":"2020","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Rahman W.","year":"2020"},{"key":"e_1_2_10_13_1","first-page":"6558","volume-title":"Proceedings of the Conference","author":"Tsai Y. H. H.","year":"2019"},{"key":"e_1_2_10_14_1","doi-asserted-by":"crossref","first-page":"992","DOI":"10.1109\/LSP.2021.3078074","article-title":"A Unimodal Reinforced Transformer With Time Squeeze Fusion for Multimodal Sentiment Analysis","volume":"28","author":"He J.","year":"2021","journal-title":"IEEE Signal Processing Letters"},{"key":"e_1_2_10_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413678"},{"key":"e_1_2_10_16_1","first-page":"7617","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Yang J.","year":"2023"},{"key":"e_1_2_10_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2023.02.028"},{"key":"e_1_2_10_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2020.2987728"},{"key":"e_1_2_10_19_1","first-page":"10790","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Yu W.","year":"2021"},{"key":"e_1_2_10_20_1","volume-title":"Proceedings of the AAAI conference on artificial intelligence","author":"Zadeh A.","year":"2018"},{"key":"e_1_2_10_21_1","doi-asserted-by":"crossref","first-page":"2924","DOI":"10.18653\/v1\/2022.emnlp-main.189","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Zeng J.","year":"2022"},{"key":"e_1_2_10_22_1","doi-asserted-by":"crossref","unstructured":"Y.Wang Y.Shen Z.Liu P. P.Liang A.Zadeh andL. P.Morency \u201cWords Can Shift: Dynamically Adjusting Word Representations Using Nonverbal Behaviors \u201d in\u2009Proceedings of the AAAI conference on artificial intelligence vol.\u200933\u2009(2019) \u20097216\u20137223.","DOI":"10.1609\/aaai.v33i01.33017216"},{"key":"e_1_2_10_23_1","doi-asserted-by":"crossref","unstructured":"G.Hu T. E.Lin Y.Zhao G.Lu Y.Wu andY.Li \u201cUniMSE: Towards unified multimodal sentiment analysis and emotion recognition \u201darXiv preprint arXiv:2211.11256(2022).","DOI":"10.18653\/v1\/2022.emnlp-main.534"},{"key":"e_1_2_10_24_1","doi-asserted-by":"crossref","unstructured":"Y.Yu M.Zhao S.Qi et al. \u201cConKI: Contrastive Knowledge Injection for Multimodal Sentiment Analysis \u201darXiv preprint arXiv:2306.15796(2023).","DOI":"10.18653\/v1\/2023.findings-acl.860"},{"issue":"3","key":"e_1_2_10_25_1","doi-asserted-by":"crossref","first-page":"2276","DOI":"10.1109\/TAFFC.2022.3172360","article-title":"Hybrid Contrastive Learning of Tri\u2010Modal Representation for Multimodal Sentiment Analysis","volume":"14","author":"Mai S.","year":"2022","journal-title":"IEEE Transactions on Affective Computing"},{"key":"e_1_2_10_26_1","doi-asserted-by":"crossref","unstructured":"J.Tang K.Li X.Jin A.Cichocki Q.Zhao andW.Kong \u201cCTFN: Hierarchical Learning for Multimodal Sentiment Analysis Using Coupled\u2010Translation Fusion Network \u201d in\u2009Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)\u2009(2021) \u20095301\u20135311.","DOI":"10.18653\/v1\/2021.acl-long.412"},{"key":"e_1_2_10_27_1","doi-asserted-by":"crossref","first-page":"4730","DOI":"10.18653\/v1\/2021.findings-acl.417","volume-title":"Findings of the Association for Computational Linguistics: ACL\u2010IJCNLP 2021","author":"Wu Y.","year":"2021"},{"key":"e_1_2_10_28_1","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1016\/j.inffus.2022.11.022","article-title":"AOBERT: All\u2010Modalities\u2010In\u2010One BERT for Multimodal Sentiment Analysis","volume":"92","author":"Kim K.","year":"2023","journal-title":"Information Fusion"},{"key":"e_1_2_10_29_1","first-page":"2554","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lv F.","year":"2021"},{"key":"e_1_2_10_30_1","doi-asserted-by":"crossref","unstructured":"A.Zadeh M.Chen S.Poria E.Cambria andL. P.Morency \u201cTensor Fusion Network for Multimodal Sentiment Analysis \u201darXiv preprint arXiv:1707.07250(2017).","DOI":"10.18653\/v1\/D17-1115"},{"key":"e_1_2_10_31_1","doi-asserted-by":"crossref","unstructured":"J.Devlin M. W.Chang K.Lee andK.Toutanova \u201cBert: Pre\u2010Training of Deep Bidirectional Transformers for Language Understanding \u201d in\u2009Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies Volume 1 (Long and Short Papers).\u2009(2019) \u20094171\u20134186.","DOI":"10.18653\/v1\/N19-1423"},{"key":"e_1_2_10_32_1","doi-asserted-by":"crossref","unstructured":"W.Han H.Chen andS.Poria \u201cImproving Multimodal Fusion With Hierarchical Mutual Information Maximization for Multimodal Sentiment Analysis \u201darXiv preprint arXiv:2109.00412(2021).","DOI":"10.18653\/v1\/2021.emnlp-main.723"},{"key":"e_1_2_10_33_1","doi-asserted-by":"crossref","DOI":"10.1016\/j.inffus.2024.102725","article-title":"AtCAF: Attention\u2010Based Causality\u2010Aware Fusion Network for Multimodal Sentiment Analysis","volume":"114","author":"Huang C.","year":"2025","journal-title":"Information Fusion"},{"key":"e_1_2_10_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2025.3606964"},{"key":"e_1_2_10_35_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3736415","article-title":"Actual Cause Guided Adaptive Gradient Scaling for Balanced Multimodal Sentiment Analysis","volume":"21","author":"Chen J.","year":"2025","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"issue":"1","key":"e_1_2_10_36_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s13735-025-00362-y","article-title":"PAMoE\u2010MSA: Polarity\u2010Aware Mixture of Experts Network for Multimodal Sentiment Analysis","volume":"14","author":"Huang C.","year":"2025","journal-title":"International Journal of Multimedia Information Retrieval"},{"issue":"4","key":"e_1_2_10_37_1","doi-asserted-by":"crossref","DOI":"10.1007\/s40747-025-01806-y","article-title":"H 2 CAN: Heterogeneous Hypergraph Attention Network With Counterfactual Learning for Multimodal Sentiment Analysis","volume":"11","author":"Huang C.","year":"2025","journal-title":"Complex & Intelligent Systems"},{"issue":"4","key":"e_1_2_10_38_1","doi-asserted-by":"crossref","DOI":"10.1007\/s00530-024-01421-w","article-title":"Text\u2010Centered Cross\u2010Sample Fusion Network for Multimodal Sentiment Analysis","volume":"30","author":"Huang Q.","year":"2024","journal-title":"Multimedia Systems"},{"key":"e_1_2_10_39_1","doi-asserted-by":"crossref","unstructured":"H.Zhang Y.Wang G.Yin K.Liu Y.Liu andT.Yu \u201cLearning Language\u2010Guided Adaptive Hyper\u2010Modality Representation for Multimodal Sentiment Analysis \u201darXiv preprint arXiv:2310.05804(2023).","DOI":"10.18653\/v1\/2023.emnlp-main.49"}],"container-title":["Concurrency and Computation: Practice and Experience"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.70593","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1002\/cpe.70593","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.70593","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,5]],"date-time":"2026-02-05T03:06:29Z","timestamp":1770260789000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/cpe.70593"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,28]]},"references-count":38,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,2]]}},"alternative-id":["10.1002\/cpe.70593"],"URL":"https:\/\/doi.org\/10.1002\/cpe.70593","archive":["Portico"],"relation":{},"ISSN":["1532-0626","1532-0634"],"issn-type":[{"value":"1532-0626","type":"print"},{"value":"1532-0634","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,28]]},"assertion":[{"value":"2025-09-16","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-09","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-28","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70593"}}