{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,12]],"date-time":"2026-07-12T02:32:49Z","timestamp":1783823569700,"version":"3.55.0"},"reference-count":43,"publisher":"American Association for the Advancement of Science (AAAS)","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No.61972437"],"award-info":[{"award-number":["No.61972437"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No.62201002"],"award-info":[{"award-number":["No.62201002"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Excellent Youth Foundation of Anhui Scientific Committee","award":["No.2208085J05"],"award-info":[{"award-number":["No.2208085J05"]}]},{"name":"Special Fund for Key Program of Science and Technology of Anhui Province","award":["No.202203a07020008"],"award-info":[{"award-number":["No.202203a07020008"]}]},{"name":"the Open Research Projects of Zhejiang Lab","award":["No.2021KH0AB06"],"award-info":[{"award-number":["No.2021KH0AB06"]}]}],"content-domain":{"domain":["spj.science.org"],"crossmark-restriction":true},"short-container-title":["Intell Comput"],"published-print":{"date-parts":[[2024,1]]},"abstract":"<jats:p>Transformer-based methods have achieved superior performance in multimodal sentiment analysis (MSA). However, recent studies have used it more to focus on the potential mutual adaptation between unimodal modalities, while ignoring intra-modal interactions and complementarity between potential fusion representations. To address this problem, this study proposes a two-stage stacked transformer framework for MSA. The framework decomposes the fusion into two stages, each of which concentrates on a subset of multimodal signals to simultaneously capture the communication information between the unimodal modalities and the interaction information between fusion representations. Further, stacked transformers are the core components of the framework. Two transformer layers are used to model the cross-modal interaction and intra-modal interaction of multimodal input. For cross-modal attention, we propose an attention weight accumulation mechanism to further improve the ability of the framework. Experimental results on 3 benchmark datasets (MOSI [Multimodal Opinion-level Sentiment Intensity], MOSEI [Multimodal Opinion Sentiment and Emotion Intensity], and SIMS) show performance superior or comparable to the state-of-the-art models, thus demonstrating the effectiveness of the proposed framework.<\/jats:p>","DOI":"10.34133\/icomputing.0081","type":"journal-article","created":{"date-parts":[[2024,4,12]],"date-time":"2024-04-12T11:11:18Z","timestamp":1712920278000},"update-policy":"https:\/\/doi.org\/10.34133\/aaas_crossmark_01","source":"Crossref","is-referenced-by-count":11,"title":["A Two-Stage Stacked Transformer Framework for Multimodal Sentiment Analysis"],"prefix":"10.34133","volume":"3","author":[{"given":"Guofeng","family":"Yi","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, \rAnhui University, Hefei, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cunhang","family":"Fan","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, \rAnhui University, Hefei, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianhua","family":"Tao","sequence":"additional","affiliation":[{"name":"Department of Automation, \rTsinghua University, Beijing, China."},{"name":"School of Artificial Intelligence, \rUniversity of Chinese Academy of Sciences, Beijing, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhao","family":"Lv","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, \rAnhui University, Hefei, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhengqi","family":"Wen","sequence":"additional","affiliation":[{"name":"Qiyuan Laboratory, Beijing, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guanxiong","family":"Pei","sequence":"additional","affiliation":[{"name":"Institute of Artificial Intelligence, Zhejiang Laboratory, Hangzhou, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Taihao","family":"Li","sequence":"additional","affiliation":[{"name":"Institute of Artificial Intelligence, Zhejiang Laboratory, Hangzhou, China."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"221","published-online":{"date-parts":[[2024,5,24]]},"reference":[{"issue":"1","key":"e_1_3_3_2_2","doi-asserted-by":"crossref","first-page":"108","DOI":"10.1109\/TAFFC.2020.3038167","article-title":"Beneath the tip of the iceberg: Current challenges and new directions in sentiment analysis research","volume":"14","author":"Poria S","year":"2020","unstructured":"Poria S, Hazarika D, Majumder N, Mihalcea R. Beneath the tip of the iceberg: Current challenges and new directions in sentiment analysis research. IEEE Trans Affect Comput. 2020;14(1):108\u2013132.","journal-title":"IEEE Trans Affect Comput"},{"key":"e_1_3_3_3_2","doi-asserted-by":"crossref","unstructured":"Morency LP Mihalcea R Doshi P. Towards multimodal sentiment analysis: Harvesting opinions from the web. Paper presented at: ICMI '11: Proceedings of the 13th International Conference on Multimodal Interfaces; 2011 Nov 11\u201318; Alicante Spain.","DOI":"10.1145\/2070481.2070509"},{"key":"e_1_3_3_4_2","doi-asserted-by":"crossref","unstructured":"Yang K Xu H Gao K. CM-BERT: Cross-modal BERT for text-audio sentiment analysis. Paper presented at: MM '20: Proceedings of the 28th ACM International Conference on Multimedia; 2020 Oct 12\u201316 Seattle WA.","DOI":"10.1145\/3394171.3413690"},{"issue":"2","key":"e_1_3_3_5_2","doi-asserted-by":"crossref","first-page":"1334","DOI":"10.1109\/TAFFC.2021.3097002","article-title":"The multimodal sentiment analysis in car reviews (MuSe-CaR) dataset: Collection, insights and improvements","volume":"14","author":"Stappen L","year":"2021","unstructured":"Stappen L, Baird A, Schumann L, Bjorn S. The multimodal sentiment analysis in car reviews (MuSe-CaR) dataset: Collection, insights and improvements. IEEE Trans Affect Comput. 2021;14(2):1334\u20131350.","journal-title":"IEEE Trans Affect Comput"},{"key":"e_1_3_3_6_2","doi-asserted-by":"crossref","unstructured":"Chen M Wang S Liang PP Baltru\u0161aitis T Zadeh A Morency LP. Multimodal sentiment analysis with word-level fusion and reinforcement learning. Paper presented at: ICMI '17: Proceedings of the 19th ACM International Conference on Multimodal Interaction; 2017 Nov 13\u201317; Glasgow Scotland.","DOI":"10.1145\/3136755.3136801"},{"key":"e_1_3_3_7_2","doi-asserted-by":"crossref","unstructured":"Tsai YHH Bai S Liang PP Kolter JZ Morency LP Salakhutdinov R. Multimodal transformer for unaligned multimodal language sequences. arXiv. 2019. https:\/\/doi.org\/10.48550\/arXiv.1906.00295","DOI":"10.18653\/v1\/P19-1656"},{"key":"e_1_3_3_8_2","doi-asserted-by":"crossref","unstructured":"Han W Chen H Gelbukh A Zadeh A Morency LP Poria S. Bi-bimodal modality fusion for correlation-controlled multimodal sentiment analysis. Paper presented at: ICMI '21: Proceedings of the 2021 International Conference on Multimodal Interaction; 2021 Oct 18\u201322 Montreal Canada.","DOI":"10.1145\/3462244.3479919"},{"issue":"2","key":"e_1_3_3_9_2","doi-asserted-by":"crossref","first-page":"423","DOI":"10.1109\/TPAMI.2018.2798607","article-title":"Multimodal machine learning: A survey and taxonomy","volume":"41","author":"Baltru\u0161aitis T","year":"2018","unstructured":"Baltru\u0161aitis T, Ahuja C, Morency LP. Multimodal machine learning: A survey and taxonomy. IEEE Trans Pattern Anal Mach Intell. 2018;41(2):423\u2013443.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"e_1_3_3_10_2","unstructured":"Hai P Manzini T Liang PP Barnab\u00e1s P. Seq2Seq2Sentiment: Multimodal sequence to sequence models for sentiment analysis. Paper presented at: Proceedings of Grand Challenge and Workshop on Human Multimodal Language (Challenge-HML); 2018 Jul; Melbourne Australia."},{"key":"e_1_3_3_11_2","doi-asserted-by":"crossref","unstructured":"Hai P Liang PP. Manzini T Morency LP P\u00f3czos B. Found in translation: Learning robust joint representations by cyclic translations between modalities. Paper presented at: 33rd AAAI Conference on Artificial Intelligence; 2019 Jan 27\u2013Feb 1; Honolulu HI.","DOI":"10.1609\/aaai.v33i01.33016892"},{"key":"e_1_3_3_12_2","unstructured":"Kenton JDMWC Toutanova LK. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv. 2018. https:\/\/doi.org\/10.48550\/arXiv.1810.04805"},{"key":"e_1_3_3_13_2","doi-asserted-by":"crossref","unstructured":"Sahay S Okur E Kumar SH Nachman L. Low rank fusion based transformers for multimodal sequences. arXiv. 2020. https:\/\/doi.org\/10.48550\/arXiv.2007.02038","DOI":"10.18653\/v1\/2020.challengehml-1.4"},{"key":"e_1_3_3_14_2","doi-asserted-by":"crossref","unstructured":"Liang PP Liu Z Zadeh A Morency LP. Multimodal language analysis with recurrent multistage fusion. arXiv. 2018. https:\/\/doi.org\/10.48550\/arXiv.1808.03920","DOI":"10.18653\/v1\/D18-1014"},{"key":"e_1_3_3_15_2","unstructured":"Qin Z Sun W Deng H Li D Wei Y Lv B Yan J Kong L Zhong Y. cosFormer: Rethinking softmax in attention. arXiv. 2022. https:\/\/doi.org\/10.48550\/arXiv.2202.08791"},{"key":"e_1_3_3_16_2","doi-asserted-by":"crossref","unstructured":"Zadeh A Chen M Poria S Cambria E Morency LP. Tensor fusion network for multimodal sentiment analysis. arXiv. 2017. https:\/\/doi.org\/10.48550\/arXiv.1707.07250","DOI":"10.18653\/v1\/D17-1115"},{"key":"e_1_3_3_17_2","doi-asserted-by":"crossref","unstructured":"Degottex G Kane J Drugman T Raitio T Scherer S. COVAREP\u2014A collaborative voice analysis repository for speech technologies. Paper presented at: 2014 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2014 May 04\u201309; Florence Italy.","DOI":"10.1109\/ICASSP.2014.6853739"},{"key":"e_1_3_3_18_2","unstructured":"Vaswani A Shazeer N Parmar N Uszkoreit J Jones L Gomez AN Kaiser L Polosukhin I Attention is all you need. arXiv. 2017. https:\/\/doi.org\/10.48550\/arXiv.1706.03762"},{"key":"e_1_3_3_19_2","doi-asserted-by":"crossref","unstructured":"Chen MX Firat O Bapna A Johnson M Macherey W Foster G Jones L Schuster M Shazeer N Parmar N. The best of both worlds: Combining recent advances in neural machine translation. Paper presented at: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); 2018 Jul 15\u201320 Melbourne Australia.","DOI":"10.18653\/v1\/P18-1008"},{"key":"e_1_3_3_20_2","doi-asserted-by":"crossref","unstructured":"Hazarika D Zimmermann R Poria S. MISA: Modality-invariant and-specific representations for multimodal sentiment analysis. arXiv. 2020. https:\/\/doi.org\/10.48550\/arXiv.2005.03545","DOI":"10.1145\/3394171.3413678"},{"key":"e_1_3_3_21_2","doi-asserted-by":"crossref","unstructured":"Lazaridou A Pham NT and Baroni M. Combining language and vision with a multimodal skip-gram model. In: Mihalcea R Chai J Sarkar A editors. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Denver (CO): Association for Computational Linguistics; 2015. p. 153\u2013163.","DOI":"10.3115\/v1\/N15-1016"},{"key":"e_1_3_3_22_2","doi-asserted-by":"crossref","unstructured":"Liu Z Shen Y Lakshminarasimhan VB Liang PP Zadeh A Morency LP. Efficient low-rank multimodal fusion with modality-specific factors. arXiv. 2018. https:\/\/doi.org\/10.48550\/arXiv.1806.00064","DOI":"10.18653\/v1\/P18-1209"},{"key":"e_1_3_3_23_2","doi-asserted-by":"crossref","first-page":"2359","DOI":"10.18653\/v1\/2020.acl-main.214","article-title":"Integrating multimodal information in large pretrained transformers","volume":"2020","author":"Rahman W","year":"2020","unstructured":"Rahman W, Hasan MK, Lee S, Zadeh A, Mao C, Morency LP, Hoque E. Integrating multimodal information in large pretrained transformers. Proc Conf Assoc Comput Linguist Meet. 2020;2020:2359\u20132369.","journal-title":"Proc Conf Assoc Comput Linguist Meet"},{"key":"e_1_3_3_24_2","doi-asserted-by":"crossref","DOI":"10.1016\/j.knosys.2022.110021","article-title":"Sentiment-aware multimodal pre-training for multimodal sentiment analysis","volume":"258","author":"Ye J","year":"2022","unstructured":"Ye J, Zhou J, Tian J, Wang R, Zhou J, Gui T, Zhang Q, Huang X. Sentiment-aware multimodal pre-training for multimodal sentiment analysis. Knowl-Based Syst. 2022;258: Article 110021.","journal-title":"Knowl-Based Syst"},{"key":"e_1_3_3_25_2","doi-asserted-by":"crossref","unstructured":"Yu W Xu H Yuan Z and Wu J. Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis. arXiv. 2021. https:\/\/doi.org\/10.48550\/arXiv.2102.04830","DOI":"10.1609\/aaai.v35i12.17289"},{"key":"e_1_3_3_26_2","doi-asserted-by":"crossref","first-page":"2015","DOI":"10.1109\/TASLP.2022.3178204","article-title":"Multimodal sentiment analysis with two-phase multi-task learning","volume":"30","author":"Yang B","year":"2022","unstructured":"Yang B, Wu L, Zhu J, Shao B, Lin X, Liu TY. Multimodal sentiment analysis with two-phase multi-task learning. IEEE\/ACM Trans. Audio Speech Lang. 2022;30:2015\u20132024.","journal-title":"IEEE\/ACM Trans. Audio Speech Lang"},{"key":"e_1_3_3_27_2","doi-asserted-by":"crossref","unstructured":"Parikh AP T\u00e4ckstr\u00f6m O Das D Uszkoreit J. A decomposable attention model for natural language inference. arXiv. 2016. https:\/\/doi.org\/10.48550\/arXiv.1606.01933","DOI":"10.18653\/v1\/D16-1244"},{"key":"e_1_3_3_28_2","unstructured":"Lin Z Feng M Santos CN dos Yu M Xiang B Zhou B Bengio Y. A structured self-attentive sentence embedding. arXiv. 2017. https:\/\/doi.org\/10.48550\/arXiv.1703.03130"},{"key":"e_1_3_3_29_2","unstructured":"Dosovitskiy A Beyer L Kolesnikov A Weissenborn D Zhai X Unterthiner T Dehghani M Minderer M Heigold G Gelly S et\u00a0al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv. 2020. https:\/\/doi.org\/10.48550\/arXiv.2010.11929"},{"key":"e_1_3_3_30_2","doi-asserted-by":"crossref","unstructured":"Dong L Xu S Xu B. Speech-transformer: A no-recurrence sequence-to-sequence model for speech recognition. Paper presented at: 2018 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP); 2018 Apr 15\u201320; Calgary AB Canada.","DOI":"10.1109\/ICASSP.2018.8462506"},{"key":"e_1_3_3_31_2","doi-asserted-by":"crossref","unstructured":"Li LH Yatskar M Yin D Hsieh CJ Chang KW. What does BERT with vision look at? Paper presented at: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics; 2020 Jul 5\u201310; online.","DOI":"10.18653\/v1\/2020.acl-main.469"},{"key":"e_1_3_3_32_2","doi-asserted-by":"crossref","unstructured":"Poria S Cambria E Hazarika D Majumder N Zadeh A Morency LP. Context-dependent sentiment analysis in user-generated videos. Paper presented at: Proceedings of the 55th annual meeting of the association for computational linguistics (volume 1: Long papers); 2017 Jul 30\u2013Aug 4 Vancouver Canada.","DOI":"10.18653\/v1\/P17-1081"},{"key":"e_1_3_3_33_2","doi-asserted-by":"crossref","unstructured":"Zadeh A Liang PP Mazumder N Poria S Cambria E Morency LP. Memory fusion network for multi-view sequential learning. arXiv. 2018. https:\/\/doi.org\/10.48550\/arXiv.1802.00927","DOI":"10.1609\/aaai.v32i1.12021"},{"issue":"6","key":"e_1_3_3_34_2","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1109\/MIS.2016.94","article-title":"Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages","volume":"31","author":"Zadeh A","year":"2016","unstructured":"Zadeh A, Zellers R, Pincus E, Morency LP. Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages. IEEE Intell Syst. 2016;31(6):82\u201388.","journal-title":"IEEE Intell Syst"},{"key":"e_1_3_3_35_2","unstructured":"Zadeh A Pu P. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph. Paper presented at: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Long Papers); 2018 Jul 15\u201320 Melbourne Australia."},{"key":"e_1_3_3_36_2","doi-asserted-by":"crossref","unstructured":"Yu W Xu H Meng F Zhu Y Ma Y Wu J Zou J Yang K. CH-SIMS: A Chinese multimodal sentiment analysis dataset with fine-grained annotation of modality. Paper presented at: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics; 2020 Jul 5\u201310; online.","DOI":"10.18653\/v1\/2020.acl-main.343"},{"key":"e_1_3_3_37_2","doi-asserted-by":"crossref","unstructured":"Han W Chen H Poria S. Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis. Paper presented at: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics; 2021 Aug 1\u20136; virtual.","DOI":"10.18653\/v1\/2021.emnlp-main.723"},{"key":"e_1_3_3_38_2","unstructured":"Tsai YHH Liang PP Zadeh A Morency LP Salakhutdinov R. Learning factorized multimodal representations. arXiv. 2018."},{"key":"e_1_3_3_39_2","unstructured":"Radford A Narasimhan K Salimans T Sutskever I. Improving language understanding by generative pre-training. 2018. https:\/\/www.mikecaptain.com\/resources\/pdf\/GPT-1.pdf"},{"key":"e_1_3_3_40_2","doi-asserted-by":"crossref","unstructured":"Liu Z Lin Y Cao Y Hu H Wei Y Zhang Z Lin S Guo B. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv. 2021. https:\/\/doi.org\/10.48550\/arXiv.2103.14030","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_3_41_2","doi-asserted-by":"crossref","unstructured":"Chen CF Fan Q and Panda R. CrossViT: Cross-attention multi-scale vision transformer for image classification. arXiv. 2021. https:\/\/doi.org\/10.48550\/arXiv.2103.14899","DOI":"10.1109\/ICCV48922.2021.00041"},{"key":"e_1_3_3_42_2","unstructured":"Akbari H Yuan L Qian R Chuang WH Chang SF Cui Y Gong B Vatt: Transformers for multimodal self-supervised learning from raw video audio and text. arXiv. 2021. https:\/\/doi.org\/10.48550\/arXiv.2104.11178"},{"key":"e_1_3_3_43_2","doi-asserted-by":"crossref","unstructured":"Li Z Xu B Zhu C Zhao T. CLMLF: A contrastive learning and multi-layer fusion method for multimodal sentiment detection. In: Findings of the Association for Computational Linguistics: NAACL 2022. Seattle (WA): Association for Computational Linguistics; 2022.","DOI":"10.18653\/v1\/2022.findings-naacl.175"},{"key":"e_1_3_3_44_2","doi-asserted-by":"crossref","first-page":"4909","DOI":"10.1109\/TMM.2022.3183830","article-title":"Cross-modal enhancement network for multimodal sentiment analysis","volume":"25","author":"Wang D","year":"2022","unstructured":"Wang D, Liu S, Wang Q, Tian Y, He L, Gao X. Cross-modal enhancement network for multimodal sentiment analysis. IEEE Trans Multimed. 2022;25:4909\u20134921.","journal-title":"IEEE Trans Multimed"}],"container-title":["Intelligent Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/spj.science.org\/doi\/pdf\/10.34133\/icomputing.0081","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,24]],"date-time":"2024-05-24T11:03:39Z","timestamp":1716548619000},"score":1,"resource":{"primary":{"URL":"https:\/\/spj.science.org\/doi\/10.34133\/icomputing.0081"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1]]},"references-count":43,"alternative-id":["10.34133\/icomputing.0081"],"URL":"https:\/\/doi.org\/10.34133\/icomputing.0081","relation":{},"ISSN":["2771-5892"],"issn-type":[{"value":"2771-5892","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1]]},"assertion":[{"value":"2023-05-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-04","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-05-24","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"0081"}}