{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,19]],"date-time":"2026-03-19T09:23:59Z","timestamp":1773912239197,"version":"3.50.1"},"reference-count":34,"publisher":"Oxford University Press (OUP)","issue":"3","license":[{"start":{"date-parts":[[2025,11,2]],"date-time":"2025-11-02T00:00:00Z","timestamp":1762041600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/pages\/standard-publication-reuse-rights"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61972267"],"award-info":[{"award-number":["61972267"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026,3,16]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>In the field of multimodal sentiment analysis, it is an important research task to fully extract modal features and perform efficient fusion. Traditional sentiment classification models usually have the problems of insufficient interaction information, easy to be affected by visual or perceptual interference, and lack of stability. To address these issues, this paper proposes an interactive gated attention network for multimodal sentiment analysis. Firstly, the feature extraction network based on the pretraining model is constructed to obtain high-dimensional feature vectors. Secondly, the mutual feature vector generator is designed to generate interactive feature vectors with high levels. Furthermore, the gating vector generator is designed to highlight the differences in semantic features. Finally, attention of the interaction mechanism is introduced to focus on the important parts between features, and the feature enhancement and fusion are realized. Comprehensive experiments and analysis on the Twitter-15 and Twitter-17 datasets show that the proposed model is superior to a series of comparative models in information interaction, semantic enhancement, and feature fusion.<\/jats:p>","DOI":"10.1093\/comjnl\/bxaf130","type":"journal-article","created":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T12:12:51Z","timestamp":1760011971000},"page":"549-560","source":"Crossref","is-referenced-by-count":0,"title":["Multimodal sentiment analysis with interactive gated attention network"],"prefix":"10.1093","volume":"69","author":[{"given":"Yu","family":"Lei","sequence":"first","affiliation":[{"name":"School of Information Science and Technology, Shijiazhuang Tiedao University , 17 North Second Ring East Road, Changan District, Shijiazhuang City, Hebei Province 050043 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bowen","family":"Song","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Shijiazhuang Tiedao University , 17 North Second Ring East Road, Changan District, Shijiazhuang City, Hebei Province 050043 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shihui","family":"Li","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Shijiazhuang Tiedao University , 17 North Second Ring East Road, Changan District, Shijiazhuang City, Hebei Province 050043 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chongxuan","family":"Su","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Shijiazhuang Tiedao University , 17 North Second Ring East Road, Changan District, Shijiazhuang City, Hebei Province 050043 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tianshuo","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Shijiazhuang Tiedao University , 17 North Second Ring East Road, Changan District, Shijiazhuang City, Hebei Province 050043 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinyu","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Shijiazhuang Tiedao University , 17 North Second Ring East Road, Changan District, Shijiazhuang City, Hebei Province 050043 ,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2025,11,2]]},"reference":[{"key":"2026031900561306900_ref1","first-page":"5105","article-title":"Multi-level attention map network for multimodal sentiment analysis","volume":"35","author":"Xue","year":"2023","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"2026031900561306900_ref2","doi-asserted-by":"publisher","first-page":"423","DOI":"10.1109\/TPAMI.2018.2798607","article-title":"Multimodal machine learning: a survey and taxonomy","volume":"41","author":"Baltru\u0161aitis","year":"2019","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"2026031900561306900_ref3","first-page":"1","article-title":"Fine-grained sentiment classification using bert","volume-title":"Proceedings of the 1st International Conference on Artificial Intelligence for Transforming Business and Society (AITB)","author":"Munikar","year":"2019"},{"key":"2026031900561306900_ref4","article-title":"RoBERTa: a robustly optimized bert pretraining approach","author":"Liu","year":"2019"},{"key":"2026031900561306900_ref5","first-page":"1103","article-title":"Tensor fusion network for multimodal sentiment analysis","volume-title":"Proceedings of the 2017 International Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Zadeh","year":"2017"},{"key":"2026031900561306900_ref6","article-title":"Neural machine translation by jointly learning to align and translate","author":"Bahdanau","year":"2014"},{"key":"2026031900561306900_ref7","first-page":"6077","article-title":"Bottom-up and top-down attention for image captioning and visual question answering","volume-title":"Proceedings of the 2018 IEEE\/CVF International Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Anderson","year":"2018"},{"key":"2026031900561306900_ref8","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/2020.challengehml-1.1","article-title":"A transformer-based joint-encoding for emotion recognition and sentiment analysis","volume-title":"Proceedings of the 58th International Conference of the Association for Computational Linguistics (ACL)","author":"Delbrouck","year":"2020"},{"key":"2026031900561306900_ref9","doi-asserted-by":"publisher","first-page":"1901","DOI":"10.1109\/TCSVT.2020.3014889","article-title":"Multimodal local-global attention network for affective video content analysis","volume":"31","author":"Yangjun","year":"2021","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"2026031900561306900_ref10","doi-asserted-by":"crossref","first-page":"157329","DOI":"10.1109\/ACCESS.2021.3126782","article-title":"Targeted aspect-based multimodal sentiment analysis: an attention capsule extraction and multi-head fusion network","volume":"9","author":"Donghong","year":"2021","journal-title":"IEEE Access"},{"key":"2026031900561306900_ref11","first-page":"770","article-title":"Deep residual learning for image recognition","volume-title":"Proceedings of the 2016 IEEE International Conference on Computer Vision and Pattern Recognition (CVPR)","author":"He","year":"2016"},{"key":"2026031900561306900_ref12","doi-asserted-by":"publisher","first-page":"212","DOI":"10.1016\/j.ins.2020.12.024","article-title":"CLAVER: an integrated framework of convolutional layer, bidirectional LSTM with attention mechanism based scholarly venue recommendation","volume":"559","author":"Pradhan","year":"2021","journal-title":"Inform Sci"},{"key":"2026031900561306900_ref13","first-page":"371","article-title":"Multi-interactive memory network for aspect based multimodal sentiment analysis","volume-title":"Proceedings of the 33rd AAAI Conference on Artificial Intelligence (AAAI)","author":"Xu","year":"2019"},{"key":"2026031900561306900_ref14","first-page":"6558","article-title":"Multimodal transformer for unaligned multimodal language sequences","volume-title":"Proceedings of the 57th International Conference of the Association for Computational Linguistics (ACL)","author":"Tsai","year":"2019"},{"key":"2026031900561306900_ref15","first-page":"5642","article-title":"Multi-attention recurrent network for human communication comprehension","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence. (AAAI)","author":"Zadeh","year":"2018"},{"key":"2026031900561306900_ref16","doi-asserted-by":"publisher","first-page":"429","DOI":"10.1109\/TASLP.2019.2957872","article-title":"Entity-sensitive attention and fusion network for entity-level multimodal sentiment classification","volume":"28","author":"Jianfei","year":"2020","journal-title":"IEEE\/ACM Trans Audio Speech Lang Process"},{"key":"2026031900561306900_ref17","doi-asserted-by":"crossref","DOI":"10.1109\/ICASSP40776.2020.9053012","article-title":"Gated mechanism for attention based multi modal sentiment analysis","volume-title":"Proceedings of the 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Kumar","year":"2020"},{"key":"2026031900561306900_ref18","article-title":"Using large pre-trained models with cross-modal attention for multi-modal emotion recognition","author":"Freshworks","year":"2021"},{"key":"2026031900561306900_ref19","doi-asserted-by":"crossref","first-page":"521","DOI":"10.1145\/3394171.3413690","article-title":"CM-BERT: cross-modal bert for text-audio sentiment analysis","volume-title":"Proceedings of the 28th ACM International Conference on Multimedia (MM)","author":"Yang","year":"2020"},{"key":"2026031900561306900_ref20","first-page":"13","article-title":"ViLBERT: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems (NeurIPS)","author":"Lu","year":"2019"},{"key":"2026031900561306900_ref21","article-title":"VAuLT: augmenting the vision-and-language transformer for sentiment classification on social media","author":"Chochlakis","year":"2022"},{"key":"2026031900561306900_ref22","doi-asserted-by":"publisher","first-page":"106136","DOI":"10.1016\/j.engappai.2023.106136","article-title":"An ALBERT-based TextCNN-Hatt hybrid model enhanced with topic knowledge for sentiment analysis of sudden-onset disasters","volume":"123","author":"Zhang","year":"2023","journal-title":"Eng Appl Artif Intel"},{"key":"2026031900561306900_ref23","first-page":"5753","article-title":"XLNet: generalized autoregressive pretraining for language understanding","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems (NeurIPS)","author":"Yang","year":"2019"},{"key":"2026031900561306900_ref24","article-title":"DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter","author":"Sanh","year":"2019"},{"key":"2026031900561306900_ref25","doi-asserted-by":"publisher","first-page":"107408","DOI":"10.1016\/j.knosys.2021.107408","article-title":"Kvl-BERT: knowledge enhanced visual-and-linguistic bert for visual commonsense reasoning","volume":"230","author":"Song","year":"2021","journal-title":"Knowledge-Based Syst"},{"key":"2026031900561306900_ref26","article-title":"Cross-modality gated attention fusion for multimodal sentiment analysis","author":"Jiang","year":"2022"},{"key":"2026031900561306900_ref27","doi-asserted-by":"crossref","first-page":"3034","DOI":"10.1145\/3474085.3475692","article-title":"Exploiting BERT for multimodal target sentiment classification through input space translation","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia (MM)","author":"Khan","year":"2021"},{"key":"2026031900561306900_ref28","doi-asserted-by":"publisher","first-page":"1966","DOI":"10.1109\/TAFFC.2022.3171091","article-title":"Hierarchical interactive multimodal transformer for aspect-based multimodal sentiment analysis","volume":"14","author":"Jianfei","year":"2023","journal-title":"IEEE Trans Affective Comput"},{"key":"2026031900561306900_ref29","doi-asserted-by":"publisher","first-page":"1966","DOI":"10.1109\/TCSVT.2022.3218018","article-title":"BAFN: bi-direction attention based fusion network for multimodal sentiment analysis","volume":"33","author":"Tang","year":"2023","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"2026031900561306900_ref30","doi-asserted-by":"publisher","first-page":"101973","DOI":"10.1016\/j.inffus.2023.101973","article-title":"Modality translation-based multimodal sentiment analysis under uncertain missing modalities","volume":"101","author":"Liu","year":"2024","journal-title":"Inf Fusion"},{"key":"2026031900561306900_ref31","doi-asserted-by":"publisher","first-page":"109731","DOI":"10.1016\/j.engappai.2024.109731","article-title":"Multimodal sentiment analysis based on multiple attention","volume":"140","author":"Wang","year":"2025","journal-title":"Eng Appl Artif Intel"},{"key":"2026031900561306900_ref32","first-page":"1","article-title":"Comparative study of BERT models and RoBERTa in transformer based question answering","volume-title":"Proceedings of the 3rd International Conference on Intelligent Technologies (CONIT)","author":"Akhila","year":"2023"},{"key":"2026031900561306900_ref33","first-page":"1","article-title":"An enhanced context-based emotion detection model using RoBERTa","volume-title":"Proceedings of the 2022 IEEE International Conference on Electronics, Computing and Communication Technologies (CONECCT)","author":"Kamath","year":"2022"},{"key":"2026031900561306900_ref34","first-page":"5343","article-title":"Adapting BERT for target-oriented multimodal sentiment classification","volume-title":"Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI)","author":"Yu","year":"2019"}],"container-title":["The Computer Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/69\/3\/549\/65092674\/bxaf130.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/69\/3\/549\/65092674\/bxaf130.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,19]],"date-time":"2026-03-19T04:56:25Z","timestamp":1773896185000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/comjnl\/article\/69\/3\/549\/8312892"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,2]]},"references-count":34,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2025,11,2]]},"published-print":{"date-parts":[[2026,3,16]]}},"URL":"https:\/\/doi.org\/10.1093\/comjnl\/bxaf130","relation":{},"ISSN":["0010-4620","1460-2067"],"issn-type":[{"value":"0010-4620","type":"print"},{"value":"1460-2067","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026,3]]},"published":{"date-parts":[[2025,11,2]]}}}