{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T02:15:06Z","timestamp":1777601706380,"version":"3.51.4"},"reference-count":31,"publisher":"Oxford University Press (OUP)","issue":"6","license":[{"start":{"date-parts":[[2024,1,31]],"date-time":"2024-01-31T00:00:00Z","timestamp":1706659200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/pages\/standard-publication-reuse-rights"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61972267"],"award-info":[{"award-number":["61972267"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2024,6,24]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>In the field of multimodal sentiment analysis, it is an important research task to fully extract modal features and perform efficient fusion. In response to the problems of insufficient semantic information and poor cross-modal fusion effect of traditional sentiment classification models, this paper proposes a composite hierarchical feature fusion method combined with prior knowledge. Firstly, the ALBERT (A Lite BERT) model and the improved ResNet model are constructed for feature extraction of text and image, respectively, and high-dimensional feature vectors are obtained. Secondly, to solve the problem of insufficient semantic information expression in cross-scene, a prior knowledge enhancement model is proposed to enrich the data characteristics of each modality. Finally, to solve the problem of poor cross-modal fusion effect, a composite hierarchical fusion model is proposed, which combines the temporal convolutional network and the attention mechanism to fuse the sequence features of each modality information and realizes the information interaction between different modalities. Experiments on MVSA-Single and MVSA-Multi datasets show that the proposed model is superior to a series of comparison models and has good adaptability in new scenarios.<\/jats:p>","DOI":"10.1093\/comjnl\/bxae002","type":"journal-article","created":{"date-parts":[[2024,2,1]],"date-time":"2024-02-01T18:06:56Z","timestamp":1706810816000},"page":"2230-2245","source":"Crossref","is-referenced-by-count":7,"title":["Multimodal Sentiment Analysis Based on Composite Hierarchical Fusion"],"prefix":"10.1093","volume":"67","author":[{"given":"Yu","family":"Lei","sequence":"first","affiliation":[{"name":"School of Information Science and Technology, Shijiazhuang Tiedao University , Shijiazhuang 050043 , China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Keshuai","family":"Qu","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Shijiazhuang Tiedao University , Shijiazhuang 050043 , China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yifan","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Shijiazhuang Tiedao University , Shijiazhuang 050043 , China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qing","family":"Han","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Shijiazhuang Tiedao University , Shijiazhuang 050043 , China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xuguang","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Shijiazhuang Tiedao University , Shijiazhuang 050043 , China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2024,1,31]]},"reference":[{"key":"2024062414160014600_ref1","doi-asserted-by":"crossref","first-page":"153072","DOI":"10.1109\/ACCESS.2021.3122025","article-title":"Urdu sentiment analysis via multimodal data mining based on deep learning algorithms","volume":"9","author":"Sehar","year":"2021","journal-title":"IEEE Access"},{"key":"2024062414160014600_ref2","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1016\/j.inffus.2021.10.013","article-title":"Multi-feature, multimodal, and multi-source social event detection: a comprehensive survey","volume":"79","author":"Afyouni","year":"2022","journal-title":"Inf. Fusion"},{"key":"2024062414160014600_ref3","doi-asserted-by":"crossref","first-page":"664","DOI":"10.26599\/TST.2021.9010055","article-title":"Cross-modal complementary network with hierarchical fusion for multimodal sentiment classification","volume":"4","author":"Peng","year":"2022","journal-title":"Tsinghua Sci. Technol."},{"key":"2024062414160014600_ref4","doi-asserted-by":"crossref","first-page":"1473","DOI":"10.1109\/SMC52423.2021.9658940","article-title":"Multimodal sentiment analysis based on attention mechanism and tensor fusion network","volume-title":"Proc. of the 2021 IEEE Int. Conf. on Systems, Man and Cybernetics (SMC)","author":"Zhang","year":"2021"},{"key":"2024062414160014600_ref5","first-page":"4360","article-title":"Multimodal sentiment analysis based on transformer and low-rank fusion","volume-title":"Proc. of the 2021 IEEE Int. Conf. on China Automation Congress (CAC)","author":"Shan","year":"2021"},{"key":"2024062414160014600_ref6","first-page":"1","article-title":"A comprehensive survey on community detection with deep learning","volume":"3","author":"Su","year":"2022","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"2024062414160014600_ref7","doi-asserted-by":"crossref","first-page":"467","DOI":"10.1109\/SLT.2016.7846305","article-title":"A prioritized grid long short-term memory RNN for speech recognition","volume-title":"Proc. of the 2016 IEEE Int. Conf. on Spoken Language Technology Workshop (SLT)","author":"Hsu","year":"2016"},{"key":"2024062414160014600_ref8","doi-asserted-by":"crossref","first-page":"212","DOI":"10.1016\/j.ins.2020.12.024","article-title":"CLAVER: an integrated framework of convolutional layer, bidirectional LSTM with attention mechanism based scholarly venue recommendation","volume":"559","author":"Pradhan","year":"2020","journal-title":"Inform. Sci."},{"key":"2024062414160014600_ref9","first-page":"9710","article-title":"Leveraging Arabic sentiment classification using an enhanced CNN-LSTM approach and effective Arabic text preparation","volume":"34","author":"Alayba","year":"2022","journal-title":"J. King Saud Univ. - Comput. Inf. Sci."},{"key":"2024062414160014600_ref10","first-page":"4171","article-title":"BERT: pre-training of deep bidirectional transformers for language understanding","volume":"6","author":"Devlin","year":"2019","journal-title":"North American Chapter of the Association for Computational Linguistics."},{"key":"2024062414160014600_ref11","first-page":"770","article-title":"Deep residual learning for image recognition","volume-title":"Proc. of the 2016 IEEE Int. Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"He","year":"2016"},{"key":"2024062414160014600_ref12","doi-asserted-by":"crossref","first-page":"1279","DOI":"10.1093\/comjnl\/bxac013","article-title":"An improved self attention mechanism based on optimized BERT-BiLSTM model for accurate polarity prediction","volume":"66","author":"Shobana","year":"2023","journal-title":"Comput. J."},{"key":"2024062414160014600_ref13","doi-asserted-by":"crossref","first-page":"1966","DOI":"10.1109\/TCSVT.2022.3218018","article-title":"BAFN: bi-direction attention based fusion network for multimodal sentiment analysis","volume":"33","author":"Tang","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"2024062414160014600_ref14","first-page":"5105","article-title":"Multi-level attention map network for multimodal sentiment analysis","volume":"35","author":"Xue","year":"2023","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"2024062414160014600_ref15","first-page":"12888","article-title":"BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation","volume-title":"Proc. of the 39th Int. Conf. on Machine Learning Research (PMLR)","author":"Li","year":"2022"},{"key":"2024062414160014600_ref16","first-page":"2325","article-title":"Retrieving users\u2019 opinions on social media with multimodal aspect-based sentiment analysis","volume-title":"Proc. of the 2023 IEEE Int. Conf. on Semantic Computing (ICSC)","author":"Ansch\u00fctz","year":"2023"},{"key":"2024062414160014600_ref17","doi-asserted-by":"crossref","first-page":"2953","DOI":"10.1093\/comjnl\/bxac091","article-title":"A semantic embedding enhanced topic model for user-generated textual content modeling in social ecosystems","volume":"65","author":"Zhang","year":"2022","journal-title":"Comput. J."},{"key":"2024062414160014600_ref18","first-page":"6000","article-title":"Attention is all you need","volume-title":"Proc. of the 31st Int. Conf. on Neural Information Processing Systems (NIPS)","author":"Vaswani","year":"2017"},{"key":"2024062414160014600_ref19","first-page":"17","article-title":"Dynamic multimodal fusion","volume-title":"Proc. of the 2022 IEEE Int. Conf. on Computer Vision and Pattern Recognition Workshops (CVPRW)","author":"Xue","year":"2022"},{"key":"2024062414160014600_ref20","first-page":"33","article-title":"Multi-interactive memory network for aspect based multimodal sentiment analysis","volume-title":"Proc. of the 31st AAAI Conf. on Artificial Intelligence (AAAI)","author":"Xu","year":"2019"},{"key":"2024062414160014600_ref21","doi-asserted-by":"crossref","first-page":"1473","DOI":"10.1109\/SMC52423.2021.9658940","article-title":"Multimodal sentiment analysis based on attention mechanism and tensor fusion network","volume-title":"Proc. of the 2021 IEEE Int. Conf. on Systems, Man and Cybernetics (SMC)","author":"Zhang","year":"2021"},{"key":"2024062414160014600_ref22","article-title":"ALBERT: a lite BERT for self-supervised learning of language representations","volume-title":"Proc. of the 2019 Int. Conf. on Learning Representations (ICLR)","author":"Lan","year":"2019"},{"key":"2024062414160014600_ref23","doi-asserted-by":"crossref","first-page":"344","DOI":"10.1109\/SLT48900.2021.9383575","article-title":"Audio Albert: a lite BERT for self-supervised learning of audio representation","volume-title":"Proc. of the 2021 IEEE Int. Conf. on Spoken Language Technology Workshop (SLT)","author":"Chi","year":"2021"},{"key":"2024062414160014600_ref24","doi-asserted-by":"crossref","first-page":"0957","DOI":"10.1016\/j.eswa.2021.115458","article-title":"Deep learning for predicting neutralities in offensive language identification dataset","volume":"185","author":"Sharma","year":"2021","journal-title":"Expert Syst. Appl."},{"key":"2024062414160014600_ref25","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.knosys.2018.06.019","article-title":"Imbalanced text sentiment classification using universal and domain-specific knowledge","volume":"160","author":"Li","year":"2018","journal-title":"Knowl.-Based Syst."},{"key":"2024062414160014600_ref26","doi-asserted-by":"crossref","first-page":"0168","DOI":"10.1016\/j.compag.2023.107622","article-title":"Improved ResNet-50 deep learning algorithm for identifying chicken gender","volume":"205","author":"Wu","year":"2023","journal-title":"Comput. Electron. Agric."},{"key":"2024062414160014600_ref27","first-page":"13708","article-title":"Coordinate attention for efficient mobile network design","volume-title":"Proc. of the 2021 IEEE Int. Conf. on Computer Vision and Pattern Recognition (CVPR)","author":"Hou","year":"2021"},{"key":"2024062414160014600_ref28","doi-asserted-by":"crossref","first-page":"239","DOI":"10.1162\/coli_a_00433","article-title":"Ethics sheet for automatic emotion recognition and sentiment analysis","volume":"48","author":"Mohammad","year":"2021","journal-title":"Comput. Linguist."},{"key":"2024062414160014600_ref29","first-page":"15","article-title":"Sentiment analysis on multi-view social data","volume-title":"Proc. of the 22nd Int. Conf. on MultiMedia Modeling (MMM)","author":"Niu","year":"2016"},{"key":"2024062414160014600_ref30","doi-asserted-by":"crossref","first-page":"4014","DOI":"10.1109\/TMM.2020.3035277","article-title":"Image-text multimodal emotion classification via multi-view attentiona network","volume":"23","author":"Yang","year":"2021","journal-title":"IEEE Trans. Multimed."},{"key":"2024062414160014600_ref31","doi-asserted-by":"crossref","first-page":"3375","DOI":"10.1109\/TMM.2022.3160060","article-title":"Multimodal sentiment analysis with image-text interaction network","volume":"25","author":"Zhu","year":"2023","journal-title":"IEEE Trans. Multimed."}],"container-title":["The Computer Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/67\/6\/2230\/58309206\/bxae002.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/67\/6\/2230\/58309206\/bxae002.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,24]],"date-time":"2024-06-24T14:55:25Z","timestamp":1719240925000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/comjnl\/article\/67\/6\/2230\/7595364"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,31]]},"references-count":31,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2024,1,31]]},"published-print":{"date-parts":[[2024,6,24]]}},"URL":"https:\/\/doi.org\/10.1093\/comjnl\/bxae002","relation":{},"ISSN":["0010-4620","1460-2067"],"issn-type":[{"value":"0010-4620","type":"print"},{"value":"1460-2067","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024,6]]},"published":{"date-parts":[[2024,1,31]]}}}