{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,2]],"date-time":"2026-08-02T20:00:35Z","timestamp":1785700835416,"version":"3.56.0"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2024,1,21]],"date-time":"2024-01-21T00:00:00Z","timestamp":1705795200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,1,21]],"date-time":"2024-01-21T00:00:00Z","timestamp":1705795200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"National Science Fundation of China","award":["62366010"],"award-info":[{"award-number":["62366010"]}]},{"name":"Project of Guangxi Key Lab of Trusted Software","award":["Kx202060"],"award-info":[{"award-number":["Kx202060"]}]},{"name":"CCF-Zhipu AI Large Model Fund"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Data Sci Anal"],"published-print":{"date-parts":[[2025,8]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Aspect-level multimodal sentiment analysis is the fine-grained sentiment analysis task of predicting the sentiment polarity of given aspects in multimodal data. Most existing multimodal sentiment analysis approaches focus on mining and fusing multimodal global features, overlooking the correlation of more fine-grained multimodal local features, which considerably limits the semantic relevance between different modalities. Therefore, a novel aspect-level multimodal sentiment analysis method based on global\u2013local features fusion with co-attention (GLFFCA) is proposed to comprehensively explore multimodal associations from both global and local perspectives. Specially, an aspect-guided global co-attention module is designed to capture aspect-guided intra-modality global correlations. Meanwhile, a gated local co-attention module is introduced to capture the adaptive association alignment of multimodal local features. Following that, a global\u2013local multimodal feature fusion module is constructed to integrate global\u2013local multimodal features in a hierarchical manner. Extensive experiments on the Twitter-2015 dataset and Twitter-2017 dataset validate the effectiveness of the proposed method, which can achieve better aspect-level multimodal sentiment analysis performance compared with other related methods.<\/jats:p>","DOI":"10.1007\/s41060-023-00497-3","type":"journal-article","created":{"date-parts":[[2024,1,20]],"date-time":"2024-01-20T21:01:46Z","timestamp":1705784506000},"page":"903-916","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":22,"title":["Aspect-level multimodal sentiment analysis based on co-attention fusion"],"prefix":"10.1007","volume":"20","author":[{"given":"Shunjie","family":"Wang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guoyong","family":"Cai","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guangrui","family":"Lv","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,1,21]]},"reference":[{"key":"497_CR1","doi-asserted-by":"crossref","unstructured":"Truong, Q.T.: Lauw H W,:Vistanet: Visual aspect attention network for multimodal sentiment analysis. In: Proceedings of the AAAI Conference on Artificial Intelligence 33(01), 305\u2013312 (2019)","DOI":"10.1609\/aaai.v33i01.3301305"},{"key":"497_CR2","doi-asserted-by":"crossref","unstructured":"Xu, N., Mao, W.: Chen G,: Multi-interactive memory network for aspect based multimodal sentiment analysis. In: Proceedings of the AAAI Conference on Artificial Intelligence. 33(01), 371\u2013378 (2019)","DOI":"10.1609\/aaai.v33i01.3301371"},{"key":"497_CR3","doi-asserted-by":"publisher","first-page":"429","DOI":"10.1109\/TASLP.2019.2957872","volume":"28","author":"J Yu","year":"2019","unstructured":"Yu, J., Jiang, J., Xia, R.: Entity-sensitive attention and fusion network for entity-level multimodal sentiment classification. IEEE\/ACM Trans Audio, Speech, Language Proc 28, 429\u2013439 (2019)","journal-title":"IEEE\/ACM Trans Audio, Speech, Language Proc"},{"key":"497_CR4","doi-asserted-by":"publisher","first-page":"157329","DOI":"10.1109\/ACCESS.2021.3126782","volume":"9","author":"D Gu","year":"2021","unstructured":"Gu, D., Wang, J., Cai, S.: Targeted aspect-based multimodal sentiment analysis: an attention capsule extraction and multi-head fusion network. IEEE Access 9, 157329\u2013157336 (2021)","journal-title":"IEEE Access"},{"key":"497_CR5","doi-asserted-by":"crossref","unstructured":"Xu, N., Mao, W., Chen, G.: A co-memory network for multimodal sentiment analysis. In: The 41st international ACM SIGIR conference on research & development in information retrieval, pp. 929\u2013932 (2018)","DOI":"10.1145\/3209978.3210093"},{"key":"497_CR6","doi-asserted-by":"publisher","first-page":"172948","DOI":"10.1109\/ACCESS.2019.2955637","volume":"7","author":"S Nemati","year":"2019","unstructured":"Nemati, S., Rohani, R., Basiri, M. E.: A hybrid latent space data fusion method for multimodal emotion recognition. IEEE Access 7, 172948\u2013172964 (2019)","journal-title":"IEEE Access"},{"issue":"2","key":"497_CR7","doi-asserted-by":"publisher","first-page":"41","DOI":"10.3390\/a9020041","volume":"9","author":"Y Yu","year":"2016","unstructured":"Yu, Y., Lin, H., Meng, J.: Visual and textual sentiment analysis of a microblog using deep convolutional neural networks. Algorithms 9(2), 41 (2016)","journal-title":"Algorithms"},{"issue":"1","key":"497_CR8","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2019.102141","volume":"57","author":"A Kumar","year":"2020","unstructured":"Kumar, A., Srinivasan, K., Cheng, W. H.: Hybrid context enriched deep learning model for fine-grained sentiment analysis in textual and visual semiotic modality social data. Inf Proc Manag 57(1), 102141 (2020)","journal-title":"Inf Proc Manag"},{"key":"497_CR9","doi-asserted-by":"crossref","unstructured":"Chen, F., Gao, Y., Cao, D.: Multimodal hypergraph learning for microblog sentiment prediction. In: 2015 IEEE International Conference on Multimedia and Expo (ICME). IEEE, pp. 1-6 (2015)","DOI":"10.1109\/ICME.2015.7177477"},{"key":"497_CR10","doi-asserted-by":"publisher","first-page":"387","DOI":"10.1016\/j.asoc.2019.04.010","volume":"80","author":"J Xu","year":"2019","unstructured":"Xu, J., Huang, F.: Sentiment analysis of social images via hierarchical deep fusion of content and links. Appl Soft Comput 80, 387\u2013399 (2019)","journal-title":"Appl Soft Comput"},{"key":"497_CR11","doi-asserted-by":"crossref","unstructured":"Dong, L., Wei, F., Tan, C.: Adaptive recursive neural network for target-dependent twitter sentiment classification. In: Proceedings of the 52nd annual meeting of the association for computational linguistics (volume 2: Short papers), pp. 49-54 (2014)","DOI":"10.3115\/v1\/P14-2009"},{"key":"497_CR12","unstructured":"Tang, D., Qin, B., Feng, X.: Effective LSTMs for target-dependent sentiment classification, in arXiv preprint arXiv:1512.01100 (2015)"},{"key":"497_CR13","doi-asserted-by":"crossref","unstructured":"Ma, D., Li, S., Zhang, X.: Interactive attention networks for aspect-level sentiment classification, in arXiv preprint arXiv:1709.00893 (2017)","DOI":"10.24963\/ijcai.2017\/568"},{"key":"497_CR14","doi-asserted-by":"crossref","unstructured":"Tang, D., Qin, B., Liu, T.: Aspect level sentiment classification with deep memory network, in arXiv preprint arXiv:1605.08900 (2016)","DOI":"10.18653\/v1\/D16-1021"},{"key":"497_CR15","doi-asserted-by":"crossref","unstructured":"Chen, P., Sun, Z., Bing, L.: Recurrent attention network on memory for aspect sentiment analysis. In: Proceedings of the 2017 conference on empirical methods in natural language processing, pp. 452-461 (2017)","DOI":"10.18653\/v1\/D17-1047"},{"key":"497_CR16","first-page":"1","volume":"2023","author":"Li Jia","year":"2023","unstructured":"Jia, Li., Ma, Tinghua, Rong, Huan, Al-Nabhan, Najla: Affective region recognition and fusion network for target-level multimodal sentiment classification[J]. IEEE Trans Emerg Topic Comput 2023, 1\u201311 (2023)","journal-title":"IEEE Trans Emerg Topic Comput"},{"key":"497_CR17","doi-asserted-by":"crossref","unstructured":"Jiang, T., Wang, J., Liu, Z.: Fusion-extraction network for multimodal sentiment analysis. In: Pacific-Asia conference on knowledge discovery and data mining. Springer, Cham, pp. 785-797 (2020)","DOI":"10.1007\/978-3-030-47436-2_59"},{"key":"497_CR18","doi-asserted-by":"crossref","unstructured":"Zadeh, A., Chen, M., Poria, S.: Tensor fusion network for multimodal sentiment analysis, in arXiv preprint arXiv:1707.07250 (2017)","DOI":"10.18653\/v1\/D17-1115"},{"key":"497_CR19","doi-asserted-by":"crossref","unstructured":"Verma, S., Wang, J., Ge, Z.: Deep-HOSeq: Deep higher order sequence fusion for multimodal sentiment analysis. In: 2020 IEEE International Conference on Data Mining (ICDM). IEEE, pp. 561-570 (2020)","DOI":"10.1109\/ICDM50108.2020.00065"},{"key":"497_CR20","doi-asserted-by":"crossref","unstructured":"Xi, C., Lu, G., Yan, J.: Multimodal sentiment analysis based on multi-head attention mechanism. In: Proceedings of the 4th international conference on machine learning and soft computing, pp. 34-39 (2020)","DOI":"10.1145\/3380688.3380693"},{"key":"497_CR21","doi-asserted-by":"crossref","unstructured":"Kim, T., Lee, B.: Multi-attention multimodal sentiment analysis. In: Proceedings of the 2020 International Conference on Multimedia Retrieval, pp. 436-441 (2020)","DOI":"10.1145\/3372278.3390698"},{"key":"497_CR22","doi-asserted-by":"crossref","unstructured":"Yu, J., Jiang, J.: Adapting BERT for target-oriented multimodal sentiment classification. In: IJCAI (2019)","DOI":"10.24963\/ijcai.2019\/751"},{"key":"497_CR23","doi-asserted-by":"crossref","unstructured":"Khan, Z., Fu, Y.: Exploiting BERT for multimodal target sentiment classification through input space translation[C]. In: Proceedings of the 29th ACM international conference on multimedia. 2021, 3034-3042 (2021)","DOI":"10.1145\/3474085.3475692"},{"key":"497_CR24","doi-asserted-by":"crossref","unstructured":"Xiao, L., Zhou, E., Wu, X., Yang, S., Ma, T., He, L.: Adaptive Multi-Feature Extraction Graph Convolutional Networks for Multimodal Target Sentiment Analysis. In: 2022 IEEE International Conference on Multimedia and Expo (ICME) (pp.1-6). IEEE (2022)","DOI":"10.1109\/ICME52920.2022.9860020"},{"issue":"5","key":"497_CR25","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2022.103038","volume":"59","author":"Li Yang","year":"2022","unstructured":"Yang, Li., Na, Jin-Cheon., Jianfei, Yu.: Cross-modal multitask transformer for end-to-end multimodal aspect-based sentiment analysis[J]. Inf Proc Manag 59(5), 103038 (2022)","journal-title":"Inf Proc Manag"},{"key":"497_CR26","doi-asserted-by":"crossref","unstructured":"Yu, Jianfei, Wang, Jieming, Xia, Rui, Li, Junjie: Targeted Multimodal Sentiment Classification based on Coarse-to-Fine Grained Image-Target Matching. IJCAI (2022)","DOI":"10.24963\/ijcai.2022\/622"},{"key":"497_CR27","doi-asserted-by":"crossref","unstructured":"Luong, M.T., Pham, H., Manning, C.D.: Effective approaches to attention-based neural machine translation. in arXiv preprint arXiv:1508.04025 (2015)","DOI":"10.18653\/v1\/D15-1166"},{"key":"497_CR28","doi-asserted-by":"crossref","unstructured":"Gkanatsios, N., Pitsikalis, V., Koutras, P.: Deeply supervised multimodal attentional translation embeddings for visual relationship detection. In: 2019 IEEE International Conference on Image Processing (ICIP),IEEE, pp. 1840-1844 (2019)","DOI":"10.1109\/ICIP.2019.8803106"},{"key":"497_CR29","unstructured":"Lu, J., Yang, J., Batra, D.: Hierarchical question-image co-attention for visual question answering. In: Advances in neural information processing systems, 29 (2016)"},{"key":"497_CR30","doi-asserted-by":"crossref","unstructured":"Jiang, M., Chen, S., Yang, J.: Fantastic answers and where to find them: Immersive question-directed visual attention. In: Proceedings of the ieee\/cvf conference on computer vision and pattern recognition, pp. 2980-2989 (2020)","DOI":"10.1109\/CVPR42600.2020.00305"},{"key":"497_CR31","unstructured":"Xu, K., Ba, J., Kiros, R.: Show, attend and tell: Neural image caption generation with visual attention. In: International conference on machine learning. PMLR, pp. 2048-2057 (2015)"},{"key":"497_CR32","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S.: Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778 (2016)","DOI":"10.1109\/CVPR.2016.90"},{"key":"497_CR33","unstructured":"Kim, J.H., On, K.W., Lim, W.: Hadamard product for low-rank bilinear pooling, in arXiv preprint arXiv:1610.04325 (2016)"},{"key":"497_CR34","unstructured":"Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization, in arXiv preprint arXiv:1412.6980 (2014)"},{"key":"497_CR35","doi-asserted-by":"crossref","unstructured":"Fan, F., Feng, Y., Zhao, D.: Multi-grained attention network for aspect-level sentiment classification. In: Proceedings of the 2018 conference on empirical methods in natural language processing, pp. 3433-3442 (2018)","DOI":"10.18653\/v1\/D18-1380"},{"key":"497_CR36","doi-asserted-by":"crossref","unstructured":"Hazarika, D., Poria, S., Zadeh, A.: Conversational memory network for emotion recognition in dyadic dialogue videos. In: Proceedings of the conference, Association for Computational Linguistics, North American Chapter,Meeting,NIH Public Access, 2122 (2018)","DOI":"10.18653\/v1\/N18-1193"}],"container-title":["International Journal of Data Science and Analytics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s41060-023-00497-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s41060-023-00497-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s41060-023-00497-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,5]],"date-time":"2025-09-05T18:28:33Z","timestamp":1757096913000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s41060-023-00497-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,21]]},"references-count":36,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,8]]}},"alternative-id":["497"],"URL":"https:\/\/doi.org\/10.1007\/s41060-023-00497-3","relation":{},"ISSN":["2364-415X","2364-4168"],"issn-type":[{"value":"2364-415X","type":"print"},{"value":"2364-4168","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,21]]},"assertion":[{"value":"27 August 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 December 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 January 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"On behalf of all authors, the corresponding author states that there is no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}