{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T23:55:13Z","timestamp":1783986913900,"version":"3.55.0"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,10,23]],"date-time":"2025-10-23T00:00:00Z","timestamp":1761177600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,10,23]],"date-time":"2025-10-23T00:00:00Z","timestamp":1761177600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Manipal Academy of Higher Education, Manipal"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Discov Computing"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Social media has become an essential platform for expressing personal experiences and emotions. Today\u2019s youth frequently share images that reflect their emotional states, including happiness, excitement, sadness, anxiety, and distress. Accurately analyzing these images using new frameworks can offer beneficial insights into the emotional well-being of individuals. Beyond mental health applications, image sentiment analysis has significant potential in marketing and advertising. Brands and Marketers can get a more comprehensive understanding of consumer sentiments and preferences by examining the emotional reactions elicited by visual content. For instance, companies can analyze images shared by customers to gauge sentiment towards their products and services. Positive or negative feedback expressed through images can offer practical insights for improving products and customer experience. Additionally, Sentiment analysis is one tool that marketers can use to gauge the effectiveness of their advertising campaigns. By analyzing the sentiments of images associated with a campaign, they can determine which aspects resonate most with the audience and adjust their strategies accordingly. Our research focuses on creating an advanced multimodal sentiment analysis system that combines BERT and Vision Transformers (ViT) to analyze textual and image data. High-precision sentiment classification is achieved by our technique using a preprocessed AllenTAN dataset from Hugging Face. It conducts sentiment analysis using BERT, creates captions for unlabeled photos, and uses OCR to retrieve embedded image text. The suggested ViT\u2009+\u2009BERT technique performs well with a variety of social network content. The proposed system achieves an accuracy of 96.91%, demonstrating its robust performance across diverse social media content and benchmark models. This technology has several uses, particularly in social media monitoring to promote mental health content, as teens frequently use visuals to describe their feelings.<\/jats:p>","DOI":"10.1007\/s10791-025-09756-2","type":"journal-article","created":{"date-parts":[[2025,10,23]],"date-time":"2025-10-23T10:32:40Z","timestamp":1761215560000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Multimodal sentiment analysis using image and text fusion for emotion detection"],"prefix":"10.1007","volume":"28","author":[{"given":"Uttam U.","family":"Deshpande","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Supriya","family":"Shanbhag","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Amit","family":"Sukhasare","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mahendra M.","family":"Dixit","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rudragoud","family":"Patil","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sangeeta","family":"Sangani","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sowmyashree H.","family":"Srinivasaiah","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Swetha","family":"Goudar","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Manjunath","family":"Managuli","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,10,23]]},"reference":[{"key":"9756_CR1","doi-asserted-by":"crossref","unstructured":"Hu R, Rui L, Zeng P, Chen L, Fan X. Text sentiment analysis: a review. In Proceedings of the 2018 IEEE 4th International Conference on Computer and Communications (ICCC), Chengdu, China, 7\u201310 December 2018; IEEE: Piscataway, NJ, USA, 2018.","DOI":"10.1109\/CompComm.2018.8780909"},{"key":"9756_CR2","unstructured":"Cai Z, Cao D, Ji R. Video (GIF) sentiment analysis using large-scale mid-level ontology. ArXiv. 2015. ArXiv:1506.00765."},{"key":"9756_CR3","unstructured":"Xu C, Cetintas S, Lee KC, Li LJ. Visual sentiment prediction with deep convolutional neural networks. ArXiv. 2014. ArXiv:1411.5731."},{"key":"9756_CR4","doi-asserted-by":"crossref","unstructured":"Tang D, Qin B, Liu T. Learning semantic representations of users and products for document level sentiment classification. In Proceedings of the 53rd annual meeting of the association for computational linguistics and the 7th international joint conference on natural language processing, Beijing, China, 26\u201331 July 2015; Volume 1: Long Papers, pp. 1014\u20131023.","DOI":"10.3115\/v1\/P15-1098"},{"key":"9756_CR5","doi-asserted-by":"crossref","unstructured":"Ibrahim M, Abdillah O, Wicaksono AF, Adriani M. Buzzer detection and sentiment analysis for predicting presidential election results in a twitter nation. In Proceedings of the 2015 IEEE International Conference on Data Mining Workshop (ICDMW), Atlantic City, NJ, USA, 14\u201317 November 2015; IEEE: Piscataway, NJ, USA; 2015. pp. 1348\u20131353.","DOI":"10.1109\/ICDMW.2015.113"},{"issue":"2","key":"9756_CR6","doi-asserted-by":"publisher","first-page":"617","DOI":"10.1007\/s10115-018-1236-4","volume":"60","author":"L Yue","year":"2019","unstructured":"Yue L, Chen W, Li X, Zuo W, Yin M. A survey of sentiment analysis in social media. Knowl Inf Syst. 2019;60(2):617\u201363.","journal-title":"Knowl Inf Syst"},{"key":"9756_CR7","doi-asserted-by":"publisher","first-page":"1785","DOI":"10.1109\/TMM.2020.3003648","volume":"23","author":"W Guo","year":"2021","unstructured":"Guo W, Zhang Y, Cai X, Meng L, Yang J, Yuan X. LD-MAN: Layout-driven multimodal attention network for online news sentiment recognition. IEEE Trans Multimedia. 2021;23:1785\u201398. https:\/\/doi.org\/10.1109\/TMM.2020.3003648.","journal-title":"IEEE Trans Multimedia"},{"key":"9756_CR8","doi-asserted-by":"publisher","first-page":"424","DOI":"10.1016\/j.inffus.2022","volume":"91","author":"A Gandhi","year":"2022","unstructured":"Gandhi A, Adhvaryu KU, Poria S, Cambria E, Hussain A. Multimodal sentiment analysis: A systematic review of history, datasets, multimodal fusion methods, applications, challenges and future directions. Inform Fusion. 2022;91:424\u201344. https:\/\/doi.org\/10.1016\/j.inffus.2022. 09.025.","journal-title":"Inform Fusion"},{"issue":"5","key":"9756_CR9","doi-asserted-by":"publisher","first-page":"1358","DOI":"10.1109\/TMM.2019.2939744","volume":"22","author":"D She","year":"2020","unstructured":"She D, Yang J, Cheng M, Lai Y, Rosin PL, Wang L. WSCNet: weakly supervised coupled networks for visual sentiment classification and detection. IEEE Trans Multimedia. 2020;22(5):1358\u201371. https:\/\/doi.org\/10.1109\/TMM.2019.2939744.","journal-title":"IEEE Trans Multimedia"},{"key":"9756_CR10","doi-asserted-by":"crossref","unstructured":"Wang M, Cao D, Li L, Li S, Ji R. Microblog sentiment analysis based on cross-media bag-of-words model. In Proceedings of the international conference on internet multimedia computing and service, Xiamen, China, 10\u201312 July 2014; pp. 76\u201380.","DOI":"10.1145\/2632856.2632912"},{"key":"9756_CR11","doi-asserted-by":"crossref","unstructured":"Zhang Y, Shang L, Jia X. Sentiment analysis on microblogging by integrating text and image features. In Proceedings of the Pacific-Asia conference on knowledge discovery and data mining, Ho Chi Minh City, Vietnam, 19\u201322 May 2015; pp. 52\u201363. 26.","DOI":"10.1007\/978-3-319-18032-8_5"},{"key":"9756_CR12","doi-asserted-by":"publisher","first-page":"41","DOI":"10.3390\/a9020041","volume":"9","author":"Y Yu","year":"2016","unstructured":"Yu Y, Lin H, Meng J, Zhao Z. Visual and textual sentiment analysis of a microblog using deep convolutional neural networks. Algorithms. 2016;9:41.","journal-title":"Algorithms"},{"key":"9756_CR13","doi-asserted-by":"crossref","unstructured":"Tsai YH, Bai S, Liang PP, Kolter JZ, Morency L, Salakhutdinov R. 2019. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the Conference. Association for Computational Linguistics Meeting, Florence, Italy, 6558\u201369. Vol. 2019. NIH Public Access.","DOI":"10.18653\/v1\/P19-1656"},{"key":"9756_CR14","doi-asserted-by":"publisher","first-page":"102141","DOI":"10.1016\/j.ipm.2019.102141","volume":"57","author":"A Kumar","year":"2020","unstructured":"Kumar A, Srinivasan K, Cheng WH, Zomaya AY. Hybrid context enriched deep learning model for fine-grained sentiment analysis in textual and visual semiotic modality social data. Inf Process Manag. 2020;57:102141.","journal-title":"Inf Process Manag"},{"issue":"1","key":"9756_CR15","doi-asserted-by":"publisher","first-page":"2000688","DOI":"10.1080\/08839514.2021.2000688","volume":"36","author":"X Yan","year":"2021","unstructured":"Yan X, Xue H, Jiang S, Liu Z. Multimodal sentiment analysis using multi-tensor fusion network with cross-modal modeling. Appl Artif Intell. 2021;36(1):2000688. https:\/\/doi.org\/10.1080\/08839514.2021.2000688.","journal-title":"Appl Artif Intell"},{"key":"9756_CR16","doi-asserted-by":"publisher","first-page":"100243","DOI":"10.1016\/j.dajour.2023.100243","volume":"7","author":"G Meena","year":"2023","unstructured":"Meena G, Mohbey K, Kumar K, Lokesh K. A hybrid deep learning approach for detecting sentiment polarities and knowledge graph representation on Monkeypox tweets. Decis Analytics J. 2023;7:100243. https:\/\/doi.org\/10.1016\/j.dajour.2023.100243.","journal-title":"Decis Analytics J"},{"issue":"1","key":"9756_CR17","doi-asserted-by":"publisher","first-page":"33","DOI":"10.1007\/s41095-021-0247-3","volume":"8","author":"Y Xu","year":"2021","unstructured":"Xu Y, Wei H, Lin M, Deng Y, Sheng K, Zhang M, Tang F, Dong W, Huang F, Xu C. Transformers in computational visual media: A survey. Comput Visual Media. 2021;8(1):33\u201362. https:\/\/doi.org\/10.1007\/s41095-021-0247-3.","journal-title":"Comput Visual Media"},{"issue":"2","key":"9756_CR18","doi-asserted-by":"publisher","first-page":"2071","DOI":"10.1609\/aaai.v36i2.20103","volume":"36","author":"S Paul","year":"2022","unstructured":"Paul S, Chen PY. Vision Transformers are robust learners. Proc AAAI Conf Artif Intell. 2022;36(2):2071\u201381. https:\/\/doi.org\/10.1609\/aaai.v36i2.20103.","journal-title":"Proc AAAI Conf Artif Intell"},{"key":"9756_CR19","doi-asserted-by":"publisher","unstructured":"Xie Y, Liao Y. 2023. Efficient-ViT: a light-weight classification model based on CNN and ViT. In Proceedings of the 2023 6th International Conference on Image and Graphics Processing, 64\u201370. https:\/\/doi.org\/10.1145\/3582649.3582676","DOI":"10.1145\/3582649.3582676"},{"key":"9756_CR20","unstructured":"Bao H, Dong L, Wei F. Beit: Bert pre-training of image transformers. 2021. ArXiv abs\/ 210608254."},{"key":"9756_CR21","doi-asserted-by":"publisher","unstructured":"Chen CF, Fan Q, Panda R. CrossViT: cross-attention multi-scale vision transformer for image classification. Proc IEEE Int Conf Comput Vis. 2021;347\u2013356. https:\/\/doi.org\/10.48550\/arxiv.2103.14899.","DOI":"10.48550\/arxiv.2103.14899"},{"key":"9756_CR22","doi-asserted-by":"publisher","first-page":"521","DOI":"10.1016\/J.NEUNET.2023.04.045","volume":"164","author":"J Chen","year":"2023","unstructured":"Chen J, Zhang Y, Pan Y, et al. A transformer-based deep neural network model for SSVEP classification. Neural Netw. 2023;164:521\u201334. https:\/\/doi.org\/10.1016\/J.NEUNET.2023.04.045.","journal-title":"Neural Netw"},{"key":"9756_CR23","doi-asserted-by":"publisher","first-page":"108861","DOI":"10.1016\/j.knosys.2022.108861","volume":"248","author":"Q Gao","year":"2022","unstructured":"Gao Q, Cao B, Guan X, Gu T, Bao X, Wu J, Liu B, Cao J. Emotion recognition in conversations with emotion shift detection based on multi-task learning. Knowl Based Syst. 2022;248:108861. https:\/\/doi.org\/10.1016\/j.knosys.2022.108861.","journal-title":"Knowl Based Syst"},{"key":"9756_CR24","doi-asserted-by":"publisher","unstructured":"Tan CH, Chan A, Haldar M, Tang J, Liu X, Abdool M, Gao H, He L, Katariya S. 2023. Optimizing Airbnb search journey with multi-task learning. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 4872\u201381. https:\/\/doi.org\/10.1145\/3580305.3599881","DOI":"10.1145\/3580305.3599881"},{"key":"9756_CR25","doi-asserted-by":"publisher","first-page":"2015","DOI":"10.1109\/TASLP.2022.3178204","volume":"30","author":"B Yang","year":"2022","unstructured":"Yang B, Wu L, Zhu J, Shao B, Lin X, Liu T. Multimodal sentiment analysis with two-phase multi-task learning. IEEE\/ACM Trans Audio Speech Lang Process. 2022;30:2015\u201324. https:\/\/doi.org\/10.1109\/TASLP.2022.3178204.","journal-title":"IEEE\/ACM Trans Audio Speech Lang Process"},{"key":"9756_CR26","doi-asserted-by":"publisher","unstructured":"Yu W, Xu H, Meng F, Zhu Y, Ma Y, Wu J, Zou J, Yang K. CH-SIMS: a Chinese multimodal sentiment analysis dataset with fine-grained annotation of modality. Proc 58th Annual Meeting Association Comput Linguistics. 2020;3718\u201327. https:\/\/doi.org\/10.18653\/v1\/2020.acl-main.343.","DOI":"10.18653\/v1\/2020.acl-main.343"},{"issue":"1","key":"9756_CR27","doi-asserted-by":"publisher","first-page":"200","DOI":"10.1109\/TETCI.2022.3224929","volume":"7","author":"S Zhang","year":"2023","unstructured":"Zhang S, Yin C, Yin Z. Multimodal sentiment recognition with multi-task learning. IEEE Trans Emerg Top Comput Intell. 2023;7(1):200\u201309. https:\/\/doi.org\/10.1109\/TETCI.2022.3224929.","journal-title":"IEEE Trans Emerg Top Comput Intell"},{"key":"9756_CR28","unstructured":"https:\/\/huggingface.co\/datasets\/AllenTAN\/image_sentiment."},{"key":"9756_CR29","unstructured":"Devlin J et al. BERT: pre-training of deep bidirectional transformers for language understanding. North American chapter of the association for computational linguistics. 2019. https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"9756_CR30","doi-asserted-by":"publisher","unstructured":"Mikolov T, Chen K, Corrado Gs, Dean J. Efficient estimation of word representations in vector space. Proceedings of Workshop at ICLR. 2013. https:\/\/doi.org\/10.48550\/arXiv.1301.3781.","DOI":"10.48550\/arXiv.1301.3781"},{"key":"9756_CR31","doi-asserted-by":"crossref","unstructured":"Pennington J, Socher R, Manning C. GloVe: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp 1532\u20131543, Doha, Qatar. Association for Computational Linguistics. 2014.","DOI":"10.3115\/v1\/D14-1162"},{"key":"9756_CR32","unstructured":"Sun C, Huang L, Qiu X. Utilizing BERT for Aspect-Based sentiment analysis via constructing auxiliary sentence. NAACL-HLT. 2019. https:\/\/arxiv.org\/abs\/1903.09588."},{"key":"9756_CR33","unstructured":"https:\/\/huggingface.co\/docs\/transformers\/model_doc\/vit."},{"key":"9756_CR34","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, Uszkoreit J, Houlsby N. An image is worth 16x16 words: transformers for image recognition at scale. 2020. ArXivabs\/2010.11929."},{"key":"9756_CR35","doi-asserted-by":"publisher","first-page":"13911","DOI":"10.1007\/s11227-021-03838-w","volume":"77","author":"I Priyadarshini","year":"2021","unstructured":"Priyadarshini I, Cotton C. A novel LSTM\u2013CNN\u2013grid search-based deep neural network for sentiment analysis. J Supercomput. 2021;77:13911\u201332. https:\/\/doi.org\/10.1007\/s11227-021-03838-w.","journal-title":"J Supercomput"},{"key":"9756_CR36","doi-asserted-by":"publisher","unstructured":"Mohbey K, Meena G, Kumar S, Lokesh K. A CNN-LSTM-based hybrid deep learning approach to detect sentiment polarities on Monkeypox tweets. 2022. https:\/\/doi.org\/10.48550\/arXiv.2208.12019.","DOI":"10.48550\/arXiv.2208.12019"}],"container-title":["Discover Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10791-025-09756-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10791-025-09756-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10791-025-09756-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,23]],"date-time":"2025-10-23T10:32:46Z","timestamp":1761215566000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10791-025-09756-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,23]]},"references-count":36,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["9756"],"URL":"https:\/\/doi.org\/10.1007\/s10791-025-09756-2","relation":{},"ISSN":["2948-2992"],"issn-type":[{"value":"2948-2992","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,23]]},"assertion":[{"value":"15 May 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 October 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 October 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"This research did not involve human participants or animals; hence ethical approval was not required.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable, as no personal data, images, or details of individual participants are included in this manuscript.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"230"}}