{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,25]],"date-time":"2026-07-25T16:46:56Z","timestamp":1784998016194,"version":"3.55.0"},"reference-count":68,"publisher":"Association for Computing Machinery (ACM)","issue":"6","funder":[{"name":"Liaoning Provincial Social Science Planning Fund Project","award":["L24CGL021, L24CGL022"],"award-info":[{"award-number":["L24CGL021, L24CGL022"]}]},{"name":"the Scientific Research Project of Liaoning Education Department","award":["LJ212410173066, LJKQZ20222444"],"award-info":[{"award-number":["LJ212410173066, LJKQZ20222444"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2025,12,31]]},"abstract":"<jats:p>Multimodal sentiment analysis has become a popular research topic in recent years. However, existing methods have two unaddressed limitations: (1) they use limited supervised labels to train models, which makes it impossible for model to fully learn sentiments in different modal data; (2) they employ text and image pre-trained models trained in different unimodal tasks to extract different modal features, so that the extracted features cannot take into account the interactive information between image and text. To solve these problems, in this paper we propose a Vision-Language Contrastive Learning network (VLCLNet). First, we introduce a pre-trained Large Language Model (LLM), which is trained from vast quantities of multimodal data, has better understanding ability for image and text contents, thus being effectively applied to different tasks while requiring few amount of labelled training data. Second, we adapt a Multimodal Large Language Model (MLLM), BLIP-2 (Bootstrapping Language-Image Pre-training) network, to extract multimodal fusion feature. Such MLLM can fully consider the correlation between images and texts when extracting features. In addition, due to the discrepancy between the pre-training task and the sentiment analysis task, the pre-trained model will output the suboptimal prediction results. We use Low-Rank Adaptation (LoRA) fine-tuning strategy to update the model parameters on sentiment analysis task, which avoids the issue of inconsistent task between pre-training task and downstream task. Experiments verify that the proposed VLCLNet is superior to other strong baselines.<\/jats:p>","DOI":"10.1145\/3709147","type":"journal-article","created":{"date-parts":[[2024,12,20]],"date-time":"2024-12-20T10:56:41Z","timestamp":1734692201000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Multimodal Large Language Model with LoRA Fine-Tuning for Multimodal Sentiment Analysis"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-1958-5110","authenticated-orcid":false,"given":"Jie","family":"Mu","sequence":"first","affiliation":[{"name":"School of Data Science and Artificial Intelligence, Dongbei University of Finance and Economics, Dalian, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8676-1190","authenticated-orcid":false,"given":"Wei","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Cyber Science and Technology, Sun Yat-sen University - Shenzhen Campus, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0834-0355","authenticated-orcid":false,"given":"Wenqi","family":"Liu","sequence":"additional","affiliation":[{"name":"WENQI LIU, School of Data Science and Artificial Intelligence, Dongbei University of Finance and Economics, Dalian, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0811-9706","authenticated-orcid":false,"given":"Tiantian","family":"Yan","sequence":"additional","affiliation":[{"name":"The National and Local Joint Engineering Laboratory of Computer Aided Design, Dalian University, Dalian, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7207-2789","authenticated-orcid":false,"given":"Guanglu","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Software, Dalian University of Technology, Dalian, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,11,24]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Samyadeep Basu Daniela Massiceti Shell Xu Hu and Soheil Feizi. 2024. Strong baselines for parameter efficient few-shot fine-tuning. arXiv:2304.01917. Retrieved from https:\/\/arxiv.org\/abs\/2304.01917"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-25207-0_14"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2022.3192728"},{"key":"e_1_3_2_5_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations (ICLR \u201923)","author":"Chen Zhe","year":"2023","unstructured":"Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao. 2023. Vision transformer adapter for dense predictions. In Proceedings of the 11th International Conference on Learning Representations (ICLR \u201923). OpenReview.Net."},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2023.3265653"},{"key":"e_1_3_2_7_2","unstructured":"Junyoung Chung \u00c7aglar G\u00fcl\u00e7ehre KyungHyun Cho and Yoshua Bengio. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv:1412.3555. Retrieved from https:\/\/arxiv.org\/abs\/1412.3555"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01055"},{"key":"e_1_3_2_9_2","volume-title":"Proceedings of the 10th International Conference on Learning Representations (ICLR \u201922)","author":"He Junxian","year":"2022","unstructured":"Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. 2022. Towards a unified view of parameter-efficient transfer learning. In Proceedings of the 10th International Conference on Learning Representations (ICLR \u201922). OpenReview.Net."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-emnlp.160"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i1.25160"},{"key":"e_1_3_2_13_2","first-page":"2790","volume-title":"Proceedings of the 36th International Conference on Machine Learning (ICML \u201919)","author":"Houlsby Neil","year":"2019","unstructured":"Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning (ICML \u201919). Proceedings of Machine Learning Research, Vol. 97, PMLR, 2790\u20132799."},{"key":"e_1_3_2_14_2","volume-title":"Proceedings of the 10th International Conference on Learning Representations (ICLR \u201922)","author":"Hu Edward J.","year":"2022","unstructured":"Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In Proceedings of the 10th International Conference on Learning Representations (ICLR \u201922). OpenReview.net."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.126992"},{"key":"e_1_3_2_16_2","unstructured":"Zeyinzi Jiang Chaojie Mao Ziyuan Huang Yiliang Lv Deli Zhao and Jingren Zhou. 2023. Rethinking efficient tuning methods from a unified perspective. arXiv:2303.00690. Retrieved from https:\/\/arxiv.org\/abs\/2303.00690"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475692"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1181"},{"key":"e_1_3_2_19_2","first-page":"19730","volume-title":"Proceedings of the International Conference on Machine Learning (ICML 2023)","volume":"202","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the International Conference on Machine Learning (ICML 2023). (Proceedings of Machine Learning Research, Vol. 202, PMLR, 19730\u201319742."},{"key":"e_1_3_2_20_2","first-page":"12888","volume-title":"Proceedings of the International Conference on Machine Learning (ICML 2022)","volume":"162","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. 2022. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine Learning (ICML 2022). Proceedings of Machine Learning Research, Vol. 162, PMLR, 12888\u201312900."},{"key":"e_1_3_2_21_2","first-page":"9694","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS \u201921)","author":"Li Junnan","year":"2021","unstructured":"Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven Chu-Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. In Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS \u201921), 9694\u20139705."},{"key":"e_1_3_2_22_2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: a robustly optimized bert pretraining approach. arXiv:1907.11692. Retrieved from https:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_2_23_2","unstructured":"Qingyu Lu Baopu Qiu Liang Ding Liping Xie and Dacheng Tao. 2023. Error analysis prompting enables human-like translation evaluation in large language models: A case study on ChatGPT. arXiv:2303.13809. Retrieved from https:\/\/arxiv.org\/abs\/2303.13809"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2018.07.041"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.02.015"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3439726"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2023.3345022"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-27674-8_2"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.208"},{"key":"e_1_3_2_30_2","unstructured":"OpenAI. 2023. GPT-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.373"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.617"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1303"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1081"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612012"},{"key":"e_1_3_2_36_2","first-page":"8748","volume-title":"Proceedings of the 38th International Conference on Machine Learning (ICML \u201921)","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML \u201921). Proceedings of Machine Learning Research, Vol. 139, PMLR, 8748\u20138763."},{"key":"e_1_3_2_37_2","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21 (2020), 140:1\u2013140:67.","journal-title":"J. Mach. Learn. Res"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46478-7_21"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_2_40_2","volume-title":"Proceedings of the Conference on Learning Representations","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In Proceedings of the Conference on Learning Representations."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00746"},{"key":"e_1_3_2_42_2","unstructured":"Haixin Wang Jianlong Chang Xiao Luo Jinan Sun Zhouchen Lin and Qi Tian. 2024. LION: Implicit vision prompt tuning. arXiv:2303.09992. Retrieved from https:\/\/arxiv.org\/abs\/2303.09992"},{"key":"e_1_3_2_43_2","doi-asserted-by":"crossref","unstructured":"Wenbin Wang Liang Ding Li Shen Yong Luo Han Hu and Dacheng Tao. 2024. WisdoM: Improving multimodal sentiment analysis by fusing contextual world knowledge. arXiv:2401.06659. Retrieved from https:\/\/arxiv.org\/abs\/2401.06659","DOI":"10.1145\/3664647.3681403"},{"key":"e_1_3_2_44_2","volume-title":"Proceedings of the 10th International Conference on Learning Representations (ICLR \u201922)","author":"Wang Zirui","year":"2022","unstructured":"Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. 2022. SimVLM: Simple visual language model pretraining with weak supervision. In Proceedings of the 10th International Conference on Learning Representations (ICLR \u201922). OpenReview.net."},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2020.3015381"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00390"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-59830-3_3"},{"key":"e_1_3_2_48_2","doi-asserted-by":"crossref","unstructured":"Yi Xin Junlong Du Qiang Wang Zhiwen Lin and Ke Yan. 2024. VMT-Adapter: Parameter-efficient transfer learning for multi-task dense scene understanding. arXiv:2312.08733. Retrieved from https:\/\/arxiv.org\/abs\/2312.08733","DOI":"10.1609\/aaai.v38i14.29541"},{"key":"e_1_3_2_49_2","unstructured":"Yi Xin Siqi Luo Haodi Zhou Junlong Du Xiaohong Liu Yue Fan Qing Li and Yuntao Du. 2024. Parameter-efficient fine-tuning for pre-trained vision models: A survey. arXiv:2402.02242. Retrieved from https:\/\/arxiv.org\/abs\/2402.02242"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3311618"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3132847.3133142"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1234"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.421"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413690"},{"key":"e_1_3_2_55_2","first-page":"328","article-title":"Multimodal sentiment detection based on multi-Channel graph neural networks","author":"Yang Xiaocui","year":"2021","unstructured":"Xiaocui Yang, Shi Feng, Yifei Zhang, and Daling Wang. 2021. Multimodal sentiment detection based on multi-Channel graph neural networks. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, 328\u2013339.","journal-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics"},{"key":"e_1_3_2_56_2","unstructured":"Bruce X. B. Yu Jianlong Chang Lingbo Liu Qi Tian and Chang Wen Chen. 2022. Towards a unified view on visual parameter-efficient transfer learning. arXiv:2210.00788. Retrieved from https:\/\/arxiv.org\/abs\/2210.00788"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2019\/751"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1115"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.49"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8682248"},{"key":"e_1_3_2_61_2","unstructured":"Qingru Zhang Minshuo Chen Alexander Bukharin Nikos Karampatziakis Pengcheng He Yu Cheng Weizhu Chen and Tuo Zhao. 2023. AdaLoRA: Adaptive budg et\u00a0al location for parameter-efficient fine-tuning. arXiv:2303.10512. Retrieved from https:\/\/arxiv.org\/abs\/2303.10512"},{"key":"e_1_3_2_62_2","unstructured":"Susan Zhang Stephen Roller Naman Goyal Mikel Artetxe Moya Chen Shuohui Chen Christopher Dewan Mona T. Diab Xian Li Xi Victoria Lin et\u00a0al. 2022. OPT: Open pre-trained transformer language models. arXiv:2205.01068. Retrieved from https:\/\/arxiv.org\/abs\/2205.01068"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.108234"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-srw.40"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2020.10.001"},{"key":"e_1_3_2_66_2","unstructured":"Henry Hengyuan Zhao Pichao Wang Yuyang Zhao Hao Luo Fan Wang and Mike Zheng Shou. 2023. SCT: A simple baseline for parameter-efficient Fine-tuning via salient channels. arXiv:2309.08513. Retrieved from https:\/\/arxiv.org\/abs\/2309.08513"},{"key":"e_1_3_2_67_2","unstructured":"Qihuang Zhong Liang Ding Juhua Liu Bo Du and Dacheng Tao. 2023. Can chatgpt understand too? A comparative study on chatgpt and fine-tuned bert. arXiv:2302.10198. Retrieved from https:\/\/arxiv.org\/abs\/2302.10198"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00918"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611974"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3709147","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,24]],"date-time":"2025-11-24T15:08:34Z","timestamp":1763996914000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3709147"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,24]]},"references-count":68,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,12,31]]}},"alternative-id":["10.1145\/3709147"],"URL":"https:\/\/doi.org\/10.1145\/3709147","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,24]]},"assertion":[{"value":"2024-02-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-11-16","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}