{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:55:15Z","timestamp":1782312915686,"version":"3.54.5"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T00:00:00Z","timestamp":1782259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"Open Fund of Tianjin Key Laboratory for Advanced Signal Processing","award":["2025ASP-TJ01"],"award-info":[{"award-number":["2025ASP-TJ01"]}]},{"name":"Beijing Natural Science Foundation","award":["L252143"],"award-info":[{"award-number":["L252143"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>Multimodal large language models (MLLMs) have achieved significant advancements in multimodal understanding, reasoning, and interaction. However, they still suffer from hallucination, where the generated text often deviates from the factual content of the input image. To mitigate this issue, prior studies have primarily employed direct preference optimization (DPO) for human preference alignment. However, these approaches treat all textual words equally, neglecting the varying significance of individual words in grounding text generation to image content. This limitation hinders fine-grained semantic alignment and consequently constrains their effectiveness in hallucination suppression. To address this limitation, we propose a vision-guided lexical DPO method, called VGL-DPO. Specifically, we quantify the significance of words in positive preference data based on their relevance to the visual input and dynamically assign different weights to different words during training. This facilitates more precise optimization by emphasizing critical words that contribute to factual grounding. Additionally, we leverage the importance differences between high-significance words in positive and negative preference data to adaptively adjust the weight of the negative preference loss. This dynamic reweighting mechanism further refines the model\u2019s ability to suppress hallucinated content while reinforcing factual accuracy. Extensive experiments across various models demonstrate that our method outperforms existing state-of-the-art methods in reducing hallucination and enhancing factual accuracy.<\/jats:p>","DOI":"10.1145\/3796715","type":"journal-article","created":{"date-parts":[[2026,5,16]],"date-time":"2026-05-16T14:19:31Z","timestamp":1778941171000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["VGL-DPO: Vision-Guided Lexical Direct Preference Optimization for Mitigating Hallucination in Multimodal Large Language Models"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7002-5081","authenticated-orcid":false,"given":"Siyuan","family":"Li","sequence":"first","affiliation":[{"name":"School of Data Science and Intelligent Media, Communication University of China, Beijing,\u00a0China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0516-1618","authenticated-orcid":false,"given":"Feng","family":"Wang","sequence":"additional","affiliation":[{"name":"Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3286-6824","authenticated-orcid":false,"given":"Simeng","family":"Qin","sequence":"additional","affiliation":[{"name":"Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-2261-4268","authenticated-orcid":false,"given":"Ranjie","family":"Duan","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3407-4318","authenticated-orcid":false,"given":"Haonan","family":"Cheng","sequence":"additional","affiliation":[{"name":"Communication University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3562-5612","authenticated-orcid":false,"given":"Long","family":"Ye","sequence":"additional","affiliation":[{"name":"Communication University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_2_3_2","unstructured":"OpenAI Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman et al. 2024. GPT-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_2_4_2","first-page":"23716","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS \u201922)","author":"Alayrac Jean-Baptiste","year":"2022","unstructured":"Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: A visual language model for few-shot learning. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS \u201922), 23716\u201323736."},{"key":"e_1_3_2_5_2","unstructured":"Jinze Bai Shuai Bai Shusheng Yang Shijie Wang Sinan Tan Peng Wang Junyang Lin Chang Zhou and Jingren Zhou. 2023. Qwen-Vl: A versatile vision-language model for understanding localization text reading and beyond. arXiv:2308.12966. Retrieved from https:\/\/arxiv.org\/abs\/2308.12966"},{"key":"e_1_3_2_6_2","unstructured":"Zechen Bai Pichao Wang Tianjun Xiao Tong He Zongbo Han Zheng Zhang and Mike Zheng Shou. 2024. Hallucination of multimodal large language models: A survey. arXiv:2404.18930. Retrieved from https:\/\/arxiv.org\/abs\/2404.18930"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3618406"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02283"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","unstructured":"Zheng Chen Xun Zhang Wenbo Li Renjing Pei Fenglong Song Xiongkuo Min Xiaohong Liu Xin Yuan Yong Guo and Yulun Zhang. 2024. Grounding-IQA: Multimodal language grounding model for image quality assessment. arXiv:2411.17237v2. Retrieved from https:\/\/arxiv.org\/abs\/2411.17237v2","DOI":"10.32388\/074NKV"},{"key":"e_1_3_2_10_2","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Cheng Tianheng","year":"2024","unstructured":"Tianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu, Xinggang Wang, and Ying Shan. 2024. YOLO-World: Real-time open-vocabulary object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_11_2","unstructured":"Glenn Jocher. 2023. Yolov8. Retrieved from https:\/\/github.com\/ultralytics\/ultralytics"},{"key":"e_1_3_2_12_2","doi-asserted-by":"crossref","first-page":"13418","DOI":"10.1109\/CVPR52733.2024.01274","volume-title":"Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Huang Qidong","year":"2024","unstructured":"Qidong Huang, Xiaoyi Dong, Pan Zhang, Bin Wang, Conghui He, Jiaqi Wang, Dahua Lin, Weiming Zhang, and Nenghai Yu. 2024. Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation. In Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13418\u201313427."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02553"},{"key":"e_1_3_2_14_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Lee Ji Soo","year":"2025","unstructured":"Ji Soo Lee, Jongha Kim, Jeehye Na, Jinyoung Park, and Hyunwoo J. Kim. 2025. VidChain: Chain-of-tasks with metric-based direct preference optimization for dense video captioning. In Proceedings of the AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_2_15_2","unstructured":"Lei Li Zhihui Xie Mukai Li Shunian Chen Peiyi Wang Liang Chen Yazheng Yang Benyou Wang and Lingpeng Kong. 2023. Silkie: Preference distillation for large visual language models. arXiv:2312.10665. Retrieved from https:\/\/arxiv.org\/abs\/2312.10665"},{"key":"e_1_3_2_16_2","unstructured":"Tsung-Yi Lin Michael Maire Serge Belongie Lubomir Bourdev Ross Girshick James Hays Pietro Perona Deva Ramanan C. Lawrence Zitnick and Piotr Doll\u00e1r. 2015. Microsoft COCO: Common objects in context. arXiv:1405.0312. Retrieved from https:\/\/arxiv.org\/abs\/1405.0312"},{"key":"e_1_3_2_17_2","unstructured":"Fuxiao Liu Kevin Lin Linjie Li Jianfeng Wang Yaser Yacoob and Lijuan Wang. 2024. Mitigating hallucination in large multi-modal models via robust instruction tuning. arXiv:2306.14565. Retrieved from https:\/\/arxiv.org\/abs\/2306.14565"},{"key":"e_1_3_2_18_2","first-page":"34892","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923)","author":"Liu Haotian","year":"2023","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruction Tuning. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923), 34892\u201334916."},{"key":"e_1_3_2_19_2","article-title":"Visual instruction tuning","volume":"36","author":"Liu Haotian","year":"2023","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. Advances in Neural Information Processing Systems 36.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_20_2","unstructured":"Shilong Liu Zhaoyang Zeng Tianhe Ren Feng Li Hao Zhang Jie Yang Chunyuan Li Jianwei Yang Hang Su Jun Zhu et al. 2023. Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection. arXiv:2303.05499. Retrieved from https:\/\/arxiv.org\/abs\/2303.05499"},{"key":"e_1_3_2_21_2","unstructured":"Ziyu Liu Yuhang Zang Xiaoyi Dong Pan Zhang Yuhang Cao Haodong Duan Conghui He Yuanjun Xiong Dahua Lin and Jiaqi Wang. 2024. Mia-DPO: Multi-image augmented direct preference optimization for large vision-language models. arXiv:2410.17637. Retrieved from https:\/\/arxiv.org\/abs\/2410.17637"},{"key":"e_1_3_2_22_2","unstructured":"Jinda Lu Junkang Wu Jinghan Li Xiaojun Jia Shuo Wang YiFan Zhang Junfeng Fang Xiang Wang and Xiangnan He. 2025. DAMO: Data- and model-aware alignment of multi-modal LLMs. arXiv:2502.01943. Retrieved from https:\/\/arxiv.org\/abs\/2502.01943"},{"key":"e_1_3_2_23_2","first-page":"124198","volume-title":"Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS \u201924)","author":"Meng Yu","year":"2024","unstructured":"Yu Meng, Mengzhou Xia, and Danqi Chen. 2024. SimPO: Simple preference optimization with a reference-free reward. In Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS \u201924), 124198\u2013124235."},{"key":"e_1_3_2_24_2","first-page":"395","volume-title":"Proceedings of the 18th European Conference on Computer Vision (ECCV \u201924)","author":"Ouali Yassine","year":"2024","unstructured":"Yassine Ouali, Adrian Bulat, Brais Martinez, and Georgios Tzimiropoulos. 2024. CLIP-DPO: Vision-language models as a source of preference for fixing hallucinations in LVLMS. In Proceedings of the 18th European Conference on Computer Vision (ECCV \u201924). Springer, 395\u2013413."},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","unstructured":"Rafael Rafailov Archit Sharma Eric Mitchell Stefano Ermon Christopher D. Manning and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. arXiv:2305.18290. Retrieved from https:\/\/arxiv.org\/abs\/2305.18290","DOI":"10.52202\/075280-2338"},{"key":"e_1_3_2_26_2","unstructured":"Joseph Redmon and Ali Farhadi. 2018. YOLOv3: An Incremental Improvement. arXiv:1804.02767. Retrieved from https:\/\/arxiv.org\/abs\/1804.02767"},{"key":"e_1_3_2_27_2","first-page":"91","volume-title":"Proceedings of the 29th International Conference on Neural Information Processing Systems (NIPS \u201915)","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Proceedings of the 29th International Conference on Neural Information Processing Systems (NIPS \u201915), 91\u201399. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2015\/file\/14bfa6bb14875e45bba028a21ed38046-Paper.pdf"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1437"},{"key":"e_1_3_2_29_2","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv:1707.06347. Retrieved from https:\/\/arxiv.org\/abs\/1707.06347"},{"key":"e_1_3_2_30_2","unstructured":"Guanzheng Chen Xin Li Shijian Lu Chunyan Miao Lidong Bing Sicong Leng and Hang Zhang. 2023. Mitigating object hallucinations in large vision-language models through visual contrastive decoding. arXiv:2311.16922. Retrieved from https:\/\/arxiv.org\/abs\/2311.16922"},{"key":"e_1_3_2_31_2","unstructured":"Jianlin Su Yu Lu Shengfeng Pan Bo Wen and Yunfeng Liu. 2021. RoFormer: Enhanced transformer with rotary position embedding. arXiv:2104.09864. Retrieved from https:\/\/arxiv.org\/abs\/2104.09864"},{"key":"e_1_3_2_32_2","unstructured":"Yinan Sun Zicheng Zhang Haoning Wu Xiaohong Liu Weisi Lin Guangtao Zhai and Xiongkuo Min. 2024. Explore the hallucination on low-level perception for MLLMs. arXiv:2409.09748. Retrieved from https:\/\/arxiv.org\/abs\/2409.09748"},{"key":"e_1_3_2_33_2","unstructured":"Zhiqing Sun Sheng Shen Shengcao Cao Haotian Liu Chunyuan Li Yikang Shen Chuang Gan Liang-Yan Gui Yu-Xiong Wang Yiming Yang et al. 2023. Aligning large multimodal models with factually augmented RLHF. arXiv:2309.14525. Retrieved from https:\/\/arxiv.org\/abs\/2309.14525"},{"key":"e_1_3_2_34_2","unstructured":"Yuan Yao Haoye Zhang Yue Zhao Chongyi Wang Shan Wang Yinxv Pan Jiao Xue Dahai Li Zhiyuan Liu Hai-Tao Zheng et al. 2023. Reformulating vision-language foundation models and datasets towards universal multimodal assistants. arXiv:2310.00653. Retrieved from https:\/\/arxiv.org\/abs\/2310.00653"},{"key":"e_1_3_2_35_2","unstructured":"Glenn Jocher Ayush Chaurasia Alex Stoken Jirka Borovec Yonghye Kwon Kalen Michael Jiacong Fang Zeng Yifu Colin Wong Diego Montes et al. 2022. Ultralytics\/yolov5: V 7.0-YOLOv5 SOTA Realtime Instance Segmentation. Retrieved May 7 2023 from https:\/\/github.com\/ultralytics\/yolov5.com"},{"key":"e_1_3_2_36_2","unstructured":"Junyang Wang Yuhang Wang Guohai Xu Jing Zhang Yukai Gu Haitao Jia Ming Yan Ji Zhang and Jitao Sang. 2023. An LLM-free Multi-dimensional benchmark for MLLMs hallucination evaluation. arXiv:2311.07397. Retrieved from https:\/\/arxiv.org\/abs\/2311.07397"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01838"},{"key":"e_1_3_2_38_2","first-page":"129944","volume-title":"Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS \u201924)","author":"Wu Junkang","year":"2024","unstructured":"Junkang Wu, Yuexiang Xie, Zhengyi Yang, Jiancan Wu, Jinyang Gao, Bolin Ding, Xiang Wang, and Xiangnan He. 2024. \\(\\beta\\) -DPO: Direct Preference Optimization with Dynamic \\(\\beta\\) . In Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS \u201924), 129944\u2013129966."},{"key":"e_1_3_2_39_2","doi-asserted-by":"crossref","first-page":"8460","DOI":"10.1609\/aaai.v39i8.32913","article-title":"Combating multimodal LLM hallucination via bottom-up holistic reasoning","volume":"39","author":"Wu Shengqiong","year":"2025","unstructured":"Shengqiong Wu, Hao Fei, Liangming Pan, William Yang Wang, Shuicheng Yan, and Tat-Seng Chua. 2025. Combating multimodal LLM hallucination via bottom-up holistic reasoning. Proceedings of the AAAI Conference on Artificial Intelligence 39 (2025), 8460\u20138468.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_3_2_40_2","first-page":"30925","volume-title":"Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS \u201924)","author":"Xiao Xin","year":"2024","unstructured":"Xin Xiao, Bohong Wu, Jiacong Wang, Chunyuan Li, Xun Zhou, and Haoyuan Guo. 2024. Seeing the image: Prioritizing visual correlation by contrastive alignment. In Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS \u201924), 30925\u201330950."},{"key":"e_1_3_2_41_2","doi-asserted-by":"crossref","unstructured":"Yuxi Xie Guanzhen Li Xiao Xu and Min-Yen Kan. 2024. V-DPO: Mitigating hallucination in large vision language models via vision-guided direct preference optimization. arXiv:2411.02712. Retrieved from https:\/\/arxiv.org\/abs\/2411.02712","DOI":"10.18653\/v1\/2024.findings-emnlp.775"},{"key":"e_1_3_2_42_2","doi-asserted-by":"crossref","unstructured":"Yun Xing Yiheng Li Ivan Laptev and Shijian Lu. 2024. Mitigating object hallucination via concentric causal attention. arXiv:2410.15926. Retrieved from https:\/\/arxiv.org\/abs\/2410.15926","DOI":"10.52202\/079017-2920"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-024-4251-x"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01230"},{"key":"e_1_3_2_45_2","first-page":"13807","volume-title":"Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Yu Tianyu","year":"2024","unstructured":"Tianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He, Yifeng Han, Ganqu Cui, Jinyi Hu, Zhiyuan Liu, Hai-Tao Zheng, and Maosong Sun. 2024. RLHF-V: Towards trustworthy MLLMs via behavior alignment from fine-grained correctional human feedback. In Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13807\u201313816."},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","unstructured":"Tianyu Yu Haoye Zhang Yuan Yao Yunkai Dang Da Chen Xiaoman Lu Ganqu Cui Taiwen He Zhiyuan Liu Tat-Seng Chua et al. 2024. RLAIF-V: Aligning MLLMs through open-source AI feedback for super GPT-4V trustworthiness. arXiv:2405.17220. Retrieved from https:\/\/arxiv.org\/abs\/2405.17220","DOI":"10.1109\/CVPR52734.2025.01861"},{"key":"e_1_3_2_47_2","first-page":"11766","volume-title":"Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Yue Zihao","year":"2024","unstructured":"Zihao Yue, Liang Zhang, and Qin Jin. 2024. Less is more: Mitigating multimodal hallucination from an EOS decision perspective. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11766\u201311781."},{"key":"e_1_3_2_48_2","first-page":"26171","volume-title":"Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS \u201924)","author":"Zhang Mengxi","year":"2024","unstructured":"Mengxi Zhang, Wenhao Wu, Lu Yu, Yuxin Song, Kang Rong, Huanjin Yao, Jianbo Zhang, Fanglong Liu, Haocheng Feng, Yifan Sun, et al. 2024. Automated multi-level preference for MLLMs. In Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS \u201924), 26171\u201326194."},{"key":"e_1_3_2_49_2","unstructured":"Zhiyuan Zhao Bin Wang Linke Ouyang Xiaoyi Dong Jiaqi Wang and Conghui He. 2023. 2023. Beyond hallucinations: Enhancing LVLMs through hallucination-aware direct preference optimization. arXiv:2311.16839. Retrieved from https:\/\/arxiv.org\/abs\/2311.16839"},{"key":"e_1_3_2_50_2","unstructured":"Kening Zheng Junkai Chen Yibo Yan Xin Zou and Xuming Hu. 2024. Reefknot: A comprehensive benchmark for relation hallucination evaluation analysis and mitigation in multimodal large language models. arXiv:2408.09429. Retrieved from https:\/\/arxiv.org\/abs\/2408.09429"},{"key":"e_1_3_2_51_2","unstructured":"Yiyang Zhou Chenhang Cui Rafael Rafailov Chelsea Finn and Huaxiu Yao. 2024. Aligning modalities in vision large language models via preference fine-tuning. arXiv:2402.11411. Retrieved from https:\/\/arxiv.org\/abs\/2402.11411"},{"key":"e_1_3_2_52_2","unstructured":"Deyao Zhu Jun Chen Xiaoqian Shen Xiang Li and Mohamed Elhoseiny. 2023. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. arXiv:2304.10592. Retrieved from https:\/\/arxiv.org\/abs\/2304.10592"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3796715","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:45:45Z","timestamp":1782312345000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3796715"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":51,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3796715"],"URL":"https:\/\/doi.org\/10.1145\/3796715","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-06-17","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-04","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}