{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T05:49:29Z","timestamp":1782452969669,"version":"3.54.5"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T00:00:00Z","timestamp":1778803200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"New Generation Artificial Intelligence-National Science and Technology Major Project","award":["2025ZD0123602"],"award-info":[{"award-number":["2025ZD0123602"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62276110 and 62172039"],"award-info":[{"award-number":["62276110 and 62172039"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Joint Laboratory of HUST and Pingan Property and Casualty Research"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>\n                    While Large Vision-Language Models (LVLMs) have exhibited remarkable capabilities across a wide range of tasks, they suffer from hallucination problems, where models generate plausible yet incorrect answers given the input image-query pair. This hallucination phenomenon is even more severe when querying the image in non-English languages, while existing methods for mitigating hallucinations in LVLMs only consider the English scenarios. In this article, we make the first attempt to mitigate this important multilingual hallucination in LVLMs. With thorough experimental analysis, we found that multilingual hallucination in LVLMs is a systemic problem that could arise from deficiencies in multilingual capabilities or inadequate multimodal abilities. To this end, we propose a two-stage Multilingual Hallucination Removal (MHR) framework for LVLMs, aiming to improve resistance to hallucination for both high-resource and low-resource languages. Specifically, in the first stage, considering that most non-English languages cannot follow instructions well and output non-sense answers given the input image, we boost multilingual instruction-following ability with a multilingual supervised fine-tuning. The second phase is aimed at enhancing the LVLM\u2019s ability to diminish multilingual hallucinations. Instead of relying on the intricate manual annotations of multilingual resources, we fully leverage the inherent capabilities of the LVLM and propose a novel cross-lingual alignment method, which generates multiple responses for each image-query input and then identifies the hallucination-aware pairs for each language. These data pairs are finally used for direct preference optimization to prompt the LVLMs to favor non-hallucinating responses. Experimental results show that our MHR achieves a substantial reduction in hallucination generation for LVLMs. Our code and model weights are available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/ssmisya\/MHR\">https:\/\/github.com\/ssmisya\/MHR<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3797025","type":"journal-article","created":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T14:49:43Z","timestamp":1774968583000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Mitigating Multilingual Hallucination in Large Vision-Language Models"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4907-3978","authenticated-orcid":false,"given":"Xiaoye","family":"Qu","sequence":"first","affiliation":[{"name":"Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-5181-1326","authenticated-orcid":false,"given":"Mingyang","family":"Song","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4488-0102","authenticated-orcid":false,"given":"Wei","family":"Wei","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8179-4508","authenticated-orcid":false,"given":"Daizong","family":"Liu","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China and Wuhan University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5244-3274","authenticated-orcid":false,"given":"Jianfeng","family":"Dong","sequence":"additional","affiliation":[{"name":"Zhejiang Gongshang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7901-8662","authenticated-orcid":false,"given":"Yu","family":"Cheng","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,5,15]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Jinze Bai Shuai Bai Shusheng Yang Shijie Wang Sinan Tan Peng Wang Junyang Lin Chang Zhou and Jingren Zhou. 2023. Qwen-VL: A frontier large vision-language model with versatile abilities. arXiv:2308.12966. Retrieved from https:\/\/arxiv.org\/abs\/2308.12966"},{"key":"e_1_3_2_3_2","unstructured":"Assaf Ben-Kish Moran Yanuka Morris Alper Raja Giryes and and Hadar Averbuch-Elor. 2023. MOCHa: Multi-objective reinforcement mitigating caption hallucinations. arXiv:2312.03631. Retrieved from https:\/\/arxiv.org\/abs\/2312.03631"},{"key":"e_1_3_2_4_2","unstructured":"Wei-Lin Chiang Zhuohan Li Zi Lin Ying Sheng Zhanghao Wu Hao Zhang Lianmin Zheng Siyuan Zhuang Yonghao Zhuang Joseph E. Gonzalez et al. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. Retrieved from https:\/\/lmsys.org\/blog\/2023-03-30-vicuna\/"},{"key":"e_1_3_2_5_2","unstructured":"Marta R. Costa-Juss\u00e0 James Cross Onur \u00c7elebi Maha Elbayad Kenneth Heafield Kevin Heffernan Elahe Kalbassi Janice Lam Daniel Licht Jean Maillard et al. 2022. No language left behind: Scaling human-centered machine translation. arXiv:2207.04672. Retrieved from https:\/\/arxiv.org\/abs\/2207.04672"},{"key":"e_1_3_2_6_2","article-title":"InstructBLIP: Towards general-purpose vision-language models with instruction tuning","volume":"36","author":"Dai Wenliang","year":"2024","unstructured":"Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N. Fung, and Steven Hoi. 2024. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_7_2","unstructured":"Chaoyou Fu Peixian Chen Yunhang Shen Yulei Qin Mengdan Zhang Xu Lin Jinrui Yang Xiawu Zheng Ke Li Xing Sun et al. 2023. MME: A comprehensive evaluation benchmark for multimodal large language models. arXiv:2306.13394. Retrieved from https:\/\/arxiv.org\/abs\/2306.13394"},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","unstructured":"Gregor Geigle Abhay Jain Radu Timofte and Goran Glava\u0161. 2023. mBLIP: Efficient bootstrapping of multilingual vision-LLMs. arXiv:2307.06930. Retrieved from https:\/\/arxiv.org\/abs\/2307.06930","DOI":"10.18653\/v1\/2024.alvr-1.2"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","unstructured":"Qidong Huang Xiaoyi Dong Pan Zhang Bin Wang Conghui He Jiaqi Wang Dahua Lin Weiming Zhang and Nenghai Yu. 2023. OPERA: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation. arXiv:2311.17911. Retrieved from https:\/\/arxiv.org\/abs\/2311.17911","DOI":"10.1109\/CVPR52733.2024.01274"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00686"},{"key":"e_1_3_2_11_2","unstructured":"Chris Kelly Luhui Hu Bang Yang Yu Tian Deshun Yang Cindy Yang Zaoshan Huang Zihao Li Jiayin Hu and Yuexian Zou. 2024. VisionGPT: Vision-language understanding agent using generalized multimodal framework. arXiv:2403.09027. Retrieved from https:\/\/arxiv.org\/abs\/2403.09027"},{"key":"e_1_3_2_12_2","first-page":"13171","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Dac Lai Viet","year":"2023","unstructured":"Viet Dac Lai, Nghia Ngo, Amir Pouran Ben Veyseh, Hiu Mn, Franck Dernoncourt, Trung Bui, and Thien Huu Nguyen. 2023. ChatGPT beyond English: Towards a comprehensive evaluation of large language models in multilingual learning. In Findings of the Association for Computational Linguistics: EMNLP 2023, 13171\u201313189."},{"key":"e_1_3_2_13_2","unstructured":"Sicong Leng Hang Zhang Guanzheng Chen Xin Li Shijian Lu Chunyan Miao and Lidong Bing. 2023. Mitigating object hallucinations in large vision-language models through visual contrastive decoding. arXiv:2311.16922. Retrieved from https:\/\/arxiv.org\/abs\/2311.16922"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01316"},{"key":"e_1_3_2_15_2","first-page":"19730","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the International Conference on Machine Learning. PMLR, 19730\u201319742."},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3182151"},{"key":"e_1_3_2_17_2","unstructured":"Bin Lin Zhenyu Tang Yang Ye Jinfa Huang Junwu Zhang Yatian Pang Peng Jin Munan Ning Jiebo Luo and Li Yuan. 2024. MoE-LLaVA: Mixture of experts for large vision-language models. arXiv:2401.15947. Retrieved from https:\/\/arxiv.org\/abs\/2401.15947"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_19_2","volume-title":"Proceedings of the 12th International Conference on Learning Representations","author":"Liu Fuxiao","year":"2023","unstructured":"Fuxiao Liu, Kevin Lin, Linjie Li, Jianfeng Wang, Yaser Yacoob, and Lijuan Wang. 2023. Mitigating hallucination in large multi-modal models via robust instruction tuning. In Proceedings of the 12th International Conference on Learning Representations."},{"key":"e_1_3_2_20_2","unstructured":"Haotian Liu Chunyuan Li Yuheng Li and Yong Jae Lee. 2024. Improved baselines with visual instruction tuning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 26296\u201326306."},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","unstructured":"Haotian Liu Chunyuan Li Qingyang Wu and Yong Jae Lee. 2023. Visual instruction tuning. Advances in Neural Information Processing Systems 36 (2023) 34892\u201334916.","DOI":"10.52202\/075280-1516"},{"key":"e_1_3_2_22_2","unstructured":"Junling Liu Ziming Wang Qichen Ye Dading Chong Peilin Zhou and Yining Hua. 2023. Qilin-Med-VL: Towards Chinese large vision-language model for general healthcare. arXiv:2310.17956. Retrieved from https:\/\/arxiv.org\/abs\/2310.17956"},{"key":"e_1_3_2_23_2","unstructured":"Kangcheng Liu Xinhu Zheng Chaoqun Wang Hesheng Wang Ming Liu and Kai Tang. 2024. Online robot navigation and manipulation with distilled vision-language models. arXiv:2401.17083. Retrieved from https:\/\/arxiv.org\/abs\/2401.17083"},{"key":"e_1_3_2_24_2","unstructured":"Yuliang Liu Biao Yang Qiang Liu Zhang Li Zhiyin Ma Shuo Zhang and and Xiang Bai. 2024. TextMonkey: An OCR-free large multimodal model for understanding document. arXiv:2403.04473. Retrieved from https:\/\/arxiv.org\/abs\/2403.04473"},{"key":"e_1_3_2_25_2","unstructured":"Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_26_2","unstructured":"Muhammad Maaz Hanoona Rasheed Abdelrahman Shaker Salman Khan Hisham Cholakal Rao M. Anwer Tim Baldwin Michael Felsberg and Fahad S. Khan. 2024. PALO: A polyglot large multimodal model for 5B people. arXiv:2402.14818. Retrieved from https:\/\/arxiv.org\/abs\/2402.14818"},{"key":"e_1_3_2_27_2","doi-asserted-by":"crossref","unstructured":"Brandon McKinzie Zhe Gan Jean-Philippe Fauconnier Sam Dodge Bowen Zhang Philipp Dufter Dhruti Shah Xianzhi Du Futang Peng Floris Weers et al. 2024. MM1: Methods analysis & insights from multimodal LLM pre-training. arXiv:2403.09611. Retrieved from https:\/\/arxiv.org\/abs\/2403.09611","DOI":"10.1007\/978-3-031-73397-0_18"},{"key":"e_1_3_2_28_2","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311\u2013318."},{"key":"e_1_3_2_29_2","first-page":"9573","article-title":"SgVA-CLIP: Semantic-guided visual adapting of vision-language models for few-shot image classification","volume":"25","author":"Peng Fang","year":"2023","unstructured":"Fang Peng, Xiaoshan Yang, Linhui Xiao, Yaowei Wang, and Changsheng Xu. 2023. SgVA-CLIP: Semantic-guided visual adapting of vision-language models for few-shot image classification. IEEE Transactions on Multimedia 25 (2023), 9573\u20139585.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3742434"},{"key":"e_1_3_2_31_2","first-page":"4428","volume-title":"Proceedings of the 31st International Conference on Computational Linguistics","author":"Qu Xiaoye","year":"2025","unstructured":"Xiaoye Qu, Jiashuo Sun, Wei Wei, Daizong Liu, Jianfeng Dong, and Yu Cheng. 2025. Look, compare, decide: Alleviating hallucination in large vision-language models via multi-view multi-path reasoning. In Proceedings of the 31st International Conference on Computational Linguistics, 4428\u20134441."},{"key":"e_1_3_2_32_2","doi-asserted-by":"crossref","unstructured":"Anna Rohrbach Lisa Anne Hendricks Kaylee Burns Trevor Darrell and Kate Saenko. 2018. Object hallucination in image captioning. arXiv:1809.02156. Retrieved from https:\/\/arxiv.org\/abs\/1809.02156","DOI":"10.18653\/v1\/D18-1437"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20074-8_9"},{"key":"e_1_3_2_34_2","first-page":"492","volume-title":"Proceedings of the Conference on Robot Learning","author":"Shah Dhruv","year":"2023","unstructured":"Dhruv Shah, B\u0142a\u017cej Osi\u0144ski, Brian Ichter, and Sergey Levine. 2023. LM-Nav: Robotic navigation with large pre-trained models of language, vision, and action. In Proceedings of the Conference on Robot Learning. PMLR, 492\u2013504."},{"key":"e_1_3_2_35_2","doi-asserted-by":"crossref","unstructured":"Uri Shaham Jonathan Herzig Roee Aharoni Idan Szpektor Reut Tsarfaty and Matan Eyal. 2024. Multilingual instruction tuning with just a pinch of multilinguality. arXiv:2401.01854. Retrieved from https:\/\/arxiv.org\/abs\/2401.01854","DOI":"10.18653\/v1\/2024.findings-acl.136"},{"key":"e_1_3_2_36_2","unstructured":"Zhiqing Sun Sheng Shen Shengcao Cao Haotian Liu Chunyuan Li Yikang Shen Chuang Gan Liang-Yan Gui Yu-Xiong Wang Yiming Yang et al. 2023. Aligning large multimodal models with factually augmented RLHF. arXiv:2309.14525. Retrieved from https:\/\/arxiv.org\/abs\/2309.14525"},{"key":"e_1_3_2_37_2","unstructured":"Xiaoyu Tian Junru Gu Bailin Li Yicheng Liu Chenxu Hu Yang Wang Kun Zhan Peng Jia Xianpeng Lang and Hang Zhao. 2024. DriveVLM: The convergence of autonomous driving and large vision-language models. arXiv:2402.12289. Retrieved from https:\/\/arxiv.org\/abs\/2402.12289"},{"key":"e_1_3_2_38_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_2_39_2","unstructured":"Junyang Wang Yuhang Wang Guohai Xu Jing Zhang Yukai Gu Haitao Jia Ming Yan Ji Zhang and Jitao Sang. 2023. An LLM-free multi-dimensional benchmark for MLLMs hallucination evaluation. arXiv:2311.07397. Retrieved from https:\/\/arxiv.org\/abs\/2311.07397"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-53302-0_3"},{"key":"e_1_3_2_41_2","article-title":"VisionLLM: Large language model is also an open-ended decoder for vision-centric tasks","volume":"36","author":"Wang Wenhai","year":"2024","unstructured":"Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, et al. 2024. VisionLLM: Large language model is also an open-ended decoder for vision-centric tasks. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_42_2","unstructured":"Weihan Wang Qingsong Lv Wenmeng Yu Wenyi Hong Ji Qi Yan Wang Junhui Ji Zhuoyi Yang Lei Zhao Xixuan Song et al. 2023. CogVLM: Visual expert for pretrained language models. arXiv:2311.03079. Retrieved from https:\/\/arxiv.org\/abs\/2311.03079"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3631710"},{"key":"e_1_3_2_44_2","first-page":"9387","article-title":"Dual modality prompt tuning for vision-language pre-trained model","volume":"25","author":"Xing Yinghui","year":"2023","unstructured":"Yinghui Xing, Qirui Wu, De Cheng, Shizhou Zhang, Guoqiang Liang, Peng Wang, and Yanning Zhang. 2023. Dual modality prompt tuning for vision-language pre-trained model. IEEE Transactions on Multimedia 25 (2023), 9387\u20139399.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_2_45_2","unstructured":"Runsen Xu Shuai Yang Xiaolong Wang Tai Wang Yilun Chen Jiangmiao Pang and Dahua Lin. 2023. PointLLM: Empowering large language models to understand point clouds. arXiv:2308.16911. Retrieved from https:\/\/arxiv.org\/abs\/2308.16911"},{"key":"e_1_3_2_46_2","unstructured":"Qinghao Ye Haiyang Xu Jiabo Ye Ming Yan Haowei Liu Qi Qian Ji Zhang Fei Huang and Jingren Zhou. 2023. mPLUG-Owl2: Revolutionizing multi-modal large language model with modality collaboration. arXiv:2311.04257. Retrieved from https:\/\/arxiv.org\/abs\/2311.04257"},{"key":"e_1_3_2_47_2","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Zhou Kun","year":"2023","unstructured":"Kun Zhou, Jinpeng Wang Wayne, Xin Zhao, Yifan Li, Yifan Du, and Ji-Rong Wen. 2023. Evaluating object hallucination in large vision-language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Retrieved from https:\/\/openreview.net\/forum?id=xozJw0kZXF"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01230"},{"key":"e_1_3_2_49_2","doi-asserted-by":"crossref","unstructured":"Tianyu Yu Yuan Yao Haoye Zhang Taiwen He Yifeng Han Ganqu Cui Jinyi Hu Zhiyuan Liu Hai-Tao Zheng Maosong Sun et al. 2023. RLHF-V: Towards trustworthy MLLMs via behavior alignment from fine-grained correctional human feedback. arXiv:2312.00849. Retrieved from https:\/\/arxiv.org\/abs\/2312.00849","DOI":"10.1109\/CVPR52733.2024.01310"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3397191"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3369699"},{"key":"e_1_3_2_52_2","unstructured":"Zhiyuan Zhao Bin Wang Linke Ouyang Xiaoyi Dong Jiaqi Wang and Conghui He. 2023. Beyond hallucinations: Enhancing LVLMs through hallucination-aware direct preference optimization. arXiv:2311.16839. Retrieved from https:\/\/arxiv.org\/abs\/2311.16839"},{"key":"e_1_3_2_53_2","unstructured":"Deyao Zhu Jun Chen Xiaoqian Shen Xiang Li and Mohamed Elhoseiny. 2023. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. arXiv:2304.10592. Retrieved from https:\/\/arxiv.org\/abs\/2304.10592"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3265842"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3673231"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3797025","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T07:52:58Z","timestamp":1778831578000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3797025"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,15]]},"references-count":54,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3797025"],"URL":"https:\/\/doi.org\/10.1145\/3797025","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,15]]},"assertion":[{"value":"2024-10-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-24","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}