{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T05:38:38Z","timestamp":1782452318901,"version":"3.54.5"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T00:00:00Z","timestamp":1778803200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/legalcode"}],"funder":[{"name":"Science and Technology Development Fund of Macau","award":["0035\/2023\/ITP1, 0021\/2023\/RIA1, and 0069\/2025\/RIB2"],"award-info":[{"award-number":["0035\/2023\/ITP1, 0021\/2023\/RIA1, and 0069\/2025\/RIB2"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62271361"],"award-info":[{"award-number":["62271361"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Hubei Provincial Key Research and Development Program","award":["2024BAB039"],"award-info":[{"award-number":["2024BAB039"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>\n                    In\n                    <jats:bold>Visual Question Answering (VQA)<\/jats:bold>\n                    , both the image and its accompanying question serve as the primary sources of information for the model. Conventional approaches typically rely heavily on dense visual representations for reasoning and answer prediction. However, when the visual and textual modalities are imbalanced or semantically misaligned, such disparities hinder effective multimodal learning and inference. To address this issue, we propose a multimodal information adjustment method, the\n                    <jats:bold>Visual Text Information Adjuster (ViTA)<\/jats:bold>\n                    . ViTA investigates the impact of embedding textual cues within images on the VQA process and promotes cross-modal balance to improve accuracy. Specifically, since image content often dominates over question content, ViTA adjusts the balance by either masking visual information or augmenting it with object-word visual cues directly embedded in the image. Experimental results validate our hypothesis and further demonstrate that ViTA can serve as an effective data augmentation strategy, yielding measurable improvements across multiple VQA models. The code will be released at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/xqx23\/ViTA\">https:\/\/github.com\/xqx23\/ViTA<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3805048","type":"journal-article","created":{"date-parts":[[2026,3,28]],"date-time":"2026-03-28T08:41:40Z","timestamp":1774687300000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Concise Object-word Visuals as Effective Cues for Visual Question Answering"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-4354-8371","authenticated-orcid":false,"given":"Quanxing","family":"Xu","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, Macau University of Science and Technology, Macau SAR, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8313-5749","authenticated-orcid":false,"given":"Ling","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Macau University of Science and Technology, Macau SAR, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5242-0467","authenticated-orcid":false,"given":"Xian","family":"Zhong","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Transportation Internet of Things, School of Computer Science and Artificial Intelligence, Wuhan University of Technology, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8153-9977","authenticated-orcid":false,"given":"Feifei","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Tianjin University of Technology, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1769-6126","authenticated-orcid":false,"given":"Rubing","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Macau University of Science and Technology, Macau SAR, China and Macau University of Science and Technology, Zhuhai MUST Science and Technology Research Institute, Zhuhai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,5,15]]},"reference":[{"key":"e_1_3_1_1_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00522"},{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00774"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_1_4_2","first-page":"2425","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Antol Stanislaw","year":"2015","unstructured":"Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: Visual question answering. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2425\u20132433."},{"key":"e_1_3_1_5_2","unstructured":"Shuai Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge Sibo Song Kai Dang Peng Wang Shijie Wang Jun Tang et al. 2025. Qwen2.5-VL technical report. arXiv:2502.13923. Retrieved from https:\/\/arxiv.org\/abs\/2502.13923"},{"key":"e_1_3_1_6_2","first-page":"1877","volume-title":"Proceedings of the International Conference on Neural Information Processing Systems","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Proceedings of the International Conference on Neural Information Processing Systems, 1877\u20131901."},{"key":"e_1_3_1_7_2","first-page":"839","volume-title":"Proceedings of the International Conference on Neural Information Processing Systems","author":"Cad\u00e8ne R\u00e9mi","year":"2019","unstructured":"R\u00e9mi Cad\u00e8ne, Corentin Dancette, H\u00e9di Ben-Younes, Matthieu Cord, and Devi Parikh. 2019. RUBi: Reducing unimodal biases for visual question answering. In Proceedings of the International Conference on Neural Information Processing Systems, 839\u2013850."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2023.110084"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2023.110706"},{"key":"e_1_3_1_10_2","article-title":"Fine-grained lexical-centric semantic network for coherent video paragraph captioning","author":"Chen Shuqin","year":"2025","unstructured":"Shuqin Chen, Xian Zhong, Xingrui Yang, Li Yang, Bin Sheng, and Alex Chichung Kot. 2025. Fine-grained lexical-centric semantic network for coherent video paragraph captioning. IEEE Transactions on Multimedia (2025).","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3679203"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1418"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00850"},{"key":"e_1_3_1_14_2","first-page":"889","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Fan Angela","year":"2018","unstructured":"Angela Fan, Mike Lewis, and Yann N. Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. Iryna Gurevych and Yusuke Miyao (Eds.), 889\u2013898."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01918"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.670"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2844175"},{"key":"e_1_3_1_18_2","first-page":"3909","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Huang Chengyue","year":"2025","unstructured":"Chengyue Huang, Brisa Maneechotesuwan, Shivang Chopra, and Zsolt Kira. 2025. FRAMES-VQA: Benchmarking fine-tuning robustness across multi-modal shifts in visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 3909\u20133918."},{"key":"e_1_3_1_19_2","first-page":"1571","volume-title":"Proceedings of the International Conference on Neural Information Processing Systems","author":"Kim Jin-Hwa","year":"2018","unstructured":"Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. 2018. Bilinear attention networks. In Proceedings of the International Conference on Neural Information Processing Systems, 1571\u20131581."},{"key":"e_1_3_1_20_2","first-page":"7918","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Kumar Dhruv","year":"2020","unstructured":"Dhruv Kumar, Lili Mou, Lukasz Golab, and Olga Vechtomova. 2020. Iterative edit-based unsupervised sentence simplification. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, 7918\u20137928."},{"key":"e_1_3_1_21_2","first-page":"4364","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Kwon Junehyoung","year":"2025","unstructured":"Junehyoung Kwon, Mihyeon Kim, Eunju Lee, Juhwan Choi, and YoungBin Kim. 2025. See-saw modality balance: See gradient, and sew impaired vision-language balance to mitigate dominant modality bias. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4364\u20134378."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475492"},{"key":"e_1_3_1_23_2","unstructured":"Feng Li Renrui Zhang Hao Zhang Yuanhan Zhang Bo Li Wei Li Zejun Ma and Chunyuan Li. 2024. LLaVA-NeXT-Interleave: Tackling multi-image video and 3D in large multimodal models. arXiv:2407.07895. Retrieved from https:\/\/arxiv.org\/abs\/2407.07895"},{"key":"e_1_3_1_24_2","first-page":"12888","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. 2022. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine Learning, 12888\u201312900."},{"key":"e_1_3_1_25_2","first-page":"5195","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Liang Xiao","year":"2023","unstructured":"Xiao Liang, Di Wang, Quan Wang, Bo Wan, Lingling An, and Lihuo He. 2023. Language-Guided visual aggregation network for video question answering. In Proceedings of the ACM International Conference on Multimedia, 5195\u20135203."},{"key":"e_1_3_1_26_2","first-page":"936","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lin Tsung-Yi","year":"2017","unstructured":"Tsung-Yi Lin, Piotr Doll\u00e1r, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie. 2017. Feature pyramid networks for object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 936\u2013944."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3498340"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00331"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01251"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3097502"},{"key":"e_1_3_1_31_2","first-page":"1562","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition: Workshops","author":"\u00d6zdemir \u00d6vg\u00fc","year":"2024","unstructured":"\u00d6vg\u00fc \u00d6zdemir and Erdem Akag\u00fcnd\u00fcz. 2024. Enhancing visual question answering through question-driven image captions as prompts. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition: Workshops, 1562\u20131571."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3355640"},{"key":"e_1_3_1_33_2","first-page":"8228","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Peng Xiaokang","year":"2022","unstructured":"Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu. 2022. Balanced multimodal learning via on-the-fly gradient modulation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 8228\u20138237."},{"key":"e_1_3_1_34_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, 8748\u20138763."},{"issue":"6","key":"e_1_3_1_35_2","doi-asserted-by":"crossref","first-page":"159","DOI":"10.1007\/s10462-025-11163-4","article-title":"GFSNet: Gaussian Fourier with sparse attention network for visual question answering","volume":"58","author":"Shen Xiang","year":"2025","unstructured":"Xiang Shen, Dezhi Han, Chin-Chen Chang, Ammar Oad, and Huafeng Wu. 2025. GFSNet: Gaussian Fourier with sparse attention network for visual question answering. Artificial Intelligence Review 58, 6 (2025), 159.","journal-title":"Artificial Intelligence Review"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2025.128857"},{"key":"e_1_3_1_37_2","first-page":"2920","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Shi Cheng","year":"2023","unstructured":"Cheng Shi and Sibei Yang. 2023. LoGoPrompt: Synthetic text images can be good visual prompts for vision-language models. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2920\u20132929."},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","first-page":"4242","DOI":"10.18653\/v1\/2024.findings-emnlp.245","volume-title":"Findings of the Association for Computational Linguistics: EMNLP","author":"Song Lingyun","year":"2024","unstructured":"Lingyun Song, Chengkun Yang, Xuanyu Li, and Xuequn Shang. 2024. A robust dual-debiasing VQA model based on counterfactual causal effect. In Findings of the Association for Computational Linguistics: EMNLP, 4242\u20134252."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1514"},{"key":"e_1_3_1_40_2","first-page":"4631","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Tapaswi Makarand","year":"2016","unstructured":"Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016. MovieQA: Understanding stories in movies through question-answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 4631\u20134640."},{"key":"e_1_3_1_41_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et al. 2023. LLaMA 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_1_42_2","unstructured":"A\u00e4ron van den Oord Yazhe Li and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv:1807.03748. Retrieved from https:\/\/arxiv.org\/abs\/1807.03748"},{"key":"e_1_3_1_43_2","first-page":"1","volume-title":"Proceedings of the International Conference on Advances in Data Engineering and Intelligent Computing Systems","author":"Varghese Rejin","year":"2024","unstructured":"Rejin Varghese and M. Sambath. 2024. YOLOv8: A novel object detection algorithm with enhanced performance and robustness. In Proceedings of the International Conference on Advances in Data Engineering and Intelligent Computing Systems, 1\u20136."},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3380259"},{"key":"e_1_3_1_45_2","first-page":"274","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Wang Haibo","year":"2024","unstructured":"Haibo Wang and Weifeng Ge. 2024. Q&A prompts: Discovering rich visual clues through mining question-answer prompts for VQA requiring diverse world knowledge. In Proceedings of the European Conference on Computer Vision, 274\u2013292."},{"key":"e_1_3_1_46_2","unstructured":"Jianfeng Wang Zhengyuan Yang Xiaowei Hu Linjie Li Lin Kevin Gan Zhe Liu Zicheng Ce Liu and Lijuan Wang. 2022. GIT: A generative image-to-text transformer for vision and language. arXiv:2205.14100. Retrieved from https:\/\/arxiv.org\/abs\/2205.14100"},{"key":"e_1_3_1_47_2","first-page":"9196","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Kaixin","year":"2019","unstructured":"Kaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou, and Jiashi Feng. 2019. PANet: Few-Shot image semantic segmentation with prototype alignment. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 9196\u20139205."},{"key":"e_1_3_1_48_2","unstructured":"Peng Wang Shuai Bai Sinan Tan Shijie Wang Zhihao Fan Jinze Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge et al. 2024. Qwen2-VL: Enhancing vision-language model\u2019s perception of the world at any resolution. arXiv:2409.12191. Retrieved from https:\/\/arxiv.org\/abs\/2409.12191"},{"key":"e_1_3_1_49_2","first-page":"12692","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Weiyao","year":"2020","unstructured":"Weiyao Wang, Du Tran, and Matt Feiszli. 2020. What makes training multi-modal classification networks hard? In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 12692\u201312702."},{"key":"e_1_3_1_50_2","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Wei Shicai","year":"2025","unstructured":"Shicai Wei, Chunbo Luo, and Yang Luo. 2025. Improving multimodal learning via imbalanced learning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision."},{"key":"e_1_3_1_51_2","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Wei Yake","year":"2024","unstructured":"Yake Wei, Siwei Li, Ruoxuan Feng, and Di Hu. 2024. Diagnosing and re-learning for balanced multimodal learning. In Proceedings of the European Conference on Computer Vision."},{"key":"e_1_3_1_52_2","first-page":"2137","volume-title":"IEEE Transactions on Multimedia","volume":"26","author":"Wen Zhiquan","year":"2024","unstructured":"Zhiquan Wen, Shuaicheng Niu, Ge Li, Qingyao Wu, Mingkui Tan, and Qi Wu. 2024. Test-time model adaptation for visual question answering with debiased self-supervisions. IEEE Transactions on Multimedia 26 (2024), 2137\u20132147."},{"key":"e_1_3_1_53_2","first-page":"3936","volume-title":"Proceedings of the International Conference on Computational Linguistics","author":"Yang Li","year":"2025","unstructured":"Li Yang, Zhiding Xiao, Wenxin Huang, and Xian Zhong. 2025. StoryLLaVA: Enhancing visual storytelling with multi-modal large language models. In Proceedings of the International Conference on Computational Linguistics, 3936\u20133951."},{"key":"e_1_3_1_54_2","first-page":"21","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yang Zichao","year":"2016","unstructured":"Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alexander J. Smola. 2016. Stacked attention networks for image question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 21\u201329."},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00644"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00681"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-022-01653-1"},{"key":"e_1_3_1_58_2","first-page":"955","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Zhu Jiawei","year":"2024","unstructured":"Jiawei Zhu, Yishu Liu, Huanjia Zhu, Hui Lin, Yuncheng Jiang, Zheng Zhang, and Bingzhi Chen. 2024. Combating visual question answering hallucinations via robust multi-space co-debias learning. In Proceedings of the ACM International Conference on Multimedia, 955\u2013964."},{"key":"e_1_3_1_59_2","unstructured":"Jinguo Zhu Weiyun Wang Zhe Chen Zhaoyang Liu Shenglong Ye Lixin Gu Hao Tian Yuchen Duan Weijie Su Jie Shao et al. 2025. InternVL3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv:2504.10479. Retrieved from https:\/\/arxiv.org\/abs\/2504.10479"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3805048","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T07:52:17Z","timestamp":1778831537000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3805048"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,15]]},"references-count":59,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3805048"],"URL":"https:\/\/doi.org\/10.1145\/3805048","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,15]]},"assertion":[{"value":"2025-09-10","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}