{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T16:42:03Z","timestamp":1783096923557,"version":"3.54.6"},"reference-count":84,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/501100004739","name":"Youth Innovation Promotion Association CAS","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004739","id-type":"DOI","asserted-by":"crossref"}]},{"name":"GPU cluster"},{"name":"MCC Lab of the Information Science and Technology Institution, USTC"},{"name":"Supercomputing Center of USTC."}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>\n                    The advent of Large Multimodal Models (LMMs) has fueled extensive research due to their sophisticated reasoning capabilities. However, for understanding text-rich images, challenges persist in fully leveraging the potential of LMMs, and existing methods struggle with effectively processing high-resolution images. Addressing this, we introduce TextCoT, a training-free Chain-of-Thought framework that improves text-rich image understanding by leveraging LMMs\u2019 captioning abilities for global context and detailed local textual analysis. TextCoT comprises three stages: Global Context Generation, Macro-Scale Positioning, and Fine-Grained Visual Inspection\u2014each contributing to a comprehensive understanding and precise information extraction needed for accurate question-answering. Our method requires no additional training, offering immediate plug-and-play functionality. We have demonstrated TextCoT\u2019s effectiveness and adaptability across various benchmarks. The source code is available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/bzluan\/TextCoT\">https:\/\/github.com\/bzluan\/TextCoT<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3785474","type":"journal-article","created":{"date-parts":[[2026,1,5]],"date-time":"2026-01-05T11:23:08Z","timestamp":1767612188000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["TextCoT: Zoom-In for Enhanced Multimodal Text-Rich Image Understanding"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-6554-1187","authenticated-orcid":false,"given":"Bozhi","family":"Luan","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8127-6639","authenticated-orcid":false,"given":"Hao","family":"Feng","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-4283-1917","authenticated-orcid":false,"given":"Hong","family":"Chen","sequence":"additional","affiliation":[{"name":"Merchants Union Consumer Finance Company Limited, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4741-8231","authenticated-orcid":false,"given":"Yonghui","family":"Wang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1690-9836","authenticated-orcid":false,"given":"Wengang","family":"Zhou","sequence":"additional","affiliation":[{"name":"EEIS Department, University of Science and Technology of\u00a0China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2188-3028","authenticated-orcid":false,"given":"Houqiang","family":"Li","sequence":"additional","affiliation":[{"name":"EEIS Department, University of Science and Technology of\u00a0China, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,3,23]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et al. 2023. GPT-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_1_3_2","unstructured":"Anthropic. 2024. The Claude 3 model family: Opus Sonnet Haiku. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:268232499"},{"key":"e_1_3_1_4_2","unstructured":"Jinze Bai Shuai Bai Yunfei Chu Zeyu Cui Kai Dang Xiaodong Deng Yang Fan Wenbin Ge Yu Han Fei Huang et al. 2023. Qwen technical report. arXiv:2309.16609. Retrieved from https:\/\/arxiv.org\/abs\/2309.16609"},{"key":"e_1_3_1_5_2","unstructured":"Jinze Bai Shuai Bai Shusheng Yang Shijie Wang Sinan Tan Peng Wang Junyang Lin Chang Zhou and Jingren Zhou. 2023. Qwen-VL: A versatile vision-language model for understanding localization text reading and beyond. arXiv:2308.12966. Retrieved from https:\/\/arxiv.org\/abs\/arXiv:2308.12966"},{"key":"e_1_3_1_6_2","unstructured":"Shuai Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge Sibo Song Kai Dang Peng Wang Shijie Wang Jun Tang Humen Zhong et al. 2025. Qwen2.5-VL technical report. arXiv:2502.13923. Retrieved from https:\/\/arxiv.org\/abs\/2502.13923"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i16.29720"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3294822"},{"key":"e_1_3_1_9_2","first-page":"4291","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Furkan Biten Ali","year":"2019","unstructured":"Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Mar\u00e7al Rusinol, Ernest Valveny, C. V. Jawahar, and Dimosthenis Karatzas. 2019. Scene text visual question answering. In Proceedings of the IEEE International Conference on Computer Vision, 4291\u20134301."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.5555\/3495724.3495883"},{"key":"e_1_3_1_11_2","unstructured":"Yuhang Cao Pan Zhang Xiaoyi Dong Dahua Lin and Jiaqi Wang. 2024. DualFocus: Integrating macro and micro perspectives in multi-modal large language models. arXiv:2402.14767. Retrieved from https:\/\/arxiv.org\/abs\/2402.14767"},{"key":"e_1_3_1_12_2","unstructured":"Keqin Chen Zhao Zhang Weili Zeng Richong Zhang Feng Zhu and Rui Zhao. 2023. Shikra: Unleashing multimodal LLM\u2019s referential dialogue magic. arXiv:2306.15195. Retrieved from https:\/\/arxiv.org\/abs\/2306.15195"},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Lin Chen Jisong Li Xiaoyi Dong Pan Zhang Conghui He Jiaqi Wang Feng Zhao and Dahua Lin. 2023. ShareGPT4V: Improving large multi-modal models with better captions. arXiv:2311.12793. Retrieved from https:\/\/arxiv.org\/abs\/2311.12793","DOI":"10.1007\/978-3-031-72643-9_22"},{"key":"e_1_3_1_14_2","doi-asserted-by":"crossref","unstructured":"Zhe Chen Weiyun Wang Hao Tian Shenglong Ye Zhangwei Gao Erfei Cui Wenwen Tong Kongzhi Hu Jiapeng Luo Zheng Ma et al. 2024. How far are we to GPT-4V? Closing the gap to commercial multimodal models with open-source suites. arXiv:2404.16821. Retrieved from https:\/\/arxiv.org\/abs\/2404.16821","DOI":"10.1007\/s11432-024-4231-5"},{"key":"e_1_3_1_15_2","unstructured":"Wei-Lin Chiang Zhuohan Li Zi Lin Ying Sheng Zhanghao Wu Hao Zhang Lianmin Zheng Siyuan Zhuang Yonghao Zhuang Joseph E. Gonzalez et al. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. Retrieved from https:\/\/vicuna.lmsys.org"},{"issue":"240","key":"e_1_3_1_16_2","first-page":"1","article-title":"PaLM: Scaling language modeling with pathways","volume":"24","author":"Chowdhery Aakanksha","year":"2023","unstructured":"Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023. PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research 24, 240 (2023), 1\u2013113.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_17_2","first-page":"49250","volume-title":"Proceedings of the 37th International Conference on Advances in Neural Information Processing Systems","author":"Dai Wenliang","year":"2024","unstructured":"Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N. Fung, and Steven Hoi. 2024. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. In Proceedings of the 37th International Conference on Advances in Neural Information Processing Systems, 49250\u201349267."},{"key":"e_1_3_1_18_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","unstructured":"Hao Feng Qi Liu Hao Liu Wengang Zhou Houqiang Li and Can Huang. 2023. DocPedia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding. arXiv:2311.11810. Retrieved from https:\/\/arxiv.org\/abs\/2311.11810","DOI":"10.1007\/s11432-024-4250-y"},{"key":"e_1_3_1_20_2","unstructured":"Hao Feng Zijian Wang Jingqun Tang Jinghui Lu Wengang Zhou Houqiang Li and Can Huang. 2023. UniDoc: A universal large multimodal model for simultaneous text detection recognition spotting and understanding. arXiv:2308.11592. Retrieved from https:\/\/arxiv.org\/abs\/2308.11592"},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","unstructured":"Pei Fu Tongkun Guan Zining Wang Zhentao Guo Chen Duan Hao Sun Boming Chen Jiayao Ma Qianyi Jiang Kai Zhou et al. 2025. Multimodal large language models for text-rich image understanding: A comprehensive review. arXiv:2502.16586. Retrieved from https:\/\/arxiv.org\/abs\/2502.16586","DOI":"10.18653\/v1\/2025.findings-acl.1023"},{"key":"e_1_3_1_22_2","unstructured":"Anwen Hu Haiyang Xu Jiabo Ye Ming Yan Liang Zhang Bo Zhang Chen Li Ji Zhang Qin Jin Fei Huang et al. 2024. mPLUG-DocOwl 1.5: Unified structure learning for OCR-free document understanding. arXiv:2403.12895. Retrieved from https:\/\/arxiv.org\/abs\/2403.12895"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611899"},{"key":"e_1_3_1_24_2","unstructured":"Mingxin Huang Yuliang Liu Dingkang Liang Lianwen Jin and Xiang Bai. 2024. Mini-monkey: Alleviating the semantic sawtooth effect for lightweight MLLMS via complementary image pyramid. arXiv:2408.02034. Retrieved from https:\/\/arxiv.org\/abs\/2408.02034"},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","unstructured":"Qidong Huang Xiaoyi Dong Pan Zhang Bin Wang Conghui He Jiaqi Wang Dahua Lin Weiming Zhang and Nenghai Yu. 2023. Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation. arXiv:2311.17911. Retrieved from https:\/\/arxiv.org\/abs\/2311.17911","DOI":"10.1109\/CVPR52733.2024.01274"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/icdar.2019.00244"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/icdarw.2019.10029"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1086"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.5555\/3600270.3601883"},{"key":"e_1_3_1_30_2","unstructured":"Ivan Krasin Tom Duerig Neil Alldrin Vittorio Ferrari Sami Abu-El-Haija Alina Kuznetsova Hassan Rom Jasper Uijlings Stefan Popov Andreas Veit et al. 2017. OpenImages: A public dataset for large-scale multi-label and multi-class image classification. Retrieved from https:\/\/storage.googleapis.com\/openimages\/web\/index.html"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_1_32_2","first-page":"36","volume-title":"Proceedings of the International Conference on Document Analysis and Recognition","author":"Kuang Jianfeng","year":"2023","unstructured":"Jianfeng Kuang, Wei Hua, Dingkang Liang, Mingkun Yang, Deqiang Jiang, Bo Ren, and Xiang Bai. 2023. Visual information extraction in the wild: Practical dataset and end-to-end solution. In Proceedings of the International Conference on Document Analysis and Recognition. ACM, 36\u201353."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612389"},{"key":"e_1_3_1_34_2","unstructured":"Xuanyu Lei Zonghan Yang Xinrui Chen Peng Li and Yang Liu. 2024. Scaffolding coordinates to promote vision-language coordination in large multi-modal models. arXiv:2402.12058. Retrieved from https:\/\/arxiv.org\/abs\/2402.12058"},{"key":"e_1_3_1_35_2","first-page":"12888","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine Learning. ACM, 12888\u201312900."},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","unstructured":"Xin Li Yunfei Wu Xinghua Jiang Zhihao Guo Mingming Gong Haoyu Cao Yinsong Liu Deqiang Jiang and Xing Sun. 2024. Enhancing visual document understanding with contrastive learning in large visual-language models. arXiv:2402.19014. Retrieved from https:\/\/arxiv.org\/abs\/2402.19014","DOI":"10.1109\/CVPR52733.2024.01472"},{"key":"e_1_3_1_37_2","unstructured":"Zhang Li Biao Yang Qiang Liu Zhiyin Ma Shuo Zhang Jingxu Yang Yabo Sun Yuliang Liu and Xiang Bai. 2023. Monkey: Image resolution and text label are important things for large multi-modal models. arXiv:2311.06607. Retrieved from https:\/\/arxiv.org\/abs\/2311.06607"},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","unstructured":"Ziyi Lin Chris Liu Renrui Zhang Peng Gao Longtian Qiu Han Xiao Han Qiu Chen Lin Wenqi Shao Keqin Chen et al. 2023. SPHINX: The joint mixing of weights tasks and visual embeddings for multi-modal large language models. arXiv:2311.07575. Retrieved from https:\/\/arxiv.org\/abs\/2311.07575","DOI":"10.1007\/978-3-031-73033-7_3"},{"key":"e_1_3_1_39_2","first-page":"15534","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu Chaohu","year":"2024","unstructured":"Chaohu Liu, Kun Yin, Haoyu Cao, Xinghua Jiang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun, and Linli Xu. 2024. HRVDA: High-resolution visual document assistant. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 15534\u201315545."},{"key":"e_1_3_1_40_2","unstructured":"Haotian Liu Chunyuan Li Yuheng Li and Yong Jae Lee. 2023. Improved baselines with visual instruction tuning. arXiv:2310.03744. Retrieved from https:\/\/arxiv.org\/abs\/2310.03744"},{"key":"e_1_3_1_41_2","first-page":"34892","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing System","author":"Liu Haotian","year":"2024","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruction tuning. In Proceedings of the 37th International Conference on Neural Information Processing System, 34892\u201334916."},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2024.103809"},{"key":"e_1_3_1_43_2","unstructured":"Yuliang Liu Zhang Li Hongliang Li Wenwen Yu Mingxin Huang Dezhi Peng Mingyu Liu Mingrui Chen Chunyuan Li Lianwen Jin et al. 2023. On the hidden mystery of OCR in large multimodal models. arXiv:2305.07895. Retrieved from https:\/\/arxiv.org\/abs\/2305.07895"},{"key":"e_1_3_1_44_2","unstructured":"Yuliang Liu Biao Yang Qiang Liu Li Zhang Zhiyin Ma Shuo Zhang and Xiang Bai. 2024. TextMonkey: An OCR-Free large multimodal model for understanding document. arXiv:2403.04473. Retrieved from https:\/\/arxiv.org\/abs\/2403.04473"},{"key":"e_1_3_1_45_2","first-page":"15630","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Luo Chuwei","year":"2024","unstructured":"Chuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng, Zhi Yu, and Cong Yao. 2024. LayoutLLM: Layout instruction tuning with large language models for document understanding. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 15630\u201315640."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3237002"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.9"},{"key":"e_1_3_1_48_2","doi-asserted-by":"crossref","unstructured":"Ahmed Masry Do Xuan Long Jia Qing Tan Shafiq Joty and Enamul Hoque. 2022. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. arXiv:2203.10244. Retrieved from https:\/\/arxiv.org\/abs\/2203.10244","DOI":"10.18653\/v1\/2022.findings-acl.177"},{"key":"e_1_3_1_49_2","first-page":"1697","volume-title":"Proceedings of the IEEE Winter Conference on Applications of Computer Vision","author":"Mathew Minesh","year":"2022","unstructured":"Minesh Mathew, Viraj Bagal, Rub\u00e8n Tito, Dimosthenis Karatzas, Ernest Valveny, and C. V. Jawahar. 2022. InfographicVQA. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision. IEEE, 1697\u20131706."},{"key":"e_1_3_1_50_2","first-page":"2200","volume-title":"Proceedings of the IEEE Winter Conference on Applications of Computer Vision","author":"Mathew Minesh","year":"2021","unstructured":"Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar. 2021. DocVQA: A dataset for VQA on document images. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision. IEEE, 2200\u20132209."},{"key":"e_1_3_1_51_2","first-page":"14420","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Mitra Chancharik","year":"2024","unstructured":"Chancharik Mitra, Brandon Huang, Trevor Darrell, and Roei Herzig. 2024. Compositional chain-of-thought prompting for large multimodal models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 14420\u201314431."},{"key":"e_1_3_1_52_2","unstructured":"OpenAI. 2022. ChatGPT. Retrieved from https:\/\/openai.com\/blog\/chatgpt"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.5555\/3600270.3602281"},{"key":"e_1_3_1_54_2","unstructured":"Zhiliang Peng Wenhui Wang Li Dong Yaru Hao Shaohan Huang Shuming Ma and Furu Wei. 2023. Kosmos-2: Grounding multimodal large language models to the world. arXiv:2306.14824. Retrieved from https:\/\/arxiv.org\/abs\/2306.14824"},{"key":"e_1_3_1_55_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, 8748\u20138763."},{"issue":"140","key":"e_1_3_1_56_2","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21, 140 (2020), 1\u201367.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_57_2","unstructured":"Fobo Shi Peijun Qing Dong Yang Nan Wang Youbo Lei Haonan Lu Xiaodong Lin and Duantengchuan Li. 2023. Prompt space optimizing few-shot reasoning success with large language models. arXiv:2306.03799. Retrieved from https:\/\/arxiv.org\/abs\/2306.03799"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3287038"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00851"},{"key":"e_1_3_1_60_2","unstructured":"Zeyi Sun Ye Fang Tong Wu Pan Zhang Yuhang Zang Shu Kong Yuanjun Xiong Dahua Lin and Jiaqi Wang. 2023. Alpha-CLIP: A CLIP model focusing on wherever you want. arXiv:2312.03818. Retrieved from https:\/\/arxiv.org\/abs\/2312.03818"},{"key":"e_1_3_1_61_2","first-page":"19071","volume-title":"Proceedings of the 37th AAAI Conference on Artificial Intelligence","author":"Tanaka Ryota","year":"2024","unstructured":"Ryota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito, and Jun Suzuki. 2024. InstructDoc: A dataset for zero-shot generalization of visual document understanding with instructions. In Proceedings of the 37th AAAI Conference on Artificial Intelligence, 19071\u201319079."},{"key":"e_1_3_1_62_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et al. 2023. LLaMA: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_1_63_2","first-page":"5309","volume-title":"Proceedings of the 37th AAAI Conference on Artificial Intelligence","author":"Wang Bin","year":"2024","unstructured":"Bin Wang, Fan Wu, Xiao Han, Jiahui Peng, Huaping Zhong, Pan Zhang, Xiaoyi Dong, Weijia Li, Wei Li, Jiaqi Wang, et al. 2024. VIGC: Visual instruction generation and correction. In Proceedings of the 37th AAAI Conference on Artificial Intelligence, 5309\u20135317."},{"key":"e_1_3_1_64_2","first-page":"2555","volume-title":"Proceedings of the 37th AAAI Conference on Artificial Intelligence","author":"Wang Jianyi","year":"2023","unstructured":"Jianyi Wang, Kelvin C. K. Chan, and Chen Change Loy. 2023. Exploring CLIP for assessing the look and feel of images. In Proceedings of the 37th AAAI Conference on Artificial Intelligence, 2555\u20132563."},{"key":"e_1_3_1_65_2","unstructured":"Xuezhi Wang Jason Wei Dale Schuurmans Quoc Le Ed Chi Sharan Narang Aakanksha Chowdhery and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv:2203.11171. Retrieved from https:\/\/arxiv.org\/abs\/2203.11171"},{"key":"e_1_3_1_66_2","unstructured":"Yonghui Wang Wengang Zhou Hao Feng Keyi Zhou and Houqiang Li. 2023. Towards improving document understanding: An exploration on text-grounding via MLLMs. arXiv:2311.13194. Retrieved from https:\/\/arxiv.org\/abs\/2311.13194"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.5555\/3600270.3600887"},{"key":"e_1_3_1_68_2","unstructured":"Haoran Wei Lingyu Kong Jinyue Chen Liang Zhao Zheng Ge Jinrong Yang Jianjian Sun Chunrui Han and Xiangyu Zhang. 2023. Vary: Scaling up the vision vocabulary for large vision-language models. arXiv:2312.06109. Retrieved from https:\/\/arxiv.org\/abs\/2312.06109"},{"key":"e_1_3_1_69_2","unstructured":"Haoran Wei Lingyu Kong Jinyue Chen Liang Zhao Zheng Ge En Yu Jianjian Sun Chunrui Han and Xiangyu Zhang. 2024. Small language model meets with reinforced vision vocabulary. arXiv:2401.12503. Retrieved from https:\/\/arxiv.org\/abs\/2401.12503"},{"key":"e_1_3_1_70_2","doi-asserted-by":"crossref","unstructured":"Jason Wei Xuezhi Wang Dale Schuurmans Maarten Bosma Fei Xia Ed Chi Quoc V. Le Denny Zhou et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the 36th International Conference on Neural Information Processing Systems 24824\u201324837.","DOI":"10.52202\/068431-1800"},{"key":"e_1_3_1_71_2","unstructured":"Penghao Wu and Saining Xie. 2023. V* : Guided visual search as a core mechanism in multimodal LLMs. arXiv:2312.14135. Retrieved from https:\/\/arxiv.org\/abs\/2312.14135"},{"key":"e_1_3_1_72_2","unstructured":"Yixuan Wu Yizhou Wang Shixiang Tang Wenhao Wu Tong He Wanli Ouyang Jian Wu and Philip Torr. 2024. DetToolChain: A new prompting paradigm to unleash detection ability of MLLM. arXiv:2403.12488. Retrieved from https:\/\/arxiv.org\/abs\/2403.12488"},{"key":"e_1_3_1_73_2","unstructured":"Zhengyuan Yang Linjie Li Kevin Lin Jianfeng Wang Chung-Ching Lin Zicheng Liu and Lijuan Wang. 2023. The dawn of LMMs: Preliminary explorations with GPT-4V(ision). arXiv:2309.17421. Retrieved from https:\/\/arxiv.org\/abs\/2309.17421"},{"key":"e_1_3_1_74_2","first-page":"11809","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923)","author":"Yao Shunyu","year":"2024","unstructured":"Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024. Tree of thoughts: Deliberate problem solving with large language models. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923), 11809\u201311822."},{"key":"e_1_3_1_75_2","unstructured":"Jiabo Ye Anwen Hu Haiyang Xu Qinghao Ye Ming Yan Yuhao Dan Chenlin Zhao Guohai Xu Chenliang Li Junfeng Tian et al. 2023. mPLUG-DocOwl: Modularized multimodal large language model for document understanding. arXiv:2307.02499. Retrieved from https:\/\/arxiv.org\/abs\/2307.02499"},{"key":"e_1_3_1_76_2","first-page":"2841","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Ye Jiabo","year":"2023","unstructured":"Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Guohai Xu, Chenliang Li, Junfeng Tian, Qi Qian, Ji Zhang, et al. 2023. UReader: Universal OCR-free visually-situated language understanding with multimodal large language model. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. ACL, 2841\u20132858."},{"key":"e_1_3_1_77_2","unstructured":"Qinghao Ye Haiyang Xu Guohai Xu Jiabo Ye Ming Yan Yiyang Zhou Junyang Wang Anwen Hu Pengcheng Shi Yaya Shi et al. 2023. mPLUG-Owl: Modularization empowers large language models with multimodality. arXiv:2304.14178. Retrieved from https:\/\/arxiv.org\/abs\/2304.14178"},{"key":"e_1_3_1_78_2","unstructured":"Shukang Yin Chaoyou Fu Sirui Zhao Tong Xu Hao Wang Dianbo Sui Yunhang Shen Ke Li Xing Sun and Enhong Chen. 2023. Woodpecker: Hallucination correction for multimodal large language models. arXiv:2310.16045. Retrieved from https:\/\/arxiv.org\/abs\/2310.16045"},{"key":"e_1_3_1_79_2","first-page":"1261","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Zeng Gangyan","year":"2023","unstructured":"Gangyan Zeng, Yuan Zhang, Yu Zhou, Bo Fang, Guoqing Zhao, Xin Wei, and Weiping Wang. 2023. Filling in the blank: Rationale-augmented prompt tuning for TextVQA. In Proceedings of the ACM International Conference on Multimedia, 1261\u20131272."},{"key":"e_1_3_1_80_2","doi-asserted-by":"crossref","unstructured":"Daoan Zhang Junming Yang Hanjia Lyu Zijian Jin Yuan Yao Mingkai Chen and Jiebo Luo. 2024. CoCoT: Contrastive chain-of-thought prompting for large multimodal models with multiple image inputs. arXiv:2401.02582. Retrieved from https:\/\/arxiv.org\/abs\/2401.02582","DOI":"10.1007\/978-3-031-78456-9_15"},{"key":"e_1_3_1_81_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i9.33076"},{"key":"e_1_3_1_82_2","unstructured":"Zhuosheng Zhang Aston Zhang Mu Li Hai Zhao George Karypis and Alex Smola. 2023. Multimodal chain-of-thought reasoning in language models. arXiv:2302.00923. Retrieved from https:\/\/arxiv.org\/abs\/2302.00923"},{"key":"e_1_3_1_83_2","first-page":"5168","article-title":"DDCoT: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models","volume":"36","author":"Zheng Ge","year":"2023","unstructured":"Ge Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou, and Sibei Yang. 2023. DDCoT: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models. Proceedings of the Advances in Neural Information Processing Systems 36 (2023), 5168\u20135191.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_84_2","unstructured":"Deyao Zhu Jun Chen Xiaoqian Shen Xiang Li and Mohamed Elhoseiny. 2023. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. arXiv:2304.10592. Retrieved from https:\/\/arxiv.org\/abs\/2304.10592"},{"key":"e_1_3_1_85_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00877"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3785474","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,23]],"date-time":"2026-03-23T15:51:51Z","timestamp":1774281111000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3785474"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,23]]},"references-count":84,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3785474"],"URL":"https:\/\/doi.org\/10.1145\/3785474","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,23]]},"assertion":[{"value":"2025-04-02","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-23","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}