{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T03:42:24Z","timestamp":1784691744844,"version":"3.55.0"},"reference-count":64,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2024,1,11]],"date-time":"2024-01-11T00:00:00Z","timestamp":1704931200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2022YFB4500600"],"award-info":[{"award-number":["2022YFB4500600"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62020106007, 62272144, U20A20183, 72188101, 62272435, and U22A2094"],"award-info":[{"award-number":["62020106007, 62272144, U20A20183, 72188101, 62272435, and U22A2094"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Major Project of Anhui Province","award":["202203a05020011"],"award-info":[{"award-number":["202203a05020011"]}]},{"name":"University Synergy Innovation Program of Anhui Province","award":["GXXT-2022-047"],"award-info":[{"award-number":["GXXT-2022-047"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,4,30]]},"abstract":"<jats:p>Effectively leveraging objects and optical character recognition (OCR) tokens to reason out pivotal scene text is critical for the challenging Text-based Visual Question Answering (TextVQA) task. Graph-based models can effectively capture the semantic relationship among visual entities (objects and tokens) and report remarkable performance in TextVQA. However, previous efforts usually leverage all visual entities and ignore the negative effect of superfluous entities. This article presents a Graph Pooling Inference Network (GPIN), which is an evolutionary graph learning method to purify the visual entities and capture the core semantics. It is observed that the dense distribution of reduplicative objects and the crowd of semantically dependent OCR tokens usually co-exist in the image. Motivated by this, GPIN adopts an adaptive node dropping strategy to dynamically downscale semantically closed nodes for graph evolution and update. To deepen the comprehension of scene text, GPIN is a dual-path hierarchical graph architecture that progressively aggregates the evolved object graph and the evolved token graph semantics into a graph vector that serves as visual cues to facilitate the answer reasoning. It can effectively eliminate object redundancy and enhance the association of semantically continuous tokens. Experiments conducted on TextVQA and ST-VQA datasets show that GPIN achieves promising performance compared with state-of-the-art methods.<\/jats:p>","DOI":"10.1145\/3634918","type":"journal-article","created":{"date-parts":[[2023,11,29]],"date-time":"2023-11-29T11:59:41Z","timestamp":1701259181000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":17,"title":["Graph Pooling Inference Network for Text-based VQA"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-4215-5464","authenticated-orcid":false,"given":"Sheng","family":"Zhou","sequence":"first","affiliation":[{"name":"HeFei University of Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2594-254X","authenticated-orcid":false,"given":"Dan","family":"Guo","sequence":"additional","affiliation":[{"name":"HeFei University of Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0201-1638","authenticated-orcid":false,"given":"Xun","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5244-3274","authenticated-orcid":false,"given":"Jianfeng","family":"Dong","sequence":"additional","affiliation":[{"name":"Zhejiang Gongshang University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3094-7735","authenticated-orcid":false,"given":"Meng","family":"Wang","sequence":"additional","affiliation":[{"name":"HeFei University of Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,1,11]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","first-page":"2552","DOI":"10.1109\/TPAMI.2014.2339814","article-title":"Word spotting and recognition with embedded attributes","author":"Almaz\u00e1n Jon","year":"2014","unstructured":"Jon Almaz\u00e1n, Albert Gordo, Alicia Forn\u00e9s, and Ernest Valveny. 2014. Word spotting and recognition with embedded attributes. IEEE Trans. Pattern Anal. Mach. Intell. 36, 12 (2014), 2552\u20132566.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_3_2","first-page":"16548","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201922)","author":"Biten Ali Furkan","year":"2022","unstructured":"Ali Furkan Biten, Ron Litman, Yusheng Xie, Srikar Appalaraju, and R. Manmatha. 2022. Latr: Layout-aware transformer for scene-text vqa. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201922). 16548\u201316558."},{"key":"e_1_3_2_4_2","first-page":"4290","volume-title":"Proceedings of the International Conference on Computer Vision (ICCV\u201919)","author":"Biten Ali Furkan","year":"2019","unstructured":"Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marcal Rusinol, C. V. Jawahar, Ernest Valveny, and Dimosthenis Karatzas. 2019. Scene text visual question answering. In Proceedings of the International Conference on Computer Vision (ICCV\u201919). 4290\u20134300."},{"key":"e_1_3_2_5_2","doi-asserted-by":"crossref","first-page":"135","DOI":"10.1162\/tacl_a_00051","article-title":"Enriching word vectors with subword information","author":"Bojanowski Piotr","year":"2017","unstructured":"Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017. Enriching word vectors with subword information. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL\u201917), 135\u2013146.","journal-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL\u201917)"},{"key":"e_1_3_2_6_2","first-page":"71","volume-title":"Proceedings of the ACM Knowledge Discovery and Data Mining (SIGKDD\u201918)","author":"Borisyuk Fedor","year":"2018","unstructured":"Fedor Borisyuk, Albert Gordo, and Viswanath Sivakumar. 2018. Rosetta: Large scale system for text detection and recognition in images. In Proceedings of the ACM Knowledge Discovery and Data Mining (SIGKDD\u201918). 71\u201379."},{"key":"e_1_3_2_7_2","first-page":"357","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (CVPR\u201921)","author":"Chen Chun-Fu Richard","year":"2021","unstructured":"Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda. 2021. Crossvit: Cross-attention multi-scale vision transformer for image classification. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (CVPR\u201921). 357\u2013366."},{"key":"e_1_3_2_8_2","first-page":"3438","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201920)","author":"Chen Deli","year":"2020","unstructured":"Deli Chen, Yankai Lin, Wei Li, Peng Li, Jie Zhou, and Xu Sun. 2020. Measuring and relieving the over-smoothing problem for graph neural networks from the topological view. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201920). 3438\u20133445."},{"key":"e_1_3_2_9_2","first-page":"1554","volume-title":"Proceedings of the International Conference on Computer Vision (ICCV\u201921)","author":"Dancette Corentin","year":"2021","unstructured":"Corentin Dancette, R\u00e9mi Cad\u00e8ne, Damien Teney, and Matthieu Cord. 2021. Beyond question-based biases: Assessing multimodal shortcut learning in visual question answering. In Proceedings of the International Conference on Computer Vision (ICCV\u201921). 1554\u20131563."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_11_2","first-page":"4171","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL\u201919)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, MingWei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL\u201919). 4171\u20134186."},{"key":"e_1_3_2_12_2","first-page":"5079","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201922)","author":"Ding Yang","year":"2022","unstructured":"Yang Ding, Jing Yu, Bangchang Liu, Yue Hu, Mingxin Cui, and Qi Wu. 2022. MuKEA: Multimodal knowledge extraction and accumulation for knowledge-based visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201922). 5079\u20135088."},{"key":"e_1_3_2_13_2","first-page":"4065","article-title":"Dual encoding for video retrieval by text","author":"Dong Jianfeng","year":"2021","unstructured":"Jianfeng Dong, Xirong Li, Chaoxi Xu, Xun Yang, Gang Yang, Xun Wang, and Meng Wang. 2021. Dual encoding for video retrieval by text. IEEE Trans. Pattern Anal. Mach. Intell. (2021), 4065\u20134080.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_14_2","article-title":"Separate and locate: Rethink the text in text-based visual question answering","author":"Fang Chengyang","year":"2023","unstructured":"Chengyang Fang, Jiangnan Li, Liang Li, Can Ma, and Dayong Hu. 2023. Separate and locate: Rethink the text in text-based visual question answering. In Proceedings of the ACM International Conference on Multimedia (ACM MM\u201923). 4378\u20134388.","journal-title":"Proceedings of the ACM International Conference on Multimedia (ACM MM\u201923)"},{"key":"e_1_3_2_15_2","article-title":"Structured multimodal attentions for textvqa","author":"Gao Chenyu","year":"2021","unstructured":"Chenyu Gao, Qi Zhu, Peng Wang, Hui Li, Yuliang Liu, Anton Van den Hengel, and Qi Wu. 2021. Structured multimodal attentions for textvqa. IEEE Trans. Pattern Anal. Mach. Intell. 44, 12 (2021), 9603\u20139614.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_16_2","first-page":"12746","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Gao Difei","year":"2020","unstructured":"Difei Gao, Ke Li, Ruiping Wang, Shiguang Shan, and Xilin Chen. 2020. Multi-modal graph neural network for joint reasoning on vision and scene text. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201920). 12746\u201312756."},{"key":"e_1_3_2_17_2","first-page":"5057","article-title":"Transform-retrieve-generate: Natural language-centric outside-knowledge visual question answering","author":"Gao Feng","year":"2022","unstructured":"Feng Gao, Q. Ping, Govind Thattai, Aishwarya N. Reganti, Yingting Wu, and Premkumar Natarajan. 2022. Transform-retrieve-generate: Natural language-centric outside-knowledge visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201922), 5057\u20135067.","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201922)"},{"key":"e_1_3_2_18_2","first-page":"6056","article-title":"Context-aware graph inference with knowledge distillation for visual dialog","author":"Guo Dan","year":"2021","unstructured":"Dan Guo, Hui Wang, and Meng Wang. 2021. Context-aware graph inference with knowledge distillation for visual dialog. IEEE Trans. Pattern Anal. Mach. Intell. 44, 10 (2021), 6056\u20136073.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_19_2","first-page":"10055","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Guo Dan","year":"2020","unstructured":"Dan Guo, Hui Wang, Hanwang Zhang, Zheng-Jun Zha, and Meng Wang. 2020. Iterative context-aware graph inference for visual dialog. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201920). 10055\u201310064."},{"key":"e_1_3_2_20_2","first-page":"3608","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918)","author":"Gurari Danna","year":"2018","unstructured":"Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham. 2018. Vizwiz grand challenge: Answering visual questions from blind people. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918). 3608\u20133617."},{"key":"e_1_3_2_21_2","first-page":"3118","volume-title":"Proceedings of the International Conference on Computational Linguistics (COLING\u201920)","author":"Han Wei","year":"2020","unstructured":"Wei Han, Hantao Huang, and Tao Han. 2020. Finding the evidence: Localization-aware answer prediction for text visual question answering. In Proceedings of the International Conference on Computational Linguistics (COLING\u201920). 3118\u20133131."},{"key":"e_1_3_2_22_2","first-page":"1584","volume-title":"Proceedings of the International Conference on Computer Vision (ICCV\u201921)","author":"Han Xinzhe","year":"2021","unstructured":"Xinzhe Han, Shuhui Wang, Chi Su, Qingming Huang, and Qi Tian. 2021. Greedy gradient ensemble for robust visual question answering. In Proceedings of the International Conference on Computer Vision (ICCV\u201921). 1584\u20131593."},{"key":"e_1_3_2_23_2","first-page":"5579","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Hegde Shamanthak","year":"2023","unstructured":"Shamanthak Hegde, Soumya Jahagirdar, and Shankar Gangisetty. 2023. Making the V in text-VQA matter. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 5579\u20135587."},{"key":"e_1_3_2_24_2","first-page":"373","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL\u201922)","author":"Heo Yu-Jung","year":"2022","unstructured":"Yu-Jung Heo, Eun-Sol Kim, Woo Suk Choi, and Byoung-Tak Zhang. 2022. Hypergraph transformer: Weakly-supervised multi-hop reasoning for knowledge-based visual question answering. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL\u201922). 373\u2013390."},{"key":"e_1_3_2_25_2","volume-title":"Proceedings of the ACM International Conference on Multimedia (ACM MM\u201920)","author":"Hu Jun","year":"2020","unstructured":"Jun Hu, Quan Fang, Shengsheng Qian, and Changsheng Xu. 2020. Multi-modal attentive graph pooling model for community question answer matching. In Proceedings of the ACM International Conference on Multimedia (ACM MM\u201920). 3505\u20133513."},{"key":"e_1_3_2_26_2","first-page":"9992","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Hu Ronghang","year":"2020","unstructured":"Ronghang Hu, Amanpreet Singh, Trevor Darrell, and Marcus Rohrbach. 2020. Iterative answer prediction with pointer-augmented multimodal transformers for textvqa. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201920). 9992\u201310002."},{"key":"e_1_3_2_27_2","first-page":"715","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201920)","author":"Kant Yash","year":"2020","unstructured":"Yash Kant, Dhruv Batra, Peter Anderson, Alexander Schwing, Devi Parikh, Jiasen Lu, and Harsh Agrawal. 2020. Spatially aware multimodal transformers for textvqa. In Proceedings of the European Conference on Computer Vision (ECCV\u201920). 715\u2013732."},{"key":"e_1_3_2_28_2","first-page":"1156","volume-title":"Proceedings of the International Conference on Document Analysis and Recognition (ICDAR\u201915)","author":"Karatzas Dimosthenis","year":"2015","unstructured":"Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, et\u00a0al. 2015. ICDAR 2015 competition on robust reading. In Proceedings of the International Conference on Document Analysis and Recognition (ICDAR\u201915). 1156\u20131160."},{"key":"e_1_3_2_29_2","first-page":"1484","volume-title":"Proceedings of the International Conference on Document Analysis and Recognition (ICDAR\u201913)","author":"Karatzas Dimosthenis","year":"2013","unstructured":"Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernandez Mota, Jon Almazan Almazan, and Lluis Pere De Las Heras. 2013. ICDAR 2013 robust reading competition. In Proceedings of the International Conference on Document Analysis and Recognition (ICDAR\u201913). 1484\u20131493."},{"key":"e_1_3_2_30_2","unstructured":"Ivan Krasin Tom Duerig Neil Alldrin Vittorio Ferrari Sami AbuElHaija Alina Kuznetsova Hassan Rom Jasper Uijlings Stefan Popov Andreas Veit et\u00a0al. 2017. Openimages: A Public Dataset for Large-scale Multi-label and Multi-class Image Classification. Retrieved from https:\/\/github.com\/openimages"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1007\/s11263-016-0981-7","article-title":"Visual genome: Connecting language and vision using crowdsourced dense image annotations","author":"Krishna Ranjay","year":"2017","unstructured":"Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, et\u00a0al. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vis. 123 (2017), 32\u201373.","journal-title":"Int. J. Comput. Vis."},{"key":"e_1_3_2_32_2","first-page":"707","volume-title":"Soviet Physics Doklady","author":"Levenshtein Vladimir I.","year":"1966","unstructured":"Vladimir I. Levenshtein et\u00a0al. 1966. Binary codes capable of correcting deletions, insertions, and reversals. In Soviet Physics Doklady. Vol. 10. Soviet Union, 707\u2013710."},{"key":"e_1_3_2_33_2","first-page":"4143","volume-title":"Proceedings of the Asian Conference on Computer Vision (ACCV\u201922)","author":"Li Bingjia","year":"2022","unstructured":"Bingjia Li, Jie Wang, Minyi Zhao, and Shuigeng Zhou. 2022. Two-stage multimodality fusion for high-performance text-based visual question answering. In Proceedings of the Asian Conference on Computer Vision (ACCV\u201922). 4143\u20134159."},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","first-page":"3367","DOI":"10.1109\/TIP.2023.3276570","article-title":"Weakly-supervised 3D spatial reasoning for text-based visual question answering","author":"Li Hao","year":"2023","unstructured":"Hao Li, Jinfa Huang, Peng Jin, Guoli Song, Qi Wu, and Jie Chen. 2023. Weakly-supervised 3D spatial reasoning for text-based visual question answering. IEEE Trans. Image Process. 32 (2023), 3367\u20133382.","journal-title":"IEEE Trans. Image Process."},{"key":"e_1_3_2_35_2","first-page":"3022","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201922)","author":"Li Juncheng","year":"2022","unstructured":"Juncheng Li, Junlin Xie, Long Qian, Linchao Zhu, Siliang Tang, Fei Wu, Yi Yang, Yueting Zhuang, and Xin Eric Wang. 2022. Compositional temporal grounding with structured variational cross-graph correspondence learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201922). 3022\u20133031."},{"key":"e_1_3_2_36_2","first-page":"108455","article-title":"Text-instance graph: Exploring the relational semantics for text-based visual question answering","author":"Li Xiangpeng","year":"2022","unstructured":"Xiangpeng Li, Bo Wu, Jingkuan Song, Lianli Gao, Pengpeng Zeng, and Chuang Gan. 2022. Text-instance graph: Exploring the relational semantics for text-based visual question answering. Pattern Recogn. 124 (2022), 108455.","journal-title":"Pattern Recogn."},{"key":"e_1_3_2_37_2","article-title":"Redundancy-aware transformer for video question answering","author":"Li Yicong","year":"2023","unstructured":"Yicong Li, Xun Yang, An Zhang, Chun Feng, Xiang Wang, and Tat-Seng Chua. 2023. Redundancy-aware transformer for video question answering. In Proceedings of the ACM International Conference on Multimedia (ACM MM\u201923). 3172\u20133180.","journal-title":"Proceedings of the ACM International Conference on Multimedia (ACM MM\u201923)"},{"key":"e_1_3_2_38_2","first-page":"4060","volume-title":"Proceedings of the ACM International Conference on Multimedia (ACM MM\u201920)","author":"Liu Fen","year":"2020","unstructured":"Fen Liu, Guanghui Xu, Qi Wu, Qing Du, Wei Jia, and Mingkui Tan. 2020. Cascade reasoning network for text-based visual question answering. In Proceedings of the ACM International Conference on Multimedia (ACM MM\u201920). 4060\u20134069."},{"key":"e_1_3_2_39_2","first-page":"3052","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI\u201919)","author":"Liu Yuliang","year":"2019","unstructured":"Yuliang Liu, Sheng Zhang, Lianwen Jin, Lele Xie, Y. Wu, and Zhepeng Wang. 2019. Omnidirectional scene text detection with sequential-free box discretization. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI\u201919). 3052\u20133058."},{"key":"e_1_3_2_40_2","first-page":"1039","article-title":"Towards end-to-end unified scene text detection and layout analysis","author":"Long Shangbang","year":"2022","unstructured":"Shangbang Long, Siyang Qin, Dmitry Panteleev, A. Bissacco, Yasuhisa Fujii, and Michalis Raptis. 2022. Towards end-to-end unified scene text detection and layout analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201922), 1039\u20131049.","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201922)"},{"key":"e_1_3_2_41_2","first-page":"2631","volume-title":"Proceedings of the International Conference on Computer Vision (ICCV\u201921) Workshop","author":"Lu XiaoPeng","year":"2021","unstructured":"XiaoPeng Lu, Zhenhua Fan, Yansen Wang, Jean Oh, and Carolyn Penstein Ros\u00e9. 2021. Localize, group, and select: Boosting text-VQA by scene text modeling. In Proceedings of the International Conference on Computer Vision (ICCV\u201921) Workshop. 2631\u20132639."},{"key":"e_1_3_2_42_2","first-page":"3040","volume-title":"Proceedings of the International Conference on Computer Vision (ICCV\u201913)","author":"Mishra Anand","year":"2013","unstructured":"Anand Mishra, Karteek Alahari, and C. V. Jawahar. 2013. Image retrieval using textual cues. In Proceedings of the International Conference on Computer Vision (ICCV\u201913). 3040\u20133047."},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58536-5_44"},{"key":"e_1_3_2_45_2","first-page":"8317","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201919)","author":"Singh Amanpreet","year":"2019","unstructured":"Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. 2019. Towards vqa models that can read. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201919). 8317\u20138326."},{"key":"e_1_3_2_46_2","article-title":"Emotion-prior awareness network for emotional video captioning","author":"Song Peipei","year":"2023","unstructured":"Peipei Song, Dan Guo, Xun Yang, Shengeng Tang, Erkun Yang, and Meng Wang. 2023. Emotion-prior awareness network for emotional video captioning. Proceedings of the ACM International Conference on Multimedia (ACM MM\u201923). 589\u2013600.","journal-title":"Proceedings of the ACM International Conference on Multimedia (ACM MM\u201923)"},{"key":"e_1_3_2_47_2","first-page":"459","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201922)","author":"Tu Zhengzhong","year":"2022","unstructured":"Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan Bovik, and Yinxiao Li. 2022. Maxvit: Multi-axis vision transformer. In Proceedings of the European Conference on Computer Vision (ECCV\u201922). 459\u2013479."},{"key":"e_1_3_2_48_2","article-title":"Coco-text: Dataset and benchmark for text detection and recognition in natural images","author":"Veit Andreas","year":"2016","unstructured":"Andreas Veit, Tomas Matera, Lukas Neumann, Jiri Matas, and Serge Belongie. 2016. Coco-text: Dataset and benchmark for text detection and recognition in natural images. arXiv:1601.07140. Retrieved from https:\/\/arxiv.org\/abs\/1601.07140","journal-title":"arXiv:1601.07140"},{"key":"e_1_3_2_49_2","volume-title":"Proceedings of the British Machine Vision Conference (BMVC\u201922)","author":"Wang Jun","year":"2022","unstructured":"Jun Wang, Mingfei Gao, Yuqian Hu, Ramprasaath R. Selvaraju, Chetan Ramaiah, Ran Xu, Joseph J\u00e1J\u00e1, and Larry Davis. 2022. TAG: Boosting text-VQA via text-aware visual question-answer generation. In Proceedings of the British Machine Vision Conference (BMVC\u201922)."},{"key":"e_1_3_2_50_2","first-page":"4337","volume-title":"Proceedings of the ACM International Conference on Multimedia (ACM MM\u201920)","author":"Wang Jing","year":"2020","unstructured":"Jing Wang, Jinhui Tang, and Jiebo Luo. 2020. Multimodal attention with image text spatial relationship for ocr-based image captioning. In Proceedings of the ACM International Conference on Multimedia (ACM MM\u201920). 4337\u20134345."},{"key":"e_1_3_2_51_2","article-title":"Vqa-gnn: Reasoning with multimodal semantic graph for visual question answering","author":"Wang Yanan","year":"2022","unstructured":"Yanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada, and Jure Leskovec. 2022. Vqa-gnn: Reasoning with multimodal semantic graph for visual question answering. arXiv:2205.11501. Retrieved from https:\/\/arxiv.org\/abs\/2205.11501","journal-title":"arXiv:2205.11501"},{"key":"e_1_3_2_52_2","first-page":"24017","volume-title":"Proceedings of the International Conference on Machine Learning (ICML\u201922)","author":"Wu Junran","year":"2022","unstructured":"Junran Wu, Xu hui Chen, Ke Xu, and Shangzhe Li. 2022. Structural entropy guided graph hierarchical pooling. In Proceedings of the International Conference on Machine Learning (ICML\u201922). 24017\u201324030."},{"key":"e_1_3_2_53_2","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201921) Workshop","author":"Yang Michael","year":"2021","unstructured":"Michael Yang, Aditya Anantharaman, Zachary Kitowski, and Derik Clive Robert. 2021. Graph relation transformer: Incorporating pairwise object features into the transformer architecture. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201921) Workshop."},{"key":"e_1_3_2_54_2","first-page":"1339","volume-title":"Proceedings of the ACM SIGIR Conference on Research and Development in Information Retrieval (ACM SIGIR\u201920)","author":"Yang Xun","year":"2020","unstructured":"Xun Yang, Jianfeng Dong, Yixin Cao, Xun Wang, Meng Wang, and Tat-Seng Chua. 2020. Tree-augmented cross-modal encoding for complex-query video retrieval. In Proceedings of the ACM SIGIR Conference on Research and Development in Information Retrieval (ACM SIGIR\u201920). 1339\u20131348."},{"key":"e_1_3_2_55_2","first-page":"1","volume-title":"Proceedings of the ACM SIGIR Conference on Research and Development in Information Retrieval (ACM SIGIR\u201921)","author":"Yang Xun","year":"2021","unstructured":"Xun Yang, Fuli Feng, Wei Ji, Meng Wang, and Tat-Seng Chua. 2021. Deconfounded video moment retrieval with causal intervention. In Proceedings of the ACM SIGIR Conference on Research and Development in Information Retrieval (ACM SIGIR\u201921). 1\u201310."},{"key":"e_1_3_2_56_2","doi-asserted-by":"crossref","first-page":"1204","DOI":"10.1109\/TIP.2022.3140611","article-title":"Video moment retrieval with cross-modal neural architecture search","author":"Yang Xun","year":"2022","unstructured":"Xun Yang, Shanshan Wang, Jian Dong, Jianfeng Dong, Meng Wang, and Tat-Seng Chua. 2022. Video moment retrieval with cross-modal neural architecture search. IEEE Trans. Image Process. 31 (2022), 1204\u20131216.","journal-title":"IEEE Trans. Image Process."},{"key":"e_1_3_2_57_2","first-page":"8751","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201921)","author":"Yang Zhengyuan","year":"2021","unstructured":"Zhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin, Dinei Florencio, Lijuan Wang, Cha Zhang, Lei Zhang, and Jiebo Luo. 2021. TAP: Text-aware pre-training for text-VQA and text-caption. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201921). 8751\u20138761."},{"key":"e_1_3_2_58_2","first-page":"376","volume-title":"Proceedings of the ACM International Conference on Multimedia (ACM MM\u201921)","author":"Zeng Gangyan","year":"2021","unstructured":"Gangyan Zeng, Yuan Zhang, Yu Zhou, and Xiaomeng Yang. 2021. Beyond OCR+ VQA: Involving OCR into the flow for robust and accurate TextVQA. In Proceedings of the ACM International Conference on Multimedia (ACM MM\u201921). 376\u2013385."},{"key":"e_1_3_2_59_2","volume-title":"Proceedings of the International World Wide Web Conferences (WWW\u201920)","author":"Zhang Liang","year":"2020","unstructured":"Liang Zhang, Xudong Wang, Hongsheng Li, Guangming Zhu, Peiyi Shen, P. Li, Xiaoyuan Lu, Syed Afaq Ali Shah, and Bennamoun. 2020. Structure-feature based graph self-adaptive pooling. In Proceedings of the International World Wide Web Conferences (WWW\u201920)."},{"key":"e_1_3_2_60_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201922)","author":"Zhang Wenqiao","year":"2022","unstructured":"Wenqiao Zhang, Haochen Shi, Jiannan Guo, Shengyu Zhang, Qingpeng Cai, Juncheng Li, Sihui Luo, and Yueting Zhuang. 2022. MAGIC: Multimodal relAtional graph adversarIal inferenCe for diverse and unpaired text-based image captioning. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201922). Vol. 36. 3335\u20133343."},{"key":"e_1_3_2_61_2","first-page":"2519","volume-title":"Proceedings of the ACM International Conference on Multimedia (ACM MM\u201921)","author":"Zhang Xuanyu","year":"2021","unstructured":"Xuanyu Zhang and Qing Yang. 2021. Position-augmented transformers with entity-aligned mesh for TextVQA. In Proceedings of the ACM International Conference on Multimedia (ACM MM\u201921). 2519\u20132528."},{"key":"e_1_3_2_62_2","article-title":"Progressive localization networks for language-based moment localization","author":"Zheng Qi","year":"2023","unstructured":"Qi Zheng, Jianfeng Dong, Xiaoye Qu, Xun Yang, Yabing Wang, Pan Zhou, Baolong Liu, and Xun Wang. 2023. Progressive localization networks for language-based moment localization. ACM Trans. Multimedia Comput. Commun. Appl. 19, 2 (2023), 1\u201321.","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_3_2_63_2","article-title":"Exploring sparse spatial relation in graph inference for text-based VQA","author":"Zhou Sheng","year":"2023","unstructured":"Sheng Zhou, Dan Guo, Jia Li, Xun Yang, and Meng Wang. 2023. Exploring sparse spatial relation in graph inference for text-based VQA. IEEE Trans. Image Process. 32 (2023), 5060\u20135074.","journal-title":"IEEE Trans. Image Process."},{"key":"e_1_3_2_64_2","first-page":"3608","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201921)","author":"Zhu Qi","year":"2021","unstructured":"Qi Zhu, Chenyu Gao, P. Wang, and Qi Wu. 2021. Simple is not easy: A simple strong baseline for TextVQA and TextCaps. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201921). 3608\u20133615."},{"key":"e_1_3_2_65_2","first-page":"11479","article-title":"Locate then generate: Bridging vision and language with bounding box for scene-text VQA","author":"Zhu Yongxin","year":"2023","unstructured":"Yongxin Zhu, Zhen Liu, Yukang Liang, Xin Li, Hao Liu, Changcun Bao, and Linli Xu. 2023. Locate then generate: Bridging vision and language with bounding box for scene-text VQA. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201923), 11479\u201311487.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI\u201923)"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3634918","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3634918","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:51:07Z","timestamp":1750287067000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3634918"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,11]]},"references-count":64,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,4,30]]}},"alternative-id":["10.1145\/3634918"],"URL":"https:\/\/doi.org\/10.1145\/3634918","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,11]]},"assertion":[{"value":"2023-07-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-21","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-11","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}