{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T04:58:34Z","timestamp":1781326714351,"version":"3.54.1"},"reference-count":64,"publisher":"Association for Computing Machinery (ACM)","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2025,12,4]]},"abstract":"<jats:p>In the big data era, Table Question Answering (Table QA) has emerged as a crucial tool for extracting insights from structured data, especially in large table scenarios. There are two main categories of methods for Table QA: Executable Code-driven methods and Language Model based (LM-based) methods. Code-driven methods, e.g. Text-to-SQL based solutions, often struggle with incomplete or mismatching schema information. LM-based methods, include Pre-trained Language Models (PLMs) and Large Language Models (LLMs), also face challenges as PLMs have limited generalization, while LLMs suffer from performance degradation and increased token cost when applied to large tables.<\/jats:p>\n                  <jats:p>To address these challenges, we propose AixelAsk, a novel LLM-based framework designed for Large Table QA. Specifically, AixelAsk incorporates a three-module architecture consisting of Decomposition module, Retrieval module and Reasoning module. The Decomposition module constructs a directed acyclic graph (DAG)-based solution plan by decomposing the question into execution nodes with explicit dependencies, making a clear reasoning path to guide the LLM through a logical process. Inspired by the Retrieval-Augmented Generation, the Retrieval Module extracts key rows and columns from the large table, reducing input token size and focusing on critical information. The Reasoning Module performs step-by-step inferences over the retrieved sub-tables, guided by each execution node in the solution plan, to generate final answer. By tackling the challenges of LLM performance degradation with large inputs and complex questions, AixelAsk achieves superior performance in Large Table QA. Extensive experiments on various baselines across three datasets demonstrate the effectiveness and efficiency of our proposed AixelAsk framework. AixelAsk outperforms the state-of-the-art baseline by 4% - 8% in the exact match score, and at the same time reduces token usage by 86.4%, achieving both high accuracy and cost efficiency in the Large Table QA task.<\/jats:p>","DOI":"10.1145\/3769831","type":"journal-article","created":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T04:32:13Z","timestamp":1764995533000},"page":"1-25","source":"Crossref","is-referenced-by-count":0,"title":["AixelAsk: A Stepwise-Guided Retrieval and Reasoning Framework for Large Table QA"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-0323-0715","authenticated-orcid":false,"given":"Chi","family":"Zhang","sequence":"first","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0752-9877","authenticated-orcid":false,"given":"Meihui","family":"Zhang","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-6694-7790","authenticated-orcid":false,"given":"Yuxin","family":"Yang","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-8324-956X","authenticated-orcid":false,"given":"Tao","family":"Chen","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7271-3999","authenticated-orcid":false,"given":"Zhaojing","family":"Luo","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,12,5]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3035930"},{"key":"e_1_2_1_2_1","volume-title":"Text2SQL is Not Enough: Unifying AI and Databases with TAG. arXiv preprint arXiv:2408.14717","author":"Biswal Asim","year":"2024","unstructured":"Asim Biswal, Liana Patel, Siddarth Jha, Amog Kamsetty, Shu Liu, Joseph E Gonzalez, Carlos Guestrin, and Matei Zaharia. 2024. Text2SQL is Not Enough: Unifying AI and Databases with TAG. arXiv preprint arXiv:2408.14717 (2024)."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3663742.3663972"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the International Conference on Machine Learning. PMLR, 2206-2240","author":"Borgeaud Sebastian","year":"2022","unstructured":"Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al., 2022. Improving language models by retrieving from trillions of tokens. In Proceedings of the International Conference on Machine Learning. PMLR, 2206-2240."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415563"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.897"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.342"},{"key":"e_1_2_1_8_1","first-page":"74899","volume-title":"Zhang (Eds.)","volume":"37","author":"Chen Si-An","year":"2024","unstructured":"Si-An Chen, Lesly Miculicich, Julian Martin Eisenschlos, Zifeng Wang, Zilong Wang, Yanfei Chen, Yasuhisa Fujii, Hsuan-Tien Lin, Chen-Yu Lee, and Tomas Pfister. 2024. TableRAG: Million-Token Table Understanding with Language Models. In Advances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.), Vol. 37. Curran Associates, Inc., 74899-74921."},{"key":"e_1_2_1_9_1","first-page":"1120","article-title":"Large Language Models are few-shot Table Reasoners","author":"Chen Wenhu","year":"2023","unstructured":"Wenhu Chen. 2023. Large Language Models are few-shot Table Reasoners. In Findings of the European Chapter of the Association for Computational Linguistics. 1120-1130.","journal-title":"Findings of the European Chapter of the Association for Computational Linguistics."},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Chen Wenhu","year":"2020","unstructured":"Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020a. TabFact: A Large-scale Dataset for Table-based Fact Verification. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Chen Wenhu","year":"2020","unstructured":"Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020b. Tabfact: A large-scale dataset for table-based fact verification. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the International Conference on Learning Representations","volume":"2210","author":"Cheng Zhoujun","year":"2023","unstructured":"Zhoujun Cheng, Tianbao Xie, Peng Shi, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir Radev, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, and Tao Yu. 2023. Binding Language Models in Symbolic Languages. Proceedings of the International Conference on Learning Representations, Vol. abs\/2210.02875 (2023)."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3637528.3672041"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3589292"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.398"},{"key":"e_1_2_1_16_1","volume-title":"NeurIPS 2023 Second Table Representation Learning Workshop.","author":"Huang Zezhou","year":"2023","unstructured":"Zezhou Huang, Pavan Kalyan Damalapati, and Eugene Wu. 2023. Data Ambiguity Strikes Back: How Documentation Improves GPT's Text-to-SQL. In NeurIPS 2023 Second Table Representation Learning Workshop."},{"key":"e_1_2_1_17_1","volume-title":"K S M Tozammel Hossain, Enamul Hoque, Shafiq Joty, and Md Rizwan Parvez.","author":"Islam Shayekh Bin","year":"2024","unstructured":"Shayekh Bin Islam, Md Asib Rahman, K S M Tozammel Hossain, Enamul Hoque, Shafiq Joty, and Md Rizwan Parvez. 2024. Open-RAG: Enhanced Retrieval Augmented Reasoning with Open-Source Large Language Models. In Findings of the Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 14231-14244."},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Izacard Gautier","year":"2021","unstructured":"Gautier Izacard and Edouard Grave. 2021. Distilling Knowledge from Reader to Retriever for Question Answering. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICAAIC60222.2024.10574972"},{"key":"e_1_2_1_20_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3314043"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654930"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3517836"},{"key":"e_1_2_1_24_1","volume-title":"Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74-81.","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74-81."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-short.133"},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the Conference, on Innovative Database Research (CIDR, )","author":"Liu Chunwei","year":"2025","unstructured":"Chunwei Liu, Matthew Russo, Michael Cafarella, Lei Cao, Peter Baile Chen, Zui Chen, Michael Franklin, Tim Kraska, Samuel Madden, Rana Shahout, and Gerardo Vitagliano. 2025. Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing. In Proceedings of the Conference, on Innovative Database Research (CIDR, ) (2025)."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00638"},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Liu Qian","year":"2022","unstructured":"Qian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, and Jian-Guang Lou. 2022. TAPEX: Table Pre-training via Learning a Neural SQL Executor. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_2_1_29_1","volume-title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR, Vol. abs\/1907.11692 (2019). arXiv:1907.11692"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2019.2916683"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i10.17067"},{"key":"e_1_2_1_32_1","volume-title":"Adaptive Lightweight Regularization Tool for Complex Analytics. In 2018 IEEE 34th International Conference on Data Engineering (ICDE). 485-496","author":"Luo Zhaojing","year":"2018","unstructured":"Zhaojing Luo, Shaofeng Cai, Jinyang Gao, Meihui Zhang, Kee Yuan Ngiam, Gang Chen, and Wang-Chien Lee. 2018. Adaptive Lightweight Regularization Tool for Complex Analytics. In 2018 IEEE 34th International Conference on Data Engineering (ICDE). 485-496."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3588936"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE51399.2021.00146"},{"key":"e_1_2_1_35_1","unstructured":"Ben Mann N Ryder M Subbiah J Kaplan P Dhariwal A Neelakantan P Shyam G Sastry A Askell S Agarwal et al. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 Vol. 1 (2020)."},{"key":"e_1_2_1_36_1","volume-title":"Bioinformatics","volume":"40","author":"Matsumoto Nicholas","year":"2024","unstructured":"Nicholas Matsumoto, Jay Moran, Hyunjun Choi, Miguel E Hernandez, Mythreye Venkatesan, Paul Wang, and Jason H Moore. 2024. KRAGEN: a knowledge graph-enhanced RAG framework for biomedical problem solving using large language models. Bioinformatics, Vol. 40, 6 (2024)."},{"key":"e_1_2_1_37_1","first-page":"5725","article-title":"TabSQLify: Enhancing Reasoning Capabilities of LLMs Through Table Decomposition","author":"Hasan Nahid Md Mahadi","year":"2024","unstructured":"Md Mahadi Hasan Nahid and Davood Rafiei. 2024. TabSQLify: Enhancing Reasoning Capabilities of LLMs Through Table Decomposition. In Proceedings of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, 5725-5737.","journal-title":"Proceedings of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00446"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.348"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1142"},{"key":"e_1_2_1_41_1","volume-title":"Semantic Operators: A Declarative Model for Rich, AI-based Data Processing. arXiv preprint arXiv:2407.11418","author":"Patel Liana","year":"2024","unstructured":"Liana Patel, Siddharth Jha, Melissa Pan, Harshit Gupta, Parth Asawa, Carlos Guestrin, and Matei Zaharia. 2024a. Semantic Operators: A Declarative Model for Rich, AI-based Data Processing. arXiv preprint arXiv:2407.11418 (2024)."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.1160"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.52202\/075280-1577"},{"key":"e_1_2_1_44_1","volume-title":"Evaluating the text-to-sql capabilities of large language models. arXiv preprint arXiv:2204.00498","author":"Rajkumar Nitarshan","year":"2022","unstructured":"Nitarshan Rajkumar, Raymond Li, and Dzmitry Bahdanau. 2022. Evaluating the text-to-sql capabilities of large language models. arXiv preprint arXiv:2204.00498 (2022)."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3725411"},{"key":"e_1_2_1_46_1","volume-title":"Docetl: Agentic query rewriting and evaluation for complex document processing. arXiv preprint arXiv:2410.12189","author":"Shankar Shreya","year":"2024","unstructured":"Shreya Shankar, Tristan Chambers, Tarak Shah, Aditya G Parameswaran, and Eugene Wu. 2024. Docetl: Agentic query rewriting and evaluation for complex document processing. arXiv preprint arXiv:2410.12189 (2024)."},{"key":"e_1_2_1_47_1","first-page":"8364","article-title":"REPLUG","author":"Shi Weijia","year":"2024","unstructured":"Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2024. REPLUG: Retrieval-Augmented Black-Box Language Models. In Proceedings of the North American Chapter of the Association for Computational Linguistics. 8364-8377.","journal-title":"Retrieval-Augmented Black-Box Language Models. In Proceedings of the North American Chapter of the Association for Computational Linguistics."},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3616855.3635752"},{"key":"e_1_2_1_49_1","volume-title":"International Conference on Very Large Data Bases.","author":"Usta Arif","year":"2023","unstructured":"Arif Usta and Semih Salihoglu. 2023. To Join or Not to Join: An Analysis on the Usefulness of Joining Tables in Open Government Data Portals. In International Conference on Very Large Data Bases."},{"key":"e_1_2_1_50_1","volume-title":"Data Ambiguity Profiling for the Generation of Training Examples. In 2023 IEEE 39th International Conference on Data Engineering (ICDE). 450-463","author":"Veltri Enzo","year":"2023","unstructured":"Enzo Veltri, Gilbert Badaro, Mohammed Saeed, and Paolo Papotti. 2023. Data Ambiguity Profiling for the Generation of Training Examples. In 2023 IEEE 39th International Conference on Data Engineering (ICDE). 450-463."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2123"},{"key":"e_1_2_1_52_1","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Wang Zilong","year":"2024","unstructured":"Zilong Wang, Hao Zhang, Chun-Liang Li, Julian Martin Eisenschlos, Vincent Perot, Zifeng Wang, Lesly Miculicich, Yasuhisa Fujii, Jingbo Shang, Chen-Yu Lee, and Tomas Pfister. 2024. Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding. Proceedings of the International Conference on Learning Representations (2024)."},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2021.3069861"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591708"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.745"},{"key":"e_1_2_1_56_1","volume-title":"Applications and Challenges for Large Language Models: From Data Management Perspective. In 2024 IEEE 40th International Conference on Data Engineering (ICDE). 5530-5541","author":"Zhang Meihui","year":"2024","unstructured":"Meihui Zhang, Zhaoxuan Ji, Zhaojing Luo, Yuncheng Wu, and Chengliang Chai. 2024b. Applications and Challenges for Large Language Models: From Data Management Perspective. In 2024 IEEE 40th International Conference on Data Engineering (ICDE). 5530-5541."},{"key":"e_1_2_1_57_1","volume-title":"Aixel: A Unified, Adaptive and Extensible System for AI-powered Data Analysis. arXiv preprint arXiv:2510.12642","author":"Zhang Meihui","year":"2025","unstructured":"Meihui Zhang, Liming Wang, Chi Zhang, and Zhaojing Luo. 2025. Aixel: A Unified, Adaptive and Extensible System for AI-powered Data Analysis. arXiv preprint arXiv:2510.12642 (2025)."},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-emnlp.131"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.14778\/3659437.3659452"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.914"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.692"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.334"},{"key":"e_1_2_1_63_1","volume-title":"Seq2sql: Generating structured queries from natural language using reinforcement learning. arXiv preprint arXiv:1709.00103","author":"Zhong Victor","year":"2017","unstructured":"Victor Zhong, Caiming Xiong, and Richard Socher. 2017a. Seq2sql: Generating structured queries from natural language using reinforcement learning. arXiv preprint arXiv:1709.00103 (2017)."},{"key":"e_1_2_1_64_1","volume-title":"Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning. CoRR","author":"Zhong Victor","year":"2017","unstructured":"Victor Zhong, Caiming Xiong, and Richard Socher. 2017b. Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning. CoRR, Vol. abs\/1709.00103 (2017)."}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3769831","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T04:47:12Z","timestamp":1781326032000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3769831"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,4]]},"references-count":64,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,12,4]]}},"alternative-id":["10.1145\/3769831"],"URL":"https:\/\/doi.org\/10.1145\/3769831","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,4]]}}}