{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T15:48:30Z","timestamp":1778082510150,"version":"3.51.4"},"reference-count":0,"publisher":"Association for the Advancement of Artificial Intelligence (AAAI)","issue":"27","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["AAAI"],"abstract":"<jats:p>Table images present unique challenges for effective and efficient understanding due to the need for question-specific focus and the presence of redundant background regions.\nExisting Multimodal Large Language Model (MLLM) approaches often overlook these characteristics, resulting in uninformative and redundant visual representations.\nTo address these issues, we aim to generate visual features that are both informative and compact for improved table understanding.\nWe first propose progressive question conditioning, which injects the question into Vision Transformer layers with gradually increasing frequency, considering each layer\u2019s capacity to handle additional information, to generate question-aware visual features.\nTo reduce redundancy, we introduce a pruning strategy that discards background tokens, thereby improving efficiency.\nTo mitigate information loss from pruning, we further propose token focusing, a training strategy that encourages the model to concentrate essential information in the retained tokens.\nBy combining these approaches, we present TabFlash, an efficient and effective MLLM for table understanding.\nTabFlash achieves state-of-the-art performance, outperforming both open-source and proprietary MLLMs, while requiring 27% less FLOPs and 30% less memory usage compared to the second-best MLLM.<\/jats:p>","DOI":"10.1609\/aaai.v40i27.39417","type":"journal-article","created":{"date-parts":[[2026,3,18]],"date-time":"2026-03-18T01:32:55Z","timestamp":1773797575000},"page":"22573-22581","source":"Crossref","is-referenced-by-count":1,"title":["TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing"],"prefix":"10.1609","volume":"40","author":[{"given":"Jongha","family":"Kim","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Minseong","family":"Bae","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sanghyeok","family":"Lee","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jinsung","family":"Yoon","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hyunwoo J.","family":"Kim","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"9382","published-online":{"date-parts":[[2026,3,14]]},"container-title":["Proceedings of the AAAI Conference on Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/ojs.aaai.org\/index.php\/AAAI\/article\/download\/39417\/43378","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ojs.aaai.org\/index.php\/AAAI\/article\/download\/39417\/43378","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,18]],"date-time":"2026-03-18T01:32:55Z","timestamp":1773797575000},"score":1,"resource":{"primary":{"URL":"https:\/\/ojs.aaai.org\/index.php\/AAAI\/article\/view\/39417"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,14]]},"references-count":0,"journal-issue":{"issue":"27","published-online":{"date-parts":[[2026,3,17]]}},"URL":"https:\/\/doi.org\/10.1609\/aaai.v40i27.39417","relation":{},"ISSN":["2374-3468","2159-5399"],"issn-type":[{"value":"2374-3468","type":"electronic"},{"value":"2159-5399","type":"print"}],"subject":[],"published":{"date-parts":[[2026,3,14]]}}}