{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,26]],"date-time":"2026-07-26T03:41:31Z","timestamp":1785037291318,"version":"3.55.0"},"publisher-location":"California","reference-count":0,"publisher":"International Joint Conferences on Artificial Intelligence Organization","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,9]]},"abstract":"<jats:p>Medical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due to inadequate localization reasoning. This work reveals a critical limitation in current medical LMMs: instead of analyzing relevant pathological regions, they often rely on linguistic patterns or attend to irrelevant image areas when responding to disease-related queries.\n\nTo address this, we introduce HEAL-MedVQA (Hallucination Evaluation via Localization MedVQA), a comprehensive benchmark designed to evaluate LMMs' localization abilities and hallucination robustness. HEAL-MedVQA features (i) two innovative evaluation protocols to assess visual and textual shortcut learning, and (ii) a dataset of 67K VQA pairs, with doctor-annotated anatomical segmentation masks for pathological regions. To improve visual reasoning, we propose the Localize-before-Answer (LobA) framework, which trains LMMs to localize target regions of interest and self-prompt to emphasize segmented pathological areas, generating grounded and reliable answers. Experimental results demonstrate that our approach significantly outperforms state-of-the-art biomedical LMMs on the challenging HEAL-MedVQA benchmark, advancing robustness in medical VQA.<\/jats:p>","DOI":"10.24963\/ijcai.2025\/853","type":"proceedings-article","created":{"date-parts":[[2025,9,19]],"date-time":"2025-09-19T08:10:40Z","timestamp":1758269440000},"page":"7670-7678","source":"Crossref","is-referenced-by-count":3,"title":["Localizing Before Answering: A Benchmark for Grounded Medical Visual Question Answering"],"prefix":"10.24963","author":[{"given":"Dung","family":"Nguyen","sequence":"first","affiliation":[{"name":"Hanoi University of Science and Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Minh Khoi","family":"Ho","sequence":"additional","affiliation":[{"name":"Hanoi University of Science and Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huy","family":"Ta","sequence":"additional","affiliation":[{"name":"Australian Institute for Machine Learning, The University of Adelaide"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thanh Tam","family":"Nguyen","sequence":"additional","affiliation":[{"name":"Griffith University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qi","family":"Chen","sequence":"additional","affiliation":[{"name":"Australian Institute for Machine Learning, The University of Adelaide"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kumar","family":"Rav","sequence":"additional","affiliation":[{"name":"SA Health"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Quy Duong","family":"Dang","sequence":"additional","affiliation":[{"name":"University of Adelaide"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Satwik","family":"Ramchandre","sequence":"additional","affiliation":[{"name":"University of Adelaide"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Son Lam","family":"Phung","sequence":"additional","affiliation":[{"name":"University of Wollongong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhibin","family":"Liao","sequence":"additional","affiliation":[{"name":"Australian Institute for Machine Learning, The University of Adelaide"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Minh-Son","family":"To","sequence":"additional","affiliation":[{"name":"Flinders University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Johan","family":"Verjans","sequence":"additional","affiliation":[{"name":"University of Adelaide"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Phi Le","family":"Nguyen","sequence":"additional","affiliation":[{"name":"Hanoi University of Science and Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Vu Minh Hieu","family":"Phan","sequence":"additional","affiliation":[{"name":"Australian Institute for Machine Learning, The University of Adelaide"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"10584","event":{"name":"Thirty-Fourth International Joint Conference on Artificial Intelligence {IJCAI-25}","theme":"Artificial Intelligence","location":"Montreal, Canada","acronym":"IJCAI-2025","number":"34","sponsor":["International Joint Conferences on Artificial Intelligence Organization (IJCAI)"],"start":{"date-parts":[[2025,8,16]]},"end":{"date-parts":[[2025,8,22]]}},"container-title":["Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence"],"original-title":[],"deposited":{"date-parts":[[2025,9,23]],"date-time":"2025-09-23T11:35:18Z","timestamp":1758627318000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.ijcai.org\/proceedings\/2025\/853"}},"subtitle":[],"proceedings-subject":"Artificial Intelligence Research Articles","short-title":[],"issued":{"date-parts":[[2025,9]]},"references-count":0,"URL":"https:\/\/doi.org\/10.24963\/ijcai.2025\/853","relation":{},"subject":[],"published":{"date-parts":[[2025,9]]}}}