{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T23:13:34Z","timestamp":1781046814819,"version":"3.54.1"},"reference-count":38,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[2026,4,5]],"date-time":"2026-04-05T00:00:00Z","timestamp":1775347200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"name":"Youth Fund of Computer Network Information Center of Chinese Academy of Sciences","award":["NO. 25YF08"],"award-info":[{"award-number":["NO. 25YF08"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["NO. 4254090"],"award-info":[{"award-number":["NO. 4254090"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Information Visualization"],"published-print":{"date-parts":[[2026,7]]},"abstract":"<jats:p>In the era of data-intensive scientific discovery, visualization serves as a crucial cognitive tool for researchers, while captions that align with visuals are essential for accurately conveying scientific intent. However, current scientific visualization workflows face significant challenges, including high technical barriers and semantic misalignment among user intent, visual output, and textual descriptions. To address these issues, this paper proposes SciVis-AGE, a visual analytics system based on multi-agent collaboration. Its core methodologies comprise an agent-based task decomposition and operator encapsulation approach for automatic visualization generation, and a multi-agent triangular debate mechanism for semantic alignment and caption optimization. The system effectively reduces the technical burden on domain experts and, through iterative debate among Intent Guardian, Visual Verifier, and Annotation Checker agents, ensures precise alignment of generated images and captions with user intent, visual content, and highlighted features, thereby enhancing the rigor and efficiency of scientific communication.<\/jats:p>","DOI":"10.1177\/14738716261434841","type":"journal-article","created":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T06:13:20Z","timestamp":1775456000000},"page":"213-231","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":1,"title":["Agentic scientific visualization generation and caption semantic alignment"],"prefix":"10.1177","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-0154-5704","authenticated-orcid":false,"given":"Xuyi","family":"Lu","sequence":"first","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China"},{"name":"University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6436-3650","authenticated-orcid":false,"given":"Guan","family":"Li","sequence":"additional","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China"},{"name":"University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2353-6567","authenticated-orcid":false,"given":"Yu","family":"Dong","sequence":"additional","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-1707-8941","authenticated-orcid":false,"given":"Ruixiao","family":"Peng","sequence":"additional","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China"},{"name":"University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1123-9925","authenticated-orcid":false,"given":"Zhe","family":"Wang","sequence":"additional","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China"},{"name":"University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6898-8973","authenticated-orcid":false,"given":"Dong","family":"Tian","sequence":"additional","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China"},{"name":"University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8283-2278","authenticated-orcid":false,"given":"Guihua","family":"Shan","sequence":"additional","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China"},{"name":"University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2026,4,5]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","unstructured":"Liang T He Z Jiao W et al. Encouraging divergent thinking in large language models through multi-agent debate. In: Proceedings of the 2024 conference on empirical methods in natural language processing 2024 pp. 17889\u201317904. https:\/\/doi.org\/10.18653\/v1\/2024.emnlp-main.992","DOI":"10.18653\/v1\/2024.emnlp-main.992"},{"key":"e_1_3_2_3_2","unstructured":"Du Y Li S Torralba A et al. Improving factuality and reasoning in language models through multiagent debate. In: Proceedings of the 41st International conference on machine learning ICML\u201924 2024 JMLR.org. https:\/\/doi.org\/10.5555\/3692070.3692537"},{"key":"e_1_3_2_4_2","unstructured":"Chan CM Chen W Su Y et al. Chateval: towards better LLM-based evaluators through multi-agent debate. arXiv preprint arXiv:230807201 2023."},{"key":"e_1_3_2_5_2","doi-asserted-by":"crossref","unstructured":"Hu K Bakker MA Li S et al. VizML: a machine learning approach to visualization recommendation. In: Proceedings of the 2019 CHI conference on human factors in computing systems 2019 pp. 1\u201312. https:\/\/doi.org\/10.1145\/3290605.3300358","DOI":"10.1145\/3290605.3300358"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","unstructured":"Cui W Zhang X Wang Y et al. Text-to-viz: automatic generation of infographics from proportion-related natural language statements. IEEE Trans Vis Comput Graph 2020; 26(1): 906\u2013916. https:\/\/doi.org\/10.1109\/TVCG.2019.2934785","DOI":"10.1109\/TVCG.2019.2934785"},{"key":"e_1_3_2_7_2","doi-asserted-by":"crossref","unstructured":"Zhang C Dong Y Wang Y et al. AuraGenome: an LLM-powered framework for on-the-fly reusable and scalable circular genome visualizations. IEEE Comput Graph Appl 2025; 45(5): 78\u201392. https:\/\/doi.org\/10.1109\/MCG.2025.3581560","DOI":"10.1109\/MCG.2025.3581560"},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","unstructured":"Narechania A Srinivasan A Stasko J. NL4DV: a toolkit for generating analytic specifications for data visualization from natural language queries. IEEE Trans Vis Comput Graph 2021; 27(2): 369\u2013379. https:\/\/doi.org\/10.1109\/TVCG.2020.3030378","DOI":"10.1109\/TVCG.2020.3030378"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","unstructured":"Ahrens J Geveci B Law C. Paraview: an end-user tool for large data visualization. Vis Handb 2005 pp. 717\u2013731. https:\/\/doi.org\/10.1016\/b978-012387582-2\/50038-1","DOI":"10.1016\/B978-012387582-2\/50038-1"},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","unstructured":"Schroeder WJ Avila LS Hoffman W. Visualizing with VTK: a tutorial. IEEE Comput Graph Appl 2000; 20(5): 20\u201327. https:\/\/doi.org\/10.1109\/38.865875","DOI":"10.1109\/38.865875"},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"Shen L Shen E Luo Y et al. Towards natural language interfaces for data visualization: a survey. IEEE Trans Vis Comput Graph 2023; 29(6): 3121\u20133144. https:\/\/doi.org\/10.1109\/TVCG.2022.3148007","DOI":"10.1109\/TVCG.2022.3148007"},{"key":"e_1_3_2_12_2","doi-asserted-by":"crossref","unstructured":"Ljung P Kr\u00fcger J Groller E et al. State of the art in transfer functions for direct volume rendering. Comput Graph Forum 2016; 35: 669\u2013691. https:\/\/doi.org\/10.1111\/cgf.12934","DOI":"10.1111\/cgf.12934"},{"key":"e_1_3_2_13_2","unstructured":"Arens S Domik G. A survey of transfer functions suitable for volume rendering. In: Proceedings of the 8th IEEE\/EG international conference on volume graphics 2010 pp. 77\u201383. https:\/\/doi.org\/10.2312\/VG\/VG10\/077-083"},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","unstructured":"Ai K Tang K Wang C. NLI4VolVis: natural language interaction for volume visualization via LLM multi-agents and editable 3D Gaussian splatting. IEEE Trans Vis Comput Graph 2026; 32: 46\u201356. https:\/\/doi.org\/10.1109\/TVCG.2025.3633888","DOI":"10.1109\/TVCG.2025.3633888"},{"key":"e_1_3_2_15_2","doi-asserted-by":"crossref","unstructured":"Haidacher M Patel D Bruckner S et al. Volume visualization based on statistical transfer-function spaces. In: 2010 IEEE pacific visualization symposium (PacificVis) 2010 pp. 17\u201324. https:\/\/doi.org\/10.1109\/PACIFICVIS.2010.5429615","DOI":"10.1109\/PACIFICVIS.2010.5429615"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","unstructured":"Correa CD Ma KL. Visibility histograms and visibility-driven transfer functions. IEEE Trans Vis Comput Graph 2011; 17: 192\u2013204. https:\/\/doi.org\/10.1109\/TVCG.2010.35","DOI":"10.1109\/TVCG.2010.35"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","unstructured":"Pan B Lu J Li H et al. Differentiable design galleries: a differentiable approach to explore the design space of transfer functions. IEEE Trans Vis Comput Graph 2024; 30: 1369\u20131379. https:\/\/doi.org\/10.1109\/TVCG.2023.3327371","DOI":"10.1109\/TVCG.2023.3327371"},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","unstructured":"Mallick T Yildiz O Lenz D et al. Chatvis: automating scientific visualization with a large language model 2024. https:\/\/doi.org\/10.1109\/SCW63240.2024.00014.","DOI":"10.1109\/SCW63240.2024.00014"},{"key":"e_1_3_2_19_2","doi-asserted-by":"crossref","unstructured":"Liu S Miao H Bremer PT. Paraview-mcp: An autonomous visualization agent with direct tool use. In: 2025 IEEE visualization and visual analytics (VIS) 2025 pp. 61\u201365. IEEE. https:\/\/doi.org\/10.1109\/vis60296.2025.00018.","DOI":"10.1109\/VIS60296.2025.00018"},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","unstructured":"Liu S Miao H Li Z et al. Ava: towards autonomous visualization agents through visual perception-driven decision-making. Comput Graph Forum 2024; 43: e15093. https:\/\/doi.org\/10.1111\/cgf.15093","DOI":"10.1111\/cgf.15093"},{"key":"e_1_3_2_21_2","doi-asserted-by":"crossref","unstructured":"Srinivasan A Drucker SM Endert A et al. Augmenting visualizations with interactive data facts to facilitate interpretation and communication. IEEE Trans Vis Comput Graph 2018; 25(1): 672\u2013681. https:\/\/doi.org\/10.1109\/TVCG.2018.2865145","DOI":"10.1109\/TVCG.2018.2865145"},{"key":"e_1_3_2_22_2","doi-asserted-by":"crossref","unstructured":"Wang Y Sun Z Zhang H et al. Datashot: automatic generation of fact sheets from tabular data. IEEE Trans Vis Comput Graph 2020; 26(1): 895\u2013905. https:\/\/doi.org\/10.1109\/TVCG.2019.2934398","DOI":"10.1109\/TVCG.2019.2934398"},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","unstructured":"Shi D Xu X Sun F et al. Calliope: automatic visual data story generation from a spreadsheet. IEEE Trans Vis Comput Graph 2021; 27(2): 453\u2013463. https:\/\/doi.org\/10.1109\/TVCG.2020.3030403","DOI":"10.1109\/TVCG.2020.3030403"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","unstructured":"Hoque E Islam MS. Natural language generation for visualizations: state of the art challenges and future directions. Comput Graph Forum 2025; 44: e15266. https:\/\/doi.org\/10.1111\/cgf.15266","DOI":"10.1111\/cgf.15266"},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","unstructured":"Poco J Heer J. Reverse-engineering visualizations: recovering visual encodings from chart images. Comput Graph Forum 2017; 36: 353\u2013363. https:\/\/doi.org\/10.1111\/cgf.13193","DOI":"10.1111\/cgf.13193"},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","unstructured":"Savva M Kong N Chhajta A et al. Revision: automated classification analysis and redesign of chart images. In: Proceedings of the 24th annual ACM symposium on user interface software and technology 2011 pp. 393\u2013402. https:\/\/doi.org\/10.1145\/2047196.2047247.","DOI":"10.1145\/2047196.2047247"},{"key":"e_1_3_2_27_2","unstructured":"Kahou SE Michalski V Atkinson A et al. FigureQA: an annotated figure dataset for visual reasoning. arXiv preprint arXiv:171007300 2017."},{"key":"e_1_3_2_28_2","doi-asserted-by":"crossref","unstructured":"Kim DH Hoque E Agrawala M. Answering questions about charts and generating visual explanations. In: Proceedings of the 2020 CHI conference on human factors in computing systems 2020 pp. 1\u201313. https:\/\/doi.org\/10.1145\/3313831.3376467.","DOI":"10.1145\/3313831.3376467"},{"key":"e_1_3_2_29_2","doi-asserted-by":"crossref","unstructured":"Lundgard A Satyanarayan A. Accessible visualization via natural language descriptions: A four-level model of semantic content. IEEE Trans Vis Comput Graph 2022; 28(1): 1073\u20131083. https:\/\/doi.org\/10.1109\/TVCG.2021.3114770","DOI":"10.1109\/TVCG.2021.3114770"},{"key":"e_1_3_2_30_2","doi-asserted-by":"crossref","unstructured":"Liu C Xie L Han Y et al. Autocaption: an approach to generate natural language description from visualization automatically. In: 2020 IEEE Pacific visualization symposium (PacificVis) 2020 pp. 191\u2013195. IEEE. https:\/\/doi.org\/10.1109\/pacificvis48177.2020.1043","DOI":"10.1109\/PacificVis48177.2020.1043"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","unstructured":"Liu C Guo Y Yuan X. Autotitle: an interactive title generator for visualizations. IEEE Trans Vis Comput Graph 2024; 30(8): 5276\u20135288. https:\/\/doi.org\/10.1109\/TVCG.2023.3290241","DOI":"10.1109\/TVCG.2023.3290241"},{"key":"e_1_3_2_32_2","unstructured":"Liu C Mei X Jiang Z et al. Autolegend: a user feedback-driven adaptive legend generator for visualizations. arXiv preprint arXiv:240716331 2024."},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","unstructured":"Hsu TY Huang CY Huang SH et al. Scicapenter: supporting caption composition for scientific figures with machine-generated captions and ratings. In: Extended abstracts of the CHI conference on human factors in computing systems 2024 pp. 1\u20139. https:\/\/doi.org\/10.1145\/3613905.3650738","DOI":"10.1145\/3613905.3650738"},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","unstructured":"Ng HYS Hsu TY Min J et al. Understanding writing assistants for scientific figure captions: a thematic analysis. In: Proceedings of the fourth workshop on intelligent and interactive writing assistants (In2Writing 2025) 2025 pp. 1\u201310. https:\/\/doi.org\/10.18653\/v1\/2025.in2writing-1.1","DOI":"10.18653\/v1\/2025.in2writing-1.1"},{"key":"e_1_3_2_35_2","doi-asserted-by":"crossref","unstructured":"Borkin MA Bylinskii Z Kim NW et al. Beyond memorability: visualization recognition and recall. IEEE Trans Vis Comput Graph 2016; 22(1). 519\u2013528. https:\/\/doi.org\/10.1109\/TVCG.2015.2467732","DOI":"10.1109\/TVCG.2015.2467732"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","unstructured":"Wanzer DL Azzam T Jones ND et al. The role of titles in enhancing data visualization. Eval Program Plann 2021; 84: 101896. https:\/\/doi.org\/10.1016\/j.evalprogplan.2020.101896","DOI":"10.1016\/j.evalprogplan.2020.101896"},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","unstructured":"Kong HK Liu Z Karahalios K. Frames and slants in titles of visualizations on controversial topics. In: Proceedings of the 2018 CHI conference on human factors in computing systems 2018 pp. 1\u201312. https:\/\/doi.org\/10.1145\/3173574.3174012","DOI":"10.1145\/3173574.3174012"},{"key":"e_1_3_2_38_2","doi-asserted-by":"crossref","unstructured":"Kong HK Liu Z Karahalios K. Trust and recall of information across varying degrees of title-visualization misalignment. In: Proceedings of the 2019 CHI conference on human factors in computing systems 2019 pp. 1\u201313. https:\/\/doi.org\/10.1145\/3290605.3300576.","DOI":"10.1145\/3290605.3300576"},{"key":"e_1_3_2_39_2","doi-asserted-by":"crossref","unstructured":"Sedlmair M Meyer M Munzner T. Design study methodology: reflections from the trenches and the stacks. IEEE Trans Vis Comput Graph 2012; 18(12). 2431\u20132440. https:\/\/doi.org\/10.1109\/TVCG.2012.213","DOI":"10.1109\/TVCG.2012.213"}],"container-title":["Information Visualization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/14738716261434841","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/14738716261434841","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/14738716261434841","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T11:02:49Z","timestamp":1781002969000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/14738716261434841"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,5]]},"references-count":38,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,7]]}},"alternative-id":["10.1177\/14738716261434841"],"URL":"https:\/\/doi.org\/10.1177\/14738716261434841","relation":{},"ISSN":["1473-8716","1473-8724"],"issn-type":[{"value":"1473-8716","type":"print"},{"value":"1473-8724","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,5]]}}}