{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T23:13:35Z","timestamp":1781046815684,"version":"3.54.1"},"reference-count":42,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[2026,4,9]],"date-time":"2026-04-09T00:00:00Z","timestamp":1775692800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"DOI":"10.13039\/501100002367","name":"Chinese Academy of Sciences","doi-asserted-by":"publisher","award":["XDB0500103"],"award-info":[{"award-number":["XDB0500103"]}],"id":[{"id":"10.13039\/501100002367","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100030917","name":"Computer Network Information Center, Chinese Academy of Sciences","doi-asserted-by":"publisher","award":["25YF07"],"award-info":[{"award-number":["25YF07"]}],"id":[{"id":"10.13039\/100030917","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004826","name":"natural science foundation of beijing municipality","doi-asserted-by":"publisher","award":["4254090"],"award-info":[{"award-number":["4254090"]}],"id":[{"id":"10.13039\/501100004826","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Information Visualization"],"published-print":{"date-parts":[[2026,7]]},"abstract":"<jats:p>Scientific visualization pipelines encode domain-specific procedural knowledge with strict execution dependencies, making their construction sensitive to missing stages, incorrect operator usage, or improper ordering. Thus, generating executable scientific visualization pipelines from natural-language descriptions remains challenging for large language models, particularly in web-based environments where visualization authoring relies on explicit code-level pipeline assembly. In this work, we investigate the reliability of LLM-based scientific visualization pipeline generation, focusing on vtk.js as a representative web-based visualization library. We propose a structure-aware retrieval-augmented generation workflow that provides pipeline-aligned vtk.js code examples as contextual guidance, supporting correct module selection, parameter configuration, and execution order. We evaluate the proposed workflow across multiple multi-stage scientific visualization tasks and LLMs, measuring reliability in terms of pipeline executability and human correction effort. To this end, we introduce correction cost as metric for the amount of manual intervention required to obtain a valid pipeline. Our results show that structured, domain-specific context substantially improves pipeline executability and reduces correction cost. We additionally provide an interactive analysis interface to support human-in-the-loop inspection and systematic evaluation of generated visualization pipelines.<\/jats:p>","DOI":"10.1177\/14738716261434848","type":"journal-article","created":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T05:26:44Z","timestamp":1775798804000},"page":"373-390","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":1,"title":["Toward reliable scientific visualization pipeline construction with structure-aware retrieval-augmented LLMs"],"prefix":"10.1177","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-3887-4893","authenticated-orcid":false,"given":"Guanghui","family":"Zhao","sequence":"first","affiliation":[{"name":"Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Hangzhou, China"},{"name":"University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1123-9925","authenticated-orcid":false,"given":"Zhe","family":"Wang","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences, Beijing, China"},{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2353-6567","authenticated-orcid":false,"given":"Yu","family":"Dong","sequence":"additional","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6436-3650","authenticated-orcid":false,"given":"Guan","family":"Li","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences, Beijing, China"},{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8283-2278","authenticated-orcid":false,"given":"Guihua","family":"Shan","sequence":"additional","affiliation":[{"name":"Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Hangzhou, China"},{"name":"University of Chinese Academy of Sciences, Beijing, China"},{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2026,4,9]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"crossref","unstructured":"Mallick T Yildiz O Lenz D et al. Chatvis: Automating scientific visualization with a large language model. In SC24-W: workshops of the international conference for high performance computing networking storage and analysis 2024 pp. 49\u201355. https:\/\/doi.org\/10.1109\/SCW63240.2024.00014","DOI":"10.1109\/SCW63240.2024.00014"},{"key":"e_1_3_3_3_2","doi-asserted-by":"crossref","unstructured":"Biswas A Turton TL Ranasinghe NR et al. Vizgenie: toward self-refining domain-aware workflows for next-generation scientific visualization. IEEE Trans Vis Comput Graph 2026; 32(1). 1021\u20131031. https:\/\/doi.org\/10.1109\/TVCG.2025.3634655","DOI":"10.1109\/TVCG.2025.3634655"},{"key":"e_1_3_3_4_2","volume-title":"The visualization toolkit","author":"Schroeder W","year":"2006","unstructured":"Schroeder W, Martin K, Lorensen B (eds). The visualization toolkit. 4th ed. Kitware, 2006.","edition":"4"},{"key":"e_1_3_3_5_2","doi-asserted-by":"crossref","unstructured":"Jourdain S O\u2019Leary P Schroeder W et al. Trame: platform ubiquitous scalable integration framework for visual analytics. IEEE Comput Graph Appl 2025; 45(2). 126\u2013134. https:\/\/doi.org\/10.1109\/MCG.2025.3540264","DOI":"10.1109\/MCG.2025.3540264"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-37639-0_1"},{"key":"e_1_3_3_7_2","doi-asserted-by":"crossref","unstructured":"Gupta Y Guerrero RED Costa C et al. Interactive web-based 3D viewer for multidimensional microscope imaging modalities. In: 2022 26th International conference information visualisation (IV) 2022 pp. 379\u2013384. https:\/\/doi.org\/10.1109\/IV56949.2022.00069.","DOI":"10.1109\/IV56949.2022.00069"},{"key":"e_1_3_3_8_2","doi-asserted-by":"crossref","unstructured":"Maddigan P Susnjak T. Chat2VIS: generating data visualizations via natural language using ChatGPT Codex and GPT-3 large language models. IEEE Access 2023; 11: 45181\u201345193. https:\/\/doi.org\/10.1109\/access.2023.3274199","DOI":"10.1109\/ACCESS.2023.3274199"},{"key":"e_1_3_3_9_2","doi-asserted-by":"crossref","unstructured":"Jiang J Wang F Shen J et al. A survey on large language models for code generation. ACM Trans Softw Eng Methodol 2026; 35: 1\u201372. https:\/\/doi.org\/10.1145\/3747588","DOI":"10.1145\/3747588"},{"key":"e_1_3_3_10_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-demo.11"},{"key":"e_1_3_3_11_2","unstructured":"Zhang W Shen Y Lu W et al. Data-copilot: Bridging billions of data and humans with autonomous workflow. arXiv preprint arXiv:230607209 2023."},{"key":"e_1_3_3_12_2","doi-asserted-by":"crossref","unstructured":"Wang C Thompson J Lee B. Data formulator: AI-powered concept-driven visualization authoring. IEEE Trans Vis Comput Graph 2024; 30(1). 1128\u20131138. https:\/\/doi.org\/10.1109\/TVCG.2023.3326585","DOI":"10.1109\/TVCG.2023.3326585"},{"key":"e_1_3_3_13_2","doi-asserted-by":"crossref","unstructured":"Zhang C Dong Y Wang Y et al. Auragenome: an LLM-powered framework for on-the-fly reusable and scalable circular genome visualizations. IEEE Comput Graph Appl 2025; 45(5). 78\u201392. https:\/\/doi.org\/10.1109\/MCG.2025.3581560","DOI":"10.1109\/MCG.2025.3581560"},{"key":"e_1_3_3_14_2","doi-asserted-by":"crossref","unstructured":"Dibia V and Demiralp \u00c7. Data2vis: automatic generation of data visualizations using sequence-to-sequence recurrent neural networks. IEEE Comput Graph Appl 2019; 39(5). 33\u201346. https:\/\/doi.org\/10.1109\/MCG.2019.2924636","DOI":"10.1109\/MCG.2019.2924636"},{"key":"e_1_3_3_15_2","doi-asserted-by":"crossref","unstructured":"Chen X Gong L Cheung A et al. Plotcoder: hierarchical decoding for synthesizing visualization code in programmatic context. In: Proceedings of the 59th annual meeting of the Association for Computational Linguistics and the 11th international joint conference on natural language processing (Volume 1: Long Papers) 2021 pp. 2169\u20132181. https:\/\/doi.org\/10.18653\/v1\/2021.acl-long.169","DOI":"10.18653\/v1\/2021.acl-long.169"},{"key":"e_1_3_3_16_2","doi-asserted-by":"crossref","unstructured":"Luo Y Tang N Li G et al. Natural language to visualization by neural machine translation. IEEE Trans Vis Comput Graph 2022; 28(1). 217\u2013226. https:\/\/doi.org\/10.1109\/TVCG.2021.3114848","DOI":"10.1109\/TVCG.2021.3114848"},{"key":"e_1_3_3_17_2","unstructured":"Han Y Zhang C Chen X et al. Chartllama: a multimodal LLM for chart understanding and generation. arXiv preprint arXiv:231116483 2023."},{"key":"e_1_3_3_18_2","doi-asserted-by":"crossref","unstructured":"Yang Z Zhou Z Wang S et al. MatPlotAgent: method and evaluation for LLM-based agentic scientific data visualization. In: Ku LW Martins A Srikumar V (eds) Findings of the Association for Computational Linguistics: ACL 2024. Association for Computational Linguistics 2024 pp. 11789\u201311804. https:\/\/aclanthology.org\/2024.findings-acl.701","DOI":"10.18653\/v1\/2024.findings-acl.701"},{"key":"e_1_3_3_19_2","doi-asserted-by":"crossref","unstructured":"Zhang F Chen B Zhang Y et al. Repocoder: repository-level code completion through iterative retrieval and generation. In: Proceedings of the 2023 conference on empirical methods in natural language processing 2023 pp. 2471\u20132484. https:\/\/doi.org\/10.18653\/v1\/2023.emnlp-main.151","DOI":"10.18653\/v1\/2023.emnlp-main.151"},{"key":"e_1_3_3_20_2","volume-title":"The eleventh international conference on learning representations","author":"Zhou S","unstructured":"Zhou S, Alon U, Xu FF, et al. Docprompting: generating code by retrieving the docs. In: The eleventh international conference on learning representations, 2023."},{"key":"e_1_3_3_21_2","unstructured":"Qin Y Liang S Ye Y et al. ToolLLM: facilitating large language models to master 16000+ real-world APIs. arXiv preprint arXiv:230716789 2023."},{"key":"e_1_3_3_22_2","first-page":"23826","article-title":"Intercode: standardizing and benchmarking interactive coding with execution feedback","volume":"36","author":"Yang J","year":"2023","unstructured":"Yang J, Prabhakar A, Narasimhan K, et al. Intercode: standardizing and benchmarking interactive coding with execution feedback. In: Advances in neural information processing systems, 2023; 36, pp. 23826\u201323854.","journal-title":"Advances in neural information processing systems"},{"key":"e_1_3_3_23_2","doi-asserted-by":"crossref","unstructured":"Koziolek H Gr\u00fcner S Hark R et al. LLM-based and retrieval-augmented control code generation. In: Proceedings of the 1st international workshop on large language models for code 2024 pp. 22\u201329. https:\/\/doi.org\/10.1145\/3643795.3648384","DOI":"10.1145\/3643795.3648384"},{"key":"e_1_3_3_24_2","volume-title":"The paraview guide: a parallel visualization application","author":"Ayachit U.","year":"2015","unstructured":"Ayachit U. The paraview guide: a parallel visualization application. Kitware, Inc, 2015."},{"key":"e_1_3_3_25_2","unstructured":"Kitware. Faster rendering of large number of actors in VTK with WebGPU https:\/\/www.kitware.com\/vtk-webgpu-on-the-desktop\/ (2023 accessed 9 December 2025)."},{"key":"e_1_3_3_26_2","unstructured":"Kitware. vtk.js: the visualization toolkit for the web https:\/\/kitware.github.io\/vtk-js\/docs\/ (2023 accessed 9 December 2025)."},{"key":"e_1_3_3_27_2","unstructured":"Kitware. Volview: all-in-one radiological viewer. https:\/\/kitware.github.io\/VolView\/ (2025 accessed 9 December 2025)."},{"key":"e_1_3_3_28_2","unstructured":"Open Health Imaging Foundation. OHIF viewer documentation (version 2). https:\/\/v2.docs.ohif.org\/ (2025 accessed 9 December 2025)."},{"key":"e_1_3_3_29_2","unstructured":"Kitware. VTK examples https:\/\/examples.vtk.org\/site\/ (2025 accessed 2025)."},{"key":"e_1_3_3_30_2","unstructured":"Yang A Li A Yang B et al. Qwen3 technical report. arXiv preprint arXiv:250509388 2025;."},{"key":"e_1_3_3_31_2","unstructured":"Achiam J Adler S Agarwal S et al.; OpenAI. GPT-4 technical report 2024. https:\/\/arxiv.org\/abs\/2303.08774.2303.08774."},{"key":"e_1_3_3_32_2","doi-asserted-by":"crossref","unstructured":"Denton JD. Lessons from rotor 37. J Therm Sci 1997; 6(1). 1\u201313. https:\/\/doi.org\/10.1007\/s11630-997-0010-9","DOI":"10.1007\/s11630-997-0010-9"},{"key":"e_1_3_3_33_2","unstructured":"Patchett J Gisler G. Deep water impact ensemble data set. Technical Report LA-UR-17-21595 Los Alamos National Laboratory 2017. https:\/\/oceans11.lanl.gov\/deepwaterimpact"},{"key":"e_1_3_3_34_2","unstructured":"The Isabel hurricane dataset. IEEE Visualization Contest 2004 provided by the National Center for Atmospheric Research (NCAR) 2004. http:\/\/vis.computer.org\/vis2004contest\/data.html"},{"key":"e_1_3_3_35_2","unstructured":"Lab KV. IEEE SciVis 2020 contest: visual analysis of lagrangian transport in fluid flows 2020. https:\/\/kaust-vislab.github.io\/SciVis2020\/index.html"},{"key":"e_1_3_3_36_2","unstructured":"Singh A Fry A Perelman A et al. OpenAI GPT-5 system card. arXiv preprint arXiv:260103267 2025."},{"key":"e_1_3_3_37_2","unstructured":"Anthropic. Claude sonnet https:\/\/www.anthropic.com\/claude\/sonnet (2026 accessed 2 March 2026)."},{"key":"e_1_3_3_38_2","unstructured":"Liu A Feng B Xue B et al. Deepseek-V3 technical report. arXiv preprint arXiv:241219437 2024."},{"key":"e_1_3_3_39_2","unstructured":"Guo D Yang D Zhang H et al. DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning. arXiv preprint arXiv:250112948 2025."},{"key":"e_1_3_3_40_2","volume-title":"Interactive transfer function specification for direct volume rendering of disparate volumes","author":"Bernardon FF","year":"2007","unstructured":"Bernardon FF, Ha LK, Callahan SP, et al. Interactive transfer function specification for direct volume rendering of disparate volumes. Technical paper UUSCI-2007-007 ed. University of Utah, 2007."},{"key":"e_1_3_3_41_2","doi-asserted-by":"crossref","unstructured":"Sun J Xie X Yu H. Rmdncache: dual-space prefetching neural network for large-scale volume visualization. IEEE Trans Vis Comput Graph 2025; 31: 4560\u20134575. https:\/\/doi.org\/10.1109\/TVCG.2024.3410091","DOI":"10.1109\/TVCG.2024.3410091"},{"key":"e_1_3_3_42_2","doi-asserted-by":"crossref","unstructured":"Ai K Tang K Wang C. NLI4VolVis: natural language interaction for volume visualization via LLM multi-agents and editable 3D Gaussian splatting. IEEE Trans Vis Comput Graph 2026; 32(1). 46\u201356. https:\/\/doi.org\/10.1109\/TVCG.2025.3633888","DOI":"10.1109\/TVCG.2025.3633888"},{"key":"e_1_3_3_43_2","unstructured":"Threejs Authors. three.js manual https:\/\/threejs.org\/manual\/ (2026 accessed 10 February 2026)."}],"container-title":["Information Visualization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/14738716261434848","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/14738716261434848","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/14738716261434848","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T11:04:06Z","timestamp":1781003046000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/14738716261434848"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,9]]},"references-count":42,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,7]]}},"alternative-id":["10.1177\/14738716261434848"],"URL":"https:\/\/doi.org\/10.1177\/14738716261434848","relation":{},"ISSN":["1473-8716","1473-8724"],"issn-type":[{"value":"1473-8716","type":"print"},{"value":"1473-8724","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,9]]}}}