{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T09:09:03Z","timestamp":1780564143575,"version":"3.54.1"},"reference-count":51,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T00:00:00Z","timestamp":1780531200000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T00:00:00Z","timestamp":1780531200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100001459","name":"Ministry of Education - Singapore","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100001459","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100010016","name":"Nottingham Trent University","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100010016","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Current Large Language Models (LLMs), especially Large Reasoning Models, can generate Chain\u2010of\u2010Thought (CoT) reasoning traces to illustrate how they produce final outputs, thereby facilitating trust calibration for users. However, these CoT reasoning traces are usually lengthy and tedious, and can contain various issues, such as logical and factual errors, which make it difficult for users to interpret the reasoning traces efficiently and accurately. To address these challenges, we develop an error detection pipeline that combines external fact\u2010checking with symbolic formal logical validation to identify errors at the step level. Building on this pipeline, we propose ReasonDiag, an interactive visualization system for diagnosing CoT reasoning traces. ReasonDiag provides 1) an integrated arc diagram to show reasoning\u2010step distributions and error\u2010propagation patterns, and 2) a hierarchical node\u2010link diagram to visualize high\u2010level reasoning flows and premise dependencies. We evaluate Reason\u2010Diag through a technical evaluation for the error detection pipeline, two case studies, and user interviews with 16 participants. The results indicate that ReasonDiag helps users effectively understand CoT reasoning traces, identify erroneous steps, and determine their root causes.<\/jats:p>","DOI":"10.1111\/cgf.70439","type":"journal-article","created":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T08:46:27Z","timestamp":1780562787000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["When the Chain Breaks: Interactive Diagnosis of LLM Chain\u2010of\u2010Thought Reasoning Errors"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-9106-7989","authenticated-orcid":false,"given":"Shiwei","family":"Chen","sequence":"first","affiliation":[{"name":"College of Computing and Data Science, Nanyang Technological University  Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-3322-3218","authenticated-orcid":false,"given":"Niruthikka","family":"Sritharan","sequence":"additional","affiliation":[{"name":"College of Computing and Data Science, Nanyang Technological University  Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8562-7640","authenticated-orcid":false,"given":"Xiaolin","family":"Wen","sequence":"additional","affiliation":[{"name":"College of Computing and Data Science, Nanyang Technological University  Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-3814-8443","authenticated-orcid":false,"given":"Chenxi","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Computing and Data Science, Nanyang Technological University  Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5693-1128","authenticated-orcid":false,"given":"Xingbo","family":"Wang","sequence":"additional","affiliation":[{"name":"Bosch Research North America  Sunnyvale CA United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0092-0793","authenticated-orcid":false,"given":"Yong","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Computing and Data Science, Nanyang Technological University  Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2026,6,4]]},"reference":[{"key":"e_1_2_10_2_2","unstructured":"AIM.:Api specs | mistral ai docs.https:\/\/docs.mistral.ai\/api 2026. Accessed: 2026-02-27. 13"},{"key":"e_1_2_10_3_2","unstructured":"Anthropic:Messages api | anthropic.https:\/\/platform.claude.com\/docs\/en\/api\/messages 2026. Accessed: 2026-02-27. 13"},{"key":"e_1_2_10_4_2","unstructured":"APIS.:Serper api. URL:https:\/\/serper.dev\/. 4"},{"key":"e_1_2_10_5_2","unstructured":"BakerB. HuizingaJ. GaoL. DouZ. GuanM. Y. MadryA. ZarembaW. PachockiJ. FarhiD.:Monitoring reasoning models for misbehavior and the risks of promoting obfuscation 2025. URL:https:\/\/arxiv.org\/abs\/2503.11926 arXiv:2503.11926. 1"},{"key":"e_1_2_10_6_2","unstructured":"BogdanP. C. MacarU. NandaN. ConmyA.: Thought anchors: Which llm reasoning steps matter?arXiv preprint arXiv:2506.19143(2025). 3 4 14"},{"key":"e_1_2_10_7_2","unstructured":"BarezF. WuT.-Y. ArcuschinI. LanM. WangV. SiegelN. CollignonN. NeoC. LeeI. ParenA. et al.: Chain-of-thought is not explainability.Preprint alphaXiv(2025) v1. 1"},{"key":"e_1_2_10_8_2","doi-asserted-by":"crossref","DOI":"10.3386\/w34255","volume-title":"How people use chatgpt","author":"Chatterji A.","year":"2025"},{"key":"e_1_2_10_9_2","doi-asserted-by":"crossref","unstructured":"ChenD.: Llmsr@ xllm25: Swrv: Empowering self-verification of small language models through step-wise reasoning and verification. InProceedings of the 1st Joint Workshop on Large Language Models and Structure Modeling (XLLM 2025)(2025) pp.322\u2013335. 2 5","DOI":"10.18653\/v1\/2025.xllm-1.29"},{"key":"e_1_2_10_10_2","unstructured":"ChenQ. QinL. LiuJ. PengD. GuanJ. WangP. HuM. ZhouY. GaoT. CheW.: Towards reasoning era: A survey of long chain-of-thought for reasoning large language models.arXiv preprint arXiv:2503.09567(2025). 1"},{"key":"e_1_2_10_11_2","unstructured":"ChengF. ZouharV. ChanR. S. M. F\u00fcrstD. StrobeltH. El-AssadyM.: Understanding large language model behaviors through interactive counterfactual generation and analysis.IEEE Transactions on Visualization and Computer Graphics(2025). 2"},{"key":"e_1_2_10_12_2","unstructured":"DeepSeek:Deepseek api | deepseek api docs.https:\/\/api-docs.deepseek.com\/api\/deepseek-api 2026. Accessed: 2026-02-27. 13"},{"key":"e_1_2_10_13_2","first-page":"337","volume-title":"International conference on Tools and Algorithms for the Construction and Analysis of Systems","author":"De Moura L.","year":"2008"},{"key":"e_1_2_10_14_2","unstructured":"Google:Gemini thinking | gemini api | google ai for developers 2026. URL:https:\/\/ai.google.dev\/gemini-api\/docs\/thinking. 13"},{"key":"e_1_2_10_15_2","unstructured":"GuanM. Y. WangM. CarrollM. DouZ. WeiA. Y. WilliamsM. ArnavB. HuizingaJ. KivlichanI. GlaeseM. et al.: Monitoring monitorability.arXiv preprint arXiv:2512.18311(2025). 1"},{"key":"e_1_2_10_16_2","unstructured":"GuoD. YangD. ZhangH. SongJ. ZhangR. XuR. ZhuQ. MaS. WangP. BiX. et al.: Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948(2025). 1"},{"key":"e_1_2_10_17_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.acl-long.905"},{"key":"e_1_2_10_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3571730"},{"key":"e_1_2_10_18_3","doi-asserted-by":"crossref","unstructured":"doi:10.1145\/3571730. 4","DOI":"10.1145\/3571730"},{"key":"e_1_2_10_19_2","doi-asserted-by":"crossref","unstructured":"JiangP. RayanJ. DowS. P. XiaH.: Graphologue: Exploring large language model responses with interactive diagrams. InProceedings of the 36th annual ACM symposium on user interface software and technology(2023) pp.1\u201320. 2","DOI":"10.1145\/3586183.3606737"},{"key":"e_1_2_10_20_2","doi-asserted-by":"crossref","unstructured":"KocielnikR. AmershiS. BennettP. N.: Will you accept an imperfect ai? exploring designs for adjusting end-user expectations of ai systems. InProceedings of the 2019 CHI conference on human factors in computing systems(2019) pp.1\u201314. 7","DOI":"10.1145\/3290605.3300641"},{"key":"e_1_2_10_21_2","unstructured":"KorbakT. BalesniM. BarnesE. BengioY. BentonJ. BloomJ. ChenM. CooneyA. DafoeA. DraganA. et al.: Chain of thought monitorability: A new and fragile opportunity for ai safety.arXiv preprint arXiv:2507.11473(2025). 1"},{"key":"e_1_2_10_22_2","first-page":"28","volume-title":"Proceedings of the 31st International Conference on Computational Linguistics: System Demonstrations","author":"Li H.","year":"2025"},{"key":"e_1_2_10_23_2","unstructured":"LightmanH. KosarajuV. BurdaY. EdwardsH. BakerB. LeeT. LeikeJ. SchulmanJ. SutskeverI. CobbeK.: Let's verify step by step. InThe Twelfth International Conference on Learning Representations(2023). 2 4"},{"key":"e_1_2_10_24_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.acl-demo.14"},{"key":"e_1_2_10_25_2","doi-asserted-by":"crossref","first-page":"29655","DOI":"10.1609\/aaai.v39i28.35357","article-title":"Llm attributor: Interactive visual attribution for llm generation","volume":"39","author":"Lee S.","year":"2025","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_2_10_26_2","unstructured":"MukherjeeS. ChintaA. KimT. SharmaT. A. TurD. H.: Premise-augmented reasoning chains improve error identification in math reasoning with LLMs. InForty-second International Conference on Machine Learning(2025). URL:https:\/\/openreview.net\/forum?id=4tYckHNVXV. 2 3 4 5 14"},{"key":"e_1_2_10_27_2","article-title":"Gaia: a benchmark for general ai assistants","volume":"9","author":"Mialon G.","journal-title":"The Twelfth International Conference on Learning Representations."},{"key":"e_1_2_10_28_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.fever-1.27"},{"key":"e_1_2_10_29_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-naacl.35"},{"key":"e_1_2_10_30_2","unstructured":"OpenAI:Chatgpt: Language model. Accessed: 2023-10-01 2023. URL:https:\/\/chatgpt.com\/. 1 3"},{"key":"e_1_2_10_31_2","unstructured":"OpenAI:Reasoning models | openai api 2026. URL:https:\/\/developers.openai.com\/api\/docs\/guides\/reasoning. 1 13"},{"key":"e_1_2_10_32_2","doi-asserted-by":"crossref","unstructured":"PanL. AlbalakA. WangX. WangW.: Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning. InFindings of the Association for Computational Linguistics: EMNLP 2023(2023) pp.3806\u20133824. 2 5","DOI":"10.18653\/v1\/2023.findings-emnlp.248"},{"key":"e_1_2_10_33_2","volume-title":"The PageRank citation ranking: Bringing order to the web","author":"Page L.","year":"1999"},{"key":"e_1_2_10_34_2","unstructured":"PangR. Y. FengK. FengS. LiC. ShiW. TsvetkovY. HeerJ. ReineckeK.: Interactive reasoning: Visualizing and controlling chain-of-thought reasoning in large language models.arXiv preprint arXiv:2506.23678(2025). 1 2"},{"key":"e_1_2_10_35_2","unstructured":"Qwen:Qwen api platform.https:\/\/qwen.ai\/apiplatform\/ 2026. Accessed: 2026-02-27. 13"},{"key":"e_1_2_10_36_2","doi-asserted-by":"crossref","unstructured":"ShneidermanB.:The eyes have it: A task by data type taxonomy for information visualizations. InThe Craft of Information Visualization BEDERSONB. B. SHNEIDERMANB. (Eds.) Interactive Technologies. Morgan Kaufmann San Francisco 2003 pp.364\u2013371. URL:https:\/\/www.sciencedirect.com\/science\/article\/pii\/B9781558609150500469 doi:https:\/\/doi.org\/10.1016\/B978-155860915-0\/50046-9. 3","DOI":"10.1016\/B978-155860915-0\/50046-9"},{"key":"e_1_2_10_37_2","doi-asserted-by":"crossref","first-page":"1123","DOI":"10.1145\/3708359.3712122","volume-title":"Proceedings of the 30th International Conference on Intelligent User Interfaces","author":"Srinivasan A.","year":"2025"},{"key":"e_1_2_10_37_3","doi-asserted-by":"crossref","unstructured":"doi:10.1145\/3708359.3712122. 2","DOI":"10.1145\/3708359.3712122"},{"key":"e_1_2_10_38_2","unstructured":"TeamG. AnilR. BorgeaudS. AlayracJ.-B. YuJ. SoricutR. SchalkwykJ. DaiA. M. HauthA. MillicanK. et al.: Gemini: a family of highly capable multimodal models.arXiv preprint arXiv:2312.11805(2023). 1"},{"key":"e_1_2_10_39_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.826"},{"key":"e_1_2_10_40_2","article-title":"Understanding reasoning in thinking language models via steering vectors","volume":"4","author":"Venhoff C.","journal-title":"Workshop on Reasoning and Planning for Large Language Models."},{"issue":"1","key":"e_1_2_10_41_2","doi-asserted-by":"crossref","first-page":"273","DOI":"10.1109\/TVCG.2023.3327153","article-title":"Commonsensevis: Visualizing and understanding commonsense reasoning capabilities of natural language models","volume":"30","author":"Wang X.","year":"2023","journal-title":"IEEE Transactions on Visualization and Computer Graphics"},{"key":"e_1_2_10_42_2","doi-asserted-by":"crossref","unstructured":"WangC. LeeB. DruckerS. M. MarshallD. GaoJ.: Data formulator 2: Iterative creation of data visualizations with ai transforming data along the way. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems(2025) pp.1\u201317. 2","DOI":"10.1145\/3706598.3713296"},{"key":"e_1_2_10_43_2","article-title":"Self-consistency improves chain of thought reasoning in language models","volume":"2","author":"Wang X.","journal-title":"The Eleventh International Conference on Learning Representations."},{"key":"e_1_2_10_44_2","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei J.","year":"2022","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_10_45_2","unstructured":"xAI:Inference api - rest api reference | xai docs.https:\/\/docs.x.ai\/developers\/rest-api-reference\/inference 2026. Accessed: 2026-02-27. 13"},{"key":"e_1_2_10_46_2","doi-asserted-by":"crossref","unstructured":"XuJ. FeiH. PanL. LiuQ. LeeM.-L. HsuW.: Faithful logical reasoning via symbolic chain-of-thought. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)(2024) pp.13326\u201313365. 2 5","DOI":"10.18653\/v1\/2024.acl-long.720"},{"key":"e_1_2_10_47_2","doi-asserted-by":"crossref","unstructured":"XieL. ZhengC. XiaH. QuH. Zhu-TianC.: Waitgpt: Monitoring and steering conversational llm agent in data analysis with on-the-fly code visualization. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology(2024) pp.1\u201314. 2","DOI":"10.1145\/3654777.3676374"},{"key":"e_1_2_10_48_2","unstructured":"YangA. LiA. YangB. ZhangB. HuiB. ZhengB. YuB. GaoC. HuangC. LvC. ZhengC. LiuD. ZhouF. HuangF. HuF. GeH. WeiH. LinH. TangJ. YangJ. TuJ. ZhangJ. YangJ. YangJ. ZhouJ. ZhouJ. LinJ. DangK. BaoK. YangK. YuL. DengL. LiM. XueM. LiM. ZhangP. WangP. ZhuQ. MenR. GaoR. LiuS. LuoS. LiT. TangT. YinW. RenX. WangX. ZhangX. RenX. FanY. SuY. ZhangY. ZhangY. WanY. LiuY. WangZ. CuiZ. ZhangZ. ZhouZ. QiuZ.:Qwen3 technical report 2025. URL:https:\/\/arxiv.org\/abs\/2505.09388 arXiv:2505.09388. 1"},{"key":"e_1_2_10_49_2","unstructured":"ZhouR. NguyenG. KharyaN. NguyenA. T. AgarwalC.: Improving human verification of llm reasoning through interactive explanation interfaces.arXiv preprint arXiv:2510.22922(2025). 1 2"},{"key":"e_1_2_10_50_2","unstructured":"ZhouZ. ZhuZ. LiX. GalkinM. FengX. KoyejoS. TangJ. HanB.: Landscape of thoughts: Visualizing the reasoning process of large language models.arXiv preprint arXiv:2503.22165(2025). 2"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70439","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70439","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70439","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T08:46:44Z","timestamp":1780562804000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70439"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,4]]},"references-count":51,"alternative-id":["10.1111\/cgf.70439"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70439","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,4]]},"assertion":[{"value":"2026-06-04","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70439"}}