{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,16]],"date-time":"2026-04-16T10:34:48Z","timestamp":1776335688405,"version":"3.51.2"},"publisher-location":"Cham","reference-count":40,"publisher":"Springer Nature Switzerland","isbn-type":[{"value":"9783032083265","type":"print"},{"value":"9783032083272","type":"electronic"}],"license":[{"start":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T00:00:00Z","timestamp":1760227200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T00:00:00Z","timestamp":1760227200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>We propose a large language model explainability technique for obtaining faithful natural language explanations by grounding the explanations in a reasoning process. When converted to a sequence of tokens, the outputs of the reasoning process can become part of the model context and later be decoded to natural language as the model produces either the final answer or the explanation. To improve the faithfulness of the explanations, we propose to use a joint predict-explain approach, in which the answers and explanations are inferred directly from the reasoning sequence, without the explanations being dependent on the answers and vice versa. We demonstrate the plausibility of the proposed technique by achieving a high alignment between answers and explanations in several problem domains, observing that language models often simply copy the partial decisions from the reasoning sequence into the final answers or explanations. Furthermore, we show that the proposed use of reasoning can also improve the quality of the answers.\n<\/jats:p>","DOI":"10.1007\/978-3-032-08327-2_1","type":"book-chapter","created":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T18:06:45Z","timestamp":1760206005000},"page":"3-18","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Reasoning-Grounded Natural Language Explanations for\u00a0Language Models"],"prefix":"10.1007","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4222-8746","authenticated-orcid":false,"given":"Vojtech","family":"Cahlik","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7458-5281","authenticated-orcid":false,"given":"Rodrigo","family":"Alves","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1433-0089","authenticated-orcid":false,"given":"Pavel","family":"Kordik","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,10,12]]},"reference":[{"issue":"2","key":"1_CR1","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3639372","volume":"15","author":"H Zhao","year":"2024","unstructured":"Zhao, H.: Explainability for large language models: a survey. ACM Trans. Intell. Syst. Technol. 15(2), 1\u201338 (2024)","journal-title":"ACM Trans. Intell. Syst. Technol."},{"issue":"1","key":"1_CR2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2022.103111","volume":"60","author":"E Cambria","year":"2023","unstructured":"Cambria, E., Malandri, L., Mercorio, F., Mezzanzanica, M., Nobani, N.: A survey on XAI and natural language explanations. Inf. Process. Manag. 60(1), 103111 (2023)","journal-title":"Inf. Process. Manag."},{"key":"1_CR3","unstructured":"Russell, S.: Human Compatible: AI and the Problem of Control. Penguin UK (2019)"},{"key":"1_CR4","doi-asserted-by":"crossref","unstructured":"Hrytsyna, A., Alves, R.: From representation to response: assessing the alignment of large language models with human judgment patterns. ACM Trans. Intell. Syst. Technol. (2024)","DOI":"10.1145\/3709148"},{"key":"1_CR5","unstructured":"Weidinger, L., et\u00a0al.: Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359 (2021)"},{"key":"1_CR6","doi-asserted-by":"crossref","unstructured":"DeYoung, J., et al.: ERASER: a benchmark to evaluate rationalized NLP models. arXiv preprint arXiv:1911.03429 (2019)","DOI":"10.18653\/v1\/2020.acl-main.408"},{"key":"1_CR7","doi-asserted-by":"crossref","unstructured":"Wu, Z., Chen, Y., Kao, B., Liu, Q.: Perturbed masking: parameter-free probing for analyzing and interpreting BERT. arXiv preprint arXiv:2004.14786 (2020)","DOI":"10.18653\/v1\/2020.acl-main.383"},{"key":"1_CR8","doi-asserted-by":"crossref","unstructured":"Sikdar, S., Bhattacharya, P., Heese, K.: Integrated directional gradients: feature interaction attribution for neural NLP models. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 865\u2013878 (2021)","DOI":"10.18653\/v1\/2021.acl-long.71"},{"key":"1_CR9","doi-asserted-by":"crossref","unstructured":"Sanyal, S., Ren, X.: Discretized integrated gradients for explaining language models. arXiv preprint arXiv:2108.13654 (2021)","DOI":"10.18653\/v1\/2021.emnlp-main.805"},{"key":"1_CR10","doi-asserted-by":"crossref","unstructured":"Enguehard, J.: Sequential integrated gradients: a simple but effective method for explaining language models. arXiv preprint arXiv:2305.15853 (2023)","DOI":"10.18653\/v1\/2023.findings-acl.477"},{"key":"1_CR11","doi-asserted-by":"crossref","unstructured":"Hao, Y., Dong, L., Wei, F., Ke, X.: Self-attention attribution: Interpreting information interactions inside transformer. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, pp. 12963\u201312971 (2021)","DOI":"10.1609\/aaai.v35i14.17533"},{"key":"1_CR12","doi-asserted-by":"crossref","unstructured":"Lyu, Q., Apidianaki, M., Callison-Burch, C.: Towards faithful model explanation in NLP: survey. Comput. Linguist., 1\u201367 (2024)","DOI":"10.1162\/coli_a_00511"},{"key":"1_CR13","unstructured":"Kokalj, E., \u0160krlj, B., Lavra\u010d, N., Pollak, S., Robnik-\u0160ikonja, M.: BERT meets shapley: extending SHAP explanations to transformer-based classifiers. In: Proceedings of the EACL Hackashop on News Media Content Analysis and Automated Report Generation, pp. 16\u201321 (2021)"},{"key":"1_CR14","doi-asserted-by":"crossref","unstructured":"Vig, J.: A multiscale visualization of attention in the transformer model. arXiv preprint arXiv:1906.05714 (2019)","DOI":"10.18653\/v1\/P19-3007"},{"key":"1_CR15","doi-asserted-by":"crossref","unstructured":"Hoover, B., Strobelt, H., Gehrmann, S.: exBERT: a visual analysis tool to explore learned representations in transformers models. arXiv preprint arXiv:1910.05276 (2019)","DOI":"10.18653\/v1\/2020.acl-demos.22"},{"key":"1_CR16","doi-asserted-by":"crossref","unstructured":"Barkan, O., et al.: Grad-SAM: explaining transformers via gradient self-attention maps. In: Proceedings of the 30th ACM International Conference on Information & Knowledge Management, pp. 2882\u20132887 (2021)","DOI":"10.1145\/3459637.3482126"},{"key":"1_CR17","doi-asserted-by":"crossref","unstructured":"Wu, T., Ribeiro, M.T., Heer, J., Weld, D.S.: Polyjuice: generating counterfactuals for explaining, evaluating, and improving models. arXiv preprint arXiv:2101.00288, 2021","DOI":"10.18653\/v1\/2021.acl-long.523"},{"key":"1_CR18","doi-asserted-by":"crossref","unstructured":"Jin, D., Jin, Z., Zhou, J.T., Szolovits, P.: Is BERT really robust? A strong baseline for natural language attack on text classification and entailment. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol.\u00a034, pp. 8018\u20138025 (2020)","DOI":"10.1609\/aaai.v34i05.6311"},{"key":"1_CR19","doi-asserted-by":"crossref","unstructured":"Garg, S., Ramakrishnan, G.: BAE: BERT-based adversarial examples for text classification. arXiv preprint arXiv:2004.01970 (2020)","DOI":"10.18653\/v1\/2020.emnlp-main.498"},{"key":"1_CR20","unstructured":"Koh, P.W., Liang, P.: Understanding black-box predictions via influence functions. In: International Conference on Machine Learning, pp. 1885\u20131894. PMLR (2017)"},{"key":"1_CR21","unstructured":"Yeh, C.-K., Kim, J., Yen, I.E.-H., Ravikumar, P.K.: Representer point selection for explaining deep neural networks. Adv. Neural Inf. Process. Syst. 31 (2018)"},{"key":"1_CR22","doi-asserted-by":"crossref","unstructured":"Jacovi, A., Goldberg, Y.: Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness? arXiv preprint arXiv:2004.03685 (2020)","DOI":"10.18653\/v1\/2020.acl-main.386"},{"key":"1_CR23","unstructured":"Huang, S., Mamidanna, S., Jangam, S., Zhou, Y., Gilpin, L.H.: Can large language models explain themselves? A study of LLM-generated self-explanations. arXiv preprint arXiv:2310.11207 (2023)"},{"key":"1_CR24","unstructured":"Chen, Y., et al.: Towards consistent natural-language explanations via explanation-consistency finetuning. arXiv preprint arXiv:2401.13986 (2024)"},{"key":"1_CR25","unstructured":"Camburu, O.-M., Rockt\u00e4schel, T., Lukasiewicz, T., Blunsom, P.: e-SNLI: natural language inference with natural language explanations. Adv. Neural Inf. Process. Syst. 31 (2018)"},{"key":"1_CR26","doi-asserted-by":"crossref","unstructured":"Rajani, N.F., Mccann, B., Xiong, C., Socher, R.: Explain yourself! Leveraging language models for commonsense reasoning. arXiv preprint arXiv:1906.02361 (2019)","DOI":"10.18653\/v1\/P19-1487"},{"key":"1_CR27","doi-asserted-by":"crossref","unstructured":"Lyu, Q., et al.: Faithful chain-of-thought reasoning. In: The 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (IJCNLP-AACL 2023) (2023)","DOI":"10.18653\/v1\/2023.ijcnlp-main.20"},{"key":"1_CR28","first-page":"22199","volume":"35","author":"T Kojima","year":"2022","unstructured":"Kojima, T., Gu, S.S., Reid, M., Matsuo, Y., Iwasawa, Y.: Large language models are zero-shot reasoners. Adv. Neural. Inf. Process. Syst. 35, 22199\u201322213 (2022)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"key":"1_CR29","unstructured":"Wang, X., et al.: Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171 (2022)"},{"key":"1_CR30","first-page":"11809","volume":"36","author":"S Yao","year":"2023","unstructured":"Yao, S.: Tree of thoughts: deliberate problem solving with large language models. Adv. Neural. Inf. Process. Syst. 36, 11809\u201311822 (2023)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"issue":"1","key":"1_CR31","doi-asserted-by":"publisher","first-page":"9","DOI":"10.1007\/s44336-024-00009-2","volume":"1","author":"X Li","year":"2024","unstructured":"Li, X., Wang, S., Siqi Zeng, Y.W., Yang, Y.: A survey on LLM-based multi-agent systems: workflow, infrastructure, and challenges. Vicinagearth 1(1), 9 (2024)","journal-title":"Vicinagearth"},{"key":"1_CR32","doi-asserted-by":"crossref","unstructured":"West, P., et al.: Symbolic knowledge distillation: from general language models to commonsense models. arXiv preprint arXiv:2110.07178 (2021)","DOI":"10.18653\/v1\/2022.naacl-main.341"},{"key":"1_CR33","unstructured":"Lightman, H., et al.: Let\u2019s verify step by step. In: The Twelfth International Conference on Learning Representations (2023)"},{"key":"1_CR34","unstructured":"Feng, X., et al.: Alphazero-like tree-search can guide large language model decoding and training. arXiv preprint arXiv:2309.17179 (2023)"},{"key":"1_CR35","unstructured":"Guo, D., et\u00a0al.: DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning. arXiv preprint arXiv:2501.12948 (2025)"},{"key":"1_CR36","unstructured":"Consumer Financial Protection Bureau: HMDA National Loan Level Dataset 2022 (2022)"},{"key":"1_CR37","unstructured":"Dong, Q., et\u00a0al.: A survey on in-context learning. arXiv preprint arXiv:2301.00234 (2022)"},{"key":"1_CR38","doi-asserted-by":"publisher","unstructured":"Cahlik, V.: Research data of the scientific paper \u201creasoning-grounded natural language explanations for language model\u201d (2025). https:\/\/zenodo.org\/records\/15149693. https:\/\/doi.org\/10.5281\/zenodo.15149693.","DOI":"10.5281\/zenodo.15149693."},{"issue":"2","key":"1_CR39","first-page":"3","volume":"1","author":"EJ Hu","year":"2022","unstructured":"Hu, E.J., et al.: LoRA: low-rank adaptation of large language models. ICLR 1(2), 3 (2022)","journal-title":"ICLR"},{"key":"1_CR40","unstructured":"Kingma, D.P., Ba, J.: Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)"}],"container-title":["Communications in Computer and Information Science","Explainable Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-032-08327-2_1","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T18:06:54Z","timestamp":1760206014000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-032-08327-2_1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,12]]},"ISBN":["9783032083265","9783032083272"],"references-count":40,"URL":"https:\/\/doi.org\/10.1007\/978-3-032-08327-2_1","relation":{},"ISSN":["1865-0929","1865-0937"],"issn-type":[{"value":"1865-0929","type":"print"},{"value":"1865-0937","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,12]]},"assertion":[{"value":"12 October 2025","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"The authors have no competing interests to declare that are relevant to the content of this article.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Disclosure of Interests"}},{"value":"xAI","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"World Conference on Explainable Artificial Intelligence","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Istanbul","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"T\u00fcrkiye","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2025","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"9 July 2025","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"11 July 2025","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"3","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"xai2025","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"https:\/\/xaiworldconference.com\/2025\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}}]}}