{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T09:21:03Z","timestamp":1775467263331,"version":"3.50.1"},"reference-count":41,"publisher":"Springer Science and Business Media LLC","issue":"10","license":[{"start":{"date-parts":[[2025,8,27]],"date-time":"2025-08-27T00:00:00Z","timestamp":1756252800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,8,27]],"date-time":"2025-08-27T00:00:00Z","timestamp":1756252800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"HORIZON EUROPE, European Union","award":["01093026"],"award-info":[{"award-number":["01093026"]}]},{"name":"National Recovery and Resilience Plan Greece 2.0","award":["MIS 5154714"],"award-info":[{"award-number":["MIS 5154714"]}]},{"DOI":"10.13039\/100031478","name":"NextGenerationEU","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100031478","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100009244","name":"Stockholm University","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100009244","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2025,10]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    This paper investigates the reliability of explanations generated by large language models\u00a0(LLMs) when prompted to explain their previous output. We evaluate two kinds of such self-explanations\u00a0(\n                    <jats:sc>SE<\/jats:sc>\n                    )\u2014extractive and counterfactual\u2014using state-of-the-art LLMs (1B to 70B parameters) on three different classification tasks (both objective and subjective). In line with Agarwal et al. (Faithfulness versus plausibility: On the (Un)reliability of explanations from large language models. 2024.\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"10.48550\/arXiv.2402.04614\" ext-link-type=\"doi\">https:\/\/doi.org\/10.48550\/arXiv.2402.04614<\/jats:ext-link>\n                    ), our findings indicate a gap between perceived and actual model reasoning: while\n                    <jats:sc>SE<\/jats:sc>\n                    largely correlate with human judgment (i.e. are\n                    <jats:italic>plausible<\/jats:italic>\n                    ), they do not fully and accurately follow the model\u2019s decision process (i.e. are not\n                    <jats:italic>faithful<\/jats:italic>\n                    ). Additionally, we show that counterfactual\n                    <jats:sc>SE<\/jats:sc>\n                    are not even necessarily\n                    <jats:italic>valid<\/jats:italic>\n                    in the sense of actually changing the LLM\u2019s prediction. Our results suggest that extractive\n                    <jats:sc>SE<\/jats:sc>\n                    provide the LLM\u2019s \u201cguess\u201d at an explanation based on training data. Conversely, counterfactual\n                    <jats:sc>SE<\/jats:sc>\n                    can help understand the LLM\u2019s reasoning: We show that the issue of validity can be resolved by sampling counterfactual candidates at high temperature\u2014followed by a validity check\u2014and introducing a formula to estimate the number of tries needed to generate valid explanations. This simple method produces plausible and valid explanations that offer a 16 times\u00a0faster alternative to SHAP\u00a0on average in our experiments.\n                  <\/jats:p>","DOI":"10.1007\/s10994-025-06838-6","type":"journal-article","created":{"date-parts":[[2025,8,27]],"date-time":"2025-08-27T17:16:38Z","timestamp":1756314998000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Mind the gap: from plausible to valid self-explanations in large language models"],"prefix":"10.1007","volume":"114","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7938-2747","authenticated-orcid":false,"given":"Korbinian","family":"Randl","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9188-7425","authenticated-orcid":false,"given":"John","family":"Pavlopoulos","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9731-1048","authenticated-orcid":false,"given":"Aron","family":"Henriksson","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7713-1381","authenticated-orcid":false,"given":"Tony","family":"Lindgren","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,8,27]]},"reference":[{"key":"6838_CR1","doi-asserted-by":"crossref","unstructured":"Abnar, S., & Zuidema, W. (2020). Quantifying attention flow in transformers. In Proceedings of ACL (pp. 4190\u20134197).","DOI":"10.18653\/v1\/2020.acl-main.385"},{"issue":"1","key":"6838_CR2","first-page":"147","volume":"9","author":"DH Ackley","year":"1985","unstructured":"Ackley, D. H., Hinton, G. E., & Sejnowski, T. J. (1985). A learning algorithm for Boltzmann machines. Cognitive Science, 9(1), 147\u2013169.","journal-title":"Cognitive Science"},{"key":"6838_CR3","doi-asserted-by":"publisher","unstructured":"Agarwal, C., Tanneru, S. H., & Lakkaraju, H. (2024). Faithfulness versus plausibility: On the (Un)reliability of explanations from large language models. https:\/\/doi.org\/10.48550\/arXiv.2402.04614","DOI":"10.48550\/arXiv.2402.04614"},{"key":"6838_CR4","doi-asserted-by":"publisher","unstructured":"Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Debbah, M., & Penedo, G. (2023). The falcon series of open language models. https:\/\/doi.org\/10.48550\/arXiv.2311.16867","DOI":"10.48550\/arXiv.2311.16867"},{"issue":"7","key":"6838_CR5","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1371\/journal.pone.0130140","volume":"10","author":"S Bach","year":"2015","unstructured":"Bach, S., Binder, A., Montavon, G., Klauschen, F., M\u00fcller, K.-R., & Samek, W. (2015). On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLOS ONE, 10(7), 1\u201346.","journal-title":"PLOS ONE"},{"key":"6838_CR6","doi-asserted-by":"publisher","unstructured":"Brandl, S., & Eberle, O. (2025). Comparing zero-shot self-explanations with human rationales in text classification. https:\/\/doi.org\/10.48550\/arXiv.2410.03296","DOI":"10.48550\/arXiv.2410.03296"},{"key":"6838_CR7","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., & Amodei, D. (2020). Language models are few-shot learners. Proceedings of NeurIPS, 33, 1877\u20131901.","journal-title":"Proceedings of NeurIPS"},{"key":"6838_CR8","doi-asserted-by":"crossref","unstructured":"Chefer, H., Gur, S., & Wolf, L. (2021). Transformer interpretability beyond attention visualization. In: 2021 IEEE\/CVF conference on computer vision and pattern recognition (CVPR) (pp. 782\u2013791).","DOI":"10.1109\/CVPR46437.2021.00084"},{"key":"6838_CR41","doi-asserted-by":"publisher","unstructured":"Chen, Y., Benton, J., Radhakrishnan, A., et al.\u00a0(2025). Reasoning Models Don\u2019t Always Say What They Think.\u00a0https:\/\/doi.org\/10.48550\/arXiv.2505.05410","DOI":"10.48550\/arXiv.2505.05410"},{"key":"6838_CR9","unstructured":"Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL (pp. 4171\u20134186)."},{"key":"6838_CR10","doi-asserted-by":"publisher","unstructured":"Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., & Ganapathy, R. (2024). The llama 3 herd of models. https:\/\/doi.org\/10.48550\/arXiv.2407.21783","DOI":"10.48550\/arXiv.2407.21783"},{"key":"6838_CR11","doi-asserted-by":"publisher","first-page":"1225093","DOI":"10.3389\/frai.2023.1225093","volume":"6","author":"S Gurrapu","year":"2023","unstructured":"Gurrapu, S., Kulkarni, A., Huang, L., Lourentzou, I., & Batarseh, F. A. (2023). Rationalization for explainable nlp: A survey. Frontiers in Artificial Intelligence, 6, 1225093.","journal-title":"Frontiers in Artificial Intelligence"},{"key":"6838_CR12","doi-asserted-by":"publisher","unstructured":"Hechtlinger, Y. (2016). Interpretation of prediction models using the input gradient. https:\/\/doi.org\/10.48550\/arXiv.1611.07634.","DOI":"10.48550\/arXiv.1611.07634"},{"key":"6838_CR13","doi-asserted-by":"publisher","unstructured":"Holtzman, A., Buys, J., Du, L., Forbes, M., & Choi, Y. (2020). The curious case of neural text degeneration. https:\/\/doi.org\/10.48550\/arXiv.1904.09751","DOI":"10.48550\/arXiv.1904.09751"},{"key":"6838_CR14","doi-asserted-by":"publisher","unstructured":"Huang, S., Mamidanna, S., Jangam, S., Zhou, Y., & Gilpin, L. H. (2023). Can large language models explain themselves? A study of llm-generated self-explanations. https:\/\/doi.org\/10.48550\/arXiv.2310.11207","DOI":"10.48550\/arXiv.2310.11207"},{"key":"6838_CR15","doi-asserted-by":"crossref","unstructured":"Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., & Zettlemoyer, L. (2020). BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of ACL (pp. 7871\u20137880).","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"6838_CR16","doi-asserted-by":"publisher","unstructured":"Li, J., Monroe, W., & Jurafsky, D. (2017). Understanding neural networks through representation erasure. https:\/\/doi.org\/10.48550\/arXiv.1612.08220","DOI":"10.48550\/arXiv.1612.08220"},{"key":"6838_CR17","unstructured":"Lin, C.-Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Text summarization branches out (pp. 74\u201381)."},{"key":"6838_CR18","doi-asserted-by":"crossref","unstructured":"Liu, S., Le, F., Chakraborty, S., & Abdelzaher, T. (2021). On exploring attention-based explanation for transformer models in text classification. In 2021 IEEE international conference on big data (big data) (pp. 1193\u20131203).","DOI":"10.1109\/BigData52589.2021.9671639"},{"key":"6838_CR19","doi-asserted-by":"publisher","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). RoBERTa: A robustly optimized Bert pretraining approach. https:\/\/doi.org\/10.48550\/arXiv.1907.11692","DOI":"10.48550\/arXiv.1907.11692"},{"key":"6838_CR20","doi-asserted-by":"publisher","unstructured":"Liu, F., Xu, P., Li, Z., Feng, Y., & Song, H. (2024). Towards understanding in-context learning with contrastive demonstrations and saliency maps. https:\/\/doi.org\/10.48550\/arXiv.2307.05052","DOI":"10.48550\/arXiv.2307.05052"},{"key":"6838_CR21","first-page":"4765","volume":"30","author":"SM Lundberg","year":"2017","unstructured":"Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Proceedings of NeurIPS, 30, 4765\u20134774.","journal-title":"Proceedings of NeurIPS"},{"key":"6838_CR22","doi-asserted-by":"crossref","unstructured":"Madsen, A., Chandar, S., & Reddy, S. (2024). Are self-explanations from large language models faithful? In Findings of ACL (pp. 295\u2013337).","DOI":"10.18653\/v1\/2024.findings-acl.19"},{"key":"6838_CR23","doi-asserted-by":"publisher","unstructured":"Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., & Kenealy, K. (2024). Gemma: Open models based on Gemini research and technology. https:\/\/doi.org\/10.48550\/arXiv.2403.08295","DOI":"10.48550\/arXiv.2403.08295"},{"key":"6838_CR24","doi-asserted-by":"crossref","unstructured":"Pang, B., & Lee, L. (2005). Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Proceedings of ACL (pp. 115\u2013124).","DOI":"10.3115\/1219840.1219855"},{"key":"6838_CR25","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). Bleu: A method for automatic evaluation of machine translation. In Proceedings of ACL (pp. 311\u2013318).","DOI":"10.3115\/1073083.1073135"},{"key":"6838_CR26","doi-asserted-by":"crossref","unstructured":"Pavlopoulos, J., Laugier, L., Xenos, A., Sorensen, J., & Androutsopoulos, I. (2022). From the detection of toxic spans in online discussions to the analysis of toxic-to-civil transfer. In Proceedings of ACL (pp. 3721\u20133734).","DOI":"10.18653\/v1\/2022.acl-long.259"},{"issue":"1","key":"6838_CR27","first-page":"1","volume":"21","author":"C Raffel","year":"2020","unstructured":"Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(1), 1\u201367.","journal-title":"Journal of Machine Learning Research"},{"key":"6838_CR28","doi-asserted-by":"crossref","unstructured":"Randl, K., Pavlopoulos, J., Henriksson, A., & Lindgren, T. (2024). CICLe: Conformal in-context learning for largescale multi-class food risk classification. In Findings of ACL (pp. 7695\u20137715).","DOI":"10.18653\/v1\/2024.findings-acl.459"},{"key":"6838_CR29","doi-asserted-by":"crossref","unstructured":"Randl, K., Pavlopoulos, J., Henriksson, A., & Lindgren, T. (2025). Evaluating the reliability of self-explanations in large language models. In Discovery science (pp. 36\u201351).","DOI":"10.1007\/978-3-031-78977-9_3"},{"key":"6838_CR30","doi-asserted-by":"crossref","unstructured":"Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). Why should i trust you?: Explaining the predictions of any classifier. In Proceedings of KDD (pp. 1135\u20131144).","DOI":"10.1145\/2939672.2939778"},{"key":"6838_CR31","doi-asserted-by":"publisher","unstructured":"Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., & Garg, S. (2024). Gemma 2: Improving open language models at a practical size. https:\/\/doi.org\/10.48550\/arXiv.2408.00118","DOI":"10.48550\/arXiv.2408.00118"},{"key":"6838_CR32","first-page":"3145","volume":"70","author":"A Shrikumar","year":"2017","unstructured":"Shrikumar, A., Greenside, P., & Kundaje, A. (2017). Learning important features through propagating activation differences. Proceedings of ICML, 70, 3145\u20133153.","journal-title":"Proceedings of ICML"},{"key":"6838_CR33","doi-asserted-by":"crossref","unstructured":"Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., & Potts, C. (2013). Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of EMNLP (pp. 1631\u20131642).","DOI":"10.18653\/v1\/D13-1170"},{"key":"6838_CR34","first-page":"3319","volume":"70","author":"M Sundararajan","year":"2017","unstructured":"Sundararajan, M., Taly, A., & Yan, Q. (2017). Axiomatic attribution for deep networks. Proceedings of ICML, 70, 3319\u20133328.","journal-title":"Proceedings of ICML"},{"key":"6838_CR35","first-page":"6000","volume":"31","author":"A Vaswani","year":"2017","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. Proceedings of NeurIPS, 31, 6000\u20136010.","journal-title":"Proceedings of NeurIPS"},{"key":"6838_CR36","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3677119","volume":"56","author":"S Verma","year":"2024","unstructured":"Verma, S., Boonsanong, V., Hoang, M., Hines, K., Dickerson, J., & Shah, C. (2024). Counterfactual explanations and algorithmic recourses for machine learning: A review. ACM Computing Surveys, 56, 1\u201342.","journal-title":"ACM Computing Surveys"},{"key":"6838_CR37","first-page":"24824","volume":"35","author":"J Wei","year":"2022","unstructured":"Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Proceedings of NeurIPS, 35, 24824\u201324837.","journal-title":"Proceedings of NeurIPS"},{"key":"6838_CR38","first-page":"27263","volume":"34","author":"W Yuan","year":"2021","unstructured":"Yuan, W., Neubig, G., & Liu, P. (2021). Bartscore: Evaluating generated text as text generation. Proceedings of NeurIPS, 34, 27263\u201327277.","journal-title":"Proceedings of NeurIPS"},{"key":"6838_CR39","doi-asserted-by":"crossref","unstructured":"Zaidan, O., & Eisner, J. (2008). Modeling annotators: A generative approach to learning from annotator rationales. In Proceedings of EMNLP (pp. 31\u201340).","DOI":"10.3115\/1613715.1613721"},{"issue":"2","key":"6838_CR40","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3639372","volume":"15","author":"H Zhao","year":"2024","unstructured":"Zhao, H., Chen, H., Yang, F., Liu, N., Deng, H., Cai, H., Wang, S., Yin, D., & Du, M. (2024). Explainability for large language models: A survey. ACM Transactions on Intelligent Systems and Technology, 15(2), 1\u201338.","journal-title":"ACM Transactions on Intelligent Systems and Technology"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06838-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-025-06838-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-025-06838-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,7]],"date-time":"2025-10-07T20:56:47Z","timestamp":1759870607000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-025-06838-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,27]]},"references-count":41,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2025,10]]}},"alternative-id":["6838"],"URL":"https:\/\/doi.org\/10.1007\/s10994-025-06838-6","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-6263278\/v1","asserted-by":"object"}]},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,27]]},"assertion":[{"value":"19 March 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 June 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 July 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 August 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}],"article-number":"220"}}