{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T06:41:05Z","timestamp":1781505665234,"version":"3.54.1"},"publisher-location":"Cham","reference-count":32,"publisher":"Springer Nature Switzerland","isbn-type":[{"value":"9783032083296","type":"print"},{"value":"9783032083302","type":"electronic"}],"license":[{"start":{"date-parts":[[2025,10,14]],"date-time":"2025-10-14T00:00:00Z","timestamp":1760400000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,10,14]],"date-time":"2025-10-14T00:00:00Z","timestamp":1760400000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>The need for interpretability in deep learning has driven interest in counterfactual explanations, which identify minimal changes to an instance that change a model\u2019s prediction. Current counterfactual (CF) generation methods require task-specific fine-tuning and produce low-quality text. Large Language Models (LLMs), though effective for high-quality text generation, struggle with label-flipping counterfactuals (i.e., counterfactuals that change the prediction) without fine-tuning. We introduce two simple classifier-guided approaches to support counterfactual generation by LLMs, eliminating the need for fine-tuning while preserving the strengths of LLMs. Despite their simplicity, our methods outperform state-of-the-art counterfactual generation methods and are effective across different LLMs, highlighting the benefits of guiding counterfactual generation by LLMs with classifier information. We further show that data augmentation by our generated CFs can improve a classifier\u2019s robustness. Our analysis reveals a critical issue in counterfactual generation by LLMs: LLMs rely on parametric knowledge rather than faithfully following the classifier.<\/jats:p>","DOI":"10.1007\/978-3-032-08330-2_8","type":"book-chapter","created":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T03:10:10Z","timestamp":1760325010000},"page":"158-176","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Guiding LLMs to\u00a0Generate High-Fidelity and\u00a0High-Quality Counterfactual Explanations for\u00a0Text Classification"],"prefix":"10.1007","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4576-9302","authenticated-orcid":false,"given":"Van Bach","family":"Nguyen","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6776-3868","authenticated-orcid":false,"given":"Christin","family":"Seifert","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3678-0390","authenticated-orcid":false,"given":"J\u00f6rg","family":"Schl\u00f6tterer","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,10,14]]},"reference":[{"key":"8_CR1","unstructured":"Bhattacharjee, A., Moraffah, R., Garland, J., Liu, H.: Towards LLM-guided causal explainability for black-box text classifiers. In: AAAI 2024 Workshop on Responsible Language Models, Vancouver, BC, Canada (2024)"},{"key":"8_CR2","doi-asserted-by":"publisher","unstructured":"Bowman, S.R., Angeli, G., Potts, C., Manning, C.D.: A large annotated corpus for learning natural language inference. In: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp. 632\u2013642. Association for Computational Linguistics, Lisbon, Portugal (2015). https:\/\/doi.org\/10.18653\/v1\/D15-1075, https:\/\/aclanthology.org\/D15-1075","DOI":"10.18653\/v1\/D15-1075"},{"key":"8_CR3","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown, T., et al.: Language models are few-shot learners. Adv. Neural. Inf. Process. Syst. 33, 1877\u20131901 (2020)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"key":"8_CR4","doi-asserted-by":"publisher","unstructured":"Calderon, N., Ben-David, E., Feder, A., Reichart, R.: DoCoGen: domain counterfactual generation for low resource domain adaptation. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 7727\u20137746. Association for Computational Linguistics, Dublin, Ireland (2022). https:\/\/doi.org\/10.18653\/v1\/2022.acl-long.533, https:\/\/aclanthology.org\/2022.acl-long.533","DOI":"10.18653\/v1\/2022.acl-long.533"},{"key":"8_CR5","doi-asserted-by":"publisher","unstructured":"Chen, H.T., Zhang, M., Choi, E.: Rich knowledge sources bring complex knowledge conflicts: Recalibrating models to reflect conflicting evidence. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 2292\u20132307. Association for Computational Linguistics, Abu Dhabi, United Arab Emirates (2022). https:\/\/doi.org\/10.18653\/v1\/2022.emnlp-main.146, https:\/\/aclanthology.org\/2022.emnlp-main.146","DOI":"10.18653\/v1\/2022.emnlp-main.146"},{"key":"8_CR6","doi-asserted-by":"publisher","unstructured":"Chen, Z., Gao, Q., Bosselut, A., Sabharwal, A., Richardson, K.: DISCO: distilling counterfactuals with large language models. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 5514\u20135528. Association for Computational Linguistics, Toronto, Canada (2023). https:\/\/doi.org\/10.18653\/v1\/2023.acl-long.302, https:\/\/aclanthology.org\/2023.acl-long.302","DOI":"10.18653\/v1\/2023.acl-long.302"},{"key":"8_CR7","doi-asserted-by":"crossref","unstructured":"Choi, Y., Choi, M., Kim, M., Ha, J.W., Kim, S., Choo, J.: StarGAN: unified generative adversarial networks for multi-domain image-to-image translation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8789\u20138797 (2018)","DOI":"10.1109\/CVPR.2018.00916"},{"key":"8_CR8","doi-asserted-by":"publisher","unstructured":"Dixit, T., Paranjape, B., Hajishirzi, H., Zettlemoyer, L.: CORE: a retrieve-then-edit framework for counterfactual data generation. In: Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 2964\u20132984. Association for Computational Linguistics, Abu Dhabi, United Arab Emirates (2022). https:\/\/doi.org\/10.18653\/v1\/2022.findings-emnlp.216, https:\/\/aclanthology.org\/2022.findings-emnlp.216","DOI":"10.18653\/v1\/2022.findings-emnlp.216"},{"key":"8_CR9","doi-asserted-by":"publisher","unstructured":"Fern, X., Pope, Q.: Text counterfactuals via latent optimization and Shapley-guided search. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 5578\u20135593. Association for Computational Linguistics, Online and Punta Cana, Dominican Republic (2021). https:\/\/doi.org\/10.18653\/v1\/2021.emnlp-main.452, https:\/\/aclanthology.org\/2021.emnlp-main.452","DOI":"10.18653\/v1\/2021.emnlp-main.452"},{"key":"8_CR10","doi-asserted-by":"publisher","unstructured":"Gardner, M., et al.: Evaluating models\u2019 local decision boundaries via contrast sets. In: Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 1307\u20131323. Association for Computational Linguistics, Online (2020). https:\/\/doi.org\/10.18653\/v1\/2020.findings-emnlp.117, https:\/\/aclanthology.org\/2020.findings-emnlp.117","DOI":"10.18653\/v1\/2020.findings-emnlp.117"},{"issue":"11","key":"8_CR11","doi-asserted-by":"publisher","first-page":"665","DOI":"10.1038\/s42256-020-00257-z","volume":"2","author":"R Geirhos","year":"2020","unstructured":"Geirhos, R., et al.: Shortcut learning in deep neural networks. Nat. Mach. Intell. 2(11), 665\u2013673 (2020)","journal-title":"Nat. Mach. Intell."},{"key":"8_CR12","unstructured":"Goodfellow, I.J., Shlens, J., Szegedy, C.: Explaining and harnessing adversarial examples. In: Bengio, Y., LeCun, Y. (eds.) 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, 7-9 May 2015, Conference Track Proceedings (2015). http:\/\/arxiv.org\/abs\/1412.6572"},{"key":"8_CR13","doi-asserted-by":"publisher","unstructured":"Guerreiro, N.M., Martins, A.F.T.: SPECTRA: sparse structured text rationalization. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 6534\u20136550. Association for Computational Linguistics, Online and Punta Cana, Dominican Republic (2021).https:\/\/doi.org\/10.18653\/v1\/2021.emnlp-main.525, https:\/\/aclanthology.org\/2021.emnlp-main.525","DOI":"10.18653\/v1\/2021.emnlp-main.525"},{"key":"8_CR14","unstructured":"Kaushik, D., Hovy, E., Lipton, Z.: Learning the difference that makes a difference with counterfactually-augmented data. In: International Conference on Learning Representations (2020). https:\/\/openreview.net\/forum?id=Sklgs0NFvr"},{"key":"8_CR15","first-page":"22199","volume":"35","author":"T Kojima","year":"2022","unstructured":"Kojima, T., Gu, S.S., Reid, M., Matsuo, Y., Iwasawa, Y.: Large language models are zero-shot reasoners. Adv. Neural. Inf. Process. Syst. 35, 22199\u201322213 (2022)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"key":"8_CR16","doi-asserted-by":"publisher","first-page":"265","DOI":"10.1016\/j.neuroimage.2013.01.060","volume":"72","author":"E Kulakova","year":"2013","unstructured":"Kulakova, E., Aichhorn, M., Schurz, M., Kronbichler, M., Perner, J.: Processing counterfactual and hypothetical conditionals: an FMRI investigation. Neuroimage 72, 265\u2013271 (2013)","journal-title":"Neuroimage"},{"key":"8_CR17","doi-asserted-by":"publisher","unstructured":"Longpre, S., Perisetla, K., Chen, A., Ramesh, N., DuBois, C., Singh, S.: Entity-based knowledge conflicts in question answering. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 7052\u20137063. Association for Computational Linguistics, Online and Punta Cana, Dominican Republic (2021). https:\/\/doi.org\/10.18653\/v1\/2021.emnlp-main.565, https:\/\/aclanthology.org\/2021.emnlp-main.565","DOI":"10.18653\/v1\/2021.emnlp-main.565"},{"key":"8_CR18","unstructured":"Lundberg, S.M., Lee, S.I.: A unified approach to interpreting model predictions. In: Proceedings of the 31st International Conference on Neural Information Processing Systems, pp. 4768\u20134777. NIPS 2017, Curran Associates Inc., Red Hook, NY, USA (2017)"},{"key":"8_CR19","unstructured":"Maas, A.L., Daly, R.E., Pham, P.T., Huang, D., Ng, A.Y., Potts, C.: Learning word vectors for sentiment analysis. In: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pp. 142\u2013150. Association for Computational Linguistics, Portland, Oregon, USA (2011). https:\/\/aclanthology.org\/P11-1015"},{"issue":"6","key":"8_CR20","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3457607","volume":"54","author":"N Mehrabi","year":"2021","unstructured":"Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., Galstyan, A.: A survey on bias and fairness in machine learning. ACM Comput. Surv. (CSUR) 54(6), 1\u201335 (2021)","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"8_CR21","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.artint.2018.07.007","volume":"267","author":"T Miller","year":"2019","unstructured":"Miller, T.: Explanation in artificial intelligence: Insights from the social sciences. Artif. Intell. 267, 1\u201338 (2019). https:\/\/doi.org\/10.1016\/j.artint.2018.07.007","journal-title":"Artif. Intell."},{"key":"8_CR22","doi-asserted-by":"crossref","unstructured":"Nguyen, V.B., Seifert, C., Schl\u00f6tterer, J.: CEval: a benchmark for evaluating counterfactual text generation. In: Mahamood, S., Minh, N.L., Ippolito, D. (eds.) Proceedings of the 17th International Natural Language Generation Conference, pp. 55\u201369. Association for Computational Linguistics, Tokyo, Japan (2024). https:\/\/aclanthology.org\/2024.inlg-main.6","DOI":"10.18653\/v1\/2024.inlg-main.6"},{"key":"8_CR23","doi-asserted-by":"crossref","unstructured":"Nguyen, V.B., Youssef, P., Schl\u00f6tterer, J., Seifert, C.: LLMS for generating and evaluating counterfactuals: a comprehensive study. In: Findings of the Association for Computational Linguistics: EMNLP 2024. Association for Computational Linguistics, Miami, USA (2024). https:\/\/arxiv.org\/html\/2405.00722v1","DOI":"10.18653\/v1\/2024.findings-emnlp.870"},{"issue":"8","key":"8_CR24","first-page":"9","volume":"1","author":"A Radford","year":"2019","unstructured":"Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al.: Language models are unsupervised multitask learners. OpenAI Blog 1(8), 9 (2019)","journal-title":"OpenAI Blog"},{"issue":"140","key":"8_CR25","first-page":"1","volume":"21","author":"C Raffel","year":"2020","unstructured":"Raffel, C., et al.: Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21(140), 1\u201367 (2020)","journal-title":"J. Mach. Learn. Res."},{"key":"8_CR26","doi-asserted-by":"publisher","unstructured":"Robeer, M., Bex, F., Feelders, A.: Generating realistic natural language counterfactuals. In: Findings of the Association for Computational Linguistics: EMNLP 2021, pp. 3611\u20133625. Association for Computational Linguistics, Punta Cana, Dominican Republic (2021). https:\/\/doi.org\/10.18653\/v1\/2021.findings-emnlp.306, https:\/\/aclanthology.org\/2021.findings-emnlp.306","DOI":"10.18653\/v1\/2021.findings-emnlp.306"},{"key":"8_CR27","doi-asserted-by":"publisher","unstructured":"Ross, A., Marasovi\u0107, A., Peters, M.: Explaining NLP models via minimal contrastive editing (MiCE). In: Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pp. 3840\u20133852. Association for Computational Linguistics, Online (2021). https:\/\/doi.org\/10.18653\/v1\/2021.findings-acl.336, https:\/\/aclanthology.org\/2021.findings-acl.336","DOI":"10.18653\/v1\/2021.findings-acl.336"},{"key":"8_CR28","unstructured":"Simonyan, K.: Deep inside convolutional networks: visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034 (2013)"},{"key":"8_CR29","doi-asserted-by":"publisher","unstructured":"Treviso, M., Ross, A., Guerreiro, N.M., Martins, A.: CREST: a joint framework for rationalization and counterfactual text generation. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15109\u201315126. Association for Computational Linguistics, Toronto, Canada (2023). https:\/\/doi.org\/10.18653\/v1\/2023.acl-long.842, https:\/\/aclanthology.org\/2023.acl-long.842","DOI":"10.18653\/v1\/2023.acl-long.842"},{"key":"8_CR30","doi-asserted-by":"publisher","unstructured":"Wen, J., Zhu, Y., Zhang, J., Zhou, J., Huang, M.: AutoCAD: automatically generate counterfactuals for mitigating shortcut learning. In: Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 2302\u20132317. Association for Computational Linguistics, Abu Dhabi, United Arab Emirates (2022). https:\/\/doi.org\/10.18653\/v1\/2022.findings-emnlp.170, https:\/\/aclanthology.org\/2022.findings-emnlp.170","DOI":"10.18653\/v1\/2022.findings-emnlp.170"},{"key":"8_CR31","doi-asserted-by":"publisher","unstructured":"Wu, T., Ribeiro, M.T., Heer, J., Weld, D.: PolyJuice: generating counterfactuals for explaining, evaluating, and improving models. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 6707\u20136723. Association for Computational Linguistics, Online (2021).https:\/\/doi.org\/10.18653\/v1\/2021.acl-long.523, https:\/\/aclanthology.org\/2021.acl-long.523","DOI":"10.18653\/v1\/2021.acl-long.523"},{"key":"8_CR32","doi-asserted-by":"publisher","unstructured":"Yu, H., Atanasova, P., Augenstein, I.: Revealing the parametric knowledge of language models: A unified framework for attribution methods. In: Ku, L.W., Martins, A., Srikumar, V. (eds.) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 8173\u20138186. Association for Computational Linguistics, Bangkok, Thailand (2024). https:\/\/doi.org\/10.18653\/v1\/2024.acl-long.444, https:\/\/aclanthology.org\/2024.acl-long.444","DOI":"10.18653\/v1\/2024.acl-long.444"}],"container-title":["Communications in Computer and Information Science","Explainable Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-032-08330-2_8","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T03:10:15Z","timestamp":1760325015000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-032-08330-2_8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,14]]},"ISBN":["9783032083296","9783032083302"],"references-count":32,"URL":"https:\/\/doi.org\/10.1007\/978-3-032-08330-2_8","relation":{},"ISSN":["1865-0929","1865-0937"],"issn-type":[{"value":"1865-0929","type":"print"},{"value":"1865-0937","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,14]]},"assertion":[{"value":"14 October 2025","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"xAI","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"World Conference on Explainable Artificial Intelligence","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Istanbul","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"T\u00fcrkiye","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2025","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"9 July 2025","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"11 July 2025","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"3","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"xai2025","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"https:\/\/xaiworldconference.com\/2025\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}}]}}