{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,19]],"date-time":"2026-08-19T09:24:47Z","timestamp":1787131487745,"version":"build-2736575974"},"publisher-location":"Cham","reference-count":23,"publisher":"Springer Nature Switzerland","isbn-type":[{"value":"9783032083234","type":"print"},{"value":"9783032083241","type":"electronic"}],"license":[{"start":{"date-parts":[[2025,10,16]],"date-time":"2025-10-16T00:00:00Z","timestamp":1760572800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,10,16]],"date-time":"2025-10-16T00:00:00Z","timestamp":1760572800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    This paper introduces an approach based on Counterfactual Shapley Values, which enhances explainability in reinforcement learning by integrating counterfactual analysis with Shapley Values. The approach aims to quantify and compare the contributions of different state dimensions to various action choices. To more accurately analyze the impacts of these contributions, we introduce new characteristic value functions, the\n                    <jats:italic>Counterfactual Difference based Characteristic Value functions<\/jats:italic>\n                    and the\n                    <jats:italic>Average Counterfactual Difference based Characteristic Value functions.<\/jats:italic>\n                    These functions help to evaluate the differences in contributions between optimal and non-optimal actions. Experiments across several RL domains, such as GridWorld, FrozenLake, and Taxi, demonstrate the effectiveness of the Counterfactual Shapley Values method. The results show that this method not only improves transparency in complex RL systems but also quantifies the differences across various decisions.\n                  <\/jats:p>","DOI":"10.1007\/978-3-032-08324-1_8","type":"book-chapter","created":{"date-parts":[[2025,10,15]],"date-time":"2025-10-15T08:50:06Z","timestamp":1760518206000},"page":"169-193","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Counterfactual Shapley Values for\u00a0Explaining Reinforcement Learning"],"prefix":"10.1007","author":[{"given":"Yiwei","family":"Shi","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weiru","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,10,16]]},"reference":[{"issue":"6","key":"8_CR1","doi-asserted-by":"publisher","first-page":"5068","DOI":"10.1109\/TITS.2020.3046646","volume":"23","author":"J Chen","year":"2021","unstructured":"Chen, J., Li, S.E., Tomizuka, M.: Interpretable end-to-end urban autonomous driving with latent deep reinforcement learning. IEEE Trans. Intell. Transp. Syst. 23(6), 5068\u20135078 (2021)","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"8_CR2","doi-asserted-by":"crossref","unstructured":"Shi, Y., Wen, M., Zhang, Q., Zhang, W., Liu, C., Liu, W.: Autonomous goal detection and cessation in reinforcement learning: a case study on source term estimation. arXiv preprint arXiv:2409.09541 (2024)","DOI":"10.1609\/aaai.v39i1.32056"},{"key":"8_CR3","doi-asserted-by":"crossref","unstructured":"Shi, Y., McAreavey, K., Liu, C., Liu, W.: Reinforcement learning for source location estimation: a multi-step approach. In: 2024 IEEE International Conference on Industrial Technology (ICIT), pp. 1\u20138. IEEE (2024)","DOI":"10.1109\/ICIT58233.2024.10540860"},{"key":"8_CR4","unstructured":"Liu, S., Ngiam, K.Y., Feng, M.: Deep reinforcement learning for clinical decision support: a brief survey. arXiv preprint arXiv:1907.09475 (2019)"},{"key":"8_CR5","unstructured":"Jiang, Z., Xu, D., Liang, J.: A deep reinforcement learning framework for the financial portfolio management problem. arXiv preprint arXiv:1706.10059 (2017)"},{"issue":"3","key":"8_CR6","doi-asserted-by":"publisher","first-page":"2411","DOI":"10.1109\/TITS.2021.3095161","volume":"23","author":"N Kumar","year":"2021","unstructured":"Kumar, N., Mittal, S., Garg, V., Kumar, N.: Deep reinforcement learning-based traffic light scheduling framework for SDN-enabled smart transportation system. IEEE Trans. Intell. Transp. Syst. 23(3), 2411\u20132421 (2021)","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"8_CR7","doi-asserted-by":"crossref","unstructured":"Ma, H., McAreavey, K., Liu, W.: TSFeatLIME: an online user study in enhancing explainability in univariate time series forecasting. In: 2024 IEEE 36th International Conference on Tools with Artificial Intelligence (ICTAI), pp. 578\u2013585. IEEE (2024)","DOI":"10.1109\/ICTAI62512.2024.00087"},{"key":"8_CR8","unstructured":"Lundberg, S.M., Lee, S.-I.: A unified approach to interpreting model predictions. In: Advances in Neural Information Processing Systems, vol.\u00a030 (2017)"},{"key":"8_CR9","doi-asserted-by":"crossref","unstructured":"Wang, J., Zhang, Y., Gu, Y., Kim, T.-K.: SHAQ: incorporating Shapley value theory into multi-agent Q-learning. In: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A., Eds., Advances in Neural Information Processing Systems, vol.\u00a035, pp. 5941\u20135954. Curran Associates, Inc. (2022). https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/file\/27985d21f0b751b933d675930aa25022-Paper-Conference.pdf","DOI":"10.52202\/068431-0430"},{"issue":"05","key":"8_CR10","first-page":"7285","volume":"34","author":"J Wang","year":"2020","unstructured":"Wang, J., Zhang, Y., Kim, T.-K., Gu, Y.: Shapley Q-value: a local reward approach to solve global reward games. Proc. AAAI Conf. Artif. Intell. 34(05), 7285\u20137292 (2020)","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"issue":"1","key":"8_CR11","doi-asserted-by":"publisher","first-page":"59","DOI":"10.1109\/MCI.2021.3129959","volume":"17","author":"A Heuillet","year":"2022","unstructured":"Heuillet, A., Couthouis, F., D\u00edaz-Rodr\u00edguez, N.: Collective explainable AI: explaining cooperative strategies and agent contribution in multiagent reinforcement learning with Shapley values. IEEE Comput. Intell. Mag. 17(1), 59\u201371 (2022)","journal-title":"IEEE Comput. Intell. Mag."},{"key":"8_CR12","unstructured":"Beechey, D., Smith, T.M., \u015eim\u015fek, \u00d6.: Explaining reinforcement learning with Shapley values. In: International Conference on Machine Learning, pp. 2003\u20132014. PMLR (2023)"},{"key":"8_CR13","doi-asserted-by":"crossref","unstructured":"Olson, M.L., Khanna, R., Neal, L., Li, F., Wong, W.-K.: Counterfactual state explanations for reinforcement learning agents via generative deep learning. ArXiv abs\/2101.12446 (2021)","DOI":"10.1016\/j.artint.2021.103455"},{"key":"8_CR14","doi-asserted-by":"crossref","unstructured":"Li, J., Kuang, K., Wang, B., Liu, F., Chen, L., Wu, F., Xiao, J.: Shapley counterfactual credits for multi-agent reinforcement learning. In: Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pp. 934\u2013942 (2021)","DOI":"10.1145\/3447548.3467420"},{"key":"8_CR15","doi-asserted-by":"crossref","unstructured":"Albini, E., Long, J., Dervovic, D., Magazzeni, D.: Counterfactual Shapley additive explanations. In: Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pp. 1054\u20131070 (2022)","DOI":"10.1145\/3531146.3533168"},{"key":"8_CR16","doi-asserted-by":"crossref","unstructured":"Shapley, L.S.: A value for n-person games. In: Contributions to the Theory of Games. Princeton University Press, pp. 307\u2013317 (1953)","DOI":"10.1515\/9781400881970-018"},{"key":"8_CR17","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511528446","volume-title":"The Shapley Value: Essays in honor of Lloyd S","author":"AE Roth","year":"1988","unstructured":"Roth, A.E.: The Shapley Value: Essays in honor of Lloyd S. Cambridge University Press, Shapley (1988)"},{"key":"8_CR18","doi-asserted-by":"crossref","unstructured":"Shapley, L.S.: A value for n-person games. In: Contribution to the Theory of Games, vol.\u00a02 (1953)","DOI":"10.1515\/9781400881970-018"},{"key":"8_CR19","doi-asserted-by":"crossref","unstructured":"Winter, E.: The Shapley Value. In: Handbook of Game Theory with Economic Applications, vol. 3, pp. 2025\u20132054 (2002)","DOI":"10.1016\/S1574-0005(02)03016-3"},{"key":"8_CR20","unstructured":"Peleg, B., Sudh\u00f6lter, P.: Introduction to the Theory of Cooperative Games, vol.\u00a034. Springer Science & Business Media (2007)"},{"key":"8_CR21","unstructured":"Brockman, G., et al.: OpenAI gym. arXiv preprint arXiv:1606.01540 (2016)"},{"key":"8_CR22","unstructured":"Dietterich, T.G.: The taxi problem: a case study in reinforcement learning. Five Open Problems in Reinforcement Learning (1998)"},{"key":"8_CR23","unstructured":"Towers, M.: Gymnasium (2023). https:\/\/zenodo.org\/record\/8127025"}],"container-title":["Communications in Computer and Information Science","Explainable Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-032-08324-1_8","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,19]],"date-time":"2026-08-19T08:30:23Z","timestamp":1787128223000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-032-08324-1_8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,16]]},"ISBN":["9783032083234","9783032083241"],"references-count":23,"URL":"https:\/\/doi.org\/10.1007\/978-3-032-08324-1_8","relation":{},"ISSN":["1865-0929","1865-0937"],"issn-type":[{"value":"1865-0929","type":"print"},{"value":"1865-0937","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,16]]},"assertion":[{"value":"16 October 2025","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"xAI","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"World Conference on Explainable Artificial Intelligence","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Istanbul","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"T\u00fcrkiye","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2025","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"9 July 2025","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"11 July 2025","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"3","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"xai2025","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"https:\/\/xaiworldconference.com\/2025\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}}]}}