{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,11]],"date-time":"2026-04-11T03:48:59Z","timestamp":1775879339972,"version":"3.50.1"},"reference-count":32,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2023,8,9]],"date-time":"2023-08-09T00:00:00Z","timestamp":1691539200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2023,8,9]],"date-time":"2023-08-09T00:00:00Z","timestamp":1691539200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"funder":[{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["2105007"],"award-info":[{"award-number":["2105007"]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["2105007"],"award-info":[{"award-number":["2105007"]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["1950491"],"award-info":[{"award-number":["1950491"]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["2105007"],"award-info":[{"award-number":["2105007"]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["1950491"],"award-info":[{"award-number":["1950491"]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["2105007"],"award-info":[{"award-number":["2105007"]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Auton Agent Multi-Agent Syst"],"published-print":{"date-parts":[[2023,12]]},"DOI":"10.1007\/s10458-023-09615-8","type":"journal-article","created":{"date-parts":[[2023,8,9]],"date-time":"2023-08-09T13:01:45Z","timestamp":1691586105000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["A novel policy-graph approach with natural language and counterfactual abstractions for explaining reinforcement learning agents"],"prefix":"10.1007","volume":"37","author":[{"given":"Tongtong","family":"Liu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joe","family":"McCalmon","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Thai","family":"Le","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Md Asifur","family":"Rahman","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dongwon","family":"Lee","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sarra","family":"Alqahtani","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,8,9]]},"reference":[{"key":"9615_CR1","doi-asserted-by":"publisher","first-page":"283","DOI":"10.1016\/S0921-8890(97)00043-2","volume":"22","author":"H Benbrahim","year":"1997","unstructured":"Benbrahim, H., & Franklin, J. A. (1997). Biped dynamic walking using reinforcement learning. Robotics and Autonomous Systems, 22, 283\u2013302.","journal-title":"Robotics and Autonomous Systems"},{"key":"9615_CR2","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., & Riedmiller, M. (2013). Playing Atari with deep reinforcement learning. arXiv preprint arXiv:1312.5603."},{"key":"9615_CR3","unstructured":"Mirhoseini, A., Goldie, A., Yazgan, M., Jiang, J., Songhori, E., Wang, S., Lee, Y.-J., Johnson, E., Pathak, O., Bae, S., Nazi, A., Pak, J., Tong, A., Srinivasa, K., Hang, W., Tuncer, E., Babu, A., Le, Q.V., Laudon, J., Ho, R., Carpenter, R., & Dean, J. (2020). Chip placement with deep reinforcement learning. https:\/\/arxiv.org\/pdf\/2004.10746.pdf."},{"key":"9615_CR4","doi-asserted-by":"publisher","unstructured":"Liu, N., Li, Z., Xu, J., Xu, Z., Lin, S., Qiu, Q., Tang, J., & Wang, Y. (2017). A hierarchical framework of cloud resource allocation and power management using deep reinforcement learning, pp. 372\u2013382 https:\/\/doi.org\/10.1109\/ICDCS.2017.123.","DOI":"10.1109\/ICDCS.2017.123"},{"key":"9615_CR5","doi-asserted-by":"publisher","unstructured":"Peters, J., & Schaal, S. (2006). Policy gradient methods for robotics. In 2006 IEEE\/RSJ International Conference on Intelligent Robots and Systems (pp. 2219\u20132225). https:\/\/doi.org\/10.1109\/IROS.2006.282564.","DOI":"10.1109\/IROS.2006.282564"},{"key":"9615_CR6","doi-asserted-by":"crossref","unstructured":"Huang, S. H., Bhatia, K., Abbeel, P., & Dragan, A. D. (2018). Establishing appropriate trust via critical states. In 2018 IEEE\/RSJ international conference on intelligent robots and systems (IROS).","DOI":"10.1109\/IROS.2018.8593649"},{"key":"9615_CR7","doi-asserted-by":"crossref","unstructured":"Hayes, B., & Shah, J. (2017). Improving robot controller transparency through autonomous policy explanation. In 2017 12th ACM\/IEEE international conference on human\u2013robot interaction (HRI) (pp. 303\u2013312).","DOI":"10.1145\/2909824.3020233"},{"key":"9615_CR8","first-page":"1899","volume-title":"Proceedings of the 33rd international conference on machine learning. Proceedings of machine learning research","author":"T Zahavy","year":"2016","unstructured":"Zahavy, T., Ben-Zrihem, N., & Mannor, S. (2016). Graying the black box: Understanding dqns. In M. .F. Balcan & K. .Q. Weinberger (Eds.), Proceedings of the 33rd international conference on machine learning. Proceedings of machine learning research (Vol. 48, pp. 1899\u20131908). PMLR."},{"key":"9615_CR9","doi-asserted-by":"crossref","unstructured":"Topin, N., & Veloso, M. (2019). Generation of policy-level explanations for reinforcement learning. In The thirty-third AAAI conference on artificial intelligence, AAAI 2019 (pp. 2514\u20132521). AAAI Press. https:\/\/aaai.org\/ojs\/index.php\/AAAI\/article\/view\/4097.","DOI":"10.1609\/aaai.v33i01.33012514"},{"key":"9615_CR10","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2021.103455","volume":"295","author":"ML Olson","year":"2021","unstructured":"Olson, M. L., Khanna, R., Neal, L., Li, F., & Wong, W.-K. (2021). Counterfactual state explanations for reinforcement learning agents via generative deep learning. Artificial Intelligence, 295, 103455. https:\/\/doi.org\/10.1016\/j.artint.2021.103455","journal-title":"Artificial Intelligence"},{"key":"9615_CR11","unstructured":"Juozapaitis, Z., Koul, A., Fern, A., Erwig, M., & Doshi-Velez, F. (2019) Explainable reinforcement learning via reward decomposition. In Proceedings at the international joint conference on artificial intelligence."},{"key":"9615_CR12","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2020.103655","volume":"113","author":"JK Aniek Markus","year":"2021","unstructured":"Aniek Markus, J. K., & Rijnbeek, P. (2021). The role of explainability in creating trustworthy artificial intelligence for health care: A comprehensive survey of the terminology, design choices, and evaluation strategies. Journal of Biomedical Informatics, 113, 103655. https:\/\/doi.org\/10.1016\/j.jbi.2020.103655","journal-title":"Journal of Biomedical Informatics"},{"key":"9615_CR13","doi-asserted-by":"crossref","unstructured":"Madumal, P., Miller, T., Sonenberg, L., & Vetere, F. (2020). Explainable reinforcement learning through a causal lens. In The thirty-fourth AAAI conference on artificial intelligence, AAAI 2020 (pp. 2493\u20132500). AAAI Press. https:\/\/aaai.org\/ojs\/index.php\/AAAI\/article\/view\/5631.","DOI":"10.1609\/aaai.v34i03.5631"},{"key":"9615_CR14","doi-asserted-by":"publisher","unstructured":"Liu, B., Xia, Y., & Yu, P. S. (2004). Clustering via decision tree construction. In Foundations and advances in data mining. Studies in fuzziness and soft computing (Vol. 180). Springer, Berlin, Heidelberg. https:\/\/doi.org\/10.1007\/11362197_5.","DOI":"10.1007\/11362197_5"},{"key":"9615_CR15","unstructured":"Schulman, J., Wolski, F., Radford, A., Dhariwal, P., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347"},{"key":"9615_CR16","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., & Riedmiller, M. (2013). Playing Atari with deep reinforcement learning. In arXiv preprint arXiv:1312.5602"},{"key":"9615_CR17","doi-asserted-by":"publisher","unstructured":"Liu, Guiliang & Schulte, Oliver & Zhu, Wang & Li, Qingcan, 2018. Toward interpretable deep reinforcement learning with linear model U-trees: European Conference, ECML PKDD 2018, Dublin, Ireland, September 10\u201314, 2018, Proceedings, Part II. https:\/\/doi.org\/10.1007\/978-3-030-10928-8_25.","DOI":"10.1007\/978-3-030-10928-8_25"},{"key":"9615_CR18","unstructured":"van der Waa, J., van Diggelen, J., van den Bosch, K., & Neerincx, M. (2018). Contrastive explanations for reinforcement learning in terms of expected consequences."},{"key":"9615_CR19","doi-asserted-by":"crossref","unstructured":"Iyer, R. R., Li, Y., Li, H., Lewis, M., Sundar, R., & Sycara, K. P. (2018). Transparency and explanation in deep reinforcement learning neural networks. In Proceedings of the 2018 AAAI\/ACM conference on AI, ethics, and society.","DOI":"10.1145\/3278721.3278776"},{"key":"9615_CR20","unstructured":"Yang, Z., Bai, S., Zhang, L., & Torr, P. H. S. (2019). Learn to interpret Atari agents."},{"key":"9615_CR21","first-page":"1792","volume-title":"Proceedings of the 35th international conference on machine learning. Proceedings of machine learning research","author":"S Greydanus","year":"2018","unstructured":"Greydanus, S., Koul, A., Dodge, J., & Fern, A. (2018). Visualizing and understanding Atari agents. In J. Dy & A. Krause (Eds.), Proceedings of the 35th international conference on machine learning. Proceedings of machine learning research (Vol. 80, pp. 1792\u20131801). PMLR."},{"key":"9615_CR22","unstructured":"Amir, D., & Amir, O. (2018) Highlights: Summarizing agent behavior to people. In AAMAS."},{"key":"9615_CR23","unstructured":"McCalmon, J., Le, T., Alqahtani, S., & Lee, D. (2022) Caps: Comprehensible abstract policy summaries for explaining reinforcement learning agents. In Proceedings of the 21st international conference on autonomous agents and multiagent systems. AAMAS \u201922 (pp. 889\u2013897). International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC."},{"key":"9615_CR24","doi-asserted-by":"publisher","first-page":"181","DOI":"10.1016\/S0004-3702(99)00052-1","volume":"112","author":"RS Sutton","year":"1999","unstructured":"Sutton, R. S., Precup, D., & Singh, S. (1999). Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning. Artificial Intelligence, 112, 181\u2013211.","journal-title":"Artificial Intelligence"},{"key":"9615_CR25","unstructured":"Moore, A. (1995). Variable resolution reinforcement learning. CMU-RI-TR-95-19. https:\/\/apps.dtic.mil\/sti\/tr\/pdf\/ADA311507.pdf."},{"key":"9615_CR26","doi-asserted-by":"publisher","first-page":"59","DOI":"10.1080\/00031305.1988.10475524","volume":"42","author":"J Rodgers","year":"1988","unstructured":"Rodgers, J., & Nicewander, A. (1988). Thirteen ways to look at the correlation coefficient. American Statistician, 42, 59\u201366. https:\/\/doi.org\/10.1080\/00031305.1988.10475524","journal-title":"American Statistician"},{"key":"9615_CR27","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.artint.2018.07.007","volume":"267","author":"T Miller","year":"2019","unstructured":"Miller, T. (2019). Explanation in artificial intelligence: Insights from the social sciences. Artif. Intell., 267, 1\u201338.","journal-title":"Artif. Intell."},{"key":"9615_CR28","unstructured":"Goyal, Y., Wu, Z., Ernst, J., Batra, D., Parikh, D., & Lee, S. (2019). Counterfactual visual explanations. In ICML (pp. 2376\u20132384). http:\/\/proceedings.mlr.press\/v97\/goyal19a.html."},{"key":"9615_CR29","doi-asserted-by":"publisher","unstructured":"Uesato, J., Kumar, A., Szepesvari, C., Erez, T., Ruderman, A., Anderson, K., Heess, N., & Kohli, P. (2018). Rigorous agent evaluation: An adversarial approach to uncover catastrophic failures. arXiv. https:\/\/doi.org\/10.48550\/ARXIV.1812.01647. arXiv:1812.01647.","DOI":"10.48550\/ARXIV.1812.01647"},{"key":"9615_CR30","unstructured":"Abolfathi, E. A., Luo, J., Yadmellat, P., & Rezaee, K. (2021). Coachnet: An adversarial sampling approach for reinforcement learning. arXiv preprint arXiv:2101.02649."},{"key":"9615_CR31","doi-asserted-by":"publisher","unstructured":"van der Waa, J., van Diggelen, J., Bosch, K. V. D., & Mark, N. (2018). Contrastive explanations for reinforcement learning in terms of expected consequences. https:\/\/doi.org\/10.48550\/ARXIV.1807.08706.","DOI":"10.48550\/ARXIV.1807.08706"},{"key":"9615_CR32","unstructured":"Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., & Zaremba, W. (2016). OpenAI Gym. arXiv preprint arXiv:1606.01540."}],"container-title":["Autonomous Agents and Multi-Agent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10458-023-09615-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10458-023-09615-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10458-023-09615-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,12,17]],"date-time":"2023-12-17T23:50:24Z","timestamp":1702857024000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10458-023-09615-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,9]]},"references-count":32,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,12]]}},"alternative-id":["9615"],"URL":"https:\/\/doi.org\/10.1007\/s10458-023-09615-8","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-2409910\/v1","asserted-by":"object"}]},"ISSN":["1387-2532","1573-7454"],"issn-type":[{"value":"1387-2532","type":"print"},{"value":"1573-7454","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,9]]},"assertion":[{"value":"9 July 2023","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 August 2023","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this article.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Our user studies were approved by the Institutional Review Board at Wake Forest University with the number IRB00024657. The approved IRB followed the <i>Exemption Category 3<\/i>: Research involving benign behavioral interventions is conjunction with the collection of information for an adult subject through verbal or written responses (including data entry) or audiovisual recording if the subject prospectively agrees to the intervention and information collected. (A) The information obtained is recorded by the investigator in such a manner that the identity of the human subjects cannot readily be ascertained, directly or through identifiers linked to the subjects; (B) Any disclosure of the human subjects\u2019 responses outside the research would not reasonably place the subjects at risk of criminal liability or be damaging to the subjects\u2019 financial standing, employability, educational advancement, or reputation; or (C) The information is recorded by the investigator in such a manner that the identity of the human subjects can readily be ascertained, directly or through identifiers linked to the subjects.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}},{"value":"You are invited to participate in a research study explaining the behavior of artificial intelligence agents. We are investigating whether summarizing the agent\u2019s behavior using English improves the end user\u2019s understanding and increases their trust in the agent\u2019s behavior. In this study, you will complete several questionnaires to measure your understanding of the agent\u2019s behavior. You may discontinue your participation at any time without penalty by closing your browser window. Any responses entered to that point will be deleted. You may also choose not to answer any question(s) you do not wish to answer for any reason. You may choose to skip any question(s) for any reason. We encourage you to print or save a copy of this page for your records (or future reference). By clicking on \u201cI agree\u201d, you indicate that you are at least 18 years old and that you agree to participate in this research project. You will advance to the experiment. If you do not wish to participate, please close your browser window. Completing the experiment should take about 15\u00a0min. You will earn 20 cents for each minute you spend on the experiment (there will be a time limit for each question) and you will get 10 cents extra for each correct answer.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Participant consent"}}],"article-number":"34"}}