{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T03:00:31Z","timestamp":1780974031798,"version":"3.54.1"},"reference-count":45,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T00:00:00Z","timestamp":1780963200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T00:00:00Z","timestamp":1780963200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Auton Agent Multi-Agent Syst"],"published-print":{"date-parts":[[2026,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Ensuring the safety of reinforcement learning (RL) policies in high-stakes environments requires more than formal verification: it needs interpretability and targeted falsification\u2014the deliberate search for counter-examples that expose potential failures before deployment. We present AEGIS-RL (Abstract, Explainable Graphs for Integrated Safety in RL), a hybrid framework that unifies (1) explainable RL, (2) probabilistic model checking, and (3) risk-guided falsification, and augments them with (4) a lightweight runtime safety shield that switches to a fallback policy when estimated risk exceeds a threshold. AEGIS-RL first builds a directed, semantically meaningful graph from offline trajectories that blends local and global explanations to make policy behavior transparent and verifier-friendly. This abstract graph is fed to a probabilistic model checker (e.g., Storm) to verify temporal safety specifications; when violations exist, the checker returns interpretable counterexample traces that pinpoint how the policy fails. When specifications appear satisfied, AEGIS-RL estimates residual risk during checking to steer falsification toward high-risk, under-explored states, broadening coverage beyond the offline data. Across safety-critical benchmarks including two MuJoCo tasks and a medical insulin-dosing scenario; AEGIS-RL uncovers significantly more violations than uncertainty- and fuzzing-based baselines and yields a broader, more novel set of failure trajectories. The resulting explanations and counterexamples provide actionable guidance to understand, debug, and repair unsafe policies while enabling runtime mitigation without retraining.<\/jats:p>","DOI":"10.1007\/s10458-026-09749-5","type":"journal-article","created":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T02:52:29Z","timestamp":1780973549000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["AEGIS-RL: Abstract, Explainable Graphs for Integrated Safety in RL"],"prefix":"10.1007","volume":"40","author":[{"given":"Tuan","family":"Le","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Risal","family":"Shefin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Debashis","family":"Gupta","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thai","family":"Le","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sarra","family":"Alqahtani","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,6,9]]},"reference":[{"key":"9749_CR1","unstructured":"Kim, D., Moon, S., Hostallero, D.E., Kang, W.J., Lee, T., Son, K., & Yi, Y. (2020). Safe reinforcement learning: A survey. arXiv:2005.00903"},{"key":"9749_CR2","doi-asserted-by":"crossref","unstructured":"Geibel, P. (2006). Reinforcement learning for MDPs with constraints. European Conference on Machine Learning, pp. 646\u2013653.","DOI":"10.1007\/11871842_63"},{"issue":"3","key":"9749_CR3","doi-asserted-by":"publisher","first-page":"4915","DOI":"10.1109\/LRA.2021.3070252","volume":"6","author":"B Thananjeyan","year":"2021","unstructured":"Thananjeyan, B., Balakrishna, A., Nair, S., Luo, M., Srinivasan, K., Hwang, M., Gonzalez, J. E., Ibarz, J., Finn, C., & Goldberg, K. (2021). Recovery RL: Safe Reinforcement Learning with Learned Recovery Zones. IEEE Robotics and Automation Letters, 6(3), 4915\u20134922.","journal-title":"IEEE Robotics and Automation Letters"},{"issue":"17","key":"9749_CR4","first-page":"15102","volume":"35","author":"O Bastani","year":"2021","unstructured":"Bastani, O. (2021). Safe Reinforcement Learning with Formal Guarantees. Proceedings of the AAAI Conference on Artificial Intelligence, 35(17), 15102\u201315110.","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"9749_CR5","doi-asserted-by":"crossref","unstructured":"Alshiekh, M., Bloem, R., Ehlers, R., K\u00f6nighofer, B., Niekum, S., & Topcu, U. (2018). Safe Reinforcement Learning via Shielding. In Proceedings of the AAAI Conference on Artificial Intelligence, 32(1).","DOI":"10.1609\/aaai.v32i1.11797"},{"issue":"2","key":"9749_CR6","first-page":"267","volume":"49","author":"O Mihatsch","year":"2002","unstructured":"Mihatsch, O., & Neuneier, R. (2002). Risk-sensitive reinforcement learning. Machine learning, 49(2), 267\u2013290.","journal-title":"Risk-sensitive reinforcement learning. Machine learning"},{"key":"9749_CR7","unstructured":"Ruggeri, F., Russo, A., Inam, R., & Johansson, K.H. (2025). Explainable reinforcement learning via temporal policy decomposition. arXiv:2501.03902"},{"key":"9749_CR8","unstructured":"Corsi, D., Marchesini, E., & Farinelli, A. (2021). Formal verification of neural networks for safety-critical tasks in deep reinforcement learning. In Proceedings of the 37th Conference on Uncertainty in Artificial Intelligence."},{"key":"9749_CR9","unstructured":"Cheng, Z., Yu, J., & Xing, X. (2025). A survey on explainable deep reinforcement learning. arXiv:2502.06869"},{"key":"9749_CR10","unstructured":"Deshmukh, J.V., Jin, X., Kapinski, J., Ueda, K., & Butts, M. (2015). Stochastic search techniques for reachability analysis of markov decision processes. In Tools and Algorithms for the Construction and Analysis of Systems, pp. 211\u2013229. Springer."},{"key":"9749_CR11","unstructured":"Karunakaran, P., & Seshia, S.A. (2020). Counterexample-guided reinforcement learning with model-based exploration. In Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"9749_CR12","doi-asserted-by":"crossref","unstructured":"Shefin, R.S., Rahman, M.A., Le, T., & Alqahtani, S. (2025). xSRL: Safety-aware explainable reinforcement learning-safety as a product of explainability. In Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems, pp. 1932\u20131940.","DOI":"10.65109\/DGKF6931"},{"key":"9749_CR13","doi-asserted-by":"crossref","unstructured":"McCalmon, J., Le, T., Alqahtani, S., & Lee, D. (2022). CAPS: Comprehensible abstract policy summaries for explaining reinforcement learning agents. In Proceedings of the 21st International Conference on Autonomous Agents and Multiagent Systems. AAMAS \u201922, pp. 889\u2013897. International Foundation for Autonomous Agents and Multiagent Systems, Richland, SC.","DOI":"10.65109\/HUQL1884"},{"key":"9749_CR14","unstructured":"Molnar, C. (2020). Interpretable Machine Learning, 1st edn. https:\/\/christophm.github.io\/interpretable-ml-book"},{"key":"9749_CR15","doi-asserted-by":"crossref","unstructured":"Hayes, B., & Shah, J.A. (2017). Improving robot controller transparency through autonomous policy explanation. In Proceedings of the 2017 ACM\/IEEE International Conference on Human-robot Interaction, pp. 303\u2013312.","DOI":"10.1145\/2909824.3020233"},{"key":"9749_CR16","doi-asserted-by":"crossref","unstructured":"Topin, N., & Veloso, M. (2019). Generation of policy-level explanations for reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, pp. 2514\u20132521.","DOI":"10.1609\/aaai.v33i01.33012514"},{"key":"9749_CR17","doi-asserted-by":"crossref","unstructured":"Amir, D., & Amir, O. (2018). Highlights: Summarizing agent behavior to people. In Proceedings of the 17th International Conference on Autonomous Agents and Multiagent Systems, pp. 1168\u20131176.","DOI":"10.65109\/WIZC1785"},{"key":"9749_CR18","unstructured":"Zahavy, T., Ben-Zrihem, N., & Mannor, S. (2016). Graying the black box: Understanding DQNs. In M.F., Balcan, & K.Q. Weinberger (Eds.), Proceedings of the 33rd International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 48, pp. 1899\u20131908. PMLR, New York, New York, USA. https:\/\/proceedings.mlr.press\/v48\/zahavy16.html"},{"key":"9749_CR19","doi-asserted-by":"publisher","DOI":"10.3389\/frai.2022.903875","volume":"5","author":"T Huber","year":"2022","unstructured":"Huber, T., Limmer, B., & Andr\u00e9, E. (2022). Benchmarking perturbation-based saliency maps for explaining atari agents. Frontiers in Artificial Intelligence, 5, Article 903875.","journal-title":"Frontiers in Artificial Intelligence"},{"key":"9749_CR20","unstructured":"Puri, N., Verma, S., Gupta, P., Kayastha, D., Deshmukh, S., Krishnamurthy, B., & Singh, S. (2019). Explain your move: Understanding agent actions using specific and relevant feature attribution. arXiv:1912.12191"},{"key":"9749_CR21","unstructured":"Juozapaitis, Z., Koul, A., Fern, A., Erwig, M., & Doshi-Velez, F. (2019). Explainable reinforcement learning via reward decomposition. In IJCAI\/ECAI Workshop on Explainable Artificial Intelligence."},{"key":"9749_CR22","doi-asserted-by":"crossref","unstructured":"Kwiatkowska, M., Norman, G., & Parker, D. (2011). Prism 4.0: Verification of probabilistic real-time systems. In International Conference on Computer Aided Verification (CAV), pp. 585\u2013591. Springer, Snowbird, USA.","DOI":"10.1007\/978-3-642-22110-1_47"},{"key":"9749_CR23","doi-asserted-by":"crossref","unstructured":"Dehnert, C., Junges, S., Katoen, J.-P., & Volk, M. (2017). A storm is coming: A modern probabilistic model checker. In International Conference on Computer Aided Verification, pp. 592\u2013600. Springer.","DOI":"10.1007\/978-3-319-63390-9_31"},{"issue":"2","key":"9749_CR24","first-page":"3","volume":"1","author":"L Li","year":"2006","unstructured":"Li, L., Walsh, T. J., & Littman, M. L. (2006). Towards a unified theory of state abstraction for MDPs. AI&M, 1(2), 3.","journal-title":"AI&M"},{"key":"9749_CR25","volume-title":"Principles of Model Checking","author":"C Baier","year":"2008","unstructured":"Baier, C., & Katoen, J.-P. (2008). Principles of Model Checking (Vol. 26202649). Cambridge, Massachusetts: The MIT Press."},{"key":"9749_CR26","unstructured":"Hahn, E.M., Hartmanns, A., Hensel, C., Klauck, M., Kretinsky, J., Parker, D., Quatmann, T., Ruijters, E., & Steinmetz, M. (2019). Storm: A modern probabilistic model checker. In International Conference on Computer Aided Verification, pp. 592\u2013600. Springer, Cham."},{"key":"9749_CR27","doi-asserted-by":"crossref","unstructured":"Liu, T., McCalmon, J., Rahman, A., Lee, T., Le, D., & Alqahtani, S. (2023). A policy-graph approach to explain reinforcement learning agents: A novel policy-graph approach with natural language and counterfactual abstractions for explaining reinforcement learning agents. Autonomous Agents and Multi-Agent Systems Journal.","DOI":"10.21203\/rs.3.rs-2409910\/v1"},{"key":"9749_CR28","doi-asserted-by":"publisher","unstructured":"Liu, B., Xia, Y., Yu, & P.S. (2000). Clustering through decision tree construction. In Proceedings of the 9th International Conference on Information and Knowledge Management. CIKM \u201900, pp. 20\u201329. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/354756.354775, https:\/\/doi.org\/10.1145\/354756.354775","DOI":"10.1145\/354756.354775"},{"key":"9749_CR29","doi-asserted-by":"crossref","unstructured":"Topin, N., & Veloso, M. (2019). Generation of policy-level explanations for reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, pp. 2514\u20132521.","DOI":"10.1609\/aaai.v33i01.33012514"},{"key":"9749_CR30","unstructured":"Bekkemoen, Y., & Langseth, H. (2024). ASAP: attention-based state space abstraction for policy summarization. In Asian Conference on Machine Learning, pp. 137\u2013152. PMLR."},{"key":"9749_CR31","doi-asserted-by":"crossref","unstructured":"Peng, S., Liu, S., Zhi, D., Wang, P., Xu, C., Chen, C., & Zhang, M. (2025). ATA: An abstract-train-abstract approach for explanation-friendly deep reinforcement learning. Neural Networks, 107749.","DOI":"10.1016\/j.neunet.2025.107749"},{"key":"9749_CR32","doi-asserted-by":"crossref","unstructured":"Rahman, M.A., Liu, T., & Alqahtani, S.M. (2023). Adversarial behavior exclusion for safe reinforcement learning. In IJCAI, pp. 483\u2013491.","DOI":"10.24963\/ijcai.2023\/54"},{"key":"9749_CR33","unstructured":"Settles, B. (2011). From theories to queries: Active learning in practice. In I. Guyon, G. Cawley, G. Dror, V. Lemaire, A. Statnikov (Eds.), Active Learning and Experimental Design Workshop In Conjunction with AISTATS 2010. Proceedings of Machine Learning Research, vol. 16, pp. 1\u201318. PMLR, Sardinia, Italy. https:\/\/proceedings.mlr.press\/v16\/settles11a.html"},{"key":"9749_CR34","doi-asserted-by":"publisher","first-page":"129","DOI":"10.1613\/jair.295","volume":"4","author":"DA Cohn","year":"1996","unstructured":"Cohn, D. A., Ghahramani, Z., & Jordan, M. I. (1996). Active learning with statistical models. Journal of Artificial Intelligence Research, 4, 129\u2013145.","journal-title":"Journal of Artificial Intelligence Research"},{"key":"9749_CR35","unstructured":"Bachman, P., Sordoni, A., & Trischler, A. (2017). Learning algorithms for active learning. In International Conference on Machine Learning (ICML), pp. 301\u2013310."},{"key":"9749_CR36","unstructured":"Levine, S., Kumar, A., Tucker, G., & Fu, J. (2020). Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv:2005.01643"},{"key":"9749_CR37","doi-asserted-by":"crossref","unstructured":"Dukkipati, A., Ayyagari, R.S., Dasgupta, B., Dutta, P., & Onteru, P.R. (2025). Active reinforcement learning strategies for offline policy improvement. In Proceedings of the 39th AAAI Conference on Artificial Intelligence (AAAI-25). AAAI Press, Philaldelpia, USA.","DOI":"10.1609\/aaai.v39i16.33803"},{"issue":"1","key":"9749_CR38","doi-asserted-by":"publisher","first-page":"55","DOI":"10.1109\/TCIAIG.2012.2188528","volume":"4","author":"S Karakovskiy","year":"2012","unstructured":"Karakovskiy, S., & Togelius, J. (2012). The mario ai benchmark and competitions. IEEE Transactions on Computational Intelligence and AI in Games, 4(1), 55\u201367.","journal-title":"IEEE Transactions on Computational Intelligence and AI in Games"},{"issue":"11","key":"9749_CR39","doi-asserted-by":"publisher","first-page":"2312","DOI":"10.1109\/TSE.2019.2946563","volume":"47","author":"VJ Man\u00e8s","year":"2019","unstructured":"Man\u00e8s, V. J., Han, H., Han, C., Cha, S. K., Egele, M., Schwartz, E. J., & Woo, M. (2019). The art, science, and engineering of fuzzing: A survey. IEEE Transactions on Software Engineering, 47(11), 2312\u20132331.","journal-title":"IEEE Transactions on Software Engineering"},{"key":"9749_CR40","unstructured":"Wang, T.-W., Huang, X., Li, B., Jha, S., Jha, S., & Chen, J. (2020). Towards verification of neural-network control systems. arXiv:2003.06139"},{"key":"9749_CR41","doi-asserted-by":"publisher","unstructured":"Wan, X., Li, T., Lin, W., Cai, Y., & Zheng, Z. (2024). Coverage-guided fuzzing for deep reinforcement learning systems. Journal of Systems and Software,\u00a0210, 111963. https:\/\/doi.org\/10.1016\/j.jss.2024.111963","DOI":"10.1016\/j.jss.2024.111963"},{"key":"9749_CR42","doi-asserted-by":"publisher","unstructured":"Tappler, M., Pferscher, A., Aichernig, B.K., & K\u00f6nighofer, B. (2024). Learning and repair of deep reinforcement learning policies from fuzz-testing data. In Proceedings of the IEEE\/ACM 46th International Conference on Software Engineering. ICSE \u201924. Association for Computing Machinery, New York, NY, USA. https:\/\/doi.org\/10.1145\/3597503.3623311, https:\/\/doi.org\/10.1145\/3597503.3623311","DOI":"10.1145\/3597503.3623311"},{"key":"9749_CR43","doi-asserted-by":"publisher","first-page":"103655","DOI":"10.1016\/j.jbi.2020.103655","volume":"113","author":"AF Markus","year":"2021","unstructured":"Markus, A. F., Kors, J. A., & Rijnbeek, P. R. (2021). The role of explainability in creating trustworthy artificial intelligence for health care: a comprehensive survey of the terminology, design choices, and evaluation strategies. Journal of Biomedical Informatics, 113, 103655.","journal-title":"Journal of Biomedical Informatics"},{"key":"9749_CR44","unstructured":"Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning, pp. 1861\u20131870. PMLR."},{"key":"9749_CR45","unstructured":"Srinivasan, K., Eysenbach, B., Ha, S., Tan, J., & Finn, C. (2020). Learning to be Safe: Deep RL with a Safety Critic. arXiv:2010.14603"}],"container-title":["Autonomous Agents and Multi-Agent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10458-026-09749-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10458-026-09749-5","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10458-026-09749-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T02:53:36Z","timestamp":1780973616000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10458-026-09749-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,9]]},"references-count":45,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,12]]}},"alternative-id":["9749"],"URL":"https:\/\/doi.org\/10.1007\/s10458-026-09749-5","relation":{},"ISSN":["1387-2532","1573-7454"],"issn-type":[{"value":"1387-2532","type":"print"},{"value":"1573-7454","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,9]]},"assertion":[{"value":"2 September 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 April 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 June 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing Interests"}}],"article-number":"33"}}