{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,20]],"date-time":"2025-09-20T20:38:14Z","timestamp":1758400694318,"version":"3.40.3"},"publisher-location":"Cham","reference-count":53,"publisher":"Springer International Publishing","isbn-type":[{"type":"print","value":"9783031040825"},{"type":"electronic","value":"9783031040832"}],"license":[{"start":{"date-parts":[[2022,1,1]],"date-time":"2022-01-01T00:00:00Z","timestamp":1640995200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,4,17]],"date-time":"2022-04-17T00:00:00Z","timestamp":1650153600000},"content-version":"vor","delay-in-days":106,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Reinforcement learning is a promising strategy for automatically training policies for challenging control tasks. However, state-of-the-art deep reinforcement learning algorithms focus on training deep neural network (DNN) policies, which are black box models that are hard to interpret and reason about. In this chapter, we describe recent progress towards learning policies in the form of programs. Compared to DNNs, such<jats:italic>programmatic policies<\/jats:italic>are significantly more interpretable, easier to formally verify, and more robust. We give an overview of algorithms designed to learn programmatic policies, and describe several case studies demonstrating their various advantages.<\/jats:p>","DOI":"10.1007\/978-3-031-04083-2_11","type":"book-chapter","created":{"date-parts":[[2022,4,16]],"date-time":"2022-04-16T17:03:23Z","timestamp":1650128603000},"page":"207-228","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Interpretable, Verifiable, and Robust Reinforcement Learning via Program Synthesis"],"prefix":"10.1007","author":[{"given":"Osbert","family":"Bastani","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jeevana Priya","family":"Inala","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Armando","family":"Solar-Lezama","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,4,17]]},"reference":[{"key":"11_CR1","doi-asserted-by":"crossref","unstructured":"Alshiekh, M., Bloem, R., Ehlers, R., K\u00f6nighofer, B., Niekum, S., Topcu, U.: Safe reinforcement learning via shielding. In: Thirty-Second AAAI Conference on Artificial Intelligence (2018)","DOI":"10.1609\/aaai.v32i1.11797"},{"key":"11_CR2","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"209","DOI":"10.1007\/3-540-57318-6_30","volume-title":"Hybrid Systems","author":"R Alur","year":"1993","unstructured":"Alur, R., Courcoubetis, C., Henzinger, T.A., Ho, P.-H.: Hybrid automata: an algorithmic approach to the specification and verification of hybrid systems. In: Grossman, R.L., Nerode, A., Ravn, A.P., Rischel, H. (eds.) HS 1991-1992. LNCS, vol. 736, pp. 209\u2013229. Springer, Heidelberg (1993). https:\/\/doi.org\/10.1007\/3-540-57318-6_30"},{"key":"11_CR3","unstructured":"Anderson, G., Verma, A., Dillig, I., Chaudhuri, S.: Neurosymbolic reinforcement learning with formally verified exploration. In: Neural Information Processing Systems (2020)"},{"key":"11_CR4","doi-asserted-by":"crossref","unstructured":"Bain, M., Sammut, C.: A framework for behavioural cloning. In: Machine Intelligence 15, pp. 103\u2013129 (1995)","DOI":"10.1093\/oso\/9780198538677.003.0006"},{"key":"11_CR5","unstructured":"Balog, M., Gaunt, A.L., Brockschmidt, M., Nowozin, S., Tarlow, D.: DeepCoder: learning to write programs. In: International Conference on Learning Representations (2017)"},{"key":"11_CR6","doi-asserted-by":"crossref","unstructured":"Bastani, H., et al.: Deploying an artificial intelligence system for COVID-19 testing at the greek border. Available at SSRN (2021)","DOI":"10.2139\/ssrn.3789038"},{"key":"11_CR7","doi-asserted-by":"crossref","unstructured":"Bastani, O.: Safe reinforcement learning with nonlinear dynamics via model predictive shielding. In: 2021 American Control Conference (ACC), pp. 3488\u20133494. IEEE (2021)","DOI":"10.23919\/ACC50511.2021.9483182"},{"key":"11_CR8","doi-asserted-by":"crossref","unstructured":"Bastani, O., Li, S., Xu, A.: Safe reinforcement learning via statistical model predictive shielding. In: Robotics: Science and Systems (2021)","DOI":"10.15607\/RSS.2021.XVII.026"},{"key":"11_CR9","unstructured":"Bastani, O., Pu, Y., Solar-Lezama, A.: Verifiable reinforcement learning via policy extraction. arXiv preprint arXiv:1805.08328 (2018)"},{"key":"11_CR10","doi-asserted-by":"crossref","unstructured":"Breiman, L., Friedman, J.H., Olshen, R.A., Stone, C.J.: Classification and Regression Trees. Routledge (2017)","DOI":"10.1201\/9781315139470"},{"key":"11_CR11","unstructured":"Brockman, G., et al.: OpenAI gym. arXiv preprint arXiv:1606.01540 (2016)"},{"key":"11_CR12","doi-asserted-by":"crossref","unstructured":"Chen, Q., Lamoreaux, A., Wang, X., Durrett, G., Bastani, O., Dillig, I.: Web question answering with neurosymbolic program synthesis. In: Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation, pp. 328\u2013343 (2021)","DOI":"10.1145\/3453483.3454047"},{"key":"11_CR13","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"587","DOI":"10.1007\/978-3-030-53291-8_30","volume-title":"Computer Aided Verification","author":"Y Chen","year":"2020","unstructured":"Chen, Y., Wang, C., Bastani, O., Dillig, I., Feng, Yu.: Program synthesis using deduction-guided reinforcement learning. In: Lahiri, S.K., Wang, C. (eds.) CAV 2020. LNCS, vol. 12225, pp. 587\u2013610. Springer, Cham (2020). https:\/\/doi.org\/10.1007\/978-3-030-53291-8_30"},{"issue":"5712","key":"11_CR14","doi-asserted-by":"publisher","first-page":"1082","DOI":"10.1126\/science.1107799","volume":"307","author":"S Collins","year":"2005","unstructured":"Collins, S., Ruina, A., Tedrake, R., Wisse, M.: Efficient bipedal robots based on passive-dynamic walkers. Science 307(5712), 1082\u20131085 (2005)","journal-title":"Science"},{"key":"11_CR15","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"337","DOI":"10.1007\/978-3-540-78800-3_24","volume-title":"Tools and Algorithms for the Construction and Analysis of Systems","author":"L de Moura","year":"2008","unstructured":"de Moura, L., Bj\u00f8rner, N.: Z3: an efficient SMT solver. In: Ramakrishnan, C.R., Rehof, J. (eds.) TACAS 2008. LNCS, vol. 4963, pp. 337\u2013340. Springer, Heidelberg (2008). https:\/\/doi.org\/10.1007\/978-3-540-78800-3_24"},{"key":"11_CR16","unstructured":"Ellis, K., Ritchie, D., Solar-Lezama, A., Tenenbaum, J.B.: Learning to infer graphics programs from hand-drawn images. arXiv preprint arXiv:1707.09627 (2017)"},{"key":"11_CR17","unstructured":"Ellis, K., Solar-Lezama, A., Tenenbaum, J.: Unsupervised learning by program synthesis (2015)"},{"issue":"6","key":"11_CR18","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1145\/2813885.2737977","volume":"50","author":"JK Feser","year":"2015","unstructured":"Feser, J.K., Chaudhuri, S., Dillig, I.: Synthesizing data structure transformations from input-output examples. ACM SIGPLAN Not. 50(6), 229\u2013239 (2015)","journal-title":"ACM SIGPLAN Not."},{"issue":"1","key":"11_CR19","doi-asserted-by":"publisher","first-page":"317","DOI":"10.1145\/1925844.1926423","volume":"46","author":"S Gulwani","year":"2011","unstructured":"Gulwani, S.: Automating string processing in spreadsheets using input-output examples. ACM Sigplan Not. 46(1), 317\u2013330 (2011)","journal-title":"ACM Sigplan Not."},{"issue":"137","key":"11_CR20","first-page":"3","volume":"45","author":"S Gulwani","year":"2016","unstructured":"Gulwani, S.: Programming by examples. Dependable Softw. Syst. Eng. 45(137), 3\u201315 (2016)","journal-title":"Dependable Softw. Syst. Eng."},{"issue":"1\u20132","key":"11_CR21","first-page":"1","volume":"4","author":"S Gulwani","year":"2017","unstructured":"Gulwani, S., Polozov, O., Singh, R., et al.: Program synthesis. Found. Trends\u00ae Program. Lang. 4(1\u20132), 1\u2013119 (2017)","journal-title":"Found. Trends\u00ae Program. Lang."},{"key":"11_CR22","first-page":"3149","volume":"25","author":"H He","year":"2012","unstructured":"He, H., Eisner, J., Daume, H.: Imitation learning by coaching. Adv. Neural. Inf. Process. Syst. 25, 3149\u20133157 (2012)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"key":"11_CR23","unstructured":"Heess, N., Hunt, J.J., Lillicrap, T.P., Silver, D.: Memory-based control with recurrent neural networks. arXiv preprint arXiv:1512.04455 (2015)"},{"key":"11_CR24","series-title":"NATO ASI Series","doi-asserted-by":"publisher","first-page":"265","DOI":"10.1007\/978-3-642-59615-5_13","volume-title":"Verification of Digital and Hybrid Systems","author":"TA Henzinger","year":"2000","unstructured":"Henzinger, T.A.: The theory of hybrid automata. In: Inan, M.K., Kurshan, R.P. (eds.) Verification of Digital and Hybrid Systems. NATO ASI Series, vol. 170, pp. 265\u2013292. Springer, Berlin (2000). https:\/\/doi.org\/10.1007\/978-3-642-59615-5_13"},{"key":"11_CR25","unstructured":"Huang, J., Smith, C., Bastani, O., Singh, R., Albarghouthi, A., Naik, M.: Generating programmatic referring expressions via program synthesis. In: International Conference on Machine Learning, pp. 4495\u20134506. PMLR (2020)"},{"key":"11_CR26","unstructured":"Inala, J.P., Bastani, O., Tavares, Z., Solar-Lezama, A.: Synthesizing programmatic policies that inductively generalize. In: International Conference on Learning Representations (2020)"},{"key":"11_CR27","unstructured":"Inala, J.P., et al.: Neurosymbolic transformers for multi-agent communication. In: Neural Information Processing Systems (2020)"},{"issue":"1\u20132","key":"11_CR28","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1016\/S0004-3702(98)00023-X","volume":"101","author":"LP Kaelbling","year":"1998","unstructured":"Kaelbling, L.P., Littman, M.L., Cassandra, A.R.: Planning and acting in partially observable stochastic domains. Artif. Intell. 101(1\u20132), 99\u2013134 (1998)","journal-title":"Artif. Intell."},{"key":"11_CR29","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"200","DOI":"10.1007\/978-3-662-46681-0_15","volume-title":"Tools and Algorithms for the Construction and Analysis of Systems","author":"S Kong","year":"2015","unstructured":"Kong, S., Gao, S., Chen, W., Clarke, E.: dReach: $$\\delta $$-reachability analysis for hybrid systems. In: Baier, C., Tinelli, C. (eds.) TACAS 2015. LNCS, vol. 9035, pp. 200\u2013205. Springer, Heidelberg (2015). https:\/\/doi.org\/10.1007\/978-3-662-46681-0_15"},{"key":"11_CR30","unstructured":"Kraska, T., et al.: SageDB: a learned database system. In: CIDR (2019)"},{"issue":"1","key":"11_CR31","first-page":"1334","volume":"17","author":"S Levine","year":"2016","unstructured":"Levine, S., Finn, C., Darrell, T., Abbeel, P.: End-to-end training of deep visuomotor policies. J. Mach. Learn. Res. 17(1), 1334\u20131373 (2016)","journal-title":"J. Mach. Learn. Res."},{"key":"11_CR32","doi-asserted-by":"crossref","unstructured":"Li, S., Bastani, O.: Robust model predictive shielding for safe reinforcement learning with stochastic dynamics. In: 2020 IEEE International Conference on Robotics and Automation (ICRA), pp. 7166\u20137172. IEEE (2020)","DOI":"10.1109\/ICRA40945.2020.9196867"},{"key":"11_CR33","unstructured":"Lillicrap, T.P., et al.: Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971 (2015)"},{"issue":"7540","key":"11_CR34","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih, V., et al.: Human-level control through deep reinforcement learning. Nature 518(7540), 529\u2013533 (2015)","journal-title":"Nature"},{"key":"11_CR35","doi-asserted-by":"crossref","unstructured":"Pepy, R., Lambert, A., Mounier, H.: Path planning using a dynamic vehicle model. In: 2006 2nd International Conference on Information & Communication Technologies, vol. 1, pp. 781\u2013786. IEEE (2006)","DOI":"10.1109\/ICTTA.2006.1684472"},{"key":"11_CR36","first-page":"331","volume":"2","author":"ML Puterman","year":"1990","unstructured":"Puterman, M.L.: Markov decision processes. Handb. Oper. Res. Manage. Sci. 2, 331\u2013434 (1990)","journal-title":"Handb. Oper. Res. Manage. Sci."},{"key":"11_CR37","unstructured":"Raghu, A., Komorowski, M., Celi, L.A., Szolovits, P., Ghassemi, M.: Continuous state-space models for optimal sepsis treatment: a deep reinforcement learning approach. In: Machine Learning for Healthcare Conference, pp. 147\u2013163. PMLR (2017)"},{"key":"11_CR38","unstructured":"Ross, S., Gordon, G., Bagnell, D.: A reduction of imitation learning and structured prediction to no-regret online learning. In: Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, pp. 627\u2013635. JMLR Workshop and Conference Proceedings (2011)"},{"key":"11_CR39","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"110","DOI":"10.1007\/978-3-030-28423-7_8","volume-title":"Numerical Software Verification","author":"S Sadraddini","year":"2019","unstructured":"Sadraddini, S., Shen, S., Bastani, O.: Polytopic trees for verification of learning-based controllers. In: Zamani, M., Zufferey, D. (eds.) NSV 2019. LNCS, vol. 11652, pp. 110\u2013127. Springer, Cham (2019). https:\/\/doi.org\/10.1007\/978-3-030-28423-7_8"},{"issue":"1","key":"11_CR40","doi-asserted-by":"publisher","first-page":"305","DOI":"10.1145\/2490301.2451150","volume":"41","author":"E Schkufza","year":"2013","unstructured":"Schkufza, E., Sharma, R., Aiken, A.: Stochastic superoptimization. ACM SIGARCH Comput. Archit. News 41(1), 305\u2013316 (2013)","journal-title":"ACM SIGARCH Comput. Archit. News"},{"key":"11_CR41","unstructured":"Schulman, J., Levine, S., Abbeel, P., Jordan, M., Moritz, P.: Trust region policy optimization. In: International Conference on Machine Learning, pp. 1889\u20131897. PMLR (2015)"},{"key":"11_CR42","unstructured":"Shah, A., Zhan, E., Sun, J.J., Verma, A., Yue, Y., Chaudhuri, S.: Learning differentiable programs with admissible neural heuristics. In: NeurIPS (2020)"},{"issue":"7587","key":"11_CR43","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","volume":"529","author":"D Silver","year":"2016","unstructured":"Silver, D., et al.: Mastering the game of go with deep neural networks and tree search. Nature 529(7587), 484\u2013489 (2016)","journal-title":"Nature"},{"key":"11_CR44","volume-title":"Reinforcement Learning: An Introduction","author":"RS Sutton","year":"2018","unstructured":"Sutton, R.S., Barto, A.G.: Reinforcement Learning: An Introduction. MIT Press, Cambridge (2018)"},{"key":"11_CR45","unstructured":"Sutton, R.S., McAllester, D.A., Singh, S.P., Mansour, Y.: Policy gradient methods for reinforcement learning with function approximation. In: Advances in Neural Information Processing Systems, pp. 1057\u20131063 (2000)"},{"key":"11_CR46","unstructured":"Tian, Y., et al.: Learning to infer and execute 3D shape programs. In: International Conference on Learning Representations (2018)"},{"key":"11_CR47","unstructured":"Valkov, L., Chaudhari, D., Srivastava, A., Sutton, C., Chaudhuri, S.: HOUDINI: lifelong learning as program synthesis. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems, pp. 8701\u20138712 (2018)"},{"key":"11_CR48","unstructured":"Vaswani, A., et al.: Attention is all you need. In: Advances in Neural Information Processing Systems, pp. 5998\u20136008 (2017)"},{"key":"11_CR49","unstructured":"Verma, A., Le, H.M., Yue, Y., Chaudhuri, S.: Imitation-projected programmatic reinforcement learning. In: Neural Information Processing Systems (2019)"},{"key":"11_CR50","unstructured":"Verma, A., Murali, V., Singh, R., Kohli, P., Chaudhuri, S.: Programmatically interpretable reinforcement learning. In: International Conference on Machine Learning, pp. 5045\u20135054. PMLR (2018)"},{"key":"11_CR51","doi-asserted-by":"crossref","unstructured":"Wabersich, K.P., Zeilinger, M.N.: Linear model predictive safety certification for learning-based control. In: 2018 IEEE Conference on Decision and Control (CDC), pp. 7130\u20137135. IEEE (2018)","DOI":"10.1109\/CDC.2018.8619829"},{"key":"11_CR52","unstructured":"Wang, F., Rudin, C.: Falling rule lists. In: Artificial Intelligence and Statistics, pp. 1013\u20131022. PMLR (2015)"},{"key":"11_CR53","unstructured":"Young, H., Bastani, O., Naik, M.: Learning neurosymbolic generative models via program synthesis. In: International Conference on Machine Learning, pp. 7144\u20137153. PMLR (2019)"}],"container-title":["Lecture Notes in Computer Science","xxAI - Beyond Explainable AI"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-031-04083-2_11","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,22]],"date-time":"2024-09-22T10:06:58Z","timestamp":1726999618000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-031-04083-2_11"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022]]},"ISBN":["9783031040825","9783031040832"],"references-count":53,"URL":"https:\/\/doi.org\/10.1007\/978-3-031-04083-2_11","relation":{},"ISSN":["0302-9743","1611-3349"],"issn-type":[{"type":"print","value":"0302-9743"},{"type":"electronic","value":"1611-3349"}],"subject":[],"published":{"date-parts":[[2022]]},"assertion":[{"value":"17 April 2022","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"xxAI","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"International Workshop on Extending Explainable AI Beyond Deep Models and Classifiers","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Vienna","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Austria","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2020","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"18 July 2020","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"18 July 2020","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"xxai2020","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"https:\/\/human-centered.ai\/xxai-icml-2020\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}}]}}