{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T07:32:26Z","timestamp":1740123146719,"version":"3.37.3"},"reference-count":48,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2021,9,5]],"date-time":"2021-09-05T00:00:00Z","timestamp":1630800000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,9,5]],"date-time":"2021-09-05T00:00:00Z","timestamp":1630800000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100006690","name":"Politecnico di Milano","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100006690","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2022,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We study the problem of identifying the policy space available to an agent in a learning process, having access to a set of demonstrations generated by the agent playing the optimal policy in the considered space. We introduce an approach based on frequentist statistical testing to identify the set of policy parameters that the agent can control, within a larger parametric policy space. After presenting two identification rules (combinatorial and simplified), applicable under different assumptions on the policy space, we provide a probabilistic analysis of the simplified one in the case of linear policies belonging to the exponential family. To improve the performance of our identification rules, we make use of the recently introduced framework of the Configurable Markov Decision Processes, exploiting the opportunity of configuring the environment to induce the agent to reveal which parameters it can control. Finally, we provide an empirical evaluation, on both discrete and continuous domains, to prove the effectiveness of our identification rules.<\/jats:p>","DOI":"10.1007\/s10994-021-06033-3","type":"journal-article","created":{"date-parts":[[2021,9,5]],"date-time":"2021-09-05T16:02:25Z","timestamp":1630857745000},"page":"2093-2145","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Policy space identification in configurable environments"],"prefix":"10.1007","volume":"111","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3424-5212","authenticated-orcid":false,"given":"Alberto Maria","family":"Metelli","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guglielmo","family":"Manneschi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marcello","family":"Restelli","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,9,5]]},"reference":[{"issue":"1","key":"6033_CR1","doi-asserted-by":"publisher","first-page":"89","DOI":"10.1007\/s10994-007-5038-2","volume":"71","author":"A Antos","year":"2008","unstructured":"Antos, A., Szepesv\u00e1ri, C., & Munos, R. (2008). Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path. Machine Learning, 71(1), 89\u2013129. https:\/\/doi.org\/10.1007\/s10994-007-5038-2.","journal-title":"Machine Learning"},{"doi-asserted-by":"crossref","unstructured":"Barnard, G. A. (1959). Control charts and stochastic processes. Journal of the Royal Statistical Society: Series B (Methodological)","key":"6033_CR2","DOI":"10.1111\/j.2517-6161.1959.tb00336.x"},{"unstructured":"Ben-Israel, A., Greville, T.N. (2003). Generalized inverses: theory and applications, vol\u00a015. Berlin: Springer Science & Business Media","key":"6033_CR3"},{"key":"6033_CR4","doi-asserted-by":"publisher","DOI":"10.1093\/acprof:oso\/9780199535255.001.0001","volume-title":"Concentration inequalities\u2014A nonasymptotic theory of independence","author":"S Boucheron","year":"2013","unstructured":"Boucheron, S., Lugosi, G., & Massart, P. (2013). Concentration inequalities\u2014A nonasymptotic theory of independence. Oxford: Oxford University Press."},{"unstructured":"Brantley, K., Sun, W., Henaff, M. (2020). Disagreement-regularized imitation learning. In: 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia. April 26-30, 2020. OpenReview.net","key":"6033_CR5"},{"doi-asserted-by":"crossref","unstructured":"Brown, L.D. (1986). Fundamentals of statistical exponential families: with applications in statistical decision theory. Ims","key":"6033_CR6","DOI":"10.1214\/lnms\/1215466757"},{"issue":"2","key":"6033_CR7","doi-asserted-by":"publisher","first-page":"156","DOI":"10.1109\/TSMCC.2007.913919","volume":"38","author":"L Busoniu","year":"2008","unstructured":"Busoniu, L., Babuska, R., & De Schutter, B. (2008). A comprehensive survey of multiagent reinforcement learning. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), 38(2), 156\u2013172.","journal-title":"IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews)"},{"key":"6033_CR8","volume-title":"Statistical inference","author":"G Casella","year":"2002","unstructured":"Casella, G., & Berger, R. L. (2002). Statistical inference (Vol. 2). Duxbury: Pacific Grove."},{"unstructured":"Dai, B., Shaw, A., Li, L., Xiao, L., He, N., Liu, Z., Chen, J., Song, L. (2018). SBEED: convergent reinforcement learning with nonlinear function approximation. In: Dy JG, Krause A (eds) Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm\u00e4ssan, Stockholm, Sweden. July 10-15, 2018, PMLR, Proceedings of Machine Learning Research, vol.\u00a080, pp. 1133\u20131142","key":"6033_CR9"},{"unstructured":"Deisenroth, M.P., Rasmussen, C.E. (2011). PILCO: A model-based and data-efficient approach to policy search. In: Proceedings of the 28th International Conference on Machine Learning, ICML 2011, Bellevue, Washington, USA, June 28\u2013July 2, 2011, pp. 465\u2013472","key":"6033_CR10"},{"doi-asserted-by":"crossref","unstructured":"Deisenroth, M.P., Neumann, G., Peters, J. (2013). A survey on policy search for robotics. Foundations and Trends in Robotics","key":"6033_CR11","DOI":"10.1109\/ICRA.2014.6907421"},{"unstructured":"Finn, C., Abbeel, P., Levine, S. (2017). Model-agnostic meta-learning for fast adaptation of deep networks. In: Proceedings of the 34th International Conference on Machine Learning, PMLR, Proceedings of Machine Learning Research, vol.\u00a070, pp. 1126\u20131135","key":"6033_CR12"},{"unstructured":"Garivier, A., Kaufmann, E. (2019). Non-asymptotic sequential tests for overlapping hypotheses and application to near optimal arm identification in bandit models. arXiv preprint arXiv:190503495","key":"6033_CR13"},{"issue":"12","key":"6033_CR14","doi-asserted-by":"publisher","first-page":"995","DOI":"10.7326\/0003-4819-130-12-199906150-00008","volume":"130","author":"SN Goodman","year":"1999","unstructured":"Goodman, S. N. (1999). Toward evidence-based medical statistics. 1: The p value fallacy. Annals of internal medicine, 130(12), 995\u20131004.","journal-title":"Annals of internal medicine"},{"unstructured":"Ho, J., Ermon, S. (2016). Generative adversarial imitation learning. In: Lee DD, Sugiyama M, von Luxburg U, Guyon I, Garnett R (eds) Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain, pp. 4565\u20134573","key":"6033_CR15"},{"doi-asserted-by":"crossref","unstructured":"Hsu, D., Kakade, S., Zhang, T., et\u00a0al. (2012). A tail inequality for quadratic forms of subgaussian random vectors. Electronic Communications in Probability, 17","key":"6033_CR16","DOI":"10.1214\/ECP.v17-2079"},{"key":"6033_CR17","doi-asserted-by":"publisher","first-page":"203","DOI":"10.1017\/S030500410001330X","volume":"31","author":"H Jeffreys","year":"1935","unstructured":"Jeffreys, H. (1935). Some tests of significance, treated by the theory of probability. Mathematical Proceedings of the Cambridge Philosophical Society, 31, 203\u2013222.","journal-title":"Mathematical Proceedings of the Cambridge Philosophical Society"},{"doi-asserted-by":"publisher","unstructured":"Jolliffe, I.T. (2011). Principal component analysis. In: Lovric M (ed) International Encyclopedia of Statistical Science. Springer, pp. 1094\u20131096. https:\/\/doi.org\/10.1007\/978-3-642-04898-2_455","key":"6033_CR18","DOI":"10.1007\/978-3-642-04898-2_455"},{"unstructured":"Lazaric, A., Restelli, M., Bonarini, A. (2007). Reinforcement learning in continuous action spaces through sequential monte carlo methods. In: Platt JC, Koller D, Singer Y, Roweis ST (eds) Advances in Neural Information Processing Systems 20. In: Proceedings of the Twenty-First Annual Conference on Neural Information Processing Systems, Vancouver, British Columbia, Canada, December 3-6, 2007, Curran Associates, Inc., pp. 833\u2013840","key":"6033_CR19"},{"key":"6033_CR20","first-page":"3041","volume":"13","author":"A Lazaric","year":"2012","unstructured":"Lazaric, A., Ghavamzadeh, M., & Munos, R. (2012). Finite-sample analysis of least-squares policy iteration. Journal of Machine Learning Research, 13, 3041\u20133074.","journal-title":"Journal of Machine Learning Research"},{"unstructured":"Lee, K., Choi, S., Oh, S. (2018). Maximum causal Tsallis entropy imitation learning. In: Bengio S, Wallach HM, Larochelle H, Grauman K, Cesa-Bianchi N, Garnett R (eds) Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3\u20138, 2018, Montr\u00e9al, Canada, pp. 4408\u20134418","key":"6033_CR21"},{"unstructured":"Levine, S., Koltun, V. (2013). Guided policy search. In: Proceedings of the 30th International Conference on Machine Learning, ICML 2013, Atlanta, GA, USA, 16\u201321 June 2013, JMLR.org, JMLR Workshop and Conference Proceedings, vol.\u00a028, pp. 1\u20139","key":"6033_CR22"},{"unstructured":"Li, L., Lu, Y., Zhou, D. (2017). Provably optimal algorithms for generalized linear contextual bandits. In: Proceedings of the 34th International Conference on Machine Learning, PMLR, Proceedings of Machine Learning Research, vol.\u00a070, pp. 2071\u20132080","key":"6033_CR23"},{"issue":"1","key":"6033_CR24","doi-asserted-by":"publisher","first-page":"e8915","DOI":"10.1371\/journal.pone.0008915","volume":"5","author":"MP Little","year":"2010","unstructured":"Little, M. P., Heidenreich, W. F., & Li, G. (2010). Parameter identifiability and redundancy: theoretical considerations. PloS ONE, 5(1), e8915.","journal-title":"PloS ONE"},{"unstructured":"Metelli, A.M., Mutti, M., Restelli, M. (2018a). Configurable Markov decision processes. In: Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm\u00e4ssan, Stockholm, Sweden, July 10-15, 2018, PMLR, Proceedings of Machine Learning Research, vol.\u00a080, pp. 3488\u20133497","key":"6033_CR25"},{"unstructured":"Metelli, A. M., Papini, M., Faccio, F., & Restelli, M. (2018b). Policy optimization via importance sampling. Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, 3\u20138 December 2018 (pp. 5447\u20135459). Canada.: Montr\u00e9al.","key":"6033_CR26"},{"unstructured":"Metelli, A.M., Ghelfi, E., Restelli, M. (2019). Reinforcement learning in configurable continuous environments. In: Chaudhuri K, Salakhutdinov R (eds) Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, PMLR, Proceedings of Machine Learning Research, vol.\u00a097, pp. 4546\u20134555","key":"6033_CR27"},{"key":"6033_CR28","first-page":"141:1","volume":"21","author":"AM Metelli","year":"2020","unstructured":"Metelli, A. M., Papini, M., Montali, N., & Restelli, M. (2020). Importance sampling techniques for policy optimization. Journal of Machine Learning Research, 21, 141:1-141:75.","journal-title":"Journal of Machine Learning Research"},{"key":"6033_CR29","doi-asserted-by":"publisher","first-page":"6","DOI":"10.1016\/j.jmp.2015.11.001","volume":"72","author":"RD Morey","year":"2016","unstructured":"Morey, R. D., Romeijn, J. W., & Rouder, J. N. (2016). The philosophy of Bayes factors and the quantification of statistical evidence. Journal of Mathematical Psychology, 72, 6\u201318.","journal-title":"Journal of Mathematical Psychology"},{"unstructured":"Neu, G., Jonsson, A., G\u00f3mez, V. (2017). A unified view of entropy-regularized Markov decision processes. arXiv preprint arXiv:170507798","key":"6033_CR30"},{"doi-asserted-by":"crossref","unstructured":"Osa, T., Pajarinen, J., Neumann, G., Bagnell, J.A., Abbeel, P., Peters, J. (2018). An algorithmic perspective on imitation learning. Foundations and Trends in Robotics","key":"6033_CR31","DOI":"10.1561\/9781680834116"},{"key":"6033_CR32","volume-title":"Monte Carlo theory, methods and examples","author":"AB Owen","year":"2013","unstructured":"Owen, A. B. (2013). Monte Carlo theory, methods and examples. Methods and Examples Art Owen: Monte Carlo Theory."},{"issue":"4","key":"6033_CR33","doi-asserted-by":"publisher","first-page":"682","DOI":"10.1016\/j.neunet.2008.02.003","volume":"21","author":"J Peters","year":"2008","unstructured":"Peters, J., & Schaal, S. (2008). Reinforcement learning of motor skills with policy gradients. Neural Networks, 21(4), 682\u2013697.","journal-title":"Neural Networks"},{"issue":"15","key":"6033_CR34","first-page":"510","volume":"7","author":"KB Petersen","year":"2008","unstructured":"Petersen, K. B., Pedersen, M. S., et al. (2008). The matrix cookbook. Technical University of Denmark, 7(15), 510.","journal-title":"Technical University of Denmark"},{"key":"6033_CR35","volume-title":"Markov Decision Processes: Discrete Stochastic Dynamic Programming","author":"ML Puterman","year":"2014","unstructured":"Puterman, M. L. (2014). Markov Decision Processes: Discrete Stochastic Dynamic Programming. London: John Wiley & Sons."},{"unstructured":"Rajeswaran, A., Lowrey, K., Todorov, E., Kakade, S. M., & (2017) Towards generalization and simplicity in continuous control. In: Guyon I, von Luxburg U, Bengio S, Wallach HM, Fergus R, Vishwanathan SVN, Garnett R (eds) Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017(December), pp. 4\u20139, . (2017). Long Beach, CA, USA (pp. 6550\u20136561).","key":"6033_CR36"},{"unstructured":"Ramponi, G., Likmeta, A., Metelli, A.M., Tirinzoni, A., Restelli, M. (2020). Truly batch model-free inverse reinforcement learning about multiple intentions. In: Chiappa S, Calandra R (eds) Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, PMLR, Online, Proceedings of Machine Learning Research, vol. 108, pp. 2359\u20132369","key":"6033_CR37"},{"unstructured":"Reddy, S., Dragan, A.D., Levine, S. (2019). Sqil: Imitation learning via regularized behavioral cloning. arXiv preprint arXiv:190511108","key":"6033_CR38"},{"unstructured":"R\u00e9nyi, A. (1961). On measures of entropy and information. Hungarian Academy of Sciences Budapest Hungary: Technical report.","key":"6033_CR39"},{"issue":"3","key":"6033_CR40","doi-asserted-by":"publisher","first-page":"577","DOI":"10.2307\/1913267","volume":"39","author":"TJ Rothenberg","year":"1971","unstructured":"Rothenberg, T. J., et al. (1971). Identification in parametric models. Econometrica, 39(3), 577\u2013591.","journal-title":"Econometrica"},{"key":"6033_CR41","volume-title":"Reinforcement learning: An introduction. Adaptive computation and machine learning","author":"RS Sutton","year":"2018","unstructured":"Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction. Adaptive computation and machine learning. Cambridge: MIT Press."},{"key":"6033_CR42","first-page":"1057","volume-title":"Advances in Neural Information Processing Systems 12, [NIPS Conference, Denver, Colorado, USA, November 29 - December 4, 1999]","author":"RS Sutton","year":"1999","unstructured":"Sutton, R. S., McAllester, D. A., Singh, S. P., & Mansour, Y. (1999). Policy gradient methods for reinforcement learning with function approximation. In S. A. Solla, T. K. Leen, & K. M\u00fcller (Eds.), Advances in Neural Information Processing Systems 12, [NIPS Conference, Denver, Colorado, USA, November 29 - December 4, 1999] (pp. 1057\u20131063). Cambridge: The MIT Press."},{"unstructured":"Sutton, R.S., Szepesv\u00e1ri, C., Geramifard, A., Bowling, M.H. (2008). Dyna-style planning with linear function approximation and prioritized sweeping. In: McAllester DA, Myllym\u00e4ki P (eds) UAI 2008, Proceedings of the 24th Conference in Uncertainty in Artificial Intelligence, Helsinki, Finland, July 9\u201312, 2008, AUAI Press, pp. 528\u2013536","key":"6033_CR43"},{"doi-asserted-by":"crossref","unstructured":"Vershynin, R. (2012). Introduction to the non-asymptotic analysis of random matrices. In: Compressed Sensing, Cambridge University Press, pp. 210\u2013268","key":"6033_CR44","DOI":"10.1017\/CBO9780511794308.006"},{"issue":"1","key":"6033_CR45","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1214\/aoms\/1177732360","volume":"9","author":"SS Wilks","year":"1938","unstructured":"Wilks, S. S. (1938). The large-sample distribution of the likelihood ratio for testing composite hypotheses. The Annals of Mathematical Statistics, 9(1), 60\u201362.","journal-title":"The Annals of Mathematical Statistics"},{"doi-asserted-by":"crossref","unstructured":"Yu, B. (1994). Rates of convergence for empirical processes of stationary mixing sequences. The Annals of Probability, pp. 94\u2013116","key":"6033_CR46","DOI":"10.1214\/aop\/1176988849"},{"unstructured":"Ziebart, B.D., Maas, A.L., Bagnell, J.A., Dey, A.K. (2008). Maximum entropy inverse reinforcement learning. In: Fox D, Gomes CP (eds) Proceedings of the Twenty-Third AAAI Conference on Artificial Intelligence, AAAI 2008, Chicago, Illinois, USA, July 13\u201317, 2008, AAAI Press, pp. 1433\u20131438","key":"6033_CR47"},{"unstructured":"Ziebart, B.D., Bagnell, J.A., Dey, A.K. (2010). Modeling interaction via the principle of maximum causal entropy. In: F\u00fcrnkranz J, Joachims T (eds) Proceedings of the 27th International Conference on Machine Learning (ICML-10), June 21\u201324, 2010, Haifa, Israel, Omnipress, pp. 1255\u20131262","key":"6033_CR48"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-021-06033-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-021-06033-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-021-06033-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,8]],"date-time":"2023-11-08T10:00:19Z","timestamp":1699437619000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-021-06033-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,5]]},"references-count":48,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2022,6]]}},"alternative-id":["6033"],"URL":"https:\/\/doi.org\/10.1007\/s10994-021-06033-3","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"type":"print","value":"0885-6125"},{"type":"electronic","value":"1573-0565"}],"subject":[],"published":{"date-parts":[[2021,9,5]]},"assertion":[{"value":"13 January 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 February 2021","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 June 2021","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 September 2021","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}