{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T16:19:37Z","timestamp":1783700377384,"version":"3.55.0"},"reference-count":54,"publisher":"Springer Science and Business Media LLC","issue":"9","license":[{"start":{"date-parts":[[2021,5,13]],"date-time":"2021-05-13T00:00:00Z","timestamp":1620864000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,5,13]],"date-time":"2021-05-13T00:00:00Z","timestamp":1620864000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"National Science Foundation","award":["CPS-1739964"],"award-info":[{"award-number":["CPS-1739964"]}]},{"name":"National Science Foundation","award":["IIS-1724157"],"award-info":[{"award-number":["IIS-1724157"]}]},{"name":"National Science Foundation","award":["NRI-1925082"],"award-info":[{"award-number":["NRI-1925082"]}]},{"DOI":"10.13039\/100000006","name":"Office of Naval Research","doi-asserted-by":"crossref","award":["N00014-18-2243"],"award-info":[{"award-number":["N00014-18-2243"]}],"id":[{"id":"10.13039\/100000006","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100017561","name":"Future of Life Institute","doi-asserted-by":"crossref","award":["RFP2-000"],"award-info":[{"award-number":["RFP2-000"]}],"id":[{"id":"10.13039\/100017561","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100006754","name":"Army Research Laboratory","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100006754","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000185","name":"Defense Advanced Research Projects Agency","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000185","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100002186","name":"Lockheed Martin","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100002186","id-type":"DOI","asserted-by":"crossref"}]},{"name":"General Motors"},{"DOI":"10.13039\/100012884","name":"Bosch","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100012884","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2021,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Reinforcement learning in simulation is a promising alternative to the prohibitive sample cost of reinforcement learning in the physical world. Unfortunately, policies learned in simulation often perform worse than hand-coded policies when applied on the target, physical system. Grounded simulation learning (<jats:sc>gsl<\/jats:sc>) is a general framework that promises to address this issue by altering the simulator to better match the real world (Farchy et al. 2013 in Proceedings of the 12th international conference on autonomous agents and multiagent systems (AAMAS)). This article introduces a new algorithm for <jats:sc>gsl<\/jats:sc>\u2014Grounded Action Transformation (GAT)\u2014and applies it to learning control policies for a humanoid robot. We evaluate our algorithm in controlled experiments where we show it to allow policies learned in simulation to transfer to the real world. We then apply our algorithm to learning a fast bipedal walk on a humanoid robot and demonstrate a 43.27% improvement in forward walk velocity compared to a state-of-the art hand-coded walk. This striking empirical success notwithstanding, further empirical analysis shows that <jats:sc>gat<\/jats:sc> may struggle when the real world has stochastic state transitions. To address this limitation we generalize <jats:sc>gat<\/jats:sc> to the <jats:italic>stochastic<\/jats:italic><jats:sc>gat<\/jats:sc> (<jats:sc>sgat<\/jats:sc>) algorithm and empirically show that <jats:sc>sgat<\/jats:sc> leads to successful real world transfer in situations where <jats:sc>gat<\/jats:sc> may fail to find a good policy. Our results contribute to a deeper understanding of grounded simulation learning and demonstrate its effectiveness for applying reinforcement learning to learn robot control policies entirely in simulation.<\/jats:p>","DOI":"10.1007\/s10994-021-05982-z","type":"journal-article","created":{"date-parts":[[2021,5,13]],"date-time":"2021-05-13T19:02:56Z","timestamp":1620932576000},"page":"2469-2499","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":26,"title":["Grounded action transformation for sim-to-real reinforcement learning"],"prefix":"10.1007","volume":"110","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7411-0398","authenticated-orcid":false,"given":"Josiah P.","family":"Hanna","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Siddharth","family":"Desai","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haresh","family":"Karnan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Garrett","family":"Warnell","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peter","family":"Stone","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,5,13]]},"reference":[{"key":"5982_CR1","doi-asserted-by":"crossref","unstructured":"Abbeel, P., Quigley, M., & Ng, A. Y. (2006). Using inaccurate models in reinforcement learning. In Proceedings of the 23rd international conference on machine learning (ICML). http:\/\/dl.acm.org\/citation.cfm?id=1143845.","DOI":"10.1145\/1143844.1143845"},{"key":"5982_CR2","doi-asserted-by":"crossref","unstructured":"Ashar, J., Ashmore, J., Hall, B., Harris, S., Hengst, B., Liu, R., Mei, Z., Pagnucco, M., Roy, R., Sammut, C., Sushkov, O., Teh, B., & Tsekouras, L. (2015). RoboCup SPL 2014 champion team paper. In R. A. C. Bianchi, H. L. Akin, S. Ramamoorthy, K. Sugiura (Eds.) RoboCup 2014: Robot World Cup XVIII, lecture notes in artificial intelligence (Vol. 8992, pp. 70\u201381). Springer International Publishing.","DOI":"10.1007\/978-3-319-18615-3_6"},{"key":"5982_CR3","doi-asserted-by":"crossref","unstructured":"Boeing, A., & Br\u00e4unl, T. (2012). Leveraging multiple simulators for crossing the reality gap. In Proceedings of the 12th international conference on control automation robotics & vision (ICARCV) (pp. 1113\u20131119). IEEE.","DOI":"10.1109\/ICARCV.2012.6485313"},{"key":"5982_CR4","doi-asserted-by":"crossref","unstructured":"Bousmalis, K., Irpan, A., Wohlhart, P., Bai, Y., Kelcey, M., Kalakrishnan, M., Downs, L., Ibarz, J., Pastor, P., Konolige, K., Levine, S., & Vanhoucke, V. (2018). Using simulation and domain adaptation to improve efficiency of deep robotic grasping. In Proceedings of the IEEE international conference on robotics and automation (ICRA).","DOI":"10.1109\/ICRA.2018.8460875"},{"key":"5982_CR5","doi-asserted-by":"crossref","unstructured":"Chebotar, Y., Handa, A., Makoviychuk, V., Macklin, M., Issac, J., Ratliff, N., & Fox, D. (2019). Closing the sim-to-real loop: Adapting simulation randomization with real world experience. In Proceedings of the IEEE international conference on robotics and automation (ICRA).","DOI":"10.1109\/ICRA.2019.8793789"},{"key":"5982_CR6","unstructured":"Christiano, P., Shah, Z., Mordatch, I., Schneider, J., Blackwell, T., Tobin, J., Abbeel, P., & Zaremba, W. (2016). Transfer from simulation to real world through learning deep inverse dynamics model. arXiv preprint arXiv:161003518."},{"issue":"7553","key":"5982_CR7","doi-asserted-by":"publisher","first-page":"503","DOI":"10.1038\/nature14422","volume":"521","author":"A Cully","year":"2015","unstructured":"Cully, A., Clune, J., Tarapore, D., & Mouret, J. B. (2015). Robots that can adapt like animals. Nature, 521(7553), 503.","journal-title":"Nature"},{"key":"5982_CR8","doi-asserted-by":"crossref","unstructured":"Cutler, M., & How, J. P. (2015). Efficient reinforcement learnng for robots using informative simulated priors. In Proceedings of the IEEE international conference on robotics and automation (ICRA).","DOI":"10.1109\/ICRA.2015.7139550"},{"key":"5982_CR9","doi-asserted-by":"crossref","unstructured":"Cutler, M., Walsh, T. J., & How, J. P. (2014). Reinforcement learning with multi-fidelity simulators. In Proceedings of the IEEE conference on robotics and automation (ICRA). http:\/\/www.research.rutgers.edu\/~thomaswa\/pub\/icra2014Car.pdf.","DOI":"10.1109\/ICRA.2014.6907423"},{"key":"5982_CR10","unstructured":"Deisenroth, M. P., & Rasmussen, C. E. (2011). PILCO: A model-based and data-efficient approach to policy search. In Proceedings of the 28th international conference on machine learning (ICML)."},{"key":"5982_CR11","doi-asserted-by":"crossref","unstructured":"Devin, C., Gupta, A., Darrell, T., Abbeel, P., & Levine, S. (2017). Learning modular neural network policies for multi-task and multi-robot transfer. In Proceedings of the IEEE international conference on robotics and automation (ICRA) (pp. 2169\u20132176). IEEE.","DOI":"10.1109\/ICRA.2017.7989250"},{"key":"5982_CR12","doi-asserted-by":"crossref","unstructured":"Fang, K., Bai, Y., Hinterstoisser, S., & Kalakrishnan, M. (2018). Multi-task domain adaptation for deep learning of instance grasping from simulation. In Proceedings of the IEEE international conference on robotics and automation (ICRA).","DOI":"10.1109\/ICRA.2018.8461041"},{"key":"5982_CR13","unstructured":"Farchy, A., Barrett, S., MacAlpine, P., & Stone, P. (2013). Humanoid robots learning to walk faster: From the real world to simulation and back. In Proceedings of the 12th international conference on autonomous agents and multiagent systems (AAMAS). http:\/\/www.cs.utexas.edu\/users\/ai-lab\/?AAMAS13-Farchy."},{"key":"5982_CR14","unstructured":"Golemo, F., Taiga, A. A., Courville, A., & Oudeyer, P. Y. (2018). Sim-to-real transfer with neural-augmented robot simulation. In Proceedings of the 2nd conference on robot learning (CORL) (pp. 817\u2013828)."},{"key":"5982_CR15","doi-asserted-by":"crossref","unstructured":"Hall, B., Harris, S., Hengst, B., Liu, R., Ng, K., Pagnucco, M., Pearson, L., Sammut, C., & Schmidt, P. (2016). RoboCup SPL 2015 champion team paper. In RoboCup 2015: Robot World Cup XIX, lecture notes in artificial intelligence (Vol. 9513, pp. 72\u201382). Springer International Publishing.","DOI":"10.1007\/978-3-319-29339-4_6"},{"issue":"1","key":"5982_CR16","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1162\/106365603321828970","volume":"11","author":"N Hansen","year":"2003","unstructured":"Hansen, N., M\u00fcller, S. D., & Koumoutsakos, P. (2003). Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (cma-es). Evolutionary Computation, 11(1), 1\u201318.","journal-title":"Evolutionary Computation"},{"key":"5982_CR17","unstructured":"Hengst, B. (2014). rUNSWift walk2014 report robocup standard platform league. Technical report. The University of New South Wales."},{"key":"5982_CR18","unstructured":"Hill, A., Raffin, A., Ernestus, M., Gleave, A., Kanervisto, A., Traore, R., Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., & Wu, Y. (2018). Stable baselines. https:\/\/github.com\/hill-a\/stable-baselines."},{"key":"5982_CR19","unstructured":"Iocchi, L., Libera, F. D., & Menegatti, E. (2007). Learning humanoid soccer actions interleaving simulated and real data. In Proceedings of the 2nd workshop on humanoid soccer robots. http:\/\/citeseerx.ist.psu.edu\/viewdoc\/summary?doi=10.1.1.139.4931."},{"key":"5982_CR20","doi-asserted-by":"crossref","unstructured":"Jakobi, N., Husbands, P., Harvey, I. (1995). Noise and the reality gap: The use of simulation in evolutionary robotics. In Proceedings of the European conference on artificial life (pp. 704\u2013720). Springer.","DOI":"10.1007\/3-540-59496-5_337"},{"key":"5982_CR21","doi-asserted-by":"crossref","unstructured":"James, S., Wohlhart, P., Kalakrishnan, M., Kalashnikov, D., Irpan, A., Ibarz, J., Levine, S., Hadsell, R., & Bousmalis, K. (2019). Sim-to-real via sim-to-sim: Data-efficient robotic grasping via randomized-to-canonical adaptation networks. In Proceedings of the IEEE international conference on computer vision and pattern recognition (CVPR) (pp. 12627\u201312637).","DOI":"10.1109\/CVPR.2019.01291"},{"key":"5982_CR22","unstructured":"Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:14126980."},{"key":"5982_CR23","unstructured":"Kober, J., Bagnell, J. A., & Peters, J. (2013). Reinforcement learning in robotics: A survey. The International Journal of Robotics Research. http:\/\/ijr.sagepub.com\/content\/early\/2013\/08\/22\/0278364913495721.abstract."},{"key":"5982_CR24","doi-asserted-by":"crossref","unstructured":"Koos, S., Mouret, J. B., & Doncieux, S. (2010). Crossing the reality gap in evolutionary robotics by promoting transferable controllers. In Proceedings of the 12th annual conference on genetic and evolutionary computation (GECCO) (pp. 119\u2013126). ACM. http:\/\/dl.acm.org\/citation.cfm?id=1830505.","DOI":"10.1145\/1830483.1830505"},{"key":"5982_CR25","unstructured":"Lee, G., Srinivasa, S. S., & Mason, M. T. (2017) GP-ILQG: Data-driven robust optimal control for uncertain nonlinear dynamical systems. arXiv preprint arXiv:170505344."},{"key":"5982_CR26","unstructured":"Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., & Wierstra, D. (2015). Continuous control with deep reinforcement learning. arXiv preprint arXiv:150902971."},{"key":"5982_CR27","doi-asserted-by":"crossref","unstructured":"Lowrey, K., Kolev, S., Dao, J., Rajeswaran, A., & Todorov, E. (2018). Reinforcement learning for non-prehensile manipulation: Transfer from simulation to physical system. In PProceedings of the IEEE international conference on simulation, modeling, and programming for autonomous robots (SIMPAR) (pp. 35\u201342). IEEE.","DOI":"10.1109\/SIMPAR.2018.8376268"},{"key":"5982_CR28","doi-asserted-by":"crossref","unstructured":"Marco, A., Berkenkamp, F., Hennig, P., Schoellig, A. P., Krause, A., Schaal, S., & Trimpe, S. (2017). Virtual vs. real: Trading off simulations and physical experiments in reinforcement learning with bayesian optimization. In Proceedings of the IEEE international conference on robotics and automation (ICRA) (pp. 1557\u20131563). IEEE.","DOI":"10.1109\/ICRA.2017.7989186"},{"issue":"4","key":"5982_CR29","doi-asserted-by":"publisher","first-page":"417","DOI":"10.1162\/artl.1995.2.4.417","volume":"2","author":"O Miglino","year":"1996","unstructured":"Miglino, O., Lund, H. H., & Nolfi, S. (1996). Evolving mobile robots in simulated and real environments. Artificial Life, 2(4), 417\u2013434.","journal-title":"Artificial Life"},{"key":"5982_CR30","doi-asserted-by":"crossref","unstructured":"Molchanov, A., Chen, T., H\u00f6nig, W., Preiss, J. A., Ayanian, N., & Sukhatme, G. S. (2019). Sim-to-(multi)-real: Transfer of low-level robust control policies to multiple quadrotors. arXiv preprint arXiv:190304628.","DOI":"10.1109\/IROS40897.2019.8967695"},{"key":"5982_CR31","unstructured":"Mozifian, M., Higuera, J.C.G., Meger, D., & Dudek, G. (2019). Learning domain randomization distributions for transfer of locomotion policies. arXiv preprint arXiv:190600410."},{"key":"5982_CR32","unstructured":"Muratore, F., Treede, F., Gienger, & M., Peters, J. (2018). Domain randomization for simulation-based policy optimization with transferability assessment. In Proceedings of the 2nd conference on robot learning (CORL) (pp. 700\u2013713)."},{"key":"5982_CR33","doi-asserted-by":"crossref","unstructured":"OpenAI, Andrychowicz, M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., Schneider, J., Sidor, S., Tobin, J., Welinder, P., Weng, L., & Zaremba, W. (2018). Learning dexterous in-hand manipulation. arXiv preprint arXiv:180800177.","DOI":"10.1177\/0278364919887447"},{"key":"5982_CR34","doi-asserted-by":"crossref","unstructured":"Peng, X. B., Andrychowicz, M., Zaremba, W., & Abbeel, P. (2017). Sim-to-real transfer of robotic control with dynamics randomization. In Proceedings of the IEEE international conference on robotics and automation (ICRA).","DOI":"10.1109\/ICRA.2018.8460528"},{"key":"5982_CR35","doi-asserted-by":"crossref","unstructured":"Pinto, L., Andrychowicz, M., Welinder, P., Zaremba, W., & Abbeel, P. (2017a). Asymmetric actor critic for image-based robot learning. In Proceedings of the robotics: Science and systems Conference (RSS).","DOI":"10.15607\/RSS.2018.XIV.008"},{"key":"5982_CR36","unstructured":"Pinto, L., Davidson, J., Sukthankar, R., Gupta, A. (2017b). Robust adversarial reinforcement learning. In Proceedings of the 34th international conference on machine learning (ICML)."},{"key":"5982_CR37","unstructured":"Puterman, M. L. (2014). Markov decision processes: Discrete stochastic dynamic programming. John Wiley & Sons"},{"key":"5982_CR38","unstructured":"Rajeswaran, A., Ghotra, S., Levine, S., & Ravindran, B. (2017). EPOpt: Learning robust neural network policies using model ensembles. In Proceedings of the international conference on learning representations (ICLR)."},{"key":"5982_CR39","doi-asserted-by":"crossref","unstructured":"Ramos, F., Possas, R. C., & Fox, D. (2019). Bayessim: Adaptive domain randomization via probabilistic inference for robotics simulators. arXiv preprint arXiv:190601728.","DOI":"10.15607\/RSS.2019.XV.029"},{"key":"5982_CR40","doi-asserted-by":"crossref","unstructured":"Rodriguez, D., Brandenburger, A., & Behnke, S. (2019). Combining simulations and real-robot experiments for bayesian optimization of bipedal gait stabilization. In RoboCup 2018: Robot World Cup XXII, lecture notes in artificial intelligence (Vol. 11374). Springer International Publishing.","DOI":"10.1007\/978-3-030-27544-0_6"},{"key":"5982_CR41","unstructured":"Ross, S., Gordon, G. J., & Bagnell, D. (2011). A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the international conference on artificial intelligence and statistics (AISTATS) (pp. 627\u2013635)."},{"key":"5982_CR42","unstructured":"Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., & Hadsell, R. (2016a). Progressive neural networks. arXiv preprint arXiv:160604671."},{"key":"5982_CR43","unstructured":"Rusu, A. A., Vecerik, M., Roth\u00f6rl, T., Heess, N., Pascanu, R., & Hadsell, R. (2016b). Sim-to-real robot learning from pixels with progressive nets. In Proceedings of the 1st conference on robot learning (CORL)."},{"key":"5982_CR44","doi-asserted-by":"crossref","unstructured":"Sadeghi, F., & Levine, S. (2017). (CAD)$$^2$$ RL: Real single-image flight without a single real image. In Proceedings of the robotics: Science and systems conference (RSS).","DOI":"10.15607\/RSS.2017.XIII.034"},{"key":"5982_CR45","unstructured":"Schulman, J., Levine, S., Moritz, P., Jordan, M., & Abbeel, P. (2015a). Trust region policy optimization. In Proceedings of the 32nd international conference on machine learning (ICML). http:\/\/jmlr.csail.mit.edu\/proceedings\/papers\/v37\/schulman15.html."},{"key":"5982_CR46","unstructured":"Schulman, J., Moritz, P., Levine, S., Jordan, M. I., & Abbeel, P. (2015b). High-dimensional continuous control using generalized advantage estimation. In Proceedings of the international conference on learning representations (ICLR)."},{"key":"5982_CR47","doi-asserted-by":"crossref","unstructured":"Sutton, R. S., & Barto, A. G. (1998). Reinforcement learning: An introduction. MIT Press.","DOI":"10.1109\/TNN.1998.712192"},{"key":"5982_CR48","doi-asserted-by":"crossref","unstructured":"Tan, J., Zhang, T., Coumans, E., Iscen, A., Bai, Y., Hafner, D., Bohez, S., & Vanhoucke, V. (2018). Sim-to-real: Learning agile locomotion for quadruped robots. In Proceedings of the robotics: Science and systems conference (RSS).","DOI":"10.15607\/RSS.2018.XIV.010"},{"key":"5982_CR49","doi-asserted-by":"crossref","unstructured":"Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., & Abbeel, P. (2017). Domain randomization for transferring deep neural networks from simulation to the real world. In Proceedings of the 30th IEEE\/RSJ international conference on intelligent robots and systems (IROS).","DOI":"10.1109\/IROS.2017.8202133"},{"key":"5982_CR50","doi-asserted-by":"crossref","unstructured":"Tobin, J., Zaremba, W., & Abbeel, P. (2018). Domain randomization and generative models for robotic grasping. In Proceedings of the 31st IEEE\/RSJ international conference on intelligent robots and systems (IROS).","DOI":"10.1109\/IROS.2018.8593933"},{"key":"5982_CR51","unstructured":"Tzeng, E., Coline, D., Hoffman, J., Finn, C., Xingchao, P., Levine, S., Saenko, K., & Darrell, T. (2016). Towards adapting deep visuomotor representations from simulated to real environments. In Proceedings of the workshop on algorithmic foundations of robotics (WAFR)."},{"key":"5982_CR52","unstructured":"Urieli, D., MacAlpine, P., Kalyanakrishnan, S., Bentor, Y., & Stone, P. (2011). On optimizing interdependent skills: A case study in simulated 3D humanoid robot soccer. In Proceedings of the 10th international conference on autonomous agents and multiagent systems (AAMAS) (pp. 769\u2013776)."},{"key":"5982_CR53","unstructured":"Zhang, F., Leitner, J., Upcroft, B., & Corke, P. (2016). Vision-based reaching using modular deep networks: From simulation to the real world. arXiv preprint arXiv:161006781."},{"key":"5982_CR54","doi-asserted-by":"publisher","unstructured":"Zhu, S., Kimmel, A., E\u00a0Bekris, K., & Boularias, A. (2018). Fast model identification via physics engines for data-efficient policy search. In Proceedings of the 27th international joint conference on artificial intelligence (IJCAI) (pp. 3249\u20133256). https:\/\/doi.org\/10.24963\/ijcai.2018\/451.","DOI":"10.24963\/ijcai.2018\/451"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-021-05982-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-021-05982-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-021-05982-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,8,30]],"date-time":"2021-08-30T18:20:04Z","timestamp":1630347604000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-021-05982-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,13]]},"references-count":54,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2021,9]]}},"alternative-id":["5982"],"URL":"https:\/\/doi.org\/10.1007\/s10994-021-05982-z","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,5,13]]},"assertion":[{"value":"9 March 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 September 2020","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 April 2021","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 May 2021","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declaration"}},{"value":"Peter Stone serves as the Executive Director of Sony AI America and receives financial compensation for this work. The terms of this arrangement have been reviewed and approved by the University of Texas at Austin in accordance with its policy on objectivity in research.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}