{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T12:24:36Z","timestamp":1783427076493,"version":"3.54.6"},"reference-count":40,"publisher":"SAGE Publications","issue":"6","license":[{"start":{"date-parts":[[2022,6,2]],"date-time":"2022-06-02T00:00:00Z","timestamp":1654128000000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000006","name":"Office of Naval Research","doi-asserted-by":"publisher","award":["N00014-21-1-2706"],"award-info":[{"award-number":["N00014-21-1-2706"]}],"id":[{"id":"10.13039\/100000006","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CNS 1837515"],"award-info":[{"award-number":["CNS 1837515"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of Robotics Research"],"published-print":{"date-parts":[[2023,5]]},"abstract":"<jats:p>We develop an approach to improve the learning capabilities of robotic systems by combining learned predictive models with experience-based state-action policy mappings. Predictive models provide an understanding of the task and the dynamics, while experience-based (model-free) policy mappings encode favorable actions that override planned actions. We refer to our approach of systematically combining model-based and model-free learning methods as hybrid learning. Our approach efficiently learns motor skills and improves the performance of predictive models and experience-based policies. Moreover, our approach enables policies (both model-based and model-free) to be updated using any off-policy reinforcement learning method. We derive a deterministic method of hybrid learning by optimally switching between learning modalities. We adapt our method to a stochastic variation that relaxes some of the key assumptions in the original derivation. Our deterministic and stochastic variations are tested on a variety of robot control benchmark tasks in simulation as well as a hardware manipulation task. We extend our approach for use with imitation learning methods, where experience is provided through demonstrations, and we test the expanded capability with a real-world pick-and-place task. The results show that our method is capable of improving the performance and sample efficiency of learning motor skills in a variety of experimental domains.<\/jats:p>","DOI":"10.1177\/02783649221083331","type":"journal-article","created":{"date-parts":[[2022,6,2]],"date-time":"2022-06-02T08:08:50Z","timestamp":1654157330000},"page":"337-355","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":30,"title":["Hybrid control for combining model-based and model-free reinforcement learning"],"prefix":"10.1177","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3095-8856","authenticated-orcid":false,"given":"Allison","family":"Pinosky","sequence":"first","affiliation":[{"name":"Department of Mechanical Engineering, Northwestern University, Evanston, IL, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0299-1760","authenticated-orcid":false,"given":"Ian","family":"Abraham","sequence":"additional","affiliation":[{"name":"Department of Mechanical Engineering and Materials Science at Yale University, New Haven, CT, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9230-7891","authenticated-orcid":false,"given":"Alexander","family":"Broad","sequence":"additional","affiliation":[{"name":"Boston Dynamics, Waltham, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Brenna","family":"Argall","sequence":"additional","affiliation":[{"name":"Department of Mechanical Engineering, Northwestern University, Evanston, IL, USA"},{"name":"Department of Electrical Engineering and Computer Science, Northwestern University, Evanston, IL, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2262-8176","authenticated-orcid":false,"given":"Todd D","family":"Murphey","sequence":"additional","affiliation":[{"name":"Department of Mechanical Engineering, Northwestern University, Evanston, IL, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2022,6,2]]},"reference":[{"key":"bibr1-02783649221083331","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2017.XIII.052"},{"key":"bibr2-02783649221083331","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2020.2972836"},{"key":"bibr3-02783649221083331","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2019.2923880"},{"key":"bibr4-02783649221083331","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2016.2596768"},{"key":"bibr5-02783649221083331","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2008.10.024"},{"key":"bibr6-02783649221083331","doi-asserted-by":"publisher","DOI":"10.1007\/s10957-007-9305-y"},{"key":"bibr7-02783649221083331","unstructured":"Bansal S, Calandra R, Chua K, et al. (2017) Mbmf: Model-based priors for model-free reinforcement learning. arXiv preprint arXiv:1709.03153."},{"key":"bibr8-02783649221083331","unstructured":"Boyan JA (1999) Least-squares temporal difference learning. In: Proceedings of the 16th International conference on machine learning, Bled, Slovenia, June 27-30, 1999, pp. 49\u201356."},{"key":"bibr9-02783649221083331","unstructured":"Brockman G, Cheung V, Pettersson L, et al. (2016) OpenAI Gym. arXiv preprint arXiv:1606.01540."},{"key":"bibr10-02783649221083331","unstructured":"Buckman J, Hafner D, Tucker G, et al. (2018) Sample-efficient reinforcement learning with stochastic ensemble value expansion. In: NeurIPS Montreal, Canada, December 3-8, 2018."},{"key":"bibr11-02783649221083331","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2017.7989384"},{"key":"bibr12-02783649221083331","unstructured":"Chua K, Calandra R, McAllister R, et al. (2018) Deep reinforcement learning in a handful of trials using probabilistic dynamics models. In: Advances in neural information processing systems, Montreal, Canada, December 3-8, 2018, pp. 4754\u20134765."},{"key":"bibr13-02783649221083331","volume-title":"Pybullet, A Python Module for Physics Simulation for Games, Robotics and Machine Learning","author":"Coumans E","year":"2016"},{"key":"bibr14-02783649221083331","unstructured":"Deisenroth M, Rasmussen CE (2011) PILCO: A model-based and data-efficient approach to policy search. In: Proceedings of the 28th International Conference on machine learning (ICML-11), pp. 465\u2013472."},{"key":"bibr15-02783649221083331","unstructured":"Feinberg V, Wan A, Stoica I, et al. (2018) Model-based value estimation for efficient model-free reinforcement learning. arXiv preprint arXiv:1803.00101."},{"key":"bibr16-02783649221083331","first-page":"1861","volume":"80","author":"Haarnoja T","year":"2018","journal-title":"Proceedings of Machine Learning Research, PMLR"},{"key":"bibr17-02783649221083331","unstructured":"Haarnoja T, Zhou A, Hartikainen K, et al. (2018b) Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905."},{"key":"bibr18-02783649221083331","unstructured":"Havens A, Ouyang Y, Nagarajan P, et al. (2019) Learning latent state spaces for planning through reward prediction. arXiv preprint arXiv:1912.04201."},{"key":"bibr19-02783649221083331","first-page":"12519","volume":"32","author":"Janner M","year":"2019","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr20-02783649221083331","unstructured":"Kingma DP, Ba J (2014) Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980."},{"key":"bibr21-02783649221083331","first-page":"761","volume":"120","author":"Lambert N","year":"2020","journal-title":"Proceedings of Machine Learning Research, PMLR"},{"key":"bibr22-02783649221083331","unstructured":"Levine S, Abbeel P (2014) Learning neural network policies with guided policy search under unknown dynamics. In: Advances in neural information processing systems, Montreal, Canada, 8-13 December, 2014, pp. 1071\u20131079."},{"key":"bibr23-02783649221083331","unstructured":"Li W, Todorov E (2004) Iterative linear quadratic regulator design for nonlinear biological movement systems. In: International Conference on informatics in control, automation and robotics, Setubal, Portugal, 25-28 August, 2004, pp. 222\u2013229."},{"key":"bibr24-02783649221083331","first-page":"4008","volume-title":"Advances in neural information processing systems","volume":"29","author":"Montgomery WH","year":"2016"},{"key":"bibr25-02783649221083331","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2018.8463189"},{"key":"bibr26-02783649221083331","volume-title":"Roboschool, Open-Source Software for Robot Simulation, Integrated With Openai Gym","author":"OpenAI","year":"2017"},{"key":"bibr27-02783649221083331","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2017.70"},{"key":"bibr28-02783649221083331","volume-title":"Advances in neural information processing systems","volume":"1","author":"Pomerleau D","year":"1998"},{"key":"bibr29-02783649221083331","unstructured":"Precup D, Sutton RS, Dasgupta S (2001) Off-policy temporal-difference learning with function approximation. In: ICML, Williamstown, MA, 28 June-1 July, 2001, pp. 417\u2013424."},{"key":"bibr30-02783649221083331","unstructured":"Ross S, Bagnell D (2010) Efficient reductions for imitation learning. In: Proceedings of the thirteenth international conference on artificial intelligence and statistics, Sardinia, Italy, 13-15 May, 2010, pp. 661\u2013668."},{"key":"bibr31-02783649221083331","unstructured":"Schulman J, Wolski F, Dhariwal P, et al. (2017) Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347."},{"key":"bibr32-02783649221083331","unstructured":"Sharma A, Gu S, Levine S, et al. (2019) Dynamics-aware unsupervised discovery of skills. arXiv preprint arXiv:1907.01657."},{"key":"bibr33-02783649221083331","unstructured":"Sutton RS, McAllester DA, Singh SP, et al. (2000) Policy gradient methods for reinforcement learning with function approximation. In: Advances in neural information processing systems, Denver, Colorado, USA, 27 November-2 December, 2000, pp. 1057\u20131063."},{"key":"bibr34-02783649221083331","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2018.8593536"},{"key":"bibr35-02783649221083331","doi-asserted-by":"crossref","unstructured":"Theodorou EA, Todorov E (2012) Relative entropy and free energy dualities: Connections to path integral and KL control. In: IEEE Conference on decision and control (CDC), Maui, Hawaii, USA, 10-13 December, 2012, pp. 1466\u20131473.","DOI":"10.1109\/CDC.2012.6426381"},{"key":"bibr36-02783649221083331","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2012.6386109"},{"key":"bibr37-02783649221083331","doi-asserted-by":"publisher","DOI":"10.1137\/120901490"},{"key":"bibr38-02783649221083331","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2016.7487277"},{"key":"bibr39-02783649221083331","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2018.XIV.042"},{"key":"bibr40-02783649221083331","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2017.7989202"}],"container-title":["The International Journal of Robotics Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649221083331","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/02783649221083331","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649221083331","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649221083331","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T10:17:05Z","timestamp":1777457825000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/02783649221083331"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,2]]},"references-count":40,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,5]]}},"alternative-id":["10.1177\/02783649221083331"],"URL":"https:\/\/doi.org\/10.1177\/02783649221083331","relation":{},"ISSN":["0278-3649","1741-3176"],"issn-type":[{"value":"0278-3649","type":"print"},{"value":"1741-3176","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,2]]}}}