{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T07:54:05Z","timestamp":1781596445190,"version":"3.54.5"},"reference-count":31,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2024,7,9]],"date-time":"2024-07-09T00:00:00Z","timestamp":1720483200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,7,9]],"date-time":"2024-07-09T00:00:00Z","timestamp":1720483200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100000038","name":"Natural Sciences and Engineering Research Council of Canada","doi-asserted-by":"publisher","award":["DG-2022-04277"],"award-info":[{"award-number":["DG-2022-04277"]}],"id":[{"id":"10.13039\/501100000038","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Intell Robot Syst"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Sparse rewards and sample efficiency are open areas of research in the field of reinforcement learning. These problems are especially important when considering applications of reinforcement learning to robotics and other cyber-physical systems. This is so because in these domains many tasks are goal-based and naturally expressed with binary successes and failures, action spaces are large and continuous, and real interactions with the environment are limited. In this work, we propose Deep Value-and-Predictive-Model Control\u00a0(DVPMC), a model-based predictive reinforcement learning algorithm for continuous control that uses system identification, value function approximation and sampling-based optimization to select actions. The algorithm is evaluated on a dense reward and a sparse reward task. We show that it can match the performance of a predictive control approach to the dense reward problem, and outperforms model-free and model-based learning algorithms on the sparse reward task on the metrics of sample efficiency and performance. We verify the performance of an agent trained in simulation using DVPMC on a real robot playing the reach-avoid game. Video of the experiment can be found here: <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/youtu.be\/0Q274kcfn4c\">https:\/\/youtu.be\/0Q274kcfn4c<\/jats:ext-link>.<\/jats:p>","DOI":"10.1007\/s10846-024-02118-y","type":"journal-article","created":{"date-parts":[[2024,7,9]],"date-time":"2024-07-09T12:01:35Z","timestamp":1720526495000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Deep Model-Based Reinforcement Learning for Predictive Control of Robotic Systems with Dense and Sparse Rewards"],"prefix":"10.1007","volume":"110","author":[{"given":"Luka","family":"Antonyshyn","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sidney","family":"Givigi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,7,9]]},"reference":[{"key":"2118_CR1","doi-asserted-by":"publisher","unstructured":"Yu, Y.: Towards sample efficient reinforcement learning. IJCAI International Joint Conference on Artificial Intelligence. 2018-July, 5739\u20135743 (2018) https:\/\/doi.org\/10.24963\/ijcai.2018\/820","DOI":"10.24963\/ijcai.2018\/820"},{"key":"2118_CR2","doi-asserted-by":"publisher","unstructured":"Ladosz, P., Weng, L., Kim, M., Oh, H.: Exploration in deep reinforcement learning: A survey. Information Fusion. 85(July 2021), 1\u201322 (2022) https:\/\/doi.org\/10.1016\/j.inffus.2022.03.003","DOI":"10.1016\/j.inffus.2022.03.003"},{"key":"2118_CR3","unstructured":"Antonyshyn, L.: Deep model-based reinforcement learning for sample efficient predictive control. Master\u2019s thesis, School of Computing, Queen\u2019s University, Kingston, ON, Canada (2022)"},{"issue":"3","key":"2118_CR4","doi-asserted-by":"publisher","first-page":"385","DOI":"10.1007\/s10994-012-5322-7","volume":"90","author":"T Hester","year":"2013","unstructured":"Hester, T., Stone, P.: TEXPLORE: Real-time sample-efficient reinforcement learning for robots. Mach. Learn. 90(3), 385\u2013429 (2013). https:\/\/doi.org\/10.1007\/s10994-012-5322-7","journal-title":"Mach. Learn."},{"issue":"3","key":"2118_CR5","doi-asserted-by":"publisher","first-page":"38","DOI":"10.1109\/37.845037","volume":"20","author":"JB Rawlings","year":"2000","unstructured":"Rawlings, J.B.: Tutorial Overview of Model Predictive Control. IEEE Control. Syst. 20(3), 38\u201352 (2000). https:\/\/doi.org\/10.1109\/37.845037","journal-title":"IEEE Control. Syst."},{"issue":"3","key":"2118_CR6","doi-asserted-by":"publisher","first-page":"577","DOI":"10.1109\/LCSYS.2019.2913347","volume":"3","author":"D Piga","year":"2019","unstructured":"Piga, D., Forgione, M., Formentin, S., Bemporad, A.: Performance-Oriented Model Learning for Data-Driven MPC Design. IEEE Control Systems Letters. 3(3), 577\u2013582 (2019)","journal-title":"IEEE Control Systems Letters."},{"key":"2118_CR7","unstructured":"Farshidian, F., Hoeller, D., Hutter, M.: Deep Value Model Predictive Control. (CoRL) (2019)"},{"issue":"6","key":"2118_CR8","doi-asserted-by":"publisher","first-page":"2713","DOI":"10.1109\/TCST.2019.2948135","volume":"28","author":"U Rosolia","year":"2020","unstructured":"Rosolia, U., Borrelli, F.: Learning How to Autonomously Race a Car: A Predictive Control Approach. IEEE Trans. Control Syst. Technol. 28(6), 2713\u20132719 (2020). https:\/\/doi.org\/10.1109\/TCST.2019.2948135","journal-title":"IEEE Trans. Control Syst. Technol."},{"issue":"4","key":"2118_CR9","doi-asserted-by":"publisher","first-page":"2767","DOI":"10.1109\/TII.2019.2940663","volume":"16","author":"H Han","year":"2020","unstructured":"Han, H., Liu, Z., Hou, Y., Qiao, J.: Data-driven multiobjective predictive control for wastewater treatment process. IEEE Trans. Industr. Inf. 16(4), 2767\u20132775 (2020). https:\/\/doi.org\/10.1109\/TII.2019.2940663","journal-title":"IEEE Trans. Industr. Inf."},{"key":"2118_CR10","unstructured":"Fujimoto, S., Van\u00a0Hoof, H., Meger, D.: Addressing Function Approximation Error in Actor-Critic Methods. 35th International Conference on Machine Learning, ICML 2018. 4, 2587\u20132601 (2018)"},{"key":"2118_CR11","unstructured":"Haarnoja, T., Zhou, A., Abbeel, P., Levine, S.: Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. 35th International Conference on Machine Learning, ICML 2018. 5, 2976\u20132989 (2018)"},{"key":"2118_CR12","unstructured":"Deisenroth, M.P., Rasmussen, C.E.: PILCO: A model-based and data-efficient approach to policy search. Proceedings of the 28th International Conference on Machine Learning, ICML 2011, 465\u2013472 (2011)"},{"key":"2118_CR13","unstructured":"Chua, K., Calandra, R., McAllister, R., Levine, S.: Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models. Advances in Neural Information Processing Systems. 2018-Decem(NeurIPS), 4754\u20134765 (2018)"},{"key":"2118_CR14","unstructured":"Janner, M., Fu, J., Zhang, M., Levine, S.: When to trust your model: Model-based policy optimization. Advances in Neural Information Processing Systems. 32(NeurIPS) (2019)"},{"issue":"11","key":"2118_CR15","doi-asserted-by":"publisher","first-page":"6912","DOI":"10.1109\/TII.2020.2974037","volume":"16","author":"H Zhao","year":"2020","unstructured":"Zhao, H., Zhao, J., Qiu, J., Liang, G., Dong, Z.Y.: Cooperative wind farm control with deep reinforcement learning and knowledge-assisted learning. IEEE Trans. Industr. Inf. 16(11), 6912\u20136921 (2020). https:\/\/doi.org\/10.1109\/TII.2020.2974037","journal-title":"IEEE Trans. Industr. Inf."},{"key":"2118_CR16","unstructured":"Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., Zaremba, W.: Hindsight Experience Replay (279 cites). Advances in Neural Information Processing Systems. 2017-Decem(Nips), 5049\u20135059 (2017)"},{"key":"2118_CR17","volume-title":"Process Dynamics & Control","author":"DE Seborg","year":"2011","unstructured":"Seborg, D.E., Edgar, T.F., Mellichamp, D.A.: Process Dynamics & Control. Sons, Hoboken, NJ (2011)"},{"key":"2118_CR18","doi-asserted-by":"publisher","unstructured":"Zhang, T., Ma, F., Peng, C., Yu, Y., Yue, D., Dou, C., O\u2019Hare, G.M.P.: A very-short-term online pv power prediction model based on ran with secondary dynamic adjustment. IEEE Transactions on Artificial Intelligence, 1\u20131 (2022) https:\/\/doi.org\/10.1109\/TAI.2022.3179353","DOI":"10.1109\/TAI.2022.3179353"},{"issue":"1","key":"2118_CR19","doi-asserted-by":"publisher","first-page":"269","DOI":"10.1146\/annurev-control-090419-075625","volume":"3","author":"L Hewing","year":"2020","unstructured":"Hewing, L., Wabersich, K.P., Menner, M., Zeilinger, M.N.: Learning-based model predictive control: Toward safe learning in control. Annual Review of Control, Robotics, and Autonomous Systems. 3(1), 269\u2013296 (2020). https:\/\/doi.org\/10.1146\/annurev-control-090419-075625","journal-title":"Annual Review of Control, Robotics, and Autonomous Systems."},{"key":"2118_CR20","doi-asserted-by":"publisher","unstructured":"Rosolia, U., Borrelli, F.: Learning model predictive control for iterative tasks. a data-driven control framework. IEEE Transactions on Automatic Control. 63(7), 1883\u20131896 (2018) https:\/\/doi.org\/10.1109\/TAC.2017.2753460","DOI":"10.1109\/TAC.2017.2753460"},{"key":"2118_CR21","doi-asserted-by":"publisher","unstructured":"Thananjeyan, B., Balakrishna, A., Rosolia, U., Li, F., McAllister, R., Gonzalez, J., Levine, S., Borrelli, F., Goldberg, K.: Safety augmented value estimation from demonstrations (saved): Safe deep model-based rl for sparse cost robotic tasks. IEEE Robotics and Automation Letters. PP, 1\u20131 (2020) https:\/\/doi.org\/10.1109\/LRA.2020.2976272","DOI":"10.1109\/LRA.2020.2976272"},{"issue":"6","key":"2118_CR22","doi-asserted-by":"publisher","first-page":"1161","DOI":"10.1109\/TNNLS.2014.2334366","volume":"26","author":"N Wang","year":"2015","unstructured":"Wang, N., Er, M.J., Han, M.: Generalized single-hidden layer feedforward networks for regression problems. IEEE Transactions on Neural Networks and Learning Systems. 26(6), 1161\u20131176 (2015). https:\/\/doi.org\/10.1109\/TNNLS.2014.2334366","journal-title":"IEEE Transactions on Neural Networks and Learning Systems."},{"key":"2118_CR23","doi-asserted-by":"publisher","unstructured":"Nagabandi, A., Kahn, G., Fearing, R.S., Levine, S.: Neural Network Dynamics for Model-Based Deep Reinforcement Learning with Model-Free Fine-Tuning. Proceedings - IEEE International Conference on Robotics and Automation, 7579\u20137586 (2018) https:\/\/doi.org\/10.1109\/ICRA.2018.8463189","DOI":"10.1109\/ICRA.2018.8463189"},{"key":"2118_CR24","doi-asserted-by":"publisher","first-page":"980","DOI":"10.1109\/LCSYS.2021.3087968","volume":"6","author":"KP Wabersich","year":"2022","unstructured":"Wabersich, K.P., Krishnadas, R., Zeilinger, M.N.: A Soft Constrained MPC Formulation Enabling Learning From Trajectories With Constraint Violations. IEEE Control Systems Letters. 6, 980\u2013985 (2022)","journal-title":"IEEE Control Systems Letters."},{"key":"2118_CR25","unstructured":"McAllister, R., Rasmussen, C.E.: Improving pilco with bayesian neural network dynamics models. (2016)"},{"key":"2118_CR26","doi-asserted-by":"publisher","unstructured":"Alessio, A., Bemporad, A.: In: Magni, L., Raimondo, D.M., Allg\u00f6wer, F. (eds.) A Survey on Explicit Model Predictive Control, pp. 345\u2013369. Springer, Berlin, Heidelberg (2009). https:\/\/doi.org\/10.1007\/978-3-642-01094-1_29","DOI":"10.1007\/978-3-642-01094-1_29"},{"issue":"12","key":"2118_CR27","doi-asserted-by":"publisher","first-page":"5522","DOI":"10.1109\/TNNLS.2020.2969215","volume":"31","author":"M Liu","year":"2020","unstructured":"Liu, M., Wan, Y., Lewis, F.L., Lopez, V.G.: Adaptive Optimal Control for Stochastic Multiplayer Differential Games Using On-Policy and Off-Policy Reinforcement Learning. IEEE Transactions on Neural Networks and Learning Systems. 31(12), 5522\u20135533 (2020). https:\/\/doi.org\/10.1109\/TNNLS.2020.2969215","journal-title":"IEEE Transactions on Neural Networks and Learning Systems."},{"issue":"8","key":"2118_CR28","doi-asserted-by":"publisher","first-page":"2874","DOI":"10.1109\/TCYB.2018.2830820","volume":"49","author":"Q Zhang","year":"2019","unstructured":"Zhang, Q., Zhao, D.: Data-Based Reinforcement Learning for Nonzero-Sum Games with Unknown Drift Dynamics. IEEE Transactions on Cybernetics. 49(8), 2874\u20132885 (2019). https:\/\/doi.org\/10.1109\/TCYB.2018.2830820","journal-title":"IEEE Transactions on Cybernetics."},{"key":"2118_CR29","doi-asserted-by":"publisher","unstructured":"Schwartz, H.: An Object Oriented Approach to Fuzzy Actor-Critic Learning for Multi-Agent Differential Games. 2019 IEEE Symposium Series on Computational Intelligence, SSCI 2019, 183\u2013190 (2019) https:\/\/doi.org\/10.1109\/SSCI44817.2019.9002707","DOI":"10.1109\/SSCI44817.2019.9002707"},{"key":"2118_CR30","doi-asserted-by":"publisher","unstructured":"Liu, W., Sun, J., Wang, G., Bullo, F., Chen, J.: Data-Driven Resilient Predictive Control under Denial-of-Service. arXiv (2021). https:\/\/doi.org\/10.48550\/ARXIV.2110.12766 . https:\/\/arxiv.org\/abs\/2110.12766","DOI":"10.48550\/ARXIV.2110.12766"},{"key":"2118_CR31","doi-asserted-by":"publisher","unstructured":"Maddalena, E.T., da S. Moraes, C.G., Waltrich, G., Jones, C.N.: A neural network architecture to learn explicit mpc controllers from data**this work has received support from the swiss national science foundation under the risk project (risk aware data-driven demand response, grant number 200021 175627. IFAC-PapersOnLine. 53(2), 11362\u201311367 (2020) https:\/\/doi.org\/10.1016\/j.ifacol.2020.12.546 . 21st IFAC World Congress","DOI":"10.1016\/j.ifacol.2020.12.546"}],"container-title":["Journal of Intelligent &amp; Robotic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10846-024-02118-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10846-024-02118-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10846-024-02118-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,26]],"date-time":"2024-09-26T13:10:58Z","timestamp":1727356258000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10846-024-02118-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,9]]},"references-count":31,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2024,9]]}},"alternative-id":["2118"],"URL":"https:\/\/doi.org\/10.1007\/s10846-024-02118-y","relation":{},"ISSN":["1573-0409"],"issn-type":[{"value":"1573-0409","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,9]]},"assertion":[{"value":"10 April 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 May 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 July 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant conflicts of interests to declare.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of Interests:"}},{"value":"The authors have no relevant financial or non-financial interests to disclose.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Financial interests"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval:"}},{"value":"Not applicable.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to Participate"}},{"value":"Not applicable.","order":6,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for Publication"}}],"article-number":"100"}}