{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,1]],"date-time":"2026-06-01T23:33:44Z","timestamp":1780356824659,"version":"3.54.1"},"reference-count":54,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,11,26]],"date-time":"2025-11-26T00:00:00Z","timestamp":1764115200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,11,26]],"date-time":"2025-11-26T00:00:00Z","timestamp":1764115200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62422605"],"award-info":[{"award-number":["62422605"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Intell."],"published-print":{"date-parts":[[2025,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Model-based reinforcement learning (MBRL) achieves significant sample efficiency in practice in comparison to model-free RL, but its performance is often limited by the existence of model prediction error. To reduce the model error, standard MBRL approaches train a single well-designed network to fit the entire environment dynamics, but this wastes rich information on multiple sub-dynamics, which can be modeled separately, allowing us to construct the world model more accurately. In this paper, we propose environment dynamics decomposition\u00a0(ED2), a novel world model construction framework that models the environment in a decomposing manner. ED2 contains two key components: sub-dynamics discovery\u00a0(SD2) and dynamics decomposition prediction\u00a0(D2P). SD2 discovers the sub-dynamics in an environment automatically and D2P constructs the decomposed world model following the sub-dynamics. ED2 can be easily combined with the existing MBRL algorithms and empirical results show that ED2 significantly reduces the model error, increases the sample efficiency, and achieves higher asymptotic performance when combined with the state-of-the-art MBRL algorithms on various continuous control tasks.<\/jats:p>","DOI":"10.1007\/s44267-025-00094-x","type":"journal-article","created":{"date-parts":[[2025,11,26]],"date-time":"2025-11-26T03:02:21Z","timestamp":1764126141000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["ED2: environment dynamics decomposition world models for continuous control"],"prefix":"10.1007","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-2194-942X","authenticated-orcid":false,"given":"Yifu","family":"Yuan","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7478-7684","authenticated-orcid":false,"given":"Hongyao","family":"Tang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8501-5778","authenticated-orcid":false,"given":"Cong","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yan","family":"Zheng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0422-8235","authenticated-orcid":false,"given":"Jianye","family":"Hao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,11,26]]},"reference":[{"key":"94_CR1","doi-asserted-by":"publisher","first-page":"350","DOI":"10.1038\/s41586-019-1724-z","volume":"575","author":"O. Vinyals","year":"2019","unstructured":"Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al. (2019). Grandmaster level in starcraft II using multi-agent reinforcement learning. Nature, 575, 350\u2013354.","journal-title":"Nature"},{"key":"94_CR2","first-page":"1","volume-title":"Proceedings of the 8th international conference on learning representations","author":"D. Hafner","year":"2020","unstructured":"Hafner, D., Lillicrap, T. P., Ba, J., & Norouzi, M. (2020). Dream to control: learning behaviors by latent imagination. In Proceedings of the 8th international conference on learning representations (pp. 1\u201320). Retrieved September 10, 2025, from https:\/\/openreview.net\/pdf?id=S1lOTC4tDS."},{"key":"94_CR3","doi-asserted-by":"publisher","first-page":"1292","DOI":"10.1109\/TETCI.2025.3529902","volume":"9","author":"X. Zhou","year":"2025","unstructured":"Zhou, X., Yuan, Y., Yang, S., & Hao, J. (2025). Mentor: guiding hierarchical reinforcement learning with human feedback and dynamic distance constraint. IEEE Transactions on Emerging Topics in Computational Intelligence, 9, 1292\u20131306.","journal-title":"IEEE Transactions on Emerging Topics in Computational Intelligence"},{"key":"94_CR4","first-page":"1","volume-title":"Proceedings of the 12th international conference on learning representations","author":"Y. Yuan","year":"2024","unstructured":"Yuan, Y., Hao, J., Ma, Y., Dong, Z., Liang, H., Liu, J., Feng, Z., Zhao, K., & Zheng, Y. (2024). Uni-RLHF: universal platform and benchmark suite for reinforcement learning with diverse human feedback. In Proceedings of the 12th international conference on learning representations (pp. 1\u201328). Retrieved September 10, 2025, from https:\/\/openreview.net\/pdf?id=WesY0H9ghM."},{"key":"94_CR5","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-022-3696-5","volume":"67","author":"F.-M. Luo","year":"2024","unstructured":"Luo, F.-M., Xu, T., Lai, H., Chen, X.-H., Zhang, W., & Yu, Y. (2024). A survey on model-based reinforcement learning. Science China. Information Sciences, 67, 121101.","journal-title":"Science China. Information Sciences"},{"key":"94_CR6","first-page":"5618","volume-title":"Proceedings of the 37th international conference on machine learning","author":"H. Lai","year":"2020","unstructured":"Lai, H., Shen, J., Zhang, W., & Yu, Y. (2020). Bidirectional model-based policy optimization. In Proceedings of the 37th international conference on machine learning (pp. 5618\u20135627). Retrieved September 10, 2025, from https:\/\/proceedings.mlr.press\/v119\/lai20b\/lai20b.pdf."},{"key":"94_CR7","first-page":"1","volume-title":"Proceedings of the 8th international conference on learning representations","author":"L. Kaiser","year":"2020","unstructured":"Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al. (2020). Model based reinforcement learning for atari. In Proceedings of the 8th international conference on learning representations (pp. 1\u201328). Retrieved September 10, 2025, from https:\/\/openreview.net\/pdf?id=S1xCPJHtDB."},{"key":"94_CR8","first-page":"20551","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"R. Kidambi","year":"2020","unstructured":"Kidambi, R., Rajeswaran, A., Netrapalli, P., & Joachims, T. (2020). MOReL: model-based offline reinforcement learning. In A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, & C. Zhang (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 20551\u201320562). Red Hook: Curran Associates."},{"key":"94_CR9","first-page":"1","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"T. Yu","year":"2020","unstructured":"Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., & Ma, T. (2020). MOPO: model-based offline policy optimization. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, & H. Lin (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 1\u201314). Red Hook: Curran Associates."},{"key":"94_CR10","first-page":"36470","volume-title":"Proceedings of the international conference on machine learning","author":"X. Wang","year":"2023","unstructured":"Wang, X., Wongkamjan, W., Jia, R., & Huang, F. (2023). Live in the moment: learning dynamics model adapted to evolving policy. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, & J. Scarlett (Eds.), Proceedings of the international conference on machine learning (pp. 36470\u201336493). Retrieved September 10, 2025, from https:\/\/proceedings.mlr.press\/v202\/wang23an.pdf."},{"key":"94_CR11","first-page":"465","volume-title":"Proceedings of the 28th international conference on machine learning","author":"M. P. Deisenroth","year":"2011","unstructured":"Deisenroth, M. P., & Rasmussen, C. E. (2011). Pilco: a model-based and data-efficient approach to policy search. In Proceedings of the 28th international conference on machine learning (pp. 465\u2013472). Retrieved September 10, 2025, from https:\/\/www.icml2011.org\/papers\/323icmlpaper.pdf."},{"key":"94_CR12","first-page":"1","volume-title":"Proceedings of the 30th international conference on machine learning","author":"S. Levine","year":"2013","unstructured":"Levine, S., & Koltun, V. (2013). Guided policy search. In Proceedings of the 30th international conference on machine learning (pp. 1\u20139). Retrieved September 10, 2025, from https:\/\/proceedings.mlr.press\/v202\/wang23an.pdf."},{"key":"94_CR13","first-page":"20737","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"G. Zhu","year":"2020","unstructured":"Zhu, G., Zhang, M., Lee, H., & Zhang, C. (2020). Bridging imagination and reality for model-based deep reinforcement learning. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, & H. Lin (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 20737\u201320748). Red Hook: Curran Associates."},{"key":"94_CR14","first-page":"8234","volume-title":"Proceedings of the 32nd international conference on neural information processing systems","author":"J. Buckman","year":"2018","unstructured":"Buckman, J., Hafner, D., Tucker, G., Brevdo, E., & Lee, H. (2018). Sample-efficient reinforcement learning with stochastic ensemble value expansion. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, & R. Garnett (Eds.), Proceedings of the 32nd international conference on neural information processing systems (pp. 8234\u20138244). Red Hook: Curran Associates."},{"key":"94_CR15","first-page":"4759","volume-title":"Proceedings of the 32nd international conference on neural information processing systems","author":"K. Chua","year":"2018","unstructured":"Chua, K., Calandra, R., McAllister, R., & Levine, S. (2018). Deep reinforcement learning in a handful of trials using probabilistic dynamics models. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, & R. Garnett (Eds.), Proceedings of the 32nd international conference on neural information processing systems (pp. 4759\u20134770). Red Hook: Curran Associates."},{"key":"94_CR16","first-page":"258","volume-title":"Proceedings of the 3rd annual conference on robot learning","author":"M. Okada","year":"2019","unstructured":"Okada, M., & Taniguchi, T. (2019). Variational inference MPC for Bayesian model-based reinforcement learning. In L. P. Kaelbling, D. Kragic, & K. Sugiura (Eds.), Proceedings of the 3rd annual conference on robot learning (pp. 258\u2013272). Retrieved September 10, 2025, from http:\/\/proceedings.mlr.press\/v100\/okada20a\/okada20a.pdf."},{"key":"94_CR17","first-page":"1","volume-title":"Proceedings of the 11th international conference on learning representations","author":"Y. Yuan","year":"2023","unstructured":"Yuan, Y., Hao, J., Ni, F., Mu, Y., Zheng, Y., Hu, Y., Liu, J., Chen, Y., & Fan, C. (2023). Euclid: towards efficient unsupervised reinforcement learning with multi-choice dynamics model. In Proceedings of the 11th international conference on learning representations (pp. 1\u201322). Retrieved September 10, 2025, from https:\/\/openreview.net\/forum?id=xQAjSr64PTc."},{"key":"94_CR18","first-page":"1","volume-title":"Proceedings of the 7th international conference on learning representations","author":"Y. Luo","year":"2019","unstructured":"Luo, Y., Xu, H., Li, Y., Tian, Y., Darrell, T., & Ma, T. (2019). Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees. In Proceedings of the 7th international conference on learning representations (pp. 1\u201327). Retrieved September 10, 2025, from https:\/\/openreview.net\/pdf?id=Auvt132abM."},{"key":"94_CR19","first-page":"1","volume-title":"Proceedings of the 6th international conference on learning representations","author":"T. Kurutach","year":"2018","unstructured":"Kurutach, T., Clavera, I., Duan, Y., Tamar, A., & Abbeel, P. (2018). Model-ensemble trust-region policy optimization. In Proceedings of the 6th international conference on learning representations (pp. 1\u201315). Retrieved September 10, 2025, from https:\/\/openreview.net\/forum?id=SJJinbWRZ."},{"key":"94_CR20","first-page":"12498","volume-title":"Proceedings of the 33rd international conference on neural information processing systems","author":"M. Janner","year":"2019","unstructured":"Janner, M., Fu, J., Zhang, M., & Levine, S. (2019). When to trust your model: model-based policy optimization. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d\u2019Alch\u00e9 Buc, E. Fox, & R. Garnett (Eds.), Proceedings of the 33rd international conference on neural information processing systems (pp. 12498\u201312509). Red Hook: Curran Associates."},{"key":"94_CR21","first-page":"387","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"F. Pan","year":"2020","unstructured":"Pan, F., He, J., Tu, D., & He, Q. (2020). Trust the model when it is confident: masked model-based actor-critic. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, & H. Lin (Eds.), Proceedings of the 34th international conference on neural information processing systems (pp. 387\u2013397). Red Hook: Curran Associates."},{"key":"94_CR22","first-page":"2555","volume-title":"Proceedings of the 36th international conference on machine learning","author":"D. Hafner","year":"2019","unstructured":"Hafner, D., Lillicrap, T. P., Fischer, I., Villegas, R., Ha, D., Lee, H., & Davidson, J. (2019). Learning latent dynamics for planning from pixels. In K. Chaudhuri & R. Salakhutdinov (Eds.), Proceedings of the 36th international conference on machine learning (pp. 2555\u20132565). Retrieved September 14, 2025, from http:\/\/proceedings.mlr.press\/v97\/hafner19a\/hafner19a.pdf."},{"key":"94_CR23","unstructured":"Asadi, K., Misra, D., Kim, S., & Littman, M. L. (2019). Combating the compounding-error problem with a multi-step model. arXiv preprint. arXiv:1905.13320."},{"key":"94_CR24","unstructured":"Yuan, Y., Zheng, Z., Dong, Z., & Hao, J. (2024). MODULI: unlocking preference generalization via diffusion models for offline multi-objective reinforcement learning. arXiv preprint. arXiv:2408.15501."},{"key":"94_CR25","unstructured":"Zhang, J., Springenberg, J. T., Byravan, A., Hasenclever, L., Abdolmaleki, A., Rao, D., Heess, N., & Riedmiller, M. (2023). Leveraging jumpy models for planning and fast learning in robotic domains. arXiv preprint. arXiv:2302.12617."},{"key":"94_CR26","first-page":"1","volume-title":"Proceedings of the 38th international conference on neural information processing systems","author":"Z. Dong","year":"2024","unstructured":"Dong, Z., Yuan, Y., Hao, J., Ni, F., Ma, Y., Li, P., & Zheng, Y. (2024). Cleandiffuser: an easy-to-use modularized library for diffusion models in decision making. In A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, & C. Zhang (Eds.), Proceedings of the 38th international conference on neural information processing systems (pp. 1\u201328). Red Hook: Curran Associates."},{"key":"94_CR27","first-page":"1","volume-title":"Proceedings of the 38th international conference on neural information processing systems","author":"Z. Dong","year":"2024","unstructured":"Dong, Z., Hao, J., Yuan, Y., Ni, F., Wang, Y., Li, P., & Zheng, Y. (2024). Diffuserlite: towards real-time diffusion planning. In A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, & C. Zhang (Eds.), Proceedings of the 38th international conference on neural information processing systems (pp. 1\u201328). Red Hook: Curran Associates."},{"key":"94_CR28","doi-asserted-by":"publisher","first-page":"300","DOI":"10.1038\/nn1010","volume":"6","author":"A. d\u2019Avella","year":"2003","unstructured":"d\u2019Avella, A., Saltiel, P., & Bizzi, E. (2003). Combinations of muscle synergies in the construction of a natural motor behavior. Nature Neuroscience, 6, 300\u2013308.","journal-title":"Nature Neuroscience"},{"key":"94_CR29","doi-asserted-by":"publisher","first-page":"622","DOI":"10.1016\/j.conb.2008.01.002","volume":"17","author":"L. H. Ting","year":"2007","unstructured":"Ting, L. H., & McKay, J. L. (2007). Neuromechanics of muscle synergies for posture and movement. Current Opinion in Neurobiology, 17, 622\u2013628.","journal-title":"Current Opinion in Neurobiology"},{"key":"94_CR30","first-page":"1940","volume-title":"Proceedings of the 2016 IEEE\/RSJ international conference on intelligent robots and systems","author":"F. Ficuciello","year":"2016","unstructured":"Ficuciello, F., Zaccara, D., & Siciliano, B. (2016). Synergy-based policy improvement with path integrals for anthropomorphic hands. In Proceedings of the 2016 IEEE\/RSJ international conference on intelligent robots and systems (pp. 1940\u20131945). Piscataway: IEEE."},{"key":"94_CR31","unstructured":"Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et\u00a0al. (2018). DeepMind control suite. arXiv preprint. arXiv:1801.00690."},{"key":"94_CR32","first-page":"1","volume-title":"Proceedings of international conference on learning representations","author":"D. Hafner","year":"2021","unstructured":"Hafner, D., Lillicrap, T. P., Norouzi, M., & Ba, J. (2021). Mastering Atari with discrete world models. In Proceedings of international conference on learning representations (pp. 1\u201326). Retrieved September 10, 2025, from https:\/\/openreview.net\/pdf?id=0oabwyZbOu."},{"key":"94_CR33","doi-asserted-by":"publisher","first-page":"604","DOI":"10.1038\/s41586-020-03051-4","volume":"588","author":"J. Schrittwieser","year":"2020","unstructured":"Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al. (2020). Mastering Atari, Go, chess and shogi by planning with a learned model. Nature, 588, 604\u2013609.","journal-title":"Nature"},{"key":"94_CR34","first-page":"1","volume-title":"Proceedings of the 6th international conference on learning representations","author":"V. Pong","year":"2018","unstructured":"Pong, V., Gu, S., Dalal, M., & Levine, S. (2018). Temporal difference models: model-free deep RL for model-based control. In Proceedings of the 6th international conference on learning representations (pp. 1\u201328). Retrieved September 10, 2025, from https:\/\/openreview.net\/pdf?id=Skw0n-W0Z."},{"key":"94_CR35","first-page":"2455","volume-title":"Proceedings of the 32nd international conference on neural information processing systems","author":"D. Ha","year":"2018","unstructured":"Ha, D., & Schmidhuber, J. (2018). Recurrent world models facilitate policy evolution. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, & R. Garnett (Eds.), Proceedings of the 32nd international conference on neural information processing systems (pp. 2455\u20132467). Red Hook: Curran Associates."},{"key":"94_CR36","unstructured":"Feinberg, V., Wan, A., Stoica, I., Jordan, M. I., Gonzalez, J. E., & Levine, S. (2018). Model-based value estimation for efficient model-free reinforcement learning. arXiv preprint. arXiv:1803.00101."},{"key":"94_CR37","first-page":"142","volume-title":"Proceedings of the 18th European conference on computer vision","author":"Q. Li","year":"2024","unstructured":"Li, Q., Jia, X., Wang, S., & Yan, J. (2024). Think2Drive: efficient reinforcement learning by thinking with latent world model for autonomous driving. In A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, & G. Varol (Eds.), Proceedings of the 18th European conference on computer vision (pp. 142\u2013158). Cham: Springer."},{"key":"94_CR38","doi-asserted-by":"publisher","first-page":"3612","DOI":"10.1109\/LRA.2020.2976272","volume":"5","author":"B. Thananjeyan","year":"2020","unstructured":"Thananjeyan, B., Balakrishna, A., Rosolia, U., Li, F., McAllister, R., Gonzalez, J. E., Levine, S., Borrelli, F., & Goldberg, K. (2020). Safety augmented value estimation from demonstrations: safe deep model-based RL for sparse cost robotic tasks. IEEE Robotics and Automation Letters, 5, 3612\u20133619.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"94_CR39","first-page":"7207","volume-title":"International conference on machine learning","author":"S. Nair","year":"2020","unstructured":"Nair, S., Savarese, S., & Finn, C. (2020). Goal-aware prediction: learning to model what matters. In I. H. Daum\u00e9 & S. Aarti (Eds.), International conference on machine learning (pp. 7207\u20137219). Retrieved September 10, 2025, from http:\/\/proceedings.mlr.press\/v119\/nair20a\/nair20a.pdf."},{"key":"94_CR40","first-page":"195","volume-title":"Proceedings of the 1st annual conference on robot learning","author":"G. Kalweit","year":"2017","unstructured":"Kalweit, G., & Boedecker, J. (2017). Uncertainty-driven imagination for continuous deep reinforcement learning. In Proceedings of the 1st annual conference on robot learning (pp. 195\u2013206). Retrieved September 10, 2025, from https:\/\/proceedings.mlr.press\/v78\/kalweit17a\/kalweit17a.pdf."},{"key":"94_CR41","first-page":"8583","volume-title":"Proceedings of the international conference on machine learning","author":"R. Sekar","year":"2020","unstructured":"Sekar, R., Rybkin, O., Daniilidis, K., Abbeel, P., Hafner, D., & Pathak, D. (2020). Planning to explore via self-supervised world models. In I. H. Daum\u00e9 & S. Aarti (Eds.), Proceedings of the international conference on machine learning (pp. 8583\u20138592). Retrieved September 10, 2025, from http:\/\/proceedings.mlr.press\/v119\/sekar20a\/sekar20a.pdf."},{"key":"94_CR42","unstructured":"Whitney, W., & Fergus, R. (2018). Understanding the asymptotic performance of model-based RL methods. arXiv preprint. arXiv:1807.04022."},{"key":"94_CR43","first-page":"2459","volume-title":"Proceedings of the international conference on machine learning","author":"N. Mishra","year":"2017","unstructured":"Mishra, N., Abbeel, P., & Mordatch, I. (2017). Prediction and control with temporal segment models. In P. Doina & T. Whye (Eds.), Proceedings of the international conference on machine learning (pp. 2459\u20132468). Retrieved September 10, 2025, from https:\/\/proceedings.mlr.press\/v70\/mishra17a\/mishra17a.pdf."},{"key":"94_CR44","unstructured":"Wu, Y.-H., Fan, T.-H., Ramadge, P. J., & Su, H. (2019). Model imitation for model-based reinforcement learning. arXiv preprint. arXiv:1909.11821."},{"key":"94_CR45","unstructured":"Zihan, D., Amy, Z., Yuandong, T., & Qinqing, Z. (2024). Diffusion world model. arXiv preprint. arXiv:2402.03570."},{"key":"94_CR46","doi-asserted-by":"crossref","first-page":"229","DOI":"10.1017\/S0263574799281520","volume":"17","author":"R. S. Sutton","year":"1999","unstructured":"Sutton, R. S., & Barto, A. G. (1999). Reinforcement learning: an introduction. Robotica, 17, 229\u2013235.","journal-title":"Robotica"},{"key":"94_CR47","first-page":"8387","volume-title":"Proceedings of the international conference on machine learning","author":"N. Hansen","year":"2022","unstructured":"Hansen, N., Su, H., & Wang, X. (2022). Temporal difference learning for model predictive control. In C. Kamalika, J. Stefanie, S. Le, S. Csaba, N. Gang, & S. Sivan (Eds.), Proceedings of the international conference on machine learning (pp. 8387\u20138406). Retrieved September 10, 2025, from https:\/\/proceedings.mlr.press\/v162\/hansen22a."},{"key":"94_CR48","first-page":"1","volume-title":"Proceedings of the 4th international conference on learning representations","author":"T. P. Lillicrap","year":"2016","unstructured":"Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., & Wierstra, D. (2016). Continuous control with deep reinforcement learning. In Proceedings of the 4th international conference on learning representations (pp. 1\u201314). Retrieved September 10, 2025, from https:\/\/openreview.net\/pdf?id=kJP8gA8BxRY."},{"key":"94_CR49","doi-asserted-by":"publisher","first-page":"1347","DOI":"10.1162\/089976602753712972","volume":"14","author":"K. Doya","year":"2002","unstructured":"Doya, K., Samejima, K., Katagiri, K., & Kawato, M. (2002). Multiple model-based reinforcement learning. Neural Computation, 14, 1347\u20131369.","journal-title":"Neural Computation"},{"key":"94_CR50","unstructured":"Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., & Zaremba, W. (2016). OpenAI gym. arXiv preprint. arXiv:1606.01540."},{"key":"94_CR51","unstructured":"Grigsby, J., & Qi, Y. (2020). Measuring visual generalization in continuous control from pixels. arXiv preprint. arXiv:2010.06740."},{"key":"94_CR52","first-page":"1332","volume-title":"Proceedings of the conference on robot learning","author":"Y. Seo","year":"2023","unstructured":"Seo, Y., Hafner, D., Liu, H., Liu, F., James, S., Lee, K., & Abbeel, P. (2023). Masked world models for visual control. In Proceedings of the conference on robot learning (pp. 1332\u20131344). Retrieved September 10, 2025, from https:\/\/openreview.net\/forum?id=Bf6on28H0Jv."},{"key":"94_CR53","first-page":"507","volume-title":"Proceedings of the 37th international conference on machine learning","author":"A. P. Badia","year":"2020","unstructured":"Badia, A. P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, Z. D., & Blundell, C. (2020). Agent57: outperforming the Atari human benchmark. In I. H. Daum\u00e9 & S. Aarti (Eds.), Proceedings of the 37th international conference on machine learning (pp. 507\u2013517). Retrieved September 10, 2025, from https:\/\/proceedings.mlr.press\/v119\/badia20a\/badia20a.pdf."},{"key":"94_CR54","first-page":"139","volume-title":"Proceedings of the 17th international conference on autonomous agents and multiagent systems","author":"S. Omidshafiei","year":"2018","unstructured":"Omidshafiei, S., Kim, D., Pazis, J., & How, J. P. (2018). Crossmodal attentive skill learner. In Proceedings of the 17th international conference on autonomous agents and multiagent systems (pp. 139\u2013146). Richland: International Foundation for Autonomous Agents and Multiagent Systems."}],"container-title":["Visual Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-025-00094-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44267-025-00094-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-025-00094-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,26]],"date-time":"2025-11-26T03:02:57Z","timestamp":1764126177000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44267-025-00094-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,26]]},"references-count":54,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,12]]}},"alternative-id":["94"],"URL":"https:\/\/doi.org\/10.1007\/s44267-025-00094-x","relation":{},"ISSN":["2097-3330","2731-9008"],"issn-type":[{"value":"2097-3330","type":"print"},{"value":"2731-9008","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,26]]},"assertion":[{"value":"10 April 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 October 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 October 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 November 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"23"}}