{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T17:35:08Z","timestamp":1782408908723,"version":"3.54.5"},"reference-count":41,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2024,11,17]],"date-time":"2024-11-17T00:00:00Z","timestamp":1731801600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Coverage Path Planning (CPP) in unknown environments presents unique challenges that often require the system to maintain a symmetry between exploration and exploitation in order to efficiently cover unknown areas. This paper introduces latent imagination-based reinforcement learning (LIRL), a novel framework that addresses these challenges by integrating three key components: memory-augmented experience replay (MAER), a latent imagination module (LIM), and multi-step prediction learning (MSPL) within a soft actor\u2013critic architecture. MAER enhances sample efficiency by prioritizing experience retrieval, LIM facilitates long-term planning via simulated trajectories, and MSPL optimizes the trade-off between immediate rewards and future outcomes through adaptive n-step learning. MAER, LIM, and MSPL work within a soft actor\u2013critic architecture, and LIRL creates a dynamic equilibrium that enables efficient, adaptive decision-making. We evaluate LIRL across diverse simulated environments, demonstrating substantial improvements over state-of-the-art methods. Through this method, the agent optimally balances short-term actions with long-term planning, maintaining symmetrical responses to varying environmental changes. The results highlight LIRL\u2019s potential for advancing autonomous CPP in real-world applications such as search and rescue, agricultural robotics, and warehouse automation. Our work contributes to the broader fields of robotics and reinforcement learning, offering insights into integrating memory, imagination, and adaptive learning for complex sequential decision-making tasks.<\/jats:p>","DOI":"10.3390\/sym16111537","type":"journal-article","created":{"date-parts":[[2024,11,18]],"date-time":"2024-11-18T05:36:45Z","timestamp":1731908205000},"page":"1537","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["LIRL: Latent Imagination-Based Reinforcement Learning for Efficient Coverage Path Planning"],"prefix":"10.3390","volume":"16","author":[{"given":"Zhenglin","family":"Wei","sequence":"first","affiliation":[{"name":"School of Electrical and Information Engineering, Zhengzhou University, Zhengzhou 450052, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tiejiang","family":"Sun","sequence":"additional","affiliation":[{"name":"School of Information Engineering, Chang\u2019an University, Xi\u2019an 710064, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mengjie","family":"Zhou","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Bristol, Bristol BS8 1QU, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,11,17]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Bormann, R., Jordan, F., Hampp, J., and H\u00e4gele, M. (2018, January 21\u201325). Indoor coverage path planning: Survey, implementation, analysis. Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia.","DOI":"10.1109\/ICRA.2018.8460566"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1258","DOI":"10.1016\/j.robot.2013.09.004","article-title":"A survey on coverage path planning for robotics","volume":"61","author":"Galceran","year":"2013","journal-title":"Robot. Auton. Syst."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"424","DOI":"10.1002\/rob.20388","article-title":"Coverage path planning on three-dimensional terrain for arable farming","volume":"28","author":"Jin","year":"2011","journal-title":"J. Field Robot."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Huang, Q. (2020, January 1\u20132). Model-based or model-free, a review of approaches in reinforcement learning. Proceedings of the 2020 International Conference on Computing and Data Science (CDS), Stanford, CA, USA.","DOI":"10.1109\/CDS49703.2020.00051"},{"key":"ref_5","unstructured":"Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., and Abbeel, P. (2018). Soft actor\u2013critic algorithms and applications. arXiv."},{"key":"ref_6","unstructured":"Sutton, R.S. (2018). Reinforcement learning: An introduction. A Bradford Book, MIT Press."},{"key":"ref_7","unstructured":"Finn, C., Abbeel, P., and Levine, S. (2017, January 6\u201311). Model-agnostic meta-learning for fast adaptation of deep networks. Proceedings of the International Conference on Machine Learning PMLR, 2017, Sydney, Australia."},{"key":"ref_8","unstructured":"Igl, M., Zintgraf, L., Le, T.A., Wood, F., and Whiteson, S. (2018, January 10\u201315). Deep variational reinforcement learning for POMDPs. Proceedings of the International Conference on Machine Learning PMLR, 2018, Stockholm, Sweden."},{"key":"ref_9","unstructured":"Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J. (2020). Mastering atari with discrete world models. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"471","DOI":"10.1038\/nature20101","article-title":"Hybrid computing using a neural network with dynamic external memory","volume":"538","author":"Graves","year":"2016","journal-title":"Nature"},{"key":"ref_11","unstructured":"Fortunato, M., Tan, M., Faulkner, R., Hansen, S., Puigdom\u00e8nech Badia, A., Buttimore, G., Deck, C., Leibo, J.Z., and Blundell, C. (2019). Generalization of reinforcement learners with working and episodic memory. Adv. Neural Inf. Process. Syst., 32."},{"key":"ref_12","unstructured":"Ha, D., and Schmidhuber, J. (2018). Recurrent world models facilitate policy evolution. Adv. Neural Inf. Process. Syst., 31."},{"key":"ref_13","unstructured":"Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R.H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., and Levine, S. (2019). Model-based reinforcement learning for atari. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"26","DOI":"10.1109\/MSP.2017.2743240","article-title":"Deep reinforcement learning: A brief survey","volume":"34","author":"Arulkumaran","year":"2017","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Faust, A., Oslund, K., Ramirez, O., Francis, A., Tapia, L., Fiser, M., and Davidson, J. (2018, January 21\u201325). Prm-rl: Long-range robotic navigation tasks by combining reinforcement learning and sampling-based planning. Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia.","DOI":"10.1109\/ICRA.2018.8461096"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Brown, S., and Waslander, S.L. (2016, January 9\u201314). The constriction decomposition method for coverage path planning. Proceedings of the 2016 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Daejeon, Republic of Korea.","DOI":"10.1109\/IROS.2016.7759499"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Cabreira, T.M., Ferreira, P.R., Di Franco, C., and Buttazzo, G.C. (2019, January 11\u201314). Grid-based coverage path planning with minimum energy over irregular-shaped areas with UAVs. Proceedings of the 2019 International Conference on Unmanned Aircraft Systems (ICUAS), Atlanta, GA, USA.","DOI":"10.1109\/ICUAS.2019.8797937"},{"key":"ref_18","first-page":"39","article-title":"An efficient complete coverage path planning in known environments","volume":"43","author":"Liu","year":"2011","journal-title":"J. Northeast Norm. Univ."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Theile, M., Bayerlein, H., Nai, R., Gesbert, D., and Caccamo, M. (2020, January 25\u201329). UAV coverage path planning under varying power constraints using deep reinforcement learning. Proceedings of the 2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA.","DOI":"10.1109\/IROS45743.2020.9340934"},{"key":"ref_20","first-page":"7424","article-title":"UAV Coverage Path Planning with Quantum-based Recurrent Deep Deterministic Policy Gradient","volume":"73","author":"Narottama","year":"2023","journal-title":"IEEE Trans. Veh. Technol."},{"key":"ref_21","unstructured":"Heydari, J., Saha, O., and Ganapathy, V. (2021). Reinforcement learning-based coverage path planning with implicit cellular decomposition. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"215","DOI":"10.32604\/csse.2023.031116","article-title":"Multi-agent dynamic area coverage based on reinforcement learning with connected agents","volume":"45","author":"Aydemir","year":"2023","journal-title":"Comput. Syst. Sci. Eng."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"359","DOI":"10.1561\/2200000049","article-title":"Bayesian reinforcement learning: A survey","volume":"8","author":"Ghavamzadeh","year":"2015","journal-title":"Found. Trends\u00ae Mach. Learn."},{"key":"ref_24","unstructured":"Brown, D., Coleman, R., Srinivasan, R., and Niekum, S. (2020, January 13\u201318). Safe imitation learning via fast bayesian reward inference from preferences. Proceedings of the International Conference on Machine Learning PMLR, Virtual Event."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"243","DOI":"10.1016\/j.inffus.2021.05.008","article-title":"A review of uncertainty quantification in deep learning: Techniques, applications and challenges","volume":"76","author":"Abdar","year":"2021","journal-title":"Inf. Fusion"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"3153","DOI":"10.1109\/LRA.2020.2974682","article-title":"A general framework for uncertainty estimation in deep learning","volume":"5","author":"Loquercio","year":"2020","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"3751","DOI":"10.1109\/LRA.2024.3368969","article-title":"Adaptive Neural Network-based Model Path-Following Contouring Control for Quadrotor Under Diversely Uncertain Disturbances","volume":"9","author":"Wei","year":"2024","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_28","unstructured":"Santoro, A., Bartunov, S., Botvinick, M., Wierstra, D., and Lillicrap, T. (2016, January 19\u201324). Meta-learning with memory-augmented neural networks. Proceedings of the International Conference on Machine Learning PMLR 2016, New York City, NY, USA."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"2468","DOI":"10.1038\/s41467-021-22364-0","article-title":"Robust high-dimensional memory-augmented neural networks","volume":"12","author":"Karunaratne","year":"2021","journal-title":"Nat. Commun."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Bae, H., Kim, G., Kim, J., Qian, D., and Lee, S. (2019). Multi-robot path planning method using reinforcement learning. Appl. Sci., 9.","DOI":"10.3390\/app9153057"},{"key":"ref_31","unstructured":"Wayne, G., Hung, C.C., Amos, D., Mirza, M., Ahuja, A., Grabska-Barwinska, A., Rae, J., Mirowski, P., Leibo, J.Z., and Santoro, A. (2018). Unsupervised predictive memory in a goal-directed agent. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Edgar, I. (2004). A Guide to Imagework: Imagination-Based Research Methods, Routledge.","DOI":"10.4324\/9780203490136"},{"key":"ref_33","unstructured":"Pascanu, R., Li, Y., Vinyals, O., Heess, N., Buesing, L., Racani\u00e8re, S., Reichert, D., Weber, T., Wierstra, D., and Battaglia, P. (2017). Learning model-based planning from scratch. arXiv."},{"key":"ref_34","unstructured":"Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M. (2019). Dream to control: Learning behaviors by latent imagination. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Liu, K., Stadler, M., and Roy, N. (August, January 31). Learned sampling distributions for efficient planning in hybrid geometric and object-level representations. Proceedings of the 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France.","DOI":"10.1109\/ICRA40945.2020.9196771"},{"key":"ref_36","unstructured":"Argenson, A., and Dulac-Arnold, G. (2020). Model-based offline planning. arXiv."},{"key":"ref_37","first-page":"30265","article-title":"The nature of temporal difference errors in multi-step distributional reinforcement learning","volume":"35","author":"Tang","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Schoknecht, R., and Riedmiller, M. (2002, January 28\u201330). Speeding-up reinforcement learning with multi-step actions. Proceedings of the Artificial Neural Networks\u2014ICANN 2002: International Conference, Madrid, Spain. Proceedings 12.","DOI":"10.1007\/3-540-46084-5_132"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"De Asis, K., Hernandez-Garcia, J., Holland, G., and Sutton, R. (2018, January 2\u20137). Multi-step reinforcement learning: A unifying algorithm. Proceedings of the AAAI Conference on Artificial Intelligence 2018, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11631"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Han, S., Chen, Y., Chen, G., Yin, J., Wang, H., and Cao, J. (2023, January 6\u20139). Multi-step reinforcement learning-based offloading for vehicle edge computing. Proceedings of the 2023 15th International Conference on Advanced Computational Intelligence (ICACI), Seoul, Republic of Korea.","DOI":"10.1109\/ICACI58115.2023.10146186"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Klamt, T., and Behnke, S. (2018, January 21\u201325). Planning hybrid driving-stepping locomotion on multiple levels of abstraction. Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia.","DOI":"10.1109\/ICRA.2018.8461054"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/16\/11\/1537\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T16:33:54Z","timestamp":1760114034000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/16\/11\/1537"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,17]]},"references-count":41,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2024,11]]}},"alternative-id":["sym16111537"],"URL":"https:\/\/doi.org\/10.3390\/sym16111537","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11,17]]}}}