{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T03:26:05Z","timestamp":1781234765295,"version":"3.54.1"},"reference-count":34,"publisher":"MDPI AG","issue":"24","license":[{"start":{"date-parts":[[2021,12,16]],"date-time":"2021-12-16T00:00:00Z","timestamp":1639612800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2016YFC0301500"],"award-info":[{"award-number":["2016YFC0301500"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Path planning technology is significant for planetary rovers that perform exploration missions in unfamiliar environments. In this work, we propose a novel global path planning algorithm, based on the value iteration network (VIN), which is embedded within a differentiable planning module, built on the value iteration (VI) algorithm, and has emerged as an effective method to learn to plan. Despite the capability of learning environment dynamics and performing long-range reasoning, the VIN suffers from several limitations, including sensitivity to initialization and poor performance in large-scale domains. We introduce the double value iteration network (dVIN), which decouples action selection and value estimation in the VI module, using the weighted double estimator method to approximate the maximum expected value, instead of maximizing over the estimated action value. We have devised a simple, yet effective, two-stage training strategy for VI-based models to address the problem of high computational cost and poor performance in large-size domains. We evaluate the dVIN on planning problems in grid-world domains and realistic datasets, generated from terrain images of a moon landscape. We show that our dVIN empirically outperforms the baseline methods and generalize better to large-scale environments.<\/jats:p>","DOI":"10.3390\/s21248418","type":"journal-article","created":{"date-parts":[[2021,12,16]],"date-time":"2021-12-16T21:32:40Z","timestamp":1639690360000},"page":"8418","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Value Iteration Networks with Double Estimator for Planetary Rover Path Planning"],"prefix":"10.3390","volume":"21","author":[{"given":"Xiang","family":"Jin","sequence":"first","affiliation":[{"name":"School of Naval Architecture and Ocean Engineering, Dalian Maritime University, Dalian 116026, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wei","family":"Lan","sequence":"additional","affiliation":[{"name":"School of Naval Architecture and Ocean Engineering, Dalian Maritime University, Dalian 116026, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tianlin","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Naval Architecture and Ocean Engineering, Dalian Maritime University, Dalian 116026, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1453-0398","authenticated-orcid":false,"given":"Pengyao","family":"Yu","sequence":"additional","affiliation":[{"name":"School of Naval Architecture and Ocean Engineering, Dalian Maritime University, Dalian 116026, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2021,12,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1109\/MRA.2014.2381359","article-title":"The Right Path: Comprehensive path planning for lunar exploration rovers","volume":"22","author":"Sutoh","year":"2015","journal-title":"IEEE Robot. Autom. Mag."},{"key":"ref_2","unstructured":"Meila, M., and Zhang, T. (2021, January 18\u201324). Path Planning using Neural A* Search. Proceedings of the 38th International Conference on Machine Learning, Virtual."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Toma, A.I., Hsueh, H.Y., Jaafar, H.A., Murai, R., Kelly, P.H., and Saeedi, S. (2021, January 26\u201328). PathBench: A Benchmarking Platform for Classical and Learned Path Planning Algorithms. Proceedings of the 18th Conference on Robots and Vision, Burnaby, BC, Canada.","DOI":"10.1109\/CRV52889.2021.00019"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1016\/j.arcontrol.2020.10.001","article-title":"A comparative review on mobile robot path planning: Classical or meta-heuristic methods?","volume":"50","author":"Wahab","year":"2020","journal-title":"Annu. Rev. Control"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1314","DOI":"10.5897\/IJPS11.1745","article-title":"Optimal path planning of mobile robots: A review","volume":"7","author":"Raja","year":"2012","journal-title":"Int. J. Phys. Sci."},{"key":"ref_6","first-page":"1334","article-title":"End-to-end training of deep visuomotor policies","volume":"17","author":"Levine","year":"2016","journal-title":"J. Mach. Learn. Res."},{"key":"ref_7","unstructured":"Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep Learning, MIT Press."},{"key":"ref_8","unstructured":"Richard, S., and Andrew, B. (2018). Reinforcement Learning: An Introduction, MIT Press."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_10","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal policy optimization algorithms. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"604","DOI":"10.1038\/s41586-020-03051-4","article-title":"Mastering Atari, Go, chess and shogi by planning with a learned model","volume":"588","author":"Schrittwieser","year":"2020","journal-title":"Nature"},{"key":"ref_12","unstructured":"Hafner, D., Lillicrap, T.P., Norouzi, M., and Ba, J. (2021, January 3\u20137). Mastering Atari with discrete world models. Proceedings of the 9th International Conference on Learning Representations, Virtual."},{"key":"ref_13","first-page":"2154","article-title":"Value iteration networks","volume":"Volume 29","author":"Tamar","year":"2016","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Niu, S., Chen, S., Guo, H., Targonski, C., Smith, M.C., and Kovaevi, J. (2018, January 2\u20137). Generalized value iteration networks: Life beyond lattices. Proceedings of the 32nd AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12081"},{"key":"ref_15","unstructured":"Deac, A., Veli\u010dkovi\u0107, P., Milinkovic, O., Bacon, P.L., Tang, J., and Nikolic, M. (2020). XLVIN: EXecuted Latent Value Iteration Nets. arXiv."},{"key":"ref_16","unstructured":"Nardelli, N., Kohli, P., Synnaeve, G., Torr, P., Lin, Z., and Usunier, N. (2019, January 6\u20139). Value propagation networks. Proceedings of the 7th International Conference on Learning Representations, New Orleans, LA, USA."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zhang, L., Li, X., Chen, S., Zang, H., Huang, J., and Wang, M. (2020, January 7\u201312). Universal value iteration networks: When spatially-invariant is not universal. Proceedings of the 34th AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i04.6157"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the 29th IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_19","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 6\u201311). Batch Normalization: Accelerating deep network training by reducing internal covariate shift. Proceedings of the 32nd International Conference on Machine Learning, Lille, France."},{"key":"ref_20","first-page":"2613","article-title":"Double Q-learning","volume":"Volume 23","author":"Hasselt","year":"2010","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Hasselt, H., Guez, A., and Silver, D. (2016, January 12\u201317). Deep reinforcement learning with Double Q-Learning. Proceedings of the 30th AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA.","DOI":"10.1609\/aaai.v30i1.10295"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"69","DOI":"10.1016\/j.neucom.2019.05.075","article-title":"A novel learning-based global path planning algorithm for planetary rovers","volume":"361","author":"Zhang","year":"2019","journal-title":"Neurocomputing"},{"key":"ref_23","unstructured":"Bachlechner, T., Majumder, B.P., Mao, H.H.H., Cottrell, G., and McAuley, J. (2021, January 27\u201330). ReZero is all you need: Fast convergence at large depth. Proceedings of the 37th Conference on Uncertainty in Artificial Intelligence, Online."},{"key":"ref_24","unstructured":"Ba, J.L., Kiros, J.R., and Hinton, G.E. (2016). Layer Normalization. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1387","DOI":"10.1109\/LRA.2019.2895892","article-title":"Rover-IRL: Inverse reinforcement learning with soft value iteration networks for planetary rover path planning","volume":"4","author":"Pflueger","year":"2019","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_26","unstructured":"Lee, L., Parisotto, E., Chaplot, D., Xing, E., and Salakhutdinov, R. (2018, January 10\u201315). Gated path planning networks. Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden."},{"key":"ref_27","unstructured":"Agarwal, R., Schuurmans, D., and Norouzi, M. (2020, January 13\u201318). An optimistic perspective on offline deep reinforcement learning. Proceedings of the 37th International Conference on Machine Learning, Virtual."},{"key":"ref_28","unstructured":"Peer, O., Tessler, C., Merlis, N., and Meir, R. (2021, January 18\u201324). Ensemble bootstrapping for Q-Learning. Proceedings of the 38th International Conference on Machine Learning, Virtual."},{"key":"ref_29","unstructured":"Allen-Zhu, Z., and Li, Y. (2020). Towards understanding ensemble, knowledge distillation and self-distillation in deep learning. arXiv."},{"key":"ref_30","unstructured":"Murphy, K.P. (2012). Machine Learning: A Probabilistic Perspective, MIT Press."},{"key":"ref_31","first-page":"8026","article-title":"PyTorch: An imperative style, high-performance deep learning library","volume":"Volume 32","author":"Paszke","year":"2019","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_32","unstructured":"He, K., Girshick, R., and Dollar, P. (November, January 27). Rethinking ImageNet pre-training. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_33","unstructured":"Smith, L.N. (2018). A disciplined approach to neural network hyper-parameters: Part 1\u2014Learning rate, batch size, momentum, and weight decay. arXiv."},{"key":"ref_34","unstructured":"Tan, M., and Le, Q.V. (2019, January 9\u201315). EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of the 36th International Conference on Machine Learning, Long Beach, CA, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/24\/8418\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:50:06Z","timestamp":1760169006000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/24\/8418"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,16]]},"references-count":34,"journal-issue":{"issue":"24","published-online":{"date-parts":[[2021,12]]}},"alternative-id":["s21248418"],"URL":"https:\/\/doi.org\/10.3390\/s21248418","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12,16]]}}}