{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T10:01:28Z","timestamp":1777716088707,"version":"3.51.4"},"reference-count":94,"publisher":"SAGE Publications","issue":"7","license":[{"start":{"date-parts":[[2025,1,19]],"date-time":"2025-01-19T00:00:00Z","timestamp":1737244800000},"content-version":"vor","delay-in-days":366,"URL":"http:\/\/www.sagepub.com\/licence-information-for-chorus"}],"funder":[{"DOI":"10.13039\/100000183","name":"Army Research Office","doi-asserted-by":"publisher","award":["W911NF-20-2-0099"],"award-info":[{"award-number":["W911NF-20-2-0099"]}],"id":[{"id":"10.13039\/100000183","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Amazon Research Award"},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["2006886"],"award-info":[{"award-number":["2006886"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["2047169"],"award-info":[{"award-number":["2047169"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of Robotics Research"],"published-print":{"date-parts":[[2024,6]]},"abstract":"<jats:p>We propose a diffusion approximation method to the continuous-state Markov decision processes that can be utilized to address autonomous navigation and control in unstructured off-road environments. In contrast to most decision-theoretic planning frameworks that assume fully known state transition models, we design a method that eliminates such a strong assumption that is often extremely difficult to engineer in reality. We first take the second-order Taylor expansion of the value function. The Bellman optimality equation is then approximated by a partial differential equation, which only relies on the first and second moments of the transition model. By combining the kernel representation of the value function, we design an efficient policy iteration algorithm whose policy evaluation step can be represented as a linear system of equations characterized by a finite set of supporting states. We first validate the proposed method through extensive simulations in 2 D obstacle avoidance and 2.5 D terrain navigation problems. The results show that the proposed approach leads to a much superior performance over several baselines. We then develop a system that integrates our decision-making framework with onboard perception and conduct real-world experiments in both cluttered indoor and unstructured outdoor environments. The results from the physical systems further demonstrate the applicability of our method in challenging real-world environments.<\/jats:p>","DOI":"10.1177\/02783649231225977","type":"journal-article","created":{"date-parts":[[2024,1,19]],"date-time":"2024-01-19T06:29:35Z","timestamp":1705645775000},"page":"1056-1080","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":3,"title":["Kernel-based diffusion approximated Markov decision processes for autonomous navigation and control on unstructured terrains"],"prefix":"10.1177","volume":"43","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7127-5093","authenticated-orcid":false,"given":"Junhong","family":"Xu","sequence":"first","affiliation":[{"name":"Indiana University-Bloomington, Bloomington, IN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kai","family":"Yin","sequence":"additional","affiliation":[{"name":"Expedia Group, Austin, TX, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zheng","family":"Chen","sequence":"additional","affiliation":[{"name":"Indiana University-Bloomington, Bloomington, IN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3929-6422","authenticated-orcid":false,"given":"Jason M","family":"Gregory","sequence":"additional","affiliation":[{"name":"Army Research Laboratory, Adelphi, MD, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ethan A","family":"Stump","sequence":"additional","affiliation":[{"name":"Army Research Laboratory, Adelphi, MD, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6796-6817","authenticated-orcid":false,"given":"Lantao","family":"Liu","sequence":"additional","affiliation":[{"name":"Indiana University-Bloomington, Bloomington, IN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2024,1,19]]},"reference":[{"key":"bibr1-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913501564"},{"key":"bibr2-02783649231225977","doi-asserted-by":"crossref","unstructured":"Al-Sabban H, Gonzalez LF, Smith RN (2013) Wind-energy based path planning for unmanned aerial vehicles using Markov decision processes. In 2013 IEEE International Conference on Robotics and Automation (ICRA), Karlsruhe, Germany, 6\u201310 May 2013.","DOI":"10.1109\/ICRA.2013.6630662"},{"key":"bibr3-02783649231225977","volume-title":"Nonlinear Model Predictive Control","author":"Allg\u00f6wer F","year":"2012"},{"key":"bibr4-02783649231225977","doi-asserted-by":"crossref","unstructured":"Althoff M, Stursberg O, Buss M (2008) Reachability analysis of nonlinear systems with uncertain parameters using conservative linearization. In 2008 47th IEEE Conference on Decision and Control, Cancun, Mexico, 9\u201311 December 2008.","DOI":"10.1109\/CDC.2008.4738704"},{"key":"bibr5-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-007-5038-2"},{"key":"bibr6-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2017.2743240"},{"key":"bibr7-02783649231225977","doi-asserted-by":"crossref","unstructured":"Baek SS, Kwon H, Yoder JA, et al. (2013) Optimal path planning of a target-following fixed-wing uav using sequential decision processes. In 2013 IEEE\/RSJ International Conference on Intelligent Robots and Systems, Tokyo, Japan, 03\u201307 November 2013.","DOI":"10.1109\/IROS.2013.6696775"},{"key":"bibr8-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1109\/CDC.2017.8263977"},{"key":"bibr9-02783649231225977","volume-title":"Adaptive Control Processes: A Guided Tour","author":"Bellman RE","year":"2015"},{"key":"bibr10-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1007\/BFb0109870"},{"key":"bibr11-02783649231225977","doi-asserted-by":"publisher","DOI":"10.3166\/ejc.11.310-334"},{"key":"bibr12-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1007\/s11768-011-1005-3"},{"key":"bibr13-02783649231225977","volume-title":"Dynamic Programming and Optimal Control","author":"Bertsekas D","year":"2012"},{"key":"bibr14-02783649231225977","volume-title":"Neuro-Dynamic Programming","author":"Bertsekas DP","year":"1996"},{"key":"bibr15-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1613\/jair.575"},{"issue":"2","key":"bibr16-02783649231225977","first-page":"631","volume":"68","author":"Braverman A","year":"2020","journal-title":"Operations Research"},{"key":"bibr17-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1109\/TIV.2017.2749181"},{"key":"bibr18-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2004.839232"},{"key":"bibr19-02783649231225977","doi-asserted-by":"crossref","unstructured":"Chen J, Su K, Shen S (2015) Real-time safe trajectory generation for quadrotor flight in cluttered environments. In 2015 IEEE International Conference on Robotics and Biomimetics (ROBIO), Zhuhai, China, 06\u201309 December 2015.","DOI":"10.1109\/ROBIO.2015.7419013"},{"key":"bibr20-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1609\/socs.v15i1.21753"},{"key":"bibr21-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2008.12.019"},{"key":"bibr22-02783649231225977","doi-asserted-by":"crossref","unstructured":"Deits R, Tedrake R (2015) Efficient mixed-integer planning for uavs in cluttered environments. In 2015 IEEE International Conference on Robotics and Automation (ICRA), Seattle, WA, USA, 26\u201330 May 2015.","DOI":"10.1109\/ICRA.2015.7138978"},{"key":"bibr23-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2017.12.012"},{"key":"bibr24-02783649231225977","unstructured":"Engel Y, Mannor S, Meir R (2003) Bayes meets bellman: the Gaussian process approach to temporal difference learning. Proceedings of the 20th International Conference on Machine Learning, Washington, DC, USA, 21\u201324 August 2003."},{"key":"bibr25-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1090\/gsm\/019"},{"key":"bibr26-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2018.2849506"},{"key":"bibr27-02783649231225977","doi-asserted-by":"crossref","unstructured":"Fu Y, Xiang Y, Zhang Y (2015) Sense and collision avoidance of unmanned aerial vehicles using markov decision process and flatness approach. In 2015 IEEE International Conference on Information and Automation, Lijiang, China, 08\u201310 August 2015.","DOI":"10.1109\/ICInfA.2015.7279378"},{"key":"bibr28-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1146\/annurev-control-061920-093753"},{"key":"bibr29-02783649231225977","doi-asserted-by":"crossref","unstructured":"Gao F, Shen S (2016) Online quadrotor trajectory generation and autonomous navigation on point clouds. In 2016 IEEE International Symposium on Safety, Security, and Rescue Robotics (SSRR), Lausanne, Switzerland, 23\u201327 October 2016.","DOI":"10.1109\/SSRR.2016.7784290"},{"key":"bibr30-02783649231225977","doi-asserted-by":"crossref","unstructured":"Gao F, Wu W, Lin Y, et al. (2018) Online safe trajectory generation for quadrotors using fast marching method and bernstein basis polynomial. In 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, QLD, Australia, 21\u201325 May 2018.","DOI":"10.1109\/ICRA.2018.8462878"},{"key":"bibr31-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1002\/rob.21842"},{"key":"bibr32-02783649231225977","volume-title":"Approximate Solutions to Markov Decision Processes","author":"Gordon G","year":"1999"},{"key":"bibr33-02783649231225977","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2015.XI.015"},{"key":"bibr34-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-27702-8_42"},{"key":"bibr35-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1109\/MITS.2010.939925"},{"key":"bibr36-02783649231225977","first-page":"2944","volume-title":"Advances in Neural Information Processing Systems","author":"Heess N","year":"2015"},{"key":"bibr37-02783649231225977","doi-asserted-by":"crossref","unstructured":"Hofmann T, Sch\u00f6lkopf B, AlexanderSmola J (2008) Kernel Methods in Machine Learning. Cambridge, MA: The Annals of Statistics, 1171\u20131220.","DOI":"10.1214\/009053607000000677"},{"key":"bibr38-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1177\/0278364906075328"},{"key":"bibr39-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1177\/0278364915616866"},{"key":"bibr40-02783649231225977","volume-title":"Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control","author":"Islam R","year":"2017"},{"key":"bibr41-02783649231225977","unstructured":"Johnson SG The NLopt nonlinear-optimization package. https:\/\/ab-initio.mit.edu\/nlopt"},{"issue":"2","key":"bibr42-02783649231225977","first-page":"102","volume":"5","author":"Kalman RE","year":"1960","journal-title":"Bol. soc. mat. mexicana"},{"key":"bibr43-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-012-5278-7"},{"key":"bibr44-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1177\/0278364911406761"},{"key":"bibr45-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-023-06419-4"},{"key":"bibr46-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913495721"},{"key":"bibr47-02783649231225977","first-page":"751","volume-title":"Advances in Neural Information Processing Systems","author":"Kuss M","year":"2004"},{"key":"bibr48-02783649231225977","first-page":"1107","volume":"4","author":"Lagoudakis MG","year":"2003","journal-title":"Journal of Machine Learning Research"},{"key":"bibr49-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1016\/j.automatica.2003.08.009"},{"key":"bibr50-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511546877"},{"key":"bibr51-02783649231225977","first-page":"262","volume":"5","author":"Likhachev M","year":"2005","journal-title":"ICAPS"},{"key":"bibr52-02783649231225977","volume-title":"Continuous Control With Deep Reinforcement Learning","author":"Lillicrap TP","year":"2015"},{"key":"bibr53-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2018.2801479"},{"key":"bibr54-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2017.2663526"},{"key":"bibr55-02783649231225977","volume-title":"Value Iteration in Continuous Actions, States and Time","author":"Lutter M","year":"2021"},{"key":"bibr56-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-36279-8_33"},{"key":"bibr57-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1177\/0278364917712421"},{"key":"bibr58-02783649231225977","volume-title":"Isaac Gym: High Performance Gpu-Based Physics Simulation for Robot Learning","author":"Makoviychuk V","year":"2021"},{"key":"bibr59-02783649231225977","volume-title":"Sparse Gaussian Process Temporal Difference Learning for Marine Robot Navigation","author":"Martin J","year":"2018"},{"key":"bibr60-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1023\/A:1022283719900"},{"issue":"5","key":"bibr61-02783649231225977","volume":"112","author":"McEwen A","year":"2007","journal-title":"Journal of Geophysical Research: Planets"},{"key":"bibr62-02783649231225977","doi-asserted-by":"crossref","unstructured":"Mellinger D, Kumar V (2011) Minimum snap trajectory generation and control for quadrotors. In 2011 IEEE International Conference on Robotics and Automation, Shanghai, China, 09\u201313 May 2011.","DOI":"10.1109\/ICRA.2011.5980409"},{"key":"bibr63-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.abk2822"},{"key":"bibr64-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1023\/A:1017992615625"},{"key":"bibr65-02783649231225977","first-page":"815","volume":"9","author":"Munos R","year":"2008","journal-title":"Journal of Machine Learning Research"},{"key":"bibr66-02783649231225977","doi-asserted-by":"crossref","unstructured":"Nair A, McGrew B, Andrychowicz M, et al. (2018a) Overcoming exploration in reinforcement learning with demonstrations. 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia, 21\u201325 May 2018.","DOI":"10.1109\/ICRA.2018.8463162"},{"key":"bibr67-02783649231225977","volume":"31","author":"Nair V","year":"2018","journal-title":"Advances in Neural Information Processing Systems"},{"key":"bibr68-02783649231225977","doi-asserted-by":"crossref","unstructured":"Oleynikova H, Burri M, Taylor Z, et al. (2016) Continuous-time trajectory optimization for online uav replanning. In 2016 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Daejeon, Korea, 09\u201314 October 2016.","DOI":"10.1109\/IROS.2016.7759784"},{"key":"bibr69-02783649231225977","doi-asserted-by":"crossref","unstructured":"Otte M, Silva W, Frew E (2016) Any-time path-planning: time-varying wind field+ moving obstacles. In 2016 IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden, 16\u201321 May 2016.","DOI":"10.1109\/ICRA.2016.7487414"},{"key":"bibr70-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1002\/rob.21472"},{"key":"bibr71-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1007\/s10479-012-1077-6"},{"key":"bibr72-02783649231225977","volume-title":"Markov Decision Processes: Discrete Stochastic Dynamic Programming","author":"Puterman ML","year":"2014"},{"key":"bibr73-02783649231225977","volume-title":"Model Predictive Control: Theory, Computation, and Design","author":"Rawlings JB","year":"2017"},{"key":"bibr74-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-28872-7_37"},{"key":"bibr75-02783649231225977","volume-title":"Proximal Policy Optimization Algorithms","author":"Schulman J","year":"2017"},{"key":"bibr76-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511809682"},{"key":"bibr77-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1109\/9780470544785"},{"key":"bibr78-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2022.3177279"},{"key":"bibr79-02783649231225977","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton RS","year":"2018"},{"key":"bibr80-02783649231225977","doi-asserted-by":"crossref","unstructured":"Taylor G, Parr R (2009) Kernelized value function approximation for reinforcement learning. Proceedings of the 26th Annual International Conference on Machine Learning, Montreal, Canada, 14 June 2009.","DOI":"10.1145\/1553374.1553504"},{"key":"bibr81-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1177\/0278364910369189"},{"key":"bibr82-02783649231225977","first-page":"3137","volume":"11","author":"Theodorou E","year":"2010","journal-title":"Journal of Machine Learning Research"},{"key":"bibr83-02783649231225977","volume-title":"Probabilistic robotics","author":"Thrun S","year":"2000"},{"key":"bibr84-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1177\/0278364911406562"},{"key":"bibr85-02783649231225977","unstructured":"Wang J, Triest S, Wang W, et al. (2021) Rough terrain navigation using divergence constrained model-based reinforcement learning. 5th Annual Conference on Robot Learning, London, UK, 8\u201311 November 2021."},{"key":"bibr86-02783649231225977","volume-title":"Kinodynamic Rrt*: Optimal Motion Planning for Systems With Linear Differential Constraints","author":"Webb DJ","year":"2012"},{"key":"bibr87-02783649231225977","doi-asserted-by":"crossref","unstructured":"Williams G, Nolan W, Goldfain B, et al. (2017) Information theoretic mpc for model-based reinforcement learning. In 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore, 29 May 2017.","DOI":"10.1109\/ICRA.2017.7989202"},{"key":"bibr88-02783649231225977","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2018.XIV.042"},{"key":"bibr89-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2007.899161"},{"key":"bibr90-02783649231225977","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2019.XV.069"},{"key":"bibr91-02783649231225977","doi-asserted-by":"publisher","DOI":"10.15607\/rss.2020.xvi.050"},{"key":"bibr92-02783649231225977","doi-asserted-by":"crossref","unstructured":"Yamauchi B (1997) A frontier-based approach for autonomous exploration. Proceedings 1997 IEEE International Symposium on Computational Intelligence in Robotics and Automation CIRA\u201997.\u2019Towards New Computational Principles for Robotics and Automation\u2019, Monterey, CA, USA, 10\u201311 July 1997.","DOI":"10.1109\/CIRA.1997.613851"},{"key":"bibr93-02783649231225977","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2019.2927938"},{"key":"bibr94-02783649231225977","doi-asserted-by":"crossref","unstructured":"Zhou B, Gao F, Pan J, et al. (2020) Robust real-time uav replanning using guided gradient-based optimization and topological paths. In 2020 IEEE International Conference on Robotics and Automation (ICRA), Paris, France, 31 May 2020.","DOI":"10.1109\/ICRA40945.2020.9196996"}],"container-title":["The International Journal of Robotics Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649231225977","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/02783649231225977","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649231225977","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649231225977","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T10:17:18Z","timestamp":1777457838000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/02783649231225977"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,19]]},"references-count":94,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2024,6]]}},"alternative-id":["10.1177\/02783649231225977"],"URL":"https:\/\/doi.org\/10.1177\/02783649231225977","relation":{},"ISSN":["0278-3649","1741-3176"],"issn-type":[{"value":"0278-3649","type":"print"},{"value":"1741-3176","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,19]]}}}