{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T09:59:38Z","timestamp":1777715978258,"version":"3.51.4"},"reference-count":57,"publisher":"SAGE Publications","issue":"4-5","license":[{"start":{"date-parts":[[2022,6,13]],"date-time":"2022-06-13T00:00:00Z","timestamp":1655078400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"name":"ANU Futures Scheme","award":["QCE20102"],"award-info":[{"award-number":["QCE20102"]}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of Robotics Research"],"published-print":{"date-parts":[[2023,4]]},"abstract":"<jats:p>Planning under partial observability is essential for autonomous robots. A principled way to address such planning problems is the Partially Observable Markov Decision Process (POMDP). Although solving POMDPs is computationally intractable, substantial advancements have been achieved in developing approximate POMDP solvers in the past two decades. However, computing robust solutions for systems with complex dynamics remains challenging. Most on-line solvers rely on a large number of forward simulations and standard Monte Carlo methods to compute the expected outcomes of actions the robot can perform. For systems with complex dynamics, for example, those with non-linear dynamics that admit no closed-form solution, even a single forward simulation can be prohibitively expensive. Of course, this issue exacerbates for problems with long planning horizons. This paper aims to alleviate the above difficulty. To this end, we propose a new on-line POMDP solver, called Multilevel POMDP Planner\u2009(MLPP), that combines the commonly known Monte-Carlo-Tree-Search with the concept of Multilevel Monte Carlo to speed up our capability in generating approximately optimal solutions for POMDPs with complex dynamics. Experiments on four different problems involving torque control, navigation and grasping indicate that MLPP\u2009substantially outperforms state-of-the-art POMDP solvers.<\/jats:p>","DOI":"10.1177\/02783649221093658","type":"journal-article","created":{"date-parts":[[2022,6,13]],"date-time":"2022-06-13T12:59:30Z","timestamp":1655125170000},"page":"196-213","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":3,"title":["Multilevel Monte Carlo for solving POMDPs on-line"],"prefix":"10.1177","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4698-5875","authenticated-orcid":false,"given":"Marcus","family":"Hoerger","sequence":"first","affiliation":[{"name":"School of Mathematics and Physics, The University of Queensland, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hanna","family":"Kurniawati","sequence":"additional","affiliation":[{"name":"School of Computing, The Australian National University, Canberra, ACT, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alberto","family":"Elfes","sequence":"additional","affiliation":[{"name":"Robotics and Autonomous Systems Group, Data61, CSIRO, Pullenvale, QLD, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2022,6,13]]},"reference":[{"key":"bibr1-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2011.6095010"},{"key":"bibr2-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1137\/110840546"},{"key":"bibr3-02783649221093658","volume-title":"Advances in Neural Information Processing Systems, volume 23","author":"Araya M","year":"2010"},{"key":"bibr4-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1109\/78.978374"},{"key":"bibr5-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1023\/A:1013689704352"},{"key":"bibr6-02783649221093658","doi-asserted-by":"crossref","first-page":"1","DOI":"10.3390\/robotics1010001","volume":"1","author":"Bai H","year":"2012","journal-title":"Robotics: Science and Systems VII"},{"key":"bibr7-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1177\/0278364914528255"},{"key":"bibr8-02783649221093658","volume-title":"Neuro-Dynamic Programming","author":"Bertsekas DP","year":"1996"},{"key":"bibr9-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2016.03.027"},{"key":"bibr10-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1177\/0278364920937074"},{"key":"bibr11-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-25566-3_32"},{"key":"bibr12-02783649221093658","first-page":"3177","volume-title":"International Conference on Machine Learning","author":"Fischer J","year":"2020"},{"key":"bibr13-02783649221093658","unstructured":"Garg NP, Hsu D, Lee WS (2019) Despot-alpha: online pomdp planning with large state and observation spaces. In: Proc. of Robotics: Science and Systems, Freiburg im Breisgau, Germany, 22\u201326 June 2019."},{"key":"bibr14-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1287\/opre.1070.0496"},{"key":"bibr15-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1017\/S096249291500001X"},{"key":"bibr16-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-45346-6_5"},{"key":"bibr17-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48506.2021.9560943"},{"key":"bibr18-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-43089-4_18"},{"key":"bibr19-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2018.8593714"},{"key":"bibr20-02783649221093658","unstructured":"Hoerger M, Kurniawati H, Elfes A (2019a) Multilevel Monte-Carlo for solving pomdps online. In: Proc. International Symposium on Robotics Research (ISRR), Hanoi, Vietnam, 6\u201310 October 2019."},{"key":"bibr21-02783649221093658","first-page":"698","volume-title":"Proc. AAAI International Conference on Autonomous Planning and Scheduling (ICAPS)","author":"Hoerger M","year":"2019"},{"key":"bibr22-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2013.6631031"},{"key":"bibr23-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1109\/ROBOT.2007.364201"},{"key":"bibr24-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1109\/ROBOT.2005.1570712"},{"key":"bibr25-02783649221093658","unstructured":"Klimenko D, Song J, Kurniawati H (2014) Tapir: a software toolkit for approximating and adapting pomdp solutions online. In: Proceedings of the Australasian Conference on Robotics and Automation, Melbourne, Australia, 2\u20134 December 2014."},{"key":"bibr26-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1007\/11871842_29"},{"key":"bibr27-02783649221093658","author":"Kurniawati H","year":"2022","journal-title":"Annual Review of Control, Robotics, and Autonomous Systems To Appear"},{"key":"bibr28-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1177\/0278364910386986"},{"key":"bibr29-02783649221093658","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2008.IV.009"},{"key":"bibr30-02783649221093658","first-page":"611","volume-title":"Robotics Research. Springer Tracts in Advanced Robotics, vol 114","author":"Kurniawati H","year":"2013"},{"key":"bibr31-02783649221093658","author":"Lim MH","year":"2020","journal-title":"arXiv Preprint arXiv:2012.10140"},{"key":"bibr32-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1137\/0311025"},{"key":"bibr33-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1177\/0278364918780322"},{"key":"bibr34-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i13.17411"},{"key":"bibr35-02783649221093658","first-page":"316","volume-title":"International Conference on Numerical Methods and Applications","author":"Mihaylova L","year":"2002"},{"key":"bibr36-02783649221093658","volume-title":"Monte Carlo Theory, Methods and Examples","author":"Owen AB","year":"2013"},{"key":"bibr37-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1287\/moor.12.3.441"},{"key":"bibr38-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1137\/17M1159208"},{"key":"bibr39-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1137\/16M1082469"},{"key":"bibr40-02783649221093658","unstructured":"Pineau J, Gordon G, Thrun S (2003) Point-based value iteration: an anytime algorithm for POMDPs."},{"key":"bibr41-02783649221093658","first-page":"17","volume-title":"Proceedings of the Winter Simulation Conference","author":"Rhee Ch","year":"2012"},{"key":"bibr42-02783649221093658","volume-title":"The Cross-Entropy Method: A Unified Approach to Combinatorial Optimization, Monte-Carlo Simulation and Machine Learning","author":"Rubinstein RY","year":"2013"},{"key":"bibr43-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2015.7139503"},{"key":"bibr44-02783649221093658","unstructured":"Silver D, Veness J (2010) Monte-carlo planning in large POMDPs. In: Advances in Neural Information Processing Systems: 23rd Annual Conference on Neural Information Processing Systems 2009, Vancouver, Canada, 7\u201310 December 2009. pp. 2164\u20132172."},{"key":"bibr45-02783649221093658","unstructured":"Smith R (2001) Open dynamics engine. http:\/\/www.ode.org\/"},{"key":"bibr46-02783649221093658","unstructured":"Smith T, Simmons R (2005) Point-based POMDP algorithms: improved analysis and implementation."},{"key":"bibr47-02783649221093658","unstructured":"Somani A, Ye N, Hsu D, et al. (2013) Despot: online pomdp planning with regularization. In: Advances in Neural Information Processing Systems: 27th Annual Conference on Neural Information Processing Systems 2013, Lake Tahoe, NV, 5\u201310 December 2013 pp. 1772\u20131780."},{"key":"bibr48-02783649221093658","volume-title":"The Optimal Control of Partially Observable Markov Decision Processes","author":"Sondik EJ","year":"1971"},{"key":"bibr49-02783649221093658","volume-title":"Robot Modeling and Control, Volume 3","author":"Spong MW","year":"2006"},{"key":"bibr50-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2014.2380273"},{"key":"bibr51-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1609\/icaps.v28i1.13882"},{"key":"bibr52-02783649221093658","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton R","year":"2012"},{"key":"bibr53-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1177\/0278364911406562"},{"key":"bibr54-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1177\/0278364912456319"},{"key":"bibr55-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8793888"},{"key":"bibr56-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1609\/icaps.v28i1.13906"},{"key":"bibr57-02783649221093658","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992698"}],"container-title":["The International Journal of Robotics Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649221093658","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/02783649221093658","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/02783649221093658","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T10:17:03Z","timestamp":1777457823000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/02783649221093658"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,13]]},"references-count":57,"journal-issue":{"issue":"4-5","published-print":{"date-parts":[[2023,4]]}},"alternative-id":["10.1177\/02783649221093658"],"URL":"https:\/\/doi.org\/10.1177\/02783649221093658","relation":{},"ISSN":["0278-3649","1741-3176"],"issn-type":[{"value":"0278-3649","type":"print"},{"value":"1741-3176","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,13]]}}}