{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T15:25:22Z","timestamp":1784906722790,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":14,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,3,5]],"date-time":"2021-03-05T00:00:00Z","timestamp":1614902400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"CAAI-Huawei MindSpore Open Fund"},{"name":"Technological Innovation 2030?New Generation Artificial Intelligence","award":["No. 2018AAA0100500"],"award-info":[{"award-number":["No. 2018AAA0100500"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,3,5]]},"DOI":"10.1145\/3461353.3461365","type":"proceedings-article","created":{"date-parts":[[2021,9,6]],"date-time":"2021-09-06T17:21:15Z","timestamp":1630948875000},"page":"151-157","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Learning Navigation Policies for Mobile Robots in Deep Reinforcement Learning with Random Network Distillation"],"prefix":"10.1145","author":[{"given":"Lifan","family":"Pan","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anyi","family":"Li","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jun","family":"Ma","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianmin","family":"Ji","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,9,4]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/3157096.3157262"},{"key":"e_1_3_2_1_2_1","volume-title":"Proceedings of the 7th International Conference on Learning Representations, ICLR","author":"Burda Yuri","year":"2019","unstructured":"Yuri Burda , Harrison Edwards , Amos\u00a0 J. Storkey , and Oleg Klimov . 2019 . Exploration by random network distillation . In Proceedings of the 7th International Conference on Learning Representations, ICLR 2019. OpenReview.net. Yuri Burda, Harrison Edwards, Amos\u00a0J. Storkey, and Oleg Klimov. 2019. Exploration by random network distillation. In Proceedings of the 7th International Conference on Learning Representations, ICLR 2019. OpenReview.net."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.3390\/s20174836"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/3157382.3157509"},{"key":"e_1_3_2_1_5_1","volume-title":"Human-level control through deep reinforcement learning. Nat. 518, 7540","author":"Mnih Volodymyr","year":"2015","unstructured":"Volodymyr Mnih , Koray Kavukcuoglu , David Silver , Andrei\u00a0 A. Rusu , Joel Veness , Marc\u00a0 G. Bellemare , Alex Graves , Martin\u00a0 A. Riedmiller , Andreas Fidjeland , Georg Ostrovski , Stig Petersen , Charles Beattie , Amir Sadik , Ioannis Antonoglou , Helen King , Dharshan Kumaran , Daan Wierstra , Shane Legg , and Demis Hassabis . 2015. Human-level control through deep reinforcement learning. Nat. 518, 7540 ( 2015 ), 529\u2013533. Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei\u00a0A. Rusu, Joel Veness, Marc\u00a0G. Bellemare, Alex Graves, Martin\u00a0A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. 2015. Human-level control through deep reinforcement learning. Nat. 518, 7540 (2015), 529\u2013533."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305890.3305962"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TEVC.2006.890271"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305890.3305968"},{"key":"e_1_3_2_1_9_1","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal Policy Optimization Algorithms. CoRR abs\/1707.06347(2017). arxiv:1707.06347  John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal Policy Optimization Algorithms. CoRR abs\/1707.06347(2017). arxiv:1707.06347"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2019.2936167"},{"key":"e_1_3_2_1_11_1","unstructured":"Bradly\u00a0C. Stadie Sergey Levine and Pieter Abbeel. 2015. Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models. CoRR abs\/1507.00814(2015). arxiv:1507.00814  Bradly\u00a0C. Stadie Sergey Levine and Pieter Abbeel. 2015. Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models. CoRR abs\/1507.00814(2015). arxiv:1507.00814"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2017.8202134"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/3294996.3295035"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2017.8206049"}],"event":{"name":"ICIAI 2021: 2021 the 5th International Conference on Innovation in Artificial Intelligence","location":"Xia men China","acronym":"ICIAI 2021"},"container-title":["2021 the 5th International Conference on Innovation in Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3461353.3461365","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3461353.3461365","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:28:35Z","timestamp":1750195715000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3461353.3461365"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,5]]},"references-count":14,"alternative-id":["10.1145\/3461353.3461365","10.1145\/3461353"],"URL":"https:\/\/doi.org\/10.1145\/3461353.3461365","relation":{},"subject":[],"published":{"date-parts":[[2021,3,5]]},"assertion":[{"value":"2021-09-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}