{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T12:36:05Z","timestamp":1778070965575,"version":"3.51.4"},"reference-count":55,"publisher":"MDPI AG","issue":"13","license":[{"start":{"date-parts":[[2020,6,30]],"date-time":"2020-06-30T00:00:00Z","timestamp":1593475200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["2018XKQYMS03"],"award-info":[{"award-number":["2018XKQYMS03"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Deep reinforcement learning (DRL) has been successfully applied in mapless navigation. An important issue in DRL is to design a reward function for evaluating actions of agents. However, designing a robust and suitable reward function greatly depends on the designer\u2019s experience and intuition. To address this concern, we consider employing reward shaping from trajectories on similar navigation tasks without human supervision, and propose a general reward function based on matching network (MN). The MN-based reward function is able to gain the experience by pre-training through trajectories on different navigation tasks and accelerate the training speed of DRL in new tasks. The proposed reward function keeps the optimal strategy of DRL unchanged. The simulation results on two static maps show that the DRL converge with less iterations via the learned reward function than the state-of-the-art mapless navigation methods. The proposed method performs well in dynamic maps with partially moving obstacles. Even when test maps are different from training maps, the proposed strategy is able to complete the navigation tasks without additional training.<\/jats:p>","DOI":"10.3390\/s20133664","type":"journal-article","created":{"date-parts":[[2020,6,30]],"date-time":"2020-06-30T09:36:04Z","timestamp":1593509764000},"page":"3664","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["Learning Reward Function with Matching Network for Mapless Navigation"],"prefix":"10.3390","volume":"20","author":[{"given":"Qichen","family":"Zhang","sequence":"first","affiliation":[{"name":"Engineering Research Center of Intelligent Control for Underground Space, Ministry of Education, China University of Mining and Technology, Xuzhou 221116 China"},{"name":"The School of Information and Control Engineering, China University of Mining and Technology, Xuzhou 221116, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Meiqiang","family":"Zhu","sequence":"additional","affiliation":[{"name":"Engineering Research Center of Intelligent Control for Underground Space, Ministry of Education, China University of Mining and Technology, Xuzhou 221116 China"},{"name":"The School of Information and Control Engineering, China University of Mining and Technology, Xuzhou 221116, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liang","family":"Zou","sequence":"additional","affiliation":[{"name":"Engineering Research Center of Intelligent Control for Underground Space, Ministry of Education, China University of Mining and Technology, Xuzhou 221116 China"},{"name":"The School of Information and Control Engineering, China University of Mining and Technology, Xuzhou 221116, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ming","family":"Li","sequence":"additional","affiliation":[{"name":"Engineering Research Center of Intelligent Control for Underground Space, Ministry of Education, China University of Mining and Technology, Xuzhou 221116 China"},{"name":"The School of Information and Control Engineering, China University of Mining and Technology, Xuzhou 221116, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yong","family":"Zhang","sequence":"additional","affiliation":[{"name":"Engineering Research Center of Intelligent Control for Underground Space, Ministry of Education, China University of Mining and Technology, Xuzhou 221116 China"},{"name":"The School of Information and Control Engineering, China University of Mining and Technology, Xuzhou 221116, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,6,30]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1309","DOI":"10.1109\/TRO.2016.2624754","article-title":"Past, Present, and Future of Simultaneous Localization and Mapping: Toward the Robust-Perception Age","volume":"32","author":"Cadena","year":"2016","journal-title":"IEEE Trans. Robot."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Ort, T., Paull, L., and Rus, D. (2018, January 21\u201325). Autonomous vehicle navigation in rural environments without detailed prior maps. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia.","DOI":"10.1109\/ICRA.2018.8460519"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Lample, G., and Chaplot, D.S. (2017, January 4\u20139). Playing fps games with deep reinforcement learning. Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.10827"},{"key":"ref_4","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013). Playing Atari with Deep Reinforcement Learning. arXiv: Learning."},{"key":"ref_5","unstructured":"Gu, S., Holly, E., Lillicrap, T., and Levine, S. (June, January 29). Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. Proceedings of the International Conference on Robotics and Automation, Singapore."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Wang, C., Zhang, Q., Tian, Q., Li, S., Wang, X., Lane, D.M., Petillot, Y., and Wang, S. (2020). Learning Mobile Manipulation through Deep Reinforcement Learning. Sensors, 20.","DOI":"10.3390\/s20030939"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Kulhanek, J., Derner, E., De Bruin, T., and Babuska, R. (2019, January 4\u20136). Vision-based navigation using deep reinforcement learning. Proceedings of the 2019 European Conference on Mobile Robots (ECMR), Prague, Czech Republic.","DOI":"10.1109\/ECMR.2019.8870964"},{"key":"ref_8","unstructured":"Zhelo, O., Zhang, J., Tai, L., Liu, M., and Burgard, W. (2018). Curiosity-driven exploration for mapless navigation with deep reinforcement learning. arXiv: Robotics."},{"key":"ref_9","unstructured":"Mirowski, P., Grimes, M.K., Malinowski, M., Hermann, K.M., Anderson, K., and Teplyashin, D. (2018, January 3\u20138). Learning to navigate in cities without a map. Proceedings of the Advances in Neural Information Processing Systems, Montr\u00e9al, QC, Canada."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Hu, Z., Wan, K., Gao, X., Zhai, Y., and Wang, Q. (2020). Deep Reinforcement Learning Approach with Multiple Experience Pools for UAV\u2019s Autonomous Motion Planning in Complex Unknown Environments. Sensors, 20.","DOI":"10.3390\/s20071890"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Hu, B., Shao, S., Cao, Z., Xiao, Q., Li, Q., and Ma, C. (2019, January 6\u20138). Learning a Faster Locomotion Gait for a Quadruped Robot with Model-Free Deep Reinforcement Learning. Proceedings of the 2019 IEEE International Conference on Robotics and Biomimetics (ROBIO), Dali, China.","DOI":"10.1109\/ROBIO49542.2019.8961651"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Hussein, A., Elyan, E., Gaber, M.M., and Jayne, C. (2017, January 14\u201319). Deep reward shaping from demonstrations. Proceedings of the 2017 International Joint Conference on Neural Networks (IJCNN), Anchorage, AK, USA.","DOI":"10.1109\/IJCNN.2017.7965896"},{"key":"ref_13","first-page":"663","article-title":"Algorithms for inverse reinforcement learning","volume":"67","author":"Ng","year":"2000","journal-title":"Int. Conf. Mach. Learn."},{"key":"ref_14","unstructured":"Finn, C., Levine, S., and Abbeel, P. (2016, January 20\u201322). Guided cost learning: Deep inverse optimal control via policy optimization. Proceedings of the International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_15","unstructured":"Fu, J., Luo, K., and Levine, S. (2017). Learning robust rewards with adversarial inverse reinforcement learning. arXiv: Learning."},{"key":"ref_16","unstructured":"Ho, J., and Ermon, S. (2016, January 5\u201310). Generative adversarial imitation learning. Proceedings of the Advances in Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_17","unstructured":"Fu, J., Singh, A., Ghosh, D., Yang, L., and Levine, S. (2018, January 3\u20138). Variational inverse control with events: A general framework for data-driven reward definition. Proceedings of the Advances in Neural Information Processing Systems, Montr\u00e9al, QC, Canada."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zhou, W., and Li, W. (2018, January 14\u201317). Safety-aware apprenticeship learning. Proceedings of the International Conference on Computer Aided Verification, Oxford, UK.","DOI":"10.1007\/978-3-319-96145-3_38"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1608","DOI":"10.1177\/0278364910371999","article-title":"Autonomous helicopter aerobatics through apprenticeship learning","volume":"29","author":"Abbeel","year":"2010","journal-title":"Int. J. Robot. Res."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Singh, A., Yang, L., Finn, C., and Levine, S. (2019, January 22\u201326). End-to-end robotic reinforcement learning without reward engineering. Proceedings of the Robotics Science and Systems, Freiburg im Breisgau, Germany.","DOI":"10.15607\/RSS.2019.XV.073"},{"key":"ref_21","unstructured":"Xie, A., Singh, A., Levine, S., and Finn, C. (2018). Few-shot goal inference for visuomotor learning and planning. arXiv: Learning."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Vecerik, M., Sushkov, O., Barker, D., Rothorl, T., Hester, T., and Scholz, J. (2019, January 20\u201324). A practical approach to insertion with variable socket position using deep reinforcement learning. Proceedings of the International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8794074"},{"key":"ref_23","unstructured":"Zou, H., Ren, T., Yan, D., Su, H., and Zhu, J. (2019). Reward shaping via meta-learning. arXiv: Learning."},{"key":"ref_24","unstructured":"Yang, Y., Caluwaerts, K., Iscen, A., Tan, J., and Finn, C. (2019, January 13\u201317). Norml: No-reward meta learning. Proceedings of the International Foundation for Autonomous Agents and Multiagent Systems, Montreal, QC, Canada."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Sung, F., Yang, Y., Zhang, L., Xiang, T., Torr, P.H., and Hospedales, T.M. (2018, January 18\u201323). Learning to compare: Relation network for few-shot learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00131"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Sun, Q., Liu, Y., Chua, T.S., and Schiele, B. (2019, January 15\u201320). Meta-transfer learning for few-shot learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00049"},{"key":"ref_27","first-page":"1-1","article-title":"Few-shot Learning for Domain-specific Fine-grained Image Classification","volume":"99","author":"Sun","year":"2020","journal-title":"IEEE Trans. Ind. Electron."},{"key":"ref_28","unstructured":"Liu, Y., Lee, J., Park, M., Kim, S., Yang, E., Hwang, S.J., and Yang, Y. (2019, January 6\u20139). Learning to propagate labels: Transductive propagation network for few-shot learning. Proceedings of the International Conference on Learning Representations, New Orleans, LA, USA."},{"key":"ref_29","unstructured":"Bertinetto, L., Henriques, J.F., Valmadre, J., Torr, P.H., and Vedaldi, A. (2016, January 5\u201310). Learning feed-forward one-shot learners. Proceedings of the Advances in Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_30","unstructured":"Vinyals, O., Blundell, C., Lillicrap, T., Kavukcuoglu, K., and Wierstra, D. (2016, January 5\u201310). Matching networks for one shot learning. Proceedings of the Advances in Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"463","DOI":"10.1017\/S0263574714000289","article-title":"Algorithms for collision-free navigation of mobile robots in complex cluttered environments: A survey","volume":"33","author":"Hoy","year":"2015","journal-title":"Robotica"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"28","DOI":"10.1109\/MCI.2006.329691","article-title":"Ant colony optimization: Artificial ants as a computational intelligence technique","volume":"1","author":"Dorigo","year":"2016","journal-title":"IEEE Comput. Intell. Mag."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1515\/mms-2017-0021","article-title":"Sequential Classification of Palm Gestures Based on A* Algorithm and MLP Neural Network for Quadrocopter Control","volume":"24","author":"Wodzinski","year":"2017","journal-title":"Metrol. Meas. Syst."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"789","DOI":"10.1016\/S0005-1098(99)00214-9","article-title":"Survey Constrained model predictive control: Stability and optimality","volume":"36","author":"Mayne","year":"2000","journal-title":"Automatica"},{"key":"ref_35","unstructured":"Shi, E., Cai, T., He, C., and Guo, J. (2007, January 18\u201321). Study of the New Method for Improving Artifical Potential Field in Mobile Robot Obstacle Avoidance. Proceedings of the International Conference on Automation and Logistics, Jinan, China."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1109\/100.580977","article-title":"The dynamic window approach to collision avoidance","volume":"4","author":"Fox","year":"1997","journal-title":"IEEE Robot. Autom. Mag."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"582","DOI":"10.1016\/j.dt.2019.04.011","article-title":"A review: On path planning strategies for navigation of mobile robot","volume":"15","author":"Patle","year":"2019","journal-title":"Def. Technol."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"132","DOI":"10.1109\/TII.2012.2198665","article-title":"Comparison of parallel genetic algorithm and particle swarm optimization for real-time UAV path planning","volume":"9","author":"Roberge","year":"2012","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"174","DOI":"10.1016\/j.rcim.2006.10.001","article-title":"Fuzzy logic path planning for the robotic placement of fabrics on a work table","volume":"24","author":"Zoumponos","year":"2008","journal-title":"Robot. Comput. Integr. Manuf."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"799","DOI":"10.1016\/j.apenergy.2018.03.104","article-title":"Continuous reinforcement learning of energy management with deep Q network for a power split hybrid electric bus","volume":"222","author":"Wu","year":"2018","journal-title":"Appl. Energy"},{"key":"ref_41","unstructured":"Li, S., Wu, Y., Cui, X., Dong, H., Fang, F., and Russell, S. (February, January 27). Robust Multi-Agent Reinforcement Learning via Minimax Deep Deterministic Policy Gradient. Proceedings of the National Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Wang, X., Zhuang, Z., Zou, L., and Zhang, W. (2019, January 27\u201330). An accelerated asynchronous advantage actor-critic algorithm applied in papermaking. Proceedings of the Chinese Control Conference, Guangzhou, China.","DOI":"10.23919\/ChiCC.2019.8866243"},{"key":"ref_43","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal policy optimization algorithms. arXiv: Learning."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Long, P., Fanl, T., Liao, X., Liu, W., Zhang, H., and Pan, J. (2018, January 21\u201325). Towards optimally decentralized multi-robot collision avoidance via deep reinforcement learning. Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia.","DOI":"10.1109\/ICRA.2018.8461113"},{"key":"ref_45","unstructured":"Ma, L., Chen, J., and Liu, Y. (2019). Using RGB Image as Visual Input for Mapless Robot Navigation. arXiv preprint."},{"key":"ref_46","unstructured":"Mirowski, P., Pascanu, R., Viola, F., Soyer, H., Ballard, A.J., Banino, A., Denil, M., Goroshin, R., Sifre, L., and Kavukcuoglu, K. (2016). Learning to Navigate in Complex Environments. arXiv preprint."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Tai, L., Paolo, G., and Liu, M. (2017, January 24\u201328). Virtual-to-real deep reinforcement learning: Continuous control of mobile robots for mapless navigation. Proceedings of the Intelligent Robots and Systems, Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8202134"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Palan, M., Landolfi, N.C., Shevchuk, G., and Sadigh, D. (2019). Learning Reward Functions by Integrating Human Demonstrations and Preferences. arXiv preprint.","DOI":"10.15607\/RSS.2019.XV.023"},{"key":"ref_49","unstructured":"Duan, Y., Schulman, J., Chen, X., Bartlett, P.L., Sutskever, I., and Abbeel, P. (2016). Rl2: Fast reinforcement learning via slow reinforcement learning. arXiv: Artificial Intelligence."},{"key":"ref_50","unstructured":"Finn, C., Abbeel, P., and Levine, S. (2017, January 6\u201311). Model-agnostic meta-learning for fast adaptation of deep networks. Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_51","unstructured":"Sun, L., Peng, C., Zhan, W., and Tomizuka, M. (October, January 30). A fast integrated planning and control framework for autonomous driving via imitation learning. Proceedings of the Dynamic Systems and Control Conference. American Society of Mechanical Engineers, Atlanta, GA, USA."},{"key":"ref_52","unstructured":"Cheng, C., Yan, X., Wagener, N., and Boots, B. (2018). Fast policy learning through imitation and reinforcement. arXiv: Learning."},{"key":"ref_53","unstructured":"Ghasemipour, S.K.S., Gu, S.S., and Zemel, R. (2019, January 8\u201314). SMILe: Scalable Meta Inverse Reinforcement Learning through Context-Conditional Policies. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_54","unstructured":"Ng, A.Y., Harada, D., and Russell, S. (1999, January 27\u201330). Policy invariance under reward transformations: Theory and application to reward shaping. Proceedings of the International Conference on Machine Learning, Bled, Slovenia."},{"key":"ref_55","unstructured":"Devlin, S., and Kudenko, D. (2012, January 4\u20138). Dynamic potential-based reward shaping. Proceedings of the 11th International Conference on Autonomous Agents and Multiagent Systems, Valencia, Spain."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/13\/3664\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:45:08Z","timestamp":1760175908000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/13\/3664"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,6,30]]},"references-count":55,"journal-issue":{"issue":"13","published-online":{"date-parts":[[2020,7]]}},"alternative-id":["s20133664"],"URL":"https:\/\/doi.org\/10.3390\/s20133664","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,6,30]]}}}