{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T16:23:35Z","timestamp":1778603015574,"version":"3.51.4"},"reference-count":39,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2018,10,27]],"date-time":"2018-10-27T00:00:00Z","timestamp":1540598400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>It is crucial for robots to autonomously steer in complex environments safely without colliding with any obstacles. Compared to conventional methods, deep reinforcement learning-based methods are able to learn from past experiences automatically and enhance the generalization capability to cope with unseen circumstances. Therefore, we propose an end-to-end deep reinforcement learning algorithm in this paper to improve the performance of autonomous steering in complex environments. By embedding a branching noisy dueling architecture, the proposed model is capable of deriving steering commands directly from raw depth images with high efficiency. Specifically, our learning-based approach extracts the feature representation from depth inputs through convolutional neural networks and maps it to both linear and angular velocity commands simultaneously through different streams of the network. Moreover, the training framework is also meticulously designed to improve the learning efficiency and effectiveness. It is worth noting that the developed system is readily transferable from virtual training scenarios to real-world deployment without any fine-tuning by utilizing depth images. The proposed method is evaluated and compared with a series of baseline methods in various virtual environments. Experimental results demonstrate the superiority of the proposed model in terms of average reward, learning efficiency, success rate as well as computational time. Moreover, a variety of real-world experiments are also conducted which reveal the high adaptability of our model to both static and dynamic obstacle-cluttered environments.<\/jats:p>","DOI":"10.3390\/s18113650","type":"journal-article","created":{"date-parts":[[2018,10,29]],"date-time":"2018-10-29T11:10:41Z","timestamp":1540811441000},"page":"3650","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":35,"title":["Learn to Steer through Deep Reinforcement Learning"],"prefix":"10.3390","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8493-0712","authenticated-orcid":false,"given":"Keyu","family":"Wu","sequence":"first","affiliation":[{"name":"School of Electrical and Electronic Engineering, Nanyang Technological University, 50 Nanyang Ave, Singapore 639798, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0709-0534","authenticated-orcid":false,"given":"Mahdi","family":"Abolfazli Esfahani","sequence":"additional","affiliation":[{"name":"School of Electrical and Electronic Engineering, Nanyang Technological University, 50 Nanyang Ave, Singapore 639798, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shenghai","family":"Yuan","sequence":"additional","affiliation":[{"name":"School of Electrical and Electronic Engineering, Nanyang Technological University, 50 Nanyang Ave, Singapore 639798, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Han","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Electrical and Electronic Engineering, Nanyang Technological University, 50 Nanyang Ave, Singapore 639798, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2018,10,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1016\/j.artint.2014.11.003","article-title":"Deliberation for autonomous robots: A survey","volume":"247","author":"Ingrand","year":"2017","journal-title":"Artif. Intell."},{"key":"ref_2","first-page":"00022","article-title":"Mobile robot navigation and obstacle avoidance techniques: A review","volume":"2","author":"Pandey","year":"2017","journal-title":"Int. Robot. Autom. J."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Yan, Z., Li, J., Zhang, G., and Wu, Y. (2018). A Real-Time Reaction Obstacle Avoidance Algorithm for Autonomous Underwater Vehicles in Unknown Environments. Sensors, 18.","DOI":"10.3390\/s18020438"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"567","DOI":"10.1007\/s10846-017-0543-4","article-title":"UAV Obstacle Avoidance Algorithm Based on Ellipsoid Geometry","volume":"88","author":"Sasongko","year":"2017","journal-title":"J. Intell. Robot. Syst."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Shim, Y., and Kim, G.W. (2018). Range Sensor-Based Efficient Obstacle Avoidance through Selective Decision-Making. Sensors, 18.","DOI":"10.3390\/s18041030"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1047","DOI":"10.1109\/LRA.2017.2656241","article-title":"Fast, on-line collision avoidance for dynamic vehicles using buffered voronoi cells","volume":"2","author":"Zhou","year":"2017","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_7","unstructured":"Zhang, X., Liniger, A., and Borrelli, F. (arXiv, 2017). Optimization-based collision avoidance, arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Al-Kaff, A., Garc\u00eda, F., Mart\u00edn, D., De La Escalera, A., and Armingol, J.M. (2017). Obstacle detection and avoidance system based on monocular camera and size expansion algorithm for UAVs. Sensors, 17.","DOI":"10.3390\/s17051061"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zhang, W., Wei, S., Teng, Y., Zhang, J., Wang, X., and Yan, Z. (2017). Dynamic Obstacle Avoidance for Unmanned Underwater Vehicles Based on an Improved Velocity Obstacle Method. Sensors, 17.","DOI":"10.3390\/s17122742"},{"key":"ref_10","unstructured":"Xie, L., Wang, S., Markham, A., and Trigoni, N. (arXiv, 2017). Towards monocular vision based obstacle avoidance through deep reinforcement learning, arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Kahn, G., Villaflor, A., Ding, B., Abbeel, P., and Levine, S. (arXiv, 2017). Self-Supervised Deep Reinforcement Learning with Generalized Computation Graphs for Robot Navigation, arXiv.","DOI":"10.1109\/ICRA.2018.8460655"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_13","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (arXiv, 2015). Continuous control with deep reinforcement learning, arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Chen, X., Ghadirzadeh, A., Folkesson, J., and Jensfelt, P. (arXiv, 2018). Deep reinforcement learning to acquire navigation skills for wheel-legged robots in complex environments, arXiv.","DOI":"10.1109\/IROS.2018.8593702"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"2905","DOI":"10.3390\/s18092905","article-title":"Intelligent Land-Vehicle Model Transfer Trajectory Planning Method Based on Deep Reinforcement Learning","volume":"18","author":"Yu","year":"2018","journal-title":"Sensors"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Sadeghi, F., and Levine, S. (arXiv, 2016). CAD2RL: Real single-image flight without a single real image, arXiv.","DOI":"10.15607\/RSS.2017.XIII.034"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Long, P., Fan, T., Liao, X., Liu, W., Zhang, H., and Pan, J. (arXiv, 2017). Towards optimally decentralized multi-robot collision avoidance via deep reinforcement learning, arXiv.","DOI":"10.1109\/ICRA.2018.8461113"},{"key":"ref_18","unstructured":"Bruce, J., S\u00fcnderhauf, N., Mirowski, P., Hadsell, R., and Milford, M. (arXiv, 2017). One-shot reinforcement learning for robot navigation with interactive replay, arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Pfeiffer, M., Shukla, S., Turchetta, M., Cadena, C., Krause, A., Siegwart, R., and Nieto, J. (arXiv, 2018). Reinforced Imitation: Sample Efficient Deep Reinforcement Learning for Map-less Navigation by Leveraging Prior Demonstrations, arXiv.","DOI":"10.1109\/LRA.2018.2869644"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Tai, L., Paolo, G., and Liu, M. (2017, January 24\u201328). Virtual-to-real deep reinforcement learning: Continuous control of mobile robots for mapless navigation. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8202134"},{"key":"ref_21","unstructured":"Zhelo, O., Zhang, J., Tai, L., Liu, M., and Burgard, W. (arXiv, 2018). Curiosity-driven Exploration for Mapless Navigation with Deep Reinforcement Learning, arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Xie, L., Wang, S., Rosa, S., Markham, A., and Trigoni, N. (2018). Learning with Training Wheels: Speeding up Training with a Simple Controller for Deep Reinforcement Learning, Institute of Electrical and Electronics Engineers.","DOI":"10.1109\/ICRA.2018.8461203"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zhu, Y., Mottaghi, R., Kolve, E., Lim, J.J., Gupta, A., Fei-Fei, L., and Farhadi, A. (June, January 29). Target-driven visual navigation in indoor scenes using deep reinforcement learning. Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore.","DOI":"10.1109\/ICRA.2017.7989381"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"3247","DOI":"10.1109\/LRA.2018.2851148","article-title":"Visual Navigation for Biped Humanoid Robots Using Deep Reinforcement Learning","volume":"3","author":"Leiva","year":"2018","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Tai, L., Zhang, J., Liu, M., and Burgard, W. (arXiv, 2017). Socially compliant navigation through raw depth inputs with generative adversarial imitation learning, arXiv.","DOI":"10.1109\/ICRA.2018.8460968"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Tai, L., and Liu, M. (arXiv, 2016). Towards cognitive exploration through deep reinforcement learning for mobile robots, arXiv.","DOI":"10.1186\/s40638-016-0055-x"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhang, J., Springenberg, J.T., Boedecker, J., and Burgard, W. (2017, January 24\u201328). Deep reinforcement learning with successor features for navigation across similar environments. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8206049"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Laina, I., Rupprecht, C., Belagiannis, V., Tombari, F., and Navab, N. (2016, January 25\u201328). Deeper depth prediction with fully convolutional residual networks. Proceedings of the 2016 Fourth International Conference on 3D Vision (3DV), Stanford, CA, USA.","DOI":"10.1109\/3DV.2016.32"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Chen, Y., Chen, R., Liu, M., Xiao, A., Wu, D., and Zhao, S. (2018). Indoor Visual Positioning Aided by CNN-Based Image Retrieval: Training-Free, 3D Modeling-Free. Sensors, 18.","DOI":"10.3390\/s18082692"},{"key":"ref_30","unstructured":"Yang, S., Konam, S., Ma, C., Rosenthal, S., Veloso, M., and Scherer, S. (arXiv, 2017). Obstacle avoidance through deep networks based intermediate perception, arXiv."},{"key":"ref_31","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (arXiv, 2013). Playing atari with deep reinforcement learning, arXiv."},{"key":"ref_32","unstructured":"Mirowski, P., Pascanu, R., Viola, F., Soyer, H., Ballard, A.J., Banino, A., Denil, M., Goroshin, R., Sifre, L., and Kavukcuoglu, K. (arXiv, 2016). Learning to navigate in complex environments, arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Van Hasselt, H., Guez, A., and Silver, D. (2016, January 12\u201317). Deep Reinforcement Learning with Double Q-Learning. Proceedings of the 2016 AAAI, Phoenix, AZ, USA.","DOI":"10.1609\/aaai.v30i1.10295"},{"key":"ref_34","unstructured":"Wang, Z., Schaul, T., Hessel, M., Van Hasselt, H., Lanctot, M., and De Freitas, N. (arXiv, 2015). Dueling network architectures for deep reinforcement learning, arXiv."},{"key":"ref_35","unstructured":"Fortunato, M., Azar, M.G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., and Pietquin, O. (arXiv, 2017). Noisy networks for exploration, arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Tavakoli, A., Pardo, F., and Kormushev, P. (arXiv, 2017). Action branching architectures for deep reinforcement learning, arXiv.","DOI":"10.1609\/aaai.v32i1.11798"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"193","DOI":"10.1023\/A:1017932429737","article-title":"A sparse sampling algorithm for near-optimal planning in large Markov decision processes","volume":"49","author":"Kearns","year":"2002","journal-title":"Mach. Learn."},{"key":"ref_38","unstructured":"Koenig, N.P., and Howard, A. (October, January 28). Design and use paradigms for Gazebo, an open-source multi-robot simulator. Proceedings of the IROS 2004, Sendai, Japan."},{"key":"ref_39","unstructured":"Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., and Isard, M. (2016, January 2\u20134). Tensorflow: A system for large-scale machine learning. Proceedings of the OSDI 2016, Savannah, GA, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/11\/3650\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T15:26:35Z","timestamp":1760196395000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/11\/3650"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,10,27]]},"references-count":39,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2018,11]]}},"alternative-id":["s18113650"],"URL":"https:\/\/doi.org\/10.3390\/s18113650","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,10,27]]}}}