{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:18:17Z","timestamp":1760239097962,"version":"build-2065373602"},"reference-count":25,"publisher":"MDPI AG","issue":"19","license":[{"start":{"date-parts":[[2020,10,1]],"date-time":"2020-10-01T00:00:00Z","timestamp":1601510400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"NSFC Grants","award":["61973311"],"award-info":[{"award-number":["61973311"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Autonomous driving with artificial intelligence technology has been viewed as promising for autonomous vehicles hitting the road in the near future. In recent years, considerable progress has been made with Deep Reinforcement Learnings (DRLs) for realizing end-to-end autonomous driving. Still, driving safely and comfortably in real dynamic scenarios with DRL is nontrivial due to the reward functions being typically pre-defined with expertise. This paper proposes a human-in-the-loop DRL algorithm for learning personalized autonomous driving behavior in a progressive learning way. Specifically, a progressively optimized reward function (PORF) learning model is built and integrated into the Deep Deterministic Policy Gradient (DDPG) framework, which is called PORF-DDPG in this paper. PORF consists of two parts: the first part of the PORF is a pre-defined typical reward function on the system state, the second part is modeled as a Deep Neural Network (DNN) for representing driving adjusting intention by the human observer, which is the main contribution of this paper. The DNN-based reward model is progressively learned using the front-view images as the input and via active human supervision and intervention. The proposed approach is potentially useful for driving in dynamic constrained scenarios when dangerous collision events might occur frequently with classic DRLs. The experimental results show that the proposed autonomous driving behavior learning method exhibits online learning capability and environmental adaptability.<\/jats:p>","DOI":"10.3390\/s20195626","type":"journal-article","created":{"date-parts":[[2020,10,1]],"date-time":"2020-10-01T09:04:12Z","timestamp":1601543052000},"page":"5626","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":12,"title":["PORF-DDPG: Learning Personalized Autonomous Driving Behavior with Progressively Optimized Reward Function"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1241-0544","authenticated-orcid":false,"given":"Jie","family":"Chen","sequence":"first","affiliation":[{"name":"The College of Intelligence Science and Technology, National University of Defense Technology, Changsha 410073, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tao","family":"Wu","sequence":"additional","affiliation":[{"name":"The College of Intelligence Science and Technology, National University of Defense Technology, Changsha 410073, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Meiping","family":"Shi","sequence":"additional","affiliation":[{"name":"The College of Intelligence Science and Technology, National University of Defense Technology, Changsha 410073, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Jiang","sequence":"additional","affiliation":[{"name":"The College of Intelligence Science and Technology, National University of Defense Technology, Changsha 410073, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,10,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"533","DOI":"10.1038\/323533a0","article-title":"Learning representations by back-propagating errors","volume":"323","author":"Rumelhart","year":"1986","journal-title":"Nature"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"504","DOI":"10.1126\/science.1127647","article-title":"Reducing the dimensionality of data with neural networks","volume":"313","author":"Hinton","year":"2006","journal-title":"Science"},{"key":"ref_3","unstructured":"Kingma, D.P., and Ba, J. (2015, January 7\u20139). Adam: A Method for Stochastic Optimization. Proceedings of the 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA."},{"key":"ref_4","unstructured":"Hinton, G.E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2012). Improving neural networks by preventing co-adaptation of feature detectors. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"Lecun","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Shahian Jahromi, B., Tulabandhula, T., and Cetin, S. (2019). Real-Time Hybrid Multi-Sensor Fusion Framework for Perception in Autonomous Vehicles. Sensors, 19.","DOI":"10.3390\/s19204357"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Sun, Y., Li, J., and Sun, Z. (2019). Multi-Stage Hough Space Calculation for Lane Markings Detection via IMU and Vision Fusion. Sensors, 19.","DOI":"10.20944\/preprints201904.0175.v1"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Le, A., Prabakaran, V., Sivanantham, V., and Mohan, R. (2018). Modified A-Star Algorithm for Efficient Coverage Path Planning in Tetris Inspired Self-Reconfigurable Robot with Integrated Laser Sensor. Sensors, 18.","DOI":"10.3390\/s18082585"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Luo, D., Rasim, Y., Li, Y., Meng, G., Xu, J., and Wang, C. (2016). A Vehicle Active Safety Model: Vehicle Speed Control Based on Driver Vigilance Detection Using Wearable EEG and Sparse Representation. Sensors, 16.","DOI":"10.3390\/s16020242"},{"key":"ref_10","unstructured":"Pomerleau, D. (1988). ALVINN: An Autonomous Land Vehicle in a Neural Network. Advances in Neural Information Processing Systems 1, Proceedings of the NIPS Conference, Denver, CO, USA, January 1988, NIPS."},{"key":"ref_11","unstructured":"LeCun, Y., Muller, U., Ben, J., Cosatto, E., and Flepp, B. (2005, January 5\u20138). Off-Road Obstacle Avoidance through End-to-End Learning. Proceedings of the Advances in Neural Information Processing Systems 18 Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_12","unstructured":"Bojarski, M., Testa, D.D., Dworakowski, D., Firner, B., Flepp, B., Goyal, P., Jackel, L.D., Monfort, M., Muller, U., and Zhang, J. (2016). End to End Learning for Self-Driving Cars. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Codevilla, F., M\u00fcller, M., Dosovitskiy, A., L\u00f3pez, A., and Koltun, V. (2017). End-to-end Driving via Conditional Imitation Learning. arXiv.","DOI":"10.1109\/ICRA.2018.8460487"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Amini, A., Rosman, G., Karaman, S., and Rus, D. (2019, January 20\u201324). Variational End-to-End Navigation and Localization. Proceedings of the International Conference on Robotics and Automation, Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8793579"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1038\/nature16961","article-title":"Mastering the game of Go with deep neural networks and tree search","volume":"529","author":"Silver","year":"2016","journal-title":"Nature"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1038\/nature24270","article-title":"Mastering the game of Go without human knowledge","volume":"550","author":"Silver","year":"2017","journal-title":"Nature"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Lange, S., Riedmiller, M.A., and Voigtl\u00e4nder, A. (2012, January 4\u20139). Autonomous reinforcement learning on raw visual input data in a real world application. Proceedings of the 2012 International Joint Conference on Neural Networks (IJCNN), Dallas, TX, USA.","DOI":"10.1109\/IJCNN.2012.6252823"},{"key":"ref_18","unstructured":"Sallab, A.E., Abdou, M., Perot, E., and Yogamani, S.K. (2016). End-to-End Deep Reinforcement Learning for Lane Keeping Assist. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Chae, H., Kang, C.M., Kim, B., Kim, J., Chung, C.C., and Choi, J.W. (2017, January 26\u201319). Autonomous braking system via deep reinforcement learning. Proceedings of the 20th IEEE International Conference on Intelligent Transportation Systems, Yokohama, Japan.","DOI":"10.1109\/ITSC.2017.8317839"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Min, K., and Kim, H. (2018, January 26\u201330). Deep Q Learning Based High Level Driving Policy Determination. Proceedings of the 2018 IEEE Intelligent Vehicles Symposium, Changshu, China.","DOI":"10.1109\/IVS.2018.8500645"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Kendall, A., Hawke, J., Janz, D., Mazur, P., Reda, D., Allen, J., Lam, V., Bewley, A., and Shah, A. (2019, January 20\u201324). Learning to Drive in a Day. Proceedings of the International Conference on Robotics and Automation, Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8793742"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_23","unstructured":"Russell, S.J., and Norvig, P. (2010). Artificial Intelligence\u2014A Modern Approach, Pearson Education. [3rd ed.]."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. (2016, January 27\u201330). The Cityscapes Dataset for Semantic Urban Scene Understanding. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.350"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Riedmiller, M.A., Montemerlo, M., and Dahlkamp, H. (2007, January 11\u201313). Learning to Drive a Real Car in 20 Minutes. Proceedings of the Frontiers in the Convergence of Bioscience and Information Technologies 2007, Jeju Island, Korea.","DOI":"10.1109\/FBIT.2007.37"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/19\/5626\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:15:41Z","timestamp":1760177741000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/19\/5626"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,1]]},"references-count":25,"journal-issue":{"issue":"19","published-online":{"date-parts":[[2020,10]]}},"alternative-id":["s20195626"],"URL":"https:\/\/doi.org\/10.3390\/s20195626","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2020,10,1]]}}}