{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T16:15:46Z","timestamp":1778256946546,"version":"3.51.4"},"reference-count":39,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2023,1,20]],"date-time":"2023-01-20T00:00:00Z","timestamp":1674172800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62062030"],"award-info":[{"award-number":["62062030"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["ZDYF2021SHFZ243"],"award-info":[{"award-number":["ZDYF2021SHFZ243"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2020-009"],"award-info":[{"award-number":["2020-009"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Key R&amp;D Project of Hainan province","award":["62062030"],"award-info":[{"award-number":["62062030"]}]},{"name":"Key R&amp;D Project of Hainan province","award":["ZDYF2021SHFZ243"],"award-info":[{"award-number":["ZDYF2021SHFZ243"]}]},{"name":"Key R&amp;D Project of Hainan province","award":["2020-009"],"award-info":[{"award-number":["2020-009"]}]},{"name":"Major Science and Technology Project of Haikou","award":["62062030"],"award-info":[{"award-number":["62062030"]}]},{"name":"Major Science and Technology Project of Haikou","award":["ZDYF2021SHFZ243"],"award-info":[{"award-number":["ZDYF2021SHFZ243"]}]},{"name":"Major Science and Technology Project of Haikou","award":["2020-009"],"award-info":[{"award-number":["2020-009"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Autonomous driving systems are crucial complicated cyber\u2013physical systems that combine physical environment awareness with cognitive computing. Deep reinforcement learning is currently commonly used in the decision-making of such systems. However, black-box-based deep reinforcement learning systems do not guarantee system safety and the interpretability of the reward-function settings in the face of complex environments and the influence of uncontrolled uncertainties. Therefore, a formal security reinforcement learning method is proposed. First, we propose an environmental modeling approach based on the influence of nondeterministic environmental factors, which enables the precise quantification of environmental issues. Second, we use the environment model to formalize the reward machine\u2019s structure, which is used to guide the reward-function setting in reinforcement learning. Third, we generate a control barrier function to ensure a safer state behavior policy for reinforcement learning. Finally, we verify the method\u2019s effectiveness in intelligent driving using overtaking and lane-changing scenarios.<\/jats:p>","DOI":"10.3390\/s23031198","type":"journal-article","created":{"date-parts":[[2023,1,20]],"date-time":"2023-01-20T06:52:41Z","timestamp":1674197561000},"page":"1198","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["Safe Decision Controller for Autonomous DrivingBased on Deep Reinforcement Learning inNondeterministic Environment"],"prefix":"10.3390","volume":"23","author":[{"given":"Hongyi","family":"Chen","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Hainan University, Haikou 570228, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Hainan University, Haikou 570228, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8743-2783","authenticated-orcid":false,"given":"Uzair Aslam","family":"Bhatti","sequence":"additional","affiliation":[{"name":"School of Information and Communication Engineering, Hainan University, Haikou 570228, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mengxing","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Information and Communication Engineering, Hainan University, Haikou 570228, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,1,20]]},"reference":[{"key":"ref_1","first-page":"127","article-title":"Analysis of the development trend of information physics systems","volume":"28","author":"Luo","year":"2012","journal-title":"Telecommun. Sci."},{"key":"ref_2","first-page":"1","article-title":"Timing Analysis of CAN FD for Security-Aware Automotive Cyber-Physical Systems","volume":"2022. 99","author":"Xie","year":"2022","journal-title":"IEEE Trans. Dependable Secur. Comput."},{"key":"ref_3","first-page":"1437","article-title":"A comprehensive surveyon safe reinforcement learning","volume":"16","year":"2015","journal-title":"J. Mach. Learn. Res."},{"key":"ref_4","unstructured":"Moldovan, T.M., and Abbeel, P. (July, January 26). Safe exploration in Markov decision processes. Proceedings of the 29th International Conference on Machine Learning, ICML 2012, Edinburgh, UK."},{"key":"ref_5","unstructured":"Tamar, A., Xu, H., and Mannor, S. (2013). Scaling up robust MDPs by reinforcement learning. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Katz, G., Barrett, C.W., Dill, D.L., Julian, K., and Kochen-Derfer, M.J. (2017, January 24\u201328). Reluplex: An efficient SMT solver for verifying deep neural networks. Proceedings of the Computer Aided V Erification\u201429th International Conference, CAV 2017, Heidelberg, Germany. Part I.","DOI":"10.1007\/978-3-319-63387-9_5"},{"key":"ref_7","unstructured":"Arnold, T., Kasenberg, D., and Scheutz, M. (2017). Value Alignment or Misalignment\u2014What Will Keep Systems Accountable? AAAI Workshops, AAAI Press."},{"key":"ref_8","unstructured":"Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S. (2018). Scalable agent alignment via reward modeling: A research direction. arXiv."},{"key":"ref_9","unstructured":"Christiano, P.F., Abate, M., and Amodei, D. (2018). Supervising strong learners by amplifying weak experts. arXiv."},{"key":"ref_10","unstructured":"Hadfield-Menell, D., Russell, S.J., Abbeel, P., and Dragan, A. (2016, January 5\u201310). Cooperative inverse reinforcement learning. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS 2016), Barcelona, Spain."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Mason, G., Calinescu, R., Kudenko, D., and Banks, A. (2017, January 24\u201326). Assured reinforcement learning with formally verified abstract policies. Proceedings of the 9th International Conference on Agents and Artificial Intelligence (ICAART 2017), Porto, Portugal.","DOI":"10.5220\/0006156001050117"},{"key":"ref_12","unstructured":"Cheng, R., Orosz, G., Murray, R.M., and Burdick, J.W. (February, January 27). End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks. Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"L\u00fctjens, B., Everett, M., and How, J.P. (2019, January 20\u201324). Safe Reinforcement Learning with Model Uncertainty Estimates. Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8793611"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Talamini, J., Bartoli, A., De Lorenzo, A., and Medvet, E. (2020). On the Impact of the Rules on Autonomous Drive Learning. Appl. Sci., 10.","DOI":"10.3390\/app10072394"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Krasowski, H., Wang, X., and Althoff, M. (2020, January 20\u201323). Safe Reinforcement Learning for Autonomous Lane Changing Using Set-Based Prediction. Proceedings of the 2020 IEEE 23rd International Conference on Intelligent Transportation Systems (ITSC), Rhodes, Greece.","DOI":"10.1109\/ITSC45102.2020.9294259"},{"key":"ref_16","unstructured":"Wachi, A., and Sui, Y. (2020, January 13\u201318). Safe Reinforcement Learning in Constrained Markov Decision Processes. Proceedings of the 37th International Conference on Machine Learning, Virtual."},{"key":"ref_17","unstructured":"Bastani, O., Pu, Y., and Solar-Lezama, A. (2018, January 3\u20138). Verifiable reinforcement learning via policy extraction. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS 2018), Montreal, QC, Canada."},{"key":"ref_18","unstructured":"De Giacomo, G., Iocchi, L., Favorito, M., and Patrizi, F. (2019, January 11\u201315). Foundations for restraining bolts: Reinforcement learning with LTLf\/LDLf restraining specifications. Proceedings of the International Conference on Automated Planning and Scheduling (ICAPS 2019), Berkeley, CA, USA."},{"key":"ref_19","unstructured":"Camacho, A., Chen, O., Sanner, S., and Mcllraith, S.A. (2017, January 16\u201317). Non-Markovian rewards expressed in LTL: Guiding search via reward shaping. Proceedings of the 10th Annual Symposium on Combinatorial Search (SoCS 2017), Pittsburgh, PA, USA."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Aksaray, D., Jones, A., Kong, Z., Schwager, M., and Belta, C. (2016, January 12\u201314). Q-learning for robust satisfaction of signal temporal logic specifications. Proceedings of the IEEE 55th Conference on Decision and Control (CDC 2016), Las Vegas, NV, USA.","DOI":"10.1109\/CDC.2016.7799279"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Balakrishnan, A., and Deshmukh, J. (2019, January 16\u201318). Structured reward functions using STL. Proceedings of the 22nd ACM International Conference on Hybrid Systems: Computation and Control (HSCC 2019), Montreal, QC, Canada.","DOI":"10.1145\/3302504.3313355"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Wen, M., Papusha, I., and Topcu, U. (2017, January 19\u201325). Learning from demonstrations with high-level side information. Proceedings of the 26th International Joint Conference on Artificial Intelligence (IJCAI 2017), Melbourne, Australia.","DOI":"10.24963\/ijcai.2017\/426"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Alshiekh, M., Bloem, R., Ehlers, R., K\u00f6nighofer, B., Niekum, S., and Topcu, U. (2018, January 2\u20137). Safe reinforcement learning via shielding. Proceedings of the 32nd AAAI Conference on Artificial Intelligence (AAAI 2018), New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11797"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Hahn, E.M., Perez, M., Schewe, S., Somenzi, F., Trivedi, A., and Wojtczak, D. (2019, January 8\u201311). Omega-regular objectives in model-free reinforcement learning. Proceedings of the International Conference on Tools and Algorithms for the Construction and Analysis of Systems (TACAS 2019), Prague, Czech Republic.","DOI":"10.1007\/978-3-030-17462-0_27"},{"key":"ref_25","unstructured":"Icarte, R.T., Klassen, T.Q., Valenzano, R.A., and Mcllraith, S.A. (2018, January 10\u201315). Using reward machines for high-level task specification and decomposition in reinforcement learning. Proceedings of the International Conference on Machine Learning (ICML 2018), Stockholm, Sweden."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Araki, B., Vodrahalli, K., Leech, T., Vasile, C.I., Donahue, M., and Rus, D. (2019, January 22\u201326). Learning to plan with logical automata. Proceedings of the Robotic: Science and Systems (RSS 2019), Breisgau, Germany.","DOI":"10.15607\/RSS.2019.XV.064"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"103452","DOI":"10.1016\/j.trc.2021.103452","article-title":"Decision making of autonomous vehicles in lane change scenarios: Deep reinforcement learning approaches with risk awareness","volume":"134","author":"Li","year":"2021","journal-title":"Transp. Res. Part C Emerg. Technol."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Muzahid, A.J.M., Rahim, M.A., Murad, S.A., Kamarulzaman, S.F., and Rahman, M.A. (2021, January 21\u201323). Optimal Safety Planning and Driving Decision-Making for Multiple Autonomous Vehicles: A Learning Based Approach. Proceedings of the 2021 Emerging Technology in Computing, Communication and Electronics (ETCCE), Dhaka, Bangladesh.","DOI":"10.1109\/ETCCE54784.2021.9689820"},{"key":"ref_29","unstructured":"Pnueli, A. (October, January 30). The temporal logic of programs. Proceedings of the 18th Annual Symposium on Foundations of Computer Science, Washington, DC, USA."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1016\/j.entcs.2004.01.029","article-title":"Monitoring Algorithms for Metric Temporal Logic Specifications","volume":"113","author":"Thati","year":"2005","journal-title":"Electron. Notes Theor. Comput. Sci."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Kober, J., and Peters, J. (2012). Reinforcement Learning in Robotics: A Survey, Springer.","DOI":"10.1007\/978-3-642-27645-3_18"},{"key":"ref_32","first-page":"1926","article-title":"Uncertainty-wise software engineering of complex systems: A systematic mapping study","volume":"32","author":"Tan","year":"2021","journal-title":"J. Softw."},{"key":"ref_33","first-page":"1999","article-title":"Formal modeling and dynamic verification for human cyber physical systems under uncertain environment","volume":"32","author":"Tan","year":"2021","journal-title":"J. Softw."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"3334","DOI":"10.1007\/s11771-013-1857-4","article-title":"Critical safe distance design to improve driving safety based on vehicle-to-vehicle communications","volume":"20","author":"Chen","year":"2013","journal-title":"J. Cent. South Univ."},{"key":"ref_35","unstructured":"Virgo, M., and Brown, A. (2022, August 20). Self-Driving Car Engineer Nanodegree Program. Available online: https:\/\/github.com\/udacity\/CarND-Path-Planning-Project."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"567","DOI":"10.1109\/JAS.2021.1004395","article-title":"Highway lane change decision-making via attention-based deep reinforcement learning","volume":"9","author":"Wang","year":"2022","journal-title":"IEEE\/CAA J. Autom. Sin."},{"key":"ref_37","unstructured":"Bojarski, M., Testa, D.D., Dworakowski, D., Firner, B., Flepp, B., Goyal, P., Jackel, L.D., Monfort, M., Muller, U., and Zhang, J. (2016). End to End Learning for Self-Driving Cars. ArXiv."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"133","DOI":"10.1186\/s10033-021-00639-3","article-title":"Planning and Decision-Making for Connected Autonomous Vehicles at Road Intersections: A Review","volume":"34","author":"Li","year":"2021","journal-title":"Chin. J. Mech. Eng."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/3\/1198\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:11:55Z","timestamp":1760119915000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/3\/1198"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,20]]},"references-count":39,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,2]]}},"alternative-id":["s23031198"],"URL":"https:\/\/doi.org\/10.3390\/s23031198","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,20]]}}}