{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T17:31:03Z","timestamp":1784136663688,"version":"3.55.0"},"reference-count":34,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2021,3,25]],"date-time":"2021-03-25T00:00:00Z","timestamp":1616630400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100015520","name":"Abu Dhabi Department of Education and Knowledge","doi-asserted-by":"publisher","award":["AQRE18-114"],"award-info":[{"award-number":["AQRE18-114"]}],"id":[{"id":"10.13039\/501100015520","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Roadway Transportation and Traffic Safety Research Center (RTTSRC) of the United Arab Emirates University","award":["31R225"],"award-info":[{"award-number":["31R225"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Recent research works on intelligent traffic signal control (TSC) have been mainly focused on leveraging deep reinforcement learning (DRL) due to its proven capability and performance. DRL-based traffic signal control frameworks belong to either discrete or continuous controls. In discrete control, the DRL agent selects the appropriate traffic light phase from a finite set of phases. Whereas in continuous control approach, the agent decides the appropriate duration for each signal phase within a predetermined sequence of phases. Among the existing works, there are no prior approaches that propose a flexible framework combining both discrete and continuous DRL approaches in controlling traffic signal. Thus, our ultimate objective in this paper is to propose an approach capable of deciding simultaneously the proper phase and its associated duration. Our contribution resides in adapting a hybrid Deep Reinforcement Learning that considers at the same time discrete and continuous decisions. Precisely, we customize a Parameterized Deep Q-Networks (P-DQN) architecture that permits a hierarchical decision-making process that primarily decides the traffic light next phases and secondly specifies its the associated timing. The evaluation results of our approach using Simulation of Urban MObility (SUMO) shows its out-performance over the benchmarks. The proposed framework is able to reduce the average queue length of vehicles and the average travel time by 22.20% and 5.78%, respectively, over the alternative DRL-based TSC systems.<\/jats:p>","DOI":"10.3390\/s21072302","type":"journal-article","created":{"date-parts":[[2021,3,25]],"date-time":"2021-03-25T21:09:45Z","timestamp":1616706585000},"page":"2302","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":56,"title":["Traffic Signal Control Using Hybrid Action Space Deep Reinforcement Learning"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5846-9897","authenticated-orcid":false,"given":"Salah","family":"Bouktif","sequence":"first","affiliation":[{"name":"Department of Computer Science and Software Engineering, University of United Arab Emirates, Al Ain 15551, Abu Dhabi, United Arab Emirates"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3288-4344","authenticated-orcid":false,"given":"Abderraouf","family":"Cheniki","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, University of Boumerdes, Boumerd\u00e8s 35000, Algeria"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ali","family":"Ouni","sequence":"additional","affiliation":[{"name":"\u00c9cole de Technologie Sup\u00e9rieure, University of Quebec, Montreal, QC H3C 1K3, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2021,3,25]]},"reference":[{"key":"ref_1","unstructured":"INRIX Scoreboard 2019 (2021, February 09). PRESS RELEASES. Available online: https:\/\/inrix.com\/press-releases\/2019-traffic-scorecard-uk\/."},{"key":"ref_2","unstructured":"Haydari, A., and Yilmaz, Y. (2020). Deep Reinforcement Learning for Intelligent Transportation Systems: A Survey. arXiv."},{"key":"ref_3","unstructured":"Lin, Y., Dai, X., Li, L., and Wang, F.Y. (2018). An Efficient Deep Reinforcement Learning Model for Urban Traffic Control. arXiv."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Greguri\u0107, M., Vuji\u0107, M., Alexopoulos, C., and Mileti\u0107, M. (2020). Application of Deep Reinforcement Learning in Traffic Signal Control: An Overview and Impact of Open Traffic Data. Appl. Sci., 10.","DOI":"10.3390\/app10114011"},{"key":"ref_5","unstructured":"Genders, W., and Razavi, S. (2016). Using a Deep Reinforcement Learning Agent for Traffic Signal Control. arXiv."},{"key":"ref_6","unstructured":"Casas, N. (2017). Deep Deterministic Policy Gradient for Urban Traffic Light Control. arXiv."},{"key":"ref_7","unstructured":"Gao, J., Shen, Y., Liu, J., Ito, M., and Shiratori, N. (2017). Adaptive Traffic Signal Control: Deep Reinforcement Learning Algorithm with Experience Replay and Target Network. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Guo, M., Wang, P., Chan, C.Y., and Askary, S. (2019). A Reinforcement Learning Approach for Intelligent Traffic Signal Control at Urban Intersections. arXiv.","DOI":"10.1109\/ITSC.2019.8917268"},{"key":"ref_9","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013). Playing Atari with Deep Reinforcement Learning. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"van Hasselt, H., Guez, A., and Silver, D. (2015). Deep Reinforcement Learning with Double Q-learning. arXiv.","DOI":"10.1609\/aaai.v30i1.10295"},{"key":"ref_11","unstructured":"Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N. (2016). Dueling Network Architectures for Deep Reinforcement Learning. arXiv."},{"key":"ref_12","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2019). Continuous control with deep reinforcement learning. arXiv."},{"key":"ref_13","unstructured":"Gu, S., Lillicrap, T., Sutskever, I., and Levine, S. (2016). Continuous Deep Q-Learning with Model-based Acceleration. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Masson, W., Ranchod, P., and Konidaris, G. (2015). Reinforcement Learning with Parameterized Actions. arXiv.","DOI":"10.1609\/aaai.v30i1.10226"},{"key":"ref_15","unstructured":"Xiong, J., Wang, Q., Yang, Z., Sun, P., Han, L., Zheng, Y., Fu, H., Zhang, T., Liu, J., and Liu, H. (2018). Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Wei, H., Zheng, G., Yao, H., and Li, Z. (2018). IntelliLight: A Reinforcement Learning Approach for Intelligent Traffic Light Control. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Association for Computing Machinery.","DOI":"10.1145\/3219819.3220096"},{"key":"ref_17","unstructured":"Liu, X.Y., Ding, Z., Borst, S., and Walid, A. (2018). Deep Reinforcement Learning for Intelligent Transportation Systems. arXiv."},{"key":"ref_18","unstructured":"Yu, C., Liu, J., and Nemati, S. (2020). Reinforcement Learning in Healthcare: A Survey. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"3751","DOI":"10.1109\/TII.2020.3014599","article-title":"Battery Thermal- and Health-Constrained Energy Management for Hybrid Electric Bus Based on Soft Actor-Critic DRL Algorithm","volume":"17","author":"Wu","year":"2021","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_20","first-page":"100020","article-title":"Decentralized network level adaptive signal control by multi-agent deep reinforcement learning","volume":"1","author":"Gong","year":"2019","journal-title":"Transp. Res. Interdiscip. Perspect."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Wei, H., Chen, C., Zheng, G., Wu, K., Gayah, V., Xu, K., and Li, Z. (2019). PressLight: Learning Max Pressure Control to Coordinate Traffic Signals in Arterial Network. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Association for Computing Machinery.","DOI":"10.1145\/3292500.3330949"},{"key":"ref_22","unstructured":"Genders, W. (2018). Deep Reinforcement Learning Adaptive Traffic Signal Control. [Ph.D. Thesis, McMaster University]."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_24","unstructured":"Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T.P., Harley, T., Silver, D., and Kavukcuoglu, K. (2016). Asynchronous Methods for Deep Reinforcement Learning. arXiv."},{"key":"ref_25","unstructured":"Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018). Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. arXiv."},{"key":"ref_26","unstructured":"Hausknecht, M., and Stone, P. (2016). Deep Reinforcement Learning in Parameterized Action Space. arXiv."},{"key":"ref_27","unstructured":"Bester, C.J., James, S.D., and Konidaris, G.D. (2019). Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces. arXiv."},{"key":"ref_28","unstructured":"Sutton, R.S., and Barto, A.G. (2018). Reinforcement Learning: An Introduction, The MIT Press."},{"key":"ref_29","unstructured":"Behrisch, M., Bieker, L., Erdmann, J., and Krajzewicz, D. (2011, January 23\u201329). SUMO - Simulation of Urban MObility: An overview. Proceedings of the SIMUL 2011, The Third International Conference on Advances in System Simulation, Barcelona, Spain."},{"key":"ref_30","unstructured":"Vidali, A., Crociani, L., Vizzari, G., and Bandini, S. (2020, August 22). A Deep Reinforcement Learning Approach to Adaptive Traffic Lights Management. Available online: http:\/\/ceur-ws.org\/Vol-2404\/paper07.pdf."},{"key":"ref_31","unstructured":"(2020, August 22). Speed Limits by Country\u2014Wikipedia, The Free Encyclopedia. Available online: https:\/\/en.wikipedia.org\/wiki\/Speed_limits_by_country."},{"key":"ref_32","unstructured":"Tieleman, T., and Hinton, G. (2021, March 24). Lecture 6.5-Rmsprop: Divide the Gradient by a Running Average of Its Recent Magnitude. Available online: http:\/\/www.cs.toronto.edu\/~hinton\/coursera\/lecture6\/lec6.pdf."},{"key":"ref_33","unstructured":"Gordon, R., and Tighe, W. (2005). Traffic Control Systems Handbook."},{"key":"ref_34","unstructured":"GEUS, S.D. (2020). Utilizing Available Data to Warm Start Online Reinforcement Learning. [Master\u2019s Thesis, Vrije Universiteit Amsterdam]."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/7\/2302\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:41:07Z","timestamp":1760161267000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/7\/2302"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,25]]},"references-count":34,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2021,4]]}},"alternative-id":["s21072302"],"URL":"https:\/\/doi.org\/10.3390\/s21072302","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,3,25]]}}}