{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,5]],"date-time":"2026-02-05T22:22:40Z","timestamp":1770330160397,"version":"3.49.0"},"reference-count":39,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2022,12,30]],"date-time":"2022-12-30T00:00:00Z","timestamp":1672358400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61861013"],"award-info":[{"award-number":["61861013"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2020GXNSFDA238001"],"award-info":[{"award-number":["2020GXNSFDA238001"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2020KY05033"],"award-info":[{"award-number":["2020KY05033"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Natural Science Foundation of Guangxi","award":["61861013"],"award-info":[{"award-number":["61861013"]}]},{"name":"Natural Science Foundation of Guangxi","award":["2020GXNSFDA238001"],"award-info":[{"award-number":["2020GXNSFDA238001"]}]},{"name":"Natural Science Foundation of Guangxi","award":["2020KY05033"],"award-info":[{"award-number":["2020KY05033"]}]},{"name":"the Middle-Aged and Young Teachers\u2019 Basic Ability Promotion Project of Guangxi","award":["61861013"],"award-info":[{"award-number":["61861013"]}]},{"name":"the Middle-Aged and Young Teachers\u2019 Basic Ability Promotion Project of Guangxi","award":["2020GXNSFDA238001"],"award-info":[{"award-number":["2020GXNSFDA238001"]}]},{"name":"the Middle-Aged and Young Teachers\u2019 Basic Ability Promotion Project of Guangxi","award":["2020KY05033"],"award-info":[{"award-number":["2020KY05033"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Software-defined networking (SDN) has become one of the critical technologies for data center networks, as it can improve network performance from a global perspective using artificial intelligence algorithms. Due to the strong decision-making and generalization ability, deep reinforcement learning (DRL) has been used in SDN intelligent routing and scheduling mechanisms. However, traditional deep reinforcement learning algorithms present the problems of slow convergence rate and instability, resulting in poor network quality of service (QoS) for an extended period before convergence. Aiming at the above problems, we propose an automatic QoS architecture based on multistep DRL (AQMDRL) to optimize the QoS performance of SDN. AQMDRL uses a multistep approach to solve the overestimation and underestimation problems of the deep deterministic policy gradient (DDPG) algorithm. The multistep approach uses the maximum value of the n-step action currently estimated by the neural network instead of the one-step Q-value function, as it reduces the possibility of positive error generated by the Q-value function and can effectively improve convergence stability. In addition, we adapt a prioritized experience sampling based on SumTree binary trees to improve the convergence rate of the multistep DDPG algorithm. Our experiments show that the AQMDRL we proposed significantly improves the convergence performance and effectively reduces the network transmission delay of SDN over existing DRL algorithms.<\/jats:p>","DOI":"10.3390\/s23010429","type":"journal-article","created":{"date-parts":[[2023,1,2]],"date-time":"2023-01-02T03:08:59Z","timestamp":1672628939000},"page":"429","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["AQMDRL: Automatic Quality of Service Architecture Based on Multistep Deep Reinforcement Learning in Software-Defined Networking"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6516-9067","authenticated-orcid":false,"given":"Junyan","family":"Chen","sequence":"first","affiliation":[{"name":"School of Computer Science and Information Security, Guilin University of Electronic Technology, Guilin 541004, China"},{"name":"School of Information and Communication, Guilin University of Electronic Technology, Guilin 541004, China"},{"name":"School of Computer Science and Engineering, Northeastern University, Shenyang 110819, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cenhuishan","family":"Liao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Security, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yong","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Security, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lei","family":"Jin","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Security, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoye","family":"Lu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Security, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaolan","family":"Xie","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Security, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rui","family":"Yao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Security, Guilin University of Electronic Technology, Guilin 541004, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,12,30]]},"reference":[{"key":"ref_1","first-page":"2184","article-title":"Overview of the application of deep learning in Software Defined Network research","volume":"31","author":"Yang","year":"2020","journal-title":"J. Softw."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Ali Khan, A., Zafrullah, M., Hussain, M., and Ahmad, A. (2017, January 19\u201322). Performance analysis of OSPF and hybrid networks. Proceedings of the International Symposium on Wireless Systems and Networks (ISWSN 2017), Lahore, Pakistan.","DOI":"10.1109\/ISWSN.2017.8250022"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"779","DOI":"10.1109\/TNET.2016.2614247","article-title":"Traffic engineering with Equal-Cost-Multipath: An algorithmic perspective","volume":"25","author":"Chiesa","year":"2017","journal-title":"IEEE\/ACM Trans. Netw."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"438","DOI":"10.1109\/LCOMM.2018.2789419","article-title":"Traffic Engineering Enhancement by Progressive Migration to SDN","volume":"22","author":"Tanha","year":"2018","journal-title":"IEEE Commun. Lett."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1386","DOI":"10.1109\/TGCN.2022.3162237","article-title":"LBSMT: Load Balancing Switch Migration Algorithm for Cooperative Communication Intelligent Transportation Systems","volume":"6","author":"Babbar","year":"2022","journal-title":"IEEE Trans. Green Commun. Netw."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1965","DOI":"10.1109\/TNSM.2022.3150978","article-title":"Dynamic Radio Access Selection and Slice Allocation for Differentiated Traffic Management on Future Mobile Networks","volume":"19","author":"Gonzalez","year":"2022","journal-title":"IEEE Trans. Netw. Serv. Manag."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"108062","DOI":"10.1016\/j.knosys.2021.108062","article-title":"DeepTSQP: Temporal-aware service QoS prediction via deep neural network and feature integration","volume":"241","author":"Zou","year":"2022","journal-title":"Knowl.-Based Syst."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"162","DOI":"10.1016\/j.comcom.2019.10.011","article-title":"SDN-based real-time urban traffic analysis in VANET environment","volume":"149","author":"Bhatia","year":"2020","journal-title":"Comput. Commun."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Nugraha, B., and Murthy, R. (2020, January 10\u201312). Deep learning-based slow DDoS attack detection in SDN-based networks. Proceedings of the IEEE Conference on Network Function Virtualization and Software Defined Networks (NFV-SDN 2020), Leganes, Spain.","DOI":"10.1109\/NFV-SDN50289.2020.9289894"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"83765","DOI":"10.1109\/ACCESS.2020.2992044","article-title":"Long short-term memory and fuzzy logic for anomaly detection and mitigation in software-defined network environment","volume":"8","author":"Novaes","year":"2020","journal-title":"IEEE Access"},{"key":"ref_11","first-page":"8354150","article-title":"ALBLP: Adaptive Load-Balancing Architecture Based on Link-State Prediction in Software-Defined Networking","volume":"2022","author":"Chen","year":"2022","journal-title":"Wirel. Commun. Mob. Comput."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"167944","DOI":"10.1109\/ACCESS.2019.2953498","article-title":"Reinforcement learning for service function chain reconfiguration in NFV-SDN metro-core optical networks","volume":"7","author":"Troia","year":"2019","journal-title":"IEEE Access"},{"key":"ref_13","first-page":"5124960","article-title":"RLMR: Reinforcement Learning Based Multipath Routing for SDN","volume":"2022","author":"Chen","year":"2022","journal-title":"Wirel. Commun. Mob. Comput."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"304","DOI":"10.1109\/TCCN.2020.2988908","article-title":"Delay-aware VNF scheduling: A reinforcement learning approach with variable action set","volume":"7","author":"Li","year":"2020","journal-title":"IEEE Trans. Cogn. Commun. Netw."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"259","DOI":"10.1109\/ACCESS.2020.3046693","article-title":"Optimizing the lifetime of software defined wireless sensor network via reinforcement learning","volume":"9","author":"Younus","year":"2020","journal-title":"IEEE Access"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"3495","DOI":"10.1109\/JIOT.2021.3102130","article-title":"Improving the software-defined wireless sensor networks routing performance using reinforcement learning","volume":"9","author":"Younus","year":"2021","journal-title":"IEEE Internet Things J."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"870","DOI":"10.1109\/TNSM.2020.3036911","article-title":"Intelligent Routing Based on Reinforcement Learning for Software-Defined Networking","volume":"18","author":"Velasco","year":"2021","journal-title":"IEEE Trans. Netw. Serv. Manag."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"851","DOI":"10.1109\/TBC.2021.3099728","article-title":"An innovative reinforcement learning-based framework for quality of service provisioning over multimedia-based SDN environments","volume":"67","author":"Shah","year":"2021","journal-title":"IEEE Trans. Broadcast."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1041","DOI":"10.1109\/JIOT.2020.3009540","article-title":"Heterogeneous task offloading and resource allocations via deep recurrent reinforcement learning in partial observable multifog networks","volume":"8","author":"Baek","year":"2020","journal-title":"IEEE Internet Things J."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Chen, J., Xiao, W., Li, X., Zheng, Y., Huang, X., Huang, D., and Wang, M. (2022). A routing optimization method for software-defined optical transport networks based on ensembles and reinforcement learning. Sensors, 22.","DOI":"10.3390\/s22218139"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2476","DOI":"10.1109\/JSAC.2021.3087270","article-title":"PnP-DRL: A Plug-and-Play Deep Reinforcement Learning Approach for Experience-Driven Networking","volume":"39","author":"Xu","year":"2021","journal-title":"IEEE J. Sel. Areas Commun."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"102865","DOI":"10.1016\/j.jnca.2020.102865","article-title":"DRL-R: Deep reinforcement learning approach for intelligent routing in software-defined data-center networks","volume":"177","author":"Liu","year":"2021","journal-title":"J. Netw. Comput. Appl."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"149","DOI":"10.23919\/JCC.2020.02.013","article-title":"EARS: Intelligence-Driven Experiential Network Architecture for Automatic Routing in Software-Defined Networking","volume":"17","author":"Hu","year":"2020","journal-title":"China Commun."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"103181","DOI":"10.1016\/j.jnca.2021.103181","article-title":"Deep Q-Network and traffic prediction based routing optimization in software defined networks","volume":"192","author":"Bouzidi","year":"2021","journal-title":"J. Netw. Comput. Appl."},{"key":"ref_25","first-page":"2669","article-title":"A SDN Routing Optimization Mechanism Based on Deep Reinforcement Learning","volume":"41","author":"Lan","year":"2019","journal-title":"J. Electron. Inf. Technol."},{"key":"ref_26","first-page":"60","article-title":"Software-defined networking QoS optimization based on deep reinforcement learning","volume":"40","author":"Lan","year":"2019","journal-title":"J. Commun."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"6242","DOI":"10.1109\/JIOT.2019.2960033","article-title":"Deep-Reinforcement-Learning-Based QoS-Aware Secure Routing for SDN-IoT","volume":"7","author":"Guo","year":"2020","journal-title":"IEEE Internet Things J."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"4308","DOI":"10.1109\/TII.2021.3132136","article-title":"Transfer Reinforcement Learning Aided Distributed Network Slicing Resource Optimization in Industrial IoT","volume":"18","author":"Mai","year":"2021","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_29","first-page":"3866143","article-title":"ALBRL: Automatic Load-Balancing Architecture Based on Reinforcement Learning in Software-Defined Networking","volume":"2022","author":"Chen","year":"2022","journal-title":"Wirel. Commun. Mob. Comput."},{"key":"ref_30","unstructured":"Fujimoto, S., Hoof, H., and Meger, D. (2018, January 10\u201315). Addressing Function Approximation Error in Actor-Critic Methods. Proceedings of the 35th International Conference on Machine Learning (ICML 2018), Stockholm, Sweden."},{"key":"ref_31","first-page":"2170","article-title":"An Intelligent Routing Technology Based on Deep Reinforcement Learning","volume":"48","author":"Sun","year":"2020","journal-title":"Acta Electron. Sin."},{"key":"ref_32","first-page":"1563","article-title":"Pinning Control-Based Routing Policy Generation Using Deep Reinforcement Learning","volume":"58","author":"Sun","year":"2021","journal-title":"J. Comput. Res. Dev."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"629","DOI":"10.1109\/TNET.2021.3126933","article-title":"Enabling scalable routing in software-defined networks with deep reinforcement learning on critical nodes","volume":"30","author":"Sun","year":"2021","journal-title":"IEEE\/ACM Trans. Netw."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"107891","DOI":"10.1016\/j.comnet.2021.107891","article-title":"ScaleDRL: A Scalable Deep Reinforcement Learning Approach for Traffic Engineering in SDN with Pinning Control","volume":"190","author":"Sun","year":"2021","journal-title":"Comput. Netw."},{"key":"ref_35","first-page":"11767","article-title":"Softmax deep double deterministic policy gradients","volume":"33","author":"Pan","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Huang, R., Guan, W., Zhai, G., He, J., and Chu, X. (2022). Deep Graph Reinforcement Learning Based Intelligent Traffic Routing Control for Software-Defined Wireless Sensor Networks. Appl. Sci., 12.","DOI":"10.3390\/app12041951"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Wang, X., Fu, L., Cheng, N., Sun, R., Luan, T., Quan, W., and Aldubaikhy, K. (2022). Joint Flying Relay Location and Routing Optimization for 6G UAV\u2013IoT Networks: A Graph Neural Network-Based Approach. Remote Sens., 14.","DOI":"10.3390\/rs14174377"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Meng, L., Gorbet, R., and Kuli\u0107, D. (2020, January 13\u201318). The effect of multi-step methods on overestimation in deep reinforcement learning. Proceedings of the 25th International Conference on Pattern Recognition (ICPR 2020), Milan, Italy.","DOI":"10.1109\/ICPR48806.2021.9413027"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"3185","DOI":"10.1109\/TNSE.2020.3017751","article-title":"RL-Routing: An SDN Routing Algorithm Based on Deep Reinforcement Learning","volume":"7","author":"Chen","year":"2020","journal-title":"IEEE Trans. Netw. Sci. Eng."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/1\/429\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:48:43Z","timestamp":1760147323000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/1\/429"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,30]]},"references-count":39,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,1]]}},"alternative-id":["s23010429"],"URL":"https:\/\/doi.org\/10.3390\/s23010429","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,30]]}}}