{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,16]],"date-time":"2026-05-16T22:07:07Z","timestamp":1778969227171,"version":"3.51.4"},"reference-count":39,"publisher":"National Library of Serbia","issue":"1","license":[{"start":{"date-parts":[[2026,1,1]],"date-time":"2026-01-01T00:00:00Z","timestamp":1767225600000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["ComSIS","COMPUT SCI INF SYST","COMPUT SCI INFORM SY","COMPUTER SCI INFORM","COMSIS J"],"published-print":{"date-parts":[[2026]]},"abstract":"<jats:p>With the rapid growth of network services, traditional static bandwidth allocation schemes can no longer meet the demands of multi-user, dynamic, and QoS-sensitive applications. Ensuring both efficiency and stability in bandwidth allocation remains a significant challenge, especially under high variability and uncertainty conditions. To address this, we propose a novel algorithm named Uncertainty- Constrained Stability-aware Deep Reinforcement Learning (UCS-DRL) for dynamic bandwidth allocation. UCS-DRL adopts a dual-policy architecture: a task policy that learns optimal bandwidth allocation decisions, and a stability policy guided by uncertainty-aware value estimation to identify and mitigate potential risky or unstable behaviors during deployment. Furthermore, the framework incorporates a curiosity-driven exploration mechanism based on Random Network Distillation, which enhances exploration efficiency by encouraging the agent to visit informative and under-explored states. Experimental results show that UCS-DRL achieves high bandwidth utilization and service quality while reducing policy volatility and risky actions, balancing performance and robustness in dynamic bandwidth allocation.<\/jats:p>","DOI":"10.2298\/csis250923008l","type":"journal-article","created":{"date-parts":[[2026,1,20]],"date-time":"2026-01-20T10:44:04Z","timestamp":1768905844000},"page":"277-297","source":"Crossref","is-referenced-by-count":0,"title":["Adaptive bandwidth allocation via uncertainty-constrained deep reinforcement learning"],"prefix":"10.2298","volume":"23","author":[{"given":"Li","family":"Wei","sequence":"first","affiliation":[{"name":"Suzhou Suneng Group Co., LTD, Suzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wu","family":"Yong","sequence":"additional","affiliation":[{"name":"Suzhou Suneng Group Co., LTD, Suzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yan","family":"Dong","sequence":"additional","affiliation":[{"name":"Suzhou Suneng Group Co., LTD, Suzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1078","reference":[{"key":"ref1","doi-asserted-by":"crossref","unstructured":"Jos\u00e9 Gaspar, Tiago Cruz, Chan-Tong Lam, and Paulo Sim\u00f5es. Smart substation communications and cybersecurity: A comprehensive survey. IEEE Communications Surveys & Tutorials, 25(4):2456-2493, 2023.","DOI":"10.1109\/COMST.2023.3305468"},{"key":"ref2","doi-asserted-by":"crossref","unstructured":"Weiti Lv. Research on network application automation system based on computer artificial intelligence technology. In 2023 IEEE 2nd International Conference on Electrical Engineering, Big Data and Algorithms (EEBDA), pages 1934-1938. IEEE, 2023.","DOI":"10.1109\/EEBDA56825.2023.10090826"},{"key":"ref3","doi-asserted-by":"crossref","unstructured":"Xiaodong Qiao, Naichen Yan, and Lu Zhang. Design and implementation of substation monitoring system based on harmonyos and broadband-narrowband wireless communication technology. In 2025 2nd International Conference on Smart Grid and Artificial Intelligence (SGAI), pages 1484-1488. IEEE, 2025.","DOI":"10.1109\/SGAI64825.2025.11009705"},{"key":"ref4","doi-asserted-by":"crossref","unstructured":"R Mohandas and D John Aravindhar. An intelligent dynamic bandwidth allocation method to support quality of service in internet of things. International Journal of Computing, 20(2):254- 261, 2021.","DOI":"10.47839\/ijc.20.2.2173"},{"key":"ref5","doi-asserted-by":"crossref","unstructured":"G Senthilkumar, KN Madhusudhan, Y Jeyasheela, and P Ajitha. A novel blockchain enabled resource allocation and task offloading strategy in cloud computing environment. Automatika, 65(3):973-982, 2024.","DOI":"10.1080\/00051144.2024.2314906"},{"key":"ref6","doi-asserted-by":"crossref","unstructured":"Nurshazlina Suhaimy, Nurul Asyikin Mohamed Radzi, Wan Siti Halimatul Munirah Wan Ahmad, Kaiyisah Hanis Mohd Azmi, and MA Hannan. Current and future communication solutions for smart grids: A review. IEEE Access, 10:43639-43668, 2022.","DOI":"10.1109\/ACCESS.2022.3168740"},{"key":"ref7","doi-asserted-by":"crossref","unstructured":"Lan-Huong Nguyen, Van-Linh Nguyen, Ren-Hung Hwang, Jian-Jhih Kuo, Yu-Wen Chen, Chien-Chung Huang, and Ping-I Pan. Toward secured smart grid 2.0: Exploring security threats, protection models, and challenges. IEEE Communications Surveys & Tutorials, 27(4):2581-2620, 2025.","DOI":"10.1109\/COMST.2024.3493630"},{"key":"ref8","doi-asserted-by":"crossref","unstructured":"Hoa Tran-Dang, Sanjay Bhardwaj, Tariq Rahim, Arslan Musaddiq, and Dong-Seong Kim. Reinforcement learning based resource management for fog computing environment: Literature review, challenges, and open issues. Journal of Communications and Networks, 24(1):83-98, 2022.","DOI":"10.23919\/JCN.2021.000041"},{"key":"ref9","unstructured":"John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, pages 1-12, 2017."},{"key":"ref10","doi-asserted-by":"crossref","unstructured":"Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. Human-level control through deep reinforcement learning. nature, 518(7540):529-533, 2015.","DOI":"10.1038\/nature14236"},{"key":"ref11","doi-asserted-by":"crossref","unstructured":"Shangding Gu, Long Yang, Yali Du, Guang Chen, Florian Walter, Jun Wang, and Alois Knoll. A review of safe reinforcement learning: Methods, theories, and applications. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12):11216-11235, 2024.","DOI":"10.1109\/TPAMI.2024.3457538"},{"key":"ref12","unstructured":"Liqing Liu, Xiaoming Yuan, Decheng Chen, Ning Zhang, Haifeng Sun, and Amir Taherkordi. Multi-user dynamic computation offloading and resource allocation in 5g mec heterogeneous networks with static and dynamic subchannels. IEEE Transactions on Vehicular Technology, 72(11):14924-14938, 2023."},{"key":"ref13","doi-asserted-by":"crossref","unstructured":"Ahmed Mohammed, Nor Fadzilah Abdullah, Sameer Alani, Othman S Alheety, Mohammed Mudhafar Shaker, Mohammed Ayad Saad, and Sarmad Nozad Mahmood. Weighted round robin scheduling algorithms in mobile ad hoc network. In 2021 3rd International Congress on Human-Computer Interaction, Optimization and Robotic Applications (HORA), pages 1-5. IEEE, 2021.","DOI":"10.1109\/HORA52670.2021.9461358"},{"key":"ref14","doi-asserted-by":"crossref","unstructured":"Mays A Mawlood and Dhari Ali Mahmood. Performance analysis of weighted fair queuing (wfq) scheduler algorithm through efficient resource allocation in network traffic modeling. Journal of Communications Software and Systems, 20(3):266-277, 2024.","DOI":"10.24138\/jcomss-2024-0009"},{"key":"ref15","doi-asserted-by":"crossref","unstructured":"Mohit Mittal, Celestine Iwendi, Suleman Khan, and Abdul Rehman Javed. Analysis of security and energy efficiency for shortest route discovery in low-energy adaptive clustering hierarchy protocol using levenberg-marquardt neural network and gated recurrent unit for intrusion detection system. Transactions on Emerging Telecommunications Technologies, 32(6):1-16, 2021.","DOI":"10.1002\/ett.3997"},{"key":"ref16","doi-asserted-by":"crossref","unstructured":"Inayat Ali, Seungwoo Hong, and Taesik Cheung. Quality of service and congestion control in software-defined networking using policy-based routing. Applied Sciences, 14(19):1-13, 2024.","DOI":"10.3390\/app14199066"},{"key":"ref17","doi-asserted-by":"crossref","unstructured":"Jianhu Gong and Hamed Nazari. A fuzzy bandwidth and delay guaranteed routing algorithm for performance enhancement of video conference over mpls networks. Journal of Ambient Intelligence and Humanized Computing, 14(6):7079-7090, 2023.","DOI":"10.1007\/s12652-021-03560-8"},{"key":"ref18","unstructured":"Yingnan Deng. A reinforcement learning approach to traffic scheduling in complex data center topologies. Journal of Computer Technology and Software, 4(3):1-6, 2025."},{"key":"ref19","doi-asserted-by":"crossref","unstructured":"Torana Kamble, Sanjivani Deokar, Vinod SWadne, Devendra P Gadekar, Hrishikesh Bhanudas Vanjari, and Purva Mange. Predictive resource allocation strategies for cloud computing environments using machine learning. Journal of Electrical Systems, 19(2):68-77, 2023.","DOI":"10.52783\/jes.692"},{"key":"ref20","doi-asserted-by":"crossref","unstructured":"Bo Wu and Yifan Hu. Analysis of substation joint safety control system and model based on multi-source heterogeneous data fusion. IEEE Access, 11:35281-35297, 2023.","DOI":"10.1109\/ACCESS.2023.3264707"},{"key":"ref21","doi-asserted-by":"crossref","unstructured":"Wei Sun, Pengyu Li, Zhi Liu, Xue Xue, Qiyue Li, Haiyan Zhang, and JunboWang. Lstm based link quality confidence interval boundary prediction for wireless communication in smart grid. Computing, 103(2):251-269, 2021.","DOI":"10.1007\/s00607-020-00816-7"},{"key":"ref22","doi-asserted-by":"crossref","unstructured":"Yang Chen, Jia Hao, Yu Peng, and Hongyan Xia. Transformer-based performance prediction and proactive resource allocation for cloud-native microservices. Cluster Computing, 28(9):568-590, 2025.","DOI":"10.1007\/s10586-025-05237-9"},{"key":"ref23","doi-asserted-by":"crossref","unstructured":"Linqiang Huang, Miao Ye, Xingsi Xue, Yong Wang, Hongbing Qiu, and Xiaofang Deng. Intelligent routing method based on dueling dqn reinforcement learning and network traffic state prediction in sdn. Wireless Networks, 30(5):4507-4525, 2024.","DOI":"10.1007\/s11276-022-03066-x"},{"key":"ref24","unstructured":"Evgenii Nikishin, Max Schwarzer, Pierluca D\u2019Oro, Pierre-Luc Bacon, and Aaron Courville. The primacy bias in deep reinforcement learning. In International Conference on Machine Learning, pages 16828-16847. PMLR, 2022."},{"key":"ref25","doi-asserted-by":"crossref","unstructured":"Jiawei Xu, Yufeng Wang, Bo Zhang, and Jianhua Ma. A graph reinforcement learning based sdn routing path selection for optimizing long-term revenue. Future Generation Computer Systems, 150:412-423, 2024.","DOI":"10.1016\/j.future.2023.09.017"},{"key":"ref26","doi-asserted-by":"crossref","unstructured":"Abdelhak Bentaleb, Mehmet N Akcay, May Lim, Ali C Begen, and Roger Zimmermann. Bob: Bandwidth prediction for real-time communications using heuristic and reinforcement learning. IEEE Transactions on Multimedia, 25:6930-6945, 2022.","DOI":"10.1109\/TMM.2022.3216456"},{"key":"ref27","unstructured":"David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. Deterministic policy gradient algorithms. In Eric P. Xing and Tony Jebara, editors, Proceedings of the 31st International Conference on Machine Learning, volume 32 of Proceedings of Machine Learning Research, pages 387-395, Bejing, China, 22-24 Jun 2014. PMLR."},{"key":"ref28","doi-asserted-by":"crossref","unstructured":"Hongliang Zeng, Ping Zhang, Fang Li, Chubin Lin, and Junkang Zhou. Ahegc: Adaptive hindsight experience replay with goal-amended curiosity module for robot control. IEEE Transactions on Neural Networks and Learning Systems, 35(11):16602-16615, 2024.","DOI":"10.1109\/TNNLS.2023.3296765"},{"key":"ref29","doi-asserted-by":"crossref","unstructured":"Yuqing Cheng, Zhiying Cao, Xiuguo Zhang, Qilei Cao, and Dezhen Zhang. Multi objective dynamic task scheduling optimization algorithm based on deep reinforcement learning. The Journal of Supercomputing, 80(5):6917-6945, 2024.","DOI":"10.1007\/s11227-023-05714-1"},{"key":"ref30","doi-asserted-by":"crossref","unstructured":"Linrui Zhang, Li Shen, Long Yang, Shixiang Chen, Bo Yuan, Xueqian Wang, and Dacheng Tao. Penalized proximal policy optimization for safe reinforcement learning. In International Joint Conference on Artificial Intelligence, pages 1-7, 2022.","DOI":"10.24963\/ijcai.2022\/520"},{"key":"ref31","unstructured":"Yue Wang and Shaofeng Zou. Online robust reinforcement learning with model uncertainty. Advances in Neural Information Processing Systems, 34:7193-7206, 2021."},{"key":"ref32","doi-asserted-by":"crossref","unstructured":"Marc Rigter, Bruno Lacerda, and Nick Hawes. Rambo-rl: Robust adversarial model-based offline reinforcement learning. Advances in Neural Information Processing Systems, 35:16082- 16097, 2022.","DOI":"10.52202\/068431-1170"},{"key":"ref33","doi-asserted-by":"crossref","unstructured":"Davide Maran, Pierriccardo Olivieri, Francesco Emanuele Stradi, Giuseppe Urso, Nicola Gatti, and Marcello Restelli. Online markov decision processes configuration with continuous decision space. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 14315-14322, 2024.","DOI":"10.1609\/aaai.v38i13.29344"},{"key":"ref34","doi-asserted-by":"crossref","unstructured":"Abdullah Al Hayajneh, Hasnain Nizam Thakur, and Kutub Thakur. The evolution of information security strategies: A comprehensive investigation of infosec risk assessment in the contemporary information era. Computer and Information Science, 16(4):1-20, 2023.","DOI":"10.5539\/cis.v16n4p1"},{"key":"ref35","unstructured":"Claire Chen, Shuze Liu, and Shangtong Zhang. Efficient policy evaluation with safety constraint for reinforcement learning. In The Thirteenth International Conference on Learning Representations, pages 1-22, 2025."},{"key":"ref36","doi-asserted-by":"crossref","unstructured":"Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell. Curiosity-driven exploration by self-supervised prediction. In 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 488-489, 2017.","DOI":"10.1109\/CVPRW.2017.70"},{"key":"ref37","unstructured":"Adri\u00e0 Puigdom\u00e8nech Badia, Pablo Sprechmann, Alex Vitvitskyi, Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martin Arjovsky, Alexander Pritzel, Andrew Bolt, and Charles Blundell. Never give up: Learning directed exploration strategies. In International Conference on Learning Representations, pages 1-28."},{"key":"ref38","unstructured":"Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov. Exploration by random network distillation. In International Conference on Learning Representations, pages 1-17, 2019."},{"key":"ref39","doi-asserted-by":"crossref","unstructured":"Zhiqun Wang, Zikai Jin, Zhen Yang, Wenchao Zhao, and Mahdi Mir. An intelligent fuzzy reinforcement learning-based routing algorithm with guaranteed latency and bandwidth in sdn: Application of video conferencing services. Egyptian Informatics Journal, 27:1-11, 2024.","DOI":"10.1016\/j.eij.2024.100524"}],"container-title":["Computer Science and Information Systems"],"original-title":[],"language":"en","deposited":{"date-parts":[[2026,5,16]],"date-time":"2026-05-16T21:34:27Z","timestamp":1778967267000},"score":1,"resource":{"primary":{"URL":"https:\/\/doiserbia.nb.rs\/Article.aspx?ID=1820-02142600008L"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026]]},"references-count":39,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026]]}},"URL":"https:\/\/doi.org\/10.2298\/csis250923008l","relation":{},"ISSN":["1820-0214","2406-1018"],"issn-type":[{"value":"1820-0214","type":"print"},{"value":"2406-1018","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026]]}}}