{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T15:45:34Z","timestamp":1781279134878,"version":"3.54.1"},"reference-count":39,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2025,9,16]],"date-time":"2025-09-16T00:00:00Z","timestamp":1757980800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2025,9,16]],"date-time":"2025-09-16T00:00:00Z","timestamp":1757980800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Science and Technology Project of State Grid Hebei Information Telecommunication Branch","award":["kj2024-017"],"award-info":[{"award-number":["kj2024-017"]}]},{"name":"Science and Technology Project of State Grid Hebei Information Telecommunication Branch","award":["kj2024-017"],"award-info":[{"award-number":["kj2024-017"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Network"],"abstract":"<jats:p>The rapid expansion of communication networks and increasingly complex service demands have presented significant challenges to the intelligent management of network resources. To address these challenges, we have proposed a network self-optimization framework integrating the predictive capabilities of the Large Language Model (LLM) with the decision-making capabilities of multi-agent Reinforcement Learning (RL). Specifically, historical network traffic data are converted into structured inputs to forecast future traffic patterns using a GPT-2-based prediction module. Concurrently, a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm leverages real-time sensor data\u2014including link delay and packet loss rates collected by embedded network sensors\u2014to dynamically optimize bandwidth allocation. This sensor-driven mechanism enables the system to perform real-time optimization of bandwidth allocation, ensuring accurate monitoring and proactive resource scheduling. We evaluate our framework in a heterogeneous network simulated using Mininet under diverse traffic scenarios. Experimental results show that the proposed method significantly reduces network latency and packet loss, as well as improves robustness and resource utilization, highlighting the effectiveness of integrating sensor-driven RL optimization with predictive insights from LLMs.<\/jats:p>","DOI":"10.3390\/network5030039","type":"journal-article","created":{"date-parts":[[2025,9,16]],"date-time":"2025-09-16T14:31:41Z","timestamp":1758033101000},"page":"39","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Integrating Reinforcement Learning and LLM with Self-Optimization Network System"],"prefix":"10.3390","volume":"5","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4564-470X","authenticated-orcid":false,"given":"Xing","family":"Xu","sequence":"first","affiliation":[{"name":"Information and Communication Branch of State Grid Hebei Electric Power Co., Ltd., Shijiazhuang 050051, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianbin","family":"Zhao","sequence":"additional","affiliation":[{"name":"Information and Communication Branch of State Grid Hebei Electric Power Co., Ltd., Shijiazhuang 050051, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou 310027, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rongpeng","family":"Li","sequence":"additional","affiliation":[{"name":"College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou 310027, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,9,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"449","DOI":"10.1029\/95WR02917","article-title":"An improved genetic algorithm for pipe network optimization","volume":"32","author":"Dandy","year":"1996","journal-title":"Water Resour. Res."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"930","DOI":"10.1109\/9.35806","article-title":"Control and optimization methods in communication network problems","volume":"34","author":"Ephremides","year":"1989","journal-title":"IEEE Trans. Autom. Control"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"55916","DOI":"10.1109\/ACCESS.2019.2913776","article-title":"Reinforcement learning based routing in networks: Review and classification of approaches","volume":"7","author":"Mammeri","year":"2019","journal-title":"IEEE Access"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"140439","DOI":"10.1109\/ACCESS.2024.3467393","article-title":"Deep reinforcement learning-based optimization method for D2D communication energy efficiency in heterogeneous cellular networks","volume":"12","author":"Pan","year":"2024","journal-title":"IEEE Access"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"563","DOI":"10.1109\/TCCN.2017.2758370","article-title":"An introduction to deep learning for the physical layer","volume":"3","author":"Hoydis","year":"2017","journal-title":"IEEE Trans. Cogn. Commun. Netw."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"3133","DOI":"10.1109\/COMST.2019.2916583","article-title":"Applications of deep reinforcement learning in communications and networking: A survey","volume":"21","author":"Luong","year":"2019","journal-title":"IEEE Commun. Surv. Tutor."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Chen, Y., and Guo, Y. (2024, January 26\u201328). Network link weight optimization based on antisymmetric deep graph networks and reinforcement learning. Proceedings of the 2024 Sixth International Conference on Next Generation Data-Driven Networks (NGDN), Shenyang, China.","DOI":"10.1109\/NGDN61651.2024.10744156"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"G\u00f3mez-delaHiz, J., and Gal\u00e1n-Jim\u00e9nez, J. (2024, January 6\u201310). Improving the traffic engineering of SDN networks by using local multi-agent deep reinforcement learning. Proceedings of the NOMS 2024 IEEE Network Operations and Management Symposium, Seoul, Republic of Korea.","DOI":"10.1109\/NOMS59830.2024.10575465"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Wu, D., Wang, X., Qiao, Y., Lu, J., Zhang, M., and Wang, K. (2024, January 4\u20138). NetLLM: Adapting large language models for networking. Proceedings of the ACM SIGCOMM 2024 Conference, Sydney, Australia.","DOI":"10.1145\/3651890.3672268"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"113","DOI":"10.23919\/JCIN.2024.10582829","article-title":"LLM4CP: Adapting large language models for channel prediction","volume":"9","author":"Liu","year":"2024","journal-title":"J. Commun. Inf. Netw."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Nascimento, N., Alencar, P., and Cowan, D. (2023, January 25\u201329). Self-adaptive large language model (LLM)-based multiagent systems. Proceedings of the 2023 IEEE International Conference on Autonomic Computing and Self-Organizing Systems Companion, Toronto, ON, Canada.","DOI":"10.1109\/ACSOS-C58168.2023.00048"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"9591","DOI":"10.1109\/JIOT.2021.3128883","article-title":"Joint Optimization Framework for Minimization of Device Energy Consumption in Transmission Rate Constrained UAV-Assisted IoT Network","volume":"9","author":"Mondal","year":"2022","journal-title":"IEEE Internet Things J."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Oliehoek, F.A., and Amato, C. (2016). A Concise Introduction to Decentralized POMDPs, Springer.","DOI":"10.1007\/978-3-319-28929-8"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"2847","DOI":"10.1109\/TCOMM.2023.3244239","article-title":"Network topology optimization via deep reinforcement learning","volume":"71","author":"Li","year":"2022","journal-title":"IEEE Trans. Commun."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Do, Q.V., and Koo, I. (2019, January 10\u201316). Dynamic bandwidth allocation scheme for wireless networks with energy harvesting using actor-critic deep reinforcement learning. Proceedings of the 2019 International Conference on Artificial Intelligence in Information and Communication, Okinawa, Japan.","DOI":"10.1109\/ICAIIC.2019.8669048"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Attiah, K., Ammar, M., Alnuweiri, H., and Shaban, K. (2020, January 10\u201313). Load balancing in cellular networks: A reinforcement learning approach. Proceedings of the 2020 IEEE 17th Annual Consumer Communications & Networking Conference, Las Vegas, NV, USA.","DOI":"10.1109\/CCNC46108.2020.9045699"},{"key":"ref_17","first-page":"21","article-title":"Security enhanced dynamic bandwidth allocation-based reinforcement learning","volume":"22","author":"Abuain","year":"2024","journal-title":"WSEAS Trans. Inf. Sci. Appl."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"74429","DOI":"10.1109\/ACCESS.2018.2881964","article-title":"Deep reinforcement learning for resource management in network slicing","volume":"6","author":"Li","year":"2018","journal-title":"IEEE Access"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Liu, Y., Ding, J., and Liu, X. (2020, January 13\u201316). A constrained reinforcement learning based approach for network slicing. Proceedings of the 2020 IEEE 28th International Conference on Network Protocols, Madrid, Spain.","DOI":"10.1109\/ICNP49622.2020.9259378"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"106","DOI":"10.1016\/j.ins.2019.05.012","article-title":"Data-driven dynamic resource scheduling for network slicing: A deep reinforcement learning approach","volume":"498","author":"Wang","year":"2019","journal-title":"Inf. Sci."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Shokouhi, M.H., and Wong, V.W.S. (2024, January 8\u201312). Large language models for wireless cellular traffic prediction: A multi-timespan approach. Proceedings of the GLOBECOM 2024 IEEE Global Communications Conference, Cape Town, South Africa.","DOI":"10.1109\/GLOBECOM52923.2024.10901784"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Yang, S., Wang, D., Zheng, H., and Jin, R. (2025, January 6\u201311). TimeRAG: Boosting LLM time series forecasting via retrieval-augmented generation. Proceedings of the ICASSP 2025 IEEE International Conference on Acoustics, Speech and Signal Processing, Hyderabad, India.","DOI":"10.1109\/ICASSP49660.2025.10889933"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"100150","DOI":"10.1016\/j.commtr.2024.100150","article-title":"Towards explainable traffic flow prediction with large language models","volume":"4","author":"Guo","year":"2024","journal-title":"Commun. Transp. Res."},{"key":"ref_24","first-page":"6000","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"26839","DOI":"10.1109\/ACCESS.2024.3365742","article-title":"A review on large language models: Architectures, applications, taxonomies, open issues and challenges","volume":"12","author":"Raiaan","year":"2024","journal-title":"IEEE Access"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"de Oliveira, R.L.S., Schweitzer, C.M., Shinoda, A.A., and Prete, L.R. (2014, January 4\u20136). Using Mininet for emulation and prototyping software-defined networks. Proceedings of the 2014 IEEE Colombian Conference on Communications and Computing, Bogota, Colombia.","DOI":"10.1109\/ColComCon.2014.6860404"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1240","DOI":"10.1109\/COMST.2022.3160697","article-title":"Applications of multi-agent reinforcement learning in future internet: A comprehensive survey","volume":"24","author":"Li","year":"2022","journal-title":"IEEE Commun. Surv. Tutor."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1054","DOI":"10.1109\/TNN.1998.712192","article-title":"Reinforcement learning: An introduction","volume":"9","author":"Sutton","year":"1998","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_29","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford","year":"2019","journal-title":"OpenAI Blog"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long Short-Term Memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Lea, C., Vidal, R., Reiter, A., and Hager, G.D. (2016, January 11\u201314). Temporal Convolutional Networks: A Unified Approach to Action Segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-49409-8_7"},{"key":"ref_32","unstructured":"Lowe, R., Wu, Y.I., Tamar, A., Harb, J., Abbeel, P., and Mordatch, I. (2017, January 4\u20139). Multi-agent actor-critic for mixed cooperative-competitive environments. Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS\u201917), Red Hook, NY, USA."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"997","DOI":"10.1109\/JSYST.2024.3402664","article-title":"preDQN-Based TAS Traffic Scheduling in Intelligence Endogenous Networks","volume":"18","author":"Li","year":"2024","journal-title":"IEEE Syst. J."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Teerapittayanon, S., McDanel, B., and Kung, H.T. (2016, January 4\u20138). Branchynet: Fast inference via early exiting from deep neural networks. Proceedings of the 2016 23rd International Conference on Pattern Recognition, Cancun, Mexico.","DOI":"10.1109\/ICPR.2016.7900006"},{"key":"ref_35","unstructured":"Louppe, G. (2014). Understanding Random Forests: From Theory to Practice. [Ph.D. Thesis, Universit\u00e9 de Li\u00e8ge]."},{"key":"ref_36","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2019). Continuous control with deep reinforcement learning. arXiv."},{"key":"ref_37","unstructured":"Kostrikov, I., Nair, A., and Levine, S. (2021). Offline reinforcement learning with implicit Q-learning. arXiv."},{"key":"ref_38","first-page":"7234","article-title":"Monotonic value function factorisation for deep multi-agent reinforcement learning","volume":"21","author":"Rashid","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"ref_39","first-page":"24611","article-title":"The surprising effectiveness of PPO in cooperative multi-agent games","volume":"35","author":"Yu","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."}],"container-title":["Network"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2673-8732\/5\/3\/39\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,21]],"date-time":"2025-11-21T12:37:19Z","timestamp":1763728639000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2673-8732\/5\/3\/39"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,16]]},"references-count":39,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2025,9]]}},"alternative-id":["network5030039"],"URL":"https:\/\/doi.org\/10.3390\/network5030039","relation":{},"ISSN":["2673-8732"],"issn-type":[{"value":"2673-8732","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,16]]}}}